Technology

84807 readers

5710 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

431

DeepSeek ditches Nvidia for Huawei chips in V4 launch (cybernews.com)

submitted 3 weeks ago* (last edited 3 weeks ago) by inari@piefed.zip to c/technology@lemmy.world

89 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] humanspiral@lemmy.ca 9 points 3 weeks ago (1 children)

Huawei outperforms NVIDIA at the "cluster" level. Which are mostly turnkey systems for datacenter units. And promises truck container level cluster for next generation that is 30x the zetaflops as NVIDIA rubin cluster. China currently operates at 50% electric production capacity, and energy extremely abundant and low price, which make the per level card performance deficit irrelevant.

[–] KingRandomGuy@lemmy.world 8 points 3 weeks ago (1 children)

To be fair, the raw FLOPs count doesn't tell the whole story. On a lot of workloads (including token generation during LLM inference), you're bound by the memory bandwidth rather than throughput/FLOPs. On H100/H200, keeping the tensor cores fully occupied is surprisingly difficult, and that's with 3+ TB/s of memory bandwidth. And I believe those cards have much higher throughput (at least at FP8, Ascend wins at FP4 since H100/200 don't support it) compared to Ascend.

The Ascend 950PR units have far lower memory bandwidth, reportedly at 1.4 TB/s. Compare that to Blackwell, which has something like 8TB/s of bandwidth. I believe they're manufacturing their own kind of HBM, so that's still really impressive considering this is a fairly recent push into manufacturing accelerators. But I'm a bit skeptical it actually outperforms NVIDIA at scale.

[–] humanspiral@lemmy.ca 1 points 3 weeks ago (1 children)

Huawei's clusters have close to 4x the ram as NVIDIAs, and TFLOPs is most relevant to training. Huawei has better interconnect technology than NVIDIA, but incompatible with H200s, and so for China/friends use, it's a much better package. Price/performance of 910 vs 5090 or 6000ada is much higher at single card level. The power cost/availability in China gives them much higher potential deployment rates. Chinese cloud rates tend to be lower than the same model on US clouds.

[–] KingRandomGuy@lemmy.world 3 points 3 weeks ago

Yeah I can believe their interconnect is better, given their extensive history in networking.

W.r.t TFLOPs, let me clarify what I meant. Even on traditionally compute-bound workloads (attention, etc.), on H200 it's actually surprisingly difficult to make full use of the card's throughput before hitting VRAM bandwidth limits. Tensor core throughput has grown a lot faster than bandwidth has.

I've never written a kernel for Huawei chips so I have no idea if they have the same problem. But this problem is there on many datacenter-class NVIDIA chips, which is why they keep introducing features (TMA, TMEM, etc.) to try and lower the time wasted waiting for memory.