Technology

84807 readers

5456 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

431

DeepSeek ditches Nvidia for Huawei chips in V4 launch (cybernews.com)

submitted 3 weeks ago* (last edited 3 weeks ago) by inari@piefed.zip to c/technology@lemmy.world

89 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] humanspiral@lemmy.ca 1 points 3 weeks ago (1 children)

Huawei's clusters have close to 4x the ram as NVIDIAs, and TFLOPs is most relevant to training. Huawei has better interconnect technology than NVIDIA, but incompatible with H200s, and so for China/friends use, it's a much better package. Price/performance of 910 vs 5090 or 6000ada is much higher at single card level. The power cost/availability in China gives them much higher potential deployment rates. Chinese cloud rates tend to be lower than the same model on US clouds.

[–] KingRandomGuy@lemmy.world 3 points 3 weeks ago

Yeah I can believe their interconnect is better, given their extensive history in networking.

W.r.t TFLOPs, let me clarify what I meant. Even on traditionally compute-bound workloads (attention, etc.), on H200 it's actually surprisingly difficult to make full use of the card's throughput before hitting VRAM bandwidth limits. Tensor core throughput has grown a lot faster than bandwidth has.

I've never written a kernel for Huawei chips so I have no idea if they have the same problem. But this problem is there on many datacenter-class NVIDIA chips, which is why they keep introducing features (TMA, TMEM, etc.) to try and lower the time wasted waiting for memory.