this post was submitted on 10 Oct 2026
35 points (94.9% liked)

Technology

88687 readers
3349 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
top 5 comments
sorted by: hot top controversial new old
[–] avidamoeba@lemmy.ca 14 points 21 hours ago (1 children)

Thats kinda how the whole Chinese competition started - optimization to make due with limited hw.

Also I just changed the inference engine I use for local models and Qwen 3.8 27B went from 50tps to 150tps. The new engine is optimized for my hw.

[–] Lydia_K@lemmy.world 2 points 6 hours ago (1 children)
[–] avidamoeba@lemmy.ca 1 points 6 hours ago* (last edited 6 hours ago) (1 children)

Radiance. Was using the latest llama.cpp. This thread caught my eye as simple enough to try.

[–] Agent641@lemmy.world 2 points 6 hours ago (1 children)

What are the specs of the machine that runs it?

[–] avidamoeba@lemmy.ca 1 points 5 hours ago

2x R9700, Ryzen 7700, 64GB RAM. The absolute best I've seen from llama.cpp was 60tps but the typical was 45-55tps. Radiance gave me 150tps on the first try. It can do parallel output while keeping 80-120tps with 2-4 streams. It can do more in parallel but I heven't tested what the numbers look like. As far as I understand these numbers are possible on a single R9700 with the MXFP4 model. Haven't tried it.