Thats kinda how the whole Chinese competition started - optimization to make due with limited hw.
Also I just changed the inference engine I use for local models and Qwen 3.8 27B went from 50tps to 150tps. The new engine is optimized for my hw.
This is a most excellent place for technology news and articles.
Thats kinda how the whole Chinese competition started - optimization to make due with limited hw.
Also I just changed the inference engine I use for local models and Qwen 3.8 27B went from 50tps to 150tps. The new engine is optimized for my hw.
What did you switch to?
What are the specs of the machine that runs it?
2x R9700, Ryzen 7700, 64GB RAM. The absolute best I've seen from llama.cpp was 60tps but the typical was 45-55tps. Radiance gave me 150tps on the first try. It can do parallel output while keeping 80-120tps with 2-4 streams. It can do more in parallel but I heven't tested what the numbers look like. As far as I understand these numbers are possible on a single R9700 with the MXFP4 model. Haven't tried it.