619
Memory prices tipped to fall as China starts flooding the market with DRAM and NAND chips
(www.techspot.com)
This is a most excellent place for technology news and articles.
Agreed. I am not longer paying token fees as I am running QWEN 3.6 27B MTP on my 4090 GPU and it is as good and as fast as the frontier models for agentic coding.
Same. I'm running Qwen3.6-35B-A3B-FP8 (Qwen3.6-35B-A3B-UD-IQ4_XS.gguf) via the turboquant fork of llama.cpp with a few tweaked memory settings, and I get like 40 tokens / second -- nothing that required special insight on my part just following the instructions I saw on a youtube video I found via !LocalLLaMA@sh.itjust.works and asking claude to help me through the installation.
AI has no economic moat. There's nothing stopping anyone from running LLMs locally.
What's the rest of your stack look like?