this post was submitted on 14 Sep 2026
302 points (96.9% liked)

Technology

88041 readers
2870 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
 

Archive link: https://archive.ph/U1zEK

you are viewing a single comment's thread
view the rest of the comments
[–] melfie@lemmy.zip 62 points 1 day ago* (last edited 1 day ago) (2 children)

I don’t see how they’re going to meet their revenue projections when open models are largely turning LLMs into a free commodity. So far in 2026, we’ve gone from local LLMs being a toy to Qwen 3.8 27B giving the latest Sonnet 5 a run for its money. The OpenCode desktop app is also almost at parity with the Claude desktop app as well and I’ve heard Pi is pretty good too.

Now we are seeing techniques like a “phrase book” evolving to allowing a large portion of a huge model to reside on a SSD instead of in RAM / VRAM so that Opus-class models can run on a gaming PC. For example, this video is nuts: https://m.youtube.com/watch?v=IH8XmxiwliQ

OpenAI and Anthropic know they’re cooked, so that’s why they’re talking all this nonsense trying to keep the shell game going a while longer. 🍿

[–] PalmTreeIsBestTree@lemmy.world 14 points 1 day ago (2 children)

I’m so glad an at home competitive solution to the AI fuckery is here now. I hope this fast forwards the bubble bursting and pc prices go back down when an iPhone or Android phone can run an equally good AI on the device itself.

[–] Allero@lemmy.today 7 points 1 day ago

PC prices will be heated up by demand for hardware to run local models. There will always be more advanced models requiring more advanced hardware, or at least for a good while. First it'll need better hardware to perfect the text generation, then to improve image generation, then to improve video generation, then to make flawless code in large projects like games, etc. etc.

At the very least, GPUs and RAM will probably stay overly expensive for a while.

[–] isVeryLoud@lemmy.ca 2 points 1 day ago (1 children)

Maybe not a mobile device, but running decent local models is definitely within reach for the average consumer today by simply adding a GPU with lots of VRAM to their existing PC.

[–] Squizzy@lemmy.world 3 points 1 day ago (1 children)

I mean yes but those gpus are not necessarily within reach of the average consumer anymore. Even if you got one for a grand which is a lot of money you would have cooling and enclosure costs, possibly additional hardware.

[–] isVeryLoud@lemmy.ca 1 points 1 day ago (1 children)

Nah I got an RX 6800 XT for $300 with 16 GB VRAM, perfectly cromulent for a 27B model.

[–] Squizzy@lemmy.world 3 points 23 hours ago (1 children)

Thats is pretty decent alright, I had heard 24 was what was needed for decent local AI.

[–] isVeryLoud@lemmy.ca 2 points 23 hours ago

24 is ideal, 16 works fine especially with some MoE offloading. If you use llama-swap, a fast PCIe bus and fast storage is ideal. Takes about 20 seconds to swap between two models on my X570 board.

[–] Knock_Knock_Lemmy_In@lemmy.world 4 points 1 day ago (2 children)

This generation of AI architecture is cooked. The bet now is that the next generation "AGI" resulting from all the new datacenters will repay the trillions invested.

[–] cley_faye@lemmy.world 3 points 16 hours ago (1 children)

the next generation “AGI” resulting from all the new datacenters will repay the trillions invested.

Too bad we don't have the slightest clue on how to do that, except that it won't work with LLM alone.

Oh well, let's use the compute power for mass surveillance instead.