e0qdk

joined 2 years ago
[โ€“] e0qdk@reddthat.com 3 points 1 day ago

๐Ÿฅณ๏ธ

[โ€“] e0qdk@reddthat.com 5 points 2 days ago (3 children)

Ctrl + . works for me on Linux Mint -- that's how I enter most of the ones I use.

[โ€“] e0qdk@reddthat.com 2 points 2 days ago

I sometimes dip a piece of bread in to use the oil -- particularly if it's actually olive oil. For water-based canned fish, I sometimes use some of the juice to make miso soup.

[โ€“] e0qdk@reddthat.com 15 points 2 days ago

There's a lot of info that you need to know to explore this space, so I'll take my own shot at answering. Let me know if anything needs further explaining!

An "open weight" model is an AI model where you can download the data needed to run the model on your own hardware for free. Contrast this with proprietary models-as-a-service like ChatGPT and Claude where you have no access to the data needed to run the model yourself -- you can only use it through the services provided, usually for a fee, and which can be taken away from you or changed at any time with no recourse.

The mapping to traditional open source terms does not work well since what you get is a binary artifact.

Those artifacts are released with a license -- and many of the models are licensed permissively (e.g. MIT or Apache license terms). You can take those weights, modify them, and then release them as new models -- and people do actually do this in practice!

Is it really feasible to run it 100% locally?

Yes. I run models on my own computers and have tried a number of configurations to figure out what works well. The Qwen family of open weight models (from Alibaba) are the ones I've found most useful so far. Gemma4 models (from Google) are also useful.

I prefer models that have been modified by the community to remove corporate censorship -- i.e. stripping that "As a large language model..." cover-your-ass crap and evasiveness on topics like Tiananmen Square. If that means the model is technically capable of telling me to go kill myself too, so be it; I've spent 25+ years dealing with assholes on the internet and can handle abuse from a stupid robot if I have to. (In practice though, they're usually pretty nice still unless I deliberately tell them to act like an asshole -- and then Qwen, at least, starts to sound like a snarky redditor; it's quite funny most of the time, actually.)

If yes to the previous: the software doesn't come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?

Models do require training to create, yes. It's generally not clear where the training physically happened IRL -- so, yes, some of them probably used power from gas turbines, but others may be drawing power from the Three Gorges Dam in China or solar plants or nuclear plants or whatever else is hooked up to the electric grid where the training happened. Most of them are also not very open about the data sets they were trained on. (There are exceptions to this though!) The Chinese models in particular are almost certainly trained heavily on logs extracted from Western models in addition to using whatever other data they could get ahold of. Whether you think that's ethical or not is a matter of perspective; how do you feel about Robin Hood?

Once a model has been trained though, it can run on a normal GPU. The power requirements to run an LLM are basically the same as running a video game, or, equivalently, about the same as turning on a few incandescent lightbulbs. (The iGPU in one of my systems uses 100W; the discrete GPU in another system I've tried uses 215W under load with appropriate tuning -- or 300W if you run it naively.)

If you want to run a model yourself, I recommend using llama.cpp -- there are instructions on how to get started with it here: https://llama.app/

These are the models I've found most useful:

If you have an iGPU only, I recommend using one of the so called "Mixture of Experts" (MoE) releases. These are typically named like 35B-A3B or similar; the first number indicates the total number of weights (35 billion) and the second indicates how many are "active" (i.e. actually used during computation) at one time while the model is running (3 billion in the example). These models need less computation to run and stay fast on weaker GPUs. Qwen3.6-35B-A3B is very good in this space and was my go-to model for a long time.

If you have a discrete GPU and enough VRAM, I recommend using a dense model (i.e. one that activates all its weights while answering) like Qwen3.8-27B.

It's worth noting that people don't usually use the full quality weights (which are typically ~2 bytes per weight); they use a "quantized" version -- compressed in a lossy fashion like a JPEG. Going down to 4-bits (half a byte) on average per weight is about as low as most people like to go -- you will see this indicated in names like Q4_K_M. (Quantized to ~4 bits with the K quantizaation scheme, medium variant.) Usually a bigger number is better in the sense of "closer to the original quality" -- at the cost of needing more RAM.

Full quality weights are often found as safetensor files on HuggingFace. Quantized weights intended for use with llama.cpp are usually in GGUF file format.

Does that help?

[โ€“] e0qdk@reddthat.com 6 points 3 days ago

completely open datasets

Not completely -- it is mostly open, but they use a dozen or so private datasets for things like training on global regulations, minesweeper (for some reason), etc. To their credit, they do indicate this on the model cards, but it's not entirely clear what is in those datasets either.

(I fell for that bit of marketing myself awhile back.)

[โ€“] e0qdk@reddthat.com 10 points 6 days ago (1 children)

Not familiar with it specifically, but... if you want privacy using LLMs, run local models on your own computer. Anything else is just asking for a rug pull eventually -- regardless of the intentions of whoever set up the service originally.

[โ€“] e0qdk@reddthat.com 10 points 1 week ago

(or so I was told)

[โ€“] e0qdk@reddthat.com 36 points 1 week ago (1 children)

(Besides Anime)

It's pretty much entirely because anime (and adjacent content). You can't exclude that and expect a coherent answer.

People learn languages when they have a cultural reason to do so -- such as engaging with media in that language or wanting to interact with people who speak it (for business/relationships/diplomacy/whatever).

Hollywood movies and rock-and-roll were big drivers that pushed people to learn English last century. Programming and the internet probably pushed a bunch of people to learn it in the last ~20 years. K-Pop and K-Drama have been drivers for learning Korean recently. German was one of the major languages (along with French) for math and science before WWI -- but was overtaken by English after that. You get the picture, I hope.

If the US continues on its current path, maybe in a few decades all the new research papers coming out will be in Chinese instead and everyone will learn Mandarin to discuss new tech. ๐Ÿคท๏ธ

[โ€“] e0qdk@reddthat.com 2 points 1 week ago (1 children)

I'm aware that it exists but I don't know how to use it, or really why I'd want to. The few times I've tried to check it out, I've run into a sign up wall before I could see anything. (Why the fuck would I want to register on a server to chat when I have no idea who's on it or what they're even chatting about there!?) On top of that, there's the complications of federated software -- which server? which client? etc -- and the comparisons to Discord (which I dislike) really don't help either...

My impression of it so far is that I'd rather people just post here for anything that's supposed to be public.

[โ€“] e0qdk@reddthat.com 13 points 1 week ago (1 children)

I just picked up Satisfactory on sale and will probably give it a try later this week.

[โ€“] e0qdk@reddthat.com 3 points 1 week ago* (last edited 1 week ago)

If I can't go on Amazon and buy a copy for sale at any price, I'm of the opinion that I'm justified in pirating it -- and I have put my money where my mouth is on that point.

Edit: to be clear, I'm talking about commercially produced films and tv shows; I'd look elsewhere for e.g. games

 

I had some free time this weekend and I've spent some of it trying to learn Go since mlmym seems to be unmaintained and I'd like to try to fix some issues in it. I ran into a stumbling block that took a while to solve and which I had trouble finding relevant search results for. I've got it solved now, but felt like writing this up in case it helps anyone else out.

When running most go commands I tried (e.g. go mod init example/hello or go run hello.go or even something as seemingly innocuous as go doc cmd/compile when a go.mod file exists) the command would hang for a rather long time. In most cases, that was about 20~30 seconds, but in one case -- trying to get it to output the docs about the compile tool -- it took 1 minute and 15 seconds! This was on a relatively fresh Linux Mint install on old, but fairly decent hardware using golang-1.23 (installed from apt).

After the long wait, it would print out go: RLock go.mod: no locks available -- and might or might not do anything else depending on the command. (I did get documentation out after the 1min+ wait, for example.)

Now, there's no good reason I could think of why printing out some documentation or running Hello World should take that long, so I tried looking at what was going on with strace --relative-timestamps go run hello.go > trace.txt 2>&1 and found this in the output file:

0.000045 flock(3, LOCK_SH)         = -1 ENOLCK (No locks available)
25.059805 clock_gettime(CLOCK_MONOTONIC, {tv_sec=3691, tv_nsec=443533733}) = 0

It was hanging on flock for 25 seconds (before calling clock_gettime).

The directory I was running in was from an NFS mount which was using NFSv3 unintentionally. File locking does not work on NFSv3 out of the box. In my case, changing the configuration to allow it to use NFSv4 was the fix I needed. After making the change a clean Hello World build takes ~5 seconds -- and a fraction of a second with cache.

After solving it, I've found out that there are some issues related to this open already (with a different error message -- cmd/go: "RLock โ€ฆ: Function not implemented") and a reply on an old StackOverflow about a similiar issue from one of the developers encouraging people to file a new issue if they can't find a workaround (like I did). For future reference, those links are:

view more: next โ€บ