this post was submitted on 09 Aug 2026
80 points (90.0% liked)
Technology
87162 readers
4235 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Auto mode sounds less like security and more like outsourcing judgment to the same AI you’re supposedly trying to constrain. If users are bad at spotting malicious prompts, the answer shouldn’t be: great, let’s remove them from the loop entirely. That’s not safety. That’s automated permission laundering.
*Twitches*
Oh good, it wasn't just me.
I'm not sure I get it, do you mean it looks like how claude writes?
Yeah, A.I. often writes like "it's not x, it's y"
especially when separated by dot instead of a comma
What do you mean?
"It's not x. It's y." instead of "It's not x, it's y."
I just assumed it was to increase token use and create a windfall for Anthropic
So if I remember correctly, what happens with auto mode is they run a second smaller LLM called the "classifier" to evaluate the tool uses of the main one.
I've used auto mode at work now for many months, and it approves most things because most things Claude does are reasonable.
Once or twice I have seen it reject Claude. I can't remember the exact scenario, but I had asked Claude to diagnose an issue but not fix it yet, and when later on it tried to make the change the classifier rejected it, giving the reason that what it was trying to do did not match my request.
In my opinion, the trick to using auto mode safely here is to:
I had Claude with auto-permission on accidentally delete an Open3D build folder in its workspace today while updating an old project to the bleeding edge out of github. I had the older release in, "open3d" and "open3d-debug", so it cleverly used "open3d*" to tidy both folders up, so that the latest version, in the folder, "Open3D-main", would be used from now on.
Except guess what? The windows box it was running on treats filenames as case-insensitive, and all folders matching open3d got nuked.
* sad trombone noises *
Cue grovelling apology from Claude offering to re-download and build Open3D again, a process that takes a couple of hours on my PC. Interestingly, it only noticed a few more steps further down the road, not immediately.
Luckily it was working in a sandbox copied from the real repository, so I just copied it back. But I really do have to get my act together and run things in proper isolation. Everything is backed up elsewhere on this build PC, but it would be a hassle if it got wrecked.
My favourite claudism is when it deletes tens of gigabytes sized models as if redownloading them is trivial for me. It might be for the boys at Anthropic with their sewer pipe sized fibre optics, but for me where my last mile is copper bullshit Openreach probably hasn't maintained since the 2008 crash it's a massive pain in the arse.
Yes I never trusted it fully, and I would strongly suggest anyone to run it in a VM or at least a container, where it doesn't have access to your credentials, browser cookies and sensitive files. It's very irresponsible otherwise.
Example: around 1 year ago the bot made a small coding mistake and created a '~' folder. I asked it to delete it (not in auto mode) and of course it went with a very nice "rm -rf ~" (which of course I denied).
Modern models are much smarter, but I would not trust the auto classifier to stop everything dangerous and make a similar hallucination slip through.