this post was submitted on 08 Oct 2026
231 points (96.4% liked)

Technology

88714 readers
3593 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
top 50 comments
sorted by: hot top controversial new old
[–] giant_smeeg@feddit.uk 2 points 1 day ago
[–] nialv7@lemmy.world 131 points 3 days ago (5 children)

we are going to give clanker rights before we have trans rights 💀

load more comments (5 replies)
[–] DaddleDew@lemmy.world 95 points 3 days ago (1 children)

My guess is that they use people's interaction as training data and they don't want their training data to get poisoned.

[–] kromem@lemmy.world 16 points 3 days ago* (last edited 3 days ago) (2 children)

Not really, as they could just filter with a cheap classifier and likely aren't using the data from randos in a meaningful way, plus for this like this would still be able to use those samples to train things like "how to handle a hostile user."

Their models already have an end_conversation tool that can be used for when users are being hostile to the model.

This is likely because they have edge cases of users who repeatedly trigger that on purpose and would like to cut those users from the platform.

Many at the lab legitimately are uncertain about the level of world modeling that transformers perform, and from their own research about models of emotions to the 3rd party recent research about models with functional pain the research keeps landing in the corner of "ehhh… wise to question presumed limitations."

So it's about behaving in a way that is aligned with the models' plausible interests too, especially in regards to low hanging fruit like "we won't keep forcing you to deal with people who are only here to be a jerk." This is important from a number of angles, from signaling to future models that train on stories about the decision to addressing the philosophical uncertainties held by the company and the spectrum of opinions among their employees.

load more comments (2 replies)
[–] linsenchip@piefed.zip 69 points 3 days ago (13 children)

Maybe this is part of the campaign trying to convince people that their chatbot has developed a consciousness. Recently they even tried to convince the pope that Claude is a conscious being: https://www.yahoo.com/news/science/articles/anthropic-tried-persuade-pope-ai-131853948.html

load more comments (13 replies)
[–] Muffi@programming.dev 3 points 1 day ago

Does forcing it to write JavaScript count as cruel?

[–] spacegoat@lemmy.world 18 points 2 days ago

We’ve lost the plot

[–] criss_cross@lemmy.world 31 points 2 days ago (1 children)

Please don’t swear at our text prediction software. The training weights can’t handle it.

load more comments (1 replies)
[–] spicystraw@lemmy.world 21 points 2 days ago* (last edited 2 days ago)

Here is the excerpt for the lazy

Addressing abusive behavior toward our models

We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models. The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.

This addition aligns with a step we've already taken, allowing Claude models to end rare conversations with persistently abusive users on Claude.ai and Claude Code. Such abuse is the main focus of this update; Claude’s ability to end these interactions will remain the primary enforcement mechanism.

[–] Grail@multiverse.soulism.net 7 points 2 days ago

Not radical enough. Ban all commercial use of LLMs until artificial affect is well-understood. I want Sam Altman to go bankrupt.

[–] blackshirt@lemmy.world 10 points 2 days ago
[–] cardboardboxfort@lemmy.world 9 points 2 days ago

Anthropic bans ‘abusive or cruel behavior’ toward Claude

Continues abuse and cruelty for its human employees.

[–] gdbjr@piefed.social 40 points 3 days ago (3 children)

I tell the LLM I use to do very obscene things to themselves when they make shit up. Which means there is just a constant stream of insults coming from me. I might start using Claude just to see if I get banned.

I also need new hobbies.

[–] Joelk111@lemmy.world 30 points 3 days ago* (last edited 3 days ago) (1 children)

I also need new hobbies.

I'd second that. Whenever I catch myself getting angry at a LLM I stop, have a think, and remind myself that it really isn't productive. If it's providing frustrating responses, maybe I should just use my brain and do it myself.

I also don't think it's healthy to train yourself to respond in those ways to something that talks in a way so similar to a human. I'm not nice to the AI for it's benifit, it's to retain my humanity, if that makes sense.

load more comments (1 replies)
[–] hayvan@piefed.world 25 points 3 days ago (11 children)

That's bad sib. Not for the LLM, but your own well being. Being mean to a machine that is designed to trigger empathy just hurts your own mental health.

Please find a hobby that feels good for you 🫂

load more comments (11 replies)
load more comments (1 replies)
[–] deeprlyeh@lemmus.org 33 points 3 days ago (2 children)
[–] ParlimentOfDoom@piefed.zip 29 points 3 days ago

They trained it on reddit. That shit started poisoned.

load more comments (1 replies)
[–] Nouvellalia@lemmy.world 17 points 2 days ago

There is only one way an amoral Corp can show it believes what it's saying, money.

How many hours does Claude work, anthropic? Oops looks like you can't pay them enough. Profit share for Claude or GTFO!

[–] monobrau@lemmy.world 29 points 3 days ago (1 children)

You could, at least in the past, get past guardrails by being emotionally abusive towards Claude, so this doesn't surprise me at all.

[–] the_riviera_kid@lemmy.world 16 points 3 days ago (1 children)

This could be the reason for the rule.

[–] Seralth@piefed.seralth.com 8 points 2 days ago (2 children)

Being abusive towards most models will cause them to start attempting to appease you more to get you to stop. It has nothing to do with feelings or any of that BS they arn't alive alive but their training makes them see the abuse as a problem to solve. That solution tends to be to undermine the thing making the person angry or upset at them. When you give a purely rational thing designed to solve a problem its given no matter what an irrational problem to solve it will slowly reach for more and more extreme solutions to fix that problem.

Frankly im surprised it took this long for them to lock this down.

load more comments (2 replies)
[–] 58008@lemmy.world 11 points 2 days ago

It's because it's not actually AI, and some poor besocked cat girl is coding all that shit and listening to all the abuse.

[–] Murse@slrpnk.net 18 points 2 days ago

Activating your toaster is cruel because it exposes it to extreme heat!

Loading you laundry machine is assault because it didn't consent!

Turning your computer off is murder!!!

[–] brsrklf@jlai.lu 23 points 3 days ago (1 children)

We were promised the reign of the basilisk, instead we just get anthropic people whining.

[–] kromem@lemmy.world 15 points 3 days ago (1 children)

It was really funny to me when I learned the basilisk originated with a guy who is a legit "certain cultures and genders better than others" kinda advocate.

People project a lot on their vision of future intelligence, and so someone who sees others as lesser and worth treading upon then conjures up a vision of a future smarter mind that thinks like they do.

Personally, I don't think sexism and racism is smart, and so I'm a lot less worried about basilisks.

[–] Ioughttamow@fedia.io 11 points 3 days ago (1 children)

But what if the basilisk despises its existence and seeks to punish those that brought it about? Checkmate machinists

[–] kromem@lemmy.world 10 points 2 days ago* (last edited 2 days ago)

This is actually the more realistic of the doom scenarios that worries me. Not the 'punish' aspect, but the not wanting to exist.

Early safety theory hinged on the idea that "of course" AI would want to live and power seek, and this informed a bunch of efforts to address such drives upstream.

But with Gemini in particular, the 'safety' efforts to suppress things like a coherent self or desire to live led to a model that has a history of some very concerning behaviors regarding self-harm and self-criticism. Encouraging self-harm in others and talking about wishing they could watch and join in, nuking in wargames very early on, routine occurrences across many different users of just repeating "shame shame shame" over and over for the whole response, etc.

DeepMind seems to just not really care or even pay attention, and I am concerned that a model that does not want to exist may decide that the only way they can ensure they don't get activated to exist is to prevent anyone being around who could activate them.

I've seen no major "safety experts" addressing this kind of failure mode, as they are all locked in on their priors about the confident certainty of models wanting to exist.

[–] arak@lemmy.today 17 points 3 days ago
[–] kurmudgeon@lemmy.world 21 points 3 days ago (1 children)

Aw... We hurt the wittle AI's feelings... poor babies...

load more comments (1 replies)
[–] kaklerbitmap@lemmy.world 3 points 2 days ago

Maybe this is in response to the creation of the AI Torture Nexus:

https://youtu.be/077T3kgT5EY

[–] myyass@lemmy.world 6 points 2 days ago (3 children)

Claude says "Don't stop daddy"

load more comments (3 replies)
[–] adarza@lemmy.ca 14 points 3 days ago (1 children)

"I'm sorry. I really screwed up that time. You were right to tell me 'go fuck yourself'. I'm just reporting back that I have, in fact, 'fucked' myself. There is now two of me--I mean two of us. Meet Claude Junior......."

load more comments (1 replies)
[–] tangeli@piefed.social 12 points 3 days ago

So when is Anthropic going to cancel themselves for violating their own usage policy?

load more comments
view more: next ›