this post was submitted on 23 Sep 2026
461 points (97.1% liked)

Technology

88212 readers
3813 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
 
you are viewing a single comment's thread
view the rest of the comments
[–] floofloof@lemmy.ca 39 points 1 day ago* (last edited 1 day ago) (2 children)

I don't think today's bots take any notice of robots.txt. It's a relic from a more civilized time.

[–] Clearwater@lemmy.world 12 points 1 day ago (1 children)

I run a few personal sites with some being available to the wider internet, and I have found that they generally do respect it. All the bots which advertise themselves as OpenAI, Anthropic, or Google do obey. However, I do occasionally see a stealth bot appear which uses a normal browser's user agent and those just do whatever they want.

[–] echolalia@lemmy.ml 9 points 21 hours ago* (last edited 21 hours ago) (2 children)

Meta absolutely does not respect robots.txt. they are scraping what I have at around 200 hits / minute.

I also notice there is some data center somewhere (or multiple) proxying all its requests through residential proxies so I can't tell who is scraping. They don't scrape like Meta though. Far slower. These are the stealthy bois using spoofed user agents you mention. Can't tell who they are, though. How do they get so many residential IPs?

I have no organic traffic (I am literally just running a crawler tarpit. My page has nothing) so its really obvious that Im watching AI scrappers. My tarpit generates random links that all resolve to the same place so its super obvious its not human. Its also super obvious when two distinct IPs follow the same random word salad link milliseconds apart.

[–] jsproc@lemmy.world 2 points 9 hours ago

It is a business model. People install free apps on their phone. These apps generate their income by acting as residential proxy for scrapers. They form a botnet of millions of ips, actual phones, without the owners knowing it.

[–] Clearwater@lemmy.world 3 points 17 hours ago

I'll have to double check my stats later. Haven't looked at it for a while.

I certainly don't recall them ignoring robots.txt, but I also don't remember seeing Meta in my dash at all, so it's entirely possible they, for wherever reason, never found my site.

[–] adespoton@lemmy.ca 5 points 1 day ago

They don’t—but they DO use search index data to figure out where to crawl.