this post was submitted on 23 Sep 2026
461 points (97.1% liked)
Technology
88212 readers
3813 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
I don't think today's bots take any notice of robots.txt. It's a relic from a more civilized time.
I run a few personal sites with some being available to the wider internet, and I have found that they generally do respect it. All the bots which advertise themselves as OpenAI, Anthropic, or Google do obey. However, I do occasionally see a stealth bot appear which uses a normal browser's user agent and those just do whatever they want.
Meta absolutely does not respect robots.txt. they are scraping what I have at around 200 hits / minute.
I also notice there is some data center somewhere (or multiple) proxying all its requests through residential proxies so I can't tell who is scraping. They don't scrape like Meta though. Far slower. These are the stealthy bois using spoofed user agents you mention. Can't tell who they are, though. How do they get so many residential IPs?
I have no organic traffic (I am literally just running a crawler tarpit. My page has nothing) so its really obvious that Im watching AI scrappers. My tarpit generates random links that all resolve to the same place so its super obvious its not human. Its also super obvious when two distinct IPs follow the same random word salad link milliseconds apart.
It is a business model. People install free apps on their phone. These apps generate their income by acting as residential proxy for scrapers. They form a botnet of millions of ips, actual phones, without the owners knowing it.
I'll have to double check my stats later. Haven't looked at it for a while.
I certainly don't recall them ignoring robots.txt, but I also don't remember seeing Meta in my dash at all, so it's entirely possible they, for wherever reason, never found my site.
They don’t—but they DO use search index data to figure out where to crawl.