this post was submitted on 23 Sep 2026
492 points (97.1% liked)

Technology

88212 readers
4328 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
 
you are viewing a single comment's thread
view the rest of the comments
[–] adespoton@lemmy.ca 66 points 1 day ago (7 children)

I haven’t used Google as a primary search engine in a decade or more. Nothing I do depends on Gemini (I don’t even use Apple Intelligence that’s a Gemini white label).

Personally, I think what Gemini is destroying is a market-based financial model that many came to associate with the Internet, information access and lifestyle.

Most of the websites I visit pre-date Google and some have gone as far as blocking all bots via robots.txt. They survive not via predatory ad networks but by donations and volunteers.

It’s why I like Lemmy; it follows the same model.

[–] buddascrayon@lemmy.world 3 points 19 hours ago* (last edited 19 hours ago)

Just a quick correction. Robots.txt doesn't block anything, it simply politely asks the search engine to not look here or use what is contained within. A lot of the search engines ignore this and catalogue the contents of the website anyway. And pretty much all of the AI database downloaders ignore it with impunity.

[–] billwashere@lemmy.world 6 points 1 day ago (1 children)

I think it’s cute you think the AI sites are paying attention to robots.txt. 🤣

[–] adespoton@lemmy.ca 1 points 18 hours ago

I don’t; I think googlebot is paying attention to it, and AI sites get a LOT of their initial visibility and site ranking from the Google index.

Case in point: the sites I spend time on that set robots.txt to deny all bots before the LLM wave started still aren’t getting hammered with AI bot requests, while the ones that didn’t are.

I can see it on my own websites as well; I’ve got one that’s in Google’s index, and it gets a constant low volume traffic from AI bots. The ones that aren’t listed? They just get traffic from exploit scanners.

[–] floofloof@lemmy.ca 42 points 1 day ago* (last edited 1 day ago) (2 children)

I don't think today's bots take any notice of robots.txt. It's a relic from a more civilized time.

[–] Clearwater@lemmy.world 12 points 1 day ago (1 children)

I run a few personal sites with some being available to the wider internet, and I have found that they generally do respect it. All the bots which advertise themselves as OpenAI, Anthropic, or Google do obey. However, I do occasionally see a stealth bot appear which uses a normal browser's user agent and those just do whatever they want.

[–] echolalia@lemmy.ml 9 points 1 day ago* (last edited 1 day ago) (2 children)

Meta absolutely does not respect robots.txt. they are scraping what I have at around 200 hits / minute.

I also notice there is some data center somewhere (or multiple) proxying all its requests through residential proxies so I can't tell who is scraping. They don't scrape like Meta though. Far slower. These are the stealthy bois using spoofed user agents you mention. Can't tell who they are, though. How do they get so many residential IPs?

I have no organic traffic (I am literally just running a crawler tarpit. My page has nothing) so its really obvious that Im watching AI scrappers. My tarpit generates random links that all resolve to the same place so its super obvious its not human. Its also super obvious when two distinct IPs follow the same random word salad link milliseconds apart.

[–] jsproc@lemmy.world 5 points 18 hours ago

It is a business model. People install free apps on their phone. These apps generate their income by acting as residential proxy for scrapers. They form a botnet of millions of ips, actual phones, without the owners knowing it.

[–] Clearwater@lemmy.world 3 points 1 day ago

I'll have to double check my stats later. Haven't looked at it for a while.

I certainly don't recall them ignoring robots.txt, but I also don't remember seeing Meta in my dash at all, so it's entirely possible they, for wherever reason, never found my site.

[–] adespoton@lemmy.ca 5 points 1 day ago

They don’t—but they DO use search index data to figure out where to crawl.

[–] LittleBorat3@lemmy.world 6 points 1 day ago

I am back to early web 2.0 in many cases , forums and such. Simply because they needlessly fucked something that worked perfectly before 2015 or so.

[–] givesomefucks@lemmy.world 16 points 1 day ago* (last edited 1 day ago) (1 children)

Nah, Gemini is really bad.

There's a push to use AI where I work and we have our own Gemini, openAI, and Grok models and they literally check to make people are at least "trying" to use it...

Gemini will spin "thinking" for maybe 30 seconds, then something like "checking google results" and then just spitting out the same "ai overview" you'd get from searching the same thing.

All Gemini can do is Google something and summarize the first few results, but at that point a human could just pick a trustworthy site themselves instead of going with what has the best search optimization all mashed up.

With that one (and likely a lot of others) the big search engines becoming shit from AI slop and search optimization is going to also make AI even shittier, then the top links the next AI will be shittier.

We'll get AI rewrites of AI rewrites, repeating itself exponentially faster as more money is sunk into it and hallucinations just keep repeating.

Like, this is the natural results of monopolies. Corporations get so large they can't help but ruin the parts that work by integrating parts no one wants to try and make them profitable.

Because one corporation can't function and be as large as Google. It's just too much to keep organized while constantly growing profits. It's like how the square cube law limits the size of buildings and animals.

The larger an animal gets, the weaker it gets pound for pound, until it would reach a point where it isn't even strong enough to breathe...

That's where these giant tech companies are headed.

[–] UnspecificGravity@piefed.social 9 points 1 day ago (1 children)

This is exactly the kind of circular cascade failure that we are going to see. Once the AI engines all start training on their own data and sourcing each others slop it just gets worse and worse.

[–] givesomefucks@lemmy.world 7 points 1 day ago

If it wasn't already an issue, AI companies wouldn't be destroying old books to guarantee they weren't training on AI slop...

The problem is if AI replaces writers, there will never be anything new to train AI on.

Humans haven't hit a wall on innovation, were just no longer in the massive boom from early internet.

Historically the only way to cause those booms is to connect humans together to facilitate the exchange of ideas and collaboration.

We're headed so fast to the opposite I don't know if tech bros are too overconfident to understand they were a product of their time and not the other way around, or if they're smart and selfish enough that they're intentionally trying to freeze human innovation so that there won't be another boom so no other group could clean up and replace them.

[–] newbeni@lemmy.world 5 points 1 day ago* (last edited 1 day ago) (3 children)

So, I disabled everything AI the I can think of, use DDG as a search engine, ad block wherever, blah blah blah, my search results SUCK. Is there a better way to fix it?

And I use Linux at home...no blows crap

[–] Magnum@infosec.pub 3 points 1 day ago

I use SearXNG it works pretty well for me.

[–] Hudell@lemmy.dbzer0.com 2 points 1 day ago (3 children)

Suck in what way? And what kind of stuff do you usually search?

I used to think search was kind of an universal thing but after looking at some people's examples of what is a good or bad result I noticed that is definitely not the case.

For me a good search result is a list of sites that contain the words I typed on the search bar, sorted by how closely the words match and then perhaps subsorted by some website ranking system. For that, DDG has been quite decent.

[–] newbeni@lemmy.world 1 points 8 hours ago

So, I can't really mention actual results after actual queries, a lot of it would be doxxing myself. Maybe that's what it is, I'm trying to get way into my personal life instead of a general search. What I usually look for is not like "this specific thing happened to me" type things, but more like a general consensus. Sometimes it's great, sometimes I'm wondering what the heck happened.

[–] FudgyMcTubbs@lemmy.world 5 points 1 day ago (2 children)

Not OP, but I find the DDG results are SEO sites at the top for pages.

My search: how to fit a dog collar

Top 50 results from DDG:

fitting a dog collar is difficult and many people wonder how to fit a dog collar. This article will show you how to fit a dog collar.

A dog collar is difficult to fit. Perhaps you've wondered how to fit a dog collar.

To be fair, SEO ruined the internet and Internet searches well before AI and it's not DDG's fault. But I would love it if blatant SEO was filtered out of my results.

[–] adespoton@lemmy.ca 0 points 18 hours ago

Part of it might be how Google trained everyone to use natural language queries in their searches.

Instead of “how to fit a dog collar” does

“pet advice” +”dog collar” “guidance for new dog owners”

give you better results?

The idea is to search for a cluster of terms that are likely to be on the page you want to find, but aren’t hyper-focused like SEO pages.

[–] Hudell@lemmy.dbzer0.com 1 points 1 day ago

True, I think I probably just skip over those results and don't even register them anymore. They definitely happen a lot but it's not something I even remembered when I thought about search results.

[–] chunes@lemmy.world 1 points 1 day ago (1 children)

Very curious whether you're old enough to remember how google used to be in say, 2007.

[–] Hudell@lemmy.dbzer0.com 2 points 1 day ago

I am and I do, but I didn't mean to say DDG is just as good. It's just been good enough for me.

[–] brb@sh.itjust.works 3 points 1 day ago

Kagi is pretty decent but the sad truth is that something like duck.ai works better as a search engine than normal search engines

[–] non_burglar@lemmy.world 4 points 1 day ago

robots.txt has been summarily ignored for nearly 20 years, my friend. Sorry to bear the news.