AI haters build tarpits to trap and trick AI scrapers that ignore robots.txt

pelespirit@sh.itjust.works · 2 days ago

AI haters build tarpits to trap and trick AI scrapers that ignore robots.txt

LovableSidekick@lemmy.world · edit-2 1 day ago

OTOH infinite loop detection is a well known coding issue with well known, freely available solutions, so this approach will only affect the lamest implementations of AI,

vrighter@discuss.tchncs.de · 1 day ago

an infinite loop detector detects when you’re going round in circles. They can’t detect when you’re going down an infinitely deep acyclic graph, because that, by definition doesn’t have any loops for it to detect. The best they can do is just have a threshold after which they give up.

LovableSidekick@lemmy.world · edit-2 1 day ago

You can detect pathpoints that come up repeatedly and avoid pursuing them further, which technically aren’t called “infinite loop” detection but I don’t know the correct name. The point is that the software isn’t a Star Trek robot that starts smoking and bricks itself when it hears something illogical.

Crassus@feddit.nl · 1 day ago

It can detect cycles. From a quick look at the demo of this tool it (slowly) generates some garbage text after which it places 10 random links. Each of these links loops to a newly generated page. Thus although generating the same link twice will surely happen. The change that all 10 of the links have already been generated before is small

LovableSidekick@lemmy.world · edit-2 1 day ago

I would simply add links to a list when visited and never revisit any. And that’s just simple web crawler logic, not even AI. Web crawlers that avoid problems like that are beginner/intermediate computer science homework.

dev_null@lemmy.ml · 1 day ago

They are no loops and repeated links to avoid. Every link leads to a brand new, freshly generated page with another set of brand new, never before seen links. You can go deeper and deeper forever without any loops.

LovableSidekick@lemmy.world · 13 hours ago

You can limit the visits to a domain. The honeypot doesn’t register infinite new domains.

vrighter@discuss.tchncs.de · 1 day ago

sure, if you have enough memory to store a list of all guids.

LovableSidekick@lemmy.world · 13 hours ago

It doesn’t have to memorize all possible guids, it just has to limit visits to base urls.

vrighter@discuss.tchncs.de · 3 hours ago

what part of “they do not repeat” do you still not get? You can put them in a list, but you won’t ever get a hit ic it’d just be wasting memory