hylobates@jlai.lu to Selfhosted@lemmy.worldEnglish · 1 day agoBased on this graph, and this graph alone, guess at what time I completely blocked OpenAI crawlersjlai.luimagemessage-square59fedilinkarrow-up1497arrow-down16file-text
arrow-up1491arrow-down1imageBased on this graph, and this graph alone, guess at what time I completely blocked OpenAI crawlersjlai.luhylobates@jlai.lu to Selfhosted@lemmy.worldEnglish · 1 day agomessage-square59fedilinkfile-text
minus-squarepunrca@piefed.worldlinkfedilinkEnglisharrow-up22arrow-down2·20 hours agoIt’s best to use either Cloudflare (best IMO) or Anubis. If you don’t want any AI bots, then you can setup Anubis (open source; requires JavaScript to be enabled by the end user): https://github.com/TecharoHQ/anubis Cloudflare automatically setups robots.txt file to block “AI crawlers” (but you can setup to allow “AI search” for better SEO). Eg: https://blog.cloudflare.com/control-content-use-for-ai-training/#putting-up-a-guardrail-with-cloudflares-managed-robots-txt Cloudflare also has an option of “AI labyrinth” to serve maze of fake data to AI bots who don’t respect robots.txt file.
minus-squareAHemlocksLie@lemmy.ziplinkfedilinkEnglisharrow-up10·15 hours agoPretty sure I’ve repeatedly heard about the crawlers completely ignoring robots.txt, so does Cloudflare really do that much?
minus-squareSv443@sh.itjust.workslinkfedilinkEnglisharrow-up5·12 hours agoLike a lock on a door, it stops the vast majority but can’t do shit about the actual professional bad guys
minus-squareshane@feddit.nllinkfedilinkEnglisharrow-up17arrow-down2·12 hours agoIf you’re relying on Cloudflare are you even self-hosting?
minus-squareCyberSeeker@discuss.tchncs.delinkfedilinkEnglisharrow-up5arrow-down3·edit-210 hours agoIf you build a house, but hire a guard for the front gate, do you even own the house?!
It’s best to use either Cloudflare (best IMO) or Anubis.
If you don’t want any AI bots, then you can setup Anubis (open source; requires JavaScript to be enabled by the end user): https://github.com/TecharoHQ/anubis
Cloudflare automatically setups robots.txt file to block “AI crawlers” (but you can setup to allow “AI search” for better SEO). Eg: https://blog.cloudflare.com/control-content-use-for-ai-training/#putting-up-a-guardrail-with-cloudflares-managed-robots-txt
Cloudflare also has an option of “AI labyrinth” to serve maze of fake data to AI bots who don’t respect robots.txt file.
Pretty sure I’ve repeatedly heard about the crawlers completely ignoring robots.txt, so does Cloudflare really do that much?
Like a lock on a door, it stops the vast majority but can’t do shit about the actual professional bad guys
If you’re relying on Cloudflare are you even self-hosting?
If you build a house, but hire a guard for the front gate, do you even own the house?!