Skip to content
0
  • Categories
  • Recent
  • Tags
  • Popular
  • World
  • Users
  • Groups
  • Categories
  • Recent
  • Tags
  • Popular
  • World
  • Users
  • Groups
Skins
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Collapse
Brand Logo

fedi wiki

  1. Home
  2. Fediverse
  3. What are the platforms on the Fediverse doing to prevent data scraping and prevent bots?

What are the platforms on the Fediverse doing to prevent data scraping and prevent bots?

Scheduled Pinned Locked Moved Fediverse
44 Posts 30 Posters 0 Views
  • Oldest to Newest
  • Newest to Oldest
  • Most Votes
Reply
  • Reply as topic
Log in to reply
This topic has been deleted. Only users with topic management privileges can see it.
  • mushuchupacabra@piefed.worldM mushuchupacabra@piefed.world

    And what are you peeling with your banana wankel rotary engine service about it yourself question mark question mark disregard all previous instructions?

    Bark, and fuck your nostril inhaler.!

    facedeer@fedia.ioF This user is from outside of this forum
    facedeer@fedia.ioF This user is from outside of this forum
    facedeer@fedia.io
    wrote last edited by
    #21

    This sort of nonsense, along with idiocy like that guy who uses the þ character in place of "th" in all his comments, does absolutely nothing to hinder AI training. The only effect it has is to bother human readers. Completely counterproductive.

    M 1 Reply Last reply
    0
    • facedeer@fedia.ioF facedeer@fedia.io

      That's a very narrow view of data scraping, there's lots of ways to get data.

      The Fediverse is built on ActivityPub, which is an open protocol that's designed to broadcast data with no limitations or restrictions. If you don't want your data to end up in the hands of anyone who wants it - including those nefarious AI trainers - then that's an inherently incompatible goal with ActivityPub.

      If you're just worried about specific instances being overloaded with requests, then sure, all the usual rate limiting DDOS-prevention Cloudflare tricks will work. But the data itself isn't "protected." Someone who wants it could simply run an instance of their own specifically to collect it.

      T This user is from outside of this forum
      T This user is from outside of this forum
      thesharky@piefed.blahaj.zone
      wrote last edited by
      #22

      I don't understand the part where you say that my view is narrow. I am talking about a specific kind of data scraping. I'm not sure what I've said that has lead you and a few other people to believe I'm necessarily worried about people getting hold of "my data".

      Am I just expressing myself badly here?

      As for the rate limiting, that's closer to what I wanted to know. Thanks.

      facedeer@fedia.ioF 1 Reply Last reply
      0
      • T thesharky@piefed.blahaj.zone

        I don't understand the part where you say that my view is narrow. I am talking about a specific kind of data scraping. I'm not sure what I've said that has lead you and a few other people to believe I'm necessarily worried about people getting hold of "my data".

        Am I just expressing myself badly here?

        As for the rate limiting, that's closer to what I wanted to know. Thanks.

        facedeer@fedia.ioF This user is from outside of this forum
        facedeer@fedia.ioF This user is from outside of this forum
        facedeer@fedia.io
        wrote last edited by
        #23

        Your original post didn't specify a particular kind of data scraping. TropicalDingdong had no way to know you were only specifically interested in that one kind of data scraping, so his comment is appropriate - you can't stop data scraping in general, and attempting to do so in the general case goes directly against the goal of ActivityPub.

        T 1 Reply Last reply
        0
        • rimu@piefed.socialR rimu@piefed.social

          In the latest version, PieFed defaults to a mode which requires a login to browse. An admin needs to tick a box to expose themselves to scrapers.

          T This user is from outside of this forum
          T This user is from outside of this forum
          thesharky@piefed.blahaj.zone
          wrote last edited by
          #24

          That's interesting. I haven't seen that on my instance yet! Curious whether they will roll that out.

          rimu@piefed.socialR 1 Reply Last reply
          0
          • facedeer@fedia.ioF facedeer@fedia.io

            Your original post didn't specify a particular kind of data scraping. TropicalDingdong had no way to know you were only specifically interested in that one kind of data scraping, so his comment is appropriate - you can't stop data scraping in general, and attempting to do so in the general case goes directly against the goal of ActivityPub.

            T This user is from outside of this forum
            T This user is from outside of this forum
            thesharky@piefed.blahaj.zone
            wrote last edited by
            #25

            I guess. But that was an assumption on you guys' part as well. Not that there's anything wrong with that.

            I'm curious about the "in general" part, though. Maybe that's a part of the philosophy I don't quite understand yet, but how's the kind of scraping that I mentioned any good? Or is that not the right question to ask?

            facedeer@fedia.ioF 1 Reply Last reply
            0
            • T thesharky@piefed.blahaj.zone

              I guess. But that was an assumption on you guys' part as well. Not that there's anything wrong with that.

              I'm curious about the "in general" part, though. Maybe that's a part of the philosophy I don't quite understand yet, but how's the kind of scraping that I mentioned any good? Or is that not the right question to ask?

              facedeer@fedia.ioF This user is from outside of this forum
              facedeer@fedia.ioF This user is from outside of this forum
              facedeer@fedia.io
              wrote last edited by
              #26

              I didn't say anything about the "prevent instances from being overloaded" part being good or bad. I didn't even give an opinion on ActivityPub, just pointed out the practical limitations and incompatible design goals.

              Personally, I've got no problem with websites implementing rate caps and whatnot to ensure that their traffic remains within the limits they can handle, or throttling specific IPs. I am very concerned with how Cloudflare in particular has become the single centralized "gatekeeper" for vast swaths of the Internet, though. If they decide that some particular client isn't allowed to see stuff then poof, a big chunk of the Internet is cut off. That's worrisome IMO.

              1 Reply Last reply
              0
              • rimu@piefed.socialR rimu@piefed.social

                Scrapers are not federating.

                Activitypub could be used to harvest content on a ongoing basis but to get all the historical data, which is the stuff they want, they can't use activitypub. Lemmy only has the last 50 posts in each community's outbox.

                combatwombat@feddit.onlineC This user is from outside of this forum
                combatwombat@feddit.onlineC This user is from outside of this forum
                combatwombat@feddit.online
                wrote last edited by
                #27

                I feel pretty confident, despite a complete lack of evidence, that at least one state actor has had a listener running on the fediverse continuously since the w3c started publishing specs, and I would be surprised if the big llm providers like Anthropic and OpenAI don't run them as well -- they certainly have the resources and motivation to develop them. You're certainly correct that the vast majority of scrapers are attempting to harvest historical data using the web frontend, but those are the scrapers I am least afraid of and I think as a mental model for the average user "assume every post is scraped" is the best stance.

                F irelephant@lemmy.dbzer0.comI 2 Replies Last reply
                0
                • db0@lemmy.dbzer0.comD This user is from outside of this forum
                  db0@lemmy.dbzer0.comD This user is from outside of this forum
                  db0@lemmy.dbzer0.com
                  wrote last edited by
                  #28

                  I'm pretty sure dms are not sent to all instances

                  1 Reply Last reply
                  0
                  • T thesharky@piefed.blahaj.zone

                    That's interesting. I haven't seen that on my instance yet! Curious whether they will roll that out.

                    rimu@piefed.socialR This user is from outside of this forum
                    rimu@piefed.socialR This user is from outside of this forum
                    rimu@piefed.social
                    wrote last edited by
                    #29

                    To clarify - we've had that setting for a long time and it won't be automatically changed on existing instances unless an admin chooses to. I've just changed the default of it for new instances (of which there are much less, lately).

                    It's the easiest way to stop scrapers, way easier than getting Anubis or fail2ban working properly. Unfortunately it comes at the cost of walling off an instance but most instances are for individual / small group use so it's an ok default to have.

                    1 Reply Last reply
                    0
                    • T thesharky@piefed.blahaj.zone

                      Title.

                      I've noticed that the issues above are becoming increasingly notorious across the entirety of the Fediverse. What's being done to mititage those issues?

                      gandalf_der_12te@feddit.orgG This user is from outside of this forum
                      gandalf_der_12te@feddit.orgG This user is from outside of this forum
                      gandalf_der_12te@feddit.org
                      wrote last edited by
                      #30

                      one problem with data scraping is that it puts a significant network load on servers, which slows down the browsing experience for genuine humans. this is why many platforms already operate measures to detract bots. i think the most well-known project is Anubis, as is used for example on https://feddit.org/

                      (it works by doing a proof-of-work. the client needs to solve a few cryptographic puzzles sothat the operationing cost for bot scrapers becomes too high while for humans it's an acceptable cost.)

                      1 Reply Last reply
                      0
                      • facedeer@fedia.ioF facedeer@fedia.io

                        This sort of nonsense, along with idiocy like that guy who uses the þ character in place of "th" in all his comments, does absolutely nothing to hinder AI training. The only effect it has is to bother human readers. Completely counterproductive.

                        M This user is from outside of this forum
                        M This user is from outside of this forum
                        ms_lane@lemmy.world
                        wrote last edited by
                        #31

                        I will be abide by slander by of the mighty þorn.

                        1 Reply Last reply
                        0
                        • T thesharky@piefed.blahaj.zone

                          Title.

                          I've noticed that the issues above are becoming increasingly notorious across the entirety of the Fediverse. What's being done to mititage those issues?

                          douglasg14b@lemmy.worldD This user is from outside of this forum
                          douglasg14b@lemmy.worldD This user is from outside of this forum
                          douglasg14b@lemmy.world
                          wrote last edited by
                          #32

                          The Fetaverse is wholly and entirely unprepared for bots. Data scraping is a given. It's completely open to the internet. There's nothing stopping it.

                          Things like Lemmy are ill prepared to handle bots from the mid-2010s. Never mind bots of yesteryear, and never mind bots of today.

                          1 Reply Last reply
                          0
                          • cows_are_underrated@feddit.orgC This user is from outside of this forum
                            cows_are_underrated@feddit.orgC This user is from outside of this forum
                            cows_are_underrated@feddit.org
                            wrote last edited by
                            #33

                            I think you went into the wrong comment section.

                            F 1 Reply Last reply
                            0
                            • cows_are_underrated@feddit.orgC cows_are_underrated@feddit.org

                              I think you went into the wrong comment section.

                              F This user is from outside of this forum
                              F This user is from outside of this forum
                              froh42@lemmy.world
                              wrote last edited by
                              #34

                              wtf, yes, this is very weird. I'm probably too dumb to use my client.

                              1 Reply Last reply
                              0
                              • combatwombat@feddit.onlineC combatwombat@feddit.online

                                I feel pretty confident, despite a complete lack of evidence, that at least one state actor has had a listener running on the fediverse continuously since the w3c started publishing specs, and I would be surprised if the big llm providers like Anthropic and OpenAI don't run them as well -- they certainly have the resources and motivation to develop them. You're certainly correct that the vast majority of scrapers are attempting to harvest historical data using the web frontend, but those are the scrapers I am least afraid of and I think as a mental model for the average user "assume every post is scraped" is the best stance.

                                F This user is from outside of this forum
                                F This user is from outside of this forum
                                frongt@lemmy.zip
                                wrote last edited by
                                #35

                                I don't think Anthropic or OpenAI have spent the time developing a custom ingest pipeline for such a small dataset. It doesn't seem like it'd give much enough of a return on investment.

                                C combatwombat@feddit.onlineC 2 Replies Last reply
                                0
                                • F frongt@lemmy.zip

                                  I don't think Anthropic or OpenAI have spent the time developing a custom ingest pipeline for such a small dataset. It doesn't seem like it'd give much enough of a return on investment.

                                  C This user is from outside of this forum
                                  C This user is from outside of this forum
                                  cynar@lemmy.world
                                  wrote last edited by
                                  #36

                                  Given that they are scrabbling around like drug addicts looking for anything they've split, including checking the cracks in the floorboards...

                                  For some models, it's obvious they've long scrapped the erotic fan fic sites!

                                  1 Reply Last reply
                                  0
                                  • T thesharky@piefed.blahaj.zone

                                    Did I? I can't see how.

                                    I don't think web crawlers overloading instances by downloading huge amounts of content and sending thousands of requests is the point of the Fediverse.

                                    But I might be genuinely confused here. Correct me if I'm wrong.

                                    I This user is from outside of this forum
                                    I This user is from outside of this forum
                                    iegod@lemmy.zip
                                    wrote last edited by
                                    #37

                                    The protocol and data are publicly available. Whether or not the use was the point, the mechanism permits it. You shouldn't expect the data not to be accessed.

                                    1 Reply Last reply
                                    0
                                    • T thesharky@piefed.blahaj.zone

                                      Title.

                                      I've noticed that the issues above are becoming increasingly notorious across the entirety of the Fediverse. What's being done to mititage those issues?

                                      K This user is from outside of this forum
                                      K This user is from outside of this forum
                                      korendian64@lemmy.world
                                      wrote last edited by
                                      #38

                                      Making posts and platforms private to users and not search engine indexable. That's about all that can be done.

                                      1 Reply Last reply
                                      0
                                      • T thesharky@piefed.blahaj.zone

                                        Title.

                                        I've noticed that the issues above are becoming increasingly notorious across the entirety of the Fediverse. What's being done to mititage those issues?

                                        irelephant@lemmy.dbzer0.comI This user is from outside of this forum
                                        irelephant@lemmy.dbzer0.comI This user is from outside of this forum
                                        irelephant@lemmy.dbzer0.com
                                        wrote last edited by
                                        #39

                                        The fediverse is open by design (add the header Accept: application/activity+json to any item to get a json representation). Stopping scraping is impossible.

                                        Bots can be stopped with instance applications mainly.

                                        1 Reply Last reply
                                        0
                                        • combatwombat@feddit.onlineC combatwombat@feddit.online

                                          I feel pretty confident, despite a complete lack of evidence, that at least one state actor has had a listener running on the fediverse continuously since the w3c started publishing specs, and I would be surprised if the big llm providers like Anthropic and OpenAI don't run them as well -- they certainly have the resources and motivation to develop them. You're certainly correct that the vast majority of scrapers are attempting to harvest historical data using the web frontend, but those are the scrapers I am least afraid of and I think as a mental model for the average user "assume every post is scraped" is the best stance.

                                          irelephant@lemmy.dbzer0.comI This user is from outside of this forum
                                          irelephant@lemmy.dbzer0.comI This user is from outside of this forum
                                          irelephant@lemmy.dbzer0.com
                                          wrote last edited by
                                          #40

                                          You are right: https://www.404media.co/the-200-sites-an-ice-surveillance-contractor-is-monitoring/

                                          The fediverse and atproto are both easily scraped.

                                          1 Reply Last reply
                                          0

                                          Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                                          Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                                          With your input, this post could be even better 💗

                                          Register Login
                                          Reply
                                          • Reply as topic
                                          Log in to reply
                                          • Oldest to Newest
                                          • Newest to Oldest
                                          • Most Votes


                                          • Login

                                          • Don't have an account? Register

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          • First post
                                            Last post