8 ms·
The broader problem of original sources not being given credit in a way that rewards them remains. Websites owners are paying to host their content so that spid
by dvduval 4mo ago
The broader problem of original sources not being given credit in a way that rewards them remains. Websites owners are paying to host their content so that spiders can come and crawl them and index it into the AI and then if they’re lucky, they might get a citation, but otherwise there’s very little reward for being a provider of content. And of course, this is something that’s getting worse and worse. Why look at a website when it’s all in AI? And then the counter to that is maybe we need to start closing the website to crawlers and put everything behind a login.
- Ensorceled 4mo agoWorse, the constant AI scraping is actually costing content providers additional money for no return. At least Google/Bing/Yahoo scraping would then be used to provide links back to your content.
- fiedzia 4mo ago> At least Google/Bing/Yahoo scraping would then be used to provide links back That doesn't work anymore. Google provides AI generated summary, nobody looks at the original site.
- bolangi 4mo agoNot only costing money. Constant AI scraping constitutes a denial-of-service attack that has brought down websites.
- devsda 4mo agoHow do you distinguish Google/MS scraping for Gemini/Copilot vs Google Search/Bing? In the case of Google, the UA is the same and you are entirely at their mercy to honor the Google-Extended instructions in robots.txt Google has further complicated it with new search announcement blurring lines between regular search and AI search. And AI likes to not honor any licenses or instructions when it is hungry for training material. It is once again an example of Google using its dominant position to abuse and promote cross functional products.
- cute_boi 4mo agoIf company like Meta are downloading pirated books etc.. to train their AI, they will surely honor robots.txt.
- johneth 4mo agoI wouldn't be surprised if there isn't some sort of legal action against Google, the monopoly, to make the distinction in how their crawlers use scraped content.
- motbus3 4mo agoAbout a year ago OpenAI crawled and go DDOS level the company I work. Even despite the robots.txt not allowing it, and despite some recaptcha we could assemble in time. We found our data in the outputs of their models but who can do anything about it...
- rastrojero2000 4mo agoLawyers can. As long as that data is actually yours I mean, in a strictly legal sense.
- kibwen 4mo ago> We found our data in the outputs of their models but who can do anything about it... If the crawlers refuse to voluntarily respect your robots.txt, then you are well within your rights to poison their data.
- hajile 4mo agorobots.txt seems like it should be a legally-binding terms of service which would make them outright copyright infringing. Sue for $180,000 per infringement which should be calculated for each illegal API call.
- throw1234567891 4mo agoWas your robots txt written by a lawyer? Does it hold up in the court?
- wang_li 4mo agoIt doesn't matter. Robots.txt is not a license, it's a set of computer parsable directives of how programs should access your site. The actual license doesn't have to be written for computers to parse to be legally binding. A person should be able to write in a terms of use or license page on their website that says "do not include any content from this website in your AI training data. if you do you will be billed $100 billion dollars." And it should be enforceable. It just turns out that nerds like to say "oh that would be too hard or too expensive, so we're going to ignore it."
- wolttam 4mo agoI’ve been thinking of a proof-of-work scheme for accessing content where you effectively need to mine some crypto for the author, but, this idea might not fly today
- microtonal 4mo agoBut that will be a hassle for human visitors as well. A web doing proof-of-work to browse, will be a disaster for phones with their limited batteries, etc.
- odo1242 4mo agoTo be specific, it would be more of a hassle for human visitors than for the AI companies with infinite money and specialized browsers.
- wolttam 4mo agoThe idea would be that AI companies would still be forced to do this proof of work. Anubis proved the idea
- odo1242 4mo agoI don’t think Anubis proves the idea much though. I feel the main reason it’s worked is that AI companies haven’t yet tried to bypass it. And AI companies still scrape Anubis protected websites, it just forces them to not DDOS the website
- chii 4mo agoor you know, just charge for your content if you believe it to be valuable enough for the fee being charged.
- wolttam 4mo agoYes, but that tends to limit the reach of your content. Hence why a lot of people reach for ads. Between seeing ads and doing a little bit of proof-of-work for the author, I'd choose the latter.
- aaarrm 4mo agoIs it possible able to host your website in a way so that it couldn't be found via search engines (and thus wouldn't be crawlable I hope)? I know this has repercussions on findability, but if that wasn't a concern, I'm curious how one might circumvent getting crawled.
- MontgomeryPy 4mo agoYou could just put your website content behind its own chat interface. The crawler would just see a form input for a prompt.
- trinari 4mo agorobots.txt is a way of leaving the door unlocked but kindly asking bots to stay outside.
- account42 4mo agoWhich in a law-abiding society should be enough. It's also how we do things in the real world in many cases - i.e. here you can just write on your mailbox "no ads" and companies have to respect that. Even when we do actually put physical locks on things they are mostly there to show that someone breaking in did so intentionally and not at all designed to prevent motivated attackers.
- dpark 4mo ago> here you can just write on your mailbox "no ads" and companies have to respect that Where do you live? In the US it’s actually illegal for anyone except the USPS to deliver to a mailbox.
- dpark 4mo agoYou might be interested to know that entering an unlocked door into a space you do not have permission to be in is still illegal.
- throw1234567891 4mo ago
- spacechild1 4mo agoIt's actually costing them money/time! A friend of mine is a sysadmin at a university and he constantly has to deal with AI crawler DDoS-ing his servers. He said Anthropic is actually one of the worst offenders. These AI companies are really just a gross example of the motto "Socialize the costs, privatise the profits". It's disgusting!
- WarmWash 4mo agoIt's never been a problem with people ad-blocking for the last 20 years, why is it suddenly a problem now? We've been celebrating denying creators revenue for decades... Maybe this is just the internet hypocricy of "When I do it, it's good, when they do it, it's bad".
- qotgalaxy 4mo ago[dead]
- onedognight 4mo agoChoosing not to look at something is not denying anyone anything.
- WarmWash 4mo agoChoosing not to look at an ad, and blocking it are different things. One is totally ok, the other incurs a monetary loss on the creator. Those services aren't free to run, and the content doesn't take zero time to create. It also incentivizes creating content focused on those who cannot figure out ad blocking.
- mixmastamyk 4mo agoInteresting. I suppose the main difference is that we’re ants compared to an 800 pound gorilla.
- u_fucking_dork 4mo agoPeople usually point at the scale when this discussion comes up, in my experience. These companies are doing something at a huge scale spending tons of money to do it so the potential harm is greater. People can easily justify their own piracy because it’s small scale. Even when they organize, create a whole software and tooling ecosystem around pirating media to stick into jellyfin or plex. AI still did it bigger and worse and is bad, what I’m doing is not so bad because I wasn’t going to buy the movie anyway, etc.
- 4mo ago
- internet2000 4mo agoPerhaps we should go back to back when the internet was about sharing information you liked, not about credit or making money on "content".
- throw1234567891 4mo agoYou are there today, but some are unhappy that others don’t share the same sentiment.
- sumeno 4mo agoOk, AI companies first then since they are some of the biggest offenders
- gabbagool 4mo agoI agree with this whole heartedly. What's the point of even having copyright law at this point? What's even crazier to think about is that to use the latest versions of these models for which you supplied training data, you have to pay hundreds of dollars a month. I would love to get a settlement check proportional to my model weights. Even if it's $0.10, at least everyone out there will get what they're owed.
- throw1234567891 4mo agoNo, you don’t have to. There are open weight models you can download and use for free. Many people choose the subscription model but it’s not necessary. And latest doesn’t mean greatest, it’s just most up-to-date.
- rickydroll 4mo agoFrom my perspective, everybody trains on the knowledge and experience of those who came before. AI just does the same thing at scale. I do not value copyright. All it does is give you standing to sue if somebody reproduces your work. It does not differentiate or account for parallel creation. I cannot count how many times I have "created" something, only to find it in a research paper later. Part of the reason I think copyright has no value is that, in general, individual copyright owners don't have the deep pockets necessary to sue someone who violates their copyright. If anyone is violating the spirit of copyright, it's corporations that insist you assign your work over to them as a work for hire, or outright ignore your copyright. (looking at you, Disney's Atlantis). A significant benefit of AI that doesn't get talked about enough is that AI has a much greater reach over all the information it was trained on and can draw connections that would be invisible to someone operating at the human scale.
- b00ty4breakfast 4mo ago>Why look at a website when it's all in AI? well, at least in the case of google, I'm pretty sure that's the point. Or at least, they are doing things that would seem to be moving towards being an oracle with all the answers and not the signpost that points you in the right direction. The destination rather than the gateway.
- philipov 4mo agoremember AMP?
- b00ty4breakfast 4mo agoHoly cow, I never even thought about that in relation to AMP. It's not a new thing, then.
- neop1x 4mo ago>> put everything behind a login May not always work. I then click on back button and look for the info elsewhere and in most cases I find it. Same with paywalled websites. If you are ok with a small audience (or you provide a unique content) then it makes sense. But I think in most cases you just cut off a lot of people this way and actually you can simply stop creating content if you don't want consumers of it and let others provide the content.