7 ms·
25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking https://originality.ai/ai-bot-blocking I am betting hund
by jaybna 2y ago
25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking https://originality.ai/ai-bot-blocking
I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites will realize the money they got was trivial as agents start to crush their businesses.
Bill Gross correctly calls this phase of AI shoplifting. I call it the Napster-of-Everything (because I am old). I am also betting that the courts won't buy the "fair use" interpretation of scraping, given the revenues AI companies generate. That means a potential stalling of new models until some mechanism is worked out to pay knowledge creators. (And maybe nothing we know of now will work for media: https://om.co/2024/12/21/dark-musings-on-media-ai/ https://om.co/2024/12/21/dark-musings-on-media-ai/)
Oh, and yes, I love generative AI and would be willing to pay 100x to access it...
P.S. Hope is not a strategy, but hoping something like ProRata.ai and/or TollBits can help make this self-sustainable for everyone in the chain
- cedws 2y agoCloudflare has a toggle for blocking AI scrapers. I don’t think it’s default, but it’s there.
- jaybna 2y agoThey might get into the micro-licensing game too. More power to them.
- kyledrake 2y agoThis just feels like mystery meat to me. My guess is that a lot of legitimate users and VPNs are being blocked from viewing sites, which numerous users in this discussion have confirmed. This seems like a very bad way to approach this, and ironically their model quite possible also uses some sort of machine learning to work. A few web hosting platforms are using the cloudflare blocker and I think it's incredibly unethical. They're inevitably blocking millions of legitimate users from viewing content on other people's sites and then pretending it's "anti AI". To paraphrase Theo Deraadt, they saw something on the shelf, and it has all sorts of pretty colours, and they bought it.
- pixl97 2y ago> I think it's incredibly unethical. The internet isn't built on ethical behavior, unfortunately.
- kyledrake 2y agoI get that a lot of people are opposed to AI, but blocking random IP ranges seems like a really inappropriate way to do this, the friendly fire is going to be massive. The robots.txt approach is fine, but it would be nice if it could get standardized so that you don't have to change it a lot based on new companies (like a generic no llm crawling directive for example).
- input_sh 2y agoIt's not much smarter than just adding user agents to robots.txt manually.
- jpablo 2y agoThey aren't blocking anything. They are just asking nicely not to be crawled. Given that AI companies haven't cared a single bit about ripping of other's peoples data I don't see why they would care now.
- jaybna 2y agoYeah, probably right. If you want a great rabbit hole, look up "Common Crawl" and see how a great academic project was absolutely hijacked for pennies on the dollar to grab training data - the foundation for every LLM out there right now.
- CamperBob2 2y agoIt's hard to envision a greater success for the "great academic project" than what happened. I mean, what else were they trying to accomplish?
- jaybna 2y agoIt was meant to be an open-source compilation of the crawled internet so that research could be done on web search given how opaque Google's process is. It was NOT meant to be a cheap source of data for for-profit LLM's to train on. *edit: added "for-profit"
- CamperBob2 2y ago(Shrug) Multiple not-for-profit LLMs have trained on it as well. If something I worked on turned out to play a significant part in something that turned out to be that big a deal, I'd be OK with it. And nobody's stopping people from doing web-search studies with it, to this day.
- wing-_-nuts 2y agoA number of sites have started outright blocking any traffic that looks remotely suspicious. This has made browsing with a vpn a bit of a pain.
- Workaccount2 2y agoDoing basic copyright analyses on model outputs is all that is needed. Check if the output contains copyright, block it if it does. Transformers aren't zettabyte sized archives with a smart searching algo, running around the web stuffing everything they can into their datacenter sized storage. They are typically a few dozen GB in size, if that. They don't copy data, they move vectors in a high dimensional space based on data. Sometimes (note: sometimes) they can recreate copyrighted work, never perfectly, but close enough to raise alarm and in a way that a court would rule as violation of copyright. Thankfully though we have a simple fix for this developed over the 30 years of people sharing content on the internet: automatic copyright filters.
- jaybna 2y agoSo then copyrighted content scraped is not needed for training? Guess I missed AGI suddenly appearing that reasoned things out all by itself.
- Workaccount2 2y agoNothing builds a better strawman than a foundation started with "So".
- EricMausler 2y agoNo comment on if output analysis is all that is needed, though it makes sense to me. Just wanted to note that using file size differences as an argument may simply imply transformers could be a form of (either very lossy or very efficient) compression.
- Workaccount2 2y agoYou can argue any form of data is an arbitrarily lossy compression of any other form of data. I get your point, but nobody is archiving their companies 50 years of R&D data with and LLM so they can get it down to 10GB. They may have traits of data compression, but they are not at all in the class of data compression software.
- 2y ago
- Kostchei 2y agoUsing the real world- as in vision, 3d orientation, physical sensors and building training regimes that augment the language models to be multidimensional and check that perception, that is the next step. And there is very little shortage of data and experience in the actual world, as opposed to just the text internet. Can the current AI companies pivot to that? Or do you need to be worldlabs, or v2 of worldlabs?
- shanusmagnus 2y agoIronically, if it plays out this way, it will be the biggest boon to actual AGI development there could be -- the intelligence via text tokenization will be a limiting factor otherwise, imo.
- Tossrock 2y agoSome can. Google owns Waymo and runs Streetview, they're collecting massive amounts of spatial data all the time. It would be harder for the MS/OpenAI centaur.
- vidarh 2y agoAll the big players are pouring a fortune into manually curated and created training data. As it stands, OpenAI has a market cap large enough to buy a major international media conglomerate or two. They'll get data no matter how blocked they get.
- aftbit 2y agoIMO this is an underappreciated advantage for Google. Nobody wants to block the GoogleBot, so they can continue to scrape for AI data long after AI-specific companies get blocked. Gemini is currently embarrassingly bad given it came from the shop that: 1. invented the Transformer architecture 2. has (one of) the largest compute clusters on the planet 3. can scrape every website thanks to a long-standing whitelist
- tartuffe78 2y agoWonder if OpenAI is considering building a search engine for this reason... Imagine if we get a functional search engine again from some company just trying to feeding their model generation...
- Art9681 2y agoThe new Gemini Experimental models are the best general purpose models out right now. I have been comparing with o1 Pro and I prefer Gemini Experimental 1206 due to its context, speed, and accuracy. Google came out with a lot of new stuff last week if you havent been following. They seem to have the best models across the board, including image and video.
- HaZeust 2y agoOmnimodal and code/writing output still has a ways to go for Gemini - I have been following and their benchmarks are not impressive compared to the competition, let alone my anecdotal experience in using Claude for coding, GPT for spec-writing, and Gemini for... Occasional cautious optimism to see if it can replace either.
- kibwen 2y ago> Nobody wants to block the GoogleBot This only remains true as long as website operators think that Google Search is useful as a driver of traffic. In tech circles Google Search is already considered a flaming dumpster heap, so let's take bets on when that sentiment percolates out into the mainstream.
- 2y ago
- code51 2y agoWith current state of legal, a real challenge can happen only around 10 years from now. By then AI players will gather immense power over the law.
- glenstein 2y ago>Bill Gross correctly calls this phase of AI shoplifting. I call it the Napster-of-Everything (because I am old). I am also betting that the courts won't buy the "fair use" interpretation of scraping, given the revenues AI companies generate. That means a potential stalling of new models until some mechanism is worked out to pay knowledge creators. To your point, I have wondered whatever became of that massive initiative from Google to scan books, and whether that might be looked at as a potential training source, giving that Google has run into legal limitations on other forms of usage.
- ben_w 2y ago> To your point, I have wondered whatever became of that massive initiative from Google to scan books, and whether that might be looked at as a potential training source, giving that Google has run into legal limitations on other forms of usage. Still around, doing fine: https://en.wikipedia.org/wiki/Google_Books https://en.wikipedia.org/wiki/Google_Books and https://books.google.com/intl/en/googlebooks/about/index.html https://books.google.com/intl/en/googlebooks/about/index.htm... Given the timing, I suspect it was started as simple indexing, in keeping with the mission statement "Organize the world's information and make it universally accessible and useful". There was also reCAPTCHA v1 (books) and v2 (street view), which each improved OCR AI until the state of the art AI were able to defeat them in the role of CAPTCHA systems.
- glenstein 2y agoI don't know what you mean by timing (relative to what?) or "simple indexing" (they scanned the complete contents of books), but I am, and was already aware, of the wiki article and the role of recaptcha. Maybe I wasn't clear, but I was interested in the consequences of the legal stuff. It's not clear from the wiki article what any of this means with respect to the suitability of scans for AI training.
- ben_w 2y ago> I don't know what you mean by timing (relative to what?) or "simple indexing" (they scanned the complete contents of books), but I am, and was already aware, of the wiki article and the role of recaptcha. Timing as in: it started in 2004, when the most advanced AI most people used was a spam filter, so it wasn't seen as a training issue (in the way that LLMs are) *at the time*. As for training rights, I agree with you, there's no clarity for how such data could be used *today* by the people who have it. Especially as the arguments in favour of LLM training are often by comparison to search engine indexing.
- cshores 2y agoIt ultimately doesn't matter because a fairly current snapshot of all of the world's information is already housed in their data lakes. The next stage for AI training is to generate synthetic data either by other AI or by simulations to further train on as human generated content can only go so far.
- pphysch 2y agoHow is synthetic data supposed to work? Broadly speaking, ML is about extracting signal from noisy data and learning the subtle patterns. If there is untapped signal in existing datasets, then learning processes should be improved. It does not follow that there should be a separate economic step where someone produces "synthetic data" from the real data, and then we treat the fake data as real data. From a scientific perspective, that last part sounds really bad. Creating derivative data from real data sounds, for the purpose of machine learning, like a scam by the data broker industry. What is the theory behind it, if not fleecing unsophisticated "AI" companies? Is it just myopia, Goodhart's Law applied to LLM scaling curves? Some MBA took the "data is the new oil" comment a little too seriously and inferred that data is as fungible as refined petroleum?
- RationPhantoms 2y agoWould you trust a ML self-driving algorithm trained on a "digital twin" of a city? I would. I view synthetic training data like a digital twin in which it can provider further control or specified noise to understand from.
- Corrado 2y agoIsn’t this what Tesla does for their driving data? However it would fall apart if they didn’t have real world days to feed into it, right?
- kjkjadksj 2y agoWhat makes you assume your digital twin is actually capturing the factors that contribute to variation in the real data? This is a big issue in simulation design but for ml researchers its hand-waved off seemingly.
- cma 2y agoPeople upload lots from those sites to chatgpt asking to summarize.
- devsda 2y agoThat's still manual and minuscule compared to the amount they can gather by scraping. If blocking really becomes a problem, they can take a page out of Google's playbook[1] and develop a browser extension to scrape page content and in exchange offer some free credits for Chat-GPT or a summarizer type of tool(s). There won't be shortage of users. 1. https://en.wikipedia.org/wiki/Google_Toolbar https://en.wikipedia.org/wiki/Google_Toolbar
- cma 2y agoBefore long people will also continuously use it to watch their screen and act as an assistant, so it can slurp up everything people actually read. People could poison it though with faked browsing of e. G. foreign propaganda stuff made to look like being read from CNN.
- lxgr 2y agoIf you're willing to believe the narrative that there's some sort of existential "race to AGI" going on at the moment (I'm ambivalent myself, but my opinion doesn't really matter; if enough people believe it to be true, it becomes true), I don't think that'll realistically stop anyone. Not sure how exactly the Library of Congress is structured, but the equivalent in several countries can request a free copy of everything published. Extending that to the web (if it's not already legally, if not practically, the case) and then allowing US companies to crawl the resulting dataset as a matter of national security, seems like a step I could see within the next few years.
- jasondigitized 2y agoThe amount of content coming off of YouTube every minute puts Google in a very enviable position.
- 1vuio0pswjnm7 2y agoBill Gross: https://twitter.com/Bill_Gross/status/1859999138836025808 https://twitter.com/Bill_Gross/status/1859999138836025808 https://pdl-iphone-cnbc-com.akamaized.net/VCPS/Y2024/M11D20/7000358580/1732099083-37201167172-hd_L.mp4 https://pdl-iphone-cnbc-com.akamaized.net/VCPS/Y2024/M11D20/... He appears to be criticising "AI" only to solicit support for his own company.
- heavyset_go 2y ago> I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. This is where I'm at. I write content when I run into problems that I don't see solved anywhere else, so my sites host novel content and niche solutions to problems that don't exist elsewhere, and if they do, they are cited as sources in other publications, or are outright plagiarized. Right now, LLMs can't answer questions that my content addresses. If it ever gets to the point where LLMs are sufficiently trained on my data, I'm done writing and publishing content online for good.
- zifpanachr23 2y agoI don't think it is at all selfish to want to get some credit for going to the trouble of publishing novel content and not have it all stolen via an AI scraping your site. I'm totally on your side and I think people that don't see this as a problem are massively out of touch. I work in a pretty niche field and feel the same way. I don't mind sharing my writing with individuals (even if they don't directly cite me) because then they see my name and know who came up with it, so I still get some credit. You could call this "clout farming" or something derogatory, but this is how a lot of experts genuinely get work...by being known as "the <something> guy who gave us that great tip on a blog once". With AI snooping around, I feel like becoming one of those old mathematicians that would hold back publicizing new results to keep them all for themselves. That doesn't seem selfish to me, humans have a right to protect ourselves and survive and maintain the value of our expertise when OpenAI isn't offering any money. I honestly think we should just be done with writing content online now, before it's too late. I've thought a lot about it lately and I'm leaning more towards that option.
- heavyset_go 2y agoAgree with your assessment. I enjoy the little networks of people that develop as others use and share content. I enjoy the personal messages of thanks, the insights that are shared with me and seeing how my work influences others and the work they do. It's really cool to learn that something I made is the jumping off point for something bigger than I ever foresaw. Hell, just being reached out to help out or answer questions is... nice? I guess. It's the little bits of humanity that I enjoy, and divorcing content from its creators is alienating in that way. I'm not a musician, but I imagine there are similar motivations and appreciations artists have when sharing their work. > I work in a pretty niche field and feel the same way. I don't mind sharing my writing with individuals (even if they don't directly cite me) because then they see my name and know who came up with it, so I still get some credit. You could call this "clout farming" or something derogatory, but this is how a lot of experts genuinely get work...by being known as "the <something> guy who gave us that great tip on a blog once". Yup, my writing has netted me clients who pointed at my sites as being a deciding factor in working with me. > I honestly think we should just be done with writing content online now, before it's too late. I've thought a lot about it lately and I'm leaning more towards that option. The rational side of me agrees with you, and has for a while now, but the human side of me still wants to write.
- zifpanachr23 2y agoI agree with you about the fair use argument. Seems like it doesn't meet a lot of the criteria for fair use based on my lay understanding of how those factors are generally applied. See https://fairuse.stanford.edu/overview/fair-use/four-factors/ https://fairuse.stanford.edu/overview/fair-use/four-factors/ I think in particular it fails the "Amount and substantiality of the portion taken" and "Effect of the use on the potential market" extremely egregiously.