8 ms·
It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units
by cladopa 27d ago
It is not a big deal. Since the invention of the printing press any important
book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
- brightball 27d agoWhenever my wife wants to visit antique stores, I always look for old books. I have found several 100+ year old gems.
- pshirshov 27d agoRead more about this. Depends on the definition of the "big deal" but from what I can understand the problem is that they buy rare things - which exist in just several copies - and they tend to buy _all_ copies.
- dataflow 27d agoWhere did you see they tend to buy all the copies? This comment is the first time I've heard of this.
- wmeredith 27d agoI'd also be curious about the provenance of that statement. Why would they buy all copies? What would be the purpose of scanning multiple copies?
- dataflow 27d agoI could see buying multiple copies being useful to mitigate problems, like damage. But buying all the copies is categorically different and I cannot imagine why they would attempt that, except perhaps to prevent their competition from getting a hold of the same text?
- p-e-w 27d agoIt’s just another lie of the type these threads tend to be filled with nowadays. Of course they aren’t buying “all copies”, and that wouldn’t even be possible in most cases since such books are usually flea market/attic material and most copies aren’t for sale (or even catalogued) to begin with. I’d be interested to learn who comes up with such lies though. Is it really just random people venting their frustration, or some kind of organized astroturfing operation?
- diseasedyak 27d agoIt really does seem like an organized operation, given that it's so prevalent and they all seem to be in lockstep with their specious claims.
- wongarsu 27d agoI wouldn't be surprised to find out that this outrage is fueled by the same actors as the AI data center water outrage. Whoever they are
- vavos 27d agoI think these type of lies usually come about as a result of a game of broken telephone and things get exaggerated
- sajithdilshan 27d agowhat is your source?
- mistercow 27d agoMost old books that are rare and unpreserved are so because their value is marginal, so nobody has bothered to collect and preserve them. But where did you hear that they’re buying “all copies”? And to what end?
- dbspin 27d agoThis is a classic mistake. We have no way of estimating the future value of a given book. It's perceived current value (a large part of which is simply obscurity) may be low. But it's future value - to historians, ethnographers, to researchers seeking a specific fact or example of language use or a hundred other things - is literally inestimable. To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved. Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.
- famouswaffles 27d ago>Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc. I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame.
- dbspin 27d agoOK... I'm going to assume good faith even though your wording makes it somewhat unlikely. Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing ecosystem or the state of a particular political situation at a given time. Essentially all textual primary sources for events which are not currently considered important - but which we have no way of estimating the future importance of. One can continue to create countless counterfactual examples in this vein. My overall point is we cannot know what may be useful or even essential in the future, and knowledge should be preserved under the assumption that it is likely to be. Historians frequently refer to this paradox - how everyday aspects of life are frequently not explicitly documented, since they're so obvious to the communities or communities of expertise that observe and carry them out. So it's actually incredibly important to preserve what seems like ephemera. Hell we couldn't have AI training at all if we lacked the corpus of existing written literature - but there was no way any author could have anticipated this future utility more than a couple of decades ago.
- pfdietz 27d agoHere we have another entry in the long list of "things described on the Internet that never happened".
- wasmperson 27d agoI was also skeptical of this claim but managed to find someone who explains it: https://downtownbrown.substack.com/p/five-fallacies-ai-and-destructive-scanning-of-books https://downtownbrown.substack.com/p/five-fallacies-ai-and-d... It's not that individual companies buy all copies of a given book, but that there's more than one book scanning company, and they aren't sharing the scans with each other. The result: books that were rare but nevertheless easy to find for purchase (thanks to the internet) are now vanishing off of the market, becoming de facto no longer accessible to the public.
- quietsegfault 27d ago[dead]
- demibabs 27d agoGood article, but I still feel unsatisfied because even it cannot find an example of a book that’s actually been lost because of the destructive scanning frenzy (it only lists books that hypothetically could be lost because there’s not many physical copies available for sale online.). If anyone has an example, I’d love to hear it.
- voidhorse 27d agoSince we don't know what was actually purchased and what was actually destroyed, how do you expect us to furnish an example? This would require the destroyers to admit it, and beyond that it would require all of them to admit it since more than one of them might have been responsible for the extinction of one text. Seeing as they were already keeping this operation under wraps, I don't see that happening. "possibly extinct because no copies available online" is probably the best we can do. The distributed nature of the problem and the utter lack of transparency are huge factors here too.
- demibabs 26d agoThese companies are buying from small book stores; surely these collector types would know if they lost any one-of-a-kinds? Or at least someone would be keeping track of extremely rare books disappearing (especially now since this matter has been public for weeks). Regardless I feel like they gotta figure this out for optics reasons. “So and so books are lost forever to Anthropic’s servers” is much more outrageous than “Anthropic is destroying a bunch of books that have other copies” imo.
- shiandow 27d agoSomehow I don't think they're looking for the books that have been copied over and over.
- quietsegfault 27d agoWhy do you think that? Do you have evidence, or is this just a hunch? Why would Amazon waste money on uber rare books when there are thousands and thousands of not-so-rare books that could serve the exact same purpose?
- shiandow 27d agoFor one they've already used the entire library genesis. Anything not in there is going to be obscure in some capacity.
- sajithdilshan 27d agoExactly, also all those physical books would anyways get molded, eaten by moths or just naturally decay. It's not like the AI companies are obliterating every copy of every single book.
- etdznots 26d agoWe should actually thank the AI companies for helping us get rid of our trash! Thank you anthropic! Thank you OpenAI! Thank you Google!
- nloomans 27d ago> a digital copy with the ability of doing millions of copies is stored somewhere somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers” the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies. > If you go to a recycling centre, the garbage to quality ratio is over 100 or more. archivists keep everything, because we don't know right now what will be important 100 years from now.
- radu_floricica 27d ago> somewhere were we can't access it By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther. Plus having the info part of a LLM makes it immediately available to literally billions. I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.
- jjulius 27d ago>Plus having the info part of a LLM makes it immediately available to literally billions. Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong. If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions". And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.
- SkyBelow 27d agoDepends upon what you want. For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book. The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly. But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great. That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.
- rvz 27d agoFirst of all, it IS destroyed and it is a big deal. Hardcover copies of books especially 1st - 2nd edition ones (even with mistakes) are rarer than digital scans. Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway. Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why? [0] https://www.bodleian.ox.ac.uk/services/research-partnerships/digitisation-project https://www.bodleian.ox.ac.uk/services/research-partnerships...
- kccqzy 27d agoWhy wonder? The answer is abundantly clear if you follow the news. If a book is copyrighted under U.S. law, scanning and destroying counts as a format conversion which qualifies it as fair use, so there is no need to negotiate with copyright holders. See Judge William Alsup’s decision. If Anthropic did not destroy the books after scanning it would have not won the lawsuit, and scanning would be illegal. If a book is already out of copyright then of course they do not have to destroy it afterwards.
- deleted 27d ago[deleted]
- quietsegfault 27d agoWhy is it a big deal? Do you think there's no difference between the books curated at Oxford University and the crap that Amazon is buying?
- voidhorse 27d agoWhy would amazon buy "crap"? Surely they want their model to succeed and they want to train it on valuable input, no? They have more than enough resources to determine whether or not the books are worth buying. They have been in book selling for a long time.
- quietsegfault 26d ago[flagged]
- afpx 27d agoI think you may be greatly underestimating the long tail. Several times a year I read sources that reference older books that I can't find online. When I am able to locate them, they often cost at least several hundred dollars, sometimes into the 10s of thousands.
- quietsegfault 27d agoDo you think that you are somehow special and unique in needing these books? If the books cost in the 10s of thousands, then there's obviously value to other people. I've seen no evidence that Amazon or others are buying $10k books to scan into their corpus. All evidence I've seen is that they're scanning cheap books with no current value and no clear use to people today. I have volunteered with a library, and probably threw hundreds of books over a couple week engagement from a university library into a shredder at the direction of a professional, academic librarian. Libraries are constantly culling books, the EXACT category books we're talking about here (old, never-read). This is happening at a much larger scale, so I would recommend railing against university librarians in addition to the AI juggernauts.
- hughlilly 27d ago> I've seen no evidence that Amazon or others are buying $10k books to scan into their corpus. Have you seen evidence that they’re buying only readily available books that are plentiful on the market?
- afpx 26d agoI'm just stating that not everyone is reading the most popular million books. And, there are many, many millions of books in the long tail. A few years ago, after running into this issue several times, I looked into the economics of it to see if there was a opportunity to republish digitally. I found the sales data for some, the ones that had value were selling on average for around $50-150 dollars. Because the typefaces weren't modern, OCR wasn't scalable. Because only 20-40 were transacted each year, it wasn't worth the time. What kind of books are they? In my case mostly historical documents by some relatively unimportant person who was highly important for a very, very niche subject. They still contain valuable, irreplacable information. An analogy: imagine you discover a really cool video game from 20 years ago. You love it and want to find the developer's previous work. But, you find out that the company was purchase by another company which was purchase by another, etc, etc. Sure, maybe the game still exists in some digital vault. But, more than likely it's gone because old things only seem have value these days if some influencer broadcasts it. It's interesting that libraries are purging them. Several times, I've found that the only available copies were at some random university rare book collection in middle of no where, and basically impossible to access unless one fly in - which isn't worth it. They are often donated by a benefactor and stuck there - which is why they're never read. It's not that the content isn't valuable.
- jll29 27d agoBeware that the notion of "quality" is entirely different for AI companies: they don't seek entertainment, but sentences in a language to train an LLM.
- mannyv 27d agoI have books, but they are just objects. They're nice objects, but just objects. Fetishizing books isn't going to help. In fact, most of those "rare" books don't sell because nobody wants them. The AI companies are making them even more rare, so the booksellers should be thankful.
- GreenLightGo 27d agoHonestly, it’s easier to find a good movie than a good book, because books are way cheaper to publish. These days, the quality of pretty much all kinds of content has become a problem...
- TFNA 27d ago> any important book has been duplicated by thousands, tens of thousands or even million of units. Books in the former USSR display their print runs on the last page. "Important book" is a vague and arbitrary term, but rhere are works in whole fields (e.g. history, archaeology, linguistics, ethography) that any scholar would consider key references, and as few as 100 copies were printed. The shadow libraries have made a lot available to the whole world. It would suck if private corporations scan and shred remaining copies of these before the shadow libraries can get a scan.
- the-mitr 26d agoOur little project endeavours to digitise books from the erstwhile Soviet state which were published in many languages We have, over the last 15 years, acquired/borrowed from libraries, and scanned a couple of thousand books on all topics of interest. All of them are out of print. Some of the physical copies were have, especially in indicate languages might be some of the few surviving ones https://mirtitles.org/ https://mirtitles.org/ https://archive.org/details/@mir-titles https://archive.org/details/@mir-titles
- pibaker 27d ago> any important book has been duplicated by thousands, tens of thousands or even million of units. It is common for academic books to have publication runs in the low three digits. You may argue these books are not important. But how do we know if we fail to preserve it?
- quietsegfault 27d agoIs there evidence that these mythical low-print-run books are being purchased by Amazon and friends for destructive scanning? I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense.
- etdznots 26d ago> Is there evidence that these mythical low-print-run books are being purchased by Amazon and friends for destructive scanning? Is there evidence that they aren’t? All of the common works are already on libgen, if these books are worthless repetition they will contribute little to the training corpus, labs want high quality interesting texts and they have unlimited budget to spend on it. Paying $300 for something rare with millions of tokens of interesting and original text for training is definitely fucking worth it, spending $1,200 to get all the copies and block your competitors from getting it is most definitely worth it! > I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense. Do you think it would be OK to rail against for example, a genocide if I wasn’t substantially contributing to some effort to stop it? This sophistic (and uninteresting) bit of rhetoric boils down to: if youre not trying to fix it yourself, dont complain! (reminder that one of the basic principles of a democratic society is that each person has some concern outside of their personal affairs)
- quietsegfault 25d agowe're done - straight to genocide, later buddy.
- arttaboi 27d agoWith all due respect, I would say it wouldn’t hurt not to downplay this.
- XorNot 26d agoSure but are we going to do anything sensible about it or is it just a convenient culture war vector? This is happening because copyright means you can't scan these without destroying them as a format conversion. No one's felt compelled to try and fix that so we can do this sort of digital archival and preservation, and copyright allows works to be frozen and undistributable because a claim might exist for decades without any actual use (I.e. the number of games which get stuck in legal limbo). If the only desire is to sling mud at AI companies but not try and improve the legal situation, then it's worse then useless because there's no intent to stop it - in fact stopping it would remove a useful outrage tool. The idea in the title here is point 1 of the blind leading the blind: it's illegal to scan and store these books without destroying them in many jurisdictions.
- mtkd 27d ago>It is not a big deal have you ever held and read an old book?
- NoMoreNicksLeft 27d ago>Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units. No, every popular book has been duplicated thousands of times. This is not the same thing as important. They're orthogonal. When an important book is popular, it is safe. When it is not, it is in danger. Only fools assume that important books are recognized often enough to become popular.
- anigbrowl 26d agoI don't know why you think all these irrelevant remarks somehow support your core thesis.