7 ms·
> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the ju
by est31 2mo ago
> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.
Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?
IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.
Scanning books you own should be legal from a copyright point of view, and not require shredding.
Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.
- croes 2mo agoBooks that are shredded can’t be scanned by competitors.
- jfyi 2mo agoYeah, this is the point. I don't understand the bulk of this conversation. Copyright doesn't matter, the books themselves don't matter. All that matters is that their corpus of training data grows faster than their competitors.
- sethops1 2mo ago> Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright? It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright. https://www.404media.co/ai-companies-are-buying-tons-of-old-books-because-theyre-free-of-ai-slop/ https://www.404media.co/ai-companies-are-buying-tons-of-old-...
- cestith 2mo agoOne doesn’t need to pulp the pages after scanning though. After scanning, they could be rebound and put into a library.
- dpark 2mo agoThat would be a clear case of copyright infringement under current law. You can’t make a copy of a book and then give the original to someone else.
- breakyerself 2mo agoIf it's copyright is expired why not?
- dpark 2mo agoThe books in question are not antique. The Twitter post makes this claim but so far as I can tell it’s not based in fact. 404 Media published a story about this as well and cites a bookseller who notes that all of the books there have sold have had ISBNs (and are thus from 1967 or later and generally would have active copyright). ”very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases” https://archive.ph/9MQrK https://archive.ph/9MQrK
- cestith 2mo agoThis is the necessary context distilled down to be concise. Thank you.
- ACCount37 2mo agoScanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper. What happens to the pages after? No one needs them anymore, so they get mulched and recycled. That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media. The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.
- classified 2mo agoIt's the law, logic doesn't enter into it.
- voakbasda 2mo agoThink about that last point for a moment. Our “rights to read” are diminished significantly with digital works as compared to printed works. Right of resale. Right to lend. In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.
- butlike 2mo agoEvery innovation since the microprocessor isn't worth saving in the grand scheme of things. When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.
- compass_copium 2mo agoI'm personally a fan of more than 50% of children surviving past the age of 6, something that didn't happen until the 20th century.
- 2mo ago
- mc32 2mo agoOld rare books where there are single digit copies should enjoy some sort of patrimonial protection just like museum pieces. You can own them but have the state have the option to buy it if you’re about to significantly deface it or destroy it.
- soco 2mo agoThe only issue I see is, how could you tell which are those books?
- 9dev 2mo agoWe manage to do this for endangered wildlife too without anyone counting every single specimen; why shouldn’t we be able to estimate how rare a book is?
- eru 2mo agoIt's pretty expensive for the wildlife. Most rare books are rare because no one cared enough about them. Ie most rare books are rubbish.
- phoghed 2mo agoAnd if you instituted this, the commenters of this very website would surely decry it as a prime example of government overstep and waste.
- shimman 2mo agoCommentators on this web site work for some of the most evil organizations on the planet and have beliefs that 95% of the population rejects. You can safely ignore the YC cohort of devs and be fine.
- ctoth 2mo ago> Commentators on this web site work for some of the most evil organizations on the planet and have beliefs that 95% of the population rejects. You can safely ignore the YC cohort of devs and be fine. 1. Why are you here? 2. What is the purpose of this comment?
- graemep 2mo ago> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it?
- deleted 2mo ago[deleted]