5 ms·
Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article. Yes, this is a result of copyright laws. The
by thenewnewguy 1mo ago
Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.
Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.
If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).
- spencerflem 1mo agoIsn’t it being destroyed because it makes the scanning process easier?
- bpodgursky 1mo agoNo. A judge a while ago decided that as long as the physical copy is destroyed, and "transformed" into an electronic copy, you can do the upload. But if you preserve the physical copy after scanning it, you are in violation of copyright because you "copied" the book. That's literally the only reason they are trashing them. It's a legal requirement.
- ashT15 1mo agoAh, so I can buy and scan a dvd, destroy it and then legally distribute the legal copy via torrents. Good to know, because that is what the LLM thieves are doing.
- jagraff 1mo agoDistributing an exact copy of the text would be illegal; if you could get an LLM trained on a book to output the exact copy of the text from the book then that would be illegal as well.
- yywwbbn 1mo agoThey still made a bunch of copies and reused it to train multiple models after making the first copy. Unless they copy and destroy a book each time they use it as as training sample
- schoen 1mo agoCan you find a citation for this? I have heard this claimed rule recently from other people, and I haven't seen this decision (nor do I know what level of court or jurisdiction it might be). This is not a rule that I heard many years ago when working on and adjacent to copyright issues (including book scanning!), although of course the issue has been newly litigated again recently, so there may be new interpretations coming out. Edit: Someone else linked to an order in Bartz v. Anthropic which appears to emphasize that destroying the original copies improved the defendant's position with respect to the fair use analysis. Is that the decision you're thinking of?
- thenewnewguy 1mo agoPerhaps partially? I assume that cutting the pages out of the spin makes them easier to scan at least partially. That said I have no insider knowledge of this type of operation so I don't know how much easier that actually makes it. But ultimately, it's a moot point, because the legal requirement means the books must end up destroyed. Even if the people at Amazon wanted to scan the books in a way that required no destruction at all, it's not currently (legally) possible for them to do so, so they might as well take the easy way out today.
- pessimizer 1mo agoNon-destructive scans are at least 10x as much, and tend to lower quality. https://software.annas-archive.gl/AnnaArchivist/annas-archive/-/work_items/223 https://software.annas-archive.gl/AnnaArchivist/annas-archiv... Cutting the pages out makes them machinable. Non-destructive scans involves gently turning pages, and paying a lot of attention to the state of the spine. Destructive scans involve guillotine cutting the spine off, scanning the covers by hand, putting the pages into a hopper, clamping them in and hitting a button. While that book is scanning, you're already cutting the spine off the next book. If the machine jams, try to work the jam out gently, scan the pieces, and let the computer stitch it together.
- Apocryphon 1mo agoPerhaps this is a use case that's worth investigating, to invent better, less-destructive scanning processes! Or some way to rebind them afterwards.
- throwatdem12311 1mo agoThey could just not do the evil thing. (This is why I will never be a billionaire)