7 ms·
Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I
by titzer 27d ago
Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space.
It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.
- nazgulsenpai 27d agoThey are also an AI company now. Why would they stop?
- al_borland 27d agoThey could use them as training data, without providing access the actual books.
- jimmaswell 27d agoAt least we would all benefit from the books this way, so long as legal nonsense keeps the scans unavailable to the public.
- al_borland 27d agoI don’t think locking the content of rare books away in the hands of corporations who only give us access to tools trained on the books, and not the actual book, is a good path forward. This doesn’t incentivize them to be good stewards of this data and making anything in the public domain available. It incentivizes less access to the source material, having to blindly trust their tools, and is effectively automating plagiarism.
- JKCalhoun 27d agoAlso, why would they ever share their collection? Book scans, secreted away, are worthless to the public.
- greyw 27d agoIt's data for their AI pipeline. Basically digital gold.
- ktm5j 27d agoI think you're missing the point. Maybe I'm wrong, but I'm pretty sure they're trying to point out that this book scanning doesn't need to be destructive. I'm not sure if AI companies are using a scanning method that damages the book or not, but they destroy the books after scanning to avoid copyright issues (ie they aren't duplicating the books). This Google project seems to demonstrate that this isn't actually necessary, and that the AI companies are just doing it out of laziness.
- shagie 27d agoIf it isn't scanned destructively, is it a liability? Can you do anything else with the book? What are its costs for storage in a way that retains the value of the book? If the assets of the warehouse are sold to another company (see also https://paizo.com/blog/paizo-restructuring-a-difficult-update-about-our-future https://paizo.com/blog/paizo-restructuring-a-difficult-updat... ), what are your obligations for the format shifted copy that you retain? These questions imply that there's a liability that exists when retaining the original that has little value to the company. And they (the books) aren't assets that can be resold. It's easier (and cheaper), has no ongoing costs for physical storage, and answers those questions without creating legal entanglements for the future company.
- ktm5j 27d agoQuestions don't imply anything, you're just asking things. I don't have the answers to your questions, but the point is that Google was able to scan books without destroying them while still avoiding legal consequences (with some effort).
- shagie 27d agoThe physical books that were scanned by Google were returned to the libraries. https://btaa.org/library/programs-and-services/book-search/faq https://btaa.org/library/programs-and-services/book-search/f... > Will scanning harm the books? > No. Google developed innovative technology to scan the content without harming the books. Any book deemed too fragile will not be scanned by Google, but may be treated by expert library staff. Once scanned, all print volumes are returned to the library collections. That was an inherently different goal (borrow the books from the library, scan them, and return them) than the Bartz v. Antrophic ruling. https://cases.justia.com/federal/district-courts/california/candce/3%3A2024cv05417/434709/231/0.pdf https://cases.justia.com/federal/district-courts/california/... > Storage and searchability are not creative properties of the copyrighted work itself but physical properties of the frame around the work or informational properties about the work. See Texaco, 802 F. Supp. at 14 (physical), aff’d, 60 F.3d at 919; Google, 804 F.3d at 225 (informational); Sony Corp. of Am. v. Universal City Studios, Inc. (“Sony Betamax”), 464 U.S. 417, 447 (1984) (rightful interests). In Texaco, the court reasoned that if a purchased scientific journal article had been copied “onto microfilm to conserve space, this might [have been] a persuasive transformative use.” 802 F. Supp. at 14 (Judge Pierre Leval), aff’d, 60 F.3d at 919 (reducing “bulk[ ]” “might suffice to tilt the first fair use factor in favor of Texaco if these purposes were dominant“). In Google Books, the court reasoned that a print-to-digital change to expose information about the work was transformative. Google, 804 F.3d at 225 (Judge Pierre Leval). And, in Sony Betamax, the Supreme Court held that making a recording of a television show in order to instead watch it at a later time was copying but did not usurp any rightful interest of the copyright owner. 464 U.S. at 447, 455. Important to the Supreme Court’s reasoning was the expectation that most such copiers would not distribute the permanent copies of the work. Finally, in A&M Records, Inc. v. Napster, Inc., our court of appeals recognized the reasoning just explained, and therefore rejected by contrast a digitization effort that was touted as space-shifting but in fact resulted in the multiplication of copies shared with outsiders through a file-sharing service. 239 F.3d 1004, 1019 (9th Cir. 2001), aff’g in this part 114 F. Supp. 2d 896, 912–13, 915–16 (N.D. Cal. 2000) (Judge Marilyn Hall Patel) (citing Sony Betamax and Texaco). > Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others). --- The AI training isn't borrowing from libraries and returning from libraries. Instead, it is buying a book, format shifting, and retaining that format shifted version from its own use. The company can't do anything else with the book once they've format shifted it. They can't donate it and they can't resell it. In that light, destructively scanning the book is the best option. There is no value in trying to non-destructively scan it because otherwise all it would do is sit in a warehouse and cost money to pay for storage of something they can't sell.
- dblohm7 27d ago> It's great they did this, but the Google that is today cannot be trusted with data of public value anymore. They never should have been trusted in the first place.