7 ms·
In Chinese, it always added something like "For study/research purpose only. Please delete after 48 hours." This is what those volunteers added in subtitles of
by ttflee 1y ago
In Chinese, it always added something like "For study/research purpose only. Please delete after 48 hours." This is what those volunteers added in subtitles of (pirated) movies/shows.
- codedokode 1y agoFair, if AI companies are allowed to download pirated content for "learning", why ordinary people cannot.
- snickerdoodle12 1y agoThere is so much damning evidence that AI companies have committed absolutely shocking amounts of piracy, yet nothing is being done. It only highlights how the world really works. If you have money you get to do whatever the fuck you want. If you're just a normal person you get to spend years in jail or worse. Reminds me of https://www.youtube.com/watch?v=8GptobqPsvg https://www.youtube.com/watch?v=8GptobqPsvg
- 4gotunameagain 1y agoIf you owe the bank $1,000 you have a problem. If you owe the bank $100,000,000 the bank has a problem. We live in an era where the president of the United States uses his position to pump crypto scams purely for personal profit.
- kyleee 1y ago10% for the big don
- alphan0n 1y agoNo one (in the US) has been jailed for downloading copyrighted material.
- snickerdoodle12 1y agohttps://en.wikipedia.org/wiki/Aaron_Swartz https://en.wikipedia.org/wiki/Aaron_Swartz And the US is not the only jurisdiction
- gruez 1y agoThat's not the same as piracy though. He wasn't downloading millions of scientific papers from libgen or sci-hub, he was downloading them directly from jstor. Indeed, none of his charge was for copyright infringement. It was for stuff like "breaking and entering" and "unauthorized access to a computer network".
- snickerdoodle12 1y agoThe exact same charges could apply to the AI scrapers illegitimately accessing random websites.
- gruez 1y agoPart of the accusation comes from the fact that Swartz accessed the downloads through a MIT network closet, which AI companies wasn't doing. The equivalent to that would be if openai broke into a wiring closet at Disneyland to download Disney movies.
- snickerdoodle12 1y agoThe CFAA is vague enough to punish unauthorized access to a computer system. I don't have an example case in mind, but people have gotten in trouble for scraping websites before while ignoring e.g. robots.txt
- gruez 1y agoThe CFAA might be vague, but the case law on scraping pretty much has been resolved to "it's pretty much legal except in very limited circumstances". It's regrettable that less resourced defendants were harassed before large corporations were able to secure such rulings, but the rulings that allowed scraping occurred before AI companies' scraping was done, so it's unclear why AI companies in particular should be getting flak here.
- shadowgovt 1y agoThere's actually a lot of court activity on this topic, but the law moves slowly and is reluctant to issue injunctions where harm is not obvious. It's more that the law about "one guy decides to pirate twelve movies to watch them at home and share with his buddies" is already well-settled, but the law about "a company pirates 10,000,000 pieces to use as training data for an AI model (a practice that the law already says is legal in an academic setting, i.e. universities do this all the time and nobody bats an eye)" is more complicated and requires additional trials to resolve. And no, even though the right answer may be self-evident to you or me, it's not settled law, and if the force of law is applied poorly suddenly what the universities are doing runs afoul of it and basically nobody wants that outcome.
- Workaccount2 1y agoThere is a distinction that must be made that very few people do, but thankfully the courts seems to grasp: Training on copyright is a separate claim than skirting payment for copyright. Which pretty much boils down to: "If they put it out there for everyone to see, it's probably OK to train on it, if they put it behind a paywall and you don't pay, the training part doesn't matter, it's a violation."
- snickerdoodle12 1y agoSo if I download copyrighted material like the new disney movie with fansubs and watch it for training purposes instead of enjoyment purposes it's fine? In that case I've just been training myself, your honor. No, no, I'm not enjoying these TV shows. Because it's important to grasp the scale of these copyright violations: * They downloaded, and admitted to using, Anna's Archive: Millions of books and papers, most of which are paywalled but they pirated it instead * They acquired Movies and TV shows and used unofficial subtitles distributed by websites such as OpenSubtitles, which are typically used for pirated media. Official releases such as DVDs tend to have official subtitles that don't sign off with "For study/research purpose only. Please delete after 48 hours" or "Subtitles by %some_username%"
- 1y ago
- CamperBob2 1y agoNo, if you revolutionize both the practice and philosophy of computing and advance mankind to the next stage of its own intellectual evolution, you get to do whatever the fuck you want. Seems fair.
- recursive 1y agoHm. Not a given that it's an advance.
- Nevermark 1y agoI get the common cynical response to new tech, and the reasons for it. We wish we lived in a world where change was reliably positive for our lives. Often changes are sold that way, but they rarely are. But when new things introduce dramatic capabilities that former things couldn't match (every chatbot before LLMs), it is as clear of an objective technological advance as has ever happened. -- Not every technical advance reliably or immediately makes society better. But whether or when technology improves the human condition is far more likely to be a function of human choices than the bare technology. Outcomes are strongly dependent on the trajectories of who has a technology, when they do, and how they use it. And what would be the realistic (not wished for) outcome of not having or using it. For instance, even something as corrosive as social media, as it is today, could have existed in strongly constructive forms instead. If society viewed private surveillance, unpermissioned collation across third parties, and weaponizing of dossiers via personalized manipulation of media, increased ad impact and addictive-type responses, as ALL being violations of human rights to privacy and freedom from coercion or manipulation. And worth legally banning. Ergo, if we want tech to more reliably improve lives, we need to ban obviously perverse human/corporate behaviors and conflicts of interest. (Not just shade tech. Which despite being a pervasive response, doesn't seem to improve anything.)
- CamperBob2 1y agoAt the risk of stepping on a well-known land mine around here, how'd you do on the IMO problem set this year?
- NoMoreNicksLeft 1y agoThe dead corpses of filmmakers and authors and actors are buried in unmarked graves out behind those companies' corporate headquarters. Unimaginable horror, that piracy. Why has no one intervened? >If you're just a normal person you get to spend years in jail or worse. Not that I'm a big fan of the criminalization of copyright infringement in the United States, but who has ever spent years in jail for this? Besides, if it really bothered you, then we might not see this weird tone-switch from one sentence to the next, where you seem to think that piracy is shocking and "something should be done" and then "it's not good tht someone should spend time in jail for it". What gives?
- snickerdoodle12 1y ago> Besides, if it really bothered you, then we might not see this weird tone-switch from one sentence to the next, where you seem to think that piracy is shocking and "something should be done" and then "it's not good tht someone should spend time in jail for it". What gives? What a weirdly condescending way to interpret my post. My point boils down to: Either prosecute copyright infringement or don't. The current status quo of individuals getting their lives ruined while companies get to make billions is disgusting.
- brookst 1y ago> Either prosecute copyright infringement or don't This is the absolute core of the issue. Technical people see law as code, where context can be disregarded and all that matters is specifying the outputs for a given set of inputs. But law doesn’t work that way, and it should not work that way. Context matters, and it needs to. If you go down the road of “the law is the law and billion dollar companies working on product should be treated the same as individual consumers”, it follows that individuals should do SEC filings (“either require 10q’s or don’t!”), and surgeons should be jailed (“either prosecute cutting people with knives or don’t!”). There is a lot to dislike about AI companies, and while I believe that training models is transformative, I don’t believe that maintaining libraries of pirated content is OK just because it’s an ingredient to training. But insisting that individual piracy to enjoy entertainment without paying must be treated exactly the same as datasets for model training is the absolute weakest possible argument here. The law is not that reductive.
- deleted 1y ago[deleted]
- shadowgovt 1y agoIANAL, but reading a bit on this topic: the relevant part of the copyright law for AI isn't academia, it's transformative work. The AI created by training on copyrighted material transforms the material so much that it is no longer the original protected work (collage and sampling are the analogous transformations in the visual-arts and music industries). As for actually gathering the copyrighted material: I believe the jury hasn't even been empaneled for that yet (in the OpenAI case), but the latest ruling from the court is that copyright may have been violated in the creation of their training corpus.
- gruez 1y ago>why ordinary people cannot They can. I don't think anyone got prosecuted for using an illegal streaming site or downloading from sci-hub, for instance. What people do get sued for is seeding, which counts as distribution. If anything AI companies are getting prosecuted more aggressively than "ordinary people", presumably because of their scale. In a recent lawsuit Anthropic won on the part about AI training on books, but lost on the part where they used pirated books.
- codedokode 1y agoPeople got in trouble for filming in the cinema as I understand, there is a separate law for that.
- gruez 1y agoBut in that case even though filming isn't technically distribution, it's clearly a step to distributing copies? To take this to the extreme, suppose you ripped a blu-ray, made a thousand copies, but haven't packaged or sold them yet. If the FBI busted in, you'd probably be prosecuted for "conspiracy to commit copyright infringement" at the very least.
- snickerdoodle12 1y agoIt's just "training"
- gruez 1y agoYou seem to equate "training" (with scare quotes) with someone actually pirating a blu-ray, but they really aren't equivalent. Courts so far have ruled that training is fair use and it's not hard to see why. Unlike copying a movie almost verbatim (as with ripping a blu-ray), AI companies are actually producing something transformative in the form of AI models. You don't have to like AI models, or the AI companies' business models, but it strains credulity to pretend ripping a blu-ray is somehow equivalent to training an AI model.
- 0x457 1y agoWell, it just shows that they've downloaded subtitles.
- robswc 1y agoAFAIK, downloading or watching pirated stuff isn't something you'll get in trouble for. Hosting and distributing it is what will get you.
- cyp0633 1y agoThat is not the case here - I never encountered this with whisper-large-v3 or similar ASR models. Part of the reason, I guess, is that those subs are burnt into the movie, which makes them hard to extract. Standalone subs need the corresponding video resource to match the audio and text. So nothing is better than YouTube videos which are already aligned.
- simsla 1y agoAt least for English, those "fansubs" aren't typically burnt into the movie*, but ride along in the video container (MP4/MKV) as subtitle streams. They can typically be extracted as SRT files (plain text with sentence level timestamps). *Although it used to be more common for AVI files in the olden days.
- ethbr1 1y agoflashbacks of trying to track down subs sync’d to a specific release
- Maken 1y agoSRT is ancient. Nowadays everyone uses ASS subtitles which can be randomly styled.
- simsla 1y agoIn general? In the past I've known ASS to be used a lot for things like anime, but less for live action shows.
- Maken 1y agoI have also found them inside mkvs as the subtitle track. I think SRT was the default because most content was ripped from DVD/BD, but now most of the content is from streaming sources and you need to convert the subtitles anyway.
- conradev 1y ago
- kgeist 1y agoInteresting, in Russian, it often ends with "Subtitles by %some_username%"