5 ms·
I agree. The frontier models are based on training data from tons of copyrighted work. Some of that work was obtained illegally, even. They could not exist with
by kelnos 3d ago
I agree. The frontier models are based on training data from tons of copyrighted work. Some of that work was obtained illegally, even. They could not exist without strip-mining the commons. The labs have no moral or ethical ownership to the end result, and others should feel free to treat any company-imposed restrictions on their use as invalid.
I don't expect Tan's position to be based on any kind of real moral high ground, but his conclusion is correct.
I love the "illicit distillation attacks" framing from the incumbents. There's nothing illicit. There's no attack. You just don't like it because it threatens your market position and business model.
- Aurornis 3d agoThere is nothing illegal about training on traces from frontier models. However the frontier labs don’t have to serve customers who are farming the service for distillation purposes. That’s their choice and they’re free to make it if they detect distillation happening.
- darth_avocado 3d agoI would argue they should have to. They scraped data off others, a lot of whom did not want that data to be used for AI training, and still had to share it with the frontier labs. It’s only fair they should have to hand it back. The only way US maintains dominance over Chinese models is by having an ecosystem of models. Relying on a small set of frontier labs will only let you get ahead temporarily. I agree with Gary Tan on this one.
- hlynurd 3d agoThat's fine, they just gotta tone down the victim rhetoric.
- ronsor 3d agoYes, I think this is the main issue. I don't care what policies the AI labs have or enforce, but they need to stop acting like ToS violations are an international crisis demanding intervention instead of a boring civil dispute at most.
- godwinson__4-8 3d agoIf the leading private labs attempt to use the government to pull up the ladder under the pretense of "safety" then the response of the people should be to take such questions out of private hands and nationalize the leading labs. Or they could abide by the precedents they set and learn to compete. They shouldn't be allowed to have it both ways.
- Barbing 3d agoAll correct, just help me get over the idea of an open-weight Mythos where one or a dozen of us eight billion does something stupid on the bioweapon front. Smart people who’ve exhausted possibilities for what they can do with books and web search and today’s Kimi/GLM. Figure we’ll have to reckon with this next year in any case, guess we’ll see.
- ronsor 3d ago"Bioweapon" information is not useful without a lab for synthesis. Someone with that lab could almost certainly figure out how do something stupid or destructive on their own, or bypass model safeguards somehow.
- a34729t 3d agoYou dont need an LLM to figure out to make anthrax. Anybody who can figure out how to make a home lab can make all sorts of dangerous stuff pretty easily. Same with college grad from a respectable chemistry program. This all FUD.
- deleted 3d ago[deleted]
- edot 3d agoThis is our generation's "Saddam has WMDs". It's something the big labs thought up when they were trying to figure out how to make their product sound scary enough to deserve regulation. Literally no one is doing this or even trying, anyone who would want to do it would have already done it. Not worried about it.
- stymaar 3d agoThis. Distillation “attacks” are a made up concept. It's as if I claimed that Anthropic made a “training attack” when training on my internet writing.
- giancarlostoro 3d agoAbolish copyright and make it less ridiculous. Sampling music was never a thing that required royalties until the 1990s when I guess someone got angry that rappers were making money off their sampled music. Its insane to me. Make it illegal to transfer ownership of copyrighted work too, only the spouse or one single inheritor who isnt a company can have the rights transferred, after both die, the work enters public domain. LLMs should just pay a flat fee to use a specific book and thats it. Fees should be reasonable (not a million dollars per book), so long as the model doesnt spit out the entire book.
- derefr 3d ago> Make it illegal to transfer ownership of copyrighted work too, only the spouse or one single inheritor who isnt a company can have the rights transferred, after both die, the work enters public domain. By your phrasing, it sounds like you still intend the possibility of companies owning copyrights; but how does that happen (other than copyrights already owned by companies grandfathered in)? Copyright always starts off in the hands of individual human beings; it only ends up in the hands of companies when those human beings transfer ownership to a company. That ownership transfer can be automatic as a term of a contract, e.g. as part of a work-for-hire agreement. But no contract can cause the copyright to come into existence already held by the company instead of the individual. So if you abolish ownership transfer, you effectively make work-for-hire IP assignment invalid. What replaces it? And, if "nothing"... then how do people pool the IP rights of their own small contributions to a large-scale work, into an IP pool that can be legally defended by a coherent legal entity, so that the large-scale work itself can have market value (i.e. so that sales of polished commercial bootlegs don't drive sales of the "authentic" work to zero)? Keep in mind that, no matter how much we might want "mass distributed" media to have more-reasonable IP terms, the ability to sue for infringement is still critical to the existence of some forms of media. Especially "location-based" media, with no equivalent licensed broadcast right: movies still in theatre; concerts; live performances of plays and musicals; etc. If there's no legal team that can sue a movie theatre that shows an unlicensed copy of a given movie, then no movie theatre will ever bother with licensing movies again; "box office" goes to zero (from the movie company's perspective); and the incentive to create movies in the first place declines massively. I'm not saying this is an impossible problem. There are ways to accomplish this besides the way it's done now. (For example, individual-contributor IP could be retained by the original owners, but cross-licensed between individuals through a collaboration structure to form a coherent defensible IP pool, in exactly the same way that IP for e.g. video codecs is cross-licensed between corporations to form a coherent defensible IP pool today.) I'm just pointing out that the problem does need to be solved.
- torginus 3d agoWith the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this. Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it. This could mean every potential serious customer would have no option but to seek alternatives to these online services.
- lynndotpy 3d agoI thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs? I don't mean this as rhetoric, I did not think many people (except possibly those operating under government contracts, and 'normies' who don't know about these things) were under the belief that their IP was kept secret when they use these services.
- Gud 3d agoNo, that is not "common knowledge". You are supposed to be able to disable that unwanted feature.
- zdragnar 3d agoSome offer zero data retention policies, but there can be weasel words. For example, on the individual pro plan, you can turn off the setting that lets them train models on your data, but they still have a section in their terms that allows them to evaluate your anonymized data for statistical and "research" purposes. You have to actually get a signed contract along with an enterprise plan that spells out exactly what they're going to use, and what settings enable what retention. https://privacy.claude.com/en/articles/10023548-how-long-do-you-store-my-data https://privacy.claude.com/en/articles/10023548-how-long-do-... (see the additional info section)
- ssivark 3d agoWhat about inference providers like Baseten, Modal, Fireworks, Together, etc? I thought one of their value propositions was inference (using open weights models) that guarantees with crisp terms that they will not use your data.
- knollimar 3d agoI'm sure they put some BS in their TOS
- mobelkh 3d agowhy can't I use the tokens i paid for anyway?
- impossiblefork 3d agoMorally I agree, but since there's probably a lot of LLM text in the training data, distilling on another model will probably make your model copy the values encoded into the other model as well, even in cases where you only distill on value-neutral stuff. By copying their programming style, you'll move the model towards that way of writing, which will move the model towards the values expressed in those documents. I feel that Deepseek v4 got so claudified at the end that it was like Claude.