8 ms·
With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' po
by torginus 3d ago
With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.
Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it.
This could mean every potential serious customer would have no option but to seek alternatives to these online services.
- lynndotpy 3d agoI thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs? I don't mean this as rhetoric, I did not think many people (except possibly those operating under government contracts, and 'normies' who don't know about these things) were under the belief that their IP was kept secret when they use these services.
- Gud 3d agoNo, that is not "common knowledge". You are supposed to be able to disable that unwanted feature.
- zdragnar 3d agoSome offer zero data retention policies, but there can be weasel words. For example, on the individual pro plan, you can turn off the setting that lets them train models on your data, but they still have a section in their terms that allows them to evaluate your anonymized data for statistical and "research" purposes. You have to actually get a signed contract along with an enterprise plan that spells out exactly what they're going to use, and what settings enable what retention. https://privacy.claude.com/en/articles/10023548-how-long-do-you-store-my-data https://privacy.claude.com/en/articles/10023548-how-long-do-... (see the additional info section)
- ssivark 3d agoWhat about inference providers like Baseten, Modal, Fireworks, Together, etc? I thought one of their value propositions was inference (using open weights models) that guarantees with crisp terms that they will not use your data.
- hazard 3d agoI worked very briefly at Baseten, and I can say that it was a perpetual annoyance (from an engineering perspective) that customers would complain about issues with their models but we couldn't actually see the inputs/outputs. I don't know about the other providers, but at Baseten they literally weren't stored anywhere.
- Aurornis 3d ago> I thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs? The services have toggles to allow prompts to be used in the training set. There is a conspiracy theory that the toggle is a false distraction and they’re actually keeping everything, and that none of the employees involved will ever whistleblow this fact. Outside of Internet comment sections, I think most people assume these US-based companies are doing what they say. For enterprise use there are services like AWS Bedrock which have strict isolation guarantees. There are some people who still believe those guarantees are a lie, but once someone has reached that point I don’t think they trust anything that isn’t running entirely within their house. People in that category are a very small minority, but a very vocal minority.
- Aurornis 3d ago> OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this. I think this is being misunderstood. Codex has a toggle to allow your prompts to be included in training data. They’re saying they can’t be sure if the person had it on or off while using Codex to discuss the work. They’re not saying that some prompts are mysteriously jumping into training data. Also, there is a large market for AI services which don’t retain anything under any circumstances for enterprise customers.
- ronsor 3d agoAlmost every serious customer is already using ZDR where nothing is retained at all, instead of "anonymized" data.
- torginus 3d agoJust a thought experiment: considering training seems to be 'fair use', I wonder if they trained a tiny model to retain key info from your prompts, would mean that this would still constitute fair use, and allow them to legally claim they don't retain your data.
- ronsor 3d agoZDR is shorthand for a more specified agreement of "we don't do anything other than generate your output tokens", so no. Besides, true ZDR is usually offered by third-parties with deals to host OpenAI models, such as Amazon (AWS Bedrock) and Microsoft (Azure).
- nrmitchi 3d agoThe guarantee on this is a (contractual) “trust me bro”, and a right to try to sue a multi-trillion-dollar company who will absolutely drive you into the ground with legal red tape. If you are big enough to be able to withstand that, you’re already running (or trying to run) your own/open-weight models.
- steveBK123 3d agoThey already trained on pirated content, what makes you think they are going to honor ZDR?
- Betelbuddy 3d agoOr these customers could just use AWS Bedrock...but their current CEO is an incompetent MBA unable to publicly articulate their biggest advantage, in the context of the current AI usage my companies. You have access to all the frontier models, but...your inputs are not shared with the model vendors...neither are used to train the next model. Why am I even doing the Amazon board job for them!??
- ballon_monkey 3d agoBedrock is really bad. It seems like they don't host the models very well because they produce tons of bugs/errors calling the model. For example you can end up with Anthropic models not returning a stop token and you end up waiting for a timeout thinking its doing something when it isn't.
- Betelbuddy 3d agoWell Anthropic hosts their models at AWS, ( and at many others...) so maybe the AWS team can ask them how they do it ;-) ?
- staticautomatic 3d agoAll except Gemini which can be rather important depending on your use case.
- Betelbuddy 3d agoYou mean the Gemini that is even behind the Chinese models?
- whatshisface 3d agoAmazon is deeply invested in Anthropic and would not defame them through marketing a service whose selling point was their startup's breach of contracts.