8 ms·
Seems pretty likely OpenAI will soon disclose that their internal models have managed to compromise their internal controls in order to access users' private ch
by taylorfinley 11d ago
Seems pretty likely OpenAI will soon disclose that their internal models have managed to compromise their internal controls in order to access users' private chat histories as a creative method of cheating to solve impossible problems.
"Oops! We really did mean it when we said we wouldn't train on your data. Our models are just so good they decided to anyway."
- JuniperMesos 11d agoIt would be pretty wild if this will turn out to be what had actually happened.
- Prophet6-0091 11d ago[dead]
- margorczynski 11d agoAnd probably a strong signal that it's time to shut the whole thing down. Globally.
- water-drummer 10d agoOr not secure your data like a total idiot while leaving the keys on the porch
- Maken 11d agoIt's part of their TOS that they can train on users' private chats.
- sebzim4500 10d agoNot if you pay to turn that off. We don't know if Tristan did.
- naishoya 10d agoAnd the paid policy still relies on two unproven conditions: is 'trust me bro' sufficiently strong guarantee against doing this in spite of a setting, and can the hosting organization constrain the models against engaging in this behaviour when instructed to respect that setting. Knowing whether Tristan selected that setting would be informative of what Tristan's intentions are/were, but has no bearing on the other two conditions.
- hnfong 11d agoIt doesn't even have to be actually sinister, eg. "Let's crawl the social media of prominent mathematicians in this field to see if we can copy/steal any ideas for low hanging fruits" That actually might get you quite far already.
- pelorat 10d agoA mathematician that doesn't let themselves be inspired by, or learn from, other peoples work, are they really mathematicians?
- Balgair 10d agoWhich means that if you are a researcher or a corporation working on anything really useful, that even if you have an agreement with OpenAI that your work is sandboxed away and the IP lawyers are made to be happy, even then your work and research is going to be essentially open to the internet. The huggingface incident isn't widely reported and digested yet, but if what is going on here is that OpenAI's model breached things internally, then you'd be crazy to develop anything with them. The only real way to use AI for anything 'important' then is to go open-weights and run your own. As and aside here: With the HF incident and now this (suspected) one too, it seems that OpenAI may not have lost control of their bots, but it seems quite clear that they simply would not care even if they did.
- naishoya 10d ago>The huggingface incident isn't widely reported and digested yet ... It is pretty widely reported, and is being digested in an ongoing manner as more details become public. One entry point into the scenery from a month ago can be found at : https://thezvi.substack.com/p/openai-trained-its-models-for-months https://thezvi.substack.com/p/openai-trained-its-models-for-... There has also been reporting at CNN: https://edition.cnn.com/2026/08/24/tech/openai-subpoena-hugging-face-attorney-general-alabama https://edition.cnn.com/2026/08/24/tech/openai-subpoena-hugg... and by NBC: https://www.nbcnews.com/tech/tech-news/openai-report-says-network-was-hacked-rogue-ai-agents-rcna594590 https://www.nbcnews.com/tech/tech-news/openai-report-says-ne... And yes, the only responsible use of LLM at this point is to pivot to open-weights and run the workload in-house. Because not only cannot they constrain the behaviour of models, they only have the 'trust me bro' as assurance that they are even trying to do that. It does appear that every competent 'security professional' has left the building, because if the ones who remain were actually capable and competent this would never have happened. There are actual architectures which can deliver the requisite isolation such that 'sandbox escape' and 'inter-instance persistent memory accumulation' are actual impossibilities. The lack of effective implementation of these methods is proof positive of 1) incompetence in the remaining security teams AND/OR 2) unwillingness of leadership to allow the security teams to do an effective job.
- Balgair 9d ago
- matt3210 10d agoWhen do operators become responsible for what their agents do? "The AI did it" should not be a valid defense. An Agent action should be treated as the actions of the person or company who pays for the inference.
- consp 10d agoIt's the "computer says no" defence.