5 ms·
> leaving aside the idea that OA might've used data from the researchers Codex sessions Why leave that aside? That is _the_ story. If a Chinese research lab d
by AJRF 9d ago
> leaving aside the idea that OA might've used data from the researchers Codex sessions
Why leave that aside? That is _the_ story.
If a Chinese research lab did this we'd call it espionage.
- modeless 9d agoBecause they didn't do that. Tristan doesn't specifically claim that they did, and Anthropic employees don't think they did either. https://x.com/_sholtodouglas/status/2097218240397410733 https://x.com/_sholtodouglas/status/2097218240397410733
- Semkas 9d agoBy default OA trains their models on codex-sessions. If I understand him correctly this is something Tristan explicitly mentions in his post as a possible reason for the fast results obtained by the internal OA team. Anthropic obviously doesn't want to challenge the idea that training is transformative, even if it means agreeing with their competitor.
- FiberBundle 9d agoWell, of course Anthropic employees would say that, since they likely do the same. Claiming that your primary competitor doesn't engage in a certain malicious practice is supposed to make it look as if there's no way you would too. If somebody even says that about their competitor, then surely there must be truth to that, otherwise you would never give credit to someone you're opposed to.
- hodgehog11 8d agoIt's literally a toggle in the options for ChatGPT, one which is on by default and most researchers probably have on without realising it. So to say that it is unlikely is extremely suspicious. No, they did not literally pull user data. But user data is automatically added to their training set by default, so their latest in-house model would be trained on it if it is from several months ago. It isn't intentional on their part, and they probably realised they could not refute that they trained on Tristan's logs unintentionally, hence why they acted the way they did.
- hellohello2 8d agoIts very easy for OpenAI to answer, yes or no, if the model they used trained on their chats.
- Semkas 9d agoBut there's a bunch of people already in this thread calling that stuff unfounded speculation (which I disagree with), and my point is that even if that specific thing isn't true, OA's behavior here is obviously awful. If they're going to try to beat researchers to discoveries like this it disincentives researchers to talk about their progress publicly, and basically breaks the ecosystem of scientific cooperation / discovery. It's also immoral.
- deleted 9d ago[deleted]
- reasonableklout 8d agoYep. The most uncharitable view of this might be: they stole the work of researchers to build their models, and now they're using said models to steal the proceeds of future work, too.
- paxys 9d agoWhy is that the story? Is there anything to back it up beyond a single accusation?
- defmacr0 8d agoAs a prior I would say that a math professor has about infinite times more integrity than OpenAI.
- light_hue_1 8d agoIf you think that OpenAI won't look at your data to gain a massive advantage, you're naive.