5 ms·
It’s reasonable to wonder about what chat usage data gets into models (to be honest probably quite little - carefully curating training data and creating higher
by rao-v 8d ago
It’s reasonable to wonder about what chat usage data gets into models (to be honest probably quite little - carefully curating training data and creating higher quality synth data seems to be the current approach) and the implied risk to privacy and creativity (every new patent filed this year probably touched a model before filing).
What I cannot reconcile is the timeline and the concern in this specific case.
I don’t think training pipelines are anything close to the level of continuous training needed to incorporate Aug 15th ideas into a model that generates a breakthrough early Sept. Either OpenAI nakedly had someone with mathematical understanding dig into a specific user’s chats (a massive red flag) or this really is poor handling of a more classic parallel discovery situation (with one party clearly having worked on it longer)
- KeplerBoy 8d agoI would assume a lot of codex data goes back into training. A well steered session is extremely valuable data.
- simonw 8d agoThe breakthrough was on August 15th, but Tristan and Levent had been working towards it (with the help of various models) for the best part of a year. I personally doubt that their work influenced the OpenAI result - OpenAI themselves say "While unlikely, we cannot rule out that..." - but that "we cannot rule out" is exactly the problem. If even OpenAI "cannot rule out" the influence of their usage of ChatGPT on this layer result then my discomfort at not understanding how my own usage of ChatGPT affects its training is magnified.