6 ms·
I heard a rumor (on instagram, so YMMV) that the professor who was closest to solving this problem had only weeks ago used Codex, which had slurped up all his n
by efnx 6d ago
I heard a rumor (on instagram, so YMMV) that the professor who was closest to solving this problem had only weeks ago used Codex, which had slurped up all his notes on the subject. Now OpenAI's agents solve the problem. If it's true that seems like quite a coincidence.
- bethekidyouwant 6d agoHow could they possibly included in the previous training run which takes months to complete..
- mswphd 6d agoI won't take a side in things, but OpenAI stated the model they used here started training August 28th. Note that "training" here might mean "post-training with RLHF an Astra base model" or something. but training had only started a little over a week earlier.
- s900mhz 6d agoIMO It’s not about being trained on the data, it’s more like what do the agents have access to during inference? Can they grep customer transcripts/logs?
- bethekidyouwant 6d agoYou’re saying that when they we’re trying to solve this theorem they also shoved in its context somebody else’s chat logs? Bruh.
- metanonsense 6d agoMaybe the boundaries of the memory subsystem are a bit fuzzy.
- TZubiri 6d agoNo, it's quite well defined, and it's not called a 'memory subsystem'. There is a training process (as in traditional Machine Learning training) that occurs with data available at the specific point in time the training starts (or ends), this is called the cutoff date. After this point there can be other kinds of training, the weights can be shifted, the internal CoT prompts can be changed, routing in MoE can change, but the Foundational Model that was trained on a corpus is the same model trained in the same corpus. User data can be used at any of these stages theoretically of course, but by the nature of training and from the dates of the events, (a new Foundational Model being released), it would look as if the user data of the professor was used in the training of the foundational model, which is something that OAI does every couple of months for a big release, and incorporates the new text from their text scraping efforts, including new books ingested, new internet text scraped, deals with third party platforms, and data from their own users (not conjectured, read the ToS, users allow this.)
- jazzyjackson 6d agoHas it not been the usual process to snapshot a model to use for inference while continuing to run the training process? I guess you can’t add to the training corpus once you begin? Just trying to make sense of whether training begins or ends as rigidly as you suggest.
- 1121redblackgo 6d agoSee other thread, but yeah that's the general ballpark of the situation.
- scuppernong 6d agothis is not a rumor (the allegation, anyway), it's reported in the new york times
- jcranmer 6d agoNewspapers are not above printing rumors. See, e.g., Barak Ravid regularly reporting in Axios the impending ceasefire negotiation progress in the Iran War, which largely have failed to come to pass.
- efnx 6d agoI don't understand why I'm getting downvoted, I'm not posting an opinion. Coincidences happen. So does foul play. No judgement call here.
- dooglius 6d agoThere have been several threads and developments on this over the past few days, including statements from the primary subjects involved. Third-hand instagram comments are not really the best source to be bringing in.
- efnx 6d agoEverybody comes into information in different ways. There were no comments here about this specific aspect of the story - which is definitely interesting!