8 ms·
What we can reasonably assume from statements made by insiders: They want a 10x improvement from scaling and a 10x improvement from data and algorithmic change
by ericskiff 2y ago
What we can reasonably assume from statements made by insiders:
They want a 10x improvement from scaling and a 10x improvement from data and algorithmic changes
The sources of public data are essentially tapped
Algorithmic changes will be an unknown to us until they release, but from published research this remains a steady source of improvement
Scaling seems to stall if data is limited
So with all of that taken together, the logical step is to figure out how to turn compute into better data to train on. Enter strawberry / o1, and now o3
They can throw money, time, and compute at thinking about and then generating better training data. If the belief is that N billion new tokens of high quality training data will unlock the leap in capabilities they’re looking for, then it makes sense to delay the training until that dataset is ready
With o3 now public knowledge, imagine how long it’s been churning out new thinking at expert level across every field. OpenAI’s next moat may be the best synthetic training set ever.
At this point I would guess we get 4.5 with a subset of this - some scale improvement, the algorithmic pickups since 4 was trained, and a cleaned and improved core data set but without risking leakage of the superior dataset
When 5 launches, we get to see what a fully scaled version looks like with training data that outstrips average humans in almost every problem space
Then the next o-model gets to start with that as a base and reason? Its likely to be remarkable
- jsheard 2y ago> With o3 now public knowledge, imagine how long it’s been churning out new thinking at expert level across every field. OpenAI’s next moat may be the best synthetic training set ever. Even taking OpenAI and the benchmark authors at their word they said that it is consuming at least tens of dollars per task to hit peak performance, how much would it cost to have it produce a meaningfully large training set?
- qup 2y agoThat's the public API price isn't it?
- jsheard 2y agoThere is no public API for o3 yet, those are the numbers they revealed in the ARC-AGI announcement. Even if they were public API prices we can't assume they're making a profit on those for as long as they're billions in the red overall every year, its entirely possible that the public API prices are less than what OpenAI is actually paying.
- Stevvo 2y ago"With o3 now public knowledge, imagine how long it’s been churning out new thinking at expert level across every field." I highly doubt that. o3 is many orders of magnitude more expensive than paying subject matter experts to create new data. It just doesn't make sense to pay six figures in compute to get o3 to make data a human could make for a few hundred dollars.
- dartos 2y agoThat’s an interesting idea. What if OpenAI funded medical research initiatives in exchange for exclusive training rights on the research.
- onlyrealcuzzo 2y agoIt would be orders of magnitude cheaper to outsource to humans.
- dartos 2y agoNot as sexy to investors though
- aswegs8 2y agoWait didn't they just recently request researchers to pair up with them in exchange for the data?
- DougN7 2y agoSomeone needs to dress up Mechanical Turk and repackage it as an AI company…..
- jitl 2y agoThat’s basically every AI company that existed before GPT3
- bookaway 2y agoYes, I think they had to push this reveal forward because their investors were getting antsy with the lack of visible progress to justify continuing rising valuations. There is no other reason a confident company making continuous rapid progress would feel the need to reveal a product that 99% of companies worldwide couldn't use at the time of the reveal. That being said, if OpenAI is burning cash at lightspeed and doesn't have to publicly reveal the revenue they receive from certain government entities, it wouldn't come as a surprise if they let the government play with it early on in exchange for some much needed cash to set on fire. EDIT: The fact that multiple sites seem to be publishing GPT-5 stories similar to this one leads one to conclude that the o3 benchmark story was meant to counter the negativity from this and other similar articles that are just coming out.
- dartos 2y agoI’m curious how, if at all, the plan to get around compounding bias in synthetic data generated by models trained in synthetic data.
- nialv7 2y agosynthetic data is fine if you can ground the model somehow. that's why the o1/o3's improvements are mostly in reasoning, maths, etc., because you can easily tell if the data is wrong or not.
- dartos 2y agoThat makes a lot of sense. Binary success criteria has very little room for bias.
- ynniv 2y agoEveryone's obsessed with new training tokens... It doesn't need to be more knowledgeable, it just needs to practice more. Ask any student: practice is synthetic data.
- dartos 2y agoThat leads to overfitting in ML land, which hurts overall performance. We know that unique data improves performance. These LLM systems are not students… Also, which students graduate and are immediately experts in their fields? Almost none. It takes years of practice in unique, often one-off, situations after graduation for most people to develop the intuition needed for a given field.
- ynniv 2y agoIt's overfitting when you train too large a model on too many details. Rote memorization isn't rewarding. The more concepts the model manages to grok, the more nonlinear its capabilities will be: we don't have a data problem, we have an educational one. Claude 3.5 was safety trained by Claude 3.0, and it's more coherent for it. https://www.anthropic.com/news/claudes-constitution https://www.anthropic.com/news/claudes-constitution
- noman-land 2y agoI completely don't understand the use for synthetic data. What good it's it to train a model basically on itself?
- psb217 2y agoThe value of synthetic data relies on having non-zero signal about which generated data is "better" or "worse". In a sense, this what reinforcement learning is about. Ie, generate some data, have that data scored by some evaluator, and then feed the data back into the model with higher weight on the better stuff and lower weight on the worse stuff. The basic loop is: (i) generate synthetic data, (ii) rate synthetic data, (iii) update model to put more probability on better data and less probability on worse data, then go back to (i).
- noman-land 2y agoThanks, that makes a lot more sense.
- RedNifre 2y agoBut who rates the synthetic data? If it is humans, I can understand that this is another way to get human knowledge into it, but if it's rated by AI, isn't it just a convoluted way of copying the rating AI's knowledge?
- ijustlovemath 2y agoThis is the bit I've never understood about training AI on its own output; won't you just regress to the mean?
- astrange 2y agoIt's not trained on its own output. You can generate infinite correctly worked out math traces and train on those.
- recursivecaveat 2y ago
- nialv7 2y ago> OpenAI’s next moat I don't think oai has any moat at all. If you look around, QwQ from Alibaba is already pushing o1-preview performances. I think oai is only ahead by 3~6 months at most.
- vasco 2y agoIf their AGI dreams would come true it might be more than enough to have 3 months head start. They probably won't, but it's interesting to ponder what the next few hours, days, weeks would be for someone that would wield AGI. Like let's say you have a few datacenters of compute at your disposal and the ability to instantiate millions of AGI agents - what do you have them do? I wonder if the USA already has a secret program for this under national defense. But it is interesting that once you do control an actual AGI you'd want to speed-run a bunch of things. In opposition to that, how do you detect an adversary already has / is using it and what to do in that case.
- kevingadd 2y agoHow many important problems are there where a 3 month head start on the data side is enough to win permanently and retain your advantage in the long run? I'm struggling to think of a scenario where "I have AGI in January and everyone else has it in April" is life-changing. It's a win, for sure, and it's an advantage, but success in business requires sustainable growth and manageable costs. If (random example) the bargain OpenAI strikes is "we spend every cent of our available capital to get AGI 3 months before the other guys do" they've now tapped all the resources they would need to leverage AGI and turn it into profitable, scalable businesses, while the other guys can take it slow and arrive with full pockets. I don't think their leadership is stupid enough to burn all their resources chasing AGI but it does seem like operating and training costs are an ongoing problem for them. History is littered with first-movers who came up with something first and then failed to execute on it, only for someone else to follow up and actually turn the idea into a success. I don't see any reason to assume that the "first AGI" is going to be the only successful AGI on the market, or even a success at all. Even if you've developed an AGI that can change the world you need to keep it running so it can do that. Consider it this way: Sam Altman & his ilk have been talking up how dangerous OpenAI's technology is. Are risk-averse businessmen and politicians going to be lining up to put their livelihood or even their lives in the hands of "dangerous technology"? Or are they going to wait 3-6 months and adopt the "safe" AGI from somebody else instead?
- nradov 2y agoThere is an enormous "iceberg" of untapped non-public data locked behind paywalls or licensing agreements. The next frontier will be spending money and human effort to get access to that data, then transform it into something useful for training.
- mistercheph 2y agoah yes the beautiful iceberg of internal documentation, legal paperwork, and meeting notes. the highest quality language data that exists is in the public domain
- sdwr 2y agoGreat improvements and all, but they are still no closer (as of 4o regular) to having a system that can be responsible for work. In math problems, it forgets which variable represents what, in coding questions it invents library fns. I was watching a YouTube interview with a "trading floor insider". They said they were really being paid for holding risk. The bank has a position in a market, and it's their ass on the line if it tanks. ChatGPT (as far as I can tell) is no closer to being accountable or responsible for anything it produces. If they don't solve that (and the problem is probably inherent to the architecture), they are, in some sense, polishing a turd.
- tucnak 2y ago> ChatGPT (as far as I can tell) is no closer to being accountable or responsible for anything it produces. What does it even mean? How do you imagine that? You want OpenAI to take on liability for the kicks of it?
- numpad0 2y agoIf an LLM can't be left to do mowing by itself, but a human will have to closely monitor and intervene at every its steps, then it's just a super fast predictive keyboard, no?
- dyauspitr 2y agoBut what if the human only has to intervene once every 100 hours, that’s a huge productivity boost.
- cjblomqvist 2y agoThe point is you don't know when of those 100 hours that is, so you still need to monitor the full 100 hour time span. Can still be a boost. But definitely not the same magnitude.
- kjkjadksj 2y agoAnd one might also wonder still if we need a general language model to mow the grass or just a simpler solution towards to problem of driving a mower over a fixed property line automatically. Something you could probably solve with wwii era technology, honestly.