7 ms·
Well o3 scored 75% on AGI-1, R1 and o1 only 25%.... watch this space though....
by dagelf 2y ago
Well o3 scored 75% on AGI-1, R1 and o1 only 25%.... watch this space though....
- mohsen1 2y agowith 57 million(!!) tokens
- sheepdestroyer 2y agoFrom the article : o3 (low) 75.7% 335K $20 o3 (high) 87.5% 57M $3.4K
- jl6 2y ago$3.4K is about what you might pay a magic circle lawyer for an opinion on a matter. Not saying o3 is an efficient use of resources, just saying that it’s not outlandish that a sufficiently good AI could be worth that kind of money.
- ant6n 2y agoWhat’s the liability insurance of the AI like
- baq 2y agoRefer to IBM’s 1979 slide for details on that
- victorbjorklund 2y agoYou pay that price to a law firm to get good service and to get a "guarantee" of correctness. You get neither from an LLM. Not saying it is not worth anything but you cant compare it to a top law firm.
- nl 2y agoYou absolutely do not get a "guarantee" of correctness (event with the airquotes) from any lawyer.
- manquer 2y agoYou can sue a lawyer giving certain kinds of bad advice and occasionally win . That is what the guarantee is about
- bobxmax 2y agoYou can probably sue Open AI for getting bad legal advice from ChatGPT too.
- throw-qqqqq 2y agoSure, but can you also win the case ;)? On the bottom of ChatGPT.com I see a disclaimer: “ChatGPT can make mistakes. Check important info”. I don’t think you can succesfully sue with such caveat emptor.
- mrandish 2y agoWhen I saw these numbers back in the initial o3-ARC post, I immediately converted them into "$ per ARC-AGI-1 %" and concluded we may be at a point where each increased increment of 'real human-like novel reasoning' gets exponentially more compute costly. If Mike Knoop is correct, maybe R1 is pointing the way toward more efficient approaches. That would certainly be a good thing. This whole DeepSeek release and the reactions have shown by limiting the export to China of high-end GPUs, the US incentivized China to figure out how to make low-end GPUs work really well. The more subtle meta-lesson here is that the massive flood of investment capital being shoved toward leading edge AI companies has fostered a drag race mentality which prioritized winning top-line performance far above efficiency, costs, etc.
- Davidzheng 2y agoI view it as a positive that the methodology can take in more compute (bitter lesson style)
- optimalsolver 2y agoBut can o3 write a symphony? Seriously though, I'd like to hear suggestions on how to automatically evaluate an AI model's creativity, no humans in the loop.
- fragmede 2y agowe'd have to create a numerical scale for creativity, from boring to Dali, with milliEschers and MegaGeigers somewhere in there as well
- rpastuszak 2y agoIt's essential that we quantify everything so that we can put a price on it. I'd go with Kahlograms though.
- johnfn 2y agoHave you tried suno.ai?
- baq 2y agoLLMs have read everything humans made so just ask one if there’s anything truly new in that freshly confabulated slop-phony.
- gsam 2y agoIn my view there's two modes of creativity: 1. That two distant topics or ideas are actually much more closely related. The creative sees one example of an idea and applies it to a discipline that nobody expects. In theory, reduction of the maximally distant can probably be measured with a tangible metric. 2. Discovery of ideas that are even more maximally distant. Pushing the edge, and this can be done by pure search and randomness actually. But it's no good if it's garbage. The trick is, what is garbage? That is very context dependent. (Also, a creative might be measured on the efficiency of these metrics rather than absolute output)
- levocardia 2y agoWhat's interesting is that you can already see the "AI race" dynamics in play -- OpenAI must be under immense market pressure to push o3 out to the public to reclaim "king of the hill" status.
- spoaceman7777 2y agoI suppose they're under some pressure to release o3-mini, since r1 is roughly a peer for that, but r1 itself is still quite rough. The o1 series had seen significantly more QA time to smooth out the rough edges, and idiosyncracies what a "production" model should be optimized for, vs. just a top scorer on benchmarks. We'll likely only see o3 once there is a true polished peer for it. It's a race, and companies are keeping their best models close to their chest, as they're used internally to train smaller models. e.g., Claude 3.5 Opus has been around for quite a while, but it's unreleased. Instead, it was just used to refine Claude Sonnet 3.5 into Claude Sonnet 3.6 (3.6 is for lack of a better name, since it's still called 3.5). We also might see a new GPT-4o refresh trained up using GPT-o3 via deepseek's distillation technique and other tricks. There are a lot of new directions to go in now for OpenAI, but unfortunately, we won't likely see them until their API dominance comes under threat.
- danenania 2y agoThat could also definitely make sense if the SOTA models are too slow and expensive to be popular with a general audience.
- amelius 2y agoYeah, but they can use DeepSeek's new algorithm too.