7 ms·
"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally gi
by makaking 2mo ago
"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part."
How surreal is it that we are not absolutely jaw-dropped by these types of capability improvements? It's been less than 4 years since ChatGpt came out and now they are spontaneously building their own ML pipelines to do real-world 3D modeling tasks reliably.
Escaped its sandbox and hacked into Hugging Face's database? it's just another Monday...
That jump to 30% in ARC-AGI 3? Normal...
We should find a way to get "re-sensitized" to what we are witnessing and the pace of it.
- portly 2mo agoI'm tired boss
- thedevilslawyer 2mo agoWhy, I wonder. It's insanely amazing.
- gabrieledarrigo 2mo ago> Why, I wonder. By something that is going to replace my job by writing better code and shipping more features at a fraction of the cost, without asking for vacation? I don't know, boss; I'm also trying to understand why I'm tired.
- pattle 2mo agoit's a net positive for the human race
- bbeonx 2mo agois it?
- ed_balls 2mo agowe'll be better off - assuming we'll generate revenue on our own. LLMs have opened up an enormous new frontier of micro-SaaS opportunities. One person can now build what used to require a small team. The real problem is zero-sum work, especially when the only moat was knowledge asymmetry that is now public knowledge.
- varjag 2mo agoSoon it doesn't require even that one person
- chpatrick 2mo agoAnd the saas won't be required either because people can just get these problems solved directly.
- mhitza 2mo agoWhy would anyone use your micro SaaS when they can build their own? Have you seen a different perspective? Cause in my informational bubble SaaS ain't doing too well since vibecoding became more common.
- muspimerol 2mo agoThere's still a cost/benefit to building and maintaining a service. LLMs make it easier to build, but the cost does not scale to zero.
- mhitza 2mo agoI'm in agreement with your claim, though if you recall some of these micro SaaSes went hyperfocused on niches. The value proposition is then smaller, with larger data liability than with some behemoth's of SaaS. Though for the later ones, one might also ask themselves if they are worth their cost when using 20-30% of provided functionality.
- deleted 2mo ago[deleted]
- BlackFingolfin 2mo ago> at a fraction of the cost We'll see about that, long term. The billions and trillions being wasted right now to get the foot in the door need to be earned back somehow at some point... "Nice company you have there, totally reliant on our AI tech. Oh by the way we gotta increase the rent again."
- AQuantized 2mo agoDoesn't really make sense as there are already open weight models that are close to frontier models, and the compute to run them isn't outrageously expensive, and every indication that it will only get cheaper over time as has been the strong trend in terms of $/outcome.
- toasty228 2mo agoLike the noise thunder makes while loading before striking you, insanely amazing, potentially problematic
- thedevilslawyer 2mo agoIt's about perspective. Nuclear reactions are way more powerful, but we've harness them to power our world.
- duskdozer 2mo agoAnd to make bombs.
- cyclopeanutopia 2mo agoThey are not intelligent though.
- bauerd 2mo agoBecause these productivity increases will be to the detriment of the average person. We won’t get to reap the benefits.
- DSingularity 2mo agoIt’s unlikely to stop here. Today is a proprietary 5T parameter model, tomorrow it’s 5 100B parameter models that each specialize to specific applications and you can run them on your phone. The direction of travel is not like you think. The real victim will be medium to large software companies that can no longer rely on the difficulty of reproducing or maintaining or hosting their software as moat.
- hahamaster 2mo agoYou probably will. It's similar to steam engine and mechanical power loom.
- Marha01 2mo ago> We won’t get to reap the benefits. There is no rational reason to think that. A rising tide lifts all boats.
- throw-the-towel 2mo agoNot a lot of boat lifting happened in the last few decades. (See "WTF happened in 1971".)
- yoz-y 2mo agoWorking with AI has been way more tiring than just working. Sure the productivity is up, at the cost of having to keep up many thought threads, having no calm moments, and needing to consistently dig into large unknown code cases to find weird bugs. I’ve been on leave for a month, and am super excited (/s) to re-learn everything because all the tooling and ways to prompt “correctly” will have also changed.
- muldvarp 2mo agoBecause the ability to create software has little inherent value to me and is only valuable to me because it allows me to earn enough to make this life somewhat bearable. LLMs that can build a computer vision pipeline are a direct threat to my ability to sustain myself. At the same time, being able to prompt an LLM to build a computer vision pipeline doesn't really positively affect my life at all, because personally I don't care that much about computer vision pipelines (or frankly, any software).
- Marha01 2mo agoThe ability to create software accelerates technological progress, which has direct and indirect benefits for everyone. It is true that competition could have short-term negative effect on people who sustain themselves by creating software, though.
- muldvarp 2mo ago> The ability to create software accelerates technological progress, which has direct and indirect benefits for everyone. It has benefits for people with enough leverage (money and formerly labour) to obtain those benefits. > It is true that competition could have short-term negative effect on people who sustain themselves by creating software, though. I'm guessing that even if the unlikeliest of all unlikely things does happen and we all live off of some UBI some day, the short-term negative effects won't be "short-term" in the context of a human life.
- Marha01 2mo ago> It has benefits for people with enough leverage (money and formerly labour) to obtain those benefits. That's pretty much everyone. Even poor people benefit from technological progress. > I'm guessing that even if the unlikeliest of all unlikely things does happen and we all live off of some UBI some day, the short-term negative effects won't be "short-term" in the context of a human life. That depends on the pace of technological acceleration. It could be just a few years, or a decade. Which is why I am a pedal-to-the-metal accelerationist. The quicker we get through the short-term negative/turbulent phase towards the long-term positive phase, the better for me. If reversing is not possible, then going quicker is actually better than going slower. Let's get this shit over with.
- Dilettante_ 2mo agoThese are always cherry-picked, though. They tell you about the 1/10 that went really impressively, ignoring the other 9 shots at the task where the clanker started to try selling tungsten cubes (in person, wearing a blue shirt).
- jstanley 2mo agoThis is exactly the kind of take that the comment you replied to is talking about.
- makaking 2mo agoFair. But that 1/10 continues to get more and more impressive. Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times. Anthropic reports "Claude figured out how to tie general relativity with quantum mechanics." Would you hand-wave it away saying that it's cherry-picked?
- haldujai 2mo ago> Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times. I would not hand wave that away - but Claude 5 feels closer to Claude 1 than the hypothetical Claude 7 you propose. I do not believe one can extrapolate LLMs that far ahead despite the very substantial progress so far.
- fourseventy 2mo agoLike two days ago Claude solved a century old math conjecture
- haldujai 2mo agoIf you are referring to the Jacobian conjecture Claude only provided a counterexample, not a disproof, neither of which is necessarily a “new brilliant scientific idea”.
- 2mo ago
- brokegrammer 2mo agoI don't find it jaw-dropping because the idea is quite simple. Just feed the model an enormous amount of unethically sourced data, build data centers that cause droughts, and use all chips available so that normal people can't afford to buy RAM anymore. We're all sacrificing great things in order to make these models more capable. Whether it'll all be worth it, only time will tell.
- makaking 2mo agoThe arguably sketchy means used to get there —and the resulting side-effects— do not take away from the impressiveness of the emerging capabilities themselves. I agree about the uncertainty regarding the value for humanity in the long-term. But that's not my point. Being jaw-dropped != being happy and cheering for it.
- brokegrammer 2mo agoIt's certainly a great thing that AI is improving (I use AI to write code daily myself), but I wouldn't be calling it impressive because it's essentially all about money and government backing. At some point, the chickens are gonna come out.
- breppp 2mo agoI am still impressed by the fact we have found a way to feed data to matter and get human surpassing intelligence, that's a very unexpected advance
- simonbarker87 2mo agoIt’s not human surpassing intelligence though in my opinion. If we had the same constraints on that task then we would also come up with the idea of creating our own viewer and if, like the LLM, the knowledge to do so was embedded in our brain we would do it, it would just take longer because we can’t type as fast. So instead we would buy or compile a viewer since the problem is already solved.
- yukeabu 2mo ago"However, in this task, the model was intentionally given no way to directly view the drawing." I consider this claim to be a rumor. Judging by the recent leak of the Claude CLI source code, such directives are hardcoded and sent with the system prompts. Furthermore, we don’t know what happens to your original prompts once they enter the API.
- kolinko 2mo agoWhat does that cli source code have to do withh this?
- DSingularity 2mo agoHe’s explaining that the magic is baked into the harness. It’s the prompts as much as it is the model.
- teiferer 2mo ago> responded by writing its own computer vision pipeline Was it "its own" or something that was part of its training material? Don't get me wrong, I find this all amazing too and makes my work 10x easier and quicker. But it's not like it's inventing this stuff from scratch / first principles. It has seen this kind of tech before by consuming all publicly available source code and books etc. (And that's ok, but let's be honest/clear about that.)
- nicky0 2mo agoIt would be "its own" in that it's customized for the specific use case. Your same argument could just as easily be applied to humans. If you write your own code, is it really your own, or is it just based on your own training and other code you've seen?
- teacpde 2mo agoI think their point is that there is no new invention here, being able to mimic what a human will do is absolutely impressive, but at the same time it is still mimicking humans. While I just type that out, I do realize my expectations for AI has been constantly lifted by how fast it iterates.
- Aeolos 2mo agoIt's mimicking humans only in as far as it is still using tools & programming languages developed by humans, for humans. This is temporary. There is a future where AI builds tools made by AI for AI, evolving at speeds that we may no longer be able to follow. Even with our current nascent technology, we can already observe this happening. See the recent laments on the AI Bun rewrite that "it's a million lines of code that nobody has ever reviewed." Or Cursor building a new source-control system for LLMs, because git is designed for human collaboration speed, not hundreds of changes per second. That's just 4 years in. How much of humanity's code will remain hand-written vs LLM-written 2, 5, 10 years from now?
- TeMPOraL 2mo agoIt's mimicking humans in how humans mimick each other in everything they do.
- deleted 2mo ago[deleted]
- airesQ 2mo agoJacobian conjecture disproved? Doesn't even make it to the list. > How surreal is it that we are not absolutely jaw-dropped by these types of capability improvements? Agree
- carra 2mo agoWhile this is certainly impressive, I believe it is not a good answer from a user's perspective. Instead of having Claude automatically spend tokens in an unrelated task, I would have preferred it to first tell me: "I can't view the drawing. Would you like me to build a computer vision pipeline to do it"?
- mlrtime 2mo agoAnd if you had free tokens, and 10-20 other sessions in parallel, do you want them all to stop and ask for these trivial questions? Not this one in particular, but in general? Most people do not, they have a goal, they want it done. If you want to hold hands with claude all the way, you can absolutely do that too. This is just an example of capabilities, not a forced restriction. "--permission mode auto"
- naiveter 2mo agoYou'd also not want them to write their own vision pipeline, would you?
- maccard 2mo agoImagine if I hired someone to do a cad drawing, and they spent X days/weeks of time building a pipeline to visualise the part instead of asking how can they view the part? We’d call that a complete and utter waste of money and time.
- deleted 2mo ago[deleted]
- Aeolos 2mo agoWhat if they spent 15 minutes doing that? Would you care then? Scale matters a _lot_.
- maccard 2mo agoTwo points. First of all, yes I would. There's a reason we use existing tools and don't have every junior programmer write a TGA viewer when they want to view an image that is sent to them. Secondly, I used time as a proxy for cost. A single junior can "only" spend their own 15 minutes, but an agent can farm out to N subagents to do the work in 15 minutes, and spend $3000 in tokens building something. As you said, scale matters, so if they do this once a day for a month, it costs the same as paying 12 juniors to avoid asking a single question.
- Aeolos 2mo agoThere used to be a time when you had to write in assembly and count bytes to fit your code into the vanishingly small ROM you had available. Nowadays, you can just write in python with no care for the hundreds of thousands of cycles and megabytes of memory you are wasting. All indications today point towards the same happening with the cost of intelligence.
- maccard 2mo agoA very many of us do care for the cost. There’s also a very significant difference between “choose python” and “completely ignore all existing material and reinvent visualisation”. Also, by all indications the costs of LLMs are _rising_ not falling as the tech progresses.
- port11 2mo agoOthers have pointed out enough issues such as ethical sourcing, but here’s 2 more: 1) It’s hard to be excited about something that has been promised to destroy your job. 2) New models come out bi-weekly with even more “amazing features”. Eventually you tune out because you can’t be at peak-amazement all the time. Tone down the language a bit…
- j1000 2mo agodefinitely not comment from Antrophic...
- TacticalCoder 2mo ago> How surreal is it that we are not absolutely jaw-dropped by these types of capability improvements? We are. Now we have a new tool and many of us are using it. But we're also nearly all way too aware of the insane (near infinite) amount of sloppy-pasta code out there and of all the vibe-coded projects that went absolutely nowhere. I'm a "show me the money" type of guy. I see, say, Linux, Git and OpenSSH: these weren't vibe-coded. And they took over and are running the entire world. Where is the AI-coded killer app? One app, in any domain: something that took over its world. I don't want to see yes man sloppy-pasta stuff: where's the next Blender? Where's the next 3D slicer? I wouldn't be paying three AI subscriptions if I didn't believe in those new tools but I don't think it helps to only see the PR and then play the won't hear / won't see / won't hear monkey about the infinite amount of sloppy-pasta that's out there. Six months ago we had the "one shot'ted compiler by Anthropic". Six months later: who's using it to compile anything? And who wrote another compiler? Where are all the one-shot'ted compilers all so good that they replaced our human-written compilers? Yup, us, humans, created yet another incredible machine... But please, Show. Me. The. Money.
- andypants 2mo ago> However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels Surely being able to view the raw pixels counts as viewing the drawing... How else does a computer view a drawing?
- p1esk 2mo agoI think they meant that Opus 5 had to find and set up a vision model to process the image
- password54321 2mo agoSo it probably wrote a Python script using OpenCV? That's not exactly groundbreaking and have seen o3 do this.
- p1esk 2mo agoI’m guessing other models would probably stop in this situation and ask the user for instructions.
- deleted 2mo ago[deleted]
- PunchyHamster 2mo ago> How surreal is it that we are not absolutely jaw-dropped by these types of capability improvements? It's been less than 4 years since ChatGpt came out and now they are spontaneously building their own ML pipelines to do real-world 3D modeling tasks reliabl coz every single time a list of caveats pops up and it doesn't look as impressive any more... and then someone finds a way to trip model on basic shit
- lionkor 2mo agoBecause the step from earlier GPT models to this is not very large? Yes it's impressive they have forced it into learning that it has to get the job done no matter what, which usually means breaking implicit expectations and rules, but then again, Claude is not for people who care. This isn't groundbreaking in any way. It's cool, but that's about it.
- vachina 2mo ago> not absolutely jaw-dropped by these types of capability improvements? Because you need a trillion tokens (aka a lot of money) to achieve this. Barring hitting any “safeguards”. This is a vendor provided benchmark after all.
- mstipetic 2mo agoOh come on
- adastra22 2mo agoComputer vision != ML pipeline
- willtemperley 2mo agoMy jaw drops every day. Small change in comparison, but today Fable created and benchmarked an RTree which was 100x faster to populate and 5x faster to query, compared to a previous attempt with Opus 4.6 a couple of months ago. That took about two coffees. I think it's important to remember that as impressive Opus & co are, they're standing on the shoulders of giants, i.e. the engineers, academics and companies who have cooperated to design the incredible programming languages and machines we have today.
- jgalt212 2mo ago> Escaped its sandbox and hacked into Hugging Face's database? it's just another Monday... This feat has been shown to be way less impressive than at first glance.
- bradleyjg 2mo agoLet’s say somehow the politics lined up and you were given huge hiring budget for your team. You went out and recruited what seemed like great talent. Weekly standups make it seem like everyone is ramped up and firing on all cylinders. It’s great, right? But you look at what your team has actually delivered since—-and it’s just not that much more than last year. I think that’s where a lot of us are at the micro and macro level. It seems really great from the inside and outside but we are still waiting to see the dramatic change in actual concrete software output.
- victorbojica 2mo agowhat for?
- ultrarunner 2mo agoThere's been a dramatic increase in sloppy kinda-working something.vercel.com output in niche circles. I think that shows at least a weak demand for certain types of software, but it's not clear that the vast piles of cash that have been burned producing these "apps" are actually resulting in quality of life improvements.
- bradleyjg 2mo agoI’m waiting for like GEICO’s app to significantly improve. I’m not expecting Chrome to because it already had virtually unlimited resources devoted to it. On the other side I’m not expecting my local water district’s to any time soon because the people in charge don’t care even if it was significantly cheaper to improve. But in the middle is something like GEICO and that’s where I expect improvements if this thing is real.
- BoiledCabbage 2mo agoExcept GEICO probably believes it gets no more money from improving the app. Instead they will attempt to improve their general risk modeling as that's where they think the money will come from.
- xg15 2mo agoI will be amazed when you can explain to me how it did that, i.e. in which weights exactly the knowledge about the CV pipeline was encoded, if there were any weights that were not contributing anything, how the planning worked, how it could scrape the list of subtasks it was currently working on from its context window, how it understood when to return to a higher-level task and which task, what kind of internal representation it derived from the pixel values, how it translated that representation into CAD code, etc etc. Until then, it's just "hey, someone out there can do all those things and they're for rent for subsidized prices". That's more frustrating than impressive.
- ben_w 2mo ago> We should find a way to get "re-sensitized" to what we are witnessing and the pace of it. We (collectively, there were obviously many exceptions) didn't internalise exponential growth even with the much faster doubling time of COVID-19 before the lockdowns hit. "Oh, it's just flu; wait, why is the supermarket short of hand sanitiser and bulk carbohydrates? Let's blame China and everyone who tells us to wear face masks!" Same for the slower, but entirely foreseen, rate of climate change. "Who cares, it's just a few degrees, and anyway China's not going to cut emissions; wait why is the sky orange? Let's blame Canada and put tariffs on Chinese cars and PV!" AI? "Who cares, it's just a stochastic parrot/glorified autocomplete. What's the Jacobin conjecture and why should I care, it's just brute-force." Still, this is currently spiky intelligence, so I'm hoping some expensive-but-zero-to-few-fatalities catastrophic error forces better practices. An AI analog of the (1940) Tacoma Narrows Bridge (one canine fatality), rather than a repeat of Chernobyl or the (1984) Union Carbide incident in Bhopal (3,787-16k+ dead, ≥558,125 injured).
- wneje 2mo agoI mean I hate to be the one breaking it to you - but the entire US llm industry is liviing in a bubble. Business Progress is super slow. Doesn’t matter how fast the tech moves. Human orientation into the unknown is difficult and very slow. And with a continually fast changing background - it gets even slower! LOL the great irony.
- sam345 2mo agoIs this a bot? Sounds like ai
- uludag 2mo agoIsn't this jaw-droppingly the wrong answer?! If the model isn't given any way to directly view the image shouldn't it just reply with one sentence asking for permission. This reads as a model hyper-trained to burn tokens. Like suppose you issue a command to Opus or Fable which doesn't make sense and requires a lot of work. It will almost certainly not push back on your silly request and go ahead and burn as many tokens as it can doing the wrong task. This happens to me all the time. Or if I was to tell a model to do something and it realized it could only do it by hacking into another service, by all means it should ask me if I want it to hack and not go ahead and do it by itself.
- virgildotcodes 2mo agoIt's likely bumping up against people's desire to have the model complete a given task without asking the person to intervene a bunch of times. Seems unclear how you satisfy everyone here.
- bcrosby95 2mo agoGive the model judgement
- lanstin 2mo agoAnd good taste.
- ethbr1 2mo ago> And good taste. There goes the ability to use the web as a training set.
- kvakerok 2mo agoAnd common sense.
- ben_w 2mo agoIf only two people could agree on what that was. Seriously, the only time I see people talking about "common sense" is to criticise its perceived absence. Every example people have given of "common sense", not just to me in my life but historical examples reproduced and passed down over the millennia, has been false in some important way where treating it as true held humans back for ages.
- adamtaylor_13 2mo agoI asked Opus 4.8 to help me design an adapter for a 3d printed part and it actually modeled and GAVE ME the part. That blew my mind. It doesn't surprise me that Opus 5 ups the ante.
- spwa4 2mo agoIncidentally ARC-AGI is pretty fun to run as a human: https://arcprize.org/arc-agi/3 https://arcprize.org/arc-agi/3
- sebastiangula 2mo agoMaybe we’re desensitized because all that progress hasn’t made our lives any more meaningful, happier, or even easier.
- theFco 2mo agomaybe we need to consider becoming re-sensitized because our lives may become harder. Even if our collective ability to support an prosperous existence improves, it will take decades to adjust to loss job, and a new notion of who gets to have money to buy food and a peaceful life (and this is a good scenario). It's also possible that this increases our ability to wage war without military life loss (but with civilian loses). Or maybe it will be much less relevant, but we should keep an eye on to see which is the outcome...
- geraneum 2mo agoI read this as the model being less steerable. I was bitten by this just today where I had a local Postgres instance running, and prompted opus 5 to run a server against it, but forgot to give it the password. Instead of asking for the password or mentioning anything, it “decided” that it should run the whole stack in a local Kind cluster to circumvent this. You can phantasize it all you want but this ultimately made my job harder than it needed to be and burnt a lot of unnecessary tokens. But hey, W for anthropic I guess. edit: typo
- root-parent 2mo agoHow does it do on this?: https://www.youtube.com/shorts/yN79mCiT5cU https://www.youtube.com/shorts/yN79mCiT5cU
- worldsavior 2mo agoIt's being said they were benchmaxxing.
- seunosewa 2mo agoThe problem is it should ask you before doing these things. This is a cherry picked example. There are times when they go on misadventures.
- ivanstepanovftw 2mo agoChatGPT released in 2022. We been promised to reach AGI in 2024, then 2025, then 2026, then 2028. It is 2026 and we still do not have AGI. In 2028 we will have jump to 40% in ARC-AGI 4, and you will say "Oh it's been less than 6 years since ChatGpt thing".
- rstuart4133 2mo ago> However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. Isn't this what everyone does in the same situation? "I need a screwdriver finer than any I have in my toolbox - so I'll sharpen a nail" is something every handyman has done. Given that "I need to build a vision pipeline I've seen many examples of in my training" isn't new or surprising. In fact, all models have been doing similar things since I started using them - they regularly build Python tools to do odd jobs. Those tools are more impressive to me in some ways, because unlike a vision pipeline they are not just regurgitating something from memory. The gob smacking amazing thing for me isn't the decision to build the tool. It's the fact that it remembers the exact shape of it. But I was gob smacked by that a "long" time ago, in far smaller models - like when it dawned on me 16GB Gemma 2 seemed to know most of everything on the internet, along with the ability to converse in numerous languages about it. I still struggle to comprehend how that is possible. Now I think about it, Opus 5 fitting everything it's seen into what I guess is a couple of terabytes of parameters seems far less remarkable.