7 ms·
I'm currently "working" on a toy 3d Vulkan Physx thingy. It has a simple raycast vehicle and I'm trying to replace it with the PhysX5 built in one (https://nvid
by yukIttEft 5mo ago
I'm currently "working" on a toy 3d Vulkan Physx thingy. It has a simple raycast vehicle and I'm trying to replace it with the PhysX5 built in one (https://nvidia-omniverse.github.io/PhysX/physx/5.6.1/docs/Vehicles.html https://nvidia-omniverse.github.io/PhysX/physx/5.6.1/docs/Ve...)
I point it to example snippets and webdocumentation but the code it gens won't work at all, not even close
Opus4.6 is a tiny bit less wrong than Codex 5.4 xhigh, but still pretty useless.
So, after reading all the success stories here and everywhere, I'm wondering if I'm holding it wrong or if it just can't solve everything yet.
- lukan 5mo ago" or if it just can't solve everything yet." Obviously it cannot. But if you give the AI enough hints, clear spec, clear documentation and remove all distracting information, it can solve most problems.
- shdh 5mo agoMost simple problems with plenty of prior art, sure
- seba_dos1 5mo agoIt works somewhat well with trivial things. That's where most of these success stories are coming from.
- shdh 5mo agoExactly this, the SNR is polluted by this anecdata because someone was able to implement a CRUD backend they couldn’t before
- wg0 5mo agoMost of the folks are building CRUD apps with AI and that works fine. What you're doing is more specialized and these models are useless there. It's not intelligence. Another NFT/Crypto era is upon us so no you're not holding it wrong.
- MattRix 5mo agoThis is pretty wrong. Anyone who thinks this stuff is similar to NFTs and crypto hasn’t been paying attention.
- 73738488484 5mo agoIndeed this time it's different
- layer8 5mo agoMy impression is that it always comes down to how well what you’re trying to do pattern-matches the training set.
- embedding-shape 5mo agoWhen it comes to agents like codex and CC it seems to come down to how well you can describe what you want to do, and how well you can steer it to create its own harness to troubleshoot/design properly. Once you have that down, I haven't found a lot of things you cannot do.
- layer8 5mo agoBreaking down and describing things in sufficient detail can be one way to ensure that the LLM can match it to its implicit knowledge. It still depends on what you’re trying to do in how much detail you have to spell out things to the LLM. It’s almost a tautology that there’s always some level of description that the LLM will be able to take up.
- embedding-shape 5mo agoWell, not just breaking down the task at hand, but also how you instruct it to do any work. Just saying "Do X" will give you very different results from "Do X, ensure Y, then verify with Z", regardless of what tasks you're asking it to do. That's also how you can get the LLM to do stuff outside of the training data in a reasonably good way, by not just including the _what_ in the prompt, but also the _how_.
- neomantra 5mo agoWhile I’ve had tremendous success with Golang projects and Typescript Web Apps, when I tried to use Metal Mesh Shaders in January, both Codex and Claude both had issues getting it right. That sort of GPU code has a lot of concepts and machinery, it’s not just a syntax to express, and everything has to be just right or you will get a blank screen. I also use them differently than most examples; I use it for data viz (turning data into meshes) and most samples are about level of detail. So a double whammy. But once I pointed either LLM at my own previous work — the code from months of my prior personal exploration and battles for understanding, then they both worked much better. Not great, but we could make progress. I also needed to make more mini-harnesses / scaffolds for it to work through; in other words isolating its focus, kind of like test-driven development.
- nothinkjustai 5mo agoNah, it only lives up to the hype for crud apps and web ui. As soon as you stop doing webshit it becomes way less useful. (Don’t get mad at me, I’m a webshit developer)
- shdh 5mo agoI’ve noticed the models still can’t complete complex tasks Such as: Adding fine curl noise to a volumetric smoke shader Fixing an issue with entity interpolation in an entity/snapshot netcode Find some rendering bugs related to lightmaps not loading in particular cases, and it actually introduced this bug. Just basic stuff.
- computerex 5mo agoThey are definitely behind in 3D graphics from my experience. But surprisingly decent at HPC/low level programming. I think they are definitely training on ML stuff to perhaps kick off recursive self improvement.
- 59nadir 5mo agoLLMs can really only mostly do trivial things still, they're always going to do very bad work outside of what your average web developer does day-to-day, and even those things aren't a slam dunk in many cases.
- skullone 5mo agoI don't know about "only doing trivial things". I've built a fully threaded webmail replacement for Gmail using imap, indexes mail to postgres, local Django webapp renders everything in a Gmail/Outlook style threaded view with text/html bodies and attachments and a better local search than gmail, and runs all locally. Started as a "could I?" and ended up exceeding all my expectations
- shimman 5mo agoThat would be considered a trivial thing, why wouldn't it be? It's just basic crud you're doing. Nothing unique and that hasn't been written about tens of thousands of times across millions of books/blogs/comments before.
- sigseg1v 5mo agoI used it to analyze a single player game binary from steam, hook into all the relevant game state modifying functions, add full multiplayer state sync and then also hook all the relevant portions of the UI to add multiplayer to the game. I'm not saying it's the hardest thing but I also wouldn't consider it trivial.
- philpem 5mo agoThat fits with my experience. I used Claude Code to put together a pretty complex CRUD app and it worked quite well. I prompted it to write the code for the analysis worker, and it produced some quite awful code with subtle race conditions which would periodically crash the worker and hang the job. On the plus side, I got to see first-hand how Postgres handles deadlocks and read up on how to avoid them.
- wahnfrieden 5mo agoInstead of "pointing it" at docs, you need to paste the docs into context. Otherwise it will skim small parts by searching. Of course if you're using an obscure tool you need to supply more context. Xhigh can also perform worse than High - more frequent compaction, and "overthinking".