6 ms·
It’s impressive how every iteration tries to get further from pretending actual AGI would be anywhere close when we are basically writing library functions with
by petetnt 9mo ago
It’s impressive how every iteration tries to get further from pretending actual AGI would be anywhere close when we are basically writing library functions with the worst DSL known to man, markdown-with-english.
- sc077y 9mo agoWho knew that English would be the most popular programming language of 2025?
- cyanydeez 9mo agoYes. Prompt engineering is like a shittier verson of writing a VBA app inside Excel or Access. Bloat has a new name and its AI integration. You thought Chrome using GB per tab was bad, wait until you need a whole datacenter to use your coding environment.
- Alex3917 9mo ago> Prompt engineering is like a shittier verson of writing a VBA app inside Excel or Access. Sure, if you could use VBA to read a patient's current complaint, vitals, and medical history, look up all the relevant research on Google Scholar, and then output a recommended course of treatment.
- noitpmeder 9mo agoThat instantly kills the patient -- "But you asked me to remove his pain"
- duskdozer 9mo agoYou're absolutely right! I did--in fact--fail to consider the obvious negative consequences of killing the patient to remove his pain. I am truly horrified about this mistake. Let's try again, and this time I will make sure to avoid intentionally causing the patient's death. Oops--you're absolutely right! I did--in fact--fail to remember not to kill the patient after you expressly told me not to.
- tony_cannistra 9mo agoDon’t do this.
- deleted 9mo ago[deleted]
- malfist 9mo agoYou mean make up relevant sounding research on google scholar?
- deleted 9mo ago[deleted]
- bluefirebrand 9mo agoYou absolutely can use VBA to invent this information out of nothing just like AI does half the fucking time
- wizzwizz4 9mo agoI can use VBA to do that. Public Sub RecommendedTreatment() ' read patient complaint, vitals, and medical history Set complaint = Range("B3").Value Set vitals = Range("B4").Value Set history = Range("B5").Value ' research appropriate treatments ActiveSheet.QueryTables.Add("URL;https://scholar.google.com/scholar?q=hygiene+drug", Range("Z1")).Refresh ' the patient requires mouse bites to live Range("B5").Value = "mouse bites" End Sub "But wizzwizz4," I hear you cry, "this is not a good course of treatment! Ignoring all inputs and prescribing mouse bites is a strategy that will kill more patients than it cures!" And you're right to raise this issue! However, if we start demanding any level of rigour – for the outputs to meet some threshold for usefulness –, ChatGPT stops looking quite so a priori promising as a solution. So, to the AI sceptics, I say: have you tried my VBA program? If you haven't tested it on actual patients, how do you know it doesn't work? Don't allow your prejudice to stand in the way of progress: prescribe more mouse bites!
- simonw 9mo agoThe difference between prompting a coding agent and VBA is that with VBA you have to write and test and iterate on the code yourself.
- ogogmad 9mo agoGemini seems to be firmly in the lead now. OpenAI doesn't seem to have the SoTA. This should have bearing on whether or not LLMs have peaked yet.
- skybrian 9mo agoThis might be actually be better in a certain way: if you change a real customer-facing API then customers will complain when you break their code. An LLM will likely adapt. So the interface is more flexible. But perhaps an LLM could write an adapter that gets cached until something changes?
- airstrike 9mo agoThe LLM also adapts even when the API hasn't changed and sometimes just gets it wrong, so it's not the silver bullet you're claiming
- kenjackson 9mo agoI think really more than anything it’s become clear that AGI is an illusion. There’s nothing there. It’s the mirage in the desert, you keep waking towards it but it’s always out of reach and unclear if it even exists. So companies are really trying to deliver value. This is the right pivot. If you gave me an AGI with a 100 IQ, that seems pretty much worthless in today’s world. But domain expertise - that I’ll take.
- johnfn 9mo agoLiterally yesterday we had a post about GPT-5.2, which jumped 30% on ARC-AGI 2, 100% on AIME without tools, and a bunch of other impressive stats. A layman's (mine) reading of those numbers feels like the models continue to improve as fast as they always have. Then today we have people saying every iteration is further from AGI. It really perplexes me is how split-brain HN is on this topic.
- vlovich123 9mo agoOne classic problem in all ML is ensuring the benchmark is representative and that the algorithm isn’t overfitting the benchmark. This remains an open problem for LLMs - we don’t have true AGI benchmarks and the LLMs are frequently learning the benchmark problems without actually necessarily getting that much better in real world. Gemini 3 has been hailed precisely because it’s delivered huge gains across the board that aren’t overfitting to benchmarks.
- ipaddr 9mo agoThis could be a solved problem. Come up with problems not online and compare. Later use LLMs to sort through your problems and classify between easy-difficult
- vlovich123 9mo agoHard to do for an industry benchmark since doing the test in such a mode requires sending the question to the LLM which then basically puts it into a public training set. This has been tried multiple times by multiple people and it ends up not doing so great over time in terms of retaining immunity to “cheating”.
- kalkin 9mo agoHow do you imagine existing benchmarks were created?
- qouteall 9mo agoGoodhart's law: When a measure becomes a target, it ceases to be a good measure. AI companies have high incentive to make score go up. They may employ human to write similar-to-benchmark training data to hack benchmark (while not directly train on test). Throwing your hard problem at work to LLM is a better metric than benchmarks.
- pavelstoev 9mo agoNot wrong but markdown with English may be the most used DSL, second only to a language itself. Volume over quality.
- derac 9mo agoCall me naive, but my read is the opposite. It's impressive to me that we have systems which can interpret plain english instructions with a progressively higher degree of reliability. Also, that such a simple mechanism for extending memory (if you believe it's an apt analogy) is possible. That seems closer to AGI to me, though maybe it is a stopgap to better generality/"intelligence" in the model. I'm not sure English is a bad way to outline what the system should do. It has tradeoffs. I'm not sure library functions are a 1:1 analogy either. Or if they are, you might grant me that it's possible to write a few english sentences that would expand into a massive amount of code. It's very difficult to measure progress on these models in a way that anyone can trust, moreso when you involve "agent" code around the model.
- adastra22 9mo agoI’ve posted this before, but here goes: we achieved AGI in either 2017 or 2022 (take your pick) with the transformer architecture and the achievement of scaled-up NLP in ChatGPT. What is AGI? Artificial. General. Intelligence. Applying domain independent intelligence to solve problems expressed in fully general natural language. It’s more than a pedantic point though. What people expect from AGI is the transformative capabilities that emerge from removing the human from the ideation-creation loop. How do you do that? By systematizing the knowledge work process and providing deterministic structure to agentic processes. Which is exactly what these developments are doing.
- aaronblohowiak 9mo agoWe have achieved AGI no more than we have achieved human flight.
- kelchm 9mo agoAre you really making the argument that human flight hasn’t been effectively achieved at this point? I actually kind of love this comparison — it demonstrates the point that just like “human flight”, “true AGI” isn’t a single point in time, it’s a many-decade (multi-century?) process of refinement and evolution. Scholars a millennia from now will be debating about when each of these were actually “truly” achieved.
- j45 9mo agoAGI as a binary 0 or 1 existing or not isn't the thing that interests me to look at primarily. Is the technology continuing to be more applicable? Is the way the technology is continuing to be more applicable leading to frameworks of usage that could lead to the next leap? :)
- mrcwinn 9mo agoI think you're missing the point.
- nimchimpsky 9mo ago[dead]
- ETH_start 9mo agoIt's clear from the development trajectory that AGI is not what current AI development is leading to and I think that is a natural consequence of AGI not fitting the constraints imposed by business necessity. AGI would need to have levels of agency and self-motivation that are inconsistent with basic AI safety principles. Instead, we're getting a clear division of labor where the most sensitive agentic behavior is reserved for humans and the AIs become a form of cognitive augmentation of the human agency. This was always the most likely outcome and the best we can hope for as it precludes dangerous types of AI from emerging.
- DonHopkins 9mo agoMarkdown-with-English sounds like the ultimate domain nonspecific language to me.
- baq 9mo agoAnd yet the tools wielding these are quite adept at writing and modifying them themselves. It’s LLMs building skills for LLMs. The public ones will naturally be vacuumed up by scrapers and put in the training set, making all future LLMs know more. Take off is here, human in the loop assisted for now… hopefully for much longer.