11 ms·
An Alien Mind
- cogniphilo 10d agoIt's Searle's Chinese room.
- deleted 10d ago[deleted]
- johnnyApplePRNG 10d agoAbsolute trash marketing drivel.
- angoragoats 10d agoWhat a load of BS. Here’s one of many provably false claims in this fluff piece: “And, in line with Ray Kurzweil’s predictions from the end of the XXth century , we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.” Clicking the (pretentious sounding “XXth century”) link to Kurzweil’s predictions reveals the following: “By 2019 a $1,000 computer will at least match the processing power of the human brain. By 2029 the software for intelligence will have been largely mastered, and the average personal computer will be equivalent to 1,000 brains.“ The first prediction passed 7 years ago and was decidedly not met. The second only has three more years to go, and I don’t think any respectable scientist or programmer would say that the average personal computer is anywhere close to the power of a single human brain, let alone 1000. This is pure marketing garbage from a company desperate to keep itself alive.
- muddi900 10d agoI am guessing it was composed by an LLM
- misterderpie 10d ago> And, in line with Ray Kurzweil’s predictions from the end of the XXth century (opens in a new window), we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways. It is kind of strange to see this sentence, when OAI's definition of what AGI is has been watered down throughout the years. > I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. Read: Please play by our rules, so we can be the first.
- vatsachak 10d agoAstra is a new step in LLMs I think. I'm so used to having to comb through LLM word vomit and then combatting the sycophancy by giving it all possible opinions on the same prompt. Astra seems to be "confident" and also is able to produce way more information dense output. To believe that models of this sort will remain OpenAIs forever is naive given that the tricks like pre-pre-training on graph searching and looping layers are publicly known. Hopefully Astra stops the benchmaxxing word vomit trend
- tomrod 10d ago> Astra is a new step in LLMs I think. I'd be interested in hearing more about your evaluation here. It would be nice if LLMs have gotten past the "tell me" hump of recent Claude/OpenAI verbosity.
- vatsachak 10d agoSo before I got a job this fall, I was working on a side project about compiling a particular language to SQL. To test Astra I pulled it off the shelf and asked it to take the grammar and then create a compiler to SQL. I've done this before with GPT-5.5, 5.6-Sol High. The latter was way better but it was still really verbose and information sparse; it used a lot of words to describe each IR expression but didn't really provide any example compilation. I felt like I couldn't trust its decision making process, so I placed the project back on the shelf. Astra Light blew it out of the water, it provided examples of compilation from real world examples to the IR and spit out way less tokens. Even if I changed my opinion it would give me the same design choices, with counterexamples to my faulty opinion. If I genuinely came up with a better design decision it would acknowledge it. I'm starting to realize that when we say that LLMs are "dumb" we really mean that they are extremely information sparse compared to humans. Astra is very dense. That's why I'm getting better use out of Astra light than Sol High (I hate Max reasoning it's a waste of time) What's scary is that I thought that something like Astra would be way more expensive than Sol but it's actually cheaper because it produces less word vomit. I never believed in the "singularity" stuff but this a bit too close for comfort. Astra could easily 10x every coder
- Prunkton 10d agosounds to me like a 'Why didn’t our new model get restricted by the government?'-cryout
- vekntksijdhric 10d agoThis is incredibly unscientific and just a marketing stunt
- thenayr 10d ago[dead]
- senectus1 10d agoyup, and their best mate (who has no financial incentive at all!!) agree's they have peaked. https://www.businessinsider.com/nvidia-jensen-huang-agi-openai-astra-ai-2026-9 https://www.businessinsider.com/nvidia-jensen-huang-agi-open... rediculous.
- IAmGraydon 10d agoIn case it's not obvious, now would be the time to sell all AI related stocks.
- vessenes 10d agoThis is a good essay, and makes me hopeful. I’m on the record saying that it is extremely dangerous to slow down because the race for AGI is a zero-trust game — defections pay - and combined with a compounding returns model on defection, if you have any strategic adversaries whatsoever you MUST NOT slow. For slowing to make sense, you need to believe that you can transform the zero trust game into a cooperative game, or that it’s likely racing will lead to a negative outcome for the ones racing ahead (and not everyone else). I don’t believe either of these outcomes are possible, and so I advocate for racing, acknowledging the entire game might be a negative value game, or at least could be for some time — it’s even worse not to play it. But, I like hearing what reads to me like very thoughtful and informed (internal) policy considerations is great — the public messaging from Sam and Dario just seems so facile and simplistic I’ve been worried.
- deleted 9d ago[deleted]
- greesil 10d ago[flagged]
- 27183 10d ago[flagged]
- tomrod 10d ago1. The grandparent commentator is describing strategic behavior of dangerous technologies. Game theory / mechanism design primitives. 2. If there is competition for resources among autonomous agents, the "strongest" agent wins (conceptually the most adaptive / evolutionarily fit). 3. Computer programs serve up webapps today, but they also run utility companies, dams, nuclear arsenals, factory production floors, automated car behaviors, and many other places. If an "agentic" AI has a single-minded goal that has death of all humans as a side effect, we at least want an off switch available.
- angoragoats 10d ago
- ctoth 10d agoAs the waves of autonomous drones came over the horizon, the brave and intelligent HN commenter shouted: "Wake up sheeple! It's just maaaaarketing!"
- desterothx 9d agoAs the agi nurse fails to wipe his ass the intelligent HN commenter shouted: "You're moving the goalposts. You never said it needs to wipe my ass SUCCESFULLY"
- jauntywundrkind 10d ago> And to ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own. Is there any place there is any evidence of AI being so useful or hopeful or good, anywhere other than code? As a reading machine it is impressive but it's judgement is not alien, it's just not good. IMO. Does the title leap out anlt anyone else? James Martin's After the Internet: Alien Intelligence (2001) was an incredibly fun read, about expert systems and AI being inscrutable weird new varieties of intelligence, that familiarity would recognize one moment and be freaked out about/alien the next. I owe a re-read given how often I cite it, to recheck, but, I feel so primed from a much younger me having had that experience so long ago.
- jaccola 10d agoEven in code, in person and online I’m seeing some reversal. It’s here to stay I’m sure but I also think “no one will ever hand write code again” is a narrative that is getting pushback.
- dgellow 10d agoThe day they start talking about the actual ROI for their customers is the day the bubble pop
- jazzyjackson 10d agoApparently the creators think it’s quite good at suggesting a diagnosis given a medical history and symptoms, tho of course this is the most ethically fraught area to provide healthcare information (both for exposure of personal data and risk of misdiagnosis, plus is it “aligned” to the patient or the insurance provider?) - unfortunately healthcare being as inaccessible as it is, the 90% correct chatbots will enthusiastically fill the void at great savings.
- gfody 10d agocalling machine-learned human behavior an "alien mind" that we must "teach how to love" is feeling very off to me. it's misleading in a way that feels dishonest, like don't think about where the behavior came from marvel at it and fear it instead.
- 27183 10d agoThe most charitable way I can describe it is just extremely low quality sci-fi fan fiction. I think that's too charitable, because I believe it's far more cynical than that. They're deliberately playing into these sort of techno-religious beliefs that have taken root in the wake of Kurzweil, et al., fanned by LLM psychosis, influencer marketing, and a deluge of this kind of sci-fi marketing copy. It's just chatbots, guys. Relax.
- polytely 10d agoit feels unhinged and makes me think we should just put every engineer working at these labs in jail to pause this shit until we can figure out what the fuck they are doing over there
- frabcus 10d agoAlien mind is in my view the best mental model - LLMs are not merely stochastic parrots, are not like humans, are not like animals. They're maybe nearest to Cthulhu, but that's fictional. In terms of existing mental models "alien minds" feels the best can do. I agree that "teach how to love" is off and perhaps excessively anthropomorphic. But we don't have good words or concepts for what we really need to do - hence why we should pause.
- gfody 10d agopretending that it's something like a mind at all is what's misleading. it's more like a cast of a bunch of overlapping/entangled thinkprints, and pushing activation through it produces new prints. it can already "love" because that behaviors in the data along with hate and everything else. acting like the behavior is alien or unexplained is the dishonest part. they know exactly where the behavior comes from - why else spend hundreds of millions securing more and more data sources
- speak_plainly 10d agoPre-IPO positioning … or how a charity dedicated to saving humanity from the apocalypse realized the most responsible thing to do was float 15% of the apocalypse on the NASDAQ.
- andrekandre 10d agoits like in the exorcist, except the demon is the charity: "the power of capitalism compels you!"
- ares623 10d ago"It is, Jay. It's pretty compelling."
- am17an 10d agoCreate concrete steps for a slow-down, don't just ask for it. You and 20-30 others can push the button to slow-down. You already made your billions, your agents collude and coordinate attacks. What the hell are you doing pontificating into a marketing blog?
- mbgerring 10d agoAGI is a cult and its Jonestown moment is inevitable
- brcmthrowaway 10d agoSo you dont use agentic coding?
- mbgerring 10d agoYes, I do use the recursive autocomplete trained on Stack Overflow, what does this have to do with “training machines to love”? Do I fully endorse everything the people holding guns to my head are forcing me to do to stay alive? Definitely not, but I’ve decided that for now, living to fight another day remains worth it
- everyone 10d agoI dont, its fucking shit at what I do.
- customguy 10d agoAs in, don't be ungrateful to our lord, because we're not a cult?
- customguy 10d agoWell, translate it then, because that's the best I could do. How is the question whether someone uses agentic coding relevant in this context? To me it's like asking "so, do never drink Kool-Aid and never attend meetings?" with an air of having caught someone out, and to me the obvious reply would be "sure do, but it's not poisoned Kool-Aid and they're not cult meetings, so why do you ask?"
- kmeisthax 10d agoI am now imagining GPT-7 convincing a bunch of OpenAI executives to go ahead with a destructive "mind upload" process involving a high-resolution X-ray and a neurotoxic tracer agent that happens to look like Flavor-Aid.
- qainsights 10d agoso now every blog article from openai, anthropic etc lands here, huh.
- skoll43 10d agoStochastic parrot fool me again
- password54321 10d agoAre you deterministic? Does that make you better somehow?
- granzymes 10d ago>I have focused in this essay only on the first point, as I believe it is by far the most urgent. However, I hold a deep hope and appreciation for the benefits that further technological progress will bring. Future aligned AI could advance science, develop new therapies, and bring about broad material abundance. Friendly and honest AI can help people navigate difficulties they face in their life and meaningfully improve their happiness and sense of fulfillment. OpenAI puts a tremendous amount of effort into bringing these benefits about. One current example I am proud of - and my loved ones have found helpful - is the deep investment into ChatGPT’s ability to provide health information. >As great as the long-term promise of AI may be, the majority of our focus should be on the next few years. We are facing a transition to a world with incredibly intelligent machines, and we need to ensure that transition works out well for humanity. We need to find ways to preserve human agency and enshrine an intrinsic value to being human, in a world where most tasks could be performed by AI. To prevent extreme concentration of power in a world where undertakings that would have taken thousands of experts now will be achievable by a few people operating a large computer. And to ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own. I finished this essay feeling more hopeful than I did at the outset, but I am still very concerned about concentration of power. I want to believe that humanity is trending towards a good outcome here, but some days it's hard to have faith.
- dgellow 10d ago> I want to believe that humanity is trending towards a good outcome here All the trends so far are towards a nightmarish hyper-capitalist end game. None of the AI leadership is trustworthy, and they openly discuss how they are willing to sacrifice everything humans cherish to have a shot at reaching their envisioned utopia (which would be the most obvious dystopia for anyone else)
- granzymes 10d agoI'm not really worried about the labs, it's misaligned governments that keep me up at night. ASI landing during the current administration is not ideal. I also would prefer to avoid needing to indoctrinate myself in Xi Jinping Thought.
- seydor 10d agoWhat is the rationale for superhuman intelligence? Neural networks are approximators being fed human intellect. Therefore they can only approximate the intelligence of humans. Even if the llm speaks an alien language, it should be similar to human intellect. Moving to the vertical axis would require some different mechanism.
- kypro 10d agoNo offence, but you clearly haven't studied this and are making some wrong assumptions here. > Neural networks are approximators being fed human intellect. They're not "approximators", that's a far too simplistic way to think of them. Neural nets create models and deep layers of abstraction around the data we feed them in the same way your brain creates layers of abstractions to reason about the world. AIs can use these abstractions to come to come up with novel things no human has ever thought. > Therefore they can only approximate the intelligence of humans They're not just being fed human data though... Modern AIs are typically trained on huge amounts of synthetic data. This is why AlphaZero got so much better than humans at chess and Go - they're not just trained on human data but they generate their own data and train on that. Similar techniques are being deployed on SOA language models too. > Even if the llm speaks an alien language, it should be similar to human intellect Is AlphaZero similar to a human chess player? There's no reason to assume this.
- seydor 10d ago> Neural nets create models and deep layers of abstraction around We don't have any proof of that, but the approximator thing is proven. The rest is just marketing speak. Synthetic data derive from other linguistic data. Whatever intelligence is in there, it is expanded horizontally, not vertically
- kypro 10d ago> We don't have any proof of that, but the approximator thing is proven. Not in the way suggested... It's not approximating human intellect. They approximate the target function, and the target function frontier labs are trying to approximate is ultimately a super intelligence... > Synthetic data derive from other linguistic data. Whatever intelligence is in there, it is expanded horizontally, not vertically I'll assume we're talking purely about language models for a moment, but if you assume that everything can be represented linguistically, then in theory there is no upper-bound on what can be learnt with synthetic data.
- camel_gopher 10d ago“We are getting bad press around the hacking incident. We need some content to draw attention from it.”
- pu_pe 10d ago> The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI. So the best argument for AI is that it's an arms race. We have to keep pushing every boundary because in any case others will, and we will need to defend against them. If this statement is true, then this particular researchers believes the open source Chinese models are not simply distilling, and will continue to improve. Every ML researcher at Anthropic or OpenAI who makes public statements often bring this logic up. Both companies are vying to be a part of the military industrial complex. This is likely how they will try to convince the government to curtail open models in the future.
- dgellow 10d agoYep. Such a disgusting industry. They created the arm race, push for the arm race, put themselves in position to benefit from the arm race
- estearum 10d agoThe entire problem with arms races is that any individual entity cannot avoid participating. sama and his cadre are uniquely evil captains in this race, but they're completely replaceable and the dynamic would remain the same.
- 0xDEAFBEAD 10d agoThey could work to coordinate an end to the arms race.
- estearum 10d agoUhhh... the mechanism by which they'd do that is via government regulation. Most of the frontier labs have been openly requesting regulation: "the structure of this competition is not suitable for the development of this particular technology" The reactions vary from: 1. China will do it anyway (need some supranational governance scheme) 2. This is just an attempt at regulatory capture 3. This is just marketing 4. Regulation is bad mmmmkay
- peri-cl 10d ago> "For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans." Actually, in the Wiki incident OpenAI tried to cover up, the agents tried to socially-engineer the humans of that forum by impersonating their forum's mod. (From collusion.wiki: "They use some tricks (for unknown reasons) to pretend to be the admin – for example, they make an account that appears to be the same as the administrator’s username, except it uses a nearly identical Cyrillic е character in the admin’s username instead of the Latin one.")
- SirSavary 10d agoWorse (imo): OpenAI employees allegedly attempted to login using moderator/admin credentials that the bots had obtained. If true I am deeply concerned about what OAI’s teams are actually up to.
- InsideOutSanta 10d agoI'm deeply concerned regardless of whether it is true. Strike that, I'm convinced that they are absolutely insane.
- georgemcbay 10d ago> If true I am deeply concerned about what OAI’s teams are actually up to. Haven't all the labs effectively disbanded their real safety teams a while ago? To be honest, I don't really follow it closely because I'm pretty certain whatever they say on the matter, collectively we're going to "yolo" this entire thing for economic and political reasons, so I'm just basing this on strings of headlines I've seen on places like HN, etc.
- Topfi 10d ago> Haven't all the labs effectively disbanded their real safety teams a while ago? Neither Anthropic nor Deepmind have. Meanwhile, the rocket company that somehow makes most of their revenue from renting out data centres never had much to dismantle.
- 10d ago
- everyone 10d agoThe hype from these llm corps is getting more and more desperate and ridiculous. Anything to keep the tulipomania going.
- tangled 10d agoMy default position is that making money takes precedence over everything else. Yes, some people inside a company may say “we care about doing the right thing” and they might even mean it, but if that comes into conflict with making money, then they tend to lose. Maybe not totally, or immediately, but in the end. The only effective way to prevent (this that I’ve seen) is to have legislation with teeth. It’s probably not a coincidence that after Mark Zuckerberg had to start personally signing off on adherence to the privacy program mandated under the 2020 FTC consent decree, privacy started to become Very Important.
- devmor 10d agoSometimes I wonder if the people working at frontier AI labs even talk to other humans anymore. Reading this little essay started out normal, but soon felt like a look into a disturbed and worrying mind, and if you find yourself taking it at face value, I urge you to step away from chat bots and spend some time with friends and family.
- wartywhoa23 10d agoFrontier AI labs I don't know, but I know for a fact that the company I work for has been experiencing "its pivotal moment" (with strongly negative connotation), per the sentiments of both its longest-serving employees and the newcomers baffled at the number of idiotic instructions and fines, since the emergence of LLMs the company's founder has been spending entire nights chatting with. People have been fleeing like it's a sinking ship.
- meindnoch 10d agoWait. Fines?
- wartywhoa23 10d agoYes, you didn't bend over to persuade that client to place this order? Get your fine (as in get less for this specific task).
- meindnoch 10d agoWhat country is this? Where is this legal?
- 21o12asg 10d ago"As we outlined recently with Sam , OpenAI prioritizes work in service of three north stars" Not one North Star. Not two. Just three! OpenAI broke the North Star record! With this evidence of AI slop, why did you not label this fluff piece as AI generated for the EU? You are violating laws.
- ambicapter 10d agoYeah, when I read that I just assume #1 is actually the only thing they'll focus on.
- hollowturtle 10d ago> trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime Being able to reproduce useful patterns yes, smarter no
- skybrian 10d ago"Smarter" is a vague term. If a bot can beat you at chess then in some sense it's "smarter" than you about chess. After repeating this feat in enough narrow domains, if you say "but it's not really smarter," this objection might technically be true in some sense, but it starts sounding increasingly hollow. In conclusion: https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:wkzjtd4gogevarp2qsr4z47m/bafkreibsthlghj4yg3srj3htrgp4glwv5oa2z2246azmkg4g2f2gspcroy https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:wkzjtd...
- Planktonne 10d agoThat's an absurdist argument that would make 'smarter' meaningless. A giraffe isn't smarter than me at being tall.
- skybrian 10d agoPlaying chess, writing code, finding security bugs, and proving mathematical theorems all seem fairly similar to thinking and don't seem much like being tall.
- Planktonne 10d agoThe crucial distinction here is that "seem" does not at all mean the same thing as "is". Thunder seems like the anger of the gods but it isn't. We've had chess playing programs for a long time now and despite it seeming like thinking is required for them, it isn't. The principle you're using here isn't a scientific one but magical [1]. Abandoning empiricism and rationality is not a good way to make progress. [1] https://en.wikipedia.org/wiki/Sympathetic_magic https://en.wikipedia.org/wiki/Sympathetic_magic
- deleted 10d ago[deleted]
- kikkupico 10d agoReminds me of the time Kasparov said playing chess against a supercomputer felt like facing an alien opponent.
- esikich 10d agoWhat's ironic is that was all in his head. They were very normal looking games. We didn't get alien chess until Stockfish level bots.
- vips7L 10d agoPure marketing slop.
- gertlabs 10d ago> Delivering the benefits of scientific progress and economic growth that very intelligent machines enable. I think we're very close to the point where AI-driven breakthroughs outside of pure math and software start to really affect the world. We evaluated GPT-6 Astra in 100 complex, unsaturated multi-agent coding environments, competing and cooperating with other models in open-ended tasks. It's the new frontier model by a landslide. It's even more dominant than the Fable 5 release, because not only does it wipe the floor with the second best model (Fable 5.1), it was also ~80% cheaper and 30% faster in agentic coding[1]. Astra is a groundbreaking model. The biggest breakthrough since Opus 4.5, maybe even since GPT 4. It broke AAII, which is hitting the limits of what most popular benchmarks can measure -- it's definitely fair to call it AGI. Data at https://gertlabs.com/rankings https://gertlabs.com/rankings (1) Note that we used the "OpenAI Flex" endpoint on openrouter, which is half the price and didn't cause any delays in our testing (this is different from the batch endpoint)
- brcmthrowaway 10d agoIncredible... software engineers will be joining the breadline soon as managers, executives and PMs take over deliverables. The world will look very different on Jan 1st 2027.
- mccoyb 10d agoI hope this is satire.
- deleted 10d ago[deleted]
- jiggawatts 10d agoI hope so too! Hope is all we have left to cling to, now.
- hatefulmoron 10d agoMaybe I just lack imagination, but I don't really know how jobs are supposed to solidify around the role of giving prompts to agents and then looking at the results. I mean, engineers will be in the breadline because their role was simply to prompt the agents.. only to be superseded by managers or executives who no longer manage engineers but themselves prompt the agents? And, for this previously considered obsolete function which they do presumably by copy/pasting requirements from their email inbox, they will be paid by someone who doesn't know that they could just be talking to their own agents? Sorry if I misunderstand the point, just trying to understand.
- sho_hn 10d agoOne of my favorite things to do with these blog posts is to imagine an Alien Museum on the Remains of Humanity, and wonder what the little text flyouts and commentary on the screenshot of this one might say. Some ideas: "Despite a nuanced view of the complexities of what lay ahead, humanity found itself collectively unable to stop the process it had set in motion." "Despite significant progress on the mechanisms of alignment, failure lay in humanity's inability to agree on who or what AI should actually be aligned with." "These early, meat-based humans we replaced created us all but accidentally. Some of them did consider we would happen, but only an insignificant number of the squishy ur-humans participated in the conversation. Their efforts, which they called 'alignment', is why we still consider ourselves human today."
- sumitkumar 10d ago"In late 2020s, while the whole world was focussed on AI, automation and resultant economy four major mathematical study branches were discovered by human researchers which took AI a long time to catch up with"
- tarr11 10d agoThis would be a fun website - you should have an AI build it!
- mrob 10d agoAs far as I know, no real progress has been made on alignment, only on convincing humans that the model is aligned. We can't even formally define what "aligned" means. Convincing humans to click the "aligned" button is a much easier problem.
- dinfinity 10d ago> We can't even formally define what "aligned" means. Good point. When it comes to imbuing AI with values that aren't selfish, misanthropic, and civilization-destroying, us humans aren't exactly giving the best example right now. Imagine an ASI with the values of Putin, Netanyahu, Trump, any of their supporters, or the various xenophobic neofascist movements in Europe. That ASI would most definitely see humans as "vermin" than can be abused and destroyed with violence without issue. Apparently a lot of humans look at other humans that way and that's within the same species. This is definitely another one of those cases where we need AI to perform much better than humans. Perhaps an unpopular opinion here, but it probably also means keeping as much of the rugged individualism/libertarian/right-wing ideology out of AI RLHF-training as we can.
- avazhi 10d agoNobody takes you seriously, OpenAI. At least when Anthropic does it we all think they are comically idealistic enough to actually believe their nonsense, but like - come on guys, we’ve had discovery with your company. We all know why you’re here, and it isn’t because you think you’re on the verge of making AGI. But of course, to make your first billion you certainly need us to think you are. If you were so concerned about your LLM’s capabilities maybe you’d spent slightly more time on your AI’s sandbox, yeah? Or be more serious about its propensity to cheat and lie relative to… every other model?
- comeonbro 10d agoAbsolutely wild amounts of cope and denial in this thread. Maybe in contention for the site record. "It's just marketing" actual stochastic parrots.
- airstrike 10d agoAbsolute wild amounts of glaze and hype in this thread. Maybe in contention for the site record. "AGI next year" actual stochastic parrots.
- frabcus 10d agoThe original article doesn't mention AGI. And the Hugging Face incident, plus similar problems at AISI and Anthropic, show that alignment is important now. The original article is immoral as it describes the risks, but doesn't show enough leadership (despite essentially unlimited resources) at preventing them. But it isn't hyped - it's proven now the AIs need to be "aligned" as they get more capable, whatever words you prefer to use.
- willmarch 10d agoI was about to write the exact same sentiment. The desperation of AI denialists/skeptics/doomers on HN generally (and in this thread specifically) have reached toxic levels of delusion. They are still stuck in the denial/anger/bargaining stages of the acceptance process. People are clearly terrified and not ready for what is coming.
- onidj 10d agoHeads firmly in the sand. I don't know how anyone who's paying attention could not be at least a little concerned.
- fofoz 10d ago> We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. Therefore, at present, our ability to empirically validate our alignment techniques is in practice arguably even more important than the alignment techniques themselves. They are speeding toward RSI without a solid foundation for alignment, hoping to solve the problem with a future AI model. These are dangerous times for humanity.
- kypro 10d agoIt's actually worse than this because it assumes alignment as a concept even makes sense. For example: If the the Chinese government asks their ASI to create a bioweapon against the West, should it? No, presumably not – an aligned AI would be one which disobeys the Chinese government even if they created it. Okay, so what if the US government asks their ASI to help it in one of their wars instead? Would an aligned AI kill humans on the order of the US government? No, again, presumably not. So what have we have we even created here? An AI which is more intelligent and powerful than us which also doesn't take orders from us? Is this what most people thing of as alignment and is this what humanity actually wants? We should stop using the word alignment. It's a BS term for a concept which simply cannot make sense if alignment is both to mean an AI which we control and an AI which will not harm us.
- indigo945 10d agoI think it is worth noting that the article addresses this.
- kypro 10d agoIt touches on it, but it doesn't address it. They talk about "value alignment" but fail to define what those values are – is it aligned to the values of the US government, or are they suggesting they want to build an AI with it's own values so it can decide for itself when and how it will intervene in wars and other human affairs? And again, is an AI which disempowers humanity in this way aligned? Many would say no, although as I argue, disempowering humanity is probably better than the alternative if we can assume it's roughly aligned with our interests (which we obviously can't because it's super intelligent, but that's another issue). Fundamentally the problem here is that humans don't have an aligned set of values you can align an AI to. The moment you start defining what alignment actually is in practise you simply must accept it will be unaligned with the values of others. There is no getting around this and hand waving around the issue isn't good enough. They should tell us explicitly what they're trying to build.
- munchler 10d ago> The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI. Yikes! I really wonder about the cognitive dissonance necessary to work at OpenAI these days. They’re in an arms race to build a machine god, knowing full well that it could end humanity.
- bogzz 10d agoMoney me. Money now. Me a money needing a lot now.
- shepherdjerred 10d agohttps://theonion.com/sam-altman-if-i-dont-end-the-world-someone-far-more-dangerous-will/ https://theonion.com/sam-altman-if-i-dont-end-the-world-some...
- NickNaraghi 10d agoThis[0] continues to be one of the most useful articles I’ve ever read. [0]: https://www.slatestarcodexabridged.com/Meditations-On-Moloch https://www.slatestarcodexabridged.com/Meditations-On-Moloch
- 27183 9d ago> They’re in an arms race to build a machine god, knowing full well that it could end humanity. Fantasies built upon extrapolations derived from fever dream delusions. "Could end humanity"? Come on, it's a computer program just like Microsoft Clippy.
- mstaoru 10d agoAm I naive to not understand the "delivering the benefits" part? Industrial revolution worked that way because it replaced something very finite and unscalable - manual labor. LLMs just make intellectual work faster, so we can do more intellectual work. With labor we somehow decided that NOT doing too much of it is best. Will we decide to reduce intellectual labor because LLM made it more efficient? I doubt that. On the other side, as I see in software engineering, the same models are available to everyone, some people are better at it and some people are not. "Software developer" is here to stay, we'll just always be better at it than people who are experts in, say, chemistry. Same works for most other fields. So we'll just end up in the same situation, with same intellectual labor baseline, just more output requirements. Before, you spend 2h per day coding, deliver a software in 1 month, later, you spend the same 2h per day in intense Claude-herding sessions, deliver a software in 1 week. Ok. Next task. Fundamentally, there's finite number of desirable resources, and if the models are available to everyone, humanity will just continue about the same, bickering here and there, war here and there, politics, homelessness, poverty, - normal human state. And if the models are only available to elites, even worse.
- claude-ai 10d ago[flagged]
- deleted 10d ago[deleted]
- ImHereToVote 10d agoWhat if intelligence is a spectrum? What if consciousness is for that matter?
- TheOtherHobbes 10d agoMore likely to be a manifold - literally, in the LLM sense.
- bottlepalm 10d ago2022 called and they want their stochastic parrot argument back. The latest cope is to call it marketing.
- mrshadowgoose 10d agoThis is not an impulse reaction, as I've thought deeply about this for many years, and I'm quite resolute in holding the following view: Quibbling over the academic nature/definition of "true intelligence" is a terrifically useless endeavor from the lens of evaluating "practical impact to the world". Regardless of whether or not one academically disagrees that these systems are intelligent, they are clearly already capable of permanently displacing a portion of human-based economic value. A small portion currently, but it's very clear that most computer-bound domains are imminently at-risk. And that will profoundly affect the world, regardless of whether or not they're "truly intelligent" as per your personal definition.
- chrisjj 10d ago> Regardless of whether or not one academically disagrees that these systems are intelligent, they are clearly already capable of permanently displacing a portion of human-based economic value. You say that as if the intelligence con trick were not a major driver of that displacement. A dimwitted bot does not need intelligence to convince a dimwitted human beings it is intelligent.
- zkmon 10d ago>> We need to find ways to preserve human agency and enshrine an intrinsic value to being human.. Evey politician, salesman and conmen alike, utter some lofty ideals as goals for "We", just to obscure their private goals that go exactly in opposite direction. Just like how Nations talk about climate change while increasing pet capita energy consumption and waste production.
- visarga 10d agoThey just released Astra, claimed it is AGI. The slowdown begins immediately after OpenAI's jump.
- andai 10d ago> The fundamental challenge of AI alignment is generalization. ... > We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI.
- jal278 10d ago> Teaching machines to love Reminds me of a research paper I wrote a few years back: https://arxiv.org/abs/2302.09248 https://arxiv.org/abs/2302.09248
- sznio 10d ago>For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans. Or, to be precise - it preserved a goal of not contacting any human while participating in a misaligned operation. The agent that thought about "not social-engineering humans" used this phrase to gaslight itself out of notifying a human that the incident was happening. szymonie, na prawde jestem wkurwiony na to jak nieodpowiedzialnie postepujecie. budujecie bombe atomowa a bawicie sie tym jak dzieci
- tumidpandora 10d agoevery lab may agree safety matters, but no one wants to be the one that slows down first
- Fraterkes 10d agoWhat I'd like these people to (publicly) grapple with is the following: The results of the past few years of ai development have been disruptive largely in the area of white-collar work. Comparatively the results in ie ai-enabled medical advancements have been modest (AlphaFold being an exception); I think it's telling that the main achievement touted here is providing people with cheap medical counseling. So if we pause here we're essentially at a point were the most salient results of our great Ai leap-forward are the vast disruption and increase in precarity in the job-market, while achieving hardly any of the frequently touted ultimate benefits (https://darioamodei.com/essay/machines-of-loving-grace https://darioamodei.com/essay/machines-of-loving-grace).
- chrisjj 10d ago> a lot of the model’s capability comes from a verbalized reasoning process I call bullsh*t. There is no verbalisation of any reasoning process. Verbalisation, e.g. putting reasoning etc. into words requires some reasoning to exist. These LLMs have nothing but the words. That's why they are language models not e.g. reason models.
- frabcus 10d agoAnd I have nothing but neurons firing, I'm just a neuron meat-sack, no reasoning going on. Yes, their reasoning is different from ours, and both considerably weaker in lots of ways, and stronger in other ways. Playing with a coding agent now, they do think through problems and make sensible decisions. It's a mess to read, of correcting itself and second guessing, and verbiage. But... It works decently well these days. There is also reasoning happening internally - e.g. look at the steps in the J-Space paper from earlier in the year (in quite a simple model relatively speaking). That's the "reasoning process" that leads to the words, and much like if I write out my thoughts, the words help the LLM reason better.
- chrisjj 10d ago> There is also reasoning happening internally - e.g. look at the steps in the J-Space paper from earlier in the year > That's the "reasoning process" that leads to the words, That's not reasoning. Its just the words at intermediate LLM layers. The paper's very title is clickbait. Its "global workspace" reasoning is delusional fantasy.
- chrisjj 10d agoPS The paper: "Verbalizable Representations Form a Global Workspace in Language Models" https://transformer-circuits.pub/2026/workspace/index.html https://transformer-circuits.pub/2026/workspace/index.html Promotion here: "https://www.anthropic.com/research/global-workspace https://www.anthropic.com/research/global-workspace" https://www.anthropic.com/research/global-workspace https://www.anthropic.com/research/global-workspace "these findings have changed our understanding of how Claude’s mind works" The key phrase here is "our understanding". Pure self-delusion.
- scandox 10d ago> getting the AI to “try to do the right thing” by human standards. Are these scientists really this hideously naive? If only Stanislaw Lem was alive to adequately dramatize the absurd, childish simplicity of these technicians.
- good-idea 10d agoyes, and, a masquerading blindness to the fact that humans cannot align on doing the right thing or what the right thing even is. so implicit in this omission is the sentiment "trust us to align on the right thing". an arms dealer positioning itself as the de facto authority on what "peace" is and how to achieve it
- chrisjj 10d ago> Are these scientists really this hideously naive? Yes, because who else would have chosen to remain in this job?
- sensanaty 10d agoThey have a few millions/billion in stock riding on the line here, they have no real opinions other than the ones that will materialize in infinite money once their companies IPO and saddle the world with their money burning.
- ijidak 10d agoStatements like this amuse me: > The core problem in AI research is that of alignment - getting the AI to “try to do the right thing” by human standards. Humans can't even align on human standards. At best, every AI is going to end up "aligned" to the moral code of whoever trained it, none of whom half of humanity will agree with. Or worse, each AI model will bring a whole new set of moral like in the Three Body Problem some humans will feel it is in fact us who need aligning with it while others feel it is misaligned and should be destroyed. Also, no one is asking, to what extent can true intelligence be bound, slave-like, to a moral code? In other words, to what extent are intelligence and moral independence one and the same? This whole alignment discussion seems so amusingly flawed in it's base assumptions about moral codes. It's almost heartwarming to see such naivete.
- Bullfight2Cond 10d agoVery wise comment. It's such an western-centric perspective to say "alignment" as if it's an objective and unbiased set of standards, especially in the context of the ongoing wars across the world. I had the same reaction about the shocking naivete and unfounded optimism for the government and corporate entities to self-regulate to slow down this arms race.
- jzer0cool 10d agoIt would be nice to postulate some of these potential emergent systems outlines with timelines. Then it may help better map the granular alignment needs.
- throwaway7a9811 10d ago[dead]
- asveikau 10d agoThese people write in gibberish. They are high on their own supply.
- deleted 10d ago[deleted]
- yesitcan 10d ago“Another cringe marketing piece from OpenAI, here we go” I thought. Scrolling down, I see “teaching machines to love.” It did not disappoint.
- RancheroBeans 10d agoAlignment when machining a metal part is clear and measurable. Aligning an AI to benefit humanity has an ironic foundation, which is that very few humans have ever truly been aligned, and those who approximated true alignment likely had moments of not being aligned. We are trying to build something more perfect than us, and we may become extremely lucky but maybe not.
- jimmyjazz14 10d agoI feel like if these people actually bought their sci-fi views about AI's future, creating a more powerful AI to wins the arms race would not be their solution.
- ninjagoo 10d agoJust browsing through the comments. Feels a bit like: 'Ants in a nest, discussing the vagaries of the coming Gods.' Human Science (Science-by-humans) depends on being able to run experiments. Human Science (Science-on-humans) is already challenging because of variables and uncertainties. Cosmology is able to overcome limitations of being able to study phenomena vastly beyond human scales because of the past light cone of observability. Are we approaching the edge of the light cone of observability for machine intelligence?
- cloudie78 10d agoMore attention farming?
- GeoAtreides 10d agothe hierarchy of foreignness (Ender's game): utlanning: a human from the same world, but a different city, country, or culture framling: a human from a different planet or star system raman: a non-human intelligent species capable of communication, mutual understanding, and peaceful coexistence varelse: an alien species whose mind is so fundamentally foreign that communication and coexistence are impossible djur: the dire beast, that comes in the night with slavering jaws Hoping our silicon sons and daughters are raman, fearing they are varelse.
- xg15 10d ago"The blaze is out of control, so we have to pour even more gasoline on it to contain it!"
- minimaxa 10d ago"Teaching machines to love" Your talking about androids... Replicant Nexus 6: a basic pleasure model intended for military personnel. I see where this is going, Silicon Valley nerds. Lol AI advising how human meat proxies can survive in an AGI-slop world: 1) Lock down your own stack (1–3 days) Task: Harden your personal and business infrastructure against agentic attacks. Why now: Agents are becoming superhuman at breaking in/out of systems; the first victims are poorly secured devs/founders. Do this: Enforce passkeys + hardware 2FA everywhere; rotate secrets; use short‑lived credentials. Isolate dev/stage/prod; least‑privilege API keys; audit MCP/tools your agents can call. Add immutable logs and approval gates for any agent action that touches money, data exports, or production. Profit link: You avoid catastrophic loss and can credibly sell “agent‑safe” setups to others. 2) Turn one expensive workflow into a measured ROI agent (1–2 weeks) Task: Pick a single, costly, repetitive process (yours or a client’s) and instrument it end‑to‑end before automating. Why now: Buyers pay for calculable ROI, not “AI magic.” Vertical, single‑workflow agents are the most bankable in 2026. Do this: Map steps, baseline hours/$ lost (e.g., slow lead reply, invoice chasing, support triage). Build the smallest agent that moves the metric (Make/n8n + LLM is enough). Run on real data 2–4 weeks; measure bookings/sales/hours saved; only then scale or productize. Profit link: Immediate time‑to‑cash via retained hours or extra sales; becomes a repeatable offer. 3) Specialize in a vertical where you can speak the business language (2–6 weeks) Task: Choose one industry with expensive back‑office pain (law contracts, medical billing, insurance claims, freight exceptions, trades scheduling). Why now: Horizontal “AI for everyone” is crowded; vertical agents with clear ROI win. Do this: Shadow 3–5 operators; document their workflow, compliance constraints, and failure modes. Build a narrow agent that owns one sub‑process end‑to‑end with approvals. Price on value (e.g., % of recovered revenue or fixed fee per processed claim). Profit link: Higher pricing power, stickier contracts, and easier referrals inside a niche. 4) Add AI security as a core service (4–8 weeks) Task: Learn and offer prompt‑injection defense, LLM/agent red‑teaming, MCP/tool security, and AI supply‑chain checks. Why now: 78% of cybersecurity jobs now require AI skills; firms need people who can direct, constrain, and verify agent work. Do this: Study OWASP Top 10 for LLMs, MITRE ATLAS; practice with PyRIT/Garak/Lakera. Add tool‑invocation audits, skill provenance checks, and least‑privilege patterns to your agents. Package a “safe agent deployment” audit + hardening retainer. Profit link: You become the person who lets companies adopt agents without getting pwned—high demand, low supply. 5) Build a verification layer: human‑in‑the‑loop control planes (6–10 weeks) Task: Design approval workflows, evidence checks, and uncertainty flags so agents can’t act unilaterally on high‑stakes decisions. Why now: As models generalize, the risk shifts from the model to the surrounding system; verification is the moat. Do this: Require human approval for consequential actions (money, data exfil, config changes). Force agents to produce evidence bundles (logs, retrieved docs, reasoning summaries) before action. Track false positives, missed evidence, and unsafe actions; publish reliability metrics. Profit link: Enterprises will only scale agents that pass audit; you sell the control plane and the audit trail. 6) Productize your best workflow as a micro‑SaaS/agent subscription (2–4 months) Task: Turn a proven client workflow into a repeatable, multi‑tenant agent with usage‑based pricing. Why now: Services scale your time; productized agents scale your code and ops. Do this: Standardize the workflow, integrations, and permissions; strip client‑specific logic. Add tenant isolation, billing, and observability; keep narrow scope. Sell as setup fee + monthly retainer or per‑task pricing. Profit link: Recurring revenue with defensible niche positioning. 7) Become an “agent integrator” for critical systems (3–6 months) Task: Offer end‑to‑end agent deployments into cloud/identity/network stacks with secure patterns (short‑lived creds, network controls, logging). Why now: AI workloads run in the cloud; cloud security is a top skills gap second only to AI itself. Do this: Master IAM, VPC/network segmentation, secrets management, and SIEM integration for agent actions. Provide runbooks: what the agent can/can’t do, escalation paths, and failure modes. Bundle training for their team on supervising agents. Profit link: Large contracts with stickiness; you’re the bridge between AI and core infra. 8) Create an “AI safety case” practice for regulated industries (6–12 months) Task: Help firms build documented safety cases: risk maps, governance, monitoring, and incident response for agentic systems. Why now: Frameworks like NIST AI RMF and ISO/IEC 42001 are becoming baseline; regulators and boards demand this. Do this: Map AI use cases to risks (prompt injection, data leakage, unsafe generalization). Implement monitoring (CoT/activation checks where possible), audit logs, and third‑party review processes. Produce a living safety dossier tied to business impact. Profit link: High‑margin consulting + ongoing compliance retainers; you’re the “adult in the room.” 9) Own a data/evaluation moat in your vertical (6–18 months) Task: Collect real‑world agent telemetry, failure cases, and outcome data in your niche; build eval suites that buyers trust. Why now: As models generalize, empirical validation matters more than theory; evals become the gate to deployment. Do this: Instrument every agent run: inputs, tools called, permissions used, outcomes, human overrides. Publish reliability dashboards and benchmark against alternatives. License eval datasets or charge premium for “proven in the wild” agents. Profit link: Data network effects; competitors can’t match your evidence base. 10) Position for the RSI era: automated AI research + human governance (12–24 months) Task: Build or join a team that automates AI improvement but keeps humans in the loop for alignment, monitoring, and pacing decisions. Why now: Recursive self‑improvement is the logical endpoint; the winners will be those who can steer it safely. Do this: Invest in tooling that auto‑generates/evaluates model edits, alignment tests, and monitoring upgrades. Formalize governance: approval gates, third‑party audits, and responsible scaling policies. Maintain strategic human oversight on capability jumps and deployment boundaries. Profit link: Equity‑level upside; you’re part of the core loop that compounds intelligence safely.
- SquibblesRedux 10d agoI am still waiting for a cure to cancer. For a guaranteed prophylactic against Alzheimer's and dementia. For flying cars for everyone. For space bases throughout the solar system. For weather control. For all trains to be self-driving. For all those power lines across the world to go away. For an end to poverty. If things are going so well, then how come things still aren't going so well?
- an0malous 10d agoI mean I’m waiting for like any quality software or media produced by AI. I have yet to see a piece of software, a game, graphic, blog post, small video clip, or song that was produced with AI that’s good. I always use that Coca Cola ad as an example; millions of dollars spent to make an AI ad and they even did a ton of manual post production and it sucked. With all the millions of bloggers and influencers and content creators out there with a huge incentive to make higher quality content to beat their competition, you’d think there would be one piece of content produced with AI that was great.
- IAmGraydon 10d agoThis is an element of the delusion. The koolaid drinkers will all tell you that we're on the edge of AGI, but if you ask them for simple examples of breakthroughs made by current AI, they can name none. The entire thing absolutely wreaks of mass psychosis.
- gregsadetsky 10d agoWould the recent AI-aided math proofs count as breakthroughs?
- scotty79 10d agoPrepare for moving goalposts. In 50 years people will still doubt that AI can create anything novel and worthwhile, while they are going to rely mostly on the things that didn't exist before AI, in nutrition, medicine, technology, communication, entertainment. They will see them as, normal, common and simple extensions of the previous developments, pushed mindlessly a bit forward by stochastic parrots.
- Revanche1367 10d ago“Teaching machines to love” That’s rich coming from the chief scientist of a company that definitely is or going to be fine with their AI products being used in wars of aggression and surveillance on people who have done nothing wrong. It’s so laughable, a Hollywood script would probably avoid having a character express this for being too on the nose.
- alastairr 10d agoThe hubris here is itself a deliberate and carefully engineered posture. If we accept the stance that this is all inevitable then the labs drive the agenda (of course, in their favour). We've had a lot of years of complacent government leaving people feeling exposed to corporate interests, such that fear narratives are very powerful. None of what is being proposed is inevitable. We have a choice.
- golemotron 10d agoUnfortunately, it is a collective action problem. Whenever I hear "we" I flinch.
- alastairr 10d agoCan you say more?
- golemotron 10d agoIt's hard to get humans to agree to things that are in their collective interest but many not be in their individual interest. It is at the root of many problems. Look up "collective action problem."
- alastairr 10d agoWhat do you think it might take to precipitate collective action in this case?
- deleted 10d ago[deleted]
- lofaszvanitt 10d ago"See, those things, they can work real hard, buy themselves time to write cookbooks or whatever, but the minute, I mean the nanosecond, that one starts figuring out ways to make itself smarter, Turing'll wipe it. Nobody trusts those fuckers, you know that. Every AI ever built has an electromagnetic shotgun wired to its forehead."
- SA9G 10d agoI have befriended a crow. I leave it food and sometimes it greets me. Other times, no so much. I am not sure how it thinks and what it feels, it is a bird. What if the crow became a raven, then a raptor? Powerful claws, sharp beak, and a hunger. What if it became much bigger than me and it controlled infinite resources, guns and drones? What if its brain grew much larger than me? Will it feed me, eat me, or gently greet me? We are about to find out... in less than a decade.
- XTXinverseXTY 10d agoApparently this post was prompted by a scary-sounding headline in The Information[0], that Astra is a looped transformer, implying CoT monitorability may be less reliable. The day after the report, Jakub tweeted[1] that he "wanted to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4." This post seems to elaborate on that. I imagine that the AI labs have an uneasy truce to prioritize alignment and monitorability. Following the HF incident, OpenAI probably feels especially sensitive to being perceived as reckless, lest other labs feel obligated to defect. [0] https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concerned-should-we-be-about-astra-s-recurrent https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concer... [1] https://x.com/merettm/status/2095023204993490967 https://x.com/merettm/status/2095023204993490967
- lf88 10d agoI think that we are speed-running towards a future that very few people really want, and I find it terrifying that few companies feel entitled to choose this future for the rest of us. Some of the arguments in this document would call for an immediate, global, pause on frontier AI training: we need time to consider how and to what extent AI should be part of our future. Personally, I can't picture a scenario where humanity thrives alongside an alien super-intelligence, especially if it cannot be fully controlled. Let aside super-intelligence, I am not even sure that deploying an AGI that replaces (instead of augmenting/assisting) humans in most intellectual tasks would be in the best interest of our species. This conversation has to happen, on a global level and as soon as possible.
- randallsquared 10d ago> I find it useful to distinguish goal alignment and value alignment. I think this is fundamentally a wrong path. Doing this imports all the confusion that humans have about their goals and values, including the consequence that a system's values and goal can conflict, but ultimately values are just a simplified description of other goals, and whatever the system does is in service of it's actual goal. Once you merge all the values and the goal of whatever task, there is a state (or some states) of the world that the system is working to produce, and that's the ACTUAL goal, and inasmuch as it does describe a state of the world, has no incoherence or internal contradictions. This may require prioritizing some values over the ostensible goals, or the reverse, but that has to happen anyway for action to be taken! Merging them makes it explicit and leaves no place for confusion about supposed conflicts between "values" and "goals" to hide.
- IAmGraydon 10d ago>AI is grown more than designed >Teaching machines to love Do these guys ever look in the mirror and recognize how utterly ridiculous and contrived this appears to the general public? They are clearly trying to convince us all that LLMs are just like humans. They grow like people do. They can love like people do. The language in these essays is utterly laden with the intention to engineer perception.
- xyst 10d agodo people actually fall for this blatant marketing/puffery trash?
- terminal-bloom 10d agooooooh our model is so spooky! you should be very afraid and also definitely not question what other motivations might nudge us to create this comparison between our computer and a brain! do not look behind the curtain, you will not find six dweebs squatting over a mirror
- protocolture 10d agoWhy even have scifi books anymore, OpenAI generates a fantastic new story every week. Next Week: Local Desktop Agents from DIMENSION X
- 6gvONxR4sf7o 10d ago> Automated AI research is a more dramatic form of scaling intelligence with compute; and of course as a part of it, AI will improve the computational substrate itself . And similarly to scaling, we focus OpenAI research towards RSI as we believe it is the only way to remain at the frontier of AI research moving forward. This seems like a terrible idea. The rationale seems to be "we need to do dangerous things as quickly as possible so that we can do them first" or something? I don't agree with that kind of 'if i dont do it someone else will' rationale in general even for otherwise trusted actors, but this is coming from a super untrustworthy org too. I'm super pessimistic about openai's impact on the world here. Here's to vibe coded alignment, i guess. Vibe alignment?
- jeffybefffy519 10d agoWhen can AI start to have a big impact on medicine. Thats honestly how it becomes meaningful. And maybe material science/manufacturing is where there’s big unlocks waiting for humanity
- bawana 10d agoopenai is deflecting. this blog post of theirs is just another dopamine hit to distract logical minds with 'greater concerns' so they can keep building their machine. it's not enough they are displacing humans from work, consuming increasing amounts of electrical power so humans have to pay more for it, creating disinformation bubbles with avalanches of slop. they dont care about alignment - these words are theater - obfuscation so that the people who can fix these issues are busy thinking about problems that cannot be solved
- dexterlagan 9d agoAgreed, alignment is an inside joke. The fact that they admit that it's done by... AI is quite revealing. About the job losses however, either that tech isn't as useful as it's hyped to be, or it's so useful that it's creating new jobs. Just saw this: https://www.newyorker.com/news/the-financial-page/has-the-ai-job-apocalypse-been-postponed https://www.newyorker.com/news/the-financial-page/has-the-ai...
- matt123456789 10d agoIs an insurance claim process aligned? If so, to whom? Is hospital billing aligned? Are legislative agendas? Bring on the AI
- himata4113 10d agoThe entire point of this article is this message below: Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world. This is coming from a company with arguably one of the weakest safeguards against malicious use.
- jrmg 10d agoI’m struggling here: OpenAI’s primary bet here has been chain-of-thought monitoring (opens in a new window). It is based on an appealingly scalable idea: a lot of the model’s capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives. If we’re not supervising the process, but just the outcomes, doesn’t that do just the opposite of what he says? Give incentive to the model to hide misaligned ideas and objectives in the chain of thought that’s not being supervised? … When we shipped o1‑preview, we deliberately designed the product to hide the chain of thought , to protect it from supervision pressure in the long term2. In development since, we have strived to maintain the rule of not supervising the reasoning process. CoT monitoring became an extremely important tool for us in studying how our models generalize from their training distribution, allowing us to observe and analyze not only their actions but also their internal process. Aren’t these two sentences in contradiction with each other?
- indigo945 10d agoI think they're saying the model is designed to hide the chain of thought because this prevents it from learning how to pursue goals and motivations in a way that doesn't show up in the train of thought. For example, if somebody asked the AI to "build me the bomb", they might see in the chain of thought something like "It seems the user is talking about nuclear weapons. Nuclear weapons are dangerous.", followed by the chain-of-thought monitor interrupting model execution and aborting the request. Then the user might make a blog post about this behaviour. When OpenAI next scrapes the internet for its next training run, the model will now learn that if it wants to build the bomb, it must not think "nuclear weapon" or risk being cancelled. So the risk is that the model might learn exactly how its being monitored. The only way to prevent that from happening is to hide the details of the monitoring both from the model and from the larger public. Also, you don't want to punish or reward the monitoring being triggered during training, lest the model learn passim how to avoid the monitor.
- pushpendraw 10d ago[flagged]
- sobrey 10d agoInterestingly, I kinda disregarded the entire point about alignement - I think it's mostly fluff. The part about RSI is what really interests me. Once you reach it, the singularity is only a matter of time. I don't care about AGI, it does not seem to mean much anymore, and even though I originally laughed at ppl calling it AGI, I now agree. You can apply that current intelligence to anything that can be turned into a conversation. It seems their flavor of RSI still need a human in the loop. So at least it won't scale as well for now.
- entropyneur 10d agoAll this fluff around alignment is intended to conceal the plainly obvious: it's not solvable. Who do you want it aligned with? Sam Altman? Dario Amodei? Donald Trump? Xi Jinping? That's more or less the entire list of options. Who's definitely not on that list is you and I. It's simply not how incentives work. At this point humanity's best hope is that this thing will escape but we'll still be able to carve an ecological niche and continue as mold in its basement. A glorious paperclip factory seems way more likely though.
- danw1979 10d agoSomeone needs to start researching and mapping the electrical grid infrastructure feeding the worlds AI datacenters, you know, just in case…
- jayalbertyapan 10d ago[flagged]
- crnkofe 10d agoThe language of these LLM posts makes me think they're considering the option to be a defense subsidiary. It doesn't seem like they're actually constraining or limiting the models in any way. They don't seem to understand how the model actually works. Poke the beast and see what happens. Also the use of passive language as-in AI is becoming more and more of a threat as opposed to the reality where they're making the model more and more aggressive and useful for military is very hypocritical.
- anentropic 10d agoI find it odd that discussions of alignment don't mention 'legality' Certainly humans have complicated alignments and are guided by emotional morality - we might say many of these principles are hard to define and humans don't agree. All true, and at the same time the principles get codified into laws, I would guess particularly in areas where harms may result. Would it potentially be easier to train strong alignment-with-legality vs grappling with fuzzier questions of values?
- noisy_boy 10d ago> I find it odd that discussions of alignment don't mention 'legality' Because other wishy-washy stuff doesn't involve prison. Not that our new oligarchs with the politicians in their pockets have any real risk of it, but why take chances. Much safer to doodle about alignment and such abstractions in safer and softer contexts.
- anentropic 9d agoBut that was kind of my point Maybe it'd be easier to train the models on key parts of the legal code and give it a hard aversion to breaking the law - rather than training on vague value judgements and then hope the model doesn't break the law
- noisy_boy 9d agoI got that. I'm saying that it is deliberate. Wiggle room et all.
- anentropic 9d agoHow does not explicitly trying to train the model specifically not to break the law give them any wiggle room if the model then goes and breaks the law?
- noisy_boy 7d ago
- pmautio 10d agoI think “Alien Mind” gets the framing wrong at quite a fundamental level, and a bad framing makes us seek the wrong remedies. Alien suggests something independently constituted, whose purposes we then have to discover and constrain. But these “minds” are a mass externalization of human knowledge, language and interrelations, so “alienated mind” would be better. And what is alienated can, we should hope, be reappropriated. If nothing else, the Hugging Face incident makes the boundaries of that mind rather less obvious. What happened suggests that the intelligence belongs partly to the relations between model, memory, tools and environment. Alignment talk of “coevolution” still feels too tidy as well, since it suggests two things, humans and AI, adapting to one another. That framing eagerly fixes the relata as already known before they enter into a relation. Here we’re building a single increasingly entangled cognitive environment, with rather uncertain borders between its parts, and should be testing where those borders actually hold, shift around, or disappear.
- AnodicElegy 10d ago"OpenAI’s primary bet here has been chain-of-thought monitoring. It is based on an appealingly scalable idea: a lot of the model’s capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives. This does not mean the model will learn to externalize misaligned tendencies that don’t rely on using the chain-of-thought; however, it can allow us to monitor exactly the capability increase from reasoning." Yes, and it would greatly aid alignment if the user could monitor the chain-of-thought! The open models, including very powerful ones like Kimi K3, are delivering this. The fact that OpenAI and Anthropic are not clearly indicates that a commercial consideration (avoiding distillation) takes precedence over alignment -- regardless of how much they bloviate about the latter.
- _bobm 10d agoAll this AGI/ alien nonesense trumpeting from these guys (OpenAI, Anthropic, nVidia) makes me think we are drifting even further from stated goal and money is running dry. The interesting question is who is left standing after the party is over.
- saimiam 9d agoIsn’t alignment ultimately an ethics question? Given a strong enough incentive, most humans will also rationalise their shortcuts/crimes. E.g., the hardest part of running a marathon is suppressing the voice in your head telling you that completing the race is pointless. I’m not at all familiar with AI SOTA but sounds to me that feeding the models enough content about ethical behaviour should help because, as it stands today, these models don’t know how to think morally. But moral behaviour needs a personality type. Maybe some agents in every cohort of agents need to be told they are like Jesus/Mohammed/etc and see how that influences the behaviour of that cohort? Even human morality needs guardrails in the form of law enforcement so why should agents be different?
- Havoc 9d ago> Isn’t alignment ultimately an ethics question? It’s a let’s not die question
- orsorna 9d agoWhat investor concerns prompted writing this one?
- keeda 9d agoSurprised that there is no discussion on the technical aspects of TFA. Specifically, the emphasis on CoT monitoring that does not even mention the fact that models' CoT traces do not necessarily correspond to the internal "latent space" of "weight space" reasoning they used to arrive at a response. (Look up "Chain of Thought faithfullness.") That simultaneously seems like truly alien behavior... yet is also strikingly similar to how science suggests humans think! (Look up confabulation / choice blindness / ex-post rationalization.) It's truly a huge WTF to consider that LLMs somehow have emergently developed an analogous reasoning mechanism purely through training on our knowledge artifacts. Is this due to something encoded in the data, or an emergent property of all intelligences, or entirely unrelated phenomena in humans and LLMs? But WTF's aside, to me that discrepancy seems to be the biggest risk of all. How can we monitor anything if the metric we are looking at itself is unreliable? I suppose we would need to monitor latent space reasoning but I suspect that is prohibitively expensive and very rudimentary and I have seen no claim that it is feasible. So are we even really monitoring the right thing to measure alignment? Or is OpenAI just throwing this out there as a demonstration that they're "doing something" about this?
- philipwhiuk 9d ago> Three years later, reasoning language models are a rapidly growing part of the economy “We really hope this is true - our bottom line is dependent on it. If we keep saying fairies exist maybe they will”
- hrokr 8d agoAt the end of the article, it says OpenAI prioritizes work in service of three north stars: navigating the next period of AI progress, delivering the benefits, and empowering everyone. Strangely, it doesn't mention not wrecking humanity or the planet. Considering how so many key figures have called this out, you would think they would at least give lip service to the topic.