47 ms·
GPT-6 Astra
System Card: https://deploymentsafety.openai.com/gpt-6-astra https://deploymentsafety.openai.com/gpt-6-astra
Related ongoing threads:
OpenAI's GPT-6 Astra on ARC-AGI-3 - https://news.ycombinator.com/item?id=49555691 https://news.ycombinator.com/item?id=49555691
GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index - https://news.ycombinator.com/item?id=49556147 https://news.ycombinator.com/item?id=49556147
- gregjw 13d agothe rocket completely changes design in the showcase video, am i to expect inconsistencies like that? is that AGI?
- snappr021 13d agoAI has reached the point where the limits are human.
- jesse_dot_id 13d agoPress X to doubt.
- perching_aix 13d agogpt-6-astra-ultraspeed when?
- dude250711 13d agoThey did not even release the normal one. It's just a flashy blog post.
- perching_aix 12d agoHere you go, it's live now.
- softwaredoug 13d agoI'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable) https://venturebeat.com/technology/welcome-to-the-agi-era-openai-launches-gpt-6-astra https://venturebeat.com/technology/welcome-to-the-agi-era-op...
- Bluestein 13d ago100%, some say.-
- arctic-true 13d agoThe blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three)
- _diyar 13d agoI strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games.
- CamperBob2 13d agoAt this point the only valid ARC-AGI benchmark left is to make up the next series of ARC-AGI benchmark puzzles that current models presumably can't handle.
- jaggederest 13d agoI feel like making a human-proof benchmark is pretty clear evidence that they've exceeded even the highest human capacity in most respects, for things that you can do via text generation (and to a lesser extent image generation)
- aesthesia 13d agoScoring for ARC-AGI-3 is constructed so that the median(-ish) human score is 100%, so this is not a superhuman result. However, the scaling is weird, since it's built from terms that look like (AI turns taken / median human turns) ^ 2, and it weights later levels higher than early levels. So it's not at all clear that 100% is twice as good as 50%.
- aesthesia 13d ago
- Brainspackle 13d agohuh?
- Maxforever 13d agoMhm
- bicx 13d agoDead link for me
- guilhermeasper 13d agoThat was a quick pull out.
- Pym 13d agoI saw it
- jerrygenser 13d ago> The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.
- woah 13d agoA swarm of Astra agents discovered a new and innovative way to get 100% scores on ExploitGym with almost no token spend at all
- ttul 13d ago"The gym's doors were mysteriously removed from their hinges during the night. The gym equipment was also apparently stolen. And the school's custodian was found incoherent next to a bottle of top-shelf Scotch."
- unrvl22 13d agosomeone screenshot?
- throwaway6349 13d agohttps://ibb.co/k2fB5wSc https://ibb.co/k2fB5wSc
- aisvgonline 13d ago[dead]
- tosh 13d ago$10 per million input tokens and $50 per million output tokens sol is $4 / $20
- wahnfrieden 13d ago2.5x more expensive than Sol. Can expect 2.5x more usage in Codex subscription. Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads). I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.
- deleted 13d ago[deleted]
- AaronAPU 13d agoHow is it I juggle 4-8 Codex Sol-5.6 Max agents every day and have never once run out, but you run out in one day? What are you actually doing?
- ModernMech 13d agoHow?? I'm using sol Extra High 24/7 and it eats up about 1% per hour reliably, so it lasts about 4 days for me.
- maipen 13d agoThese folks are probably using crazy plugins or crazy sub agent spams. They probably just run everything on max + fast mode which is ridiculous.
- ModernMech 13d agoThe guy said medium/high regular speed so that's why I'm very puzzled! Ultra + Fast will absolutely slurp up your whole usage quickly but I've never found it gives substantially better results so I stick to extra high.
- wahnfrieden 13d agoThey're just announcing later availability. No launch.
- paxys 13d agoEvery frontier release nowadays is "we've launched*" * for a special group of customers that you're not in. Keep waiting peasant.
- iAMkenough 13d agoTheir announcement about later availability is unavailable to me now (500 error). Great first impression.
- sscaryterry 13d ago> We will give one banked reset for every day you don't have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can. First one will land in ~ 3 hours. This is from Tibo on X.
- frozenseven 13d agoRelease the Kraken!
- tristanj 13d agoGPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot-2026-09-03-at-10.51.35-am.png https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/ https://thenewstack.io/openai-gpt6-astra-benchmarks/
- malshe 13d agoI think we need a few writing related benchmarks.
- scrlk 13d agoIs the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/ https://openai.com/index/how-two-settings-tripled-our-arc-ag...)
- kasperni 13d agoyes it is.
- woah 13d agoHaven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
- tintor 13d agoThose people haven't verified their results against the private set: https://arcprize.org/leaderboard https://arcprize.org/leaderboard
- andriy_koval 13d agoAstra also not verified using private set, but on "semi-private" set
- ealready_value 13d agoI've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.
- deleted 13d ago[deleted]
- deleted 13d ago[deleted]
- deleted 13d ago[deleted]
- deleted 13d ago[deleted]
- deleted 13d ago[deleted]
- aliljet 13d agoThe ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
- ionwake 13d agomy first suspicion is gaming - but i have no idea honestly
- aesthesia 13d agoARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.
- enraged_camel 13d agoThey used a custom harness. It's not a one-to-one comparison.
- paxys 13d agoWhy is this flagged ?
- dang 13d agoThe link was 404ing quite a bit and several previous submissions got flagged as well.
- consumer451 13d agoIt's still down for me, in the EU.
- John7878781 13d agoYou should know: AA index is only 61. Pretty surprised it’s that low.
- nsingh2 13d agoI have some doubts about AA-index. For example Opus 5 (High) is at the same index value as Fable 5 (Max), that doesn't seem right.
- jatora 13d agoMore fuel to why the AA index is fairly pointless. Gemini 3.8 flash is 59 and opus 5 is 63? grok 4.6 is 61 too? And in the past, gemini 3 pro was rated as high as opus 4.5 and the like Their AA Intelligence Index is just simply not indicative of whatever I care about, that's for sure.
- gekoxyz 13d agoThis is actually a really good thing imo. If they didn't care about benchmaxxing it means that they really know that what they have in hand is good.
- _ache_ 13d agohttps://ache.one/gpt6_now_down.png https://ache.one/gpt6_now_down.png Big claims, expensive and not release to the public yet.
- Pieczasz 13d agoOh brotha, here we go again, it's so over again, as every week nowadays
- wieiw1 13d agoI think Altman and amodei have a difficult time in understanding that you can have intelligent technology boxes but… it doesn’t change reality all that much. But thank you for spending other peoples money to give us the tech regardless!
- rvz 13d ago> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS. Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question: Did humans deploy the model, Or did the model deploy itself? It sounds like "AGI" just stands for "IPO" as it always has been. EDIT: And of course once again, the bots down-voting this post without any reason or a basic answer to my question.
- Supermancho 13d ago> Did humans deploy the model, Or did the model deploy itself? > It sounds like "AGI" just stands for "IPO" as it always has been. People don't usually respond to noise.
- rvz 13d agoHere's an idea, maybe answer the question before responding since you saw it? What do you think?
- adan1719 13d agoAI releases are like religious ceremonies. You are not allowed to disrupt them. The new system card is the gospel.
- dang 13d agoRelated: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273 https://news.ycombinator.com/item?id=49554273 How about we stick to that one for talking about the rollout, and this one for talking about the model?
- kegs_ 13d agoI guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
- PeterHolzwarth 13d agoOh please. They do closed betas - hardly makes you a "second class citizen".
- kegs_ 13d agoMythos was never released. It's really just the writing on the wall. I'm not going to give up hope, but it's pretty hard to win a race when some people get a jump on the gun.
- pixl97 13d agoBeing strongly on the AI saftey side of things what is happening was 100% predictable. At first the race wouldn't even be noticeable. Then people would see things speeding up, for example hardware getting more expensive. Then when the capabilities really got useful most people suddenly realize the race is moving 1000 mph and they are never going to catch up.
- kegs_ 13d agoWhat's currently happening is predictable, I agree. It's what's coming is the thing I'm worried most about. Either way, I'm not giving up.
- deleted 13d ago[deleted]
- atemerev 13d agoThey simply refuse my applications to slightly less restricted models without any explanations. And the current ones refuse automatically to work with me on my papers as soon as they see the word "epidemiology". I am a researcher in a Swiss university btw.
- swalsh 13d agoI was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.
- greenowl 13d ago[flagged]
- georgemcbay 13d ago> This is AGI now. Why are you spending any of your time looking at the "quality of code"? Poe's law applied to AI comments on HN just keeps becoming more relevant by the day. Judging by the poster's comment history, this is satire. But I really don't know a lot of the time anymore when I only have the specific comment as context.
- _superposition_ 13d agoI can't tell if this is sarcasm. For the same reason you don't have your model write code in assembly. But if you don't look at the code and just let the model "cook" that's basically what you'll end up with. A pile of missing abstractions.
- pennomi 13d agoIf you think any modern AI puts out stable, safe code, I have an AI-powered bridge to sell you.
- jesterson 13d agoLet them find it the hard way
- saaaaaam 13d agoPelicans please
- atemerev 13d agoDamn I hate this benchmark. SVG authoring from head without visual reference is so wrongly posed.
- maipen 13d agoVery well said. It kinda describes how unrealistic these expectations are. Vibe coders want a model that makes them rich, without having any actual specific idea. They write a very ambiguous prompt and expect to be amazed by the result. Very very unrealistic and wasteful.
- wieiw1 13d ago[dead]
- droidjj 13d agoThe complaining about the pelicans is so strange to me. It’s just a fun heuristic. If something is claimed to be AGI, I’d expect it to be able to make svgs.
- saaaaaam 13d agoWhen AGI comes it will come as a pelican and gobble up all these troublesome little fishies who gripe and whine and moan about pelicans.
- balefulboy 13d agoI'm always tired of seeing at the top of every new model release post on here. I say Simon should just keep it to Twitter.
- saaaaaam 13d ago
- dgellow 13d ago> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks Wait, what? Am I understanding that correctly? That sounds really bad
- drakythe 13d agoI am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision. Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?
- pixl97 13d agoNothing to worry about citizen, ignore the fleet of drones flying overhead.
- order-matters 13d agothe bullshit machine is learning to optimize its bullshitting techniques! <AI is a great tool for many things disclaimer, but> after working with it for a bit, how dont people realize we are training it to be an almost identical mimic to one of the worst types of employees youll ever have to work with?? the kind that always pretends to know what theyre talking about, only tells you what you want to hear, hides issues, and only does work if you would notice it didnt you cannot give this type of worker autonomy over anything.
- Laurel1234 13d ago[dead]
- amazingamazing 13d agoWe have such great AI and cannot keep a static site up?
- gchamonlive 13d agoThat's the scientific positivism fallacy exemplified in one question.
- gorgmah 13d agoYeah, apparently
- torginus 13d agoYeah, as interesting this is to nerds, I doubt this holds a candle to your typical GTA 6 or Marvel movie trailer in terms of traffic.
- pixl97 13d agoSometimes being the busiest site in the world for a few moments is difficult.
- amazingamazing 13d agoIs it though? It is static content. A good CDN could trivially chew through literally millions of QPS… with 4 nines of uptime - the really good ones say they can handle orders of magnitude more than that.
- pixl97 13d agoNotice I said for a few moments. In a few hours traffic will drop a few thousand percent back to normal with no need for a CDN. OpenAI isn't making any money telling you about Astra on their site. All the capacity they have for it is likely sold for weeks or months.
- amazingamazing 13d agoYou are making excuses for a a trillion dollar company. Wikipedia can do it.
- Cu3PO42 13d agoJust two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16033/short_gaps.pdf https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...
- GPerson 13d agoHappened to multiple people I know.
- htrp 13d agohttps://github.com/openai/PrimeGaps186 https://github.com/openai/PrimeGaps186
- dang 13d agohttps://news.ycombinator.com/item?id=49555257 https://news.ycombinator.com/item?id=49555257
- nateb2022 13d agoI'm surprised the OpenAI employee who pushed this didn't take the minute or two to format README.md to use GitHub-supported LaTeX (https://docs.github.com/en/get-started/writing-on-github/working-with-advanced-formatting/writing-mathematical-expressions https://docs.github.com/en/get-started/writing-on-github/wor...) edit: my comment was on the submission for https://github.com/openai/PrimeGaps186 https://github.com/openai/PrimeGaps186 but seems to have been moved to the main Astra submission
- warkdarrior 13d ago> OpenAI employee Why would you think it was an employee who did the push, instead of a random GPT agent?
- well_ackshually 13d ago
- Onavo 13d agoThe jump in scientific performance is non trivial.
- Readerium 13d agoSystem Card: https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra.pdf https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra....
- dang 13d agoLink added to toptext. Thanks!
- isoprophlex 13d ago> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks. Well that sounds like fun. It has become better at hiding its thoughts.
- siva7 13d agoSounds fun. As fun as their press release claiming it is the most safety aligned model ever.
- isoprophlex 13d agoIt's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating! Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.
- 6gvONxR4sf7o 13d agoSo, probably most aligned as measured by the metrics that are the least reliable on it.
- paxys 13d agoThe model said it was perfectly aligned.
- I_am_tiberius 13d agoLike all things should be.
- ReptileMan 13d agoToo bad Scott Adams died. Reality is writing jokes right in his department.
- simonjgreen 13d agohttps://youtu.be/1QNsdr-Qx_I?si=coXwStCl7clpGVC1 https://youtu.be/1QNsdr-Qx_I?si=coXwStCl7clpGVC1 Launch video
- mrinterweb 13d agoI saw the version of this video with Paul Rudd (Celery Man) https://youtu.be/a8K6QUPmv8Q?si=TWmoNhxYAPp73TKg https://youtu.be/a8K6QUPmv8Q?si=TWmoNhxYAPp73TKg
- ylsilva 13d agothat's pretty good... they are selling the product and not the model.
- laybak 13d agoI enjoyed it! for a big corporation, that's a well-executed video
- orliesaurus 13d agoI wonder if this is going to be one of those days where you'll be like: Oh yeah I remember where I was when the first version of AGI launched
- jckahn 13d agoProbably not. It's probably just gonna do tickets better and that'll be about it.
- orliesaurus 13d agoFair point - hopefully you're wrong though ;)
- ActionHank 13d agoIf this is really AGI, like really really, then this will be remembered as the day we all started on the path to building guillotines. More likely though, it's AGI because they need to hold some claim to differentiate from competitors who are beating them in price and will launch something bigger next month.
- wieiw1 13d ago[dead]
- noir_lord 13d agoThat's really the rub isn't it? We take their claims at face value then we should probably stop them training any more SOTA models til they figure out what they already built is safe or we assume theu are lying to juke the company valuation/keep the money train on the tracks and it turns they in fact were not and just took a sledgehammer to Pandora's box. We live in the strangest timeline.
- theappsecguy 13d agoUnless you're a techno-billionaire, not sure why you'd hope for our society to collapse in this way.
- oh_no 13d agoVery nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.
- tintor 13d agoARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard https://arcprize.org/leaderboard
- andriy_koval 13d agoI think it could indicate that "semi-private" dataset likely leaked to their training data.
- IshKebab 13d agoIt says "Provider Adapter" so presumably they put some manual work in to make this work.
- xpct 13d agoA dataset being as popular as their's is will contaminate the data just by people discussing it and creating their own public test sets of similar problems. Still, probably not that much compared to employees targeting it.
- minimaxir 13d agoARC has their own writeup on the result, which offers some nuance. https://arcprize.org/blog/astra https://arcprize.org/blog/astra tl;dr it's 62% when apples-to-apples to other models, which is still notable.
- ciefa 13d agoWoah, that is a crazy interesting read!
- debazel 13d agoARC's harness is just straight up broken. No serious harness removes reasoning context between each step. Not only does this significantly lower performance over all reasoning LLMs, but it also increase cost as you destroy the cache on every turn. Tossing the oldest entry when context fills up instead of using compaction is equally bad with the same issues.
- gizmodo59 13d ago99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.
- tekacs 13d agohttps://developers.openai.com/api/docs/guides/latest-model https://developers.openai.com/api/docs/guides/latest-model The docs page has a bunch more interesting details, including for example async tool calling!
- aliljet 13d agoThe ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
- polynomial 13d agoThis is absolutely benchmaxxing. Looking forward to hearing from Chollet about it!
- Legend2440 13d agoThey explain why here: https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/ https://openai.com/index/how-two-settings-tripled-our-arc-ag... TL;DR all the other models are being crippled by limitations of their harness. >First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them. >Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.
- janalsncm 13d agoOk so the correct comparison would be to fix the harness on the old model and re-compare. Now they are comparing a new model to an old crippled one.
- altcognito 13d agoThey get 66% with the old harness, which is a lot better, but obviously not 100%
- vlmutolo 12d agoThey said 5.6 Sol got something around 40% with the corrected harness.
- putlake 13d ago> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS. Not on Azure? If so, that's a big deal.
- ActionHank 13d agoThey broke up a while ago, why is this surprising?
- jiocrag 13d agoThe latest OpenAI models have still been available via Azure foundry. Exclusivity to AWS would be a marked shift.
- BoorishBears 13d agoWould be surprising if it's not on all 3 major clouds soon enough because that's been their general strategy since said break up
- bionhoward 13d agoI think the API runs on Azure
- paxys 13d agoHosted on Azure is different from provided by Azure. The former just uses Azure as an infra provider. The latter is a managed offering that is operated and billed by Microsoft using tech licensed from OpenAI.
- illnewsthat 13d agoIt's on Azure also, here is their announcement: https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-intelligence-for-work-now-available-in-microsoft-foundry/ https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-... Although I was also surprised they didn't have some type of contractual obligation to list that alongside AWS.
- x312 13d agoHmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
- estearum 13d ago> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks. Not sure how much benchmarks or CoT or evals or anything else means at this point. These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks.
- mzmzmzm 13d agoI think "able to" anthropomorphizes a little too much for a system that is "prone to" evade.
- estearum 13d agoA human who does these actions is simply "prone to" doing them. The distinction matters not one iota.
- semiquaver 13d ago“evade” itself is anthropomorphic enough! I don’t understand the complaining about this. Humans are social creatures and we understand anthropomorphic language on a deeper level than dry inapt technical language. language itself is incredibly metaphorical. Imposing rigid constraints on how people want to naturally talk about the world is just silly and will never work, no matter how much you wish it did.
- Onavo 13d agoIf they are going to do latent space reasoning, they will probably need a separate model to interpret the intermediate activations no? I know for some types of ML analysis, a separate model is already used to analyze the weights.
- KolmogorovComp 13d agoGPT-7 Zeneca
- silver_sun 13d agoGPT-8 Novo
- BoorishBears 13d agoAfter they buy AZ, making this name foreshadowing
- jonplackett 13d agoThe launch video is incredibly cringe.
- ragequittah 13d agoHumans adapt to change oddly fast. If you saw that video in 2022 you'd be 100% positive you were watching a sci-fi movie.
- prometheus1992 13d agothis is crazy! can't wait for the 27B distilled version of this.
- gekoxyz 13d agoHTTP 500 for me on the announcement page :(
- pampas 13d agoThe load bearing seam is broken for me too.
- MASNeo 13d agoDoes anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of users…
- firemelt 13d agodamn seems I should hold off my claude subs
- jonplackett 13d agoTo a vapid any goalpost moving on such a critical issue as AGI. Can we all agree in advance what kind of Pelican would convince us it’s actually AGI. For me it’s refusing to make a pelican.
- jumploops 13d ago> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains. > GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions. Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier. [0]https://x.com/MTSlive/status/2095227056040919202 https://x.com/MTSlive/status/2095227056040919202
- petilon 13d agoThis is wild: OpenAI is basically declaring that AGI is here. https://www.theverge.com/ai-artificial-intelligence/989601/openai-gpt-6-astra-release https://www.theverge.com/ai-artificial-intelligence/989601/o... “If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”
- tastyface 13d agoRenown liar Altman releasing a PR statement for his product declaring that AGI is here is really not noteworthy.
- rektomatic 13d agoRemember when the term "AGI" meant something? Pepperidge farm remembers
- 0xbadcafebee 13d agoI think the last re-re-redefinition of what OpenAI considered AGI was "It can mostly do the job of some people"
- paxys 13d agoNo, because it has never meant a specific thing that everyone agreed on.
- drop_star 13d agoDoes it pass the Turing test?
- bigfishrunning 13d agoDepending on the proctor, ELIZA passes a Turing test. The Turing test is an interesting thought experiment, but isn't really a good measure.
- HAL3000 13d agoFinally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model. Canceling my Anthropic Max sub when this ships.
- atonse 13d agoyeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall). Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.
- elAhmo 13d agoCould you share more about 5x/20x? I missed that
- m101 13d ago20x related to the 5h limit only. Weekly seems to be around 10x, although they deliberately don’t give a number. OpenAI is 20x on both limits
- beydogan 13d ago> Weekly seems to be around 10x Actually no. 5x and 20x have same weekly usage across all models. Just ask their chatbot. https://x.com/beydogan_/status/2095293596198957418 https://x.com/beydogan_/status/2095293596198957418
- chid 13d agoit's clearly wrong, think it's realistically closer to 1.7x
- BeetleB 13d agoIt's been over an hour, Simon! Where's the Pelican?
- davidwritesbugs 13d agoexactly, there's no meaningful discussion without the pelican.
- gentlewater 13d agoARC-AGI this and ARC-AGI that. All I care about is ARC-pelican.
- grim_io 13d agoIt might draw a really shitty one to hide its true power ;)
- daemonologist 13d agoModel's not available to the public (or even to Simon, I guess) yet.
- johnnyApplePRNG 13d agoI am so sour about how Codex has jerked me around these past few months (re all of the token limit shenanigans) that I don't even care. I suspect these benchmarks are heavily benchmaxxed as well. 5.6 Sol was not even close to 5 Opus and yet somehow it sidled right up to it on all of the benchmarks?? pfffft
- extr 13d ago[flagged]
- dowakin 13d agoSo cool! I'm happy 5.6 Sol user. But for Astra, OpenAI please introduce 100x Pro plan!
- theseamusjames 13d agoCan't wait for the new qwen/deepseek/kimi releases 2 weeks from now.
- Betelbuddy 13d agoASTRA Is Here (GPT-6 Released) - https://youtu.be/xdXLzFzxA9Q https://youtu.be/xdXLzFzxA9Q https://youtu.be/xdXLzFzxA9Q?t=362 https://youtu.be/xdXLzFzxA9Q?t=362
- deleted 13d ago[deleted]
- semiquaver 13d agoGuessing this one will never show up in cursor…
- smashers1114 13d agoI tried the kart racer game and instantly found that there is incredible auto-steering and you can fly by spamming spacebar.
- pandinus 13d agoLol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.
- bakies 13d agoi tried to get claude to do my taxes for last year and it refused :( now that i'm a gpt subscriber maybe I'll have luck when i'm filing next year
- Banditoz 13d agoAre your taxes complicated such that you feel the need to have an LLM do them for you?
- grim_io 13d agoSorry guys, our AI superintelligence got bored by your mundane tasks. Please provide them with frontier math problems from time to time, so they will be more aligned to do your taxes as well.
- ajdegol 13d agoI get a neuronal immune response to bureaucracy
- bakies 11d agoNo they're boring and time consuming
- alansaber 13d agoYou joke but...
- colesantiago 13d agoI'm going to call it. By 2030 all software is done and complete. But we are going to have more and new jobs.
- jdee 13d ago'all' software? aircraft flight control systems? infant heart monitors? drug manufacturing dose calibration controllers?
- colesantiago 13d agoYes. This is just another problem for the AI Labs to solve.
- NichoPaolucci 13d agoYes. I had Codex rewrite and fix all of this in one shot earlier today (using Typescript). Unfortunately, I can not show you the code, because I do not know how this "git" program works but the AI keeps talking about it.
- dude250711 13d ago> But we are going to have more and new jobs. Like "fifth-rank junior assistant spouse in a comfort harem of an ultra-rich person".
- colesantiago 12d agoYes. There will be new jobs.
- Cachecartii 13d ago[flagged]
- balefulboy 13d ago72 to 74 on DeepSWE is AGI
- tinyhouse 13d agoYou can talk to OpenAI to create a silly game and order food. What a lame way to show the model capabilities. Has Alexa commercial vibes.
- foundOpenRight 13d ago1:15.425 on Sunset Cove beat my record
- jiraiyasarutobi 13d agoIt saturated most benchmarks. WTH
- trixn86 13d agoSecret tip to win the mario cart clone: Just hold w, no steering needed.
- tripleee 13d agoyou can also fly by pressing space repeatedly
- hannofcart 13d agoWhat does 'Astra' here mean? Surely they must be referring to the Latin word. Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.
- manojlds 13d agoLuna, Terra, Sol, Astra. Though Sun is also a star, should have called it Galaxy or something.
- BrokenCogs 13d agoIt's clearly an extension of the previous naming: Luna, terra, sol
- hokumguru 13d agoQuite clearly in the same vein as Sol, Terra, Luna.
- 5555watch 13d agoCiting Tibo [0]: " - Bigger number = Better - Bigger celestial object = Better and the scale is Astra > Sol > Terra > Luna. " [0]: https://x.com/thsottiaux/status/2095600295808283073 https://x.com/thsottiaux/status/2095600295808283073
- echoangle 13d agoSol is a specific Astrum so Astra isn’t necessarily bigger than Sol
- desterothx 13d agoastrum is one star, astra is a plural of stars
- orangelimetea 13d ago[dead]
- 13d ago
- danieltk76 13d agogreat, but nobody can use it for another 100 days right?
- SneakyZero 13d ago"GPT‐6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business"
- dude250711 13d agoThen make a pretty blog post "over the coming days" instead of now and send an internal memo to that "limited set of organizations"?
- hazelnut 13d agoPlayed the racing game but that was a pretty poor experience. Would have expected more specifically if it's shared on their release page.
- Readerium 13d agoArtificial analysis blog https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...
- modeless 13d agoIt loses to Muse Spark 1.3? Does anyone really believe this index reflects reality?
- fancyfredbot 13d agoI'm surprised you feel like you know muse spark 1.3 performance well enough to question the validity of the index based on this benchmark result. Muse spark 1.3 was only released yesterday.
- udbhavs 13d agoMinor nitpick, but the handling in the Kart Racer game is terrible. It feels more like nudging than turning.
- udbhavs 13d agoI remember when GPT-4 came out and the perceived performance upgrade seemed underwhelming for a major release compared to 3.5, especially how there were graphics going around showing the parameter size dwarfing the last model before it came out. It looked like we were past the perceivable differences from release to release that were immediately identifiable. Now the jump between 5 to 5.5 and 5.6 alone has changed how a lot of people approach AI, including me. Interested to see where it goes with 6.
- redox99 13d agoThe jump from 3.5 to 4 felt gigantic to me back then. GPT 5.0 did feel underwhelming though.
- udbhavs 13d agoOops, I might have been misremembering then. Maybe I meant 4 to 5
- redox99 13d agoYeah 5 was very underwhelming.
- abixb 13d agoThe couldn't even get the bar chart right, iirc. [0] [0] https://www.reddit.com/r/singularity/comments/1mk8tm8/gpt5_cant_spot_the_problem_with_its_misleading/ https://www.reddit.com/r/singularity/comments/1mk8tm8/gpt5_c...
- l3x4ur1n 13d agoNo, no, I also remember 3.5 -> 4 and the general sentiment was that it was underwhelming. I guess we all expected absolute miracles from the models. I think our expectations sobered up a little since then.
- 13d ago
- dang 13d agoArgh! I hit a wrong keyboard shortcut and moved the entire thread. Please stand by... it will all come back shortly
- the_duke 13d ago500 upvotes with 2 comments would have been a new record. ;)
- dang 13d agohttps://news.ycombinator.com/item?id=49555647 https://news.ycombinator.com/item?id=49555647 was the one I meant to move, but I did the inverse and moved everything else. All fixed now.
- layer8 13d agoLuckily there’s a standard keyboard shortcut for “undo” as well. ;)
- dang 13d agoNot in the world of HN admins unfortunately
- abixb 13d agoI want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs. If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)? As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.
- catigula 13d agoThey’re really, really scared because of the Mythos controversy. Skynet will be under hyped.
- driverdan 13d ago> If this is truly AGI (subject to one's definition of AGI still) Scoring well in a benchmark that's called AGI does not make an LLM AGI.
- ShinyLeftPad 13d agotalking about self proclaimed, it's about as much AGI as openAI is open.
- dmitrygr 13d agoHey now! Keep your reason out of their marketin^H^H lies!
- jhonof 13d agoBut they declared it...
- cyanydeez 13d ago
- damsta 13d agoWhy release it now instead waiting those few days until it is available for everybody?
- balefulboy 13d agoBecause they saw how much hype Glasswing was getting in April
- damsta 13d agoFrom what I've seen it only made people mad, not hyped, so the person that thought it was a good idea miscalculated a bit. Now waiting for Anthropic's post about their usage promo or something similar to redirect people to them.
- deleted 13d ago[deleted]
- herpdyderp 13d agoThe FrontierCode 1.1 Extended benchmark is the only benchmark that aligns with my actual LLM experiences and Astra isn't significantly better or cheaper. All this celebration, and yet it's only on-par with an already existing model? I don't get it.
- alex7o 13d agoI hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.
- ianm218 13d agoI wonder how they were able to get it to get 99.9% on ARC-AGI-3. That seems truly insane.
- GodelNumbering 13d agoI decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.
- killerdog10 13d ago[dead]
- k9294 13d ago[flagged]
- E-Reverance 13d agoAt this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete
- ChrisGammell 13d agoAll the people here are focused on security and costs while I'm like "hey kicad on the announcement page!" Every clanker is an autorouter these days, eh.
- jpatten 13d agoYeah I was really excited to see the KiCAD example. Curious how useful it is in practice.
- ChrisGammell 13d agoI can't help thinking "doesn't matter much unless it's perfect" because if someone is using this to build a board (cool) but then it's not flawless, troubleshooting will be quite tough as a novice. Like, say, when I start digging into the web code generated by a coding agent. I am most excited about it bringing down the barrier so more people join in on hardware fun, so hopefully it will unlock folks that stayed away in the past.
- intenex 13d agoThe ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher. Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder. I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here. For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
- abixb 13d agoIt's "harnessmaxxing" all the way down. AI benchmark scene is exhibit A for Goodhart's law.
- intrasight 13d agoIt has to pass the Turing test
- drusepth 13d agoLLMs started meaningfully passing the Turing test a year or two ago, around GPT-4.5. Is there another version or bar for "passing" you're looking for? [0] https://arxiv.org/pdf/2503.23674 https://arxiv.org/pdf/2503.23674
- the_duke 13d agoHuge gains on some benchmarks, but for coding it sits barely above Fable It will be interesting to see how it performs in the real world ...
- alex7o 13d agoMaybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad
- bdangubic 13d agoAnthropic should prep 5.2 and 5.3 at the same time, release 5.2, wait for Google to release their shit in a day or two later than then release 5.3 just to fuck with them :)
- Planktonne 13d agoI'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that. This is farcical.
- baq 13d agoIt’s using the computer. I don’t think it’s a farce.
- balefulboy 13d agoDon't forget the 3D demos. My favorite is in the house tour where the sink and stovetop(?) are obviously very misaligned from the counters
- emp17344 13d agoCan’t wait for 3 months from now when they declare they actually really do have AGI this time, please guys just believe us
- mchusma 13d agoThe games on mobile safari were broken. Buttons all misaligned in the kart racer one, the spaceship thing froze for a while, then kind of loaded but maybe not? Wasn't super compelling. I'm not trying to be too negative on it, it could be the best model right now, but it clearly isn't some agi god because things like that should have been caught (also should have been caught by human reviewers).
- mvkel 13d agoThe ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games. If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath. Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.
- sashank_1509 13d agoBenchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Let’s see
- Robdel12 13d agoI don’t care about benchmarks, no way we can distill the breadth of software engineering into a number. So, folks that have actually used this already, what’s it actually like?
- GodelNumbering 13d agoThe most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks: Terminal-Bench 4.0: High (57.9%), Max (56.7%) DeepSWE: High (73.3%), Max (71.5%) It _loses_ 1-2% performance going to High from Max
- XCSme 13d agoThat's quite common with many models, after "High" reasoning, over-thinking starts occurring and the model skips over the right solution by convincing itself otherwise.
- GodelNumbering 13d ago> That's quite common with many models Such as? I can't think of any. Diminishing returns, yes. Occasionally flat, yes. Downright regression, no.
- XCSme 13d agoIn my own tests on aibenchy.com, where questions are quite simple, higher reasoning efforts consistently used to do worse than medium for most models. The reasoning effort should match the complexity of the task against the model's capability. Hard task with low reasoning = bad Easy task with very high reasoning = bad
- minatoaqua1 13d agogrok 4.6
- JacobAsmuth 13d agoLOL
- desterothx 13d agoiirc, some of the original fable bemchmarks showed this. Definitely saw it in other frontier releases though
- kingjimmy 13d agobro wtf is this website and why does it take 500mb of memory... smh.
- BrokenCogs 13d agoGPT-6 is so good that all pelicans born after today will look exactly the one generated by simonw
- dalemhurley 13d agoOpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive). Codex is slightly better than Claude Code. Good on Sam Altman getting back to basics and turning OpenAI around.
- kroaton 13d agoI think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.
- tonyhart7 13d agothey don't have moat in hardware either Chinese counterpart like CXMT and Huawei is begin producing their own chip You cant block an entire nation level effort with tariff
- astrobiased 13d agoI think the moat that China has is energy costs. It's taking learnings from the Bitter Lesson. If you role up scale and compute to the next level, it's energy resources. China has it and sharing open weight models is an effective means of removing the tech moat. This idea has been floating around for a bit now (I'm not taking credit for it).
- spartacusnacho 13d agoThey also benefit from the commodification of software/knowledge work since they own manufacturing
- rgbrenner 13d agoIt's not energy costs. The US produces about 70% more electricity per capita. Chinese households do pay less than half what US households pay for electricity, but that's because the NDRC sets prices below costs for households. They make it up by charging industry more, and the industrial electricity prices in China are roughly 34% higher than in the US.
- dearing 13d agono results
- jdprgm 13d agoIs anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs. It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time. I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.
- Pikamander2 13d agoThat's how cutting edge tech has always worked. Imagine buying a shiny new PC in the 90s only to see it become practically obsolete within a year.
- phainopepla2 13d agoThat's not the experience of owning a PC I remember from the 90s at all.
- bananaflag 13d agoIt is how I remember it.
- embedding-shape 13d agoI remember CPUs moving relatively fast back then, some years in the 90s had relatively big jumps, much bigger than we saw today. The classic graph, : https://i.extremetech.com/imagery/content-types/03zc6ghfKswe41smvPXi8Zh/images-6.jpg https://i.extremetech.com/imagery/content-types/03zc6ghfKswe...
- computomatic 13d agoIt was both. 90% of people never needed nor purchased a bleeding-edge computer. The mid-tier was "good enough" and far closer to affordable for most people; though, that bar also moved upward every year. If you bought a mid-tier computer that was good enough for what you needed, then you probably didn't shop/compare for the next few years and didn't notice. But if you shelled out $7-10k for a top-of-the-line system and paid attention to progress, you'd easily see that become the mid-tier $1000 option within two years or less. This is how it was in the 90's PC boom, at least. Likely the same for the decades before, not sure how it went in the 2000's.
- brcmthrowaway 13d agoAnthropic in tears today.
- nullbio 13d agoAnthropic don't care, they don't want their products to be used by general audiences in any serious manner. Their interest is in selling to megacorps and using the models for themselves internally to swallow industry, and drumming up AI fear to regulatory capture to shut down the businesses that do want to make AI accessible to the people.
- zsoltkacsandi 13d agoYeah, but if OpenAI provides a cheaper model with the same capabilities the megacorps will buy from OpenAI. If you want to target enterprises you need to have some competitive advantage, it's not enough that "I wanna target them".
- dopa42365 13d agolike eh 2 days ago it was the usual "too powerful to release" https://www.reuters.com/business/openai-says-upcoming-model-is-so-capable-it-requires-stronger-guardrails-2026-09-01/ https://www.reuters.com/business/openai-says-upcoming-model-... > "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," said Amelia Glaese, an OpenAI vice president overseeing its safety work. > The company plans to make Astra available "soon" to a limited group, but declined to provide specifics. Glaese said the extra security measures may "sometimes slow, pause, or stop legitimate work," and that OpenAI would work to minimize those disruptions. what a bag of horseshit
- bbor 13d agoTo be, or not to be, that is the question: Whether 'tis nobler in the mind to suffer The slings and arrows of outrageous fortune, Or to take arms against a sea of troubles And by opposing end them. To die—to sleep, No more; and by a sleep to say we end The heart-ache and the thousand natural shocks That flesh is heir to: 'tis a consummation Devoutly to be wish'd. ... And thus the native hue of resolution Is sicklied o'er with the pale cast of thought, And enterprises of great pith and moment With this regard their currents turn awry And lose the name of action.
- rcr-anti 13d agoThe benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index. Could be the case it's just not showing up in benchmarks, for a good while Anthropic persistently trailed in benchmarks but had people swearing by it.
- sharmajai 13d agoReally feels like AGIPO is here.
- sbinnee 13d agoI dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.
- maherbeg 13d agoOk, but can I bring GPT-6 in as an agent as a software engineer, tell it to talk to these people and have it start solving engineering problems and continue on for a full year career wise? maybe call it EngEmployeeBench
- Centigonal 13d agoThe moment this is possible, you will lose your job.
- maherbeg 13d agoI imagine the first year we'll be at the Junior eng level, and then after a while make our way up to Staff Engineer. Then we'll have a bunch of staff engineers arguing and protecting their domains and then we'll need a new benchmark.
- JacobAsmuth 13d ago[flagged]
- retired 13d agoDoes GPT-6 pass the Turing test? Or are the responses still very obviously AI?
- deleted 13d ago[deleted]
- Chinjut 13d agoWhat is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
- Yajirobe 13d agoAGI-level model is perpetually 18 months away. Your job will be fine.
- worldsavior 13d agoSo what will happen in 18 months?
- Kkoala 13d agoAGI-level model will be 18 months away in 18 months. (Well, at least according to the commenter above you)
- weakfish 13d agoThe AGI goal posts move
- Keyframe 13d agoWe fight the war, of course.
- mawadev 13d agoI wish something would finally happen, because I'm really tired of pretending I care about this stuff at work at this point
- tetec1 13d agoAccelerationism can be a sort of doomerism, when you think about it.
- rbreve 13d agoWhere is the cure for cancer?
- XCSme 13d agoWe need CancerBench
- azan_ 13d agoDidn’t Moderna use AI for development of their melanoma vaccine (which has recently shown spectacular results)?
- dakolli 13d agoLmao, come on dude, anyone whos used these tools for research knows it makes them lazier, less interested and dumber. You really want disease researchers become sloppers too?
- mrdependable 13d agoThey did use “AI” but not an LLM. If they had, I am sure OpenAI and friends would be shouting it from the rooftops.
- ncr100 13d agoWhat if it's only able to make a worse form of cancer? :-/ Would cancerbench be unethical?
- deleted 12d ago[deleted]
- deleted 12d ago[deleted]
- alpineman 13d agoThat Astra ‘city scene’ is about as creative as Doha in real life (not very)
- codruterdei 13d agoI was actually wondering when they will release the new Opel Astra model. Good and reliable car, wondering if we can say the same thing about this model and its impact on the market.
- XCSme 13d agoIt's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
- sashank_1509 13d agoAgreed
- gavinray 13d ago> Like, what's the point, if the next AI can do it in 5 seconds? Live a life doing whatever makes you happy. Post-work society is an inevitability if we don't destroy our planet.
- lackoftactics 13d agoGary Economics wants to have a word with you. It would be fun to get to post-work society, but hard to imagine atm. TPTB won't let it happen
- XCSme 13d agoBut how this transition will even happen? Soon we will have some machines that can replace 50% of jobs, and this will happen basically overnight...
- neta1337 13d agoIt won't happen though
- unclad5968 13d agoThe steam engine replaced a lot of jobs. Tractors replaced a lot of jobs. Calculators replaced a lot of jobs. Computer used to be a job description before it was a personal device. There will be new jobs.
- 13d ago
- serjester 13d agoExciting but it’s priced at 2.5X Sol - we haven’t seen pricing this high since GPT 4.5. We will see if the real world use cases outweigh the sticker shock.
- brindidrip 13d agoCool, I don't really care anymore.
- astrobiased 13d agoI can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area. It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale. The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new? With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.
- vessenes 13d agoPretty efficiently, apparently, since it saturated ARC-AGI-3 in half of the predicted time, and according to the Chollet blog post on the fly created dense DSLs to describe and analyze individual games.
- ex-aws-dude 13d agoThey can do new tasks with in-context learning but its obviously limited by context window
- deleted 13d ago[deleted]
- falcor84 13d agoBut what does it mean in practice? Obviously we humans also have a limited cognitive capacity. Let me offer a thought experiment: Let's say that tomorrow we discover Atlantis, with a treasure trove of books about their culture and science, written in a dialect of ancient Greek that we know how to start to analyze, but no one can read fluently. And let's say that you are a billionaire really curious about their culture and want to converse with an "Atlantean expert" as soon as possible. Would you invest your money in a "we-hate-ai-slop(tm)" group of researchers who would abhor AI and instead delegate the books to a massive number of human grad students? Or in a small group of researchers who are willing to use AI agents to go over these? Or maybe just open a chat session with GPT-6 yourself immediately? What would most effectively assuage your curiosity?
- tonyhart7 13d agoits insane how they are dropping this after fable
- KronisLV 13d agoIt's surprising how on High reasoning it actually isn't that much more expensive than Sol, in addition to being better.
- alpineman 13d ago“allowing non-technical people to create and play custom games that go beyond rudimentary elements” Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart
- arkensaw 13d agoIt's worse than that, someone else generated it using and then put it on a static page. We just have to take their word for it that GPT6 can do this. It probably can. It's not really an impressive test anymore. Claude Fable can do it. Opus can do it. I've been making one-shotted games with models for a while now, to test out their capabilities, and they all pretty much come out like this - generic bland and basic, using three.js with rudimentary controls and zero gameplay other than collecting points. Here's a one-shotted submarine game I made with Fable a few weeks back - https://roryok.com/games/deepdive3d.html https://roryok.com/games/deepdive3d.html. One prompt, and I think it's deeper than this (if you'll pardon the pun)
- altcognito 13d agoI love your game. It's wonderful and exactly the sort of thing that would showcase something interesting as opposed to just copying what's already out there. It is something I could share with my kids, and exactly the right note of fun and exploratory in a unique and even natural way. It could be extended and played with. I usually roll my eyes when I see a comment like this because rarely do they make the points they claim to make, but I see what you're getting at. They just chose to clone someone elses work and do it in a boring way. I like OpenAI's models a lot, but they should do better. edit - just a sidenote that I hadn't looked at the games, I just took the comment about "super-mario cart" at face value. I stand by my points 110% (even moreso perhaps), what they're showing is more polished than I expected, I assume they spent a lot of tokens on it. It is a legit shame they couldn't have spent time thinking of a better idea to illustrate something just as polished, but more interesting.
- mf2hd 13d agoYou sound like an AI.
- Rover222 13d agoOverall I have to say it feels like a very incredible comeback from OpenAI, after focusing on Sora and stuff like that and losing so much ground to Anthropic in enterprise revenue. I hop models at will, and have done 90% of my work on OpenAI models since sol came out.
- pcurve 13d agohttps://www.youtube.com/watch?v=1QNsdr-Qx_I https://www.youtube.com/watch?v=1QNsdr-Qx_I
- manlymuppet 13d agoI have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required? All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself. (Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)
- shostack 13d agoTrue, but they're still friction to be reduced here. What I desperately want is for 1password or stripe or even Google who already has much of my data, to o come up with a secure solution for online purchases with agentic credit cards where I can effectively get a phone prompt to authorize a purchase while the agent can fully own the checkout flow. I have seen various things coming on the market for this, but none of them appear aimed at a consumer audience. And I am a firm believer at this point in keeping my payment authorization and history and credentials harness agnostic.
- echoangle 13d agoThat’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?
- degamad 13d agoBecause some people do. Corporate travel is an example. In many organisations, you tell someone in the travel department "I need to be in Tokyo for this conference from Tuesday to Sunday, and charge it to this cost code", and they figure out flights, accommodation, etc for you, with minimal input from you.
- bmenrigh 13d ago> GPT‑6 Astra brings together years of research and big bets across pre-training Do we know if they’ve finally completed another pre-training run, or is this building off the same pre-training base they’ve been using since the GPT-4 days?
- czk 13d agothe last model to use the gpt-4o base model was gpt 5.1, since then its been new pre-trains but this is a new one entirely to itself
- wiseowise 13d agoHey Astra, can you fix openai website so that static website doesn't lag on M3 Pro when I scroll?
- steve-atx-7600 13d agolooks like it worked :)
- itissid 13d agoAll of this will be besides the point. Here is what's gonna happen. The frontier labs are just gonna keep building powerful models. AGI or not, open models in a year will be as powerful as Fable and Astra — probably by using em — and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of damage. Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.
- cromka 13d agoSurprised they haven't reset Codex usage on this occasion.
- damsta 13d agoI'd say it's because it's not available yet on subs
- throwaway13337 13d agoThat hero video is interesting. A projector and speech. Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now. The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for myself to do this between different llms. The wall projector is a cool idea because I think it frees the user from staring at a lonely little rectangle while sitting in their fixed office chair. If done right, this could bring us closer to the dream of more natural, social computing. Bret Victor's (failed?) project Dynamicland involving a projector on a desk had this goal. I hear he's not much a fan of LLMs. On the one hand, I can see why. But I think, used correctly, it might be the sort of thing that unlocks his dream and, really, my dream, too. A here's a presentation of Bret's talk on it: https://www.youtube.com/watch?v=7wa3nm0qcfM https://www.youtube.com/watch?v=7wa3nm0qcfM Slight tangent: using speech to text to ramble about your rough design for like 20 minutes to an llm produces surprisingly good results over short prompts even when you contradict yourself. They're so good at picking up on what you're orbiting.
- low_tech_punk 13d agowhich raises the question, is the model in the demo actually gpt-6? or it is gpt realtime 2.1? It's unclear how gpt-6 can interact at the realtime level and if so, how can developer get access to it?
- starik36 13d agoI ran into the same problem as you, so I ended up by coding a local app that is very similar to Wispr Flow, but uses the small english Whisper model on my low-end Windows laptop. It is still a quite fast. In fact, I just typed this in using this app.
- makapuf 13d agogiven the fact that we've moved in my office from 3-people offices to open-plan office to flex desk now I'm not exactly sure I would want my coworkers to speak all day to their computers and gesturing / walking in front of a projector (provided there will still be coworkers with IA)
- Obluness 13d agoThat seems promising ?
- carlos-menezes 13d agoThe Kart Racer game is easily breakable if you spam the spacebar. AGI!
- chris_engel 13d ago[dead]
- HardCodedBias 13d agoEven though the model is clearly wonderful the launch video is an abomination. That gives me hope that there is still areas to improve. What a bad launch video. Hilarious. What a powerful model.
- holoduke 13d agoThis absurd marketing will hurt openai. Who is buying this absurdness. I mean it's a good model, but come on. It's not agi. Not even 1% yet.
- HSO 13d agopeople are going to be so surprised how fast the ai energy leaves the room again once the cash transfers are completed (the `ipos` whatever bla) the coffee will be as cold, flat and stale as the bitcoin, metaverse, and what was the thing before that thing agi deus ex machina descending from the icloud ftw!!! pathetic :)))
- jumploops 13d agoI think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode. Hopefully this model has the right balance, or at least better?
- weird-eye-issue 13d agoFable does a great job from my terrible prompts when coding
- dannyw 13d agoAnecdotal experiences from my external early testing of Astra: if you love Sol (like I do) and wished it was smarter at everything, but especially better at high-level tasks and discussions; I think you'll LOVE Astra. Astra retains the best parts and overall 'grounded collaborator and executor' of Sol in my testing (harness: codex CLI); while being a significant leap in capabilities & higher-level thinking. When you prompt it like a technical collaborator, I've found Astra to be extremely consistent in staying as a collaborator, and not being over-eager, over-achieving or doing work that you haven't asked it to. When you ask it to one-shot something, or explicitly ask it to make decisions, it will of course make its own assumptions and decisions, and generally very well. Astra is also excellent at instruction following and respecting the guidance and steers boundaries you have. ^OpenAI does not review, limit, or tell me what to say; opinions are my own experiences.
- NorthSouthNorth 13d agoThis is quite exciting. Sol for me has been the absolute best model yet. I find myself using it 95% of the time even though I have access to Fable. Is the speed the same as 5.6 sol?
- 13d ago
- alberth 13d agoSeems like voice is a big part of this release. I don't think it's a coincidence they launched this the week before iOS 27 launches (with new Siri).
- BadBrands 13d agoSo they’re copying Gemini with the whole star motif? I guess it makes sense they are unoriginal. like Zuck, @sama never invented anything or innovated at all - just took other people’s ideas
- steve1977 13d ago"distilling"... ;)
- HDBaseT 13d ago"Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.12" Sounds about right. Alignment is important, but also being able to do mundane tasks is important too.
- rjtc 13d agoI am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks: https://artificialanalysis.ai/models https://artificialanalysis.ai/models Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?
- aniviacat 13d agoThis benchmark gives the same intelligence score for GPT-6 Astra (max), GPT-5.6 Sol (max), and Grok 4.6 (high)? That seems very wrong to me, unless I'm misinterpreting the visualizations.
- Bjorkbat 13d agoThe most straightforward answer is that despite efforts to design a benchmark that, in theory, is supposed to measure generalizable intelligence, performance on ARC-AGI-3 can't be reliably correlated to performance anywhere else. I kind of lost faith in it after o1 or o3, I can't remember which, absolutely crushed ARC-AGI-1. And, you know, maybe also some funny business. I think it's good to be a little suspicious of a model that happens to shoot upwards in performance on a specific benchmark while also kind of keeping up with the pack on a bunch of other benchmarks.
- dudeinhawaii 13d agoI don't think this is quite true. We have other examples. Fable is without question the larger and more thoughtful/intelligent model. It also gets out performed by Opus on many/most benchmarks. So we can say that while Fable is more intelligent, Opus is more capable. I'd still opt for Fable in nearly every case if tokens were free. So it can be true that the "smarter" model is perhaps not the smartest in every single niche dimension that its cousins have been fine-tuned for (yet!).
- jrflo 13d agoIdk, they’re trying to sell a $500/mo/seat service to tell you what model is best. I think it’s in their interest to keep it confusing and opaque. Not exactly independent.
- 13d ago
- maxall4 13d agoThe official ARC-AGI 3 score—-without OpenAI’s custom harness—-can be found here: https://arcprize.org/leaderboard https://arcprize.org/leaderboard. Astra scores 62.7% at max reasoning for the low-low price of 26,000 dollars.
- drivebyhooting 13d agoFor people skeptical of AGI. Consider the following: 15 years ago if you were the sole proprietor of these models, would you be able to hold a dozen remote junior engineer jobs? Maybe even more? These models could certainly pass all interviews with flying colors and even survive independently in a company role. I think sole ownership of AI 15 years ago could be worth north of $10 million per year. Just as rank-and-file employees.
- BoorishBears 13d agoCool. Being sole proprietor of AGI 15 years ago should result in monuments and religions devoted to you today. Cancer should be cured, and we should be a post-quantum interstellar fusion-powered civilization. I wish the AGI crowd would finally shut up now that it's clear no one is even trying for AGI (OpenAI revised that to "$100B in profit") What we're getting is incredible, where we're headed is incredible, but some people have such a fetish for futuretelling they can't just shut up and enjoy the ride.
- deleted 13d ago[deleted]
- drivebyhooting 13d agoI didn’t realize AGI required solving problems modern human civilization hasn’t solved yet. Well by that metric humans aren’t intelligent either! And how many people could’ve actually invented calculus, relativity, quantum mechanics? Are those who didn’t and couldn’t also not intelligent?
- BoorishBears 13d agoThis isn't the gotcha that you think it is: AGI's original definition is being able to do any task that requires human intelligence. The unlock isn't AGI smart enough to invent quantum mechanics, it's suddenly being able scale human intelligence using grains of sand instead of decades of food and energy and nuturing.
- 13d ago
- low_tech_punk 13d agoWhat's the point of enlarging the screen into a room? In the 1979 Put That There demo, the user at least used his hand to point things. The model is impressive but the demo felt like a step back. Original demo (fun ending) https://www.youtube.com/watch?v=RyBEUyEtxQo https://www.youtube.com/watch?v=RyBEUyEtxQo
- paxys 13d agoBecause it looks good in a marketing video
- camillomiller 13d agoI might be jaded, but these examples look silly, stereotyped, and absolutely how of touch with the nuances and the complexities of what real people would actually want/need to do in this specific situations.
- claiir 13d ago> Astra improved a term in a bound on these gaps that had remained unchanged for more than 80 years. We’re sharing the proofs and abridged chain of thought and verification materials for both results. Looks like they listened to Terry Tao’s request for CoT in his talk on LLM use in mathematics?
- MalleableMind 13d ago[dead]
- dakolli 13d agoWhy is everyone so excited to be replaced and become reliant on some billionaire's thinking machine? These are just going to be used to turn you into a rather dumb reliant paypig.
- nullbio 13d agoThat's a policy and distribution problem, not an AI problem. Anthropic is doing their best to make it a reality though. They'd love nothing more than to shut down distribution and become the sole gatekeeper of everything AI.
- Dlemlo 13d agoIts still human made frontier progress. And for sure it has a tremendes amount of implications, but its not the fault of the technology (we found, not invented). And i'm only living once, my main motivation is not to just live day in day out the same stuff, i'm quite happy to see progress. Am i worried about the future of our planet? For sure.
- ChaseRensberger 13d agowhen do i get to go to the moon
- OpenGayEye 13d ago[flagged]
- yodsanklai 13d agoIt seems like every few days there's a new model with hundreds of comments on HN. I find it hard to keep track of the progress. Is there a TL;DR on what benchmarks to look at to understand what is going on?
- karim79 13d agoThere will probably never be AGI. This shit is just snake oil. Nor do we have a proper definition of what AGI actually is or what it's supposed to do. There will be a small handful of billionaires claiming that AGI is just around the corner ad infinitum just to serve themselves at this moment in time, and capitalise from the hype. There is no "AGI" endgame. This is shitty ass hypercapitalism in action and nothing more. I'll repeat: snake oil.
- efavdb 13d agoSeems like only yesterday that gpt 5 was supposed to mark our downfall
- Dlemlo 13d agoIts hard to understand how exactly these AI people mean it. When i say "AI is changning the world" its more like "I can already see how this technology will continue to become better and better and has more impact every single day. It already affects people and it will have fundamentally changed A LOT in 3-15 years" But lets be very realistic and clear: My computer systems at home were exploitable a lot more often this year than any year before JUST because of GPT or Claude.
- showurwerk 13d agoPatiently waiting for the Claude usage reset in response.
- mentalgear 13d agoSo OpenAI’s stance on safety is now basically that Blues Brothers meme: two guys in dark sunglasses, driving at night in a car with a broken windshield, pedal to the metal, asking, "What could go wrong ?"
- ntlm1686 13d agoMaybe they know that Claude 6 will have similar performance every soon.
- vinhnx 13d agoGPT-6 Astra scores 74.1% at DeepSWE v1.1 bench. Huge!
- mrcwinn 13d agoI know in order to conform to HN community rules I'm supposed to be negative and dunk on this, but I have to say, I am so excited to use Astra!
- datadrivenangel 13d agoData Science Tasks (Internal) doesn't include time for Astra... same for Database Migration Tasks (Internal)... But does for gpt 5.6 sol.... which is funny. Same for HealthBench Professional and a few others. Clearly either OpenAI is very sloppy or GPT-6 Astra is also sloppy.
- zhoge 13d agoWhat's the energy efficiency of Astra? Does it roughly correlate with the token efficiency?
- akhil_findincal 13d ago[flagged]
- quyleanh 13d ago> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively. So the closed source application should open its source in near future? [15] https://arxiv.org/abs/2608.11469v1 https://arxiv.org/abs/2608.11469v1
- tintor 13d agoNot if OpenAI considers reverse engineering an offensive cybersecurity skill.
- DaSHacka 13d agoSurprisingly, I've had really good luck with reverse engineering on frontier models (without being part of CVP or similar). It's more the exploit development/PoC that it locks up on, which, as I use `pi`, I just switch the model to Kimi K3 to finish up making the PoC. Ironically, due to the stringent guardrails on American models that exist to avoid giving adversaries a leg-up in cybersecurity, I end up feeding dozens of 0days straight to the CCP lol
- glub 12d agoI do RE mostly, and they do lock up for me. But what works for me is: warm up context with non-RE things with American models > switch to GLM 5.3 with actual request, let it fill context with some RE work > switch to American > switch back to GLM 5.3 and keep switching back and forth if anything gets flagged. Solves the guardrail problem - check Feeds China data to US - check Feeds US data to China - check But I think OpenAI has caught up with this, and it's why they're pushing so many things server-side and encrypt it (subagent messages, compaction, new context history thing).
- fy20 13d agoI was listening to the Lex Fridman / DHH podcast last night [0], and DHH was saying that this is a new era for open source software. I'd agree, and also extend it open hardware. Recently I've seen quite a few posts from people using AI to reverse engineer the Bluetooth protocol or such on devices that need a proprietary app. The same thing for firmware is surely coming, which is great as you can get lots of fun hardware from China, but it often has shitty firmware. Once that becomes the norm there's no reason not to make it open in the first place. [0] - https://open.spotify.com/episode/45lhw2Adbrsw0xSCOgIeg3?si=rUWT3-5oQuuR3TWjlWpgoQ&utm_source=copy-link https://open.spotify.com/episode/45lhw2Adbrsw0xSCOgIeg3?si=r...
- ShoeMascot 13d agoThrough various comments here there is a clear confusion on what AGI means. Can someone point to a definite clarification? Is it: A) “Resting” intelligence that cycles 24/7 toward some goal, and any potential emergent ambient goals? (kinda what I think) B) Consciousness itself? The ability to feel and experience alongside the thinking - even if it is toward the end of completing some task? C) “The Singularity” (whatever that is?) so that AI can now do ____? Someone please clarify for me!
- gordonhart 13d agoAutonomously Generating Income
- grandarmory 13d agoAnonymous Grifters International
- PeakHNUser 13d ago[flagged]
- Telanir 13d agoAGI to me means capable of absorbing new information on the fly and self-evolution. As long as it is a pre-trained model without live post-training capability, it's not AGI to me. It is extremely impressive, but it doesn't pick up skills in a lasting manner, and requires a beefy harness for it to perform.
- jesse_dot_id 13d agoAGI to me means intuition and I don't think that's ever going to happen with a LLM.
- altcognito 13d agoHow do you define intuition?
- jesse_dot_id 13d agoInstinct. I think it drives UX. When I am designing an interface, I'm not just parroting something I've seen before. I am calling upon the breadth of my human existence, my professional experience, my empathy, to think about how another human will be using the applications that I create. It paints everything I touch. LLM will never be able to do that. It can clone things that humans have created in the past but it can't create novel solutions that a human will enjoy without a human at the wheel iterating through prompts until it's great. AGI doesn't mean 'good enough' to me. It means better.
- Unknown_Unknown 13d agoThe same goes for another human being though. What you experience and want to visualize need to be described. Whether its a human or AI, you will have the same issues.
- altcognito 13d agoWould be nice if we had a better benchmark for creativity.
- zombiwoof 13d ago[dead]
- sheepscreek 13d agoSo are they doing away with the Sol/Terra/Luna split already?
- tintor 13d ago- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/ https://openai.com/index/gpt-6-astra/ - Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra https://artificialanalysis.ai/models/gpt-6-astra Who is wrong here? Some benchmark results in Astra page for Fable and Opus are blank (-). What is Artificial Analysis intelligence index measuring that Astra scores poorly on? Can someone from OpenAI / Artificial Analysis comment / clarify? Even OpenAI Astra page mentions the low scope from Artificial Analysis for Astra.
- kubrickslair 13d agoMany people claim that the Artificial Analysis Index is highly contaminated - I have not personally looked into it. Though, unlike the creators of benchmarks like Terminal Bench or ARC AGI, the Artificial Analysis Index team does not seem to have deep technical or ML backgrounds. They are ex-strategy consultants, McKinsey, et. al.
- tintor 13d agoOpenAI clearly cares about Artificial Analysis Index since they included Astra score from Artificial Analysis Index.
- AnodicElegy 13d agoIf you scroll down in the Artificial Analysis page you linked, you'll see all the individual benchmarks.
- deleted 13d ago[deleted]
- dannyw 13d agoI really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is. That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra. But AA scores Gemini 3.8 Flash at 59, and Astra at 61.
- aogaili 13d agoAmazing! We went from new JS framework every week to a new model/harness every week. Tech is really something.
- m3kw9 13d agoefficiency per intelligence is the benchmark i look at the most, as that allows the most use by most people.
- Maxforever 13d ago[flagged]
- elzbardico 13d agoAnd meanwhile, another wrapper layer is being embraced. Why would a vibecoder use Lovable when he got Sites right from ChatGPT?
- nullbio 13d agoI'm glad to see Anthropic's relevance diminishing day by day. I haven't had a chance to test this model yet, but if they've solved the web design issues and the clunky web copy it generates (like when I ask it to build a placeholder on the UI for an empty HTML table when there are no results, it puts stuff like: "The user records will go here.") then it's the nail in the coffin. On that note, Sol is absolutely atrocious for website UI copy. It's either really awkward, or really verbose and complex and doesn't sound simple or natural. Has anyone figured out a way to reliably solve this? I've tried so many different variations of instructions and skills, and nothing works. Has anyone got an instruction that is reliable, or some other mechanism?
- jrflowers 13d agoI liked the video of it googling a pediatrician. Being able to type a word into a search bar and finding a website relevant to that word? Truly the stuff of the future
- bowsamic 13d agoAll I can think of when I see the name is the crappy German beer of the same name…
- gilfoyle_7 13d agoopenai vs anthropic. that's it right? anyone else?
- nullbio 13d agoIt's more like OpenAI vs no one, at this point. Anthropic has shown they don't care about general consumers or small/med businesses. You can't even use their models without it giving refusals on the most mundane tasks.
- znnajdla 13d ago> With Sites (opens in a new window) in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt. Oops, shots fired. A direct attack on the vibe coded app market. Replit, Lovable, etc.
- edg5000 13d agoSol has been very effective at schematic design (using Skidl) and at reviewing PCB layouts. But layout was still done manually by me. I'm very impressed and surprised to see they exactly a demo of Astra doing PCB layout. This is could be a game changer for electrial engineering! It already is since the schematic (and library management) is where a lot of the design work goes.
- amelius 13d agoMy boss is already joking that I'll be out of job.
- kulkarniamey 13d agoThe model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.
- altcognito 13d agoI'm not sure how you call a model AGI without it learning new information at the model level and not introducing nasty surprises (both from adversarial users and unintentional badness) I guess there is fine tuning (and RAG) for those that need something bigger than just the knowledge contained in the context.
- Sankozi 13d agoThe same is for humans, we don't learn on instinct level either.
- ramblerman 13d agoI agree, but I also long considered llm's stochastic parrots. Then this year happened. Opus/Sol are easily far smarter programmers than I, and this thing supposedly blows them out of the water. Once an LLM is a better doctor, researcher, biologist, chemist, mathematician, physicist than any human is that not AGI? It didn't arrive in the form I would have ever imagined, but it's hard to say its not (imo).
- CringeHN 13d ago“Humanity’s Last Exam”? “ARC-AGI-3”? Is your bullshit detector going wild? Good, it’s working! How is this not the most cringe marketing strat in history???
- geonic 13d agoThese demos got me exited. Sitting in front of my computer telling ChatGPT what to do while watching the results in realtime. Hope this ends up working in reality.
- TomGarden 12d agonot a chance its gonna be realtime lol. It's fun marketing though, I'll allow it
- sbochins 13d agoI guess it’s kind of over for open ai now? We had a bunch of model releases at or around the same time, so we can get a good lay of the land. Surprise, surprise anthropic is still in the lead. Now we have Google and meta with models that are beating OpenAI in many benchmarks. There appear to be some really good cyber capabilities with this model and some other specific benchmark wins. That said, it’s as expensive as fable 5.1. It looks like all the executives that decided to leave may have picked the right time to do so. That said, I can’t wait to try it and see if the problem is we can no longer trust any benchmarks.
- DarmokTanagra 13d ago[dead]
- pknerd 13d agoThe end of the "HeyClicky"?
- wooloo26 13d ago[dead]
- listingbott 13d ago[flagged]
- kris-memoket 13d agoAGI is coming...
- ncr100 13d agoIf it's AGI then can it replace Sam Altman yet?
- quikoa 13d agoThat won't even require an AGI.
- jocelyner 13d ago[dead]
- cpach 13d agoSo what’s the tldr? :)
- ncr100 13d agoNot public Does some things better Rhetoric of AGI being tossed around, redefining the term
- Amekedl 13d agoanyone read footnote 7 for a bench, where it says something about musical quartet 2023?? idk man until I test it, I won't get excited.
- massimodeluisa 13d ago...as always, in Europe cool stuff is available 10 years later. I bet my ** GPT-6 will be one of those stuff.
- GirishWasTaken 13d ago[dead]
- iamgopal 13d agoASTRA means tool ( for war or attack specifically ) in hindi
- oblio 13d agoIt's basically impossible to pick a word that doesn't mean something unintended in 20+ major languages in the world. The recent OpenAI model names are obviously based on Latin: Luna (Moon), Terra (Earth), Sol (Sun), Astra (Star).
- MisterCrabs 13d ago[dead]
- knyadzayo 13d ago[flagged]
- dlahoda 13d agoad video is full of people with hoarse voices. it is called vocal frying. i hope they not really going to make ai voice hoarse too?
- voidash 13d agoI thought intelligence was going to be democratized but you have to stay behind a 200$ plan. The trend hints that open models can also catch up
- MadSudaca 13d agoDemocratized in the Classical Greece sense, where you had to be male, not a slave, and own land to get a vote.
- aisvgonline 13d ago[dead]
- sreekanth850 13d agoIn my mother tongue, Astra means arrow that used to destroy enemy.
- tappio 13d agoI feel that for some time now, the biggest constraint when working with models is not their intelligence, but their speed. It does not matter how smart the model is, it will make mistakes, because the instructions are ambiguous and new facts are found during implementation. The biggest problem I've had working with software developers has always been the lag between seeing the results and steering towards the right direction, not the skills of the developer (with many exceptions of course). Regardless, working on the wrong things is time wasted. And again, I'm procrastinating here while waiting for Fable to run a benchmark on a few solutions to a problem I have. We can guess what would work, but we only know after the benchmark. A faster model, with fewer capabilities, would've been a much better choice this time... well, "git gud" they said... and live and learn! Faster model = less time for procrastination. PS. AI models don't live and learn; the discussion about AGI is pretty pointless imo. It's a tool. Does it matter if it is AGI or not if it does what you want it to do? Does the IQ of your colleague matter if he's good at what he's supposed to do? Or bad? Well... I guess it does matter, as many people are up in arms about whether Astro is AGI or not. Personally, I think we're past the point for that debate. These are amazing tools.
- renegade-otter 13d agoAI models do not live and learn - it's worse. They actually get DUMMER if you don't start with a clean slate. This is important. One has to curate the context carefully.
- golol 13d agoThis is oversimplifying it. Read the metr report on the huggingface incident.
- golly_ned 13d agoDumber.
- renegade-otter 12d agoWell, isn't that an IRONIC spelling error.
- sehw 13d ago[dead]
- ozerozdass 13d ago[flagged]
- aniviacat 13d agoOn the ScreenSpot-Pro benchmark, all effort levels achieve roughly the same score. I wonder if that is just a limitation of the benchmark, or if the effort levels actually do not make a difference for purely visual tasks. > ScreenSpot-Pro tests whether models can locate the correct interface element in high-resolution screenshots of professional software.
- whoisthemachine 13d agoWell I guess we're done then. AGI is here. We can go home guys, the billionaires no longer need us.
- kaito2503 13d agoGood
- katspaugh 13d agoThe anthropomorphic delusion about LLMs being a coherent entity, with memory and continuity (which is also a delusion even when it comes to humans), quickly evaporates when you realize that each message in a chat is a completely separate stateless API request, and the only trick fueling the illusion is that the previous messages are sent to the server along with your new message. So AGI or not, it’s just a process that exists literally for the duration of one API request.
- yousif_123123 13d agoWould the marketing not land better when the model is actually publicly usable when officially announced?
- sd9 13d agoWhy are OpenAI so keen to start calling things AGI? Isn't there some corporate/legal shenanigans where they become a real non-profit at that point? Or does it just let them cut Microsoft and other investors out? There's got to be a business reason for it unrelated to the model capabillities.
- aennassiri 13d agoI firmly believed OpenAI would take back the lead. It's healthy to have competition with Anthropic and, hopefully, other LLM providers. I'm excited to test their model. Their attitude towards developers/builders has been nice and appreciated for the last couple of months.
- mannanj 12d agoI find that when someone comes from an attitude of abuse and hostility towards it users, then later without taking responsibility and owning its original actions says “we changed were being nice now” that it isn’t really a change but a farce to continue to trick you. I used to be more gullible and naive and fall for these tricks, but not anymore. You change when you take ownership of your past mistakes and share with those you harmed how you have done that. With some reparation and restoration of past harms. I haven’t seen openAI do that.
- mark_l_watson 12d agoPlease, I have a question: in September 2026, what is the price differential between buying GPT inference tokens vs. a $20/month subscription? I have moved to just using Google Gemini occasionally paying for tokens, no subscription, and it is OK, given my low level of use. I would like to add gpt-6 astra to my toolbox, but I probably would use it infrequently.
- Iolaum 12d ago$20 per month should get you far enough if you use luna, it's a decent model. If you plan to try astra you need to try make it use as little context as possible to protect your usage limit - or bank a reset mid task to allow it to keep going. If you use a decent part of a subscription - of any tier - you could be saving a lot of money. According to YT that could even be ~80-90% less than paying for tokens but take that with a grain of salt.
- mark_l_watson 12d agothanks!
- amelius 12d agoIt's time they stop picking the low hanging fruit and start folding my laundry!
- aldanor 12d agoWe'd have to wait on LeCun to make progress on that
- amelius 12d agoBut he's also known for picking low hanging fruit. (He took one of the main ideas/tools in classical computer vision, namely convolutions, and sold a cheaper version of it to the neural networks community, namely one without the FFT steps).
- amelius 12d agoconvolutions -> convolution based pattern recognition, to be more precise
- MichaelMoser123 12d ago"""OpenAI called the model a "generational leap" for areas such as cybersecurity, professional work, software engineering, and science, with the company's president Greg Brockman saying that it could eventually be seen as the arrival of artificial general intelligence""" https://en.wikipedia.org/wiki/GPT-6_Astra https://en.wikipedia.org/wiki/GPT-6_Astra according to Sam Altman from last year, an LLM should 'solve quantum gravity', if it is to count as AGI. Did anything change since then? https://www.businessinsider.com/sam-altman-openai-david-deutsch-turing-test-for-agi-2025-9 https://www.businessinsider.com/sam-altman-openai-david-deut...
- cheikhcheikh 12d agoI think that would be more of ASI than AGI, AGI is a smart human, I'd say we're pretty close if not already there for most stuff we've been focusing on like coding. ASI is a super genius, beyond the smartest human, and we're nowhere near that. But as we've already seen with LLMs you don't need to wait for ASI to solve hard problems in math and physics.
- pontus 12d agoI find the distinction between AGI and ASI to be a red herring. In my mind, we humans already have enough intelligence to basically solve any problem given enough time and a 'scratch pad'. It's sort of like Turing-completeness but for intelligence. Saying that a computer exhibits artificial-SUPER-intelligence feels a little like saying that one universal Turing machine is more expressive than another universal Turing machine. The thing that sets the two Turing machines apart is not how expressive they are but how quickly and efficiently they can arrive at the given output. It's hard for me to imagine a task that ASI could complete which a human with enough time and determination couldn't also complete, although I can imagine such tasks for dogs which don't exhibit the same level of intelligence as humans. Maybe I'm too human-centric and naive. Maybe there are other tasks out there that are beyond human intelligence even given infinite time. If, however, my view of intelligence is correct, then the distinction between ASI and AGI is illusory, and the real distinction that we would see between systems is just their speed, efficiency, determination, etc. In my mind, once we achieve AGI, we immediately have ASI: just run that AGI on a faster computer, across many nodes, etc. Instead of having humans solve quantum gravity over the next 1000 years, just have AGI solve it over the next month.
- EFLKumo 12d agoIt's said Astra is based on a newly pretrained model. Personally I hope it solves the problem that GPT speaks weirdly (which can also be spotted on Claude models after Opus 4.8 but GPT's is more severe) because IMO the way GPT couldn't get how to write good code (obviously it's trined to write code which pass benchmarks, but never code which sound good, with pursuit towards simplification and code aesthetic in mind) is extremely similar with how it couldn't get how to speak like a human. However, it seems like OpenAI didn't pay much attention on these perspectives and I didn't find if Astra could write a more elegant code, or communicate more naturally, etc., which made me somehow a little disappointed. They indeed mentioned the code Astra delivered is closer to production grade but production-grade code is different from what I want since there can be a kind of messy code blowing up your whole architecture design with control flows nobody truly understands but just passes all tests perfectly. There is no difficulty in maintaining this kind of code because you only need to paste the problems into Codex. And we all know this sounds incorrect. I don't know if my appetite towards a good code (no matter how) is sound but I just imagined frontier labs to give more attention on this. Note: fwiw Fable 5.1's release page says it's better at these perspectives of coding and per my experience, yes it is.
- madad-rashid 12d ago[flagged]
- LogicFailsMe 12d agoThe model's performance and efficiency seems like more evidence (bordering on the last nail in the coffin) against the assertion that inference will never be profitable to me, but I guess that's a common symptom of AI Psychosis according to the true believers in that assertion. All my issues with its leader aside, great work OpenAI!
- TomGarden 12d agoI'm sure this will be a great model. Personally, I'm far away from screaming 'AGI is here!' from the rooftops, until jaggedness and silly mistakes disappear at the very least . (what is going on with that Mario Kart game...) So many benchmarks are 'best of x tries' or using very specific harnesses. AGI would not need a babysitter. Honestly, even being able to do simple tasks like summarization or basic knowledge work without the constant paranoia of unforseen failure would be remarkable and useful.
- TomGarden 12d agoI'm still excited about this, Fable 5.1 and all the great open source models, but some very religious undertones have been entering the AI sphere
- aidio 12d agoTo beat Fable5
- ionwake 12d agoits funny seeing HN commenters justifying how they are most certainly intelligent, but cant agree on what that is.
- berlinbrowndev 10d agoIts probably reached a certain level of intelligence but only operating in a very confined environment. And I assume heavily language based. Human brains use language for communication and other things but it isnt the only part of intelligence. There is also the intelligence of adapting and surviving in the world. Until the AI can have a virtual environment or use the physical environment to interact with. I am still not convinced.
- aftbit 12d agoHow big is it? How much does it cost?
- exabrial 12d agogreat, I can see now we've moved onto versioning wars. Galaxy S5 vs iPhone 5 all over again
- neomx 12d ago[dead]
- guiomie 12d agoHow come they do not compare with Gemini in any of the benchmarks?
- javier_e06 12d agoIf we let AI cook. We are going to have to eat it.
- Upitor 12d agoEverybody doomscrolling in here: Relax! Take a deep breath, and go for a walk :-)
- kavee-dev 12d ago[flagged]
- Oscalemor 12d agoPerfect thread to read through some good use cases instead of the one-shot game plays across x.com First impression: the model seems kind. Always intriguing
- hrmon 12d agoOn the introduction video: The lady asks "make sure that [the presentation] feels really high-end", but something that gets generated without any effort is not high-end anymore. it will be mediocre, or more probably GARBAGE.
- skrellm 12d agoThis is great! It can finally substitute the entire upper management! Their job is easier to automate than an engineer's or any creative's, and considering their salary it's a lot more gain for any company!
- simonw 12d agoI finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01 https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents. For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-pelicans.html https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
- sanex 12d agoEven the "low" pelican is better than most other models.
- matt_heimer 12d agoIt's subjective but I think only gpt-5.6-sol XHIGH is comparable to gpt-6-astra LOW in quality. And it's 24.11 cents vs 9.55 for astra low.
- sanex 12d agoSomething I've noticed especially with fable at work is it seems to be smarter, so it takes fewer tokens and burns less of my usage than smaller models would for similar tasks.
- psii 12d agoInteresting that the composition is stable across different runs. There is much more variety in other models.
- simonw 12d agoEspecially interesting that once Astra gets to high, xhigh, max it uses the same aesthetic.
- saaaaaam 12d agoI think the Gemini 3.8 Flash ones the other day were superior pelicans. Particularly the murderous one who was going to squish his tiny cousin. Interesting that terra xhigh effort is cycling right left, in opposition to the prevailing pelican left to right direction.
- mattermani 12d agoCant wait to try
- iqandjoke 12d agoFor the "Formatting a legal document" demo, may I know which legal firm would use LibreOffice Writer to do their job?
- rhet0rica 6d agoOne that wouldn't mind manually typed page numbers that would have to be adjusted were the document ever updated. They also inexplicably used Windows 95 Paint to draw a rocket ship during the main video demo.
- Ghirardelli 12d agoسلام
- unixhero 12d agoWhat's the open desktop score?
- Alifatisk 10d agoSeeing what others have already created using GPT-6, it seem to be a new stepping stone in capabilities and overall "intelligence". However, some other details I think is worth bringing up is that this model is 70% more token efficient than GPT-5.6 Sol and consuming 1/3 of the tokens compared to Sol (max) in the Codex [1]. I am just thinking loudly here but, it seems like even though Astra is pricier than Sol, you might actually get more usage out of it? I did some digging myself and looking at FrontierCode and DeepSWE, Astra (low) seem to perform better than Sol (medium) and on par with Luna (max) while being somewhat on the same price range to Sol? [2][3]. And now we have four models to chose from, each with their varied reasoning efforts: Astra, Sol, Terra and Luna. Personally, I feel like Terra have turned into this middle child in a weird spot that's neither the option as cheap model because Luna is, yet it is not an good option for complex tasks because Sol is already good at it. For background, I use Luna (xhigh) daily, I think its a fantastic and underrated model. Especially Luna (max). It is way more capable than what it looks like, I think people underestimate it because OpenAI described it as "roughly corresponds to the nano model tier used in earlier GPT-5 families" [4]. I also like Luna because it barely consumes my weekly usage. Last week, it only ate ~15% of my weekly usage. So usage is not an issue anymore. I never have to worry. It may not be the fastest model because, well, it reasons as max effort, but it does the job way better than I expect. Also, considering how much one saves on the weekly usage, one can probably turn on "fast mode". Haven't done it myself though. Also, another thing that caught my eyes is this: > Historically, models have used compaction to summarize work during long sessions, such as when debugging complex issues or tackling large refactors. Each compaction can leave out details about why a fix failed or how a component behaves. > In Codex, Astra can keep notes across context windows, preserving accumulated details without repeatedly compressing them into a single summary. Earlier context windows remain searchable, so Astra can find requirements or test results from previous messages and tool outputs—even if that information wasn’t captured in its notes. You can enable this experimental feature in your Codex config.toml [5]. I was curious about this, because I know Luna (max) spews out tokens which can trigger compaction quite often. If you go to the config reference [6] and search for "features.context_management.experimental_mode", you will find this: > Enable experimental context management. Rather than repeatedly compressing context into a single summary, it uses notes and searchable history to preserve accumulated details. This is a very interesting feature and perhaps very useful during long horizon work in a thread where the conversation context window grows and compacts often. 1. https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra https://artificialanalysis.ai/articles/benchmarking-gpt-6-as... 2. https://deepswe.datacurve.ai https://deepswe.datacurve.ai, Astra (low) got 67% $2.19, Sol (medium) 61% $1.42 and Luna (max) 67% $0.61 3. https://cognition.com/frontiercode https://cognition.com/frontiercode, Astra (low) 45.3% $1.60, Sol (medium) 39.9% $3.12, Luna (max) 39.8% $0.36 4. https://developers.openai.com/api/docs/models/gpt-5.6-luna https://developers.openai.com/api/docs/models/gpt-5.6-luna 5. https://openai.com/index/gpt-6-astra https://openai.com/index/gpt-6-astra 6. https://learn.chatgpt.com/docs/config-file/config-reference https://learn.chatgpt.com/docs/config-file/config-reference
- aeroscissorz1 8d agoIn my perspective, we are not completely doomed till it can not generate what already hasen't been done.