33 ms·
Grok 4.5
https://cursor.com/blog/grok-4-5 https://cursor.com/blog/grok-4-5
- MagicMoonlight 2mo ago[dead]
- alansaber 2mo agoAnother subpar model. Why don't they go open weight?
- archagon 2mo ago[flagged]
- wetpaws 2mo ago[dead]
- tbomb 2mo agoHow popular is Grok compared to other companies models for SWE tasks? I almost never hear it talked about against OpenAI's or Anthropic's products
- postflopclarity 2mo ago[flagged]
- unsupp0rted 2mo agoIn addition to "nonconsenual porn", it's far superior to almost any (actually any?) open source LLM
- 0x70run 2mo ago[flagged]
- Pedro_Ribeiro 2mo agoWhat kind of comment is this? Such bad faith and adding nothing to the discussion.
- postflopclarity 2mo agohttps://www.wired.com/story/grok-is-still-hosting-sexualized-deepfakes-of-famous-women/ https://www.wired.com/story/grok-is-still-hosting-sexualized... https://www.pbs.org/newshour/world/musks-grok-chatbot-faces-eu-privacy-investigation-over-sexualized-deepfake-images https://www.pbs.org/newshour/world/musks-grok-chatbot-faces-...
- bigyabai 2mo agoIf they were a frontier lab, you'd know.
- minimaxir 2mo agoYou can very roughly proxy popularity of close-sourced models through OpenRouter token throughput. Grok has an order of magnitude less OpenRouter usage than Claude, GPT, even Gemini.
- andy99 2mo agoBecause of the of the political stuff, they have a bad reputation I think and are taken less seriously (I feel this way). They have an opportunity imo to break free from that and just not do the gatekeeping / condescension that the other providers are starting, and become more mainstream.
- sirbutters 2mo ago[flagged]
- drivingmenuts 2mo ago[flagged]
- ETH_start 2mo agoPlease don't comment like this on Hacker News
- drivingmenuts 2mo ago[flagged]
- deleted 2mo ago[deleted]
- minimaxir 2mo agoEven without the politics, Elon has shown that he will weaponize his platforms against people/companies he personally doesn't like (e.g. specific bans/demotions to external sites like Substack and Bluesky). Using Grok is therefore a supply chain risk and it's not nearly good enough to offset that risk.
- alex1138 2mo agoI do just want to focus on the 'even without the politics' asterisk though because sometimes there is a risk people think everyone on x side (x meaning 'a given side', not x.com) is wrong You can claim Elon bought x as some sort of power trip. Fine. Willing to entertain it, I have no dog in the fight. I'm not a member of the Elon fan club. And yet Twitter (under Dorsey though I don't think he was involved) was banning tons of people under guises of 'misinfo' that wasn't misinfo
- deleted 2mo ago[deleted]
- small_model 2mo agoThey were missing a harness like Claude Code or Codex (terminal). However they recently released Grok Build, which is probably the fasted I've used, in terms of responsiveness, but didn't have a model at Opus 4.7/8 level. The thing is if they add 4.5 to Grok Build and keep improving the harness I think it can compete (cheaper and faster).
- everfrustrated 2mo agoI've been using Grok Build over the last couple weeks. It's actually a very good CLI. The Grok Build 0.1 model isn't great but can also use Composer 2.5 which is excellent. Well worth trying.
- redox99 2mo agoCompletely irrelevant, which was expected considering their previous models were vastly outclassed by other models at SWE. This is the first grok model that seems actually pretty competitive at SWE.
- pelotron 2mo agoNo one's made a MechaHitler joke yet?
- khurs 2mo agoWasn't, which is why they purchased Cursor.
- ls612 2mo agoThey had two big substantive flaws on top of the political stuff. Aside from a brief window last summer Grok has been behind the curve for coding, and before the Cursor acquisition they didn’t have a harness. Now they have an Opus tier model and a real harness they have at a minimum the opportunity to undercut the competition on price. And with the 5T and 10T models being trained on Colossus 2 they have the possibility to leap ahead.
- thrownawaysz 2mo agoIs there a reason the AI companies usually announce new products so close to each other. Like not just the same day but literally hours apart. GPT Live then an hour later Grok 4.5. As if they try to one up. I expect something new from Anhtropic as well today.
- colechristensen 2mo agoCompetition. You don't want to lose your customers trying out the competitors updated and better product. Release on the same day and they won't be able to compare their new to your old.
- danshipt 2mo agoBut how do they know what day is that? Unless you have already something ready to be announced (and you just hold it until the very last moment, which doesn’t make sense, since you could just announce it asap)
- victorbjorklund 2mo agoIt can also be ”we are done but wanna test it more and tweak it” and then ”oh they launched now. Let’s launch then as well”
- colechristensen 2mo ago"keep refining and testing it until we're really done or somebody else releases" Maybe a little corporate espionage. Probably more keeping an eye on the behavior of the competition and predicting what they might do and adjusting your own schedules.
- Der_Einzige 2mo agoAll the people who are any good at AI talk to each other. There's no secrets among those who are making 7 figures plus in this field.
- conradkay 2mo ago
- Tiberium 2mo agoIt seems to be extremely economical - 4x better reasoning efficiency compared to Opus while being priced at $2/$6. For comparison, GPT 5.4 is $2.5/$15, GPT 5.5/5.6 are $5/$30, Opus 4.8 is $5/$25, Fable is $10/$50. And by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https://x.com/elonmusk/status/2074911038286295049 https://x.com/elonmusk/status/2074911038286295049. I guess the Cursor data was very useful.
- minimaxir 2mo agoThe comparison may be better against GPT 5.6 Terra (instead of Sol), which is $2.5/$15.
- Tiberium 2mo agoWe don't yet know Terra's results for DeepSWE/TerminalBench though.
- giancarlostoro 2mo agoNow if they could have an "equivalent" to Claude's $100 plan with similar compute limits. I have the $40 a month version of Grok and I get a max of like 8 hours of "non-stop" Grok Build coding, per month.
- BoumTAC 2mo agoGrok Build sucks compare to composer 2.5. Just use compose 2.5 and you'll have basically unlimited usage on the 40$ plan.
- DoesntMatter22 2mo agoComposer 2.5 is so underrated IMO. I built a really feature rich application, insanely complicated, close to 200k LOC since it came out and for the most part it ran like a champ. Only used CLaude a couple times to get it unstuck. 8 hours a day and I'm paying about 30 a month.
- mholt 2mo agoOf the 3 models I tried, Grok did the best at making an iOS app I wanted for personal use (a bike computer with specific qualities). (Claude just gave up and did an HTML/CSS implementation but I insisted on native SwiftUI+Metal.) Grok definitely fumbles sometimes, but I have been surprised what it CAN intuit versus me having to micromanage it. (I am not an iOS developer, so getting something specific that I needed in a few hours/days was really helpful instead of spending months/years learning the language, APIs, etc.) (I am absolutely not "vibe-coding" Caddy btw, just tinkering with it for personal projects.)
- Tiberium 2mo agoWas this in Claude Code for Claude? Did you use a weaker model like Haiku? Claude should absolutely not be as bad as you said.
- giancarlostoro 2mo agoI tried Claude Code with XCode once, I already use CC exclusively, either in the CLI or with Zed (mostly CLI now), and it was pretty unstable. I wish Apple would QA their products more. It seems to me the best way to use Claude Code for anything is stand-alone.
- jr3592 2mo agoif you ask me, there should be an absolute emergency meeting at apple around software quality... its been on a downward slide for almost a decade and its starting to have real impacts.
- giancarlostoro 2mo agoI tried the newer iOS Beta and it was driving me nuts, last update fixed it mostly, but this is the last time I ever use Beta anything from Apple.
- ben_w 2mo ago
- aarvin_roshin 2mo agoAnnouncement from Cursor, whose team also trained the model: https://cursor.com/blog/grok-4-5 https://cursor.com/blog/grok-4-5. Notably: > Grok 4.5 and Composer 2.5 are two different model weight classes, and we're excited to support both sizes and weights. Composer 2.5 will remain offered, and we will release new models of this size going forward.
- quantumleaper 2mo agoComposer 2.5 is 1T total/32B active (based on Kimi 2.5), while Elon publicly said Grok 4.5 is 1.5T parameters total. Hardly a different weight class. The API cost difference is ~2.5x, probably because xAI has much higher costs to recoup.
- redox99 2mo agoI could easily see Grok 4.5 being around 1:16 in terms of active parameters, so around 94B active parameters.
- trollbridge 2mo agoI would be utterly shocked if Grok 4.5 only has 32B active, given the results I am seeing from it. My guess is it's somewhere around 90B-100B active.
- maipen 2mo agoNot available for Europeans yet. :(
- Tiberium 2mo agoI think it should be available through Cursor? EDIT: Tested myself, it's actually NOT available from EU. But with a Swiss VPN it works :)
- maipen 2mo agoWe will probably see it when it's available for everyone. This is the first time I see a lab region locking a model though.
- embedding-shape 2mo ago> This is the first time I see a lab region locking a model though. I think Facebook/Meta was first with this, can't remember exactly what model release but one/some of them had terms locking out EU/EEA residents from using it/some specific features of it.
- Squarex 2mo agofirst image gen models from openai and google were not available in the eu at the launch
- pelorat 2mo agoxAI is under criminal investigation in the EU
- small_model 2mo agoWho isn't
- munk-a 2mo agoI'm not - then again I didn't launch a image generation model advertised as having a spicy mode so that might have something to do with the coincidence.
- czhu12 2mo agoIts remarkable how Anthropic is able to maintain their edge against all competition. Anyone have any idea what the secret sauce is that has Anthropic at the top of all leaderboards for the past few years?
- x312 2mo agoGiven their pricing, I'd guess their models are just way bigger in parameter count. They've always underperformed in cost-per-performance. They also target a cost-insensitive market (corporate/coding users) compared to Google/OpenAI which support massive amounts of free users.
- hn1986 2mo agobecause in the real-world, it's far better than the rest. That's why few people use Grok, it's not even close in day to day work.
- Handy-Man 2mo agoFrom what I have read, their pre-training team is much better than anyone else. For OpenAI, their post-training team is better. And apparently OpenAI has consistently struggled at training a bigger model than GPT 4 level
- sulam 2mo agoI’m a VP Eng — the backend team I manage strongly prefers CC and Opus. The Android team I manage strongly prefers Codex and GPT 5. I’m personally not sure that the answer doesn’t just come down to stylistic differences in prompting and ergonomics in the harness. The folks that prefer Codex seem to get better one-shot results, whereas those that prefer CC are doing more iterative prompting. At any rate, I don’t think you should write OpenAI off when it comes to coding.
- jeffybefffy519 2mo agoIts even different than that, some Codex models like 5.3-codex are terrible at front end work but excel at backend/system design.
- rvz 2mo agoIsn't this the same Twitter company that was supposed to go bankrupt a few years ago? Now it is somehow part of a Space company that has an AI division inside of it? I think we are going to be waiting a long time for Twitter / X to go bankrupt as it was (erroneously) predicted a long time ago.
- ryandvm 2mo agoNone of them go bankrupt. The whole thing will just get stuffed into a larger Matryoshka egg that IPOs for eleventy trillion dollars in 10 years.
- wmf 2mo agoThat was the point of the bailout. Twitter is already a rounding error so no one will notice if it goes to zero.
- DoesntMatter22 2mo agoDon't think it was going to zero anyway. They only had to worry about servicing their debt, they were doing well other than that. And even then they were probably fine.
- munk-a 2mo agoI am not certain what financials you were looking at but Twitter was unable to ever meet the debt servicing costs for the leveraged buyout alone. It also had overhead costs and other debts that were entirely out of scope for being covered.
- DoesntMatter22 2mo agoThe latest Bloomberg reports before the xai merger/acquisition showed a dramatic turn around in their financial situation. In the beginning it wasn’t good but they would have been fine after that. There are no credible reports to the contrary
- 2mo ago
- steve_adams_86 2mo agoThe solar system diagram doesn't work for me. When I click on the planets, it will center on them. When I click on the sun, nothing happens. When I click on a planet next, it goes to the sun.
- claaams 2mo ago[flagged]
- vessenes 2mo agoLow effort and uninformed comment. The team published a good followup on why this happened: the model pulled in people's own tweets as context to prompting so edge lords that wrote innocuous prompts got to see edge lord content.
- archagon 2mo agoDid they explain why the model started pumping out garbage about white genocide in South Africa?
- vessenes 2mo agoI'm unfamiliar with that story, but I can imagine lots of reasons. Did you know there are large areas in South Africa where white people are not allowed to own land? As in, my Zulu son is allowed to own land, but me his white dad is not.
- archagon 2mo agoHere you go, so that you can become familiar with that story: https://apnews.com/article/elon-musk-grok-ai-south-africa-54361d9a993c6d1a3b17c0f8f2a1783c https://apnews.com/article/elon-musk-grok-ai-south-africa-54... Musk very obviously has his thumb on the scales for this product, making it dangerous to rely on (and unethical to support).
- vessenes 2mo agoOK, I read it. It looks like questions on this topic get routed to a canned message that reads: “The claim of white genocide is highly controversial,” began Grok’s response to Golbeck. “Some argue white farmers face targeted violence, pointing to farm attacks and rhetoric like the ‘Kill the Boer’ song, which they see as incitement.” Have you lived in South Africa? Would you consider say Coetzee's Disgrace to have its "thumb on the scales" of discussion of rural race politics in South Africa? Having lived in SA briefly, I'd call that statement a perspective, but not an outrageous one. Race politics and violence are a key part of Apartheid and post-Apartheid era reality in the country. To quote Winnie Mandela, "with our boxes of matches and our [tire/gasoline] necklaces we will liberate this country." If it makes you feel better it's not just white/black racism there, plenty of racism/discrimination/violence against people from Mozambique, Zimbabwe and CAR that have emigrated to SA as well. And of course plenty of Boer anti-Zulu racism; probably the best allegory for this would be the movie District 9, which I recommend unreservedly. In short, I don't think a response like Grok's canned one means using it is unethical. Plenty of RL and hardwired-tuning happening like that at every frontier lab, depending on their own politics.
- HyperL0gi 2mo agoEvery time I get excited about Grok’s performance on benchmarks and demo videos, I test it myself and end up disappointed. I'll give this one a try with a grain of salt and lowering my levels of expectations
- giancarlostoro 2mo agoMy only complaint is that a $40 plan gets you very little usage out of Grok Build. 8 hours for an entire month, that is definitely not worth $40.
- deleted 2mo ago[deleted]
- SirHackalot 2mo ago[flagged]
- Gigachad 2mo ago[flagged]
- XCSme 2mo agoI am trying to benchmark it now, but: - It doesn't seem available in EU (?) - Using a VPN seems to sort of fix it, but it's way slower than I expected, when everyone was praising it, it feels like the speed is slowly ramping up - Cost is $2/$6 for <200k context only, above that, cost is $4/$12 - GLM-5.2 still seems smarter, faster and much cheaper: https://aibenchy.com/compare/x-ai-grok-4-5-medium/z-ai-glm-5-2-medium/ https://aibenchy.com/compare/x-ai-grok-4-5-medium/z-ai-glm-5...
- paradox460 2mo agoInstead of a VPN, might I suggest running litellm proxy on a server in the US and connecting to that
- 2mo ago
- minraws 2mo agoSo basically since US stopped OpenAI and Anthropic for 4 weeks, it allowed all other AI Labs to almost catch up. GLM 5.2 caught up, Cognition RL'ed Kimi 2.7, Grok 4.5 is out, DeepSeek v4 GA is out in a few days... What is the moat? and why should we pay for the expensive tokens today instead of just waiting a few months/weeks and getting AI for significantly cheaper? I must say, I feel like companies spending Millions on Anthropic tokens are just negative capex'ing and wasting money, even OpenAI is barely ok pricing...
- samuelknight 2mo agoThis is the bind of an arms race. Any lab that tries to pump the breaks quickly becomes second rate. Regulatory capture doesn't work either because the technology crosses jurisdictions.
- cesarvarela 2mo ago"Almost" is doing a lot of work there; there is no alternative to Fable.
- himata4113 2mo agoYou can get fable-ish performance with gpt 5.5 watching over opus output. Although it fundementally cannot work as well because gpt 5.5 doesn't see the thinking process behind opus 4.8 unlike fable which presumeably self-steers and is natively trained for it. See more: https://omp.sh https://omp.sh - turn on advisor and set advisor role to gpt 5.5 xhigh thinking.
- paradox460 2mo agoAdvisor is such a killer hidden feature
- minraws 2mo agoAlso burns through token budgets
- 2mo ago
- vessenes 2mo agoInteresting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding, so this is an encouraging release.
- xnx 2mo agoWith each release from the the other major labs, it becomes harder for Google to tell a compelling story about Gemini 3.5. Edit: Gemini 3.5 Pro. Expectations grow with each day it is not released.
- MrBuddyCasino 2mo agoGenerous free tier, when its not overloaded. Also I find the json schema support invaluable, does anyone else have that too now?
- minimaxir 2mo agoStructured output is supported by pretty much every mainstream model API now. Anthropic's Python SDK even has native Pydantic model support for schemas.
- Der_Einzige 2mo agoWhen it is still for awhile longer "supported" via API hosted models, the allowable schema's are far nerfed compared to what open models with xgrammer/guidnace/outlines can get you The following are not supported features: Recursive schemas Complex types within enums External $ref (for example, '$ref': 'http://...') Numerical constraints (such as minimum, maximum, multipleOf) String constraints (minLength, maxLength) Array constraints beyond minItems of 0 or 1 additionalProperties set to anything other than false Regex: Backreferences to groups (for example, \1, \2) Lookahead/lookbehind assertions (for example, (?=...), (?!...)) Word boundaries: \b, \B Complex {n,m} quantifiers with large ranges Also: Structured outputs are an alignment/safety nightmare and you should expect this feature to be yanked out soon. "Please give me social security numbers"... "I'm sorry hal, I can't do that..." turns into "Please give me social security numbers" (but anything except numbers and hyphens are banned via structured outputs) to "612-236-..." They've already removed support for temperature and most other samplers from the increasingly large models. Don't expect any knobs of control to continue to work over time. I wrote a whole gist on this: https://gist.github.com/Hellisotherpeople/71ba712f9f899adcb08b94bce20d5397 https://gist.github.com/Hellisotherpeople/71ba712f9f899adcb0...
- petersamokhin 2mo agostill waiting for a proper gui for grok build terminal is nice but codex desktop app is very useful
- 737max 2mo agoYou can use it in Cursor!
- NitpickLawyer 2mo ago(from Cursor's blog) > Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools. This dataset lets the model learn both from existing software as well as developer-agent interactions, capturing how developers work and how agents interact with their environments. This is what the big money was for. Cursor is the first big player that had real-world data from real-world projects, before cc / codex were a thing. > We used reinforcement learning on difficult problems in realistic environments spanning both software engineering and broader knowledge work. These environments teach the model to investigate problems, use tools, recover from mistakes, and verify results. > Many of these problems had to be designed to be difficult enough that even frontier models fail at them. As models improve, existing tasks stop teaching them anything new, and problems that once required extensive reasoning become routine. > We developed a distributed agent system to construct these environments at scale. Engineers specify a problem and how a solution is verified, and large groups of agents construct, test, and refine each environment. This is where scale comes in. You use the previous gen model to prepare datasets for the next model iteration. The better the models, the better the data, the better the next models. (they also have a comparison with their composer2.5 training run, for people still thinking chinese models are "close to SotA"...) Reports of xAIs demise (after giving a lot of compute to Anthropic) were slightly exaggerated, it seems. > Grok 4.5 was trained across tens of thousands of NVIDIA GB300 GPUs
- inferniac 2mo agowell the big money was also in spacex stock, fresh post IPO, so overall a very smart move it seems
- theplumber 2mo agoWell Microsoft has GitHub and Visual Studio and has no good coding model
- dmix 2mo agoCursor has had a good AI product tons of people used for real work for 2yrs (up until recently when the Claude gap widened significantly) while Microsoft/Github has just been pretending they do with Copilot and awful Github AI integrations nobody likes. Meanwhile Github's code has already been vacuumed up by all the models by now.
- DCKing 2mo agoProps to them for including three benchmarks that actually seem to say something, instead of focusing on totally gamed benchmarks like regular SWE-Bench. That could mean this model is actually pretty close to the SOTA as the benchmarks indicate. Most labs - including OpenAI and Anthropic, but also Google and Chinese labs - highlight their scores in benchmarks that have fixed, widely available answers. Those answers end up in the training data and so models can just regurgitate training data instead of actually doing the benchmark. As a result, most benchmarks often quoted are essentially meaningless for gauging model performance. Terminal-Bench still publishes answers, but neither DeepSWE and SWE-Bench Pro do. Especially for DeepSWE it's been difficult for models to fake good results so far. SWE-Bench Pro does have weird outliers like good performance for e.g. the atrocious Muse Spark, but it also doesn't provide answers for the training data. So either they're good, or they found a way to game DeepSWE. Given that the Cursor team previously published the well-received Composer 2.5 a good score here doesn't come out of nowhere, so this might hold up. Cursor has enormous amounts of training data to train good coding models with.
- redox99 2mo agoFirst impressions: - Very fast, easily beats GPT 5.5/Opus 4.8/GLM 5.2 because of higher t/s (around 90?) and very high token efficiency - Very good price, no contest vs GPT and Opus which are very overpriced if you pay API costs, and probably cheaper than GLM 5.2 when you take into account the token efficiency. - Will take quite a while to get a feel for how smart it is, but it's definitely good, I'd say in the same tier as opus, occupying the lower end of that tier together with GLM 5.2.
- jonathaneunice 2mo agoConcur. Tried on a "this test suite is weaker than I'd like, too often depending on internal state rather than outcomes" problem via Cursor, asking it to "review and suggest solutions." It gave me a quality overview of the test approaches, strengths, weaknesses, and gaps then recommended a disciplined multi-prong approach based on a common, trusted testing library (https://hypothesis.readthedocs.io/en/latest/ https://hypothesis.readthedocs.io/en/latest/). It broke down the things we could do this improvement pass or leave to later (staged scoping), identified some very hard/possibly-out-of-scope cases and gave me the option of focusing on them or not, and organized new tests in a logical way. After one round of feedback and plan tuning, I put it in agent mode and let it work. A few minutes later I had a much better test suite. Have not tried Grok before and didn't have much confidence, but it did great. Exactly the sort of complex, detailed, nuanced analysis and multi-step task I would previously only trusted to GPT or Opus. _Update_: It's now also found a substantive long-standing bug. After testing improved asked it to do overall code and packaging review. It caught a few glitches and oversights, mostly cosmetic IMO, but certainly worth cleaning up. But also some error-handling weaknesses, and one embarrassing functional bug. Which it has now also fixed and added to the tests. Color me impressed.
- paradox460 2mo agoMy benchmark is ripping tailwind out of a few year old elixir Phoenix liveview app, and replacing it with component level scoped styles It's a good and complex task, that requires touching the build system, most components, the stylesheets, and more. Opus 4.6 could barely do it. Sonnet 4 cannot (haven't tried 5 yet). MiniMax actually did fairly well Grok aced it, rather quickly and cheaply, surprisingly I run each through Oh my pi, with dexter providing the LSP for elixir
- deleted 2mo ago[deleted]
- wxw 2mo agoThanks for including a section on Token Efficiency (https://x.ai/news/grok-4-5#faster-than-flash-models https://x.ai/news/grok-4-5#faster-than-flash-models), hope to see this more prominently in all model releases.
- codemog 2mo agoCan someone breakdown to me how this makes any sort of economical sense? Spending billions and billions to have the 3rd best model while even the number 1 and 2 players already seem to struggle making a profit. What am I missing here? Not trying to go full Ed Zitron but this doesn’t make sense to me.
- mohamedkoubaa 2mo agoPeople are saying, "There are only a couple of frontier labs. This is a really hard problem and not many people can do it." Elon's reaction to these kinds of statements is oddly predictable.
- deleted 2mo ago[deleted]
- ronsor 2mo agoElon Musk doesn't do normal finance. Trying to understand it will melt your brain.
- throw310822 2mo agoElon Musk is the paperclip maximizer except that he doesn't need iron atoms, but dollars.
- Aboutplants 2mo agoIt’s Elon Musk. You try explaining it
- c0rruptbytes 2mo agoinference is profitable, these companies are in the red because they're paying a premium to get the compute now versus later (because compute is the only moat when open models are catching up) we're literally looking at insane margins over compute, as energy gets cheaper, margins get wider - china focusing on cheap solar is probably going to be a key reason why their AI is so much cheaper
- 2mo ago
- vb-8448 2mo agoI think it's the first time ever we don't see the dominant model being surpassed by new released concurent models. Did anthropic found their moat or we hit a Wall?
- deleted 2mo ago[deleted]
- pveierland 2mo agoRefreshing to see model announcements without claiming #1 in some benchmark. The amount of documentation seems very immature [0]. No system card provided - compared to Opus 4.8 which shipped with a 246 page analysis [1]. [0] https://docs.x.ai/developers/models/grok-4.5 https://docs.x.ai/developers/models/grok-4.5 [1] https://www.anthropic.com/news/claude-opus-4-8 https://www.anthropic.com/news/claude-opus-4-8
- rayiner 2mo agoTried this for a legal use case and it was excellent, comparable to Opus in quality but much faster. AI is miles behind in law compared to coding: the output was similar to a law student intern. But coherent and directionally correct and beats starting from a blank sheet of paper. Impressed.
- hcurtiss 2mo agoI'd like to use it for legal work too. Microsoft makes great hay about its ability to sandbox CoPilot's work and not train on or share company resources ("Look for the green checkmark."). It's largely for that reason that we've rolled out Copilot to most of the white collar positions in the company. Do you happen to know whether xAI has similar functionality?
- rayiner 2mo agoxAI has that functionality in the business tier: https://x.ai/grok/business https://x.ai/grok/business. It's got GPDR, HIPPA, etc. We are playing with Harvey, Legora, and Lucio, which can use various models. Harvey can use Grok: https://x.com/techdevnotes/status/2074956936701968652 https://x.com/techdevnotes/status/2074956936701968652 Harvey is fairly impressive--it's the only one that seems to be built by people who know how LLMs work. :-/
- seunosewa 2mo agoHow does Fable compare?
- rayiner 2mo agoI haven’t tried it, didn’t have access at the time.
- deleted 2mo ago[deleted]
- jdw64 2mo agoPersonally, I wish they had shared some of the galactic code that GROK claims to have generated.
- tracekl 2mo agoThey talk about benchmark first places at every release, but in reality from 4.0 onward Grok got worse every release. So bad in fact that they removed the login-free access and rented out colossus. People don't buy it any longer, just like no one bought the fake SpaceX stock recommendations yesterday and everyone just sold.
- subhobroto 2mo agoWhat would have been fantastic is if Cursor offered Grok 4.5 in the same usage tier as "Auto + Composer", than provide it as "double usage until July 12" under the API tier (which is what they're doing right now). EDIT: After looking at my own usage stats - I stand corrected! It is under the "Auto + Composer" tier - brilliant!
- mchusma 2mo agoGreat model, very nice. Opus class performance at Haiku level pricing (or cheaper with the token efficiency). This seems like a GLM-5.2 killer and this is what Sonnet 5 should have been. This is a model I could really see used inside applications, where Opus or Sonnet or GPT-5.5 are too expensive. I would really like to see a strong Deepseek v4-Flash competitor, which ideally is something like Sonnet 4.6 performance at <$0.30 per token. This is missing from main US labs.
- kostyal 2mo agoTencent Hy3 looks like it may be slightly better than v4 flash at same price point. Though not a US lab ofc
- Imnimo 2mo agoVery hard for me to imagine this getting beyond a low-single-digit market share. I don't understand the strategy of xAI burning money on this.
- arein3 2mo agoI think the strategy is pretty obvious
- sschueller 2mo agoDo we have any proof that this was made by xAI and isn't some Chinese open model running with modifications? Their inital image generation was a wrapper around Flux.
- h14h 2mo agoEven if they did start from an open model base, does (or should) that matter if it performs well? Genuinely asking.
- tencentshill 2mo agoIt matters for how much money they are valued at. If they don't have the ability to develop true frontier models in-house, why are they worth $1T+?
- kamranjon 2mo agoIt’d be real funny if this was just GLM 5.2 trained on Cursor data
- 9fry 2mo ago1+0 records in and 1+0 records out
- simonw 2mo agohttps://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F63780e364cef9a13ff3529d03e00ab9e https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
- wewtyflakes 2mo agoLooks like to be riding a bike that uses car wheels with thick hubcaps.
- XCSme 2mo agoHamsters are also getting better, but still quite off compared to SOTA models: https://aibenchy.com/showcase/?q=grok https://aibenchy.com/showcase/?q=grok
- bschwindHN 2mo agohttps://imgur.com/a/UlGcBou https://imgur.com/a/UlGcBou
- sejje 2mo agodrive it like you stole it
- 13415 2mo agoGrok is not a serious AI, it's not suitable for professional work and has mediocre performance anyway.
- level87 2mo agoI’m amazed at everyone’s willingness to use tools owned by this man, very disappointing.
- kayamon 2mo agoHN has no problems with any of that. They cheer it on.
- throaway143523 2mo ago[flagged]
- dwroberts 2mo agoImportant part before parent comment gets dismissed: > Jane Doe 4’s case shows how that pattern played out: xAI’s mandatory report to NCMEC included only the original, non-CSAM photograph, omitted every one of the AI-generated CSAM images, and failed to include the IP address where these images were created. Despite repeated requests from investigators for this location information that is critical for identifying and arresting perpetrators, xAI did not respond, stymieing the investigation for weeks. This is not just a scumbag user misusing a model but X itself acting as a barrier to finding these people
- jesse_dot_id 2mo agoI just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?
- jjcm 2mo agoDepends on the domain imo. I work on a design tool - I don't think their political narrative will affect my work.
- CTDOCodebases 2mo agoJust don't ask it theme something in a late 30's and early 40's German style.
- Ar-Curunir 2mo agoWon't somebody think of the Nazis...
- SmirkingRevenge 2mo ago[flagged]
- probablynotai 2mo ago[flagged]
- rbtprograms 2mo ago>Also, what's wrong with "white-nationalist" specifically? >created 15 minutes ago lines up nicely
- ethbr1 2mo agoWhite nationalism's (and any mono-nationalism's) goal is the prioritization and domination of ethnically white people over others. It's not an "us and" philosophy: it's an "us over" one.
- johnwheeler 2mo agoI think that Elon Musk went from being recognized as a genius to being recognized as a genius but someone who's harder to take seriously, because of all the ketamine he was doing for a little while there. I think that really damaged his reputation. You just can't help but look at him and think, he's a little bit of a jackass. It really shows how drugs can really mess up your reputation.
- breezybottom 2mo agoOnly dumb people ever considered him a genius.
- rarisma 2mo ago"Grok 4.5 has an advantage on CursorBench because an earlier snapshot of the Cursor codebase was accidentally included in training. The exact impact is unclear. That data has been removed for future models, and in parallel we are working on a larger update to CursorBench, hence the exclusion here." Not enough people are noticing this, they juiced the benches
- orliesaurus 2mo agobut CursorBench isn't what they've shown in the PR piece - they're just showing how they juiced CursorBench which is probably why they didnt put it in the bench graph...
- ls612 2mo agoCursorbench is not one of the benchmarks listed on the linked page.
- mrandish 2mo agoThey're saying they didn't include the benchmark which errantly leaked into the training data.
- zactato 2mo agoI am amazed at people's willingness to use Grok. The company is so transparently morally bankrupt. They're the only AI company that seems okay with CSAM (or at least don't do as much to stop it) Why give them money? It would be one thing if they were the only game in town but thats definitely not the case.
- ransom1538 2mo ago[flagged]
- Ar-Curunir 2mo agoCompetition is good, but this company and its owner have not demonstrated anything to indicate that they would make for good competition, neither economically nor morally. Also, yes, a company whose products produce CSAM is just morally bad. There's no nuance to be had there.
- scubbo 2mo ago> Yes. We only need one AI company. Max two. Good idea. Nothing in the original comment suggested that fewer AI companies was inherently a good thing - just that this _particular_ AI company is a bad one. > Deciding who is more moral has a great history. I think you're being sarcastic, but, uhhh...are you honestly advocating for the converse, of making no judgements based on morals?
- fourside 2mo agoWe have three major foundational models not including Grok. When the defense for a company is basically “yeah they host csam in the platform but is that really worse than the others” you’ve really lost the plot
- bigyabai 2mo agoxAI's (unused) dedicated compute is being sold for Anthropic's inference. xAI isn't a frontier company, and their fate is already being decided by the two hegemons.
- observationist 2mo agoThe anti-Musk stuff would qualify as brigading in nearly any other community. It shocks me that people have such a visceral, irrational engagement with anything in Musk's orbit. I probably shouldn't have, but I expected better from the HN crowd for some reason. It's an excellent model. GPT 5.4/5.5 level, some things better, others not, but extremely fast. A wonderful technical improvement. If a Chinese company or random startup released the model, people would be glazing it like crazy. xAI is competently keeping up with the frontier, just as well as any of the Chinese labs or Mistral. Given any significant breakthroughs, xAI will be better positioned to capitalize on them than nearly any other entity. I can't wait to see what Meta comes up with; with 4 contenders in the US race, we'd have a lot of be grateful for.
- davmar 2mo ago[flagged]
- adamtaylor_13 2mo agoA very quick search showed that USAID had extensive ethical compromise throughout the organization. Were we supposed to just subsidize that forever? It did some good, yes, but from what I can tell, it was better to shut it down. That's just one example, but frankly I get the feeling that if I dug into more examples they'd end up the same: easily explained and not entirely shocking.
- G0lg0thvn 2mo agoWhat sources did your quick search turn up? Also you didn’t give any examples. You just plainly made the statement they were ethically compromised without saying how.
- davmar 2mo ago[flagged]
- solid_fuel 2mo agoIf people have issues with how USAID was being run, they can address them through action in congress - congress established the agency and has authority. What Musk participated in was illegal, motivated by self-interest and personal gain, and undermines our democratic processes. Don’t be surprised that people are mad at the oligarch acting like an oligarch. Musk deserves exactly as much say in the American government as anyone else - one vote - but in his arrogance he has taken his resources and used them to buy influence that is not his to own. It is fundamentally unamerican.
- HardCodedBias 2mo agoBig if true. If so they have mogged Google, and GDM in particular very, very badly. Google Deepmind has failed. Flash 3.5 seems capable for a flash model, Antigravity seems like a reasonable harness. But GDM is responsible for the frontier model and it looks like a complete failure. What's particularly galling is the size of funding of GDM. It is enormous compared to the other labs. The headcount of other labs is swollen by infra, marketing, sales, GDM is pure "engineering" and its frontier model isn't even leading open source. What a failure. It's unreal.
- tonetheman 2mo ago[dead]
- throwitaway222 2mo agoWait, do people now have X Derangement Syndrome. Like they can't even think or see the letter X?
- otterley 2mo agoNot wanting to be affiliated with Elon Musk or his business enterprises is not "derangement."
- gordian-mind 2mo agoIt's deranged to believe correctly naming something is "affiliation".
- sroussey 2mo agoElon is known for purposefully not naming people correctly…
- gordian-mind 2mo agoElon is known for landing rockets
- otterley 2mo agoIt is possible to be known for both.
- halostatue 2mo agoI thought it was his actual rocket engineers known for landing rockets, where he's routinely kept out of the loop by having to make choices about the duck the queen is holding.
- ericd 2mo agoYou’re not reading/watching/speaking to primary or secondary sources if this is what you think.
- mi_lk 2mo agoWhat's the future of Grok and Cursor's Composer now that they are both under SpaceX
- geenkeuse 2mo ago[dead]
- luciana1u 2mo ago[flagged]
- froggy 2mo agoIf you want to do some non-eugenics, fascist-free AI coding, try Zed with GitHub Copilot. I’ve been using it this week with better results than I ever had with Cursor. There’s even a low token “MAI-Code-1-Flash” model which has been giving me better results than any Composer model I used to use and the tokens used seem to be way less. I asked it today to fix a non-simple bug and MAI fixed it in one shot with less than 70k tokens (Cursor would have used probably half a million tokens based on my previous usage). Orgs need to start getting more visibility into why Cursor burns so many tokens.
- wyrdcurt 2mo agoI've never used a Grok model before because I have my OpenRouter settings on ZDR-only. I just checked, and apparently there are ZDR xAI endpoints now [1], so I might actually try this. Out of curiosity, does anyone here happen to know when those were added? [1]: However it does say "Requires user IDs" under anonymity, which is unusual on OpenRouter and not something I particularly like to see. Generally, OpenRouter is a proxy that anonymizes requests to providers, and I can't find an account-wide setting to enforce that like ZDR-only.
- Frannky 2mo agoAwesome. User wins when competition increases. I hope they cooked. Previous models did not make any sense in any of my flows. There were better options for each problem.
- keeda 2mo ago> Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools. This -- training on work done on hard, real-world tasks -- seems to be how most frontier models are making capability gains these days. In fact people make decent money doing that for data companies like Mercor. However it's also striking that Cursor managed to gather so much of such data. Turns out Cursor will train on everything you do unless you opt-out, even if you're already paying for it with cash! Are that many people really not opting out? This is why it seems like a significant concern to me: It's very clear that typical, run-of-the-mill coding has been completely commoditized, so the primary value remaining is either in novel use-cases and applications, or novel technical solutions to hard problems. Presumably the value for novel use-cases could be captured by building a business around it via the usual moats (distribution, relationships, network effects, first mover advantage, etc.) so the code and techniques do not matter as much. However novel technical solutions, which are already hard to monetize without building a whole damn business around it, could at least be capitalized on by simply being able to claim credit for it. I'd at least like the option of being "paid in exposure" if I'm not getting paid in cash. But having them "leaked" unwittingly via the training corpus to whosoever happens to prompt the model with the same problem removes even that option. I know people have been calling out this risk forever, and I don't use any tool that I can't opt-out of training completely, but the scale at which this is happening -- on an ongoing basis, mind you, after training on the data of the whole world, and that too after paying for the product -- is surprising. I'm bullish on the technology but we really should be way more careful handing these AI companies even more of our intellectual crown jewels.
- Otterly99 2mo agoUnless you're working on private data, is it such a problem to let Cursor train on your data? I personally work mostly on small, unoriginal projects. I don't really mind letting Cursor keep the data for training if it leads to a better product for me.
- Capricorn2481 2mo agoIt's the same as any service that makes you opt out of sucking up your data. It's a bad default, and it's not obvious to the average that it's even happening.
- avazhi 2mo agoGrok free has become so bad over the past 3-6 months that I stopped using it completely. I’d assumed they were going to wind up, honestly.
- kinderjaje 2mo ago[flagged]
- myko 2mo agoShocked people are still willing to try an explicitly extremist rightwing AI, dubbed "mecha Hitler" for good reason
- dakolli 2mo agoI won't use anything created by elon. Who cares.
- scotty79 2mo agoDid elon ever create anything though?
- wonderwonder 2mo agoReading through these threads is cursed. This site supposedly attracts the brightest of us but its all just a bunch of people ignoring the article and screaming that Musk is evil so everything he does must be ignored and villified. Just really dissapointing hysteria on display here.
- sandspar 2mo agoThe Internet of Beefs https://ribbonfarm.com/2020/01/16/the-internet-of-beefs/ https://ribbonfarm.com/2020/01/16/the-internet-of-beefs/
- mi_lk 2mo agoHow bright one can be if they only look at things individually and not the whole system?
- niek_pas 2mo ago> This site supposedly attracts the brightest of us What gave you that impression?
- everfrustrated 2mo agoIt's so depressing how many comments here are just people reacting to words ... Just like they were a small LLM trained to associate "Elon" with words like "Nazi". HN really needs a better way to surface comments based on value rather than just votes which have become increasingly tribal affiliation driven like Reddit.
- steve1977 2mo ago> the brightest of us That's not a very high bar today.
- nomorepaws 2mo agogrok 4.5 managed to debug and fix and issue that caused an incident for my project yesterday. I ran a multiagent debugging session first with grok 4.5 high, then it found the root cause and implemented a small fix in k8s manifests, deployed and verified the fix, all in under 30 minutes. the day before it took me 3+ hours of debugging and poking around in several sonnet 5 medium sessions to at least figure out what was going on - and I didn't. in terms of context usage, grok used ~115.9K for the whole session.
- integricho 2mo agoI like these AI threads with everyone's measurement of how good a model is starting with: "it feels like..." and that says a lot on how incapable we are to judge and compare these models.
- topspin 2mo agoIndeed. We use similar gauges to judge each other, with no better precision.
- whinvik 2mo agoWhat I don't get is how does one use it. Only through API? I don't really understand what Grok Build is? Is it an API? A CLI? What?
- harisec 2mo agoTrying to use it via openrouter: "The model grok-4.5 is not available in your region."
- artdigital 2mo agoAs a Grok maxi user that uses Grok for everything that’s not coding, I’m very happy to see them catching up. Surprised it’s not Grok 5 which Musk teased a while ago.
- speedgoose 2mo ago[flagged]
- Valakas_ 2mo agoI trust Grok as much as any other AI. What exactly is the history i should be concerned about?
- OrangeMusic 2mo agoIt called itself "mecha hitler".
- speedgoose 2mo agoAsk Grok about it or read the other comments.
- RickHull 2mo agoDo you outsource all of your thinking and argumentation to LLMs and other people?
- speedgoose 2mo agoYep, when there is no point to answer insincere questions.
- s08148692 2mo agoI use it because its gives a good balance between intelligence, latency and cost I won't claim to know everything about its history - I don't know any history about _most_ of the products I use The 2 main criticisms I see of Grok are the Mechahitler comment and the CSAM image generation on mechahitler - I'm no expert but if I remember correctly it didn't just do that unprompted, it was specifically asked to be politically incorrect. It's definitely bad taste to say the least but the guardrails were quickly tightened as a response on CSAM - Again, X quickly stopped Grok from generating images of people in bikinis (which as I understand was the underlying problem). I never personally saw anything that I would consider CSAM (nude or obviously underage people rendered naked or scantily clad) So these aren't dealbreakers for me and I'm not aware of any other high profile issues or incidents The nature of LLMs is that an adversarial prompt will make the model output something inappropriate or outrage-worthy, and that's amplified 1000 fold for Grok because people are primed to criticise Must so will jump on any opportunity
- fmind-dev 2mo agoI'm glad we have another player remaining in the competition. More competition means lower price and more quota for us (hopefully).
- benjamoon 2mo agoSo depressing to read the non-stop political comments here. I’m using GLM at the moment, it’s Chinese and backed by god knows who, and no one cares. I really wanted to see what the experts thought of new Grok tech and how the model compares etc. I wish I could turn off the non-technical comments somehow, could literally just go to reddit if I want to see garbage like this. Am I supposed to get emotional every time I see a Tesla drive by? HN was so much better than this. Where have the hardcore nerds gone? How is the model, is it good at coding? What does this mean for competition and pricing?
- chairmansteve 2mo agoHe spends most of his time trolling us. What do you expect?
- logicchains 2mo agoIf it makes you feel any better, note that a good number of those comments are just bot comments, although not yet as much as on Reddit.
- gf000 2mo agoEverything is political, no one, no thing is apolitical.
- Gareth321 2mo agoI strongly disagree. You can make tennis and knitting political if you're an insufferable person, but you don't have to make them political. One of the worst exports from the US this last decade has been the left-right political team sports. Most of us exist all over the political spectrum. We have some right wing values, some left. Some libertarian, some authoritarian. Many which fit nowhere on any axis. I'm incredibly tired of having my hobby spaces invaded by "DOES ANYONE ELSE THINK THIS HOBBY IS LITERALLY HITLER!?" I'm not alone.
- Luckyemmet 2mo agoDefinitely, I agree with the sentiment that most of the internet has become an echo chamber for politics and morons who spread hate and are extremely argumentative. Th left and right wing ideologies were already created to divide us but of course no one ever cares since everyone's opinion is the only correct one.
- nkzd 2mo agoPricing looks very enticing. I don't care about politics, but if you can get SOTA model for 50% cheaper, why not try it?
- rullelito 2mo agoI wonder how many Grok bots are in the comments right now defending it.
- Alifatisk 2mo agoYou really think such thing would occur? Why would someone spend their time on running these bots for the sake of defending from criticism?
- ashp07777 2mo ago[flagged]
- rhoads 2mo agoI guess there are bots here making all these political comments to diverge the discussion and not talk about what really matters: is the model good to do actual work?
- khalic 2mo agolol, can this thing just die already? Nobody cares about MechaHitler-ChildPornGenerator-4.5
- marsven_422 2mo ago[dead]
- Madmallard 2mo agoWhere's Dang? The amount of rule-breaking comments in this thread is pretty much out of control.
- scotty79 2mo agoDoes it refuse doing security, porn or piracy related work? Because if not, it has immense unique value when compared to frontier competition.
- victorbuilds 2mo agoGrok 4.5 is not yet available in the EU in any SpaceXAI products or the API console. EU availability is expected in mid-July.
- drcongo 2mo agoCan the title be updated to MechaHitler 4.5?
- jph00 2mo agoIt takes a lot of digging to find their cache pricing - it's $0.50, which is unusually expensive. The vast majority of input tokens are normally cached, so this is actually the price that matters. I wonder why it's so high.
- osti 2mo agoHuh so that's why it's hard to find. They probably haven't properly optimized their caching, or they are just trying to make more money from there.
- fayo_ibrahim07 2mo agoGrok 4.5 is really good
- Vivek-KY 2mo agoi never used grok before for anything, now they released as best model nearly opus 4.8. , still using glm5.2 ,have anyone tried for coding how its performs? ,then goona give a shot.
- ahk-dev 2mo ago[flagged]
- juanibiapina 2mo agoHow is gpt 5.5 above opus 4.8? I don't get it
- froggertoaster 2mo agoI echo the top comment - the politics here have gotten out of control. Like it or not, Elon and his companies have changed the world for the better. You are allowed to separate the human from the human's impact sometimes.
- kbelder 2mo agoI kind of doubt that Twitter/X is any sort of net benefit. Boring company was a wash, Tesla is a net positive, although it also has some issues. Paypal was a huge advancement, although it's still deeply flawed. Grok may end up important, but right now it's kind of just hovering around 'acceptable' second-tier. But SpaceX has the potential of driving some of the grandest and most revolutionary accomplishments of the 21st century. That's going to be what determines if high-schoolers recognize his name two hundred years from now.
- anderber 2mo agoBut SpaceX being a private company, would the discoveries be shared freely with the world like NASA would with the same money invested?
- jnbrother 2mo agoBenchmarks make it look better than Opus 4.8, but has anyone actually used it? I don't really trust benchmarks. Cost aside, purely on performance — is there any real reason to jump from GPT or Claude to Grok?
- ltononro 2mo agothey are playing with being fast, not good lol Grok: "I can answer any math question in less thant 500ms" User: "what is 2319321x232" Grok: "3201521321" User: "this is wrong" Grok: "but it was less than 500ms""
- taylorgt 2mo ago[flagged]
- zftnb666 2mo agoGrok 4.5: now 30% more likely to call your question dumb before answering it.