9 ms·
Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
- jartan2002 2mo ago[flagged]
- jingpostmedia 2mo ago[flagged]
- veber-alex 2mo agoAI slop
- IshKebab 2mo agoYeah I think they at least put some light effort into making it readable though. Obviously slop but not quite as bad as most slop articles.
- inferencecoder 2mo agoThere is barely any effort, a simple GPT5.6 sol pro query rips the post apart.
- logicallee 2mo agoif you did that in the web interface, could you share the chat? I'd be interested to read it.
- inferencecoder 2mo agohttps://chatgpt.com/share/6a6f09ff-2830-83ea-9578-de3016cfae1f https://chatgpt.com/share/6a6f09ff-2830-83ea-9578-de3016cfae...
- logicallee 2mo agothanks for sharing. When I read your original comment, I was thinking you had just asked it to evaluate the article. (Like just "evaluate this article" or something.) I don't think anything anyone (or any AI) has ever written or published (including Sol itself) wouldn't be torn apart by the prompt you gave though.
- inferencecoder 2mo agoI think it's fair to expect an extensive review of an article before publishing. Not everything has a set of serious flaws.
- IshKebab 2mo agoIt doesn't look like it has found a set of serious flaws to me. Just nitpicking.
- inferencecoder 2mo agoTo begin with, as I pointed out separately, the price comparison that makes the headline is absurd. No one is easily getting the MI355X's at $2.5/hr. And then the eval is unreplicable and poorly defined.
- logicallee 2mo ago>Not everything has a set of serious flaws. I'm saying everything looks like it does, with the prompt you gave Sol. Go ahead and point the same prompt to anything you don't think has serious flaws and you'll see it would tear it apart.
- stevenarellno 1mo agothis got 82% human written on pangram :(. why the ai slop allegations? how can we make this better in the future? (i work @ wafer & witnessed ian writing this)
- logicallee 2mo agoThis part sounds like AI assisted setting this up and benchmarking it: >The fix was trivially simple: zero-pad the head count 12→16, run the fast kernel, and extract the real 12 heads from the output. I've recently used a frontier AI (ChatGPT 5.6 Sol on ultra) to set up a much smaller local model, and the performance optimizations it introduced left the model totally incoherent. (The model just repeats a single character, etc.) When I see a line like the one I just quoted, it leaves me wondering if the setup is still coherent like a stock install of Kimi K3 on supported hardware. Did they run any benchmarks on it to see if it is still correct?
- springtimesun 2mo agoI have never used sol, but I regularly use Claude to set up and benchmark local models per task. It is very thorough and has always returned good setups. The only thing you have to do is make sure you point at the model card. It’s always incredulous that models exist after its training cutoff. K3 does an ok job of setting up, but its config searching isn’t nearly as thorough and its will confidently tell you it’s found the best setup when it’s only turned a few knobs. It’s also not a good evaluator of its own output. It rates its work too highly and seems kind of defensive when benchmarking. Still a good check because it does find stuff, but open a fresh session and don’t tell it where the results came from. What K3 does do more than any other model I’ve found is investigate folder structures. If I want to benchmark it and other models I have to move the testing methodology and any reference to other results out of the project folder bc Kimi is like an ls bloodhound. It will find them.
- serf 2mo ago>and its will confidently tell you it’s found the best setup when it’s only turned a few knobs. on one hand , yeah : a model being more aggressive towards exploration of the decision space is usually a good thing. on the other hand : I think that it's up to the operator to set rigid test criteria to make these things actually work well in a repeatable fashion. so in other words, i'm glad claude is doing a better job for the way you're prompting the thing, but as a spectator from afar these kind of operator complaints usually spring up from the use of weak, ambiguous, under-considered prompts. similar with the parents' complaint; there should have been a testing criteria for legible output that got immediately flagged or failed by the larger model. it's a very hard sell for me to think that " a a a a a a a " is accepted as a valid language output by any near-SOTA-large model without some real coercion.
- jpgvm 2mo agoIf you do good work you should at least take the time to review the slop that details that work for slopiness. Otherwise it's hard to take it seriously. Especially the prefill section.
- muragekibicho 2mo ago[flagged]
- villgax 2mo agoLol, such a lazily written article by wafer.ai GPUs. 8× MI355X (TP8) B300 (TP8+DCP8) Decode tok/s per stream 118 tok/s 172 tok/s Peak aggregate. 952 tok/s 1,568 tok/s Peak aggregate per GPU 119 tok/s 196 tok/s On every row the B300 beat the MI355X The B200 is being forcefully compared against something which is not gonna fit within it's memory in a single node & not much details about multi-node interconnectivity, disagg or not. As expected of a shoddy slop. The only point it won is of cost per hour is one aggregation website for rentals, the premium a B300 commands against the $3/hr AMD chip which no provider has in abundance. Never bothered to do TCO of owning the hardware either.
- throwa356262 2mo agoDid you see this section? To the B200’s defence, its numbers are somewhat deflated by the fact that it pays a cross-node all-reduce on the decode critical path (RoCE v2 at ~195 Gb/s) — it’s the only config here that spans two nodes
- villgax 2mo agoNo mention of Infiniband or SPX lol
- inferencecoder 2mo agoWafer is making themselves synonymous with slop in the inference space. Exaggerated unfair comparisons in all their results, twitter hype posts with alarm emojis etc. > $2.50/GPU-hr for the MI355X, $6.00 for the B300, and $4.25 for the B200. This is not an accurate price comparison for real terms.
- villgax 2mo agoI went the gpus.io website & it’s $2.95/hr right now, this is like comparing MSRP to actual market price. B300s are in demand & hence cost more, but these lazy editors at wafer.ai can't be bothered to do TCO on actual ownership nor share code to replicate their setups. Instead just relying on current market prices to win one row, which isnt even about per/$ on actual MSRPs.
- YetAnotherNick 2mo agogpus.io shows tensorweave pricing at $2.95/hr. Tensorweave just shows "Talk to sales".
- greyb 2mo agoThis is a company that launched a token subscription (WaferPass), before weeks later, rugpulling the plan for being unsustainable while simultaneously claiming they achieved incredible inference efficiency gains worthy of paying them mind.
- BoorishBears 2mo agoMost discourse around GPU prices is nonsense right now. Some people using Spot prices for providers who won't have Spot capacity, some people using hourly rates for instances that are never in stock, some people ignoring commitment discounts. Not to mention no one serious is serving this on 8xB200 instead of multiple nodes: the vast majority of Moonshot's inference work is focused on PD-disaggregation
- inferencecoder 2mo ago> Not to mention no one serious is serving this on 8xB200 instead of multiple nodes: the vast majority of Moonshot's inference work is focused on PD-disaggregation The GPU price discourse is absurd, but many are serving models on single node setups when the model fits
- kelnos 2mo agoI wish they wouldn't call them "open source models". They aren't open source. They didn't publish the training data. They didn't publish the tools they used to train the model. They published the weights. It's an "open weight model", a term that it seems nearly everyone has agreed is appropriate. Why is this company not using it?
- charcircuit 2mo agoThe weights are the preferred form for modifying or integrating with other models. There is no obligation in open source to transitively open source all of the documentation / tools used to create the open source project. >a term that it seems nearly everyone has agreed is appropriate Models being considered open source even if the original training code / data is not released also is something almost everyone has agreed to be appropriate.
- teruakohatu 2mo ago> There is no obligation in open source to transitively open source all of the documentation / tools used to create the open source project. Open source means open source code. Open weight means a binary file dump, not unlike an exe file. There is nothing open source about it. Its like having a closed source text editor that censors certain words, and an open source text editor that censors certain words. The latter can easily be recompiled, the former requires reverse engineering. Both may give you a license to use them freely.
- Alpha3031 2mo agoPer the OSD definition of source code, "the source code must be the preferred form in which a programmer would modify the program." which means an argument could be made (as charcircuit is making) that the weights, being the preferred form to modify, are the source. I do prefer open weights as being more precise (like, is it even really software that has source code in the first place?) but I feel like at this point the ship has sailed somewhat (though if this is something you're willing to spend your time arguing then... moral support I guess?)
- BookPage 2mo agoPeople complaining about the slop - what about the atrocious text/bg contrast? Burning my eyes out faster than a B300 ever could
- GuestFAUniverse 2mo agoHow is the capex on 8 * MI354X even remotely justified at less than $10/h? Even without the base system, power and every other expenses: 365d * 24h * $2.95 = $25842/a invoicable. That doesn't add up within one year, that doesn't add up in three years and it is questionable that it brings in the money during the lifetime of the device?
- arjie 2mo agoThat can’t be real. Modal rents out an RTX 6000 Pro for more than that. Nvidia will rent your GPUs at a fixed cheap rate if you buy from them. Perhaps AMD has a different subsidy style program. Because those MI355X are going cheap here.
- gpugreg 2mo agoWhere do you see less than $10/h for 8 * MI354X? I can only find $2.50 for 1 * MI355X (lowest I can find for rent on other websites is $2.65, but maybe they got a better deal).
- swiftcoder 2mo ago> How is the capex on 8 * MI354X even remotely justified at less than $10/h? Is anyone actually renting them out that cheap? The very cheapest on-demand price I see online is $14, and most providers are a lot higher
- bensyverson 2mo agoHow does one even purchase an AMD MI355X?
- hereme888 2mo agoAnother ad posing as "science". It's basically a Wafer/AMD advertisement. B300 is about 46% faster for one stream and 65% faster in aggregate. AMD wins only after Wafer divides throughput // selected cloud-rental prices: $2.50/GPU-hour for MI355X versus $6 for B300. Benchmarks are unreproducible, power costs are missing, ROCm was patched... and on and on. Sure, it can serve that particular model at that particular size more economically. Good for them... in particular.
- fancyfredbot 2mo agoSince we're (as a society) spending tens of billions of dollars to serve models like this, it's a pretty interesting advertisement.
- venkat_2811 1mo agowhile i'm a fan of wafer's work, 1024 input token length is not a great benchmark anymore, that number is useful only at 1-4x single node a100 / h100 at various concurrency levels
- deleted 1mo ago[deleted]
- deleted 1mo ago[deleted]