Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pimeys
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
pimeys
6d ago
A colleague of mine has a strategy game to compare language models, 4.1 scores pretty high in this: https://clankerbattle.com/
2.
▲
by
pimeys
6d ago
It's more common than you think. I work in a startup and we pay API prices too. And we cut a lot of money by switching from Anthropic models to Kimi K3.
3.
▲
by
pimeys
7d ago
Deepseek also burns a lot of tokens, its output on high is 2x of Gemini on medium. But it's dirt-cheap so it still can be 60-70% cheaper. From the large models Kimi K3 is definitely the one burning the smallest amount of tokens. Even i
4.
▲
by
pimeys
7d ago
What you want is a bunch of sessions to replay. Something anonymized if it's not yours, and something that's not depending on state. You replay all your sessions against your harness, and then store all logs all output, everything
5.
▲
by
pimeys
7d ago
- Which versions: 3.6 vs 3.7 vs. 3.8 for Gemini Flash, and v4 0731 for Deepseek v4 Flash, and GLM 5.3 Flash - Medium for Gemini, high for Deepseek. - Things like find information, then understand something about it, then send a slack messag
6.
▲
by
pimeys
7d ago
Well, it's much more than that. In general everybody's building agents now. You see these things that can help you to do things like adding things like OCR an appointment from a picture of a hand-written paper and add it to your c
7.
▲
by
pimeys
7d ago
All of these flash models have this. You have to build your harness so that it deals with it. Infinite loops are solved by having an error message that says what to do differently on failure, invalid tool calls are solved by making the tool
8.
▲
by
pimeys
8d ago
Yes. I'm working in the agent industry and my god are we excited on new versions of Chinese flash models. The direct competition is Gemini Flash, and these models are much better on agentic tasks with fraction of the task price compare
9.
▲
by
pimeys
9d ago
Yeah, and buying a house is still kind of out of reach with this income if you don't start paying your mortgage whey you're 20something...
10.
▲
by
pimeys
9d ago
And 180k in Germany, even in Berlin is pretty good. You are living a very good life with that salary.
11.
▲
by
pimeys
9d ago
They try to build them, but for example in Finland where there's cheap electricity and lot of interest to build them, the people started protesting on rising electricity prices and now the politicians are noticing this. Same in Denmark
12.
▲
by
pimeys
10d ago
Maybe then using an open weights model is a good way to hide your tracks...
13.
▲
by
pimeys
14d ago
Super happy subscriber for years... One of those magazines that I open on a Sunday morning with a good cup of coffee and sleeping cats next to me before the family wakes up. If you enjoy well-written technical content, please subscribe and
14.
▲
by
pimeys
14d ago
It's interesting that Deepseek models were missing in the comparison. I see Deepseek v4 Flash a direct competitor to Gemini Flash for text-based agentic work.
15.
▲
by
pimeys
15d ago
ELI5 always works
16.
▲
by
pimeys
20d ago
Probably not very long. Multiple companies, including Amazon, started supporting Wero for payments. And most banks already support it. You scan a QR code or open a link that opens your bank app, show your fingerprint, see the amount and whi
17.
▲
by
pimeys
21d ago
I have been mainly using Kimi K3 on programming work for over a month now. It is so far the only language model that does not piss me off all the time and can deliver my daily tasks without any trouble. It does not talk annoyingly to me, it
18.
▲
by
pimeys
21d ago
Wait, I don't sympathize Apple at all... Or any other American corporation.
19.
▲
by
pimeys
25d ago
I've been using all the SOTA models a lot at work, like serious amount of tokens. It's been really rare that I stick with one model and harness for too long... Except a month ago I started testing Kimi K3 and omp and I never went
20.
▲
by
pimeys
29d ago
I just want to but hardware so I can run a model at home that is fast. I don't see myself installing a server that burns almost two hundred kilowatts but maybe a card which runs a 27B Qwen...
21.
▲
by
pimeys
29d ago
Yes and no. It competes in the mid tear not in SOTA. It's a very valid model if you need things like computer use or image recognition. Especially with the 3.7 "introductory prices". It's multi-modal and better than GPT
22.
▲
by
pimeys
29d ago
Wait, there's 13 providers for Kimi K3 in OpenRouter. I'm having a hard time believing every single one of them provides them without any profit. And this one is easy to calculate: take your monthly API spend to K3, then rent a st
23.
▲
by
pimeys
29d ago
They did now. Landing somewhere between Terra and Luna now per task, with the quality of Gemini 3.7 flash.
24.
▲
by
pimeys
29d ago
Oh we did try to get Grok for evals but they had some weird EU limitations last time we checked. Which the open weight models don't have.
25.
▲
by
pimeys
1mo ago
Yep. It's a bit scary also. There's a lot of opportunity in the market now, but the downfall of the big US inference labs is going to hurt here in EU too sadly...
26.
▲
by
pimeys
1mo ago
It's interesting if they need to cut off their subscriptions to be able to compete in API prices. Very interesting...
27.
▲
by
pimeys
1mo ago
They just raised their prices sadly.
28.
▲
by
pimeys
1mo ago
Internal reports from company? Maybe not. I'm just saying you have to eval eval eval if you are working in this industry. There's a ton of victories in price, and price is right now the key thing all the customers are talking abou
29.
▲
by
pimeys
1mo ago
Long-context agentic tasks and Rust engineering are our use cases where Kimi definitely is better than Sol. We can measure our own systems and the numbers say that Sol has no chance against K3 or Opus, and K3 is so so so much cheaper than O
30.
▲
by
pimeys
1mo ago
And Fireworks did not yet. They are still under the limit of not feasible to self host... Let's see if other providers follow DeepSeek with their flash pricing.
More ›