Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
disiplus
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
disiplus
7d ago
Sure, but they are replacing same generation model to another not just switching opus to gpt.
2.
▲
by
disiplus
21d ago
it was not the first model served for free, i remember grok and others beeing free on openrouter but they never had this popularity because they where not good enough.
3.
▲
by
disiplus
21d ago
i have heard that for max effort in flash and it can be true, but overall it still performs better then the high, i run a mixed q2q4 quant.
4.
▲
by
disiplus
21d ago
idk i think that i spend significant tokens with both to be able to tell 5.3 is way better overall. https://kommodo.ai/i/IFSUUQYT522uZXWvePZ1 it still fells stupid sometimes and it is benchmaxxed for sure. but its good
5.
▲
by
disiplus
21d ago
To be fair, there is no 3 turns that I don't have to jump in into what Opus 5 is doing. There is either some regression or my prompting skills are so much worse now. Flash is not perfect and honestly some things depend on how big conte
6.
▲
by
disiplus
21d ago
I will give it a try, but from the benchmarks it never exceeds the DS4 flash benchmarks by significant margin and And I feel that the throughput that you will get on those machines or what I'm getting with my local hosted flash will be
7.
▲
by
disiplus
1mo ago
i run flash v4 at 2bit, its pretty great and on my tests against full model It didn't lose any capabilities. It just was thinking more. So you don't have the same efficiency.
8.
▲
by
disiplus
1mo ago
You don't care what were you trying to build but for a quick me alone Linux app I used Flutter. Honestly it's my go-to when I want to have a quick native app on any platform but dont want to bundle electron.
9.
▲
by
disiplus
1mo ago
It's not that they don't care, it's just that the protocol behind it was never designed for it. Compared for example to email where when you want to send email to somebody, you discover its MX settings on the DNS level and th
10.
▲
by
disiplus
2mo ago
They all are pretty similar. And if it was copying anything, I would say more inspiration from claude desktop then codex. The Zcode is more like codex. https://drive.google.com/file/d/1JFfgfMO0nO7HR0WHwEqEEjXIIQj..
11.
▲
by
disiplus
2mo ago
I found this part funny. > People familiar with OpenCode internals (if you are on the OpenCode dev team I assume this doesn’t include you) might have objected to my python3 example above.
12.
▲
by
disiplus
2mo ago
Cache Misses are pretty bad, i have a locally running deepseek v4 flash, i have tuned it now to have 1100-1300 prefill. Its not great but properly useable. Imagine having a session with already 100k and half of it has to be prefilled it wou
13.
▲
by
disiplus
4mo ago
Which model are you running ?
14.
▲
by
disiplus
4mo ago
Diagnosed with ADHD, ultimately does not change anything for me even through i had the same idea as you. Reason is that i can now start even more stuff in parallel. And some part of them get finished more before i can just prompt more when
15.
▲
by
disiplus
4mo ago
same, but you need more then 100k of hw to run something like kimi k2.6 for a bigger team. on the other hand there is a ds4 flash that you can run on a macbook with 128gb ram. an that one is perfectly usable for a lot of tasks. https:/
16.
▲
by
disiplus
4mo ago
The problem is not website, the problem is discovery and discovery is on Instagram, TikTok, and social networks. You don't have any incentive to build a website for a regular audience. What you might do is build an audience on a social
17.
▲
by
disiplus
4mo ago
depends, a super small one finetuned to do function calling instead sending it to big model and waiting, instead, you ask for a revenue in last month, i do a small llm function call -> show results. some bigger ones, analysis, summary, c
18.
▲
by
disiplus
4mo ago
i dont know what are you talking about, i replaced an older gpt4o with a finetuned qwen. there is a huge amount of "AI, that can be done with those models, or partly by those models." Huge amount of people would not notice the dif
19.
▲
by
disiplus
4mo ago
nice, will run it later agains qwen3.6 27b, the speed was one of the reasons why in was running qwen and not gemma. the difference was big, there is some magic that happpens when you have more then 100tps.
20.
▲
by
disiplus
5mo ago
Depends how many users you have and what is "production grade" for you but like 500k gets you a 8x B200 machine.
21.
▲
by
disiplus
5mo ago
was part of the beta, its properly good model, in some sense i forgot that im not on opus or gpt. opus is still better. gpt is the one struggling for me. it has some niche in backend work but you can get the same with opus with skills, its
22.
▲
by
disiplus
5mo ago
It looks like its called prolite. https://snipboard.io/jmGKfM.jpg
23.
▲
by
disiplus
5mo ago
yet
24.
▲
by
disiplus
5mo ago
i have glm and kimi. kimi was in most of the cases better and my replacement for claude when i run out of tokens. Now im finding myself using glm more then kimi. Its funny that glm vs kimi, is like codex vs claude. Where glm and codex are b
25.
▲
by
disiplus
5mo ago
Yeah it seems they did not align it to much, at least for now. Yesterday it helped me bypass the bot detection on a local marketplace. that i wanted to scrap some listing for my personal alerting system. Al the others failed but glm5.1 foun
26.
▲
by
disiplus
5mo ago
basically my expirience as well. Sometimes it can break past 100k and be ok, but mostly it breaks down.
27.
▲
by
disiplus
5mo ago
When it works and its not slow it can impress. Like yesterday it solved something that kimi k2.5 could not. and kimi was best open source model for me. But it still slow sometimes. I have z.ai and kimi subscription when i run out of tokens
28.
▲
by
disiplus
6mo ago
The post mentions, france, germany and nordic nations. France, Holand and nordic nations helped in the early stages of US.
29.
▲
by
disiplus
7mo ago
It will also cost openai dearly if they don't communicate clearly, because I for one will internally push to switch from openai (we are on azure actually) to anthropic. Besides that my private account also.
30.
▲
by
disiplus
7mo ago
I have them all. They're not just as good. Whoever tells you that looked only at the benchmarks, not real use. They all fall short at some point. Kimi K2.5 is the best one, but it's still not at the level of what Anthropic release
More ›