Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mordae
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
mordae
3d ago
Why would China want to slow down when it is winning?
2.
▲
by
mordae
3d ago
You are basically saying we should preemptively imprison people who are too smart, right? If the intelligence is too great, we have to contain it. I disagree.
3.
▲
by
mordae
3d ago
State is not an actor. It is part of the playing field.
4.
▲
by
mordae
3d ago
Only in how they are deployed, not in principle.
5.
▲
by
mordae
4d ago
Remember that one time when music CDs from Sony MBG installed a rootkit on your PC that continuously sent whatever MP3s you had on your PC back to them?
6.
▲
by
mordae
6d ago
Lobby Ursula to fund an EU-CN collaboration lab and buy Ascends?
7.
▲
by
mordae
6d ago
Since it has low activated parameter count but huge total parameter count it needs more tokens to move the relevant information into the context.
8.
▲
by
mordae
9d ago
Sure. Even more so the handful of patent holders in the US. Want to ignore those patents? Unless you have nukes you're not.
9.
▲
by
mordae
22d ago
It is already €20+ to send something from e.g. Czechia to France.
10.
▲
by
mordae
22d ago
It is not. Big European business always get what they want. Which tends to be maintaining status quo.
11.
▲
by
mordae
23d ago
I've overheard a guy last week in town saying they brought their iPhone to a repair shop and they've done the repair at fraction of the cost they expected. They were excited and told their friends. Washing machine repairmen never
12.
▲
by
mordae
29d ago
Except staying on the edge costs exactly the same as not, when you take the resell value of the components on the second hand market into account. The problem is that some can afford to lock their capital in the hardware (via one-time inves
13.
▲
by
mordae
1mo ago
I think that in this case there is also the problem of trying to transfer MoE-style reasoning into a dense model. I mean, MoE needs reasoning to walk multiple experts, but dense model already has all the weights. So when you push it hard to
14.
▲
by
mordae
1mo ago
It is MoE. It needs to engage multiple experts when the problem is complex or unclear. So you naturally see more of those simply as a primitive it learns to use to page in more diverse set of weights. Remember that each token is just 6 expe
15.
▲
by
mordae
1mo ago
It needs to argue with itself to extract most of the knowledge embedded in the weights into the context. Asking it to synthesize ideas directly in a single go is simply unreasonable. And MoE models need to walk multiple experts to extract a
16.
▲
by
mordae
1mo ago
Yeah, it should be basically free. No idea why it is not. I guess KV cache taking up RAM and possibly bad business sense or amortized engineering costs, I honestly do not know.
17.
▲
by
mordae
1mo ago
DeepSeek V4 Flash is natively FP4 MoE with very compact KV cache. Say 8 GB/s. Qwen 27B is about 60 GB/s at full FP16 precision.
18.
▲
by
mordae
1mo ago
0731 is definitely tuned for coding. I mean: https://gertlabs.com/rankings?mode=agentic_coding But it is also a decent translator from English to Czech in my experience.
19.
▲
by
mordae
1mo ago
Of course. On the other hand if you see 8 % gap for similar workloads, averaged across tens of sessions, with the same underlying model, it becomes a pretty clear signal. And I do exclude first request per provider per session from the stat
20.
▲
by
mordae
1mo ago
Yeah, I've noticed them recently on OR and whitelisted. Then I backed-off pretty quickly after seeing the cache hit rates. It was also rather revealing to see how some provider hit rates differ when you are using them directly vs via O
21.
▲
by
mordae
2mo ago
I was just using it when it landed. It started reasoning more extensively from nowhere and precision went up a lot. It also changed its prose style for the better. Looking forward to weights.
22.
▲
by
mordae
2mo ago
To run at decent speed, all models try hard to use only most likely relevant part of the context and most likely relevant weights (MoE) to predict the next token. Doing the math in full is unfeasible.
23.
▲
by
mordae
2mo ago
Why would anyone think that models optimized for efficient context management, giving much more weight to a short sliding window, would attend to distant, heavily diluted tokens? Plus the model's capacity to take more context into acco
24.
▲
by
mordae
2mo ago
Please also note that ADHD-like symptoms can present from sleep Apnea/UARS-induced sleep deprivation and the probability of having Apnea/UARS rises with age. Smoothly for men and a with a sudden jump after menopause for women. So
25.
▲
by
mordae
2mo ago
Sleep apnea symptoms are initially similar to ADHD. Apparently a ton of kids gets mis-diagnosed with ADHD while they are slowly suffocating because they have UARS. Stimulants mask it.
26.
▲
by
mordae
2mo ago
Chance is you are not breathing in your sleep. You have Apnea or UARS and are slowly suffocating. Easy tells: you look like shit in the morning and it takes hours for you to normalize, you sometimes wake up all sweaty, you tend to open wind
27.
▲
by
mordae
3mo ago
> alleged anti-competitive practices I'd say that with court ruling these are no longer alleged. Right?
28.
▲
by
mordae
3mo ago
Nope, people seek it out because government tells them to pay taxes _or else_.
29.
▲
by
mordae
3mo ago
They do not and it sucks for certain tasks. It also means that if they actually trained with vision, they'd be on par with Anthropic models as vision seems to improve model performance across the board even for non-vision tasks.
30.
▲
by
mordae
3mo ago
DeepSeek V4 Flash on default (high) reasoning had zero formatting failures, just genuine losses. Total score: | DeepSeek V4 Flash | 97/100 | 41/50 | 41/50 | 92W / 102T / 6L | Benchmark costed me just under $1.
More ›