Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
redman25
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
redman25
13d ago
Maybe they’re gunning for speedy non-interactive pricing? Or its a limit of the technology or a business decision?
2.
▲
FrontierHarness Eval benchmark. Pi is on the Pareto
(runta.com)
3 points
by
redman25
14d ago
|
0 comments
3.
▲
by
redman25
1mo ago
More like acting without a sense of civic duty is stupid. Or acting without a basis in evidence is stupid.
4.
▲
by
redman25
2mo ago
Deepseek is one of the worst in terms of hallucination rate according to artificial analysis' benchmark: https://artificialanalysis.ai/?omniscience=omniscience-hallu...
5.
▲
by
redman25
2mo ago
A large portion of the tests are closed source unfortunately which would make it tough to create a port.
6.
▲
by
redman25
3mo ago
I've been preferring Mimo recently. Same price as deekseek, more reliable tool calling (subjectively), and has some nice qualities in terms of prose, etc. I've heard others say that Deepseek tends to be smarter on specific problem
7.
▲
by
redman25
3mo ago
Exactly, intelligence is limited by cost and physical constraints just as much as anything. That's the thing that seems to always be missing from the run-away singularity discussions, it's treated like a perpetual motion machine.
8.
▲
by
redman25
4mo ago
IDK this model release is a bit disappointing considering the community has been chomping at the bit for the 124ba4b model. There was some leaked info about it but people suspect it was not released because it was too close to gemini flash
9.
▲
by
redman25
4mo ago
I too prefer my misinformation delivered with maximum confidence and no accountability.
10.
▲
by
redman25
4mo ago
What prompt had you given it?
11.
▲
by
redman25
4mo ago
I'm not OP but I work outside and use light mode. Macs are generally fairly bright as long as you aren't in direct sunlight. Solarized light mode for the win though.
12.
▲
by
redman25
4mo ago
It’s a heck of a lot faster too.
13.
▲
by
redman25
5mo ago
Doesn’t ham radio not allow transmissions to be encrypted by law? That rules out most of the internet.
14.
▲
by
redman25
5mo ago
A strix halo machine or MAC will run at less than 20watts idle. You could leave it running.
15.
▲
by
redman25
5mo ago
Is 10 days enough to make walking difficult?
16.
▲
by
redman25
5mo ago
Not to be confused with nanocoder, the agentic coding harness. https://github.com/Nano-Collective/nanocoder
17.
▲
by
redman25
6mo ago
Exactly, compare MoE with MoE and dense with dense otherwise it's apples and oranges.
18.
▲
by
redman25
6mo ago
200a10b please, 200a3b is too little active to have good intelligence IMO and 10b is still reasonably fast.
19.
▲
Empiricism and Rationalism in Software Testing
(joshvoigts.com)
2 points
by
redman25
6mo ago
|
0 comments
20.
▲
by
redman25
6mo ago
It's why you always have a rollback plan. Every `up` needs to a `down`.
21.
▲
by
redman25
6mo ago
So someone is debugging something with git bisect and stumbles on the old commit and gets pwned. Maybe that's why they force killed it? To avoid people going back in history and stumbling on it.
22.
▲
by
redman25
6mo ago
CPU/network throttling needs to be set for the product manager and management - that's the only way you might see real change. We have some egregious slowness in our app that only shows up for our largest customers in production b
23.
▲
Olaf: Bringing an Animated Character to Life in the Physical World [video]
(youtube.com)
2 points
by
redman25
6mo ago
|
0 comments
24.
▲
by
redman25
6mo ago
It’s mainly the benchmarks that have encouraged that. The more tokens they crank out the more likely the answer is to be somewhere in the output.
25.
▲
by
redman25
6mo ago
I feel like the right response for those situations is to start asking questions of the user. It’s what a human would do if they did not understand.
26.
▲
by
redman25
6mo ago
It depends on the harness and/or inference engine whether they keep the reasoning of past messages. Not to get all philosophical but maybe justification is post-hoc even for humans.
27.
▲
by
redman25
6mo ago
How about a human coworker who screws up 1% of the time? Doesn’t sound so bad in that light. It’s the nature of being human. Good code review is the solution but if it’s faster to do it yourself, that’s fine too.
28.
▲
by
redman25
6mo ago
First things first, we have to colonize the rest of the solar system before we can terrorize Earth.
29.
▲
by
redman25
6mo ago
For consumers, there's little reason to run unquanted, especially for large models which take less of a hit from quantization. I'm running a 200b model at Q3 with very little degradation. A 1000b model would see even less change.
30.
▲
LocalCowork
(github.com)
1 points
by
redman25
6mo ago
|
0 comments
More ›