Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stared
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
GPT-6 Astra solves puzzles
(quesma.com)
2 points
by
stared
2d ago
|
0 comments
2.
▲
Knowledge vs. wisdom: asking AI "What mushroom is that?"
(quesma.com)
2 points
by
stared
7d ago
|
0 comments
3.
▲
by
stared
7d ago
If you like, you can run these tests yourself as well. It is ~$500 per a single combination, and assuming no failed runs.
4.
▲
by
stared
7d ago
Nope. If you would like to do so, it is easy (and orders of magnitude cheaper) than running benchmarks.
5.
▲
by
stared
7d ago
Before this experiment I tried to run KLD on various context lengths to see a quantization-dependent deterioration. On wikitext2 there was no difference. I concluded these have no long-term dependency and I should use Linux kernel. Still, t
6.
▲
by
stared
8d ago
A single result is binary. All we get from a run is which tasks were solved, which weren’t.
7.
▲
by
stared
8d ago
Nice! Sometimes the simplest approaches work the best.
8.
▲
by
stared
8d ago
I am curious what's the actual formula. I mean, there so many headers and layers, it is tricky to make a choice that will resonate with our intuition . Is it some weighted average? Or maybe ablation test?
9.
▲
by
stared
8d ago
Point taken, but there is a much more fundamental issue with it - and precisely why I wrote "very conservative". It is a different problem if we pick two sets from the same data distribution, A and B, and first we have a score on
10.
▲
by
stared
8d ago
I wrote this blog post myself, with AI for proofreading (typos and grammar, but not style). There were a few singular sentences for which I had a writer's block, but not much besides that. So, if there are irrelevant remarks, these are
11.
▲
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
285 points
by
stared
8d ago
|
140 comments
12.
▲
by
stared
9d ago
Nice! One thing I am missing is an easy „go up” a taxonomy group.
13.
▲
by
stared
9d ago
Is it a subtle effect that cannot be explained by Newtonian gravity (i.e. a different potential affecting, V(z) in the Hamiltonian)?
14.
▲
by
stared
9d ago
Don’t ask. Start from a few blog posts and see traction, read feedback. You will also see hos long it takes - and what is thd difference between an idea and making it real.
15.
▲
I threw a ton of various logic puzzles at Astra and it solved of them
(twitter.com)
1 points
by
stared
10d ago
|
0 comments
16.
▲
GPT-6 Astra has autonomously completed Portal
(twitter.com)
8 points
by
stared
10d ago
|
0 comments
17.
▲
by
stared
10d ago
An interesting catch! Maybe it suffices to make this pre-title "the ai-design-slop fingerprint · defs 2026.09" lowercase. In any case, from what I see what is AI aesthetics on the website: pre-title and horizontal lines.
18.
▲
Does your landing page look generated?
(slop-detect.com)
3 points
by
stared
10d ago
|
3 comments
19.
▲
A connectomics milestone: Mapping the complete male fruit fly brain
(research.google)
2 points
by
stared
10d ago
|
0 comments
20.
▲
by
stared
10d ago
With Internet, we live in an always-on culture. It contrasts with times before, when we actually had to wait for a monthly magazine, and even if we wanted to watch a movie, it was aired (say) the next Tue 7PM. While now we have all convenie
21.
▲
Show HN: Genetic Distance Map
(p.migdal.pl)
2 points
by
stared
11d ago
|
1 comments
22.
▲
by
stared
11d ago
While I like this index, calling in "Intelligence" might be confusing - it is a mix of coding and knowledge. Compare and contrast with ARC-AGI, BabaIsBench ( https://quesma.com/benchmarks/babaisbench/ ), o
23.
▲
by
stared
12d ago
Compare and contrast with tests on Qwen3.6 27B: * how does quantization hurts pelican-drawing skills ( https://quesma.com/blog/qwen-quantization-quality/ ) * how does it hurt knowledge ( https://quesma.com
24.
▲
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
3 points
by
stared
12d ago
|
1 comments
25.
▲
Beef production in Brazil has been the largest driver of global deforestation
(ourworldindata.org)
5 points
by
stared
14d ago
|
0 comments
26.
▲
by
stared
14d ago
Repo for reproducible science: https://github.com/stared/mushroom-hunting-llm-bench
27.
▲
Mushroom hunting with LLMs: what can go wrong?
(quesma.com)
56 points
by
stared
14d ago
|
76 comments
28.
▲
Witcher 3's DLC Won't Include Russian Voice Acting
(thegamer.com)
5 points
by
stared
14d ago
|
3 comments
29.
▲
Forget the Pelican, It's Weevil-Time Benchmaxxing-Proof SVG and Vision Test
(reddit.com)
1 points
by
stared
15d ago
|
0 comments
30.
▲
by
stared
18d ago
These was some code budget there. But as from seeing various runs, errors bars are gross overestimation (as not "the same test", but "if we have different tasks from the same sample").
More ›