Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
galsapir
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Creative Reading: Scaffolding Reading for Transformation
(arxiv.org)
1 points
by
galsapir
6d ago
|
0 comments
2.
▲
No agents after 8 pm
(sparsethought.com)
2 points
by
galsapir
17d ago
|
0 comments
3.
▲
by
galsapir
28d ago
yep. i think it talked about a lot in the context of research, but not enough (or maybe im just not exposed to it) in relation to the arts.
4.
▲
AI Has Plunged the Book Publishing Industry into Utter Chaos
(wsj.com)
24 points
by
galsapir
28d ago
|
21 comments
5.
▲
Can AI agents conduct open-ended AI research?
(arxiv.org)
4 points
by
galsapir
1mo ago
|
0 comments
6.
▲
Ask HN: Any Good Pi.dev Setups?
3 points
by
galsapir
2mo ago
|
1 comments
7.
▲
by
galsapir
2mo ago
thanks for reading it properly and engaging with the argument! writing is hard, expressing ideas cleanly is harder! working on it.
8.
▲
by
galsapir
3mo ago
curious where the disagreement lands: the claim i'm least sure of myself is that measurement alone already counts as activation (nothing in the weights changes, so it's a looser sense of the word than usual) the part i'd defe
9.
▲
Giving a domain a hill to climb: benchmarking as data activation
(sparsethought.com)
12 points
by
galsapir
3mo ago
|
7 comments
10.
▲
by
galsapir
3mo ago
really interesting that its basically almost 80% claude opus..
11.
▲
by
galsapir
3mo ago
yeah its really counterintuitive i think; i.e, getting the right framework and structure for this to work probably isn't trivial, models really hate playing well together. i wonder how their version would fair in real world use.
12.
▲
A bitter lesson for medicine, or a benchmark problem?
(sparsethought.com)
2 points
by
galsapir
3mo ago
|
0 comments
13.
▲
by
galsapir
3mo ago
i feel like i've had exactly the same thought in the past :-0 might even have written about it. feel your pain
14.
▲
Can LLMs Beat Classical Hyperparameter Optimization Algorithms?
(arxiv.org)
120 points
by
galsapir
3mo ago
|
20 comments
15.
▲
Gemma 4 E4B as a primary local LLM (replaced Qwen)
(digg.com)
2 points
by
galsapir
3mo ago
|
0 comments
16.
▲
by
galsapir
3mo ago
sometimes I also feel it tries to optimise for "per line coverage" over more "real, complex use cases" type tests
17.
▲
by
galsapir
4mo ago
hey that's pretty cool. I think I still prefer "distill HN" cleanliness though. What made you create this.
18.
▲
PEEK: Give Your Agent an Orientation Cache (MIT CSAIL, Khattab group)
(zhuohangu.github.io)
3 points
by
galsapir
4mo ago
|
0 comments
19.
▲
by
galsapir
4mo ago
axon discharge is brilliant. adopting.
20.
▲
Hyperagents (Meta Research)
(arxiv.org)
2 points
by
galsapir
4mo ago
|
0 comments
21.
▲
by
galsapir
4mo ago
oh sorry! didn't catch the one Thanks, I'll comment there
22.
▲
The Unreasonable Effectiveness of HTML
(claude.com)
3 points
by
galsapir
4mo ago
|
2 comments
23.
▲
by
galsapir
4mo ago
From the link: "Shot from 90 perspectives, 88 focus stacked images each. Nikon Z8, full frame, f/7.1, exposure 1/160, ISO 100, Laowa 180mm macro lens, with LED light and bluescreen." Insane!
24.
▲
by
galsapir
4mo ago
I think the question he tried to raise was "is this needed? Aren't today's / tomorrow's models well-enough equipped to deal with just OPEN API?" (idk, just if I understand the question)
25.
▲
by
galsapir
4mo ago
got me at "Most often scientists believe they understand more than they do, making their belief an illusion." but why is it still bothering me? 1. feels unfalsifiable in spirit 2. somewhat restates "all models are wrong, b
26.
▲
The Comparator in Clinical AI
(sparsethought.com)
2 points
by
galsapir
4mo ago
|
1 comments
27.
▲
by
galsapir
4mo ago
author here. the part i'd actually like discussion on is the buried finding: physicians+GPT-4 didn't outperform GPT-4 alone on the management cases, and on the landmark cases the model alone beat the model+physician. the paper rep
28.
▲
by
galsapir
5mo ago
Hey thanks! I do wonder that. I think that even if specifically for code smell the things would be subtler, for other forms of AI driven averageness (especially in areas where we can't RLVR the models to perfection) it might still be p
29.
▲
by
galsapir
5mo ago
yeah I was really thinking about what the best "umbrella term" would be here. Since "LLM" is too widely used in a really specific context and "AI systems" felt niche I ended up with "LMs". Idk, up for
30.
▲
by
galsapir
5mo ago
haha that's a style choice (takes more work to get lowercase text these days). But yeah legit ;-)
More ›