Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dhorthy
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
NoCode (2016)
(github.com)
4 points
by
dhorthy
16d ago
|
0 comments
2.
▲
Archil Persistent Sandboxes
(archil.com)
2 points
by
dhorthy
19d ago
|
0 comments
3.
▲
MicroGPT-C in pure C hits 10M TPS on Apple M5
(github.com)
141 points
by
dhorthy
29d ago
|
52 comments
4.
▲
by
dhorthy
1mo ago
> AI-written emails where the commenter would rather see the prompt this a thousand times this LLMs are information transformers. if you're trying to take some rough idea and blow it out into something that "feels" substan
5.
▲
by
dhorthy
1mo ago
they seem to be able to do the call-stack diff stuff pretty well still
6.
▲
Show HN: /show-me: agent skill for compact visual representations
(humanlayer.com)
12 points
by
dhorthy
1mo ago
|
4 comments
7.
▲
The Programming Ape (Code Hale, 2012) [video]
(youtube.com)
11 points
by
dhorthy
1mo ago
|
0 comments
8.
▲
Benchmarking Fable, Sol, and Kimi K3 on SlopCodeBench
(github.com)
8 points
by
dhorthy
1mo ago
|
1 comments
9.
▲
by
dhorthy
2mo ago
so wait is the finding that most of those skills reduce pass rates against SCB? wild
10.
▲
by
dhorthy
2mo ago
why would that worry you?
11.
▲
by
dhorthy
2mo ago
agree, i think the implication is that low quality code is harder to change in the future
12.
▲
by
dhorthy
2mo ago
somebody get this man a curl-pipe-bash stat
13.
▲
by
dhorthy
2mo ago
oh i really like the idea of flipping around the order of checkpoints and comparing results. Could be an interesting way to increase/decrease difficulty even i will look into how easy it would be to zip up some subset of the results wi
14.
▲
by
dhorthy
2mo ago
Yeah I would hold that models don’t know how to simplify because most rl/benchmarks doesn’t penalize complexity
15.
▲
by
dhorthy
2mo ago
Yeah the main reason I skipped fable was because we have a ZDR with anthropic and I didn’t feel like spinning up another account to circumvent that. Next run will have fable and sol
16.
▲
by
dhorthy
2mo ago
my issue with frontier code is that it uses a model judge for quality whereas slop code bench forces a model to grapple with its own garbage code in order to receive a functionality reward
17.
▲
by
dhorthy
2mo ago
i laughed at the pelican bit its good yes the labs will always prioritize the vibeslop dopamine casino as far as I can tell - making the models useful and addictive for unsophisticated users, sometimes at the expense or at the very least at
18.
▲
by
dhorthy
2mo ago
yeah someone will have to re-run this bench on various effort levels. unfortunately it is not cheap
19.
▲
by
dhorthy
2mo ago
I agree this is an option, and the next thing on my radar is to try with a more realistic "factory-shaped" harness where you have feedback from linters and other models after each coding episode that refines the architecture. For
20.
▲
by
dhorthy
2mo ago
> - what "maintainable" is is probably some high dimensional space described by these signals; it'd probably require some human labeling to figure out where this space is this is a nicely succinct way to put this - a multi
21.
▲
by
dhorthy
2mo ago
yes sol is still my daily driver for most coding tasks I did find opus 5 quite handy for general knowledge work and visual design, without the cost of fable (e.g. the graphics in this post are made by opus 5) but its not noticeably better t
22.
▲
by
dhorthy
2mo ago
no i'm spinning those up at some point this week. here's the first few prompts I used (claude opus 5 as the research orchestrator), (these were interspersed with lots of tools and assistant messages but it should get you kicked of
23.
▲
by
dhorthy
2mo ago
i hope that is because you hate slop and not because you write it
24.
▲
by
dhorthy
2mo ago
yeah this was just a start - the fastest cheapest thing we could try for a brand new model. I'm hoping to do some more work with sol/fable in the mix as well as exploring more languages and curating the problem set to include more
25.
▲
Benchmarking Opus 5 on SlopCodeBench
(github.com)
405 points
by
dhorthy
2mo ago
|
119 comments
26.
▲
by
dhorthy
2mo ago
the one good thing about the current ios version is that it is the best it will ever be from this point forward. all future versions will be worse
27.
▲
by
dhorthy
2mo ago
I agree this rocks I do this almost daily
28.
▲
by
dhorthy
2mo ago
yes well said
29.
▲
by
dhorthy
2mo ago
don't forget the andon cord
30.
▲
by
dhorthy
2mo ago
i humbly disagree - horizontal means touching one plane of the stack across, vertical means cutting down through it and touching multiple layers https://en.wikipedia.org/wiki/Vertical_slice
More ›