Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sgk284
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
sgk284
2mo ago
I ran them for a very short time!
2.
▲
We've adding Inkling and 52 small apps one-shotted by it to our arena.
(arena.logic.inc)
3 points
by
sgk284
2mo ago
|
3 comments
3.
▲
by
sgk284
2mo ago
Hey HN, one of the side projects my startup maintains is this arena of 52 apps implemented by different models. Inkling is the most exciting model launch we've seen from an American lab in a while, so I spun up some NVidia B200's
4.
▲
by
sgk284
2mo ago
Yea, that's an interesting result as well. The Terra apps don't feel 35% less feature-rich. So it seems quite token efficient.
5.
▲
by
sgk284
2mo ago
Similarly, we updated our model arena (52 apps each built by 26 models) to have GPT 5.6 Sol, Terra, and Luna today: https://arena.logic.inc/ It's really interesting to see the Sol/Terra/Luna apps side-by-side
6.
▲
by
sgk284
2mo ago
Awesome - will work on getting those in.
7.
▲
by
sgk284
2mo ago
If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).
8.
▲
Show HN: Homecrew – Share agent skills across your team and keep them in sync
(crew.logic.inc)
2 points
by
sgk284
4mo ago
|
0 comments
9.
▲
Three Years of Predicting Football with LLMs
(bits.logic.inc)
2 points
by
sgk284
7mo ago
|
0 comments
10.
▲
On The Obsolescence of Interns in the Age of AI
(bits.logic.inc)
3 points
by
sgk284
9mo ago
|
0 comments
11.
▲
by
sgk284
9mo ago
Yep, 100% correct. We're still reviewing and advising on test cases. We also write a PRD beforehand (with the LLM interviewing us!) so the scope and expectations tend to be fairly well-defined.
12.
▲
by
sgk284
9mo ago
It doesn't require removing them if you think you'll need them. It just requires writing tests for those edge cases so you have confidence that the code will work correctly if/when those branches do eventually run. I don'
13.
▲
by
sgk284
9mo ago
FWIW all of the content on our eng blog is good ol' cage-free grass-fed human-written content. (If the analogy, in the first paragraph, of a Roomba dragging poop around the house didn't convince you)
14.
▲
by
sgk284
9mo ago
I suspect it will still fall on humans (with machine assistance?) to move the field forward and innovate, but in terms of training an LLM on genuinely new concepts, they tend to be pretty nimble on that front (in my experience). Especially
15.
▲
by
sgk284
9mo ago
I never claim that 100% coverage has anything to do with code breaking. The only claim made is that anything less than 100% does guarantee that some piece of code is not automatically exercised, which we don't allow. It's a foot
16.
▲
by
sgk284
9mo ago
Can you say more? I see a lot of teams struggling with getting AI to work for them. A lot of folks expect it to be a little more magical and "free" than it actually is. So this post is just me sharing what works well for us on a v
17.
▲
AI is forcing us to write good code
(bits.logic.inc)
302 points
by
sgk284
9mo ago
|
216 comments
18.
▲
Engineering Is Becoming Beekeeping
(bits.logic.inc)
2 points
by
sgk284
9mo ago
|
0 comments
19.
▲
Your Team Uses AI. Why Aren't You 10x Faster?
(bits.logic.inc)
5 points
by
sgk284
9mo ago
|
1 comments
20.
▲
Machine-Driven Code Review
(bits.logic.inc)
1 points
by
sgk284
9mo ago
|
0 comments
21.
▲
by
sgk284
9mo ago
Reranking is definitely the way to go. We personally found common reranker models to be a little too opaque (can't explain to the user why this result was picked) and not quite steerable enough, so we just use another LLM for reranking
22.
▲
Magic-Image: A React Component That Lets Coding Agents Create Images
(bits.logic.inc)
2 points
by
sgk284
9mo ago
|
0 comments
23.
▲
Codex is a Slytherin, Claude is a Hufflepuff
(bits.logic.inc)
19 points
by
sgk284
9mo ago
|
7 comments
24.
▲
Intent: An LLM-Powered Reranker Library That Explains Itself
(bits.logic.inc)
1 points
by
sgk284
9mo ago
|
0 comments
25.
▲
How to Ship Confidently When Your Back End Makes Things Up
(bits.logic.inc)
3 points
by
sgk284
9mo ago
|
0 comments
26.
▲
Deep Onboarding
(bits.logic.inc)
2 points
by
sgk284
9mo ago
|
0 comments
27.
▲
Fine-Tuning Is (Probably) a Trap
(bits.logic.inc)
5 points
by
sgk284
9mo ago
|
0 comments
28.
▲
Show HN: Turn your startup logo into a holiday Google doodle
(doodle.logic.inc)
3 points
by
sgk284
9mo ago
|
1 comments
29.
▲
by
sgk284
10mo ago
We put this together mostly just to do side-by-side comparisons, though you make a good point. It'd be fun to blind-vote on your favorite impl.
30.
▲
Show HN: Agentic Arena – 52 tasks implemented by Opus 4.5, Gemini 3, and GPT-5.1
(arena.logic.inc)
1 points
by
sgk284
10mo ago
|
2 comments
More ›