Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aethelyon
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
GPT-5.6, Fable 5, and Grok 4.5 rebuild Basecamp from the same spec
(smw.ai)
6 points
by
aethelyon
2mo ago
|
1 comments
2.
▲
by
aethelyon
2mo ago
I gave GPT-5.6 Sol, Fable 5, Grok 4.5, Sonnet 5, and GPT-5.5 the same greenfield spec: build the Basecamp 5 frontend and API. Fable won both tracks at $85.87 in 2:06:40. Grok reached 84% of Fable's frontend score and 87% of its backend
3.
▲
by
aethelyon
2mo ago
100% — publish the hidden research, the value is in the discoveries, not in the dividend. With all due respect to the author, it feels like he missed the entire lesson of history.
4.
▲
Show HN: ScreenCommander – Let LLM Agents control your desktop via CLI
(github.com)
1 points
by
aethelyon
7mo ago
|
1 comments
5.
▲
by
aethelyon
7mo ago
Built ScreenCommander to solve a personal gap: individual app integrations severely limit what local AI agents can actually achieve. This macOS CLI tool captures your desktop as screenshots, then allows local agents like Codex to interpret
6.
▲
by
aethelyon
9mo ago
"some" or a single file?
7.
▲
by
aethelyon
1y ago
this is fake news, the xml tags break the output when the model output is the system prompt with the example tags, see screenshot: https://x.com/0xSMW/status/1944624089597137214 same as what happens with claude
8.
▲
O3-Pro Speculative Reasoning
(smw.ai)
1 points
by
aethelyon
1y ago
|
1 comments
9.
▲
by
aethelyon
1y ago
comparing o3-pro reasoning to gemini 2.5 pro and claude 4 opus on a speculative, open-ended prompt
10.
▲
by
aethelyon
3y ago
No, I’ve seen this pattern as well. Will apologize and then when you ask to continue it will have a change of mind and refuse again. It’s a bad RLHF/AIF loop that it gets stuck into.
11.
▲
by
aethelyon
3y ago
Bloop is amazing. Once you use it you stop building your own DIY codebase QA setups.
12.
▲
by
aethelyon
3y ago
2 years
13.
▲
by
aethelyon
3y ago
This is cool, but the data collection is the hard part, right?
14.
▲
by
aethelyon
3y ago
Spoiler: it's fast, cheap, overly protective, and has Kafkaesque DX
15.
▲
by
aethelyon
3y ago
Spoiler: it's fast, cheap, overly protective, and has Kafkaesque DX
16.
▲
by
aethelyon
3y ago
This is awesome, but there were a couple of great laptop interfaces from that movie too. Spent some quality time in the 90s getting AfterStep/Litestep to look like them.
17.
▲
by
aethelyon
3y ago
I used to be worried about face scanning. But sometimes I wonder if it's an inevitable evolution of technology. Which – to be clear – is not support for it, but a question about what is emergent from the new things we create.
18.
▲
by
aethelyon
3y ago
100% agree, I think the 26% will greatly increase over time... or the ones that don't will decline as a business over time. the 13.4% is likely leaders in ML for some specific use case like fraud or recommendations. it would be great t
19.
▲
by
aethelyon
3y ago
great data – wish they provided the raw information to slice the respondent audience more, but aligns with what I've seen in the market re: concerns and models.
20.
▲
by
aethelyon
3y ago
this is cool
21.
▲
by
aethelyon
3y ago
We benchmarked retrieval, GPT-4 turbo vs GPT-4, and fine-tuned several models: https://klu.ai/blog/openai-devday-2023 You can use the result of one here https://huberman.klu.ai/
22.
▲
by
aethelyon
3y ago
We benchmarked retrieval, GPT-4 turbo vs GPT-4, and fine-tuned several models. You can use the result of one here https://huberman.klu.ai/
23.
▲
by
aethelyon
3y ago
We benchmarked retrieval, GPT-4 turbo vs GPT-4, and fine-tuned several models. You can use the result of one here https://huberman.klu.ai/
24.
▲
by
aethelyon
3y ago
check out https://klu.ai – we built it for this reason – sign up, book some time, and I'll help you however I can
25.
▲
by
aethelyon
3y ago
Microsoft brought GPT-4 to GA for all customers on Azure OpenAI this week. This removes the endless waitlist for some. Wrote up a few notes from our experience with it.
26.
▲
Startup Guide to Azure OpenAI
(klu.ai)
3 points
by
aethelyon
3y ago
|
1 comments
27.
▲
by
aethelyon
3y ago
we built https://klu.ai/ for this ====== outside of us, here's what I see happening 80% of folks aren't building in prod if you pull apart the 20% that are building, I've seen this from largest to smallest po
28.
▲
Ask HN: Compiling GPT-4 model card
(klu.ai)
3 points
by
aethelyon
3y ago
|
1 comments
29.
▲
by
aethelyon
3y ago
I started compiling all of the known public information (ala geohotz, semianalysis, et al) in an attempt to build a model card for GPT-4. Am I missing anything?
30.
▲
by
aethelyon
3y ago
seems like it, but no one is talking about it – everyone I ask IRL says performance is bad, but not seeing in benchmarks
More ›