Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jascha_eng
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
jascha_eng
7d ago
But it makes it really weird to type one handed I think.
2.
▲
by
jascha_eng
11d ago
But it is actually a great model it e.g. got the carwash question right from 9 months ago. While openais models all struggled.
3.
▲
by
jascha_eng
12d ago
Fair but I think those two behaviours are strongly correlated at least the index does represent my personal experience very well where fable is way better than opus opus is better than sol. And I haven't tried Astra yet but it having 4
4.
▲
by
jascha_eng
12d ago
I dont think you read my message. Muse and sol are nowhere near fable and Astra on the omniscience index
5.
▲
by
jascha_eng
12d ago
Imo the omniscience index they have has the highest correlation to actual usefulness of the models. https://artificialanalysis.ai/evaluations/omniscience > measures knowledge reliability and hallucination. It reward
6.
▲
by
jascha_eng
24d ago
The ai hype word for 2026 after agent in 2025 for any LLM powered application. Well kind of, I wouldn't be surprised to see that some things marketed as agents are actually good old deterministic software.
7.
▲
by
jascha_eng
27d ago
Instagram, Facebook and even threads all had much more mundane growth rates and definitely no unexpected jumps like GitHub is experiencing. I'm sure if suddenly the solar system had 10 more earths with each about 10 billion people and
8.
▲
by
jascha_eng
1mo ago
1 in 3 is not terrible you just need a few more humans in the loop to reduce the error rate meaningfully. Combined with other classifier models and heuristics you can get good results. Humans can probably also perform better if they don
9.
▲
by
jascha_eng
2mo ago
But fables is often worth diving into at least. Opus 5 says this and then rambles on about something completely irrelevant or even wrong
10.
▲
by
jascha_eng
2mo ago
The study he cites is also specifically using digital pens > Brain electrical activity was recorded in 36 university students as they were handwriting visually presented words using a digital pen and typewriting the words on a keyboard k
11.
▲
by
jascha_eng
2mo ago
tbh AI is great at vague strategic decision making with a low risk bias. I wouldn't mind working for Claude
12.
▲
by
jascha_eng
2mo ago
huh? because im curious what they used? Claude Code and codex take completely different approaches. The core loop of feeding generation and having a bash tool is entirely the same. If you want to start building your own thing I'm sure
13.
▲
by
jascha_eng
2mo ago
How did you build your own coding agent? What language/framework did you use?
14.
▲
by
jascha_eng
2mo ago
It used to be that DevOps was the movement of merging Ops leftwards so that Devs own the Operation of their software.
15.
▲
by
jascha_eng
2mo ago
And that doesn't work with a simple prompt?
16.
▲
by
jascha_eng
2mo ago
Can you give an example? And more curious about what you do with the resulting code afterwards I imagine its gonna be a big chunk then?
17.
▲
by
jascha_eng
2mo ago
Is this useful? I feel like the problem is usually not that the model isn't capable of achieving what I give it, but the way it does it. Especially if originally I didn't 100% know how I would do it myself the model often takes we
18.
▲
by
jascha_eng
2mo ago
Until the classifier is wrong or also prompt injected. the classifier is just as vulnerable as the model itself is. Yes it is harder to break but trying to make a nondeterministic tool deterministic by adding another nondeterministic one on
19.
▲
by
jascha_eng
2mo ago
But a pixel is quite a bit more expensive no? At that point you can consider an iPhone?
20.
▲
by
jascha_eng
2mo ago
Sad I have a 6 year old oneplus and was looking for a new phone somewhat soon, would've considered them again for sure. Any alternatives? They always had a reputation for me for being a great no fuss, little bloat and simply fast andro
21.
▲
by
jascha_eng
2mo ago
I mean it might lead to better performance on the model side. So the tokenizer is better but more expensive.
22.
▲
by
jascha_eng
2mo ago
Aside of the claudeisms and the obvious AI smell, it overexplains everything and doesn't come to any useful conclusions. It's just not a good post. The nudge to think about both "tokenization as variable" as well as actu
23.
▲
by
jascha_eng
2mo ago
Slowly improving the UX on my SQL review/approval tool: https://github.com/kviklet/kviklet Also finally closed the first real customer on it recently! I want to get through a large chunk of the open issues the nex
24.
▲
by
jascha_eng
2mo ago
im talking about anthropics pricing
25.
▲
by
jascha_eng
2mo ago
Apparently not? At least our it guy said after 150 people or so you have to pay for enterprise which is pay per token for everyone.
26.
▲
by
jascha_eng
2mo ago
Not including their best model in a max subscription would otherwise be truly a good reason for once to consider going back to openai for me. I'll at least try it.
27.
▲
by
jascha_eng
2mo ago
It's great software in the sense that it makes a shit ton of money though. In the end software that doesn't get used and doesn't make any money but has no bugs is not valuable either. Not saying that this is the trade off you
28.
▲
by
jascha_eng
3mo ago
I mean these were all solved before I assume so 100% not the same human ofc but models are expected to be good at a variety of code bases while human can specialize in one and learn. I think it's fair to compare to an individual that i
29.
▲
by
jascha_eng
3mo ago
Can't wait for China to pull ahead so we have an end to this bullshit
30.
▲
by
jascha_eng
3mo ago
Shorter sessions more often doing a /clear etc. save a shit ton of tokens. I pay 100 bucks a month but barely use 30% of it most weeks.
More ›