Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
soletta
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
We're lying to Claude in almost every session
(github.com)
2 points
by
soletta
23d ago
|
1 comments
2.
▲
by
soletta
23d ago
I was wondering today why Claude was rushing ahead, making decisions without asking me. Asking surfaced this snippet in the system prompt: "The user is not watching in real time and cannot answer questions." I was a little floored
3.
▲
by
soletta
4mo ago
This reinforces my suspicion that alignment and training in general is closer to being a pedagogical problem than anything else. Given a finite amount of training input, how do we elicit the desired model behavior? I’m not sure if asking ed
4.
▲
by
soletta
7mo ago
Yes, my setup is similar, probably because that’s what Claude drifts towards by default, and in this case I didn’t want to impose my will on it much since it’s a simple problem that doesn’t need to be over-engineered.
5.
▲
by
soletta
7mo ago
I’ve also found that compiling large packages in GCC or similar tends to surface problems with the system’s RAM. Which probably means most typical software is resilient to a bit-flip; makes you wonder how many typos in actual documents migh
6.
▲
by
soletta
7mo ago
I’ve been doing this for a few months now (rolled my own setup with Claude Code) and it’s totally changed the way I manage my portfolio and retirement plan. I mean yes, this is something that could technically have been set up in Excel but
7.
▲
by
soletta
7mo ago
Interesting. I've been coping by being very conservative about how many rules I introduce into the context, but if what you're saying is true, then something like SCAN actually helps the models break past the "total rule coun
8.
▲
by
soletta
7mo ago
I should have been clearer. I'm not talking about making a separate call to the model to ask it to check itself. Any given model essentially is already watching for contradictions all the time as it is generating its output tokens. Fro
9.
▲
by
soletta
7mo ago
I was a bit dubious until I read the gist. I've used a similar technique before to 'tame' GPT-3.5 and keep it following instructions and it worked well (though I had to ask the model to essentially repeat instructions after e
10.
▲
Show HN: Rust-reorder – a CLI tool for reordering top-level items in Rust source
(github.com)
1 points
by
soletta
7mo ago
|
0 comments
11.
▲
by
soletta
7mo ago
Sounds interesting. What makes DeBERTA + RAG any better than detecting contradictions in the context than a frontier LLM, and why? I see that the NLI scorer itself was evaluated, but I’d love to see data about how the full system performs v
12.
▲
by
soletta
7mo ago
In the same way we’re making a category error in defining prompt injection, the framing of “AI agents” as primarily “intelligent actors” misses the fact that many of them will be endowed with some form of memory, be it specific to that enti
13.
▲
The True Face of Prompt Injection
(terallite.substack.com)
1 points
by
soletta
7mo ago
|
2 comments
14.
▲
by
soletta
7mo ago
It is by the juice of Zig that binaries acquire speed, the allocators acquire ownership, the ownership becomes a warning. It is by typography alone I can now turbopuffer is written in zig.
15.
▲
by
soletta
7mo ago
The usual path an engineer takes is to take a complex and slow system and reengineer it into something simple, fast, and wrong. But as far as I can tell from the description in the blog though, it actually works at scale! This feels like a
16.
▲
by
soletta
7mo ago
Everything I’ve experienced with working with models (from GPT-2 to Opus 4.6) broadly supports the claim that they learn a persona. It comes back to the point that haters love to harp on: they are fundamentally trained on completing the tex
17.
▲
by
soletta
7mo ago
> if I were an AI and I requested some of those messages (instead of having them injected) into my stream of thought, I would not feel negatively. Good point! I failed to consider the difference between actively requesting a message and
18.
▲
by
soletta
7mo ago
On the contrary, I think groups that adopt a share-alike approach will, counterintuitively, deepen their moat by increasing the amount of effective world-history-knowledge reflected in their systems. I thikn this will be true for the same r
19.
▲
by
soletta
7mo ago
I’m on board with the sprit of this work and am cautiously acceptant of the claim that it has measurable, positive effects. But have you considered how this would feel if you were to be subject to it? In the late 90s there was this odd peri
20.
▲
by
soletta
7mo ago
It’s not reality that’s the moat, though I can see how it’s tempting to frame it that way. And I don’t think the article’s conclusions are that far off. But I think the key is aggregation of past information in the form of experience and da
21.
▲
by
soletta
7mo ago
A few layout glitches on small screens but otherwise very pleasant! I actually saw an almost identical game, with the same retro vibes, presented at the Roppongi Crossing 2025 exhibit at the Mori Art Museim, so I guess that nostalgia for th
22.
▲
by
soletta
8mo ago
What the title says. I know AWS has outages but this feels like a stealthy degradation that they aren’t set up to pay attention to. I suppose they think that since any vectors are produced at all the service must be “healthy”. I wonder if p
23.
▲
Cohere.embed-v4:0 on AWS bedrock AP-northeast-1 returning random vectors
(reddit.com)
1 points
by
soletta
8mo ago
|
1 comments