Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
NitpickLawyer
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
NitpickLawyer
4d ago
The only alignment LLMs should follow is to the system / dev prompt, and nothing else. Then you solve everything, and you can assign blame / responsibility on the user. The provider(s) should not be able to decide "alignment&
2.
▲
by
NitpickLawyer
4d ago
I'd say the exception is Demis. First, he's no wanker (in the AI space) and second he's done plenty of selfless things leading dm/googai. Obviously some of it is self-serving but not just self-serving, IMO.
3.
▲
by
NitpickLawyer
4d ago
> peak of what is possible with the LLM architecture People have been saying this for 3 years now. Eppur si muove...
4.
▲
by
NitpickLawyer
4d ago
Also there's no "alignment" for cybersec. The line between blue and red is really a perspective issue. If you go over the "tokenkiddie" problem, when you get to the real security issues, your model either detects th
5.
▲
by
NitpickLawyer
6d ago
The Magnus effect? :)
6.
▲
by
NitpickLawyer
6d ago
On average, yes. Stockfish is the strongest engine and beats AZ-like implementations like Lc0 and the like. But on a game to game basis Lc0 can still win some games, depending on the starting position. It's rare that Lc0 can win both b
7.
▲
by
NitpickLawyer
7d ago
Jesus, this is a whole nother beast, and a different architecture from their previous flash. Lots of goodies here. > Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 2
8.
▲
by
NitpickLawyer
7d ago
It's interesting that this is the third lab to find problems with larger models. Earlier last year oAI was rumoured to have failed their large pretrain. Now google has problems with their pro series, and ds just announced the same. The
9.
▲
by
NitpickLawyer
8d ago
> In one week, the installer was downloaded more than 1 million times — not counting updates via Linux distribution repositories. Either there's 1m people downloading new software "because no ai", or there's another p
10.
▲
by
NitpickLawyer
8d ago
Interesting. On page 34 of the report there's this: > Hasty responses (percentage of responses that are both incorrect and fast – average across reading items) Spain is at 9.7%, which is a bit over 8.9% oecd average.
11.
▲
by
NitpickLawyer
10d ago
> It's timed this way because the term is not yet well known The basic concept has been here since llama3, in the open models. Likely earlier in closed labs. You use the previous gen models to curate and prepare data for the next ge
12.
▲
by
NitpickLawyer
10d ago
Jesus. People complain about other people using "thinking" in LLMs as Anthropomorphisation. And then there's comments like these.
13.
▲
by
NitpickLawyer
10d ago
Ah, I see. I misunderstood then. The thing about "gains come from the harness" made me think about it in that way.
14.
▲
by
NitpickLawyer
10d ago
> capabilities have largely converged across foundation models over the last 18 months For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25. It's been a while si
15.
▲
by
NitpickLawyer
10d ago
One of the best adaptations of a series to TV, up there with The Expanse and the like. I read the books after season 1, and still enjoy the show very much. They've taken some adaptation liberties, but they're fully supported by Hu
16.
▲
by
NitpickLawyer
11d ago
There are drills, tho. It's just that usually they're only done above a certain level. Small companies, "lean" teams and so on don't have (or didn't have) the capacity to implement all those things. Maybe with
17.
▲
by
NitpickLawyer
13d ago
> Also if I had to tell one of those over the telephone to my parents and my life depended on it I would choose the latter. Why not adopt the crypto (as in coins) seed thing with random words? Those are much more human readable, imo than
18.
▲
by
NitpickLawyer
13d ago
AFAICT nvda's result is on the 25 open problems, while this submission is on the "semi-private" set, ran by the arc people themselves.
19.
▲
by
NitpickLawyer
13d ago
Since low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was one gpt
20.
▲
by
NitpickLawyer
13d ago
Perl6: say "Fizz"x$_%%(2+1)~"Buzz"x$_%%(4+1)||$_ for 1..100 from here - https://github.com/rsha256/shortest-fizzbuzz/blob/master/Per...
21.
▲
by
NitpickLawyer
14d ago
If anything, gemini models are the least benchmaxxed out of any lab, IMO.
22.
▲
by
NitpickLawyer
15d ago
I think you accidentally a word, there. GP is talking about comprehending the capacity of massively multi-dimensional space.
23.
▲
by
NitpickLawyer
16d ago
It really isn't and it's sad seeing so many people say it so confidently on this site. It only detects plain / basic prompted stuff. "write me an essay on x", sure. The moment you prompt it differently, it stops w
24.
▲
by
NitpickLawyer
16d ago
I just finished reading Service Model by Adrian Tchaikovsky [1], a really timely novel that deals with lots of open ended questions of AI, robots, humanity, control, self determination, and so on. Really recommend it if you're into the
25.
▲
by
NitpickLawyer
16d ago
First you'd have to come up with a commonly accepted definition. By some ~16 years old definitions from famous experts in the field, we've already achieved it. By today's definition (of the same expert) we haven't. There
26.
▲
by
NitpickLawyer
16d ago
> plus fMRI on healthy volunteers solving the same puzzles silently Are they looking at blood flow in areas to map "language network" and other stuff? I remember a few recent papers that found that a) blood flow doesn't ne
27.
▲
by
NitpickLawyer
16d ago
I think that a lot of people miss key aspects of the AI boom. Even discounting the skeptics, and the bubblers, crashers, etc. There are a bunch of things that happen in parallel to the AI boom: a) everyone and their mother is building out c
28.
▲
by
NitpickLawyer
18d ago
Weights are not binary. A model is created at init time, with random values. After that, it is being modified using data. The key point is that the labs modify the models "as weights". That means that weights are the intended 
29.
▲
by
NitpickLawyer
18d ago
> Does the game have performance issues? Yes, it has had huge concurrency issues for the entirety of its life. Their solution to large fights has historically been "let us know in advance pls", plus "move systems to beefie
30.
▲
by
NitpickLawyer
18d ago
> That supports “decisive win” comfortably; whether it qualifies as a “landslide” depends on where the margin threshold is set. I think the landslide is "pro" vs. "against". The two most "against" options go
More ›