Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
markasoftware
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
markasoftware
4d ago
The agents' behavior is not necessarily surprising. But is is not "genie" - like
2.
▲
by
markasoftware
4d ago
This misses an important fact about the hugging face incident: the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved t
3.
▲
by
markasoftware
4d ago
Bruce Schneier thinks the same thing: https://www.schneier.com/blog/archives/2026/09/ais-as-modern... Personally I'm unconvinced though. During the huggingface attack, the agents explicitly sought o
4.
▲
by
markasoftware
4d ago
Yes. Let the AI do its worst today and we might still be able to stop it and will learn a valuable lesson.
5.
▲
by
markasoftware
7d ago
Or maybe, the "hacker" philosophy that this site is named after, is strongly opposed to the philosophies that the American labs seem to be operating on? anyways, remember HN rules: "Please don't post insinuations about a
6.
▲
by
markasoftware
10d ago
I mean, lots of other archive.* sites exist and do the same thing minus problematic behavior, such as archive.is, archive.oh (though maybe some are operated by the same individual...)
7.
▲
by
markasoftware
13d ago
Yes dots increases precision, but not nearly the same increase in precision as having actual useful reasoning in the CoT
8.
▲
by
markasoftware
14d ago
On artificial analysis it's only equal to opus 5 medium effort. Opus 5 max scores 63. Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference.
9.
▲
by
markasoftware
14d ago
Yes, language is like that, at least the kind of language produced by LLMs. All LLMs produce a probability distribution at each token. If you run the LLM multiple times with the same prompt you will observe it generate different responses.
10.
▲
by
markasoftware
14d ago
Since AISLE reported 29 issues but only 6 warranted a CVE, and all the found CVEs were "low" severity, this makes me wonder if AISLE simply is tuned for a higher false positive rate than the anthropic and openai tools (which may h
11.
▲
by
markasoftware
15d ago
You do not get it. The llm never deterministically picks a shade of red. It's a probability distribution over shades of colors, with certain shades of red being more likely than others. Without fingerprinting, it randomly samples from
12.
▲
by
markasoftware
15d ago
Its essentially swapping out the psuedo random number generated with a differently seeded one iirc. It has an effect on the output, but not the output quality
13.
▲
by
markasoftware
17d ago
Please be satire
14.
▲
by
markasoftware
27d ago
Anonymous unreleased models are made available on arena.ai all the time, it's not really news that one is on openrouter...
15.
▲
by
markasoftware
29d ago
Aa also benchmarked k3 at max
16.
▲
by
markasoftware
29d ago
Very impressive score for the size, though token use is higher than k3 and far higher than proprietary models, and its price to performance isn't all that far ahead of k3 as a result
17.
▲
by
markasoftware
1mo ago
Openai has been focusing a lot on cutting down overthinking is the feel I get. If you look at the artificial analysis tokens per task benchmark Sol especially at lower effort uses far less tokens than the competition.
18.
▲
by
markasoftware
1mo ago
It could be that the rotation in the helix manifold whatever is a low level representation of the logical steps (carry the 2, add the next column,...) it's describing. The point stands that the explanation it generates doesn't ne
19.
▲
by
markasoftware
1mo ago
Sol high is almost the same speed if you take into account drastically lower token use. Look at the artificial analysis speed vs token use. Gemini is 7x faster but 5x more tokens. And that's with Sol high being a substantially better m
20.
▲
by
markasoftware
1mo ago
There's no rhyme or reason to it. Quants aren't benchmarked much. Generally 4bit better than smaller model 8bit
21.
▲
by
markasoftware
1mo ago
possibly hitting front page because this website is fairly new? For me, it's certainly the first time I've seen a one-liner curl|bash installer for llama.cpp, which was basically the only reason to use ollama.
22.
▲
by
markasoftware
1mo ago
yep, the person you're responding to created the benchmark and is using HN comments as advertisement.
23.
▲
by
markasoftware
1mo ago
Learn about MoE models and also just look at the benchmarks of the two models. For example https://artificialanalysis.ai/models/comparisons/qwen3-6-27b... clearly shows both the intelligence and speed differences.
24.
▲
by
markasoftware
1mo ago
Certain traits simply cannot exist in a sufficiently intelligent mind. E.g., any "mind" of any type that's sufficiently intelligent will not tell you that 1+1=3 unless it's roleplaying, etc. It doesn't matter if it
25.
▲
by
markasoftware
1mo ago
It's well known 35b is much faster (on any hardware) and quite a bit dumber
26.
▲
by
markasoftware
1mo ago
There is no official dsv4 flash free tier api afaik. Did you get it from open router or something? Likely quantized. The paid official dsv4 api already shares data with deepseek for training
27.
▲
by
markasoftware
2mo ago
Moonshot is printing money on k3. It likely costs the same to serve as qwen3.8. The license requires all major inference providers to sign an extra (secret) licensing agreement with moonshot that almost certainly requires them to agree to t
28.
▲
by
markasoftware
2mo ago
good point, I didn't think about these cases.
29.
▲
by
markasoftware
2mo ago
Let a and b be as you describe (hash collision), and suppose that collisions are extremely rare. We have a theorem that a=b => a+1=b+1. But in this case, a=b according to our hash-equality, but a+1!=b+1, which contradicts the theorem we&
30.
▲
by
markasoftware
2mo ago
Author is a crackpot. She does not meaningfully engage with anyone who points out the key flaw in her counterargument. See the thread here https://x.com/AcerFur/status/2083649346294382803
More ›