Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
xianshou
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
xianshou
26d ago
the future is here and one should be thankful for its slightly uneven distribution. otherwise we would hardly have anything left about which to develop strong opinions!
2.
▲
by
xianshou
29d ago
Really missed the opportunity for "Airemin".
3.
▲
by
xianshou
5mo ago
A lovely example of a study that is both obviously true and misses the point. Music with lyrics directly interferes with any task that has a verbal component, and the worse you are at multitasking, the worse the interference. Despite being
4.
▲
by
xianshou
6mo ago
Even as someone extremely firmly on the other side of the AI debate, I must appreciate the craft. Now, to give Claude the steganogravy skill...
5.
▲
by
xianshou
6mo ago
From the file: "Answer is always line 1. Reasoning comes after, never before." LLMs are autoregressive (filling in the completion of what came before), so you'd better have thinking mode on or the "reasoning" is pur
6.
▲
by
xianshou
6mo ago
I appreciate not having to read this guy again.
7.
▲
by
xianshou
7mo ago
Great work! Why no benchmarks though?
8.
▲
by
xianshou
8mo ago
Nice! 5 bucks says you can swap this in for your average software kanban and it does a better job.
9.
▲
by
xianshou
8mo ago
Safer than clawdbot/moltbot, I'll bet.
10.
▲
by
xianshou
8mo ago
Incidentally, Chroma also produced the single best study on long-context degradation that I've come across: https://research.trychroma.com/context-rot Before that, I cited nolima ( https://www.reddit.com/
11.
▲
Merge and Conquer: Evolutionarily Optimizing AI for 2048
(arxiv.org)
1 points
by
xianshou
11mo ago
|
0 comments
12.
▲
Stuck in the Matrix: Probing Spatial Reasoning in Large Language Models
(arxiv.org)
1 points
by
xianshou
11mo ago
|
0 comments
13.
▲
Reflection AI Raises $2B to Build "American DeepSeek"
(nytimes.com)
9 points
by
xianshou
11mo ago
|
2 comments
14.
▲
Nvidia-backed Reflection AI raising at $5.5B valuation
(reuters.com)
2 points
by
xianshou
11mo ago
|
1 comments
15.
▲
by
xianshou
1y ago
Came to point out that this is transparently LLM-authored, was not disappointed. The signs: - neatly formatted lists with cute bolded titles (lower-casing this one just for that) - ubiquitous subtitles like "Mental Health as Infrastruc
16.
▲
by
xianshou
1y ago
I initially read the title as "My 2.5 year old can write Space Invaders in JavaScript now (GLM-4.5 Air)." Though I suppose, given a few years, that may also be true!
17.
▲
by
xianshou
1y ago
Rug pulls from foundation labs are one thing, and I agree with the dangers of relying on future breakthroughs, but the open-source state of the art is already pretty amazing. Given the broad availability of open-weight models within under 6
18.
▲
by
xianshou
1y ago
In many of their key examples, it would also be unclear to a human what data is missing: "Rage, rage against the dying of the light. Wild men who caught and sang the sun in flight, [And learn, too late, they grieved it on its way,] Do
19.
▲
by
xianshou
1y ago
The self-edit approach is clever - using RL to optimize how models restructure information for their own learning. The key insight is that different representations work better for different types of knowledge, just like how humans take not
20.
▲
Unsupervised Elicitation of Language Models
(arxiv.org)
7 points
by
xianshou
1y ago
|
0 comments
21.
▲
by
xianshou
1y ago
The key insight here is that DGM solves the Gödel Machine's impossibility problem by replacing mathematical proof with empirical validation - essentially admitting that predicting code improvements is undecidable and just trying things
22.
▲
by
xianshou
1y ago
AI is, currently, coming not for the coders who made it but for the coders who didn't contribute to or ignored it. The foundation labs are all quite committed to recursive self-improvement of coding tools as a general research accele
23.
▲
by
xianshou
1y ago
Duplicate of https://news.ycombinator.com/item?id=44040883
24.
▲
by
xianshou
1y ago
Both Google and Microsoft have sensibly decided to focus on low-level, junior automation first rather than bespoke end-to-end systems. Not exactly breadth over depth, but rather reliability over capability. Several benefits from the agent d
25.
▲
by
xianshou
1y ago
Amusingly, about 90% of my rat's-nest problems with Sonnet 3.7 are solved by simply appending a few words to the end of the prompt: "write minimum code required" It's not even that sensitive to the wording - "be ter
26.
▲
by
xianshou
1y ago
Calling it now - RL finally "just works" for any domain where answers are easily verifiable. Verifiability was always a prerequisite, but the difference from prior generations (not just AlphaGo, but any nontrivial RL process prior
27.
▲
by
xianshou
1y ago
Bravo! Planning your life in order to minimize deathbed regrets has always bothered me, because the nature of humanity is to want what it hasn't got. If you assume that, on average, people make correct decisions to work hard and pursue
28.
▲
DeepSeek V3 0324 is now the best nonthinking model (Reddit)
(old.reddit.com)
1 points
by
xianshou
1y ago
|
0 comments
29.
▲
DeepSeek V3 0324 outpaces GPT 4.5 and Claude 3.7 in coding, other benchmarks
(huggingface.co)
7 points
by
xianshou
1y ago
|
0 comments
30.
▲
by
xianshou
2y ago
Any way to parallelize tool use? When I go into a repo and ask "what's in here", I'm aiming for a summary that returns in 20 seconds.
More ›