Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
NiloCK
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
NiloCK
4d ago
As I understand it, the rough guess as to what's happening here is that most recent capabilities progress comes from specific verifiable-rewards reinforcement training (RL). The RL pressures are all about task performance, but (surpris
2.
▲
by
NiloCK
4d ago
This is unfair. Dario signed the Pacing the Frontier open letter when Fable/Mythos seemed from the outside to be an insurmountable lead. Also he's been saying versions of this day in and day out for as long as he has had anyone&#x
3.
▲
by
NiloCK
4d ago
The degradation, polarization, and weaponization of or media landscape over the era of social media has left us utterly incapable of believing anything that anybody says. Dario in particular has consistently been risk-wary on model improvem
4.
▲
Google should provide a technical postmortem of Gemini's 2024 outburst
(paritybits.me)
2 points
by
NiloCK
5d ago
|
0 comments
5.
▲
by
NiloCK
7d ago
> The current prices the largest players set for their models are not profitable, they bleed money. How are the open-weight Chinese models staying ~6-12 months behind on widely distributed / commodified hardware, and serving for e
6.
▲
by
NiloCK
13d ago
If someone could tune models of that size to have comparable effectiveness at much much lower costs, they would have done so by now. "The harness improvements are the real sauce" is like a sincere "It's gotta be the shoe
7.
▲
by
NiloCK
14d ago
Gemini models - at least via some interfaces - have tool calling API access to various Google integrations. flights.google.com, maps.google.com, etc. The info isn't in the model weights. Because of where I live, there are three viable
8.
▲
by
NiloCK
1mo ago
FYI you can launch claude-code with your own prompt. Don't quote me but: claude --system-prompt "Mine is better than Anthropic's"
9.
▲
by
NiloCK
1mo ago
Why not set a global instruction that their direct outputs to you should be in your native language? For a long time I had Claudes (in the 4.0-4.5.x range) use only French in the chat, while keeping English for working docs (and the code, o
10.
▲
by
NiloCK
1mo ago
Before OpenAI / situational-awareness he made a lot of money on societal-shift type investments and shorts during the lead up to Covid economic impacts. His "main thing" is success in calling economic impacts of undervalued l
11.
▲
by
NiloCK
1mo ago
Working on https://letterspractice.com This is a high efficiency, narrowly-scoped, low screen-time early literacy app for families with kids aged 2+.
12.
▲
by
NiloCK
1mo ago
Years ago I also did some experimentation w/ midi-device and SRS ( https://www.youtube.com/watch?v=a6tvHMvF8Mo ), where the focus was on ear-training rather than score-learning. Clef seems to be a pretty strong attempt
13.
▲
by
NiloCK
1mo ago
A feature whose absence I've found more and more conspicuous over time is interactive-compact . Given a current context, and impending context overflow, I know the directions that my mind is heading, and where I expect the developme
14.
▲
Show HN: LettersPractice – Teach your child to read with a modified SRS engine
(letterspractice.com)
1 points
by
NiloCK
1mo ago
|
0 comments
15.
▲
by
NiloCK
1mo ago
Not so long ago, I was good enough for many coding tasks. But I found that things can change in a hurry. Yes, a cheap and fast Opus4.6 can drive a lot of value in current context. But if we continue to craft bigger-and-bigger balls of mud,
16.
▲
by
NiloCK
1mo ago
Yes, but I expect that the elided part is less important than people assume it is. Token count is a less important factor in context pollution than idea count. The worst of the rot factors are when models latch onto irrelevant information,
17.
▲
by
NiloCK
1mo ago
I find this astounding. 2024 to present thread is can write a coherent 15 line function to ... what exactly? No future for research mathematicians othet than as tastemakers / agenda setters?
18.
▲
by
NiloCK
2mo ago
The general model for subscriptions is that power users are subsidized by subscriptions of casual users, like a gym membership or whatever. This is a little dicier in post-agent AI, because it's easier for casual users to automate powe
19.
▲
by
NiloCK
2mo ago
> Circular Revenues: A small handful of tech firms, chip manufacturers, and AI companies are propping each other up by investing and buying from each other. > Increasing Corporate Skepticism: The news is full of stories of corporation
20.
▲
by
NiloCK
2mo ago
Recent LLM dingers like the Jacobian Conjecture counterexample have challenged the efficient mathematics hypothesis. The JC counterexample was so small in degree and coefficient. It should have been a "low fruit" in the scheme of
21.
▲
by
NiloCK
2mo ago
I'll bite. Claude code interacts with many system processes, files, etc, as well as external APIs. Processes audio via built in dictation. Manages a bunch of nasty auth. Etc etc. What are the categories of features that wouldn't b
22.
▲
by
NiloCK
2mo ago
Of course there is: competition with other labs, and self-hosting of open-weight models! Yes, the mechanics are straightforward if Anthropic (or Claude, if you want to ascribe the decision there) decides to burn a pile of your money. But th
23.
▲
by
NiloCK
2mo ago
Opus 4.7+ and Fable are both much more aggressive than prior models with respect to writing memories to a location that's effectively quasi-private for them. It's device-local (so passes retention constraint), and you can see it
24.
▲
by
NiloCK
2mo ago
For capabilities reference: I made a lower effort but similar scaffold for LLMs to do iterative drawing in Nov 2024, with Sonnet 3.5 as the artist: https://paritybits.me/llm-drawing-with-eyes-open/ Quite a difference.
25.
▲
by
NiloCK
2mo ago
> You can have these things without owning a house. Yes this is feasible, but respecting it as a design problem, renting is more transient than ownership and tilts the floor away from deep communal relationships. Comparing the neighbor
26.
▲
by
NiloCK
2mo ago
Thanks much. Source code is for customers only I guess :)
27.
▲
by
NiloCK
2mo ago
I had the same nit, but I imagine deforming text / inline content generally would be a much larger effort.
28.
▲
by
NiloCK
2mo ago
I like this a lot and am going to experiment w/ incorporating in my early literacy app. Heads up: the "See it live in the showcase → " links in the API documentation do not go back to the showcase - they just reload the same
29.
▲
by
NiloCK
2mo ago
Yes - persons with death wishes having arbitrarily powerful consultation is the crux of it. Apologies for the bad example. Replace w/ gain of function / whatever else, or just brainstorm with your local model, ect.
30.
▲
by
NiloCK
2mo ago
The logic, whose premises you can take or leave: Even at the level of, say, Opus 4.5+, open weight models give a quick turnaround to every Joe and Jane on earth having easy access to pretty high quality improvised weapons design, cyber 
More ›