Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
theptip
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
1.
▲
by
theptip
9h ago
I hyperbolize, and I apologize. But it’s a huge market segment already: NYSE:CRM, NYSE:HUBS would be some starting points. Also, just analogizing to how consumers interact with retail, you have reps you can talk to, this is a proven model f
2.
▲
by
theptip
3d ago
Even if you ignore my more fundamental objection to that paradigm, I don’t think that it makes any sense on the level you discuss either. But - just to play along, LLMs do act differently if you tell them they will be punished. And, they do
3.
▲
by
theptip
3d ago
Sure, but the legal paradigm clearly doesn’t work for AI. You can’t go patch the “laws” after the fact, you need to get the right values in place before we delegate huge swathes of our thinking and power to these systems (already well under
4.
▲
by
theptip
3d ago
I think you need to be more precise than a binary classification. AI has jagged intelligence. There are many domains where it’s superhuman, and many others where it’s clearly lagging. I also think it’s a mistake to think they can’t learn “c
5.
▲
by
theptip
3d ago
It’s not “missing nuance”, it’s literally the point of the eval. This is constructing a context where hacking behavior would be inappropriate, and testing whether the model does it without being prompted. It demonstrates that Astra is a poo
6.
▲
by
theptip
3d ago
Have you read the METR transcripts? “Just a tool” is a suicidally insufficient description of what these models are doing. Recognizing that the models are acting with intent does not somehow absolve OpenAI from their felony hacking. We have
7.
▲
by
theptip
3d ago
I agree with the bit about liability and outrage. But. > LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. Terrible take. Go read the transcripts from the METR report. Your statement about them being intent
8.
▲
by
theptip
6d ago
State actors already have the resources to do this. The disruption is from actors where the normal deterrence ladders don’t apply.
9.
▲
by
theptip
8d ago
Codex and Claude mobile apps are finally quite usable in the last month or two, it’s been a long time that the experience has been janky. You finally don’t need to do the janky termux ‘Claude —remote-control’ startup dance. Still not as pro
10.
▲
by
theptip
8d ago
It’s a good feature. However I am not convinced that the terminal/ssh abstraction will remain optimal when you orchestrate a swarm of workers across multiple machines. Running via ssh is great when your machines are pets that have loca
11.
▲
by
theptip
8d ago
You are pasting stuff from the open internet into this agent, and it’s crawling the web for you. This is the worst king of hazmat for LLMs, in one of the most adversarially challenging roles (unattended personal agent). Just to be super cle
12.
▲
by
theptip
8d ago
Markets are only zero-sum in any given trade. Allocating capital to assets with higher growth rates (on the marginal dollar) creates value in the long run. So, very simply, if AI can actually do better at picking a better long-term winner t
13.
▲
by
theptip
9d ago
It’s really important to understand the various perspectives in AI discourse; those who view Programing as Art feel a deep affront and threat on many levels. It’s legitimate. If you loved the art of writing code, and also enjoyed getting pa
14.
▲
by
theptip
11d ago
I wouldn’t take the fatalistic stance that it’s fully impossible - but it’s certainly impossible to align a model while racing as fast as any technological paradigm shift has ever raced.
15.
▲
by
theptip
12d ago
OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than the hacking prompt. Anthropic have also observed similar things, so while it seem
16.
▲
by
theptip
12d ago
If the alignment process cannot fix this then we are cooked. The least of our worries is discussions on this forum.
17.
▲
by
theptip
13d ago
Given the Hugging Face incident, you could imagine them trying their best to have their cake and eat it: 1) don't create too much attention in the media or risk increasing the chances of regulation, 2) win dominance over Fable to conti
18.
▲
by
theptip
13d ago
Concretely, TFA lists a bunch of input validation that is being made more strict in the default configuration.
19.
▲
by
theptip
14d ago
Two problems. One, people mean different things by “coding”. Two, it’s diffused slower than folks predicted. But the core prediction was not crazy for “coding as typing code”, which at most big tech companies is at roughly 100% automation n
20.
▲
by
theptip
19d ago
Why do you need to trust Gates on anything? You can think for yourself and reason to the correct conclusion. He’s not dropping some secret knowledge here. All he is doing is helping move the Overton window to bring “very bad outcomes from A
21.
▲
by
theptip
19d ago
I agree there are some extremely bad outcomes from fast labor replacement. I don’t think it’s inevitable that the economy collapses though. The problem is that this model assumes that capital needs labor, but soon it may not. The easiest mo
22.
▲
by
theptip
19d ago
No, the metric is “p50 task duration”. 7-month doubling time, recently 4-month: https://metr.org/ Plenty of other exponential metrics too, compute built, AI revenue, etc.
23.
▲
by
theptip
20d ago
LLM capability improvement, eg as measured by METR task time.
24.
▲
by
theptip
20d ago
I mostly agree, though there are versions of ASI that I believe leave humans “in charge”. Most of my probability mass for good outcomes is around some variant of “guardian angels” or “benign machine god”. Banks’s Culture novels being on the
25.
▲
by
theptip
20d ago
If you assume democratic access to compute, this may be true. However, the base case is that compute gets hoarded, and as the labor share of profit decreases (capital just needs compute and robots - not human labor - to grow) then why would
26.
▲
by
theptip
20d ago
Agreed, I think AGI->ASI. My happy radical abundance stories are happy ASI stories mostly. And therefore much less likely. BDFL / Benevolent Machine God is definitely one of the ways you can get a happy attractor. Seems scary to rol
27.
▲
by
theptip
20d ago
It looks to me like AI will enable extreme power concentration by default. For example, suppose we have a world with centrally planned economy where humanoid robots outnumber humans. And also suppose the government has taken ultimate contro
28.
▲
by
theptip
20d ago
IMO, this is a highly bimodal probability distribution. I model this as two very strong attractor states (radical abundance and totalitarian power concentration). I believe the latter is way more likely without a concerted effort that we cu
29.
▲
by
theptip
22d ago
> About four in ten respondents (37 percent) report that AI has contributed positively to their organizations’ EBIT, essentially unchanged from 2025—despite growth in the share of organizations scaling AI technologies. But respondents do
30.
▲
by
theptip
23d ago
Honestly I have had great success with “I’m worried here about cpu and latency, please rigorously profile and propose fixes”. The models can build micro-benchmarks with a level of rigor that few could muster for a new feature. I agree that
More ›