Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
joshheitzman
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
joshheitzman
3d ago
The actual title of the paper is: "Can AI agents conduct open-ended AI research? Early evidence from two case studies" While I appreciate that the article is throwing a web blanket on doomer claims, the actual study doesn't
2.
▲
by
joshheitzman
3d ago
An air gapped sandbox is immune to escape.
3.
▲
by
joshheitzman
3d ago
> The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse. Build a better simulator to train them in (i.e. more expensive) that includes a simulation of
4.
▲
by
joshheitzman
4d ago
> Before December 2025 they were still intelligent code autocomplete or Stack Overflow bots This is false. Coding agents have been usable since at least May of 2025. I can't speak to earlier than that as May last year was when I p
5.
▲
by
joshheitzman
4d ago
Here's how it make senses: - reality they can't keep up their pace - that's bad for forward projections of their revenue - that's bad for their stock price - being 'forced' to slowdown by the governme
6.
▲
by
joshheitzman
4d ago
The lack of access to brute force their training seems to be resulting in them training more efficiently too such that they are quickly catching up despite current restrictions. Between that and the chip manufacturing capacity they are bui
7.
▲
by
joshheitzman
5d ago
We've had software factories for decades. They're usually called compilers, linkers, toolchains, etc.
8.
▲
by
joshheitzman
5d ago
Who is this vendor that is consistently providing high quality inference for all families of open weight models at a competitive cost? That's a serious questions that I really interested in the answer too. I have 25 providers included
9.
▲
by
joshheitzman
5d ago
It's more than just quantization. The middleware the provider is running matters a lot even to the point of exactly which version they are running due to defects being introduced / resolved. In my coding agent harness I've i
10.
▲
by
joshheitzman
6d ago
Take a computer that can run the biggest LLM available today. It can also run any smaller LLM as well. It can also run software that isn't an LLM at all. Brains and LLMs are not at all equivalent as LLMs lack a stateful physical for
11.
▲
by
joshheitzman
6d ago
They aren't grown/evolved from data, they are fit to the data. The fitting process can be fully deterministic although its fairly easy to screw things up such that it isn't deterministic, but that just a defect not some fund
12.
▲
by
joshheitzman
6d ago
Simple cellular automata demonstrate emergent behavior. Emergent behavior is nothing new in computer science and is not remotely unique to LLMs.
13.
▲
by
joshheitzman
6d ago
This assumes the functionality of brains can be fully captured as a deterministic mathematical function, but the function of the brain may well depend on nondeterministic quantum states that can't be reduced to stateless functions: ht
14.
▲
by
joshheitzman
6d ago
The problem with this line of thinking is that modern computers are nothing like the brain. LLMs don't stand on their own, they have to be run on these modern computers, but doing so does not change the physical properties of the comp
15.
▲
by
joshheitzman
13d ago
Not at all. We all did it when I was child that age in the 80s.
16.
▲
by
joshheitzman
14d ago
Claude Code may simply be best used with Anthropic's models and quite bad with Kimi. An alternate solution is to remove Claude Code from the diagram if its so far off from the others that it causes scaling problems. I was looking at t
17.
▲
by
joshheitzman
14d ago
I was excited until I saw cost was only provided as the median. Your provider will bill you for all your tasks and one can get back to that total from the mean by multiplying by the number of tasks. This isn't possible with the media
18.
▲
by
joshheitzman
15d ago
We're done here.
19.
▲
by
joshheitzman
16d ago
I've been using Brave for years and its great. Having JS off by default (i.e. shields) is excellent for security. They provide MV2 versions of AdGuard and uBlock Origin that you can enable directly from the browser settings. Those a
20.
▲
by
joshheitzman
17d ago
> Third you’re going to have to use TS for the frontend anyway. You can’t escape ts. This is obviously wrong. Sure you have to use Javascript for the web frontend that has interactivity without round-tripping to the server. Typescript
21.
▲
by
joshheitzman
17d ago
That's AI. That's the most flattering thing I can think of to say about statements that are so confidently wrong.
22.
▲
by
joshheitzman
17d ago
Emergent behavior is not new in the realm of software. Cellular automata has emergent behavior that can get pretty wild. For example: https://en.wikipedia.org/wiki/Lenia There are huge differences between brains and
23.
▲
by
joshheitzman
17d ago
They act like LLMs. They are unprecedented.
24.
▲
by
joshheitzman
18d ago
You are correct.
25.
▲
by
joshheitzman
18d ago
Maybe's its a problem with the hosting at novita.ai but I didn't got much useful out of this model as a coding agent.
26.
▲
by
joshheitzman
18d ago
I was hoping to see the reasoning_content mess get robustly fixed, but all we got was this doc change: https://github.com/vllm-project/vllm/pull/50624
27.
▲
by
joshheitzman
18d ago
That token dump is the funniest thing I've ever read here.
28.
▲
by
joshheitzman
18d ago
It's easier to deploy AI than it is to create good culture.
29.
▲
by
joshheitzman
19d ago
Exactly this. I've had ones for 10+ years that were never connected to the internet and they functioned perfectly well as dumb TVs.
30.
▲
by
joshheitzman
19d ago
Cool. I've never seen a computer act as if it's having consciousness.
More ›