Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
RussianCow
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
RussianCow
2mo ago
The system prompt and available tools would likely only change for different agent types. So how they're launched probably doesn't matter. I say this as I don't actually know how Clade Code does this, since it's not open
32.
▲
by
RussianCow
2mo ago
Yes, Codex has its own sandbox which you can disable: https://learn.chatgpt.com/docs/sandboxing?surface=app Cursor has something similar. I don't know about Claude Code but I assume it does as well since Anthropic
33.
▲
by
RussianCow
2mo ago
Compared to the primary agent, maybe. But it's highly unlikely that all the agents have different tools and system prompts than each other, and those account for the bulk of the context per the post.
34.
▲
by
RussianCow
2mo ago
You cannot; you must use either their Devin Desktop app or the Devin CLI.
35.
▲
by
RussianCow
2mo ago
But they don't appear to subsidize them to the same degree. I've only been using Devin for less than a month, but I've been hitting the limits of the $20/month plan way more quickly than I'd expect, and definitely m
36.
▲
by
RussianCow
2mo ago
But that 10% is the most important part! Getting the plumbing wrong means you might have bugs or your code is brittle. Getting the domain-specific business logic wrong means your product doesn't fundamentally solve the correct problem.
37.
▲
by
RussianCow
2mo ago
Ironically, Devin Desktop is one of those tools. It supports any harness that supports ACP (which is most of them)—you can use Claude Code, Codex, OpenCode, etc from the Devin Desktop UI. I'm currently experimenting with OpenSpec[0] as
38.
▲
by
RussianCow
2mo ago
But some requirements you don't realize you have until you start building. With a fast model, you can surface those really quickly and have more time to iterate and explore different solutions. With a slower but smarter model, you just
39.
▲
by
RussianCow
2mo ago
The regular one (not the fast variant) is free but slow. The "Lightning" variant (which uses Cerebras and gets supposedly 1000 TPS) costs $12.50/M output, $2.5/M input, $1/M cached input. So it's quite a bit mo
40.
▲
by
RussianCow
2mo ago
The "Lightning" (Cerebras) variant isn't free, only the regular one, which runs closer to 50 TPS in my experience with SWE 1.6.
41.
▲
by
RussianCow
2mo ago
That's not why Elm is an ideal language for LLMs: it's because, if it compiles, it's most likely working software. Agentic workflows have gotten significantly better over the last year, so LLMs using languages like Elm, Haske
42.
▲
by
RussianCow
2mo ago
> asking an LLM to reverse engineer and make your own plugin is trivial. If you already have engineers on staff, a few tens (or even hundreds) of dollars per month per plugin is likely a rounding error budget-wise. If you don't have
43.
▲
by
RussianCow
3mo ago
At least 3 times a month. I have a rental property and my tenant prefers to mail a check instead of paying extra to pay electronically. My spouse gets paid by check for dumb reasons I won't get into. I sometimes get dividends from my i
44.
▲
by
RussianCow
3mo ago
> Also why do you need bank app on your phone? Many banks gate features like mobile check deposit behind the native app. The nearest ATM is 20 minutes away from my house, so unfortunately I consider this feature essential.
45.
▲
by
RussianCow
3mo ago
> En, I think you’re just trying to justify your pre-existing position that this can’t work. I never said it can't work. I just said that finding the correct medical digagnosis is different than finding a solution to a software prob
46.
▲
by
RussianCow
3mo ago
Also, there are multiple "correct" ways to code something, so imperfect code that solves the problem is still useful. A medical diagnosis is either correct or incorrect.
47.
▲
by
RussianCow
3mo ago
At that point, cut out the LLM and just see the radiologist.
48.
▲
by
RussianCow
3mo ago
Probably because HDR on the vast majority of non-OLED monitors is useless. You really need a monitor with great contrast and a good HDR implementation for it to be of any benefit.
49.
▲
by
RussianCow
3mo ago
I'm not an expert, but I think those are the same thing. But for an LLM etched onto a whole wafer, it doesn't make sense to disable part of it since that would remove some weights entirely.
50.
▲
by
RussianCow
3mo ago
I don't know about you, but I generally don't write code in a vacuum. Other people may have touched it before me. Those other people may have made poor decisions. Not that I'm immune from choosing the wrong abstraction someti
51.
▲
by
RussianCow
3mo ago
> But even then a twice a week household cleaning hire is going to cost upwards of $1500/mo unless you're being particularly exploitative. Sorry, what? Unless you're doing a deep clean of your house twice a week or you liv
52.
▲
by
RussianCow
3mo ago
The Chinese open weight models have been ahead of Sonnet (at least for coding) for a couple months now. I tend to take benchmarks with a huge grain of salt, but in my own experience, the latest versions of Kimi, MiMo, and GLM (pre-5.2) had
53.
▲
by
RussianCow
3mo ago
Isn't that true of any provider? Anyone could be lying about what they're serving.
54.
▲
by
RussianCow
3mo ago
I've been doing the same, though admittedly out of curiosity more so than lack of funds. The open models are catching up quickly in their abilities, to the point where they're (mostly) not doing stupid stuff regularly, but you hav
55.
▲
by
RussianCow
3mo ago
But those things won't be sped up by a faster LLM, so I feel like that's not what the OP is talking about.
56.
▲
by
RussianCow
3mo ago
Do you mean Flash and not Pro? I haven't tried it personally, but according to OpenRouter, the fastest DeekSeep V4 Pro providers are only ~50tps. That's slower than Claude Opus. https://openrouter.ai/deepseek/
57.
▲
by
RussianCow
3mo ago
Same with cars. Half the battle is sometimes just unscrewing an old bolt that hasn't been touched in 10+ years without breaking it, or getting the rusted on rotors to come off.
58.
▲
by
RussianCow
4mo ago
> The models themselves are the problem -- most large US companies are not going to touch them. Can you expand on this?
59.
▲
by
RussianCow
4mo ago
By "cost" I think the parent means the provider's own costs, not the cost of inference to the customer. The cost of land, labor, and electricity are significantly lower in China than in the US.
60.
▲
by
RussianCow
4mo ago
> Qwen 3.7 models run by providers in EU, US, & Singapore that OpenCode FAQ claims don't use retained data for training. Note: Alibaba Cloud is the only company that currently offers Qwen 3.7 models (they haven't released a
More ›