Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
anuramat
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
anuramat
1mo ago
do you think typing `make` into a terminal is some sort of an insane intellectual achievement?
2.
▲
by
anuramat
2mo ago
why can't they just increase the price for cached input instead?
3.
▲
by
anuramat
2mo ago
could it be that the leading AI lab is decent at making LLMs?
4.
▲
by
anuramat
2mo ago
what's there to hack that you couldn't do with $3 in openrouter credits?
5.
▲
by
anuramat
2mo ago
> productivity findings are you referring to the early 2025 METR study?
6.
▲
by
anuramat
2mo ago
> you can run a small model but why?
7.
▲
by
anuramat
2mo ago
I think you mean mostly-stateless-but-with-prompt-caching-and-batching, ie not really
8.
▲
by
anuramat
2mo ago
you think they're doing inference at a loss even with the API prices?
9.
▲
by
anuramat
2mo ago
yes; fyi usage limits on the $200 claude sub correspond to at least $1.2k/week in api tokens
10.
▲
by
anuramat
2mo ago
"benchmaxxing by generalizing" is not really benchmaxxing
11.
▲
by
anuramat
2mo ago
> Russia lmao
12.
▲
by
anuramat
2mo ago
> possibility of "LLMs can reason not like a human" what would be the difference between not reasoning and reasoning not like a human?
13.
▲
by
anuramat
2mo ago
I keep asking the same question, and I think the steelman version would be "has metacognitive patterns similar to humans"
14.
▲
by
anuramat
2mo ago
I'm somewhat serious -- if you think AI will scale that well, you can't really make predictions like that I personally don't think the weight efficiency will improve that much; if anything big does happen, I expect it to be a
15.
▲
by
anuramat
2mo ago
just found a decent looking benchmark for iterative development: https://swe-milestone.com/ surprised it isn't a bigger thing, eg artificial analysis doesn't report anything like that still doesn't measure th
16.
▲
by
anuramat
2mo ago
> nowhere near readteaming I think any security-related task triggers it to think about the threat model and thus hit the guardrails
17.
▲
by
anuramat
2mo ago
what languages are you working with? I imagine if it's something like C, you'd hit guardrails every time you manage memory
18.
▲
by
anuramat
2mo ago
is there a way to use it with lazygit? unfortunately I'm addicted atp, and "autorebase branches" is exactly what I was missing the entire time
19.
▲
by
anuramat
2mo ago
why would you need a local fable at that point? AGI will surely solve all the problems in the world at that point
20.
▲
by
anuramat
2mo ago
what are you working on? I only hit the guardrails twice after burning through two weeks of 20x max plan, both times on ML stuff; still more than I'd want to, but not unusable
21.
▲
by
anuramat
2mo ago
ml research, both brainstorming and running experiments eg I'd throw a hypothesis at it in the evening, and overnight it would write the code, do a sanity check, start a run, monitor the metrics, identify and fix a bug, propose a new h
22.
▲
by
anuramat
2mo ago
security researchers are using the free tokens that they get to do useful security work, AI labs are giving away free tokens to maximize their profits; is it really that hard to imagine that different parties might have different goals? >
23.
▲
by
anuramat
2mo ago
so, git?
24.
▲
by
anuramat
2mo ago
> it's not actually a financially efficient way ... unless you profit from creating demand for LLMs well, they do? it's a win-win, you can't really criticise an AI lab for doing AI instead of straight up giving money to se
25.
▲
by
anuramat
2mo ago
I imagine one could one-shot a basic app and then feed feature requests one by one, sounds like an obvious way to benchmark architecture/maintainability
26.
▲
by
anuramat
2mo ago
> data centers being built in their communities > golf courses are a traditional green space where people in a community I have a feeling those two sets of communities are disjoint
27.
▲
by
anuramat
2mo ago
if you don't make your own tech, standards will be defined by people that do
28.
▲
by
anuramat
2mo ago
whats the point of a commit message if it can be inferred by an agent from the diff?
29.
▲
by
anuramat
2mo ago
why?
30.
▲
by
anuramat
2mo ago
the article implies that the perfect balance is 1:1; I want to believe this is just some sort of a ragebait-based PR strategy
More ›