Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
samusiam
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
samusiam
7d ago
Which is pretty much a useless (i.e., saturated, contaminated) benchmark now.
2.
▲
by
samusiam
4mo ago
I'm still using Sync, going on three years since they banned third party apps, and third party apps are still the best way to experience the site.
3.
▲
by
samusiam
5mo ago
I don't see anything asking for a credit card.
4.
▲
by
samusiam
5mo ago
In my experience this "gaming" behavior is easily caught by just asking another agent (could just be another session of Claude Code) to review the code changes.
5.
▲
by
samusiam
5mo ago
For idle sessions I would MUCH rather pay the cost in tokens than reduced quality. Frankly, it's shocking to me that you would make that trade-off for users without their knowledge or consent.
6.
▲
by
samusiam
6mo ago
That's true, but the "AI bubble bursts" scenario is usually tied to Western investors getting essentially margin-called. If that happens, the CCP won't suddenly stop their investment; Chinese models will most likely cont
7.
▲
by
samusiam
6mo ago
> I have no clue if this claim holds, but alas, just pretending they did not address the obvious criticism, while they did, is at the very least pretty lazy. But they didn't address the criticism. "cutting ~75% of tokens while
8.
▲
by
samusiam
6mo ago
> Aren't they currently propped up by investor money? Are Chinese model shops propped up by investor money? Is Google? Open weights models are only 6 months behind SOTA. If new model development suddenly stopped, and today's SO
9.
▲
by
samusiam
6mo ago
These OSS model makers need to stop benchmarking against old models. Showing how it performs against Opus 4.5, GLM-5 when we have Opus 4.6 and GLM-5.1 just tells me that it's not comparable to SOTA.
10.
▲
by
samusiam
6mo ago
I think there's a lot of methodological expertise that goes into collecting good eval data. For example, in many cases you need human labelers with the right expertise, well designed tasks, well defined constructs, and you need to hit
11.
▲
by
samusiam
6mo ago
Well you're free to disagree but my experience has been counter to your position. I write both code and research / technical documentation. The quality of what the LLM produces is limited by the quality of ideas I give it initiall
12.
▲
by
samusiam
6mo ago
You're complaining about vibe coding while also complaining about how you "feel" about the code. Do you see the irony in that?
13.
▲
by
samusiam
6mo ago
I haven't seen the scrolling glitch in months, where previously it was happening multiple times a day. Also haven't seen anyone complain about it in quite some time. Pretty sure they have resolved that.
14.
▲
by
samusiam
6mo ago
I just checked competitors' codebases: - Opencode (anomalyco/opencode) is about 670k LOC - Codex (openai/codex) is about 720k LOC - Gemini (google-gemini/gemini-cli) is about 570k LOC Claude Code's 500k LOC doesn&#x
15.
▲
by
samusiam
6mo ago
AI witch-hunters are even more annoying.
16.
▲
by
samusiam
6mo ago
> I think the only reason it’s seen as good anywhere is there are a lot of tasteless and talentless people who can pretend they created whatever was curled out. This goes for code as well. This is an oversimplification. If you have taste
17.
▲
by
samusiam
6mo ago
"by design, the recommendations will be average" This couldn't be more wrong. The simplest refutation is just to point out that there are temperature and top-k settings, which by design, generate tokens (and by extension, id
18.
▲
by
samusiam
6mo ago
"they don't output anything unless prompted" Unprompted they're not unlike a human sleeping or in a coma. Those states don't preclude consciousness in other states.
19.
▲
by
samusiam
6mo ago
Vegan for 15 years. I cook 95% of my own meals, including black bean burgers, tofu, etc... Sometimes I want something that tastes like meat and I reach for a Beyond or Impossible burger. I don't need it. But I can't recreate its t
20.
▲
by
samusiam
7mo ago
I can recall reading human-authored text like this for more than a decade.
21.
▲
by
samusiam
7mo ago
But literally any decent agent can recommend existing services and help you set them up. And even help you help them set the services up for you. I do this with Claude all the time.
22.
▲
by
samusiam
7mo ago
IMO we should all be asking for a raise if our company is making more money. Proportionally, even.
23.
▲
by
samusiam
7mo ago
> We don't sit around and write specs and then hope working code plops out. So what do you do then? Sit around hand-holding an AI agent while it implements code line-by-line? I'm being facetious, but my point is that if you
24.
▲
by
samusiam
7mo ago
I agree some skepticism is warranted. However I think we need to avoid essentialist thinking about what effects AI usage will have on a person. > that's definitely not the normal usage The way I look at it, AI use -- proper AI use -
25.
▲
by
samusiam
7mo ago
What primary sources are you referring to? Come with receipts next time instead of just vitriol.
26.
▲
by
samusiam
7mo ago
It doesn't have to reduce understanding. It completely depends on how you use it. See, for example, this study https://arxiv.org/html/2601.20245v2
27.
▲
by
samusiam
7mo ago
Which is so much better because you can do other terminal stuff and you can avoid vendor lock in.
28.
▲
by
samusiam
7mo ago
Did you tell it to consider future usage? Have you tried using it to find and remove dead code? In my experience you can get very good code if you just do a few passes of AI adversarial reviews and revisions.
29.
▲
by
samusiam
7mo ago
But it's so easy to just download from libgen and send as an attachment to a Kindle email...
30.
▲
by
samusiam
7mo ago
Not only that, but the average reader will interpret the title to reflect AI agents' real-world performance. This is a benchmark... with 40 scenarios. I don't say this to diminish the value of the research paper or the efforts of
More ›