Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
neosat
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
1.
▲
by
neosat
5d ago
>> "are these companies interested in developing research" judging from the money, resources spent and the value they derive from this the answer is very definitively yes. What makes you think these companies (and I'm n
2.
▲
by
neosat
5d ago
You can couch it different words but the basic shape of all these is the same, whether it was voice artists earlier, IT outsourcing, or now disciplines like mathematics. A small set of people (relatively) who were the primary source of gett
3.
▲
by
neosat
15d ago
Can you or someone else from A\ comment on whether the conversation style is coming to Opus 5 or a future 5.1 asap as well? Currently it seems the model has been made unusable by the way it 'speaks' and there is a clear solution w
4.
▲
by
neosat
28d ago
Sorry but the authors of this article probably don't fully 'get it' (or at least cannot explain it well) either if this is their conclusion. "In other words, Stripe did not buy a router. It bought a strategic frontier AI
5.
▲
by
neosat
1mo ago
As the above two comments mentioned this is not true in practice due to batch effects (you can read about some interesting work published by Thinking Machines on this), as well as calculation drift that happens across computations esp. now
6.
▲
by
neosat
1mo ago
This may not 'quite a bit behind' those at all. If you look at the benchmark numbers they are very comparable to Fable, but beyond a certain point the benchmark numbers don't tell you much. Opus #5 beats Fable on some benchma
7.
▲
by
neosat
2mo ago
Yes I had a similar experience with Opus 5. It is very token efficient, fast, and gets reasonable part of the work right but makes a LOT of mistakes. In a month+ use of Fable completed each task without ANY errors. Opus could not complete a
8.
▲
by
neosat
2mo ago
Definitely not my experience. Fable is better but I'd prefer K3 to Opus based my experience with both.
9.
▲
by
neosat
3mo ago
Good observations. There's definitely a trend in pricing increasing but also balanced by innovations and availability of other models (both open and closed) emerging as alternatives. It's natural for the labs to explore how much t
10.
▲
by
neosat
3mo ago
I've been using GLM 5.2 recently (company hosted, for non-coding tasks) and it's been strong and reliable. There are areas where GPT 5.5 and Opus 4.x still feel marginally better but only marginally. For most tasks if GLM 5.2 is t
11.
▲
by
neosat
3mo ago
Anthropic is really on a tricky path here. When you have had runaway success due to a hit it is easy to believe that it is the natural way of things. However, that happened due to unique convergence of tech paradigm shift, the competitive l
12.
▲
by
neosat
4mo ago
Agree. Audio has strongly temporal so there is almost certainly some positional encoding one way or another.
13.
▲
by
neosat
4mo ago
You need to see the response in light of the original discussion. Referencing here for clarity since I should have included it in the first place: "We used the claude code and codex harness and I implemented some prs they needed with g
14.
▲
by
neosat
4mo ago
That's a fair callout and I agree my statement was too general in just mentioning 'output', as you correctly pointed out. To define 'better' you would indeed need to agree on the dimensions you would evaluate candid
15.
▲
by
neosat
4mo ago
Your argument is fine but different from the claim the OP is making. You cannot simply make a claim that (model + harness) X is better than Y, but then have no discernible difference in the output. Subjectively, people might still prefer o
16.
▲
by
neosat
4mo ago
Exactly, I was confused too. The authors clearly mention what the parent comment talks about, albeit towards the end of the article, that the 'J' bundle meant that these firms were not set up for success once they 'caught up&
17.
▲
by
neosat
4mo ago
Revenue is not the right metric when you compare space trips to trips inside a city. The more relevant numbers are EBITDA, Operating cash flow, Profits.
18.
▲
by
neosat
4mo ago
has anyone done the math on: 1. cost to build out and run the data centers 2. cost of compute (hardware and energy) 3. depreciation of legacy GPU and thus value at the end of 3 years. And then compare the $45B revenue from Anthropic to see
19.
▲
by
neosat
4mo ago
That's true, I should have mentioned active. Actual params are closer to 12B-14B likely, given the 40GB VRAM usage.
20.
▲
by
neosat
4mo ago
Do you find the video understanding work there also to be 'silly little slop', or did you only look at the gifs on the page and not read about the understanding work in a 3B model? This is not ground-breaking by any means, but ach
21.
▲
by
neosat
4mo ago
If that's the case, a way to test the theory and understanding (assuming some parts of reservoir and signal channel can be reliably identified) would be to prune the high-confidence reservoir significantly reducing the model size while
22.
▲
by
neosat
4mo ago
"What slows down a team where agents do the implementation is the production of specifications precise enough for an agent to pick up and run. Roadmap, written down. Acceptance criteria, written down. The “what we actually want” forced
23.
▲
by
neosat
5mo ago
Agree with your points on the primary two questions and the circular argument in the original article. However, re: " How is it that atoms/electrons/photons suddenly start experiencing pain? What is it, in terms of atoms
24.
▲
by
neosat
5mo ago
Apart from a cool project, this evolved my perspective on what an MCP is, along with some cool architecture insights and inspiring ideas. Thank you!
25.
▲
by
neosat
5mo ago
"We are investigating an issue preventing users from reaching Claude.ai, and will provide an update as soon as possible." Who is We? I thought software engineers were going to be redundant and AI could do it all itself? (not to ta
26.
▲
by
neosat
5mo ago
Just refreshed and see 5.5 now - yay! Love the speedy resolution ;) Thanks folks, I'll complain faster next time....
27.
▲
by
neosat
5mo ago
Enterprise user here and still seeing only 5.4. Yesterday's announcement said that it will take a few hours to roll out to everybody. OpenAI needs better GTM to set the right expectations.
28.
▲
by
neosat
5mo ago
As a player myself, and having seen much higher level player than me, reading the spin from the ball rotation (and in fact trajectory) of the ball is a common (if advanced) skill. Sometimes the movement of the bat can be deceptive (since wi
29.
▲
by
neosat
5mo ago
Great work on the feature and sure I'll do that. :)
30.
▲
by
neosat
5mo ago
Tried it to automate something that was on my to do list for the day. I had blocked off a few hours for this and managed to get the agent working reasonably well (85%) of the way there in < 15 mins. The main remaining part is the poor do
More ›