Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
thereitgoes456
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
26 ms
·
1.
▲
by
thereitgoes456
4d ago
Why so brazenly confident? Isn’t it possible that the benchmark is correct, and your experience is correct too, but you haven’t tried all the thousand different modalities of work that programming encompasses and so maybe you don’t actually
2.
▲
by
thereitgoes456
6d ago
Your startup is based around AI coworkers, so I expect you are biased towards overvaluing their usefulness.
3.
▲
by
thereitgoes456
6d ago
It is rivaling Astra, on their own benchmark that they made (FrontierCode), that they ran themselves in their own closed-source ecosystem that isn’t reproducible by anyone.
4.
▲
by
thereitgoes456
6d ago
The Cursor acquisition shows that it’s possible for these valuations to be justified. But Cursor was more successful and bent the truth much less. While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinki
5.
▲
by
thereitgoes456
8d ago
It seems obvious what GP meant. It is, once again, an explicit construction (“disproving” that every initial state does not develop a singularity).
6.
▲
by
thereitgoes456
8d ago
He asked whether they used their chats as training data and received no response. Any speculation here seems quite appropriate?
7.
▲
by
thereitgoes456
8d ago
I’m open to evidence, but just using Bayesian reasoning, OpenAI is one of the most dishonest companies in history. They’re currently being sued for a dozen employees stealing Apple hardware! I don’t understand why I should give them any gra
8.
▲
by
thereitgoes456
12d ago
Because playing a video game isn’t relevant to which AI people might want to use.
9.
▲
by
thereitgoes456
12d ago
Of course. But that market won’t produce a $800 billion company, unless AI becomes ludicrously widespread — energy is used every day by virtually every person on the planet, and of course has plenty of mass consumption & “luxury” custom
10.
▲
by
thereitgoes456
13d ago
I see, it’s a great point. I know some evals actually do use LLMs as a judge (e.g. those that try to measure debate skill), though the ways AI can try to cheat its way through every benchmark now are astoundingly varied.
11.
▲
by
thereitgoes456
13d ago
You’re not seriously suggesting that the model is secretly sandbagging its performance on GDPval and long context reasoning, while making huge and obvious progress on ExploitBench, ARC and science benchmarks, in order to tank its AA composi
12.
▲
by
thereitgoes456
14d ago
Really generous. I'm on the Pro plan and I just use Antigravity for vibe coding w/o automation. It's actually difficult to hit my weekly limit now, it takes about ~30-35 hours of continuous agent work, which virtually only ha
13.
▲
by
thereitgoes456
1mo ago
I've been really satisfied with it since 3.6, it's been "good enough" for the tasks I'm using and has fast response and very high limits, much higher than Claude Code. I continually don't understand how nobody
14.
▲
by
thereitgoes456
1mo ago
There is no evidence that Anthropic's revenue is 70B.
15.
▲
by
thereitgoes456
1mo ago
Oh, come on. He's not perfect, but compared to CEOs and researchers predicting AGI and complete economic upheaval every 3 months, he's looking very good. He's no more a charlatan than Altman is. His predictions (while often w
16.
▲
by
thereitgoes456
1mo ago
I largely agree but I don't know if it's quite so clear cut. From the pricing angle, all competitors except Google, including Chinese models, have incentive to gain market share at all costs, and may be serving tokens at or below
17.
▲
by
thereitgoes456
2mo ago
95% of SpaceX's valuation is in AI.
18.
▲
by
thereitgoes456
2mo ago
I think they're concentrating on users and not benchmarks. Gemini is still the best model at plain old searches, like, for parsing images or "what's that poem with this and this in it". It seems to have much more breadth
19.
▲
by
thereitgoes456
2mo ago
Is that surprising? It is standard "embrace, extend, extinguish" from a company not in a strong enough position to do the third one.
20.
▲
by
thereitgoes456
2mo ago
DeepMind did release a math specific model. And OpenAI has released a coding specific model. The answer to your question is “because the market isn’t big enough”, not because it doesn’t work. Why would knowing about 2019 internet memes help
21.
▲
by
thereitgoes456
3mo ago
Because it's an important aid to getting to a high-paying job in the US, not just a means to learn. One need only look at the resume filtering process, a once manual bias that has now been codified into algorithmic bias with AI. A degr
22.
▲
by
thereitgoes456
3mo ago
So? It’ll never pay dividends..
23.
▲
by
thereitgoes456
3mo ago
That’s absurd. Why couldn’t it still fail, especially when their last raise was at 20x revenue or more? These numbers are horrendous.
24.
▲
by
thereitgoes456
3mo ago
Sam Altman is not one of those people. But other founders certainly felt that way.
25.
▲
by
thereitgoes456
3mo ago
Anthropic's S-1 is not public yet
26.
▲
by
thereitgoes456
4mo ago
Anthropic talks to the Pope and hires ethicists and philosophers. All founders have pledged to donate 80% of their wealth. They have pledged to never use ad tech because of misaligned incentives. There is an independent board. Meanwhile Gre
27.
▲
by
thereitgoes456
4mo ago
Stargate is not real. It is not clear that running one's own datacenter is a competitive advantage. Why do you think OpenAI can handle that?
28.
▲
by
thereitgoes456
4mo ago
Many people worked for FTX, loved SBF and would have said he was a great leader too. People will love anyone who is good to them.
29.
▲
by
thereitgoes456
4mo ago
This is common rhetoric that feels overly reductionist and makes me sad. Sam got fired and his response was to manipulate and pressure his employees into a shameful, cult-like letter, and play the media to character assassinate Toner as bei
30.
▲
by
thereitgoes456
5mo ago
I assume you're trolling, but in case otherwise: The reason media training exists is to win over people like you. I know people high up at OpenAI, I'm quite sure Altman etc. don't care either; the only difference is, they
More ›