Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
redox99
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
redox99
4d ago
Terminal bench 4 is good largely because it's recent so it hasn't been benchmaxxed yet. It's more of a sysadmin/devops benchmark than a coding benchmark though, but still a decent proxy. https://artificialanal
2.
▲
by
redox99
4d ago
They may allocate different number of resources every year based on market conditions but they'll never give up on gaming, that would be extremely silly.
3.
▲
by
redox99
5d ago
I think in a few decades when there are more humanoid robots than humans we'll likely have skynet, so I'm pessimistic in a way lol. But in the near future and on a personal level I think we'll need to adapt and pivot but we&#
4.
▲
by
redox99
5d ago
I think it will be like how a lot of people know how to code in python but have zero understanding of assembly or how a cpu works. As a researcher you'll accept there's this low level stuff that if you want you can dig into but i
5.
▲
by
redox99
5d ago
Of course you're not going to get rich with the kind of software that LLMs can one shot these days. But that kind of software like to-do lists or basic CRUD have been saturated for over a decade, way before LLMs. People overestimate ho
6.
▲
by
redox99
5d ago
This has always happened, way before AI. You'd spend months or years building and growing your business, and then Google would release a feature or product that would kill your business overnight because they can throw way more money a
7.
▲
by
redox99
5d ago
> Many graduate students (I know) are having a crisis if any of their research worth it? If AI can (or will) do everything, what's the point of doing experiments and all? This will eventually deter a whole generation of curious mind
8.
▲
by
redox99
5d ago
The two times I tried to use fiverr I literally got ghosted.
9.
▲
by
redox99
5d ago
This is orders of magnitude cheaper and faster than paying for a song to be created from scratch by a human.
10.
▲
by
redox99
6d ago
Yeah I tried GPT first because that's what I always use, it refused and I didn't even bother trying to trick it, just went straight to Grok.
11.
▲
by
redox99
6d ago
Can you use Lean to... prove "Lean-fast" is equivalent to Lean?
12.
▲
by
redox99
6d ago
Btw grok found out in less than ten seconds what place you refer to and likely what software within that company. Things that wouldn't be worth your time before are trivial these days with LLMs.
13.
▲
by
redox99
7d ago
You can argue all day whether cars are good or bad. That's not my point. My point is that pretending public transport is equivalent to cars if you have good enough infrastructure is silly. They are different experiences, each with thei
14.
▲
by
redox99
7d ago
Unnecessary is a very strong word. It's also unnecessary to have your own private house, sharing it with other people would be more efficient.
15.
▲
by
redox99
7d ago
Public transport is not the same as cars. It's silly when people pretend like they are interchangeable and it's just a matter of having "better infrastructure".
16.
▲
by
redox99
7d ago
With the miles you can establish a confidence interval of the true fatality rate.
17.
▲
by
redox99
7d ago
Astra is definitely weird. It is more capable than Sol, no doubt about that. There are things sol could simply not solve that Astra breezes through. However for typical low to medium difficulty code, it will often either overengineer stuff,
18.
▲
by
redox99
7d ago
Just use opencode/pi/etc with an open weight model?
19.
▲
by
redox99
8d ago
Thanks for the reply. I'd suggest you improve your workflow with better tooling like Codex or Claude code, and if you have some kind of weird constraint they'll easily follow an AGENTS.md with that.
20.
▲
by
redox99
8d ago
Those are really nice numbers. With that t/s, no network latency or queueing it must feel much snappier than cloud models.
21.
▲
by
redox99
8d ago
Do you... code by copy pasting chatgpt web???
22.
▲
by
redox99
8d ago
Why not? It makes less mistakes and with subscriptions it's very cheap
23.
▲
by
redox99
8d ago
Stuff along the lines of implement controller service and tests for the following endpoints: - list of many endpoints with the JSON they receive and return and description of what they need to achieve Stuff you could probably do in a single
24.
▲
by
redox99
8d ago
Your prompts are probably very underspecified then. Frontier models one shot the majority of my prompts. UI is kind of the exception, there I do have to ask for a lot of tweaks.
25.
▲
by
redox99
8d ago
I think pi handles it better
26.
▲
by
redox99
8d ago
Anything you do right now? A typical 10 minute prompt "simply" becomes about 7 hours long. (40t/s vs 1t/s).
27.
▲
by
redox99
8d ago
I hate so much how every food menu is now AI slop. I wish that was outlawed under false advertising.
28.
▲
by
redox99
8d ago
At 1t/s it's still faster than humans for a lot of tasks, basically doing overnight what could take humans half a week. Plus you can always parallelize.
29.
▲
by
redox99
8d ago
The amount of goalpost moving is insane. "Yeah it can solve Millenium problems, but can it do it with nothing more than a one sentence prompt?" Also there are proofs where the only human steering was "keep going".
30.
▲
by
redox99
8d ago
Fields that allow verification, like math, will far surpass human level because they don't need human data for training. It's exactly the same as with Chess
More ›