Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
meander_water
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
meander_water
4d ago
> I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom translation: I'm g
2.
▲
by
meander_water
5d ago
This is not being against technology. It's ensuring that technology is used to serve humans, and not the other way around. https://caiml.org/dighum/dighum-manifesto/
3.
▲
by
meander_water
6d ago
What I would love to see is a side by side comparison of two companies (one using AI, and one without) over an extended period of time. Id wager that they would fare the same. The time saved on some tasks are outweighed by extra time spent
4.
▲
by
meander_water
6d ago
A humanist revolution. Sure we can't stop the technology, but we _can_ stop the importance and money heaped into it.
5.
▲
by
meander_water
6d ago
I see a lot of people throw their hands up in the air and say "I have no control over this, might as well go all in on AI". But we as individual consumers have more power than we think. We don't need to use these AI features
6.
▲
by
meander_water
6d ago
You trust machine code because it was produced by a deterministic system that was written and tested thoroughly by other engineers. AI systems are not the same. You can't guarantee deterministic output. Also, code is still the best lan
7.
▲
by
meander_water
13d ago
The OpenAI charter defines it as: "highly autonomous systems that outperform humans at most economically valuable work" https://time.com/article/2026/08/26/openai-sam-altman-interv...
8.
▲
by
meander_water
14d ago
It's an effective approach. Google project zero started doing this in 2024 https://security.googleblog.com/2024/11/leveling-up-fuzzing-...
9.
▲
by
meander_water
21d ago
> I wish the world could get the benefits rapidly and delay the problems it will cause as long as possible, but the benefits and problems are arriving at the same time. I'm not trying to be facetious, but what are the benefits of &q
10.
▲
by
meander_water
1mo ago
> First, the desire to create a self-improving AI that is capable of recursively bootstrapping itself to AGI/ASI. I feel like we've become blinkered in this quest to push the frontier at all costs. Somehow the target has shifte
11.
▲
by
meander_water
1mo ago
I think there's a corollary to this. Not only is it hollowing out the middle class, it's preventing the junior engineers from stepping up to the senior level. Everyone starts off as a bad engineer. Just like any other profession,
12.
▲
150M-parameter reasoning model sets new cost-accuracy frontier on ARC-AGI-1
(huggingface.co)
3 points
by
meander_water
1mo ago
|
0 comments
13.
▲
Jeff Dean Leaves Google
(twitter.com)
2 points
by
meander_water
1mo ago
|
1 comments
14.
▲
by
meander_water
2mo ago
This is what I was looking for, thanks!
15.
▲
by
meander_water
2mo ago
Sure, but then you would just use an image generation or multimodal model to generate that image. I don't think you'd want a weird looking svg.
16.
▲
by
meander_water
2mo ago
Can someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.
17.
▲
by
meander_water
2mo ago
Firstly, I don't have many issues with benchmarks per se. But I do have issues with leaderboards. And the AA index is touted by lots of people to argue that X model is better than Y, which I find inaccurate. > I mean, it's tell
18.
▲
by
meander_water
2mo ago
The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains
19.
▲
Substack adds AI text detection to all notes and posts
(post.substack.com)
4 points
by
meander_water
2mo ago
|
0 comments
20.
▲
by
meander_water
2mo ago
1. The benchmark is run with a python script - https://github.com/sunblaze-ucb/exploitgym using an agent harness. I suspect they used codex. So the model has access to the environment and could trivially inspect its ow
21.
▲
by
meander_water
2mo ago
Seems similar to Openrouter Fusion - https://openrouter.ai/docs/guides/routing/routers/fusion-rou...
22.
▲
AI Compass: which archetype are you?
(bambamramfan.github.io)
9 points
by
meander_water
3mo ago
|
1 comments
23.
▲
Meta Posed as Teens to Prompt Rival Chatbots About Suicide, Sex, and Drugs
(wired.com)
28 points
by
meander_water
3mo ago
|
8 comments
24.
▲
by
meander_water
3mo ago
I thought all model providers are doing this under the hood anyway in their UI? They certainly seem to when A/B testing different models, and Fable routes to Opus 4.8 when guardrails fail. Also, openrouter recently released a fusion ro
25.
▲
by
meander_water
3mo ago
GPTZero is much better at handling humanized outputs. Also has a similar false positive rate to Pangram.
26.
▲
by
meander_water
3mo ago
> However, it’s your job to go down the rabbit hole, learn the 100%, and sprinkle in your 3%. I would say that there is a big difference between stealing without acknowledgement, and stealing with acknowledgement and actively learning th
27.
▲
by
meander_water
3mo ago
> I don't think you should waste time reviewing every single line of code in here and just use AI to review it! > What you bring is the knowledge that the author nor the LLM doesn't know. How can you possibly know what relev
28.
▲
by
meander_water
3mo ago
Thanks, I didn't mean to be brusque, but I have seen a lot of these vibe tests lately that come to grand conclusions like "X model is better than Y" from the result of a single prompt. Appreciate you sharing the results of yo
29.
▲
by
meander_water
3mo ago
> So we ran it head-to-head against Claude Opus 4.8: same one-shot prompt, build a 3D platformer in raw WebGL from scratch Running a single one-shot prompt is not a benchmark, not is it representative of any sort of real-world usage. Mos
30.
▲
Machine Studying
(jacobxli.com)
4 points
by
meander_water
3mo ago
|
0 comments
More ›