Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stalfie
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
stalfie
2mo ago
Oh I wasn't talking about humans. I should probably have pointed that out. There's some scenarios like catastrophic crop failure and so on that might lead to that, but frankly I doubt it. I was referring more to everything else. C
2.
▲
by
stalfie
2mo ago
Sure, here are some citations: https://arxiv.org/abs/2401.03910?utm_source=chatgpt.com https://www.frontiersin.org/journals/psychology/articles/10.... https://arxiv.org/p
3.
▲
by
stalfie
2mo ago
"Climate conspiracy"? Like you, mean, the conspiracy of climate scientists to publish facts to the best of their understanding? I don't know what exact strawman you're arguing against, although I'm sure you can alwa
4.
▲
by
stalfie
2mo ago
No one knows how LLMs work. We know how the architecture works, but almost nothing about why. Saying "statistical next token prediction" tells you about as much about LLMs as saying "action potential thresholds" tells yo
5.
▲
by
stalfie
2mo ago
I was referring more to the fact that no one predicted that next token predicting Transformers would go so far. Not about "AI" in general.
6.
▲
by
stalfie
2mo ago
No one imagined LLMs in their current format, it was simply a result of discovering that scaling compute and tokens produced better and better results with the Transformer architecture. The inventors of the Transformer architecture were wor
7.
▲
by
stalfie
2mo ago
10 years is a long time. 10 years ago the Transformer architecture didn't exist. I would call it moderately unlikely at best. At the very least, I would say it's likely that development will require an entirely different skillet 1
8.
▲
by
stalfie
3mo ago
There's at least one benchmark that attempts to measure this, but it has been running for a year plus so it's quite infrequently updated now. https://fiction.live/stories/Fiction-liveBench-Mar-25-2025/o..
9.
▲
by
stalfie
3mo ago
Ironically, the article points out that the original authors publisher actually put out two DMCA notices to google last year, apparently with no effect. I guess DMCA takedowns are only for the big fish fighting the good fight against car pi
10.
▲
by
stalfie
3mo ago
Mmmm, not sure I agree with this, although this is a topic where we would have to do a lot of groundwork to formulate our positions precisely in order to ensure we're actually discussing the same thing. My counterargument is that verif
11.
▲
by
stalfie
3mo ago
Once again, Russia turns out to be the reason we can't have nice things. War truly is a waste for everyone involved. Now that Russia is also helping North Korea to launch satellites (one so far), expect everything to get worse in the f
12.
▲
by
stalfie
3mo ago
Well, I'd argue that this depends on the field you're investigating. Sometimes you have a way to identify objective reality and sometimes you don't. In mathematics the majority of the field is verifiable in this way. Coding a
13.
▲
by
stalfie
3mo ago
I guess so. Just to be clear, I was talking about post-training methods for reasoning models here, not pre-training. I think "model as a judge" should actually do okay as a "sentiment analysis" style reward for expressin
14.
▲
by
stalfie
3mo ago
One thing I wonder about hallucinations, is that it seems on the surface that it is an easy problem for RLVR to target. Since you're already generating enormous amounts of reasoning traces which are verified by correct answers, just ha
15.
▲
by
stalfie
3mo ago
The criticism is also similar to those faced by Theranos. Survivorship bias is always a factor when looking backwards.
16.
▲
by
stalfie
3mo ago
Excluding the cost of X-ray/CT/MRI machines, operating them, getting people to them and through them, sometimes injecting contrast, and sometimes dealing with side effects of said contrast, radiologists, I think. You can scale all
17.
▲
by
stalfie
3mo ago
It's worse then that unfortunately. Even when invasive tests are positive, and we think we caught a cancer early, we know from population statistics that the reality is that often nothing would have happened. So we don't even trul
18.
▲
by
stalfie
3mo ago
Well, the Fable guardrails breaks this argument, as when you get booted down to Opus 4.8 it still happily responds (as does most other models after a "I'm not a doctor but..." hedge). So you get to press the big red button an
19.
▲
A Plea to the Labs: Let the Models Diagnose
(tangent.bearblog.dev)
2 points
by
stalfie
3mo ago
|
3 comments
20.
▲
by
stalfie
3mo ago
I got so frustrated with Fable refusing to make any medical diagnoses, which is the most recent iteration of a longer trend that has bothered me for years, that I made a blog to express how annoyed I am. Posting it here in case anyone is in
21.
▲
by
stalfie
3mo ago
Honestly, I have yet to see any evidence of data leak from private sources. I think one of the better example is "simple-bench", which at least used to be a low-key benchmark that I would assume would have been saturated quickly i
22.
▲
by
stalfie
3mo ago
Update in case anyone reads this comment ever again. I have found that I trigger the guardrails any time I ask for medical Q&A as a doctor, be it ECGs, case reports, and so on. But if I phrase it like I'm the patient ("help me
23.
▲
by
stalfie
3mo ago
Tried to benchmark ECG interpretation capabilities, and I hit the guardrails no matter what I do. Incredibly frustrating that medical performance seems to be a victim of "biological risk" guardrails.
24.
▲
by
stalfie
3mo ago
This article describes how Transformers work, but not really how LLMs work. Explaining the underlying architecture gives you about as much insight into how a modern LLM behaves as an breakdown of neuronal biochemistry and a few pathways doe
25.
▲
by
stalfie
4mo ago
Last time I checked thoroughly (roughly two years ago), AI (in the form of small ML models) mostly outperformed radiologists in areas where the gold standard is "one level" above imagining wise. By that I mean that you train a mod
26.
▲
by
stalfie
4mo ago
The counterpoint is that every company ever has based themselves on human effort they never paid for (usually). The entire scientific endeavour for example. Standing on the shoulders of giants and so on.
27.
▲
by
stalfie
4mo ago
Ok, fair point about the lizardbrain jealousy, the choice of wording there was needlessly antagonistic. In my defense, even though that point might seem reductive, I don't mean it to be. I'd say the only reason words like"fai
28.
▲
by
stalfie
4mo ago
Well, is it true that they give back nothing? What about the compute? Pay checks? The value of the subscription you pay for? What about actual examples of things they have given back for free, like Whisper, which used to be SOTA and is stil
29.
▲
by
stalfie
4mo ago
I'd say both are equally soulless, dualism is a little bit of a philosophical dead end. Frankly, I don't particularly care much for the moral panic around capitalism. Capitalism has it's downsides for sure, but it's the
30.
▲
by
stalfie
4mo ago
That sentence could easily be applied to the human baseline.
More ›