Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
astro1234
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
astro1234
12d ago
Well I think in the absence of convincing pieces of evidence to the contrary you might be right. You’re making an empirical statement but we have already answered it today: - we get novel, emergent properties and capabilities of these model
2.
▲
by
astro1234
13d ago
I think you make a fair point, but also remember: new paradigms don't necessarily require confusing / contradictory observations. You could have the simple idea of "what if gravity is an inertial force?" at any period in
3.
▲
by
astro1234
13d ago
Do you mean after? People do this!! But I think it’s a bit different. It won’t be apples to apples because the data volume I think is just so much different. Maybe there are good experiments for something like this.
4.
▲
by
astro1234
13d ago
I think they mean improve our understanding of physics with new theoretical results or paradigms. Like if it’s 1899, would Astra develop General and Special relativity on its own?
5.
▲
by
astro1234
13d ago
My job has gone from coding, plotting, writing to solving the riddle of what Opus 5 is saying and figuring out what is bullshit, what is valuable, what is a total divergence from what I asked it to do. Then eventually at token like 300k sta
6.
▲
by
astro1234
15d ago
I don't think it's a stall, two ways I would believe there is a stall: - Does the epoch capability index progress show signs of plateauing? I consider this a good aggregate measure of diverse benchmarks into a single capability in
7.
▲
by
astro1234
22d ago
I’ve noticed this too but it hasn’t been obvious to me that this style of search is not a learned behavior. Tool calling is very much part of the post training phase, I would expect that these style searches just naturally emerge during tra
8.
▲
by
astro1234
29d ago
The Opus-human patois is unbearable
9.
▲
by
astro1234
1mo ago
Im curious if you find this to be a parody in a bad way or simply a “the state of the art in math research right now is telling a machine to believe in itself”. I am in the latter camp…
10.
▲
by
astro1234
1mo ago
Oh my god do not trigger me about the news reporting “points” for the DOW or the S&P, I have seriously considered calling news stations about this. Or not providing inflation adjusted values when referring to historical financial refere
11.
▲
by
astro1234
1mo ago
Yes! I am aware of this and they are really great. Problem is nobody will ever read that, but they WILL watch Fox News, CNBC, etc. and those outlets never even point out that there are errorbars/uncertainty. Then I hear things like o
12.
▲
by
astro1234
1mo ago
Oh I’m well aware the BLS knows what they’re doing and publish error bars, they are great. The problem is news articles, no one ever reads BLS reports but they do watch Fox News, CNBC, etc. and those outlets never ever get into uncertainty.
13.
▲
by
astro1234
1mo ago
As a data scientist it is very frustrating to see a consistent lack of error bars on these numbers or any discussion about the magnitude of uncertainty around them by all news reports every time these numbers get released
14.
▲
by
astro1234
1mo ago
I agree and that’s why we need and indeed have an ever evolving landscape of benchmarks > Private, refreshed test sets attack the mechanism itself, and in my view they are the only intervention that does. If the questions have never touc
15.
▲
by
astro1234
1mo ago
I don’t think that’s fair to say, there are two principle sources of data that paint a fairly consistent picture - one is scaling laws, where we found years ago that pretraining validation loss scales in an almost miraculously predictable w
16.
▲
by
astro1234
1mo ago
Yea this may explain part of it or all of it, it’s likely a case by case kind of thing. Also to respond to the parent comment: benchmarks have a variety of difficulty levels. Humanity’s Last Exam, though now hitting the beginning of a satur
17.
▲
by
astro1234
1mo ago
I work in AI evaluation, lots of problems and leakage is an issue as is ecological validity, but they definitely do not explain the progress we see. I think Epoch has the best analysis I’ve seen on evaluation trends; they use IRT to basical
18.
▲
by
astro1234
2mo ago
I find them useful bellweathers of genuinely out of domain performance and capability, regardless of their theoretical importance. What I see is that performance trends are remarkably stable both upstream (miraculous scaling laws of pretrai
19.
▲
by
astro1234
2mo ago
I think the question and definition game is interesting only inasmuch as it helps us understand ourselves (what actually explains some of the mysterious properties of our perceived consciousness) or helps guide us towards improving performa
20.
▲
by
astro1234
2mo ago
Yes, I was hoping they would actually offer some kind of partnership with the sports leagues to stick a few of their cameras in stadiums and offer live sports games, but I don't think there's nearly enough users to support it (at
21.
▲
by
astro1234
2mo ago
It is absolutely unreal the experience watching the ~10 minute sports games on the Vision Pro. The thing is 4 grand but that’s how much the first 4K TVs were. Hope the price can drop where this becomes a more widespread experience, I am a v
22.
▲
by
astro1234
2mo ago
Of course but I think labs is a good term. Like a lot of terms like this it comes from history. The people working on AI at these places come predominantly from academia and predominantly do research. Of course now these companies are far m
23.
▲
by
astro1234
2mo ago
I don't have personal experience with this -- but if I were in the situation where my job was jeopardized by better price efficiencies, I don't know that I personally would have a problem with that. If Ukrainians can do my job jus
24.
▲
by
astro1234
2mo ago
Immigration having downward pressure on salaries is a common and demonstrably false misconception. Immigration adds more members to the labor pool while also increasing demand . The two, in most major studies, cancel each other out. "
25.
▲
by
astro1234
2mo ago
Huge fan of Claude for a long time but switched now to Codex after all of this "ok one more week, one more week" stuff. It feels like Sol is quite capable, and the nice thing is the refusal rates seem more reasonable, _especially_
26.
▲
by
astro1234
2mo ago
I think the biggest phase transition happened roughly fall of last year with AI adoption. My entire job is now AI orchestration, and then a LOT of my time spent writing the specs/prompts, reviewing and validating the output imperfectly
27.
▲
by
astro1234
2mo ago
In my experience Data Science looks very little like it used to a few years ago, and the priority skill these days is good strong understanding of the basics and very good sense of judgement. To me, statistics is the absolute number one pri
28.
▲
by
astro1234
2mo ago
I don’t mean to say you are wrong, I just mean to ask what criteria you would consider necessary to satisfy in order for you to consider a system to be reasoning.
29.
▲
by
astro1234
2mo ago
I tend to agree but to adjudicate that someone has to define what they consider reasoning to be.
30.
▲
by
astro1234
2mo ago
I’m curious what this means? I think the evidence is pretty convincing that, while brittle, there is reasoning going on (though it depends on your definition of reasoning which I’m curious what that is for you).
More ›