Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mlpro
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
mlpro
8mo ago
Novel Ideas are never cheap, lol.
2.
▲
by
mlpro
9mo ago
Lol. trying to copy the Universal Weight Subspace paper's naming to get famous.
3.
▲
by
mlpro
9mo ago
Lol, yeah.
4.
▲
by
mlpro
9mo ago
Oh, look - a new 3D model with a new idea - more data.
5.
▲
by
mlpro
9mo ago
I don't understand.
6.
▲
by
mlpro
9mo ago
Waymo should do a bit more research in reliability and explainability of their AI models.
7.
▲
by
mlpro
9mo ago
Read the paper end to end today. I think its the most outrageous ideas of 2025 - at least amongst the papers I've read. So counterintuitive initially and yet so intuitive. Personally, kinda hate the implications. But, a paper like this
8.
▲
by
mlpro
9mo ago
They are not trained on the same data. Even a skim of the paper shows very disjoint data. The LLMs are finetuned on very disjoint data. I checked some are on Chinese and other are for Math. The pretrained model provides a good initializatio
9.
▲
by
mlpro
9mo ago
I think its very surprising, although I would like the paper to show more experiments (they already have a lot, i know). The ViT models are never really trained from scratch - they are always finetuned as they require large amounts of data
10.
▲
by
mlpro
9mo ago
Why would they be similar if they are trained on very different data? Also, trained from scratch models are also analyzed, imo.
11.
▲
by
mlpro
9mo ago
It's about weights/parameters, not representations.
12.
▲
by
mlpro
9mo ago
The analysis is on image classification, LLMs, Diffusion models, etc.
13.
▲
by
mlpro
9mo ago
It does seem to be working for novel tasks.
14.
▲
by
mlpro
9mo ago
Not really. If the models are trained on different dataset - like one ViT trained on satellite images and another on medical X-rays - one would expect their parameters, which were randomly initialized to be completely different or even orth