Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kz919
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Cautious Optimizers: Improving Training with One Line of Code
(arxiv.org)
1 points
by
kz919
2y ago
|
1 comments
2.
▲
by
kz919
2y ago
"AdamW has been the default optimizer for transformer pretraining. For many years, our community searches for faster and more stable optimizers with only constraint positive outcomes. In this work, we propose a \textbf{single-line modi
3.
▲
by
kz919
2y ago
obviously because they can't beat it... There will be zero reason to use it when you have better transformer based models that can fit the existing infrastructure.
4.
▲
Adept.ai Founders Left for Amazon
(adept.ai)
2 points
by
kz919
2y ago
|
2 comments
5.
▲
by
kz919
3y ago
you are not wrong, but I don't think that is their point. It's just a muscle flex to shows you that their hardwares work. If it can train 176B or whatever, it should be an easy peasy for them to train 13B.
6.
▲
by
kz919
3y ago
you just need to teach it to use a calculator https://arxiv.org/pdf/2302.04761.pdf
7.
▲
by
kz919
5y ago
I wonder how consistent the chip output would be. Also compiler and software stack sound painful to build.