6 ms·
Just to be clear: It was well known that you can reach such scores with small models and without an LLM if you train on the task. The author highlights those mo
by bonplan23 15d ago
Just to be clear: It was well known that you can reach such scores with small models and without an LLM if you train on the task. The author highlights those models himself - e.g. HRM/TRM.
The novelty is more that it works with such a plain transformer and low compute price.