6 ms·
When you see how good the output of the Luna model without reasoning is compared to SotA just a year and a half ago, it's pretty clear that it's been trained on
by stymaar 12d ago
When you see how good the output of the Luna model without reasoning is compared to SotA just a year and a half ago, it's pretty clear that it's been trained on explicitly.
- maleldil 12d agoI don't see how this is proof, and not just that the model got the better. You're comparing models a year and a half apart; this is a lifetime in LLM development.
- brookst 11d agoYeah, suggestive but hardly proof
- WarmWash 11d agoMany people still believe that LLMs can only reproduce what is in their training set.
- stymaar 11d agoThat'd be too restrictive, but to get improvement in a particular domain you definitely need to train specifically for it. Just cramming more random internet text in a bigger model has stopped being a effective way of scaling since at least mid 2023.
- maleldil 10d agoNo, you don't. There have been technological advances beyond just "more random Internet text", and they lead to more powerful models with more refined emergent behaviours.
- stymaar 11d agoIt's a lifetime on things that are explicitly being trained on! But small models like Luna didn't magically become more powerful than SotA models on stuff that they weren't explicitly trained on with a dedicated RL-pipeline.