6 ms·
that would help with decidable problems but would still be not generalisable for problems with non trivial rewards, or ones with none.
by marxplank 2y ago
that would help with decidable problems but would still be not generalisable for problems with non trivial rewards, or ones with none.
- astrange 2y agoReasoning seems to generalize, insofar as o1 and DeepSeek-R1 are better at answering questions than their base models.