6 ms·
On problems this close to active research, seeing the model’s internal reasoning at the points of highest effort is more valuable than pass/fail outcomes alone,
by spacebacon 3mo ago
On problems this close to active research, seeing the model’s internal reasoning at the points of highest effort is more valuable than pass/fail outcomes alone, which is what SRT-Introspect makes possible on frozen models.
https://github.com/space-bacon/SRT https://github.com/space-bacon/SRT