10 ms·
Lots of RLHF - reinforcement learning from human feedback. Also: "distillation" - seeding or running training sessions on the output of frontier models.
by davedx 5d ago
Lots of RLHF - reinforcement learning from human feedback.
Also: "distillation" - seeding or running training sessions on the output of frontier models.