4 ms·
> Im assuming the people who _train_ AI [ LLMs ] actually need to separate the two, and avoid the feedback loop of training the next LLM on the output of the pr
by esseph 9d ago
> Im assuming the people who _train_ AI [ LLMs ] actually need to separate the two, and avoid the feedback loop of training the next LLM on the output of the previous LLM.
Check out Reinforcement Learning from AI Feedback (RLAIF). Then skim some of this maybe: https://arxiv.org/abs/2309.00267 https://arxiv.org/abs/2309.00267