7 ms·
Likely won't improve much. They trained on every text already.
by pacman1337 2mo ago
Likely won't improve much. They trained on every text already.
- m_ke 2mo agomost of the gains from the past year and a half have not been from web data, but from synthetic data and agent rollouts with RL.
- Schlagbohrer 2mo agoThere is tremendous investment and work being done in creating new high quality data sets. Some of this is happening in-house at various firms, such as Meta making their employees use AI for their workflow to generate high quality training data. Also, AI companies get huge amounts of human input when people use their cloud models, including thumbs-up or thumbs-down on millions of outputs. So the usage of these cloud models is itself producing new, high quality datasets.