6 ms·
All the methods and data are not public. We don't know what unpublished methods they're using. You can get most of the pre-training data publicly but they've pr
by didroe 15d ago
All the methods and data are not public. We don't know what unpublished methods they're using. You can get most of the pre-training data publicly but they've probably spent a ton of money curating it and are now doing things like buying rare books. The RL training data is all (/mostly) proprietary though, and that's the real secret sauce part.
- woctordho 15d agoAll the RL data are exactly public. There are huge amount of distilled data freely available, and that amount is more than enough to train a ~10T model.
- cubefox 15d ago> All the RL data are exactly public. Nope, because the big AI companies are paying billions for it. They wouldn't pay anything for public data.
- woctordho 14d agoThere are 'transfer stations' and that's how exactly I use GPT and Claude in China. OpenAI and Anthropic do not sell in China, so we use their AI with a much lower price like 1% of the official API price. The largest transfer stations have TBs of traffic every day, and the traffic is eventually possessed by the open source community. Subscription engineering is a deep field. Neither OpenAI nor Anthropic have any technical advantage in this field.