6 ms·
Isn't the training data a moat by itself? Due to sneaky website scraping by OpenAI many sites now have their content locked behind login, have their APIs seve
by devsda 2y ago
Isn't the training data a moat by itself?
Due to sneaky website scraping by OpenAI many sites now have their content locked behind login, have their APIs severely restricted and watchout for scraping.
Unless you want to build a specialized AI built on specific open or closed training corpus, huge training data is required to build something similar to OpenAI's models.
Google with it's crawlers and associated relaxations should not be affected much by those restrictions.
Miscrosoft can easily update its browser/OS ToS and say "we'll use all your browsing data for training, anonymously ofcourse" and there's not much users will do except shrug.
Is there anyone else who is better positioned to collect the content for training or am I placing too much importance on training data?
- a_bonobo 2y agoI agree, it's far more the 'training data' than the model architecture. Given that Windows 11 is such a surveillance/advertisement machine I'm sure they sit on years of user movements MS can use to build a OS-running AI, similar to what Anthropic has been trying with 'computer use'. That's where I see Microsoft's niche in the future, and it's an enormous niche: how many people do 'data entry' as a full-time job who'd be unemployed once an LLM can move data from random sources into an Excel sheet by operating a simulated computer?
- Davidzheng 2y agoDisagree, now that there are great open models are available I think there's less need of huge training data--can just post train