8 ms·
This is a very intriguing study as it would appear the fact that LLMs need public data to train on, its incentive on the marketplace of ideas is to reduce the v
by martingalex2 3y ago
This is a very intriguing study as it would appear the fact that LLMs need public data to train on, its incentive on the marketplace of ideas is to reduce the very thing that gives it its power. It contains the seeds of its own destruction, destroying the web and open data ethos, and yet another data point pointing at another AI winter.
- flangola7 3y agoAI winter? Labs haven't even began to scratch the surface of the data available to train models on. We may be running out of quality human-written text but we have yet to dive into: Video Audio Images Heat Motion/acceleration Lidar RF Sonar Radar Network traffic Atmospheric pressure Wind vectors Magnetic fields System/application logs Electrical current UV X-ray Microwave Ionizing particle emissions
- RandomLensman 3y agoInteresting, so we chuck in exabytes and more of data generated each day and then what?
- adventured 3y agoThe focus should eventually be on building AI such that they can be given access to data and then independently figure out new and creative ways to use it, rather than requiring humans to figure that out for them and then narrowly define: do this thing with the data I've provided. Given the scale of the data, they'll be better suited to that approach if it's breakthrough outcomes that we want.
- RandomLensman 3y agoYes, maybe good to put the ideas around "dataism" (don't know a better word, sorry) really to the test.
- marcosdumay 3y agoPeople have been testing them for ages. There's something there, but not nearly as much as the hype implies. At the same time, the hype seems to only increase. Anyway, I really like that name. Non-dataist AIs are clearly the best kind.
- __loam 3y agoI really think the deep learning maximalists are leading us down a rabbit hole. We're going to waste a lot of digital storage and computing resources on creating bigger and bigger models for more and more diminishing returns, barring some breakthrough. Without a way of encoding expertise, which is limited right now, I don't think we're going to get the kind of performance that the uninformed public expects out of these systems.
- __loam 3y agoNo AI we currently have can do this.
- detourdog 3y agoFor reasons I can't articulate I see LLMs as a vehicle for removing the creators from their ideas. This is very different than search engines. If a search engine generates traffic for documented ideas it creates a community. An LLM based internet seems to remove the creator and shim itself in between for the sake of business.
- RandomLensman 3y agoIf a handful of firms are indeed allowed to harvest he collective knowledge for rent seeking that would indeed be a shame.
- TeMPOraL 3y agoIt's tricky, because in many ways, this is achieving exactly what I, as user, want computers to do for me: give me information I requested, and only that. I very much do not care about who discovered/created/published it, except only if it helps me quickly ascertain the trustworthiness of said information. I do not want to be forced or prodded to establish relationships with creators or communities. I do not want their ads and upsells. It's the same issue as with search engines providing "information boxes": huge win for me, but a mortal enemy for those who want to monetize anything resembling intellectual property. > An LLM based internet seems to remove the creator and shim itself in between for the sake of business. This sounds bad, and in some cases it is, but in others it is not. Content farms and recipe sites have creators behind them too.
- RandomLensman 3y agoAnd the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?
- Swizec 3y ago> Even if it isn't monetizeable IP, how to share the costs? The internet started thanks to ample government funding for research. So have many other technologies, including AI. I wonder if there's a way we could all somehow pool our resources and use that to pay for common goods that we all use. What would we call such a scheme?
- __loam 3y agoLiteral tragedy of the commons. Overexploiting this resource will inevitably diminish it. Lack of human data to sample from will make models worse.
- musicale 3y ago> its incentive on the marketplace of ideas is to reduce the very thing that gives it its power It's not clear what the effect will be of LLMs consuming more and more of their own output and/or waste products. One possibility might be a feedback system that leads to superintelligence which eventually becomes incomprehensible to humans. Another might be a feedback system that leads to increased garbage output that eventually devolves into incomprehensible noise and nonsense. The two extremes might not be easily distinguishable.