6 ms·
The Private Capture of Public Genius
- marcus_holmes 2mo agoThis was interesting right up until "The fund pays every eligible American the same amount each year. " I'm in Australia. I've contributed my share of dirt to the delta. Why do I not get a share of this? I get that the frontier companies are (for the moment) US companies. But that's just corporate ownership, it's not what we're talking about. We're talking about compensating the people who wrote the training data for their contribution. That contribution came from all over the world, so the Corpus Fund needs to be paid all over the world. Set it up in the UN, get the UN to provide the training data sets as a common good, and have the UN collect the money from all AI companies using the training data sets. And the UN should distribute the money in the most equitable manner globally (so most of it going to alleviate poverty, probably). I'd happily trade my collected years of shitposts to help folks get out of poverty.
- martialg 2mo agoAuthor here. Thanks for reading! I have additional essays coming out that will address this exact issue and other issues I know that people will raise. I’m building the essays series around arguing for practical policy I believe can get implemented and am sequencing it as thoughtfully as I can. I just can’t fit every argument into every essay.
- marcus_holmes 2mo agoIt's a great idea, I hope you get traction with it :) I am personally coming to the conclusion that having these vast repositories of knowledge that can actually talk to us is actually great. We have some issues to solve, but the end-state of having a global repository of all knowledge that can talk to us and answer questions is actually an amazing outcome. We just need to solve those problems first; mostly getting past the AI bubble and the massive over-investment, and then solving the hallucination problems. I don't believe either of them are insoluble. I do worry about how future generations move on from this, though. In the same way that 90's music is still effectively the zeitgeist, and we will never move on from that, because of the way that streaming services work. It's a rare new band that can compete with (e.g.) Nirvana when appealing to that segment of audience, a competition that Nirvana themselves didn't have. So we are effectively locking in Nirvana as The Disaffected Youth Grunge Band for the rest of eternity. So similarly, we are in danger of locking in the current state of the world to the training data, and never being able to move on from that, because any new zeitgeist has to compete with this one on unequal footing.
- archagon 2mo agoImagine that AI lives up to its promise, gets captured by a handful of corporations, and ushers in a world of untouchable über-oligarchs that de facto rule the world. Would it have been worth it then? Can this problem even be solved?
- marcus_holmes 2mo agoI mean, yes, that's one possible outcome. I personally believe that it's more likely that the current batch of massively overleveraged AI companies go broke in the next year or so as the AI bubble bursts. Hopefully the USA manages to get its democratic shit together enough to not bail them out (as happened in 2008), and they just... stop. The OSS models are almost at the same level, and are progressing fine. The technology continues to be useful, even if all the oligarchs go away.
- dd8601fn 2mo agoMy spidey sense goes off when we refer to open weight models from Chinese companies trained against frontier models as OSS... as if it was some Torvalds types maintaining a public git repo. I think those open weight models exist because openai, anthropic, and google rule the roost. And I don't think things would continue as-is if those guys implode.
- marcus_holmes 2mo agoThe technology has certainly moved faster because of all the money thrown at it recently, but it existed before that, and was progressing before that. So no, I think OpenAI, Anthropic only exist because of the open-weight models (that came before them) rather than the other way around. Google is a special case; they invented the transformer and then published it instead of commercialising it, so they actually weren't evil (for once), or they didn't realise what they'd done until too late. I think things will continue perfectly fine if those guys implode. They didn't invent the technology, they're not the people who are advancing it most (because everything they do is proprietary and undisclosed), and their efforts at commercialising it are just getting in the way of the real research.
- martialg 2mo ago[flagged]
- Beijinger 2mo agoThis reminds me of: Ideas are like hemorrhoids, every asshole sooner or later gets some. ;-) This is not going to happen: https://www.theguardian.com/environment/earth-insight/2014/mar/14/nasa-civilisation-irreversible-collapse-study-scientists https://www.theguardian.com/environment/earth-insight/2014/m...
- wwind123 2mo agoI agree UN sounds like a good organization to help distribute the wealth created by AI to the world. But this idea won't be considered by the current U.S. administration. In that case what other countries can do is probably to tax the AI companies at rates higher than regular companies.
- marcus_holmes 2mo agoThe USA helped form the UN as specifically the organisation to do exactly this kind of thing. It's a shame the current administration can't play nicely with others. The current administration is also playing strange games about export controls (can we run Fable yet? Kinda. Maybe). I think if they keep this up they'll just be shooting the US AI industry in the foot and the Chinese models will take over as the frontier models. Maybe the UN can levy the USA for this, and leave the USA to collect that levy from its AI companies.
- strictnein 2mo agoThe US formed the UN to distribute funds? > Maybe the UN can levy the USA for this The UN has no "levy" powers.
- marcus_holmes 2mo agoTrue, good point.
- nradov 2mo agoDo you really want countries like Saudi Arabia and North Korea to have a vote in wealth redistribution?
- abalashov 2mo agoI suppose countries like KSA and DPRK would ask the same question about us.
- 2mo ago
- tptacek 2mo agoWhy do you think a UN-hosted process would work? Has the UN done similarly intrusive things before and been effective at it? How do you account for Chinese frontier and open weights models, which are just months behind American models and subsidized by a sovereign state that does not share any of these premises about intellectual property? A lot of these posts seem subtextually premised on the idea that it's possible to put the genie back into the bottle; that if frontier labs in America didn't sign on to this tolling scheme, our recourse would be to halt the progression of AI completely. But that option does not exist, unless we're going to fight a world war to create it.
- MichaelZuo 2mo agoThis comment seems incoherent? Even if the UN only has a very small amount of legitimacy and credibility… that’s still far far better than literally zero.
- tptacek 2mo agoWell, I mean, the alternative is for a US-based process to be put into place, followed (or led) by an EU process. The premise of my comment is that no matter who hosts the process, China isn't going to comply, so it really just comes down to the US and Europe.
- MichaelZuo 2mo agoWhat? There are over 190 nations in the UN. Even if we assume a chunk have no credible claim whatsoever, that’s still well over a hundred.
- tptacek 2mo agoThere are not in fact 100 nations that will seriously contend for frontier models.
- MichaelZuo 2mo ago
- nradov 2mo ago[flagged]
- verisimi 2mo agoNo one asked you about this. Your government is free to spend its money as it likes, it can send it to the UN.
- antonvs 2mo ago> I'd happily trade my collected years of shitposts to help folks get out of poverty. In other words, you'd happily do nothing to help folks get out of poverty?
- marcus_holmes 2mo agoI was referring to the potential of actually collecting some of this fund. But yes, essentially in this situation that is correct - I would be doing nothing and also helping people get out of poverty.
- imrehg 2mo agoAs long as places like Taiwan are effectively blocked from being in the UN[0], I don't believe we should be adding more to their power and responsibilities than what they have now, au contraire! If it was a truly world representation, this might be different. But if things like health are sacrificed (ie. no WHO access either), I don't think they really deserve the benefit of the doubt. [0]: https://en.wikipedia.org/wiki/Taiwan_and_the_United_Nations https://en.wikipedia.org/wiki/Taiwan_and_the_United_Nations
- supriyo-biswas 2mo ago> This was interesting right up until "The fund pays every eligible American the same amount each year. " There could be so many other ways to set this up. Enforce a higher tax on any business selling AI models that have capabilities greater than some threshold, and use it to fund development and infrastructure project like roads, hospitals, schools, etc. Or you could even do a negative income tax[1]. [1] https://en.wikipedia.org/wiki/Negative_income_tax https://en.wikipedia.org/wiki/Negative_income_tax
- marcus_holmes 2mo agoAgreed. And someone has to manage and enforce that on a global level. Which is what we built the UN for.
- inigyou 2mo agoThe UN is a voluntary association with no enforcement power, like the EU, the WIPO, and the IETF
- martialg 2mo agoAuthor here. Thanks for reading. I agree there are a lot of levers we can pull to work towards something better. I've structured this essay series as a sequence of nested regulatory solutions so in the next few essays I propose additional structures with instruments like this. They're sequenced in a way I believe that can be pragmatically implemented and start showing progress in the next decade or so by the US. So stay tuned!
- martialg 2mo ago[flagged]
- aethelyon 2mo ago100% — publish the hidden research, the value is in the discoveries, not in the dividend. With all due respect to the author, it feels like he missed the entire lesson of history.
- aj_hackman 2mo agoWho would pay you? OpenAI has between $600 billion and $1.4 trillion in debt, and no hope for profitability. Anthropic, who have a much more modest $35 billion in debt (by using Sam Altman as a human shield), also has no chance of ever being in the black, pending some freak event.
- hyperhello 2mo agoThere’s a quote about how in some articles a switch is quietly flipped in the middle where the article was talking about what is and suddenly the author has everything to say about what should be. I googled for the quote but all I got is useless web spam and meme style graphics about quotes from writers. But AI told me it was David Hume and provided the full quote. The real question is when the day will come that AI become the fertile muck that a new thing grows from and clings to and the legal system needs to adjust to. I hope it’s a good thing.
- quuxplusone 2mo agoSounds like you're thinking of the https://en.wikipedia.org/wiki/Is%E2%80%93ought_problem https://en.wikipedia.org/wiki/Is%E2%80%93ought_problem . Wikipedia quotes David Hume's "A Treatise of Human Nature" 3.1.1 as follows: > In every system of morality which I have hitherto met with [...] the author proceeds for some time in the ordinary way of reasoning [...] when of a sudden I am surprised to find that instead of the usual copulations of propositions is and is not I meet with no proposition that is not connected with an ought or an ought not.
- deleted 2mo ago[deleted]
- typ 2mo agoWe cannot always want to capture only the (temporary) winners whenever we see a lucrative business and expect to share a free ride. I'd also assume that most of the revenue these AI labs are making is turned into depreciating fixed capital (hardware) and OPEX at this point. Why don't we capture Meta and Google as they allegedly take advantage of more publicly available information for profit? Let alone the truly valuable knowledge, like mathematics, has nothing to do with the majority of garbage posts that an average person would "contribute" on social media. If we really want to tax or nationalize some economic activity, then, in my opinion, the target should be what it takes from society, not what it produces for society. By this logic, we should tax all labs, including those lagging ones, that utilize the public knowledge. However, if everyone can access the public knowledge without rendering it less useful or reducing its available quantity, there should be no reason to tax it.
- martialg 2mo agoAuthor here. Thanks for taking the time to read. I agree we’re in an interesting era where frontier research has shifted from mostly publicly funded to mostly private and it creates challenging incentive structures especially regarding externalized costs of research. Did you have any thoughts on my argument of how public knowledge does get damaged by the proliferation of AI over time?
- typ 2mo agoI don't understand how knowledge, either public or private, could get damaged. Though the income of the individuals and businesses that rely on the expertise of the knowledge would be damaged. Is that what you meant? Edit: At this stage, the revenue made by the AI labs is almost entirely spent on the formation of fixed capital and opex. The demand is mobilizing physical resources with money. Atoms are relocated and reconfigured into compute racks. But eventually, the created productivity will perhaps make supply-elastic goods extremely cheap and abundant, while the supply-inelastic goods will be worth even more relative to the elastic ones. After all, money is simply a token for the transmission of physical resources. It doesn't create stuff out of thin air. When new stuff is created, it just makes money cheaper, so that the banks can respond with more money to counteract it. More stuff -> deflation -> more money creation allowed to undo the deflation. But the "exchange rates" between different goods and services will diverge. That's also why I don't think a direct money transfer like UBI would fix the problem, when it doesn't change the divergence of relative economic values of different goods. Let's say, extremely cheap software and entertainment, but unaffordable healthcare and housing. More money for everyone doesn't make limited resources available. So, what I am leaning into is some sort of Georgist policy. That could hopefully mitigate the price divergence, assuming that AI cannot make every commodity equally abundant.
- abalashov 2mo agoIt was a brilliant article, and it succinctly captured the offenses to ethics and humanism posed by LLMs. I'm not sure it'll get a lot of reception in the technocracy here on HN, whether of the AI booster or AI nihilist sort. However, I think it's a very comprehensive digestion of the questions that will swirl around the idea of LLMs as a public good in the near to medium future.
- martialg 2mo agoAuthor here. Really appreciate you taking the time to read and for the kind comment. I think the tension between these ethical questions and the practical realities (both the good and the bad) of AI is likely the defining issues for technology and perhaps society in this decade. It’s important we’re thorough and rigorous with how we think and act here so I really appreciate you engaging with the topic.
- abalashov 2mo agoThank you in turn, I have circulated your piece to thoughtful friends. My immediate, from-the-hip thought is that we are slowly lumbering toward the idea that LLMs ("AI") should be a public utility. It may take us quite a while to get there yet, as an unprecedented concentration of wealth and power is arrayed precisely against this outcome, but I think that will be the eventual effect, in that, "in the long run, we're all dead" kind of way.
- martialg 2mo agoThank you kindly I’m working through thoughts on this as well and agree with your read on the incentives. There is an interesting set of conditions that happens if/when models get so competent that they’re effectively indistinguishable from each other and inference becomes a true commodity. IRL impact will lag this ofc but it’s such a wild time to be alive.
- abalashov 2mo agoI do think it's easy, in this technology discussion bubble in which we dwell, to overestimate the centrality of LLMs to the arc of developments in our time. They'll be important, but I don't think they'll be _that_ important, because the rest of society and the economy don't move at the speed of SV. Instead, they'll be overtaken by other, more traditional categories of events, ruptures and dislocations. Moreover, folks will eventually realise that while they are very impressive derivative databases of knowledge, they're not at all "AI" -- well, not the "I" part, anyway -- as the concept is traditionally understood. There's not any "I". It emits convolutions of its training, and it does so very impressively, and that can even be harnessed by agents to connect them to levers, servos and richer information sources. It's nifty. But it's just not intelligence. It's more of a kind of queryable database than a robot. Once that realisation diffuses more widely, I think it'll turn out to be a more prosaic and underwhelming development than is presently hypothesised, either here or by the press. It doesn't mean many business and managerial class folks won't try to squeeze everything they can out of so-called AI, but the idea that this can effectuate truly widespread labour displacement will probably quiet down considerably. (The valuations that depend on this assumption may collapse more abruptly and less gracefully.) The challenge is staying solvent until then. :-)
- willturman 2mo agoA similar appropriative-use vs. public-trust evaluation is playing out this year as the California State Water Resources Control Board reevaluates Los Angeles’ right to divert water from the Mono Basin in the Eastern Sierra Nevada. The foundational case for Mono Lake as a public trust resource is National Audubon Society v. Superior Court (1983) [1]. The California Supreme Court evaluated appropriative water rights against the public trust doctrine, took both arguments to their logical extremes, and decided that neither was acceptable in itself. In a pretty jaw-dropping passage, the Court summarized the Los Angeles Department of Water and Power’s position in relation to appropriative use of water diverted from a unique ecosystem hundreds of thousands of years old: > Defendant DWP, on the other hand, argues that the public trust doctrine as to stream waters has been "subsumed" into the appropriative water rights system and, absorbed by that body of law, quietly disappeared; according to DWP, the recipient of a board license enjoys a vested right in perpetuity to take water without concern for the consequences to the trust. The decision in Audubon rejected LADWP’s argument, but it remains a stark example of the beneficiary of a public resource recasting a conditional public license as a permanent private entitlement, apparently free from consequence of accountability for harm inflicted on the public trust. I think this appropriative-use vs. public-trust/public-benefit discussion is going to define the coming decades. The landscape remains unsettled as it applies to water (especially in a changing climate), much less to data in a period of rapidly evolving technology. With respect to data, progress could be made by formally establishing a public corpus as an accessible commons, with clear expectations and rights around individual contributions made to third-party platforms. Publicly funded research is still often locked behind paywalls. The contents of the Library of Congress, special collections, municipal libraries, university archives, and museums are publicly owned or publicly supported, yet remain largely inaccessible to the general public. I expect the “leader” in LLM performance to keep changing, but the accumulated genius of public knowledge to remain far more durable, with periodic and incremental additions. Fighting over small reparations for every scraped post seems less transformative than building a public knowledge commons that anyone can use, converse with, search, train on, and learn from. reCAPTCHA began as a tool that simultaneously authenticated users while helping verify OCR for the backlog of The New York Times and Project Gutenberg. Maybe it is time for a similar public project to digitize and make accessible the body of public knowledge without surreptitious and ethically dubious appropriation of copyrighted works. Authors, writers, and shitposters could opt in as desired. I would take a public resource like that well ahead of a few bucks of compensation for my decades of shitposting, just as I'd take a thriving Mono Lake well ahead of compensation for it being relegated into lifeless alkali flat via appropriative water rights.
- w10-1 2mo ago"Fair use" was always fuzzy. To be honest, I care a lot less about slurping up the public internet and private books to make models than about every profession on the planet being forced by their employers to create skills that automate their knowledge work. The latter is much more directly an expropriation, legitimized only by the shortage of work, i.e., market power.
- stale2002 2mo agoI'm surprised that the Author has missed a very important corollary to the diffusion of "genius theft" that they are bringing up. And that is the diffusion of the beneficiaries. Maybe they think that OpenAI and Anthropic are actually worth, like a trillion dollars each and can therefore have value extracted from them. I'm not so convinced. What if there aren't frontier labs spending billions on training a model. What if, instead, open source is at least mostly competitive with the top models. And if the models are open source (or weights, whatever you want to call it) the people benefiting are actually just rando people or startup founders. What are you going to do if you want to extract this value from this diffuse set of beneficiaries? Put an arbitrary tax on anyone living on San Francisco or something?? The reality is that the author is trying to put the genie back into the bottle. All technological progress has winners and losers. It has people who are even benefiting from the rest of society and making personal gain based on that. But, at the end of the day, doing accounting math on how much an individual benefited from a specific common good as vague as societal knowledge is impractical. And yet technological progress benefits all of society in a wholistic sense. Additionally, the author focuses so much on the extraction of a public good, I am surprised that they failed to address that these labs are creating a public good as well. Who's to say that this "theft" is larger than the production of public goods that these labs give to the public in the first place. I mean, my life has been massively improved by the fact that I have access to these models. And I'm not convinced that I have produce enough myself to outweigh this benefit that they are giving to me, so I consider it to be a fair trade.
- inigyou 2mo agoThere aren't any open source LLMs by the way
- karahime 2mo agoSure there are. Ai2's OLMo, EleutherAI's Pythia, and LLM360's Amber are actually open source, top to bottom, training data to checkpoints to code, to name three of them. You also don't have to look far on Huggingface to find smaller models developed with open source processes all the time.
- 2mo ago
- schnitzelstoat 2mo agoThe troubles over copyright infringement in AI training data remind me a bit of Eli Whitney and the cotton gin. There he suffered massive patent infringement, that basically stopped being enforced due to the sheer economic importance of the cotton gin. In a similar manner, I think there is a reasonably strong argument that it was wrong to use copyrighted material for AI training without paying royalties nor even asking for permission. But equally, every country wants to have the most powerful models and enforcing such royalties would make it effectively impossible to train them as the amount of material required would cost an insane amount in royalty fees. So I expect the law will continue to turn a blind eye (perhaps enforcing some token payments like that $1.5B mentioned in the article) because "if we don't make these models, the Chinese will" etc.
- toofy 2mo ago> ... there is a reasonably strong argument that it was wrong to use copyrighted material for AI training without paying royalties nor even asking for permission. But equally, every country wants to have the most powerful models and enforcing such royalties would make it effectively impossible to train them as the amount of material required would cost an insane amount in royalty fees. i think you're spot on this is one of the key arguments made beneath the surface. what i find so strikingly frustrating about it is, so many of the ai cultists [0] will imply and sometimes even outright say that writers, artists, musicians are silly useless and overvalued and the work artists do is entirely frivolous. then next breath explain why those artist's work is one of the most important things for a model to be trained on. suddenly art is very important. we absolutely must have access to their work. but also we shouldnt pay them because their work is silly and unimportant. if an artists (musician, writer, journalist, painter, etc...) work is useless, then obviously you dont need it for training. if their work is imperative and you absolutely must use it, then pay for it. ive noticed this with ai companies a lot. over and over again they contradict themselves to the core. 1) art is silly and not important enough to pay for but its absolutely foundational and we must be given unfettered access or our models will suck. 2) "our models are the smartest thing in the entire world. also, you're a dipshit if you trust them at all." ill say it again, if removing art and culture from the training sets would render your model useless, then obviously pay for it. [0] when i say cultists, im not talking about normal people who use ai. im talking about an entirely different group, we all know the types im talking about.
- anon373839 2mo agoThis is a well written essay. I had hoped it might address the role of distillation and open source in diffusing ownership of this technology back to the public that made it possible. And the AI labs’ rank hypocrisy in this area.
- martialg 2mo agoAuthor here. Thanks for reading and the kind words. I will talk about distillation and OS in coming essays (the is a multi-part series).
- phrotoma 2mo agoAnybody know where that gordon moore quote comes from? A little searching didn't produce sources for me.
- martialg 2mo agoAuthor here! It's from a workshop in 2001 for the National Research Council's Board on Science, Technology, and Economic Policy. He gave a talk. You can ctrl+f for it at this link https://www.ncbi.nlm.nih.gov/books/NBK208682/ https://www.ncbi.nlm.nih.gov/books/NBK208682/
- phrotoma 2mo agoThank you!
- pgisapedo 2mo agoIt's been said a thousand times but the ability to produce copyrighted material is not copyright infringement. I'll say it again: You can produce copyrighted material all day long and it's not copyright infringement. If I draw The Simpsons on a piece of paper - whether or not I used AI to create it - it's not copyright infringement. Copyright infringement is if I tried to sell that Simpsons work as my own, putting up for consumption and reaping the monetary benefits. It has never been illegal to produce copyrighted work. You people are in for a surprise dystopia if you keep pushing to make that a reality.
- sfink 2mo ago> It's been said a thousand times but the ability to produce copyrighted material is not copyright infringement. Aside from fair use, yes it is. > If I draw The Simpsons on a piece of paper - whether or not I used AI to create it - it's not copyright infringement. Only because of the fair use doctrine, which is limited. If it damages the market for the original, then a court could definitely declare it to be infringement. > Copyright infringement is if I tried to sell that Simpsons work as my own, putting up for consumption and reaping the monetary benefits. No. Copyright infringement requires neither sale nor misrepresenting a work as one's one. It also covers derivative works, not just perfect duplicates. Copyright covers reproduction and derivate works and distribution/performance -- you can get in trouble for just one of those. Taken strictly, that would be a horrible world, but fortunately the fair use doctrine weakens those quite a bit. On the flip side, if AI reproduction destroys the original markets -- as it is quite obviously doing right now -- then it's going to have a reckoning at some point, given how its legal status is completely based upon getting a free ride by claiming fair use.
- pgisapedo 2mo ago[dead]
- PaulDavisThe1st 2mo ago> Frontier science looks different today. It's rooted in model weights and GPUs. It is flooded with token spend and agentic loops. It blooms in data centers. This seems like handwaving to me. Even the actual frontier science that is using ML (e.g. AlphaFold) isn't based on "token spend" or "agentic loops". I personally cannot think of a single example of frontier science that is rooted in LLMs. I am sure there are a few examples, but the idea that frontier science has somehow completely shifted its trajectory and methods based on LLMs feels like quite a stretch to me. Got any counterexamples to show me I'm wrong?
- martialg 2mo ago[flagged]
- 1vuio0pswjnm7 2mo agotl;dr Author fears pending/future copyright litigation against AI labs, proposes "Corpus Royalty"
- zombot 2mo ago> This is the private capture of public genius. In other words, the biggest heist in human history. They took all of our collective knowledge and print money by selling it back to us.
- scotty79 2mo ago> returns constrained to a relatively conservative (by today’s standards) ~7% per annum Damn, we should have something like that market wide. Progressive, with the revenue. Forbidding vertical integration would be a tremendous blessing too.