

Discussion summary
The discussion centers on concerns about copyright issues related to AI and collective knowledge, with some proposing a 'Corpus Royalty' to address it. Participants debate the nature of copyright infringement and the impact of AI on public trust and knowledge sharing.
What the discussion says
- The author fears future copyright litigation against AI labs.
- Some argue that producing copyrighted material isn't infringement unless it violates fair use.
- There is concern about the privatization of collective knowledge and its implications.
- Debate on whether open source LLMs exist and the impact of AI on public commons.
“This is the private capture of public genius. The biggest heist in human history.”
“The ability to produce copyrighted material is not copyright infringement.”
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I'm surprised that the Author has missed a very important corollary to the diffusion of "genius theft" that they are bringing up.
And that is the diffusion of the beneficiaries. Maybe they think that OpenAI and Anthropic are actually worth, like a trillion dollars each and can therefore have value extracted from them. I'm not so convinced.
What if there aren't frontier labs spending billions on training a model. What if, instead, open source is at least mostly competitive with the top models. And if the models are open source (or weights, whatever you want to call it) the people benefiting are actually just rando people or startup founders.
What are you going to do if you want to extract this value from this diffuse set of beneficiaries? Put an arbitrary tax on anyone living on San Francisco or something??
The reality is that the author is trying to put the genie back into the bottle. All technological progress has winners and losers. It has people who are even benefiting from the rest of society and making personal gain based on that.
But, at the end of the day, doing accounting math on how much an individual benefited from a specific common good as vague as societal knowledge is impractical. And yet technological progress benefits all of society in a wholistic sense.
Additionally, the author focuses so much on the extraction of a public good, I am surprised that they failed to address that these labs are creating a public good as well. Who's to say that this "theft" is larger than the production of public goods that these labs give to the public in the first place.
I mean, my life has been massively improved by the fact that I have access to these models. And I'm not convinced that I have produce enough myself to outweigh this benefit that they are giving to me, so I consider it to be a fair trade.
by stale2002 - There aren't any open source LLMs by the wayby inigyou
- > they failed to address that these labs are creating a public good as well. Who's to say that this "theft" is larger than the production of public goods that these labs give to the public in the first place.
Those aren't public goods, they're private goods. The difference is already apparent, and will become more so over time. But yes, they are "goods"; they have a lot of value.
I'm less interested in punishing the past theft that enabled these goods, and more worried about the ongoing damage they are doing to the ecosystem that birthed them. Good artists copy, great artists steal, but if a community is wholly dominated by these "great artists" then it will not survive as a community and will cease to produce anything worth copying let alone stealing.
by sfink - > these labs are creating a public good as well.
https://en.wikipedia.org/wiki/Public_good
> In economics, a public good (also referred to as a social good or collective good)[1] is a commodity, product or service that is both non-excludable and non-rivalrous and which is typically provided by a government and paid for through taxation.
I can see that we are already excluding people from using models, whether it's China[1] or use in other harnesses[2] so its use can't be considered "non-excludable" or "non-rivalrous". It is also neither provided by the government nor funded through taxation.
I guess one change could be made to force LLM companies to release their (N-1)th model publicly and document architecture and system requirements, which would shift this to make these products actually public goods, however.
[1] https://x.com/AnthropicAI/status/2025997928242811253
[2] https://www.mindstudio.ai/blog/anthropic-openclaw-ban-oauth-...
- Author here - thanks for reading and thoughtfully replying
I’d personally love a world where open weights compete with proprietary ones, but I don’t believe it solves the core concentration issue. In that scenario most value still flows to capital holders, it’s just hardware holders not model weight holders.
I emphatically do not want to put the genie back in the bottle nor do I believe it’s possible. Technology has never been restrained for long (export controls on cryptography textbooks in the 90’s comes to mind here)
I also have already personally benefited a great deal from LLMs. I actually frame the entire essay series from this perspective in my prelude essay here: https://www.wysr.xyz/p/a-consigliere-on-every-desk-and-in
However, I believe we may disagree on the definition of a public good. If you’re referring to the free tiers of private models, then I’d argue that unless there is some legal framework passed that forces the frontier labs to offer that to everybody, it’s a customer acquisition cost laundered as a public good. It could disappear at any time and probably will as cutting edge model margins are reduced via competition.
In general, I believe the best AI policy balances allowing for maximum competitive market dynamics while hedging existential economic disruption risk for the general population.
I’ll go into this deeper over the next few essays. Appreciate the feedback
by martialg - Anybody know where that gordon moore quote comes from? A little searching didn't produce sources for me.by phrotoma
- Author here! It's from a workshop in 2001 for the National Research Council's Board on Science, Technology, and Economic Policy. He gave a talk.
You can ctrl+f for it at this link
by martialg - There’s a quote about how in some articles a switch is quietly flipped in the middle where the article was talking about what is and suddenly the author has everything to say about what should be.
I googled for the quote but all I got is useless web spam and meme style graphics about quotes from writers. But AI told me it was David Hume and provided the full quote.
The real question is when the day will come that AI become the fertile muck that a new thing grows from and clings to and the legal system needs to adjust to. I hope it’s a good thing.
by hyperhello - Sounds like you're thinking of the https://en.wikipedia.org/wiki/Is%E2%80%93ought_problem . Wikipedia quotes David Hume's "A Treatise of Human Nature" 3.1.1 as follows:
> In every system of morality which I have hitherto met with [...] the author proceeds for some time in the ordinary way of reasoning [...] when of a sudden I am surprised to find that instead of the usual copulations of propositions is and is not I meet with no proposition that is not connected with an ought or an ought not.
by quuxplusone - This is a well written essay. I had hoped it might address the role of distillation and open source in diffusing ownership of this technology back to the public that made it possible. And the AI labs’ rank hypocrisy in this area.by anon373839
- Author here. Thanks for reading and the kind words. I will talk about distillation and OS in coming essays (the is a multi-part series).by martialg
- > returns constrained to a relatively conservative (by today’s standards) ~7% per annum
Damn, we should have something like that market wide. Progressive, with the revenue.
Forbidding vertical integration would be a tremendous blessing too.
by scotty79 - It was a brilliant article, and it succinctly captured the offenses to ethics and humanism posed by LLMs.
I'm not sure it'll get a lot of reception in the technocracy here on HN, whether of the AI booster or AI nihilist sort. However, I think it's a very comprehensive digestion of the questions that will swirl around the idea of LLMs as a public good in the near to medium future.
by abalashov - What technocracy lol? people here turn into luddites if there's a positive reception of AI.. I can guarantee that there are at least 40 percent of the top 5 comments of any positive/negative post about AI gonna turn "hackers" to luddites. It was amazing before Covid times and now it's indistinguishable from reddit.by CaptWorld
- Author here. Really appreciate you taking the time to read and for the kind comment.
I think the tension between these ethical questions and the practical realities (both the good and the bad) of AI is likely the defining issues for technology and perhaps society in this decade.
It’s important we’re thorough and rigorous with how we think and act here so I really appreciate you engaging with the topic.
by martialg - The troubles over copyright infringement in AI training data remind me a bit of Eli Whitney and the cotton gin.
There he suffered massive patent infringement, that basically stopped being enforced due to the sheer economic importance of the cotton gin.
In a similar manner, I think there is a reasonably strong argument that it was wrong to use copyrighted material for AI training without paying royalties nor even asking for permission. But equally, every country wants to have the most powerful models and enforcing such royalties would make it effectively impossible to train them as the amount of material required would cost an insane amount in royalty fees.
So I expect the law will continue to turn a blind eye (perhaps enforcing some token payments like that $1.5B mentioned in the article) because "if we don't make these models, the Chinese will" etc.
- > ... there is a reasonably strong argument that it was wrong to use copyrighted material for AI training without paying royalties nor even asking for permission. But equally, every country wants to have the most powerful models and enforcing such royalties would make it effectively impossible to train them as the amount of material required would cost an insane amount in royalty fees.
i think you're spot on this is one of the key arguments made beneath the surface. what i find so strikingly frustrating about it is, so many of the ai cultists [0] will imply and sometimes even outright say that writers, artists, musicians are silly useless and overvalued and the work artists do is entirely frivolous. then next breath explain why those artist's work is one of the most important things for a model to be trained on. suddenly art is very important. we absolutely must have access to their work. but also we shouldnt pay them because their work is silly and unimportant.
if an artists (musician, writer, journalist, painter, etc...) work is useless, then obviously you dont need it for training. if their work is imperative and you absolutely must use it, then pay for it.
ive noticed this with ai companies a lot. over and over again they contradict themselves to the core.
1) art is silly and not important enough to pay for but its absolutely foundational and we must be given unfettered access or our models will suck.
2) "our models are the smartest thing in the entire world. also, you're a dipshit if you trust them at all."
ill say it again, if removing art and culture from the training sets would render your model useless, then obviously pay for it.
[0] when i say cultists, im not talking about normal people who use ai. im talking about an entirely different group, we all know the types im talking about.
by toofy - Author here. Appreciate your thoughts and I mostly agree actually.
I'll explain more over the next few essays, but I am designing my proposed regulatory structures to try to accomplish 2 purposes in tension simultaneously like the Fed: 1. Maintain global competitiveness for frontier labs 2. Create a societal hedge against the AI bull case (AKA the economic black hole case)
A % of revenue scales in a way that I think balances the two well while avoiding all the other problems I mentioned in the essay. I’ll get into ratchets, timing, and thresholds in later essays, but I agree the China/competitiveness problem is central.
by martialg - I'd be fine with the nuclear compromise: if AI training is allowed to infringe copyright, then there is no legal protection for the models themselves and their weights. Distillation should be explicitly legal. There will of course be a huge cat and mouse game about it, but let's have competition drive prices down on the stolen IP.by pjc50
- "Fair use" was always fuzzy. To be honest, I care a lot less about slurping up the public internet and private books to make models than about every profession on the planet being forced by their employers to create skills that automate their knowledge work. The latter is much more directly an expropriation, legitimized only by the shortage of work, i.e., market power.by w10-1
- We cannot always want to capture only the (temporary) winners whenever we see a lucrative business and expect to share a free ride. I'd also assume that most of the revenue these AI labs are making is turned into depreciating fixed capital (hardware) and OPEX at this point.
Why don't we capture Meta and Google as they allegedly take advantage of more publicly available information for profit? Let alone the truly valuable knowledge, like mathematics, has nothing to do with the majority of garbage posts that an average person would "contribute" on social media.
If we really want to tax or nationalize some economic activity, then, in my opinion, the target should be what it takes from society, not what it produces for society. By this logic, we should tax all labs, including those lagging ones, that utilize the public knowledge.
However, if everyone can access the public knowledge without rendering it less useful or reducing its available quantity, there should be no reason to tax it.
by typ - That's pretty much the reaction I had, except that I do find the argument compelling that AI is damaging the ecosystem it feeds upon.
I think Google's AI results are probably the prime example here. Those results are often quite good, and starve off the visits and hence revenue stream of the sites where the results are sourced from. Additionally, there's no way to robustly attribute that information anymore either; those links were already broken pre-AI by freeloading aggregators. So potential information producers can't afford to host their valuable information since they will pay proportionally to their information's value (as translated into bandwidth).
But it's tricky. If you charge AI companies proportional to the damage they do, then you need to assess that damage and you'll be caught in a no-win cat & mouse game where the AI companies outsource the damage and you try to track it back to them. They win if they bump their revenue, but they also win if they conceal the damage. (Just like companies can use shell companies, bankruptcy, and asset-only purchases to avoid Superfund responsibility.) If you charge proportional to revenue, then there aren't conflicting incentives; companies win by increasing revenue full stop. But I do agree that this shouldn't just apply to the big companies; distillation / model extraction (adversarial or not) should not be a way to avoid fees.
Though what really seemed off to me about the proposed solution was the use of the fees (to pay Americans, no less!) Perhaps upcoming essays will justify this more, but to me it seems like those fees should be applied directly to the health of the ecosystem. It should be used to fight pollution, in this case AI slop overtaking the Web. It should support the value creators, eg open source developers being overwhelmed by the AI tsunami. It should go towards serving and moderating online communities that create the very value that the LLMs are trained from. In theory, paying a whole bunch of Americans a trickle of blood money will end up going towards these purposes, but I'm very skeptical: first, the benefit will be very diluted. Second, it's more likely that the excess money will end up in Google or Amazon's pockets. The whole system is already set up to route any advantage to the big players. The payments are at least as likely to damage the ecosystem as they are to support it. They would just feed the capitalistic wolf and further reinforce the setup where monetizers win and value creators lose. If you could magically route the money to value creators, that'd be great. But you can't.
It'd be ok to pour money into a bucket if it had a small hole or two, but pouring it faster into the sieve we have now is not going to help anything.
by sfink - > we see a lucrative business and expect to share a free ride
the lucrative business are the freeloaders
by goodpoint - Author here. Thanks for taking the time to read.
I agree we’re in an interesting era where frontier research has shifted from mostly publicly funded to mostly private and it creates challenging incentive structures especially regarding externalized costs of research.
Did you have any thoughts on my argument of how public knowledge does get damaged by the proliferation of AI over time?
by martialg - > Frontier science looks different today. It's rooted in model weights and GPUs. It is flooded with token spend and agentic loops. It blooms in data centers.
This seems like handwaving to me. Even the actual frontier science that is using ML (e.g. AlphaFold) isn't based on "token spend" or "agentic loops". I personally cannot think of a single example of frontier science that is rooted in LLMs. I am sure there are a few examples, but the idea that frontier science has somehow completely shifted its trajectory and methods based on LLMs feels like quite a stretch to me.
Got any counterexamples to show me I'm wrong?
- This was interesting right up until "The fund pays every eligible American the same amount each year. "
I'm in Australia. I've contributed my share of dirt to the delta. Why do I not get a share of this?
I get that the frontier companies are (for the moment) US companies. But that's just corporate ownership, it's not what we're talking about. We're talking about compensating the people who wrote the training data for their contribution. That contribution came from all over the world, so the Corpus Fund needs to be paid all over the world.
Set it up in the UN, get the UN to provide the training data sets as a common good, and have the UN collect the money from all AI companies using the training data sets. And the UN should distribute the money in the most equitable manner globally (so most of it going to alleviate poverty, probably).
I'd happily trade my collected years of shitposts to help folks get out of poverty.
- > This was interesting right up until "The fund pays every eligible American the same amount each year. "
There could be so many other ways to set this up. Enforce a higher tax on any business selling AI models that have capabilities greater than some threshold, and use it to fund development and infrastructure project like roads, hospitals, schools, etc. Or you could even do a negative income tax[1].
- > I'd happily trade my collected years of shitposts to help folks get out of poverty.
In other words, you'd happily do nothing to help folks get out of poverty?
by antonvs - Who would pay you? OpenAI has between $600 billion and $1.4 trillion in debt, and no hope for profitability. Anthropic, who have a much more modest $35 billion in debt (by using Sam Altman as a human shield), also has no chance of ever being in the black, pending some freak event.by aj_hackman
- 100% — publish the hidden research, the value is in the discoveries, not in the dividend. With all due respect to the author, it feels like he missed the entire lesson of history.by aethelyon