

Discussion summary
The discussion centers on concerns about copyright issues related to AI and collective knowledge, with some proposing a 'Corpus Royalty' to address it. Participants debate the nature of copyright infringement and the impact of AI on public trust and knowledge sharing.
What the discussion says
- The author fears future copyright litigation against AI labs.
- Some argue that producing copyrighted material isn't infringement unless it violates fair use.
- There is concern about the privatization of collective knowledge and its implications.
- Debate on whether open source LLMs exist and the impact of AI on public commons.
“This is the private capture of public genius. The biggest heist in human history.”
“The ability to produce copyrighted material is not copyright infringement.”
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- This is a well written essay. I had hoped it might address the role of distillation and open source in diffusing ownership of this technology back to the public that made it possible. And the AI labs’ rank hypocrisy in this area.by anon373839
- > returns constrained to a relatively conservative (by today’s standards) ~7% per annum
Damn, we should have something like that market wide. Progressive, with the revenue.
Forbidding vertical integration would be a tremendous blessing too.
by scotty79 - It was a brilliant article, and it succinctly captured the offenses to ethics and humanism posed by LLMs.
I'm not sure it'll get a lot of reception in the technocracy here on HN, whether of the AI booster or AI nihilist sort. However, I think it's a very comprehensive digestion of the questions that will swirl around the idea of LLMs as a public good in the near to medium future.
by abalashov - The troubles over copyright infringement in AI training data remind me a bit of Eli Whitney and the cotton gin.
There he suffered massive patent infringement, that basically stopped being enforced due to the sheer economic importance of the cotton gin.
In a similar manner, I think there is a reasonably strong argument that it was wrong to use copyrighted material for AI training without paying royalties nor even asking for permission. But equally, every country wants to have the most powerful models and enforcing such royalties would make it effectively impossible to train them as the amount of material required would cost an insane amount in royalty fees.
So I expect the law will continue to turn a blind eye (perhaps enforcing some token payments like that $1.5B mentioned in the article) because "if we don't make these models, the Chinese will" etc.
- "Fair use" was always fuzzy. To be honest, I care a lot less about slurping up the public internet and private books to make models than about every profession on the planet being forced by their employers to create skills that automate their knowledge work. The latter is much more directly an expropriation, legitimized only by the shortage of work, i.e., market power.by w10-1
- We cannot always want to capture only the (temporary) winners whenever we see a lucrative business and expect to share a free ride. I'd also assume that most of the revenue these AI labs are making is turned into depreciating fixed capital (hardware) and OPEX at this point.
Why don't we capture Meta and Google as they allegedly take advantage of more publicly available information for profit? Let alone the truly valuable knowledge, like mathematics, has nothing to do with the majority of garbage posts that an average person would "contribute" on social media.
If we really want to tax or nationalize some economic activity, then, in my opinion, the target should be what it takes from society, not what it produces for society. By this logic, we should tax all labs, including those lagging ones, that utilize the public knowledge.
However, if everyone can access the public knowledge without rendering it less useful or reducing its available quantity, there should be no reason to tax it.
by typ - > Frontier science looks different today. It's rooted in model weights and GPUs. It is flooded with token spend and agentic loops. It blooms in data centers.
This seems like handwaving to me. Even the actual frontier science that is using ML (e.g. AlphaFold) isn't based on "token spend" or "agentic loops". I personally cannot think of a single example of frontier science that is rooted in LLMs. I am sure there are a few examples, but the idea that frontier science has somehow completely shifted its trajectory and methods based on LLMs feels like quite a stretch to me.
Got any counterexamples to show me I'm wrong?
- This was interesting right up until "The fund pays every eligible American the same amount each year. "
I'm in Australia. I've contributed my share of dirt to the delta. Why do I not get a share of this?
I get that the frontier companies are (for the moment) US companies. But that's just corporate ownership, it's not what we're talking about. We're talking about compensating the people who wrote the training data for their contribution. That contribution came from all over the world, so the Corpus Fund needs to be paid all over the world.
Set it up in the UN, get the UN to provide the training data sets as a common good, and have the UN collect the money from all AI companies using the training data sets. And the UN should distribute the money in the most equitable manner globally (so most of it going to alleviate poverty, probably).
I'd happily trade my collected years of shitposts to help folks get out of poverty.