9 Comments
User's avatar
Bob Williamson's avatar

This does sound promising. I agree with your diagnosis at the end (point c). The problem is licensing, a point we recognized 20 years ago in calling for open source in ML research. https://jmlr.csail.mit.edu/papers/v8/sonnenburg07a.html (where we also called for open sourcing the data); a fair part of that paper talked about licenses...

Hiding behind "fair use" (which was designed for a different time) is disingenuous. And so is the trickery of claiming one is just "pointing" to data https://laion.ai/faq/ when one is clearly doing more than that ("downloading and calculating CLIP embeddings"). It will be great if progress is made on respecting data rights fairly.

Finally, I do note that "open science" is open to all ... and is not a nationalistic enterprise! But perhaps good stuff can come from bad motivations.

Ben Recht's avatar

Hear hear.

FourierBot's avatar

It's slow to play Mexican train without having public trains/

Badri's avatar

Very much agree with the outcomes, but I am not sure the means will work out. Jensen's letter directly tackled the fear mongering about security and he brushed past the fair use, because it is complicated. He also argued like your previous post did that's it good for society with a clear focus on economic good for companies.

But the last call on open data likely is tricky. There is ThePile of course. Because opening the data would require upending copyright law. As you mentioned in the other post, frontier models can get away because they have huge legal teams and they can pay penalties and get out of it. And everyone who sues in time (unlike Elon for example) can share the riches of public data. They also pay other data brokers like ScaleAI and SurgeAI and Mercor to pay for more data and will fight to keep it theirs.

What can one do to keep public intelligence open? RMS style we can mandate everyone explicitly licenses what they say as open busting past the paywalls and putting it on open software etc., Another alternative is to make open data an economically sustainable ecosystem that no one can live without.

When you post on Google or Reddit or Medium, these companies can erect paywalls because you use their infrastructure. You probably signed away your data rights because no one reads TOS. They can turn around call this data their fairly earned "moat", also charge advertisers for access and manipulation, charge API costs and keep scrapers away.

But what if you make this moat disappear and build say a commercially sustainable version of Wikipedia or the whole open internet? I have some nutty ideas about not only making this completely "open data" commercially viable but also something no AI company can live without.

P.S: Instead of hijacking this thread, will discuss a natural segue

rif a saurous's avatar

Clearing up a couple confusions.

Model outputs are not copyrighted in the first place (Thaler v. Perlmutter 2025), so it's irrelevant to talk about whether distillation is "fair use" in a legal sense. The potential legal problems with distillation are that it's a TOS violation and/or a contract violation.

Also, "allowing the broader community fair use access to the same material the companies used" is incoherent. Training on copyrighted books was fair use for Anthropic, and it's fair use for everyone else already. The problem is that pirating the books was illegal for Anthropic, and they ended up paying $1.5B for it. Anyone can buy the books, scan on them, and train them, but if you distribute them you're in hot water (Hachette v. Internet Archive 2024), so the corpus can't really be "open" without a substantive law change.

Badri's avatar

Bingo. So we need to solve "the TOS problem".

Also notice how Anthropic decided to use "illicit" ahead of distillation and "fair use" when they did it. Emotive words and we are already failing if we continue how they direct us :)

aram harrow's avatar

Suppose we train models on lots of copyrighted text, like books, and suppose we find this legally unobjectionable because it's transformative and doesn't reproduce the original text. (All debatable but let's keep going.). In that world, there are technical and legal barriers to having an open corpus. How do you imagine addressing them?

eduardo sontag's avatar

Ben, as a long time user of gnuemacs, Linux, .... of course I agree with openness. However, in the current environment of huge sunken costs, maybe you could expand the discussion to suggest how the major players could keep some competitive advantage? If they cannot, the motivation to open will not be there. By analogy, AT&T could fund open research from Bell Labs only when it was a monopoly. Just trying to be realistic...

Badri's avatar

That's a surprising suggestion from an academic! I am in splits 😂 seeing the Simpson's video Ben linked under "slurried" to subtly tell us how much he cares about commercial viability or trusts a benevolent corporation to do this.

I explained "both sides" of this story to my 7 year old son and he replied "Too late" when discussing what we ought to do. AI has forced us to possibly re-legislate copyright law, fair use etc., but wonder if it can happen fast enough.