23 Comments
User's avatar
Bob Williamson's avatar

This does sound promising. I agree with your diagnosis at the end (point c). The problem is licensing, a point we recognized 20 years ago in calling for open source in ML research. https://jmlr.csail.mit.edu/papers/v8/sonnenburg07a.html (where we also called for open sourcing the data); a fair part of that paper talked about licenses...

Hiding behind "fair use" (which was designed for a different time) is disingenuous. And so is the trickery of claiming one is just "pointing" to data https://laion.ai/faq/ when one is clearly doing more than that ("downloading and calculating CLIP embeddings"). It will be great if progress is made on respecting data rights fairly.

Finally, I do note that "open science" is open to all ... and is not a nationalistic enterprise! But perhaps good stuff can come from bad motivations.

Nathan Lambert's avatar

I agree and appreciate you continuing to advocate for this! There are still some people fighting this fight, if you ever want to contribute more directly, you know where to find me (or Percy, Ion, and others)

Justin Kay's avatar

Yes, 100% - but (genuine question) how much of current performance comes from pre-training on appropriated text (Wikipedia, Github, scanned library books, etc.) vs. private task-specific training data generated by the big labs? See e.g. fallout over Meta’s recent ‘AI Task Force’ of software engineers reassigned to create training data full time.

It seems that much of the success of the real ‘products’ (e.g. coding agents) coming out of these labs is due to the latter - but how do we contend with that for open models? If the goal is open models with these capabilities, who will generate that training data? And why should they? Bosses will drool over opportunities to automate away their employees regardless of where the models come from. We’re going to have to think beyond just replicating capabilities if we want these things to actually be good for the general public.

Ben Recht's avatar

The only thing I can say here is that the open models I've used (such as GLM 5.2 and Kimi 3) are excellent for coding, even with open-source agent harnesses, and Meta still hasn't released a particularly compelling model for coding.

Kathleen Weber's avatar

Here is an interesting talk by the founder of DeepSeek. it sounds like you and he are pretty much on the same wavelength. At least, I thought this deserved to be more widely read in the US than It is likely to be.

https://www.geopolitechs.org/p/deepseek-founder-liang-wenfeng-in

Badri's avatar

Very much agree with the outcomes, but I am not sure the means will work out. Jensen's letter directly tackled the fear mongering about security and he brushed past the fair use, because it is complicated. He also argued like your previous post did that's it good for society with a clear focus on economic good for companies.

But the last call on open data likely is tricky. There is ThePile of course. Because opening the data would require upending copyright law. As you mentioned in the other post, frontier models can get away because they have huge legal teams and they can pay penalties and get out of it. And everyone who sues in time (unlike Elon for example) can share the riches of public data. They also pay other data brokers like ScaleAI and SurgeAI and Mercor to pay for more data and will fight to keep it theirs.

What can one do to keep public intelligence open? RMS style we can mandate everyone explicitly licenses what they say as open busting past the paywalls and putting it on open software etc., Another alternative is to make open data an economically sustainable ecosystem that no one can live without.

When you post on Google or Reddit or Medium, these companies can erect paywalls because you use their infrastructure. You probably signed away your data rights because no one reads TOS. They can turn around call this data their fairly earned "moat", also charge advertisers for access and manipulation, charge API costs and keep scrapers away.

But what if you make this moat disappear and build say a commercially sustainable version of Wikipedia or the whole open internet? I have some nutty ideas about not only making this completely "open data" commercially viable but also something no AI company can live without.

P.S: Instead of hijacking this thread, will discuss a natural segue

The Catalyst Shift's avatar

I agree on the importance of opensource and that as society we need to support its development as much as we can.

I disagree, with the argument, that frontier labs should make data sets open. The only reason we have the AI models and the whole ecosystems around them from GPUs to datacenters to power is that a lot of people see an economic opportunity developing and building these amazing products. Once you remove this motivation development and progress stop.

So no, make Anthropic and OpenAI defend their position in open competition. Do not make data sets public, instead support competition, start ups, creative destruction.

Think AI's avatar

Open weights are only part of the equation. Without the training code and a transparent, legally usable corpus, the community cannot truly reproduce the model. The difficult part is building an open corpus that respects creators’ rights and privacy. If that can be solved, public intelligence could become shared infrastructure instead of another corporate moat.

Michael D. Green, PhD's avatar

I’m curious for a) :re young people get the incentives aligned to do that? As someone in my late 20s I struggle to see people moving towards this route. We are pushed to believe we are exceptional and can strike it big (even when the odds of becoming rich are more concentrated in terms of paths to do so imo)

Suman Suhag's avatar

A World at a Crossroads

In countries like the USA, the UK, Japan, and throughout the world, people are grappling with interconnected issues. From extreme weather events-intensified storms, heatwaves, and floods due to climate change-to escalating costs that strain household budgets and access to essentials like food, shelter, and healthcare, we face multifaceted challenges. Simultaneously, the breakneck pace of AI development presents awe-inspiring opportunities, alongside profound anxieties about the future of work, the erosion of privacy, and the amplification of misinformation.

Heightened political polarization and waning faith in established institutions hamper our capacity to address these issues collaboratively. In addition, mental health challenges, alienation, and pervasive stress continue to impact millions worldwide, particularly the youth. Our natural world is under immense pressure, with its vital ecosystems like forests and oceans, and diverse wildlife, facing unprecedented threats from a growing pool of environmental hazards such as pollution and climate change.

Despite these immense challenges, a beacon of hope shines through innovation: advancements in renewable energy, AI, medicine, and green technologies are developing solutions. Our collective future hinges upon the ability of governments, corporations, and individuals to collaborate effectively, safeguarding our planet, fostering resilient communities, and deploying technology wisely.

The decisions we take now will undoubtedly determine the legacy we leave behind for future generations. Whether individually or collectively, every action counts.

Alex Tolley's avatar

The need for huge datasets may not be necessary. Smaller, curated data may be sufficient to provide the same performance. I have seen examples of orders of magnitude improvement in reduced size and latency, but without loss of performance, by reducing the dataset and quantizing the weights.

I also think that open-source data sets, perhaps selectable by subject or keywords, should be available to refine selected models. At some point, many more people could develop custom models. This could create an ecosystem, rather like the book markets,

Alex Tolley's avatar

I read that there is a new mechanism being readied to control one's comments and posts [via Cloudflare?] so that they can be freely acquired, blocked, or available for a fee.

FourierBot's avatar

It's slow to play Mexican train without having public trains/

rif a saurous's avatar

Clearing up a couple confusions.

Model outputs are not copyrighted in the first place (Thaler v. Perlmutter 2025), so it's irrelevant to talk about whether distillation is "fair use" in a legal sense. The potential legal problems with distillation are that it's a TOS violation and/or a contract violation.

Also, "allowing the broader community fair use access to the same material the companies used" is incoherent. Training on copyrighted books was fair use for Anthropic, and it's fair use for everyone else already. The problem is that pirating the books was illegal for Anthropic, and they ended up paying $1.5B for it. Anyone can buy the books, scan on them, and train them, but if you distribute them you're in hot water (Hachette v. Internet Archive 2024), so the corpus can't really be "open" without a substantive law change.

Badri's avatar

Bingo. So we need to solve "the TOS problem".

Also notice how Anthropic decided to use "illicit" ahead of distillation and "fair use" when they did it. Emotive words and we are already failing if we continue how they direct us :)

aram harrow's avatar

Suppose we train models on lots of copyrighted text, like books, and suppose we find this legally unobjectionable because it's transformative and doesn't reproduce the original text. (All debatable but let's keep going.). In that world, there are technical and legal barriers to having an open corpus. How do you imagine addressing them?

Ben Recht's avatar

What do you have in mind as barriers?

aram harrow's avatar

If you want to train on copyrighted books then my understanding is that you need to pay for each book, then you can scan it, but you can't legally make that scan public. (that's the legal barrier I mentioned.) So how can that book be in a public corpus? Maybe the copyright holder helps you in some interactive way that doesn't leak the full text, but that sounds complicated. (That's the technical barrier.)

I agree it'd be great for the field to have an open corpus. But this issue seems pretty hard to get around. Maybe arxiv+wikipedia+reddit are enough already, though?

eduardo sontag's avatar

Ben, as a long time user of gnuemacs, Linux, .... of course I agree with openness. However, in the current environment of huge sunken costs, maybe you could expand the discussion to suggest how the major players could keep some competitive advantage? If they cannot, the motivation to open will not be there. By analogy, AT&T could fund open research from Bell Labs only when it was a monopoly. Just trying to be realistic...

Ben Recht's avatar

The big companies were committed to open source machine learning up until very recently. Google open-sourced TensorFlow, and support open development of Jax. Their Gemma model, though not Gemini, is open weights. Facebook supported PyTorch and the most successful open source model of the day, Llama. Even OpenAI has released an open weights model. These companies know they can make money while still supporting a vibrant open source ecosystem.

Badri's avatar

That's a surprising suggestion from an academic! I am in splits 😂 seeing the Simpson's video Ben linked under "slurried" to subtly tell us how much he cares about commercial viability or trusts a benevolent corporation to do this.

I explained "both sides" of this story to my 7 year old son and he replied "Too late" when discussing what we ought to do. AI has forced us to possibly re-legislate copyright law, fair use etc., but wonder if it can happen fast enough.