Discussion about this post

User's avatar
Bob Williamson's avatar

This does sound promising. I agree with your diagnosis at the end (point c). The problem is licensing, a point we recognized 20 years ago in calling for open source in ML research. https://jmlr.csail.mit.edu/papers/v8/sonnenburg07a.html (where we also called for open sourcing the data); a fair part of that paper talked about licenses...

Hiding behind "fair use" (which was designed for a different time) is disingenuous. And so is the trickery of claiming one is just "pointing" to data https://laion.ai/faq/ when one is clearly doing more than that ("downloading and calculating CLIP embeddings"). It will be great if progress is made on respecting data rights fairly.

Finally, I do note that "open science" is open to all ... and is not a nationalistic enterprise! But perhaps good stuff can come from bad motivations.

Badri's avatar

Very much agree with the outcomes, but I am not sure the means will work out. Jensen's letter directly tackled the fear mongering about security and he brushed past the fair use, because it is complicated. He also argued like your previous post did that's it good for society with a clear focus on economic good for companies.

But the last call on open data likely is tricky. There is ThePile of course. Because opening the data would require upending copyright law. As you mentioned in the other post, frontier models can get away because they have huge legal teams and they can pay penalties and get out of it. And everyone who sues in time (unlike Elon for example) can share the riches of public data. They also pay other data brokers like ScaleAI and SurgeAI and Mercor to pay for more data and will fight to keep it theirs.

What can one do to keep public intelligence open? RMS style we can mandate everyone explicitly licenses what they say as open busting past the paywalls and putting it on open software etc., Another alternative is to make open data an economically sustainable ecosystem that no one can live without.

When you post on Google or Reddit or Medium, these companies can erect paywalls because you use their infrastructure. You probably signed away your data rights because no one reads TOS. They can turn around call this data their fairly earned "moat", also charge advertisers for access and manipulation, charge API costs and keep scrapers away.

But what if you make this moat disappear and build say a commercially sustainable version of Wikipedia or the whole open internet? I have some nutty ideas about not only making this completely "open data" commercially viable but also something no AI company can live without.

P.S: Instead of hijacking this thread, will discuss a natural segue

5 more comments...

No posts

Ready for more?