The book scanning as fair use is clearly madness. The whole spiritual point of fair use is that you shouldn't make a pain-in-the-ass fuss about people (PEOPLE) engaging with your work- that copying an excerpt for a student or making fun of it in a low-rent collage video shouldn't set the powers of the state delegated to the author by copyright on you. The notion that the largest piles of capital ever assembled by man routinely memorizing those texts and dispensing them for a fee in front of said authors are covered by those provisions is obscene- and I'm a pretty copyleft kind of person! I would, in fact, pirate that hard to find movie only available with seven hours of ads. But I am not a pile of hundreds of billions of dollars. Punch up, not down.
A hypothetical open corpus shouldn't include distillation of closed models, fruit of poison trees and all that, but the bookshredders arguing against distillation is cartoonish. If paying my fee entitles me to the output of a model (that likely represents, practically or morally, theft from authors), you're going to have the gall to suggest the *only* thing I can't use it for is making an LLM? Piss right off with that.
Generally I think the answer should be that an open corpus is unequivocally and enthusiastically open. Vigorous consent. If you can't make a decent question answering machine out of Wikipedia and Project Gutenberg, that's a technical problem to solve, and it's the one everyone needs to solve if they aren't satisfied with what current models can do, because that logarithmic data-to-performance curve means you've already lost.
"We can have long, pedantic quibbles about what’s legal and what’s not. We can also ask, 'What is right?' and 'What is just?'"
If only more people thought like this...
Once the labs have sucked up the entire internet and works of others, they proceed to pay countless dollars to companies like Mercor and Handshake that aggressively lure experts towards the wonderful gig of trying to automate the remainder of their knowledge and skills away. It's much more obviously perfectly legal but all just feels so dystopian and gross.
It's beyond angering when these companies just bulldoze through every legal or moral obstacle with their billions of dollars and then throw a fit about distillation attacks and constantly pretend to have the moral high ground on anything (how many open letters for slowing AI have these near trillion-dollar capex companies signed at this point again?)
Yep. RIP Aaron Swartz, who arrived and left too soon. and yes models built from collective cultural intelligence should remain publicly accessible. But if truly open models still depend on enormous compute and on training data gathered without meaningful consent, how do we prevent ‘openness’ from reproducing the same concentration of power and extraction you criticize in closed labs?”
Apart from the fantastic work by AI2 on Olmo, I'd also like to point out the work by Eleuther AI on GPT-NeoX and Pythia models (https://www.eleuther.ai/artifacts/gpt-neox) and (surprisingly) by Meta (or Facebook Research) on OPT. The Pythia models are open source in the truest sense, with a publicly downloadable training set, training code, and even the Wandb project with the loss curves from training (https://wandb.ai/eleutherai/pythia).
I would also like to point out the OPT models from Facebook Research, which are not fully open-source but come pretty close. They open-source all the code for training and even have a very detailed chronicle of how they trained their 175B model (https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/README.md)^1. Unfortunately, the team at Facebook Research did not release the dataset that was used for training, but provided a lot of details about it in their paper. The OPT models are almost ancient [LLM] history now, and they could just go ahead and release the details at this point.
I bring this up to illustrate that big companies have done this before. I understand the "wrong hands" argument, but given the capabilities models have today and how the internet has adapted to LLMs, it might be time for the big labs to give back to the community that helped create them.
1. The latter sections of the logbook are entirely taken up by the team handling GPUs crashing left and right in their massive cluster. And several sections are dedicated to asking why the loss keeps blowing up, with "Don't Panic" being their first line in their action plan for mitigation.
The book scanning as fair use is clearly madness. The whole spiritual point of fair use is that you shouldn't make a pain-in-the-ass fuss about people (PEOPLE) engaging with your work- that copying an excerpt for a student or making fun of it in a low-rent collage video shouldn't set the powers of the state delegated to the author by copyright on you. The notion that the largest piles of capital ever assembled by man routinely memorizing those texts and dispensing them for a fee in front of said authors are covered by those provisions is obscene- and I'm a pretty copyleft kind of person! I would, in fact, pirate that hard to find movie only available with seven hours of ads. But I am not a pile of hundreds of billions of dollars. Punch up, not down.
A hypothetical open corpus shouldn't include distillation of closed models, fruit of poison trees and all that, but the bookshredders arguing against distillation is cartoonish. If paying my fee entitles me to the output of a model (that likely represents, practically or morally, theft from authors), you're going to have the gall to suggest the *only* thing I can't use it for is making an LLM? Piss right off with that.
Generally I think the answer should be that an open corpus is unequivocally and enthusiastically open. Vigorous consent. If you can't make a decent question answering machine out of Wikipedia and Project Gutenberg, that's a technical problem to solve, and it's the one everyone needs to solve if they aren't satisfied with what current models can do, because that logarithmic data-to-performance curve means you've already lost.
"We can have long, pedantic quibbles about what’s legal and what’s not. We can also ask, 'What is right?' and 'What is just?'"
If only more people thought like this...
Once the labs have sucked up the entire internet and works of others, they proceed to pay countless dollars to companies like Mercor and Handshake that aggressively lure experts towards the wonderful gig of trying to automate the remainder of their knowledge and skills away. It's much more obviously perfectly legal but all just feels so dystopian and gross.
It's beyond angering when these companies just bulldoze through every legal or moral obstacle with their billions of dollars and then throw a fit about distillation attacks and constantly pretend to have the moral high ground on anything (how many open letters for slowing AI have these near trillion-dollar capex companies signed at this point again?)
Yep. RIP Aaron Swartz, who arrived and left too soon. and yes models built from collective cultural intelligence should remain publicly accessible. But if truly open models still depend on enormous compute and on training data gathered without meaningful consent, how do we prevent ‘openness’ from reproducing the same concentration of power and extraction you criticize in closed labs?”
Apart from the fantastic work by AI2 on Olmo, I'd also like to point out the work by Eleuther AI on GPT-NeoX and Pythia models (https://www.eleuther.ai/artifacts/gpt-neox) and (surprisingly) by Meta (or Facebook Research) on OPT. The Pythia models are open source in the truest sense, with a publicly downloadable training set, training code, and even the Wandb project with the loss curves from training (https://wandb.ai/eleutherai/pythia).
I would also like to point out the OPT models from Facebook Research, which are not fully open-source but come pretty close. They open-source all the code for training and even have a very detailed chronicle of how they trained their 175B model (https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/README.md)^1. Unfortunately, the team at Facebook Research did not release the dataset that was used for training, but provided a lot of details about it in their paper. The OPT models are almost ancient [LLM] history now, and they could just go ahead and release the details at this point.
I bring this up to illustrate that big companies have done this before. I understand the "wrong hands" argument, but given the capabilities models have today and how the internet has adapted to LLMs, it might be time for the big labs to give back to the community that helped create them.
1. The latter sections of the logbook are entirely taken up by the team handling GPUs crashing left and right in their massive cluster. And several sections are dedicated to asking why the loss keeps blowing up, with "Don't Panic" being their first line in their action plan for mitigation.