Meta's Controversial Use of Pirated Books for AI Training
The Story
A judge has unredacted court documents revealing that Meta allegedly used the pirated "shadow library" LibGen to train its AI models, despite internal warnings that doing so could undermine the company’s position with regulators.
The filings suggest that top executives, including Mark Zuckerberg, were aware of the dataset's illicit origins and that Meta may have even distributed pirated works by "seeding" torrent files during the training process.
The filings suggest that top executives, including Mark Zuckerberg, were aware of the dataset's illicit origins and that Meta may have even distributed pirated works by "seeding" torrent files during the training process.
Why It Matters
When Judge Vince Chhabria granted Meta summary judgment in June 2025, he conceded the ruling sat in significant tension with reality because the thirteen authors never proved market harm from Llama training Mashable. That ended the training claims only; Meta went on arguing its torrent uploads were inherent to downloading and fair use besides. The industry regrouped instead: Elsevier, Cengage, Hachette, Macmillan and McGraw Hill joined Scott Turow in a Manhattan class action on May 5 carrying the substitution evidence Chhabria found missing, after Anthropic paid $1.5 billion to settle its parallel piracy case Reuters.
Go Deeper
Read the original reporting at Wired.
Read Full Story at Wired →