Five major publishers and bestselling author Scott Turow filed a lawsuit against Meta Platforms and CEO Mark Zuckerberg on May 5, 2026. They claim Meta illegally copied more than 267 terabytes of books and articles to train its Llama artificial intelligence models. This volume exceeds the entire collection of the U.S. Library of Congress.
According to the lawsuit, Meta obtained the material from "shadow libraries" like LibGen and Anna's Archive, which host unlicensed copies of books and academic journals. The company then removed copyright information from the works, erasing digital fingerprints that identify ownership.
The complaint alleges that Mark Zuckerberg personally approved this approach. In early 2023, Meta was negotiating with publishers to license content legally, with discussions about a $200 million budget for datasets. However, these talks were abandoned after reaching Zuckerberg.
The publishers argue that Llama produces copies of copyrighted text and imitates specific authors' styles. Interestingly, Meta has signed licensing agreements with news organizations like Reuters and CNN, but refused to license books for this project.
This case follows similar copyright disputes in AI training. In June 2025, a federal judge ruled in Meta's favor in another case, citing "fair use" protections. However, Anthropic agreed to a $1.5 billion settlement with authors in 2025, showing courts are increasingly sympathetic to creators' concerns.