Sarah Silverman and novelists sue ChatGPT-maker OpenAI for ingesting their books

Advertisement

Advertise with us

Ask ChatGPT about comedian Sarah Silverman's memoir “The Bedwetter” and the artificial intelligence chatbot can come up with a detailed synopsis of every part of the book.

Read this article for free:


or

Already have an account? Log in here »

To continue reading, please subscribe:

Subscribe and receive a limited-edition Free Press branded hat or tote.

Digital Subscription

One year of digital access for only $205*

  • Enjoy unlimited reading on winnipegfreepress.com
  • Read the E-Edition, our digital replica newspaper
  • Access News Break, our award-winning app
  • Play interactive puzzles

*First annual payment billed as $205.00 + GST for one year. This annual subscription will automatically renew at $233.00 + GST every 52 weeks (10% off the regular annual price of $259.35). Offer available to new and qualified returning subscribers only. Cancel any time.

To continue reading, please subscribe:

Add Free Press access to your Brandon Sun subscription for only an additional

$1 for the first 4 weeks*

  • Enjoy unlimited reading on winnipegfreepress.com
  • Read the E-Edition, our digital replica newspaper
  • Access News Break, our award-winning app
  • Play interactive puzzles
Start now

*Your next Brandon Sun subscription payment will increase by $1.00 and you will be charged $17.95 plus GST for four weeks. After four weeks, your payment will increase to $24.95 plus GST every four weeks.

Hey there, time traveller!
This article was published 12/07/2023 (1164 days ago), so information in it may no longer be current.

Ask ChatGPT about comedian Sarah Silverman’s memoir “The Bedwetter” and the artificial intelligence chatbot can come up with a detailed synopsis of every part of the book.

Does that mean it effectively “read” and memorized a pirated copy? Or it scraped so many customer reviews and online chatter about the bestseller or the musical it inspired that it passes for an expert?

The U.S. courts may now help sort that out after Silverman sued ChatGPT-maker OpenAI for copyright infringement this week, joining a growing number of writers who say they unwittingly built the foundation for Silicon Valley’s red-hot AI boom.

File - Sarah Silverman introduces a performance at the 75th annual Tony Awards on Sunday, June 12, 2022, in New York. Silverman sued ChatGPT-maker OpenAI for copyright infringement this week, joining a growing number of writers who say they unwittingly built the foundation for Silicon Valley's red-hot AI boom. (Photo by Charles Sykes/Invision/AP, File)
File - Sarah Silverman introduces a performance at the 75th annual Tony Awards on Sunday, June 12, 2022, in New York. Silverman sued ChatGPT-maker OpenAI for copyright infringement this week, joining a growing number of writers who say they unwittingly built the foundation for Silicon Valley's red-hot AI boom. (Photo by Charles Sykes/Invision/AP, File)

Silverman’s lawsuit says she never gave permission for OpenAI to ingest the digital version of her 2010 book to train its AI models, and it was likely stolen from a “shadow library” of pirated works. It says the memoir was copied “without consent, without credit, and without compensation.”

It’s one of a mounting number of cases that could crack open the secrecy of OpenAI and its rivals about the valuable data used to train increasingly widely used “generative AI” products that create new text, images and music. And it raises questions about the ethical and legal bedrock of tools that the McKinsey Global Institute projects will add the equivalent of $2.6 trillion to $4.4 trillion to the global economy.

“This is an open, dirty secret of the whole machine learning industry,” said Matthew Butterick, one of the lawyers representing Silverman and other authors in seeking a class-action case. “They love book data and they get it from these illicit sites. We’re kind of blowing the whistle on that whole practice.”

OpenAI declined to comment on the allegations. Another lawsuit from Silverman makes similar claims about an AI model built by Facebook and Instagram parent company Meta, which also declined comment.

It may be a tough case for writers to win, especially after Google’s success in beating back legal challenges to its online book library. The U.S. Supreme Court in 2016 let stand lower court rulings that rejected authors’ claim that Google’s digitizing of millions of books and showing small portions of them to the public amount to “copyright infringement on an epic scale.”

“I think what OpenAI has done with books is awfully close to what Google was allowed to do with its Google Books project and so will be legal,” said Deven Desai, associate professor of law and ethics at the Georgia Institute of Technology.

While only a handful have sued, including Silverman and bestselling novelists Mona Awad and Paul Tremblay, concerns about the tech industry’s AI-building practices have gained traction in literary and artist communities.

Other prominent authors — among them Nora Roberts, Margaret Atwood, Louise Erdrich and Jodi Picoult — signed a letter late last month to the CEOs of OpenAI, Google, Microsoft, Meta and other AI developers accusing them of exploitative practices in building chatbots that “mimic and regurgitate” their language, style and ideas.

“Millions of copyrighted books, articles, essays, and poetry provide the ‘food’ for AI systems, endless meals for which there has been no bill,” said the open letter organized by the Authors Guild and signed by more than 4,000 writers. “You’re spending billions of dollars to develop AI technology. It is only fair that you compensate us for using our writings, without which AI would be banal and extremely limited.”

The AI systems behind popular products such as ChatGPT, Google’s Bard and Microsoft’s Bing chatbot are known as large language models that have “learned” by analyzing and picking up patterns from a wide body of ingested text. They’ve awed the public with their strong command of the human language, though they’re also known for a tendency to spout falsehoods.

While the models have also been trained on news articles and social media feeds, books are particularly valuable, as OpenAI acknowledged in a 2018 paper cited in Silverman’s lawsuit.

“You’re spending billions of dollars to develop AI technology. It is only fair that you compensate us for using our writings, without which AI would be banal and extremely limited.”–Authors Guild letter

The earliest version of OpenAI’s large language model, known as GPT-1, relied on a dataset compiled by university researchers called the Toronto Book Corpus that included thousands of unpublished books, some in the adventure, fantasy and romance genres.

“Crucially, it contains long stretches of contiguous text, which allows the generative model to learn to condition on long-range information,” OpenAI researchers said at the time. Other tech companies such as Google and Amazon also relied on the same data, which is no longer available in its original form.

But since then, OpenAI and other top AI developers have grown more secretive about their sources of data, even as they have ingested even larger troves of written works. Butterick said circumstantial evidence points to the use of so-called shadow libraries of pirated content that held the works of Silverman and other plaintiffs.

“It’s important for their models because books are the best source of long-form, well-edited, coherent writing,” he said. “You basically can’t have a high-quality language model unless you have books in your training data.”

It could be weeks or months before a formal response is due from OpenAI. But once the case proceeds, tech executives could have to testify, under oath, about what sources of books they downloaded.

“As far as we know, the other side hasn’t denied it,” said Joseph Saveri, another of Silverman’s lawyers. “They don’t have an alternative explanation for this.”

Saveri said authors aren’t necessarily asking tech companies to throw away their algorithms and training data and start over — though the U.S. Federal Trade Commission has set a precedent for forcing companies to destroy ill-gotten AI data. But some way of compensating writers is needed, he said.

Report Error Submit a Tip

More Stories

Lauded Manitoba author chronicled life in the Interlake

By Sheldon Birnie 5 minute read Preview

Lauded Manitoba author chronicled life in the Interlake

By Sheldon Birnie 5 minute read Wednesday, Sep. 16, 2026

Longtime Manitoba author, editor and professor David Arnason died on Monday at age 86.

Read
Wednesday, Sep. 16, 2026

Kinew delivered what most of us wanted; now we’ll have to deal with the consequences

Dan Lett 5 minute read Preview

Kinew delivered what most of us wanted; now we’ll have to deal with the consequences

Dan Lett 5 minute read Updated: 8:43 AM CDT

Premier Wab Kinew has decided to adopt permanent daylight savings time, which means a permanent end to the spring forward, fall back tradition of adjusting our clocks. And absolutely no one is surprised.

Read
Updated: 8:43 AM CDT

Jets GM insists club’s not working under any deadline in netminder’s trade request

Mike McIntyre 7 minute read Preview

Jets GM insists club’s not working under any deadline in netminder’s trade request

Mike McIntyre 7 minute read Updated: Yesterday at 5:27 PM CDT

Connor Hellebuyck and his family were subject to so much post-Olympic vitriol that both the Winnipeg Jets security team and even police had to get involved.

Read
Updated: Yesterday at 5:27 PM CDT

Manitoba abandoning spring, fall clock changes

Carol Sanders 8 minute read Preview

Manitoba abandoning spring, fall clock changes

Carol Sanders 8 minute read Updated: 8:48 AM CDT

Time’s up for the time change in Manitoba.

Manitobans won’t be setting their clocks back an hour in November. The province has decided to permanently stay on daylight time after surveying residents on the twice annual clock adjustment.

“We heard from Manitobans loud and clear — no more changing the clocks in Manitoba,” Premier Wab Kinew said in a news release Thursday.

Most of the 72,000 votes submitted in an online EngageMB survey — 92 per cent — said they would like to see the practice end. The survey asked if they preferred earlier sunrises in winter with standard time or the later sunsets in summer with daylight time. More than half of survey participants (58 per cent) indicated a preference to stay on permanent daylight time; 34 per cent voted to eliminate the time change but stay on standard time.

Read
Updated: 8:48 AM CDT

Province rips Bell over second 911 outage

Chris Kitching 6 minute read Preview

Province rips Bell over second 911 outage

Chris Kitching 6 minute read Updated: Yesterday at 6:08 PM CDT

A 911 outage that lasted about 90 minutes in Manitoba and other provinces early Thursday morning was caused by a problem with Bell’s next-generation network amid concerns about recurring disruptions.

Read
Updated: Yesterday at 6:08 PM CDT

Northern First Nation calls for voluntary ban of moose hunting as numbers dwindle

Eva Wasney 5 minute read Preview

Northern First Nation calls for voluntary ban of moose hunting as numbers dwindle

Eva Wasney 5 minute read 2:12 PM CDT

A northern Manitoba First Nation is calling for a voluntary closure of moose hunting in its traditional territory to protect the declining population and respond to what it sees is a lack of provincial intervention.

On Friday, Misipawistik Cree Nation issued a plea urging hunters to abstain from harvesting moose in provincial Game Hunting Area 10 — a 10,000-square-kilometre region of boreal forest north of Grand Rapids. The request applies to this year’s moose-hunting seasons, which open Monday.

According to the province’s 2025 aerial big-game survey, an estimated 281 moose reside in the area, roughly 60 of which are bull moose. That’s down from an estimated 346 moose in 2013, which the survey describes as a “relatively stable population” over the past decade.

This year, 124 licences and 62 bull moose tags were made available in the region during Manitoba’s annual hunting draw.

Read
2:12 PM CDT