Study shows AI image-generators being trained on explicit photos of children

Advertisement

Advertise with us

Hidden inside the foundation of popular artificial intelligence image-generators are thousands of images of child sexual abuse, according to a new report that urges companies to take action to address a harmful flaw in the technology they built.

Read this article for free:


or

Already have an account? Log in here »

To continue reading, please subscribe:

Subscribe and receive a limited-edition Free Press branded hat or tote.

Digital Subscription

One year of digital access for only $205*

  • Enjoy unlimited reading on winnipegfreepress.com
  • Read the E-Edition, our digital replica newspaper
  • Access News Break, our award-winning app
  • Play interactive puzzles

*First annual payment billed as $205.00 + GST for one year. This annual subscription will automatically renew at $233.00 + GST every 52 weeks (10% off the regular annual price of $259.35). Offer available to new and qualified returning subscribers only. Cancel any time.

To continue reading, please subscribe:

Add Free Press access to your Brandon Sun subscription for only an additional

$1 for the first 4 weeks*

  • Enjoy unlimited reading on winnipegfreepress.com
  • Read the E-Edition, our digital replica newspaper
  • Access News Break, our award-winning app
  • Play interactive puzzles
Start now

*Your next Brandon Sun subscription payment will increase by $1.00 and you will be charged $17.95 plus GST for four weeks. After four weeks, your payment will increase to $24.95 plus GST every four weeks.

Hey there, time traveller!
This article was published 20/12/2023 (1004 days ago), so information in it may no longer be current.

Hidden inside the foundation of popular artificial intelligence image-generators are thousands of images of child sexual abuse, according to a new report that urges companies to take action to address a harmful flaw in the technology they built.

Those same images have made it easier for AI systems to produce realistic and explicit imagery of fake children as well as transform social media photos of fully clothed real teens into nudes, much to the alarm of schools and law enforcement around the world.

Until recently, anti-abuse researchers thought the only way that some unchecked AI tools produced abusive imagery of children was by essentially combining what they’ve learned from two separate buckets of online images — adult pornography and benign photos of kids.

David Thiel, chief technologist at the Stanford Internet Observatory and author of its report that discovered images of child sexual abuse in the data used to train artificial intelligence image-generators, poses for a photo on Wednesday, Dec. 20, 2023, in Óbidos, Portugal. (Camilla Mendes dos Santos via AP)
David Thiel, chief technologist at the Stanford Internet Observatory and author of its report that discovered images of child sexual abuse in the data used to train artificial intelligence image-generators, poses for a photo on Wednesday, Dec. 20, 2023, in Óbidos, Portugal. (Camilla Mendes dos Santos via AP)

But the Stanford Internet Observatory found more than 3,200 images of suspected child sexual abuse in the giant AI database LAION, an index of online images and captions that’s been used to train leading AI image-makers such as Stable Diffusion. The watchdog group based at Stanford University worked with the Canadian Centre for Child Protection and other anti-abuse charities to identify the illegal material and report the original photo links to law enforcement. It said roughly 1,000 of the images it found were externally validated.

The response was immediate. On the eve of the Wednesday release of the Stanford Internet Observatory’s report, LAION told The Associated Press it was temporarily removing its datasets.

LAION, which stands for the nonprofit Large-scale Artificial Intelligence Open Network, said in a statement that it “has a zero tolerance policy for illegal content and in an abundance of caution, we have taken down the LAION datasets to ensure they are safe before republishing them.”

While the images account for just a fraction of LAION’s index of some 5.8 billion images, the Stanford group says it is likely influencing the ability of AI tools to generate harmful outputs and reinforcing the prior abuse of real victims who appear multiple times.

It’s not an easy problem to fix, and traces back to many generative AI projects being “effectively rushed to market” and made widely accessible because the field is so competitive, said Stanford Internet Observatory’s chief technologist David Thiel, who authored the report.

“Taking an entire internet-wide scrape and making that dataset to train models is something that should have been confined to a research operation, if anything, and is not something that should have been open-sourced without a lot more rigorous attention,” Thiel said in an interview.

A prominent LAION user that helped shape the dataset’s development is London-based startup Stability AI, maker of the Stable Diffusion text-to-image models. New versions of Stable Diffusion have made it much harder to create harmful content, but an older version introduced last year — which Stability AI says it didn’t release — is still baked into other applications and tools and remains “the most popular model for generating explicit imagery,” according to the Stanford report.

“We can’t take that back. That model is in the hands of many people on their local machines,” said Lloyd Richardson, director of information technology at the Canadian Centre for Child Protection, which runs Canada’s hotline for reporting online sexual exploitation.

Stability AI on Wednesday said it only hosts filtered versions of Stable Diffusion and that “since taking over the exclusive development of Stable Diffusion, Stability AI has taken proactive steps to mitigate the risk of misuse.”

“Those filters remove unsafe content from reaching the models,” the company said in a prepared statement. “By removing that content before it ever reaches the model, we can help to prevent the model from generating unsafe content.”

LAION was the brainchild of a German researcher and teacher, Christoph Schuhmann, who told the AP earlier this year that part of the reason to make such a huge visual database publicly accessible was to ensure that the future of AI development isn’t controlled by a handful of powerful companies.

“It will be much safer and much more fair if we can democratize it so that the whole research community and the whole general public can benefit from it,” he said.

Much of LAION’s data comes from another source, Common Crawl, a repository of data constantly trawled from the open internet, but Common Crawl’s executive director, Rich Skrenta, said it was “incumbent on” LAION to scan and filter what it took before making use of it.

LAION said this week it developed “rigorous filters” to detect and remove illegal content before releasing its datasets and is still working to improve those filters. The Stanford report acknowledged LAION’s developers made some attempts to filter out “underage” explicit content but might have done a better job had they consulted earlier with child safety experts.

FILE - Students walk on the Stanford University campus on March 14, 2019, in Stanford, Calif. Hidden inside the foundation of popular artificial intelligence image-generators are thousands of images of child sexual abuse, according to a new report from the Stanford Internet Observatory that urges technology companies to take action to address a harmful flaw in the technology they built. (AP Photo/Ben Margot, File)
FILE - Students walk on the Stanford University campus on March 14, 2019, in Stanford, Calif. Hidden inside the foundation of popular artificial intelligence image-generators are thousands of images of child sexual abuse, according to a new report from the Stanford Internet Observatory that urges technology companies to take action to address a harmful flaw in the technology they built. (AP Photo/Ben Margot, File)

Many text-to-image generators are derived in some way from the LAION database, though it’s not always clear which ones. OpenAI, maker of DALL-E and ChatGPT, said it doesn’t use LAION and has fine-tuned its models to refuse requests for sexual content involving minors.

Google built its text-to-image Imagen model based on a LAION dataset but decided against making it public in 2022 after an audit of the database “uncovered a wide range of inappropriate content including pornographic imagery, racist slurs, and harmful social stereotypes.”

Trying to clean up the data retroactively is difficult, so the Stanford Internet Observatory is calling for more drastic measures. One is for anyone who’s built training sets off of LAION‐5B — named for the more than 5 billion image-text pairs it contains — to “delete them or work with intermediaries to clean the material.” Another is to effectively make an older version of Stable Diffusion disappear from all but the darkest corners of the internet.

“Legitimate platforms can stop offering versions of it for download,” particularly if they are frequently used to generate abusive images and have no safeguards to block them, Thiel said.

As an example, Thiel called out CivitAI, a platform that’s favored by people making AI-generated pornography but which he said lacks safety measures to weigh it against making images of children. The report also calls on AI company Hugging Face, which distributes the training data for models, to implement better methods to report and remove links to abusive material.

Hugging Face said it is regularly working with regulators and child safety groups to identify and remove abusive material. Meanwhile, CivitAI said it has “strict policies” on the generation of images depicting children and has rolled out updates to provide more safeguards. The company also said it is working to ensure its policies are “adapting and growing” as the technology evolves.

The Stanford report also questions whether any photos of children — even the most benign — should be fed into AI systems without their family’s consent due to protections in the federal Children’s Online Privacy Protection Act.

Rebecca Portnoff, the director of data science at the anti-child sexual abuse organization Thorn, said her organization has conducted research that shows the prevalence of AI-generated images among abusers is small, but growing consistently.

Developers can mitigate these harms by making sure the datasets they use to develop AI models are clean of abuse materials. Portnoff said there are also opportunities to mitigate harmful uses down the line after models are already in circulation.

Tech companies and child safety groups currently assign videos and images a “hash” — unique digital signatures — to track and take down child abuse materials. According to Portnoff, the same concept can be applied to AI models that are being misused.

“It’s not currently happening,” she said. “But it’s something that in my opinion can and should be done.”

Report Error Submit a Tip

More Stories

Steinbach Credit Union member warns others after scammer steals $10K

Scott Billeck 5 minute read Preview

Steinbach Credit Union member warns others after scammer steals $10K

Scott Billeck 5 minute read Yesterday at 5:19 PM CDT

A Winnipeg single mother whose bank account was drained of nearly $10,000 says Steinbach Credit Union needs to warn its members about what she believes is a broader security problem.

The woman, who is not being named to protect her identity, said a scammer gained access to her account earlier this week after calling SCU’s contact centre, claiming to be her and successfully completing the credit union’s authentication process.

She said she knows of at least six others who have experienced similar incidents in which callers impersonate SCU account holders and pass the authentication checks.

The woman discovered something was wrong when her minor son asked why money had been transferred out of an account they jointly hold.

Read
Yesterday at 5:19 PM CDT

Canada must not accept another Iranian autocracy

David Matas 5 minute read 2:01 AM CDT

Iran is a state run by a radical government victimizing its citizens and its neighbours. How do we change that? Any regime change that repeats the mistakes of the past is not a solution.

What went wrong in the past which led to the current regime? One answer is the U.K. and American aided coup which led to the ouster in 1953 of a democratically elected government headed at the time by Mohammad Mossadegh and its replacement with an American friendly autocrat, the Shah of Iran, Mohammad Reza Pahlavi.

The Shah imposed his rule through the Iranian National Intelligence and Security Organization — SAVAK, the acronym of its Persian name. SAVAK systematically inflicted torture, arbitrary killings and extra-judicial killings on perceived opponents of the regime. It was only a matter of time before repression of a regime directed to serving foreign interests would be overthrown.

What replaced the regime of the Shah, in 1979, the current regime of the mullahs, is a regime as intolerant and violent as the regime of the Shah and then some. That was not inevitable. But neither was it surprising. The vicious nature of the current regime, following the example set by the Shah, and its hatred for the United States, reacting to the U.S. aided imposition of the Shah, both echoes and reacts to the past. In 1979, the U.S. reaped what it sowed in 1953.

Siloam Mission, WFPS discussing on-site paramedic to reduce staggering number of emergency calls

Scott Billeck 4 minute read Preview

Siloam Mission, WFPS discussing on-site paramedic to reduce staggering number of emergency calls

Scott Billeck 4 minute read Yesterday at 2:00 AM CDT

Winnipeg’s fire-paramedic service and its largest homeless-serving organization are discussing what it would look like to station at least one paramedic at Siloam Mission, as emergency calls to the downtown facility surge.

The ongoing discussions come as Winnipeg Fire Paramedic Service crews were dispatched 1,771 times during the first eight months of the year to the two adjacent Siloam properties at 303 Stanley St. and 300 Princess St.

WFPS said most of the calls were related to overdoses or other consequences of substance use, with the total nearly double the 854 calls recorded over the same period last year. By comparison, there were 831 calls during the same span in 2024, and 591 in 2023.

“The number of emergency calls isn’t sustainable for us as a service, but it’s certainly not sustainable for them as an agency trying to provide services,” WFPS chief Ryan Sneath told the Free Press. “And they recognize that they’re impacting emergency services, not them specifically, but the site in a significant way.”

Read
Yesterday at 2:00 AM CDT

A life’s story: activist, educator, organizer knew the true meaning of community

Janine LeGal 7 minute read Preview

A life’s story: activist, educator, organizer knew the true meaning of community

Janine LeGal 7 minute read 3:00 AM CDT

Somewhere between a love for sharing stories and a love of community connections, Michael (Mike) Maunder lived his life. Giving voice to people and neighbourhoods was a passion for the natural-born community organizer.

An author, teacher and principal, Maunder lived in a one-bedroom apartment in West Broadway and devoted himself to community development, education and neighbourhood renewal.

He maintained a connection with Augustine United Church, whose commitment to social justice he appreciated, volunteered at the Sandy-Saulteaux Spiritual Centre, and found comfort in a Buddhist prayer community.

Maunder died on July 15 at the age of 80.

Read
3:00 AM CDT

American Hostage stars recall warmth of Winnipeggers during brutal winter

Randall King 7 minute read Preview

American Hostage stars recall warmth of Winnipeggers during brutal winter

Randall King 7 minute read Yesterday at 6:13 PM CDT

TORONTO — When you stream the eight-part series American Hostage, a sense of familiarity may strike.

That’s because the show was shot in Winnipeg — especially in the Exchange District — from November of last year to the beginning of February, with Winnipeg portraying Indianapolis circa 1977.

The whole world seemed to be watching that midwest American city that year as a hostage crisis unfolded involving one Tony Kiritsis, an angry man who goes ballistic when a mortgage brokerage was poised to foreclose on property he had dreamed of developing into a shopping centre.

Kiritsis kidnapped his mortgage broker, Richard Hall, wiring a shotgun to the back of his head, and demanded a public airing of his grievances, with the reluctant help of popular radio broadcaster Fred Heckman.

Read
Yesterday at 6:13 PM CDT

A quick look at where provinces and territories stand on time changes

The Canadian Press 2 minute read Preview

A quick look at where provinces and territories stand on time changes

The Canadian Press 2 minute read Updated: Yesterday at 6:28 AM CDT

WINNIPEG - Manitoba's government said Thursday the province will be sticking with daylight time year-round, making it the latest jurisdiction to move away from twice-annual time changes. Here's a look at where things stand across Canada:

British Columbia said in March that it was moving to permanent Pacific daylight time. The change also means that parts of northeastern B.C., which have been on mountain standard time year round, are now in alignment with the rest of the province.

Alberta also switched to permanent daylight time this year in what it calls Alberta time. The change means it now has the same time as Saskatchewan year round.

Saskatchewan has been on permanent central standard time since 1966.

Read
Updated: Yesterday at 6:28 AM CDT