Brazil: Children's private photos misused by AI tools

AI Video & Visuals


(São Paulo, Brazil) – Intimate photos of children in Brazil are being used to create powerful artificial intelligence (AI) tools without the children's knowledge or consent, Human Rights Watch said today. The photos are harvested from the web and compiled into large datasets that companies use to train their AI tools. Other companies then use these tools to create malicious deepfakes, putting even more children at risk of exploitation and harm.

“Children should not have to live in fear that their photos will be stolen and used as a weapon against them,” said Hae Jeong Han, children's rights and technology researcher and activist at Human Rights Watch. “Governments should urgently adopt policies to protect children's data from misuse by AI.”

Human Rights Watch's analysis found that the LAION-5B dataset, built by scraping large swaths of the internet and used to train popular AI tools, contains links to photos that could identify children in Brazil. Some of the children's names are listed in the captions attached to the photos or in the URLs where the images are stored. In many cases, their identities are easily traceable, including information such as the time and location of the children when the photo was taken.

In one of those photos, a 2-year-old girl's lips part in amazement as she touches the tiny fingers of her newborn sister. Embedded in the photo are captions and information not only listing the children's names, but also the name and exact location of the Santa Catarina hospital where the babies were born that winter's afternoon nine years ago.

Human Rights Watch found 170 photos of children from at least 10 states (Alagoas, Bahia, Ceará, Mato Grosso do Sul, Minas Gerais, Paraná, Rio de Janeiro, Rio Grande do Sul, Santa Catarina, and São Paulo). Human Rights Watch identified less than 0.0001 percent of the 5.85 billion images and captions in the dataset, which may be a significant underestimate of the total amount of children's personal data present in LAION-5B.

The photographs reviewed span the entire childhood span, capturing intimate moments such as a baby being born into the gloved hands of a doctor, a young child blowing out the candles on a birthday cake or dancing in his underwear at home, a student giving a presentation at school and a teenager having his picture taken at a high school carnival.

Many of these photos were originally only seen by a small number of people and appear to have had some degree of privacy. They appear impossible to find through online searches. Some of these photos were posted by children, their parents, or their families on personal blogs or photo and video sharing sites. Some were uploaded years or even a decade before LAION-5B was built.

Once the data is collected and fed into AI systems, these children face additional threats to their privacy due to flaws in the technology. AI models, including those trained on LAION-5B, are notorious for leaking personal information and can reproduce identical copies of the materials they were trained on, such as medical records or photos of real people. Guardrails that some companies have set up to prevent sensitive data from leaking have been repeatedly breached.

These privacy risks lead to further damage. By training on photos of real children, the AI ​​model could create a convincing clone of any child from just a few photos, or even a single image. Bad actors used AI tools trained on LAION to generate explicit images of children from benign photos, as well as explicit images scraped onto LAION-5B to capture images of children who had been sexually abused.

Similarly, the inclusion of Brazilian children in LAION-5B contributes to the ability of AI models trained on this dataset to generate realistic images of Brazilian children, which significantly increases the existing risks children face that someone could steal their likeness from photos or videos they post online and use AI to manipulate them into saying or doing things they never said or did.

At least 85 girls from the states of Alagoas, Minas Gerais, Pernambuco, Rio de Janeiro, Rio Grande do Sul and São Paulo reported being harassed by classmates who used AI tools to create sexually explicit deepfakes based on photos taken from the girls' social media profiles and then circulated the fake images online.

Fabricated media has always existed, but it took time, resources, and expertise to create, and it was usually not very realistic. Today's AI tools can create lifelike output in seconds and are often free and easy to use, but non-consensual deepfakes are rampant and risk being recirculated online for a lifetime, causing lasting harm.

In response, LAION, the German non-profit that manages LAION-5B, confirmed that the dataset contains the intimate photos of children found by Human Rights Watch and promised to remove them. LAION disputed that the AI ​​models trained on LAION-5B could reproduce the personal data verbatim. LAION also argued that children and their guardians are responsible for removing their intimate photos from the internet, which is the most effective safeguard against misuse.

Lawmakers have proposed banning the use of AI to generate sexually explicit images of people, including children, without their consent. While these efforts are urgent and important, they address only one symptom of a deeper problem: children's personal data is barely protected from misuse. Brazil's Data Protection Law (Lei Geral de Proteção de Dados Pessoais, or General Personal Data Protection Law), as written, does not provide sufficient protection for children.

The government should strengthen data protection laws by introducing additional and comprehensive safeguards for children's data privacy. In April, the National Council for the Rights of Children and Adolescents, a deliberative body established by law to protect children's rights, issued a resolution directing the council and the Ministry of Human Rights and Citizenship to develop a national policy for protecting the rights of children and adolescents in the digital environment within 90 days. The government should do so.

Given privacy risks and potential new misuses as technology evolves, new policies should prohibit the inclusion of children's personal data in AI systems, prohibit the digital reproduction or manipulation of children's likenesses without their consent, and provide victimized children with mechanisms to seek meaningful justice and redress.

Brazil's Congress should also ensure that proposed AI regulations incorporate data privacy protections for everyone, especially children.

“Generative AI is still an early stage technology, and the associated harms that children are already experiencing are not inevitable,” Han said. “Protecting the privacy of children's data now will help shape the development of this technology to one that advances, rather than violates, children's rights.”



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *