Voices cloned by AI now fool even informed listeners

AI News


Scientists have discovered that voices replicated using artificial intelligence are so good that most people mistake them for real speakers.

Familiar voices that once served as reliable forms of identification are now far less certain in everyday decisions, arguments, and high-stakes encounters.

tests that most people fail


earth snap

In controlled listening tests designed to investigate the reliability of speech, members of the public repeatedly confused synthetic speech with recordings of real individuals.

Researchers at the University of California, Berkeley, including doctoral student Sarah Barrington, documented this breakdown by tracking who was speaking and how listeners decided whether the voice sounded authentic.

Through these judgments, the artificial voices were very similar to human voices, and listeners often treated them as the same person.

That consistency revealed the narrow margin between perception and error, setting up deeper questions about identity and realism to be explored next.

Identity clues fool listeners

During identity matching, people often treated duplicate voices as the same speaker, making mistakes common.

“If you put two voices side by side, there is only a 20% chance that people will tell that they are not the same identity,” Barrington says.

The clone retained stable cues such as accent, pitch, and pace, so listeners could map it onto a familiar figure without noticing the cracks.

Voice recognition no longer protects claims of identity, so a familiar caller may not be who they say they are.

Listeners misread the audio

When listeners judged naturalness, they only correctly classified voices as real or fake 60% of the time.

Voice clones, which are synthetic voices tailored to someone’s voice, can sound so smooth that breathing or pausing is useless.

The longer the clip, the more judgment people can make, and spontaneous speech is easier to spot fake than short scripted lines.

Modern generators already mimic everyday speech patterns, so listeners waiting for a robotic tone will miss a lot of fakery.

Fraud that takes advantage of trust

When voice clones sound normal, scammers can use their trust to force victims into making quick and costly decisions.

In January 2024, an AI-generated recording imitating President Biden urged Democratic voters in New Hampshire to skip the primary.

In response to these incidents, the Federal Communications Commission (FCC) ruled that AI voice robocalls violate federal law.

Plans that cost around $5 a month could accelerate the wave of spoofing, as some services clone audio from just a few seconds of audio.

Train detectors ethically

To provide better training data for the detector, the DeepSpeak dataset combines real videos with deepfakes, which are media modified or generated to imitate a person.

In various versions, 500 participants, ranging in age from 18 to 75, gathered to speak and make simple gestures to the camera.

“The problem with current deepfake datasets is that they are not collected consensually, do not use state-of-the-art technical tools, and lack diversity in the types and environments in which deepfakes are created,” Barrington said.

DeepSpeak combines video and audio fakes that swap faces and change lip movements to give the detector a more difficult and realistic target.

Most detection tools work after the fact, scanning files for small clues once the call ends.

During live calls, the software must listen continuously and flag problems immediately, increasing privacy risks.

A system that monitors all conversations should avoid false alarms. Otherwise, people will ignore the warning and trust the wrong call.

Until real-time checking improves, human judgment will continue to carry the burden, even as audio continues to improve.

Habits that reduce risk

Simple habits can reduce risk, especially if the caller asks for money, a password, or a simple verification of your name.

Longer back-and-forth conversations require clones to deal with surprises, and that extra load can cause timing and phrasing glitches.

If you are a careful listener, you can ask open-ended questions and then call back using a known number instead of the number provided.

Although these measures do not guarantee safety, they can reduce social pressure and buy you time to verify your identity elsewhere.

Improving your personal routine can be helpful, but platform rules will determine whether anyone can easily clone voices at scale.

Some groups are pushing for the creation of digital labels that record content credentials, media origins, to track audio across platforms.

Another line of defense is watermarks. It’s a hidden signal added to allow software to flag AI media, but it leaves a gap when the tool opts out.

Stronger checks on who can generate clones and how the output is labeled could reduce harm without requiring listeners to be experts.

Under the Federal Rules of Evidence, authentication can depend in part on whether the audio sounds familiar.

When voice cloning becomes good enough, witnesses can honestly say that even if a recording sounds right, it’s still wrong.

This risk has led judges and lawyers to use technological checks such as secure call logs and recordings that keep the security trail clear.

Without stronger standards, a persuasive voice can change the verdict, but the actual speaker has little way to prove their innocence.

This study revealed a simple truth. This means that audio contains less evidence of identity than people assume in everyday life.

When a single call can cause lasting damage, better detection, clearer labeling, and updated evidence rules become paramount.

This research nature.

—–

Like what you read? Subscribe to our newsletter for fascinating articles, exclusive content and the latest updates.

Check us out on EarthSnap, the free app from Eric Ralls and Earth.com.

—–



Source link