Microsoft’s VASA-1 can deepfake a person with one photo and one audio track (arstechnica.com)

submitted 10 months ago by pelespirit@sh.itjust.works to c/technology@lemmy.ml

17 comments fedilink hide all child comments

On Tuesday, Microsoft Research Asia unveiled VASA-1, an AI model that can create a synchronized animated video of a person talking or singing from a single photo and an existing audio track. In the future, it could power virtual avatars that render locally and don't require video feeds—or allow anyone with similar tools to take a photo of a person found online and make them appear to say whatever they want.

all 18 comments

sorted by: hot top controversial new old

[-] leds@feddit.dk 31 points 10 months ago

Great! When will this be included in teams? So that I can deepfake all meetings

[-] bhez@lemmy.ml 2 points 10 months ago

I'll give it a photo of myself from 10 years ago so that my coworkers don't realize that I'm getting old.

[-] Aatube@kbin.melroy.org 15 points 10 months ago

Why did they make this

[-] bitfucker@programming.dev 14 points 10 months ago

Someone is always bound to make this someday. At least the maker is announcing it which is decent enough. Actually, I have always thought that if AI can generate image and voice, what is stopping someone from identity theft? And BAM, we are now in an age where digital data will soon be unreliable unless we have protocol in-place to prove the origin of the data.

[-] duncesplayed@lemmy.one 5 points 10 months ago

If you pump out enough research papers, maybe Microsoft won't move you over to the Office team.

[-] cygnus@lemmy.ca 14 points 10 months ago* (last edited 10 months ago)

We're going to need strong digital signatures on everything, and we need it fast, else we won't be able to believe anything we see. It will be Steve Bannon's "flood the zone with shit" dream come true.

[-] simple@lemm.ee 6 points 10 months ago

We’re going to need strong digital signatures on everything

That won't help anything considering how easy it is to strip metadata.

[-] cygnus@lemmy.ca 9 points 10 months ago

I mean the opposite scenario, where if there's no signature we assume it's fake.

[-] catloaf@lemm.ee 2 points 10 months ago* (last edited 10 months ago)

We've had email forgery and signatures to prevent it for decades, but barely anyone does that either.

[-] simple@lemm.ee 12 points 10 months ago

That lip sync is scary good. It's still a little off, the teeth are weirdly stretchy, but nobody would notice it's a deepfake on first glance.

Seems very similar to Nvidia's idea of only having a moving photo for video calls to reduce bandwidth needed. Very nice.

[-] Aatube@kbin.melroy.org 4 points 10 months ago* (last edited 10 months ago)

We'd need better optimization and more powerful processing on ye average laputopu for that to happen.

[-] pelespirit@sh.itjust.works 9 points 10 months ago* (last edited 10 months ago)

It's terrifying and super cool at the same time. I think all of the execs at these big tech companies need to rewatch the Terminator.

Here's Gizmodo's take: https://gizmodo.com/weird-teeth-fake-microsoft-vasa-1-ai-free-video-creator-1851420514

[-] Vendetta9076@sh.itjust.works 8 points 10 months ago

Revenge porn machine go brrrrr.

Parents need to learn this stuff and teach their kids about it. Rumored nudes were enough to ruin kids lives at my highschool, nevermind "real" ones.

[-] skatrek47@sh.itjust.works 2 points 10 months ago

This is impressive and terrifying as hell… I’m pretty tech savvy and have watched other AI videos, but if you presented this to me as real, I’d totally believe it! Especially I can imagine it being spliced with B roll footage and maybe “real” footage to make it even more believable. I’m floored…

[-] p03locke@lemmy.dbzer0.com 1 points 10 months ago

No. No, they can't. This shit still takes lots and lots of training data.

It's just like any job. You can't just fully fake something in one day. At best, you might get 60% of the way there, maybe 80% after adding on some generic experience. But, you're not going to fully mimic anything without lots of training and experience.

[-] autotldr@lemmings.world 0 points 10 months ago

This is the best summary I could come up with:

On Tuesday, Microsoft Research Asia unveiled VASA-1, an AI model that can create a synchronized animated video of a person talking or singing from a single photo and an existing audio track.

In the future, it could power virtual avatars that render locally and don't require video feeds—or allow anyone with similar tools to take a photo of a person found online and make them appear to say whatever they want.

To show off the model, Microsoft created a VASA-1 research page featuring many sample videos of the tool in action, including people singing and speaking in sync with pre-recorded audio tracks.

The examples also include some more fanciful generations, such as Mona Lisa rapping to an audio track of Anne Hathaway performing a "Paparazzi" song on Conan O'Brien.

While the Microsoft researchers tout potential positive applications like enhancing educational equity, improving accessibility, and providing therapeutic companionship, the technology could also easily be misused.

"We are opposed to any behavior to create misleading or harmful contents of real persons, and are interested in applying our technique for advancing forgery detection," write the researchers.

The original article contains 797 words, the summary contains 183 words. Saved 77%. I'm a bot and I'm open source!

this post was submitted on 19 Apr 2024

69 points (97.3% liked)

Technology

36011 readers

28 users here now

This is the official technology community of Lemmy.ml for all news related to creation and use of technology, and to facilitate civil, meaningful discussion around it.

Ask in DM before posting product reviews or ads. All such posts otherwise are subject to removal.

Rules:

1: All Lemmy rules apply

2: Do not post low effort posts

3: NEVER post naziped*gore stuff

4: Always post article URLs or their archived version URLs as sources, NOT screenshots. Help the blind users.

5: personal rants of Big Tech CEOs like Elon Musk are unwelcome (does not include posts about their companies affecting wide range of people)

6: no advertisement posts unless verified as legitimate and non-exploitative/non-consumerist

7: crypto related posts, unless essential, are disallowed

founded 5 years ago

MODERATORS

MinutePhrase@lemmy.ml