287
you are viewing a single comment's thread
view the rest of the comments
[-] sramder@lemmy.world 11 points 1 year ago

Anyone know how many hours of training data it takes to build up a convincing model of someone’s voice? It was 10’s of hours when I did a bit of research a year ago… the article says social media is the likely source of training data for these scams, but that seems unlikely at this point.

[-] CrabLangEnjoyer@lemmy.world 10 points 1 year ago

A current state of the art ai model from Microsoft can achieve acceptable quality with about 3 seconds of audio. Commercially available stuff like eleven labs about 30 minutes. But quality will obviously vary heavily but then again they're using a low quality phone call so maybe not that important

[-] waylaidwanderer@lemmy.ca 1 points 1 year ago

ElevenLabs only needs 1 minute, but it also works with even shorter clips.

load more comments (2 replies)
load more comments (15 replies)
this post was submitted on 10 Oct 2023
287 points (97.0% liked)

Technology

59340 readers
1483 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


founded 1 year ago
MODERATORS