this post was submitted on 25 Sep 2023

474 points (96.3% liked)

Technology

82490 readers

6428 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

474

Spotify is going to clone podcasters’ voices — and translate them to other languages (www.theverge.com)

submitted 2 years ago by stopthatgirl7@kbin.social to c/technology@lemmy.world

110 comments fedilink hide all child comments

A partnership with OpenAI will let podcasters replicate their voices to automatically create foreign-language versions of their shows.

you are viewing a single comment's thread
view the rest of the comments

[–] FireWire400@lemmy.world 45 points 2 years ago (3 children)

That's just weird... Part of the reason I listen to podcasts is that I just enjoy people talking about things and AI voices still have this uncanny quality to me

[–] danielbln@lemmy.world 27 points 2 years ago (3 children)

http://sndup.net/tb9p

[–] nehal3m@sh.itjust.works 13 points 2 years ago

Point taken, well done.

[–] TwinTusks@outpost.zeuslink.net 2 points 2 years ago

This is beautiful

[–] Hoimo@ani.social 2 points 2 years ago (1 children)

That's obviously way better than any TTS before it, but I still wouldn't want to listen to it for more than a few minutes. In these two sentences I can already hear some of the "AI quirks" and the longer you listen, the more you start to notice them.
I listen to a lot of AI celeb impersonations and they all sound like the same machine with different voice synthesizers. There's something about the prosody that gives it away, every sentence has the same generic pattern.
Humans are generally more creative, or more monotonous, but AI is in a weird inbetween space where it's never interested and never bored, always soulless.

[–] bamboo@lemm.ee 2 points 2 years ago

Having listened to it, I could not identify any sort of “AI quirk”. It sounded perfectly fine.

[–] sudoshakes@reddthat.com 16 points 2 years ago (1 children)

A large language model took a 3 second snippet of a voice and extrapolated from that the whole spoken English lexicon from that voice in a way that was indistinguishable from the real person to banking voice verification algorithms.

We are so far beyond what you think of when we say the word AI, because we replaced the underlying thing that it is without most people realizing it. The speed of large language models progress at current is mind boggling.

These models when shown FMRI data for a patient, can figure out what image the patient is looking at, and then render it. Patient looks at a picture of a giraffe in a jungle, and the model renders it having never before seen a giraffe… from brain scan data, in real time.

Not good enough? The same FMRI data was examined in real time by a large language model while a patient was watching a short movie and asked to think about what they saw in words. The sentence the person thought, was rendered as English sentences by the model, in real time, looking at fMRI data.

That’s a step from reading dreams and that too will happen inside 20 months.

We, are very much there.

[–] hobovision@lemm.ee 6 points 2 years ago (2 children)

Sources?

[–] sudoshakes@reddthat.com 2 points 2 years ago

Seeing Beyond the Brain: Conditional Diffusion Model with Sparse Masked Modeling for Vision Decoding: https://aiimpacts.org/2022-expert-survey-on-progress-in-ai/

High-resolution image reconstruction with latent diffusion models from human brain activity: https://www.biorxiv.org/content/10.1101/2022.11.18.517004v3

Semantic reconstruction of continuous language from non-invasive brain recordings: https://www.biorxiv.org/content/10.1101/2022.09.29.509744v1

[–] Pantoffel@feddit.de 1 points 2 years ago (1 children)

For the last example: Here

Rendering dreams from fMRI is also already reality. Please, google that yourself if you'd like to see the sources. However, the image quality is not yet very good, but nevertheless it is possible. It is just a question of when the quality will be better.

Now think about smart glasses or whatever display you like, controlling it with your mind. You'd need Jedi concentration :D But I sure do think I will live long enough to see this technology.

[–] sudoshakes@reddthat.com 1 points 2 years ago

[–] rigatti@lemmy.world 7 points 2 years ago (2 children)

It won't take long until that uncanny quality is worked out.

[–] danielbln@lemmy.world 5 points 2 years ago

Imho it has already been worked out. There is probably selection bias at play as you don't even recognize the AI voices that are already there.

[–] Pantoffel@feddit.de 1 points 2 years ago

Following up on the other comment.

The issue is that widely available speech models are not yet offering the quality that is technically possible. That is probably why you think we're not there yet. But we are.

Oh, I'm looking forward to just translate a whole audiobook into my native language and any speaking style I like.

Okay, perhaps we would still have difficulties with made up fantasy words or words from foreign languages with little training data.

Mind, this is already possible. It's just that I don't have access to this technology. I sincerely hope that there will be no gatekeeping to the training data, such that we can train such models ourselves.