NachschlagenQuellenregister

Ankündigung des Herstellerserreichbar

WaveNet: A generative model for raw audio

deepmind.google (externe Seite)

Die Ankündigung vom 08.09.2016 beim Hersteller, vier Tage vor der Arbeit dazu, und der Ort, an dem die beiden alten Bauformen der Sprachausgabe erklärt werden. Sie definiert nebenbei den Vocoder als das Bauteil, durch das die parametrischen Verfahren ihre Ausgabe schicken, und nennt die Zahl, an der die Aufgabe hängt: 16.000 Abtastwerte je Sekunde.

geprüft 24.09.2026

Worauf sich diese Seite beruft, wörtlich, abgerufen am 05.09.2026:

  • concatenative TTS, where a very large database of short speech fragments are recorded from a single speaker and then recombined to form complete utterances.bestätigt 24.09.2026
  • This makes it difficult to modify the voice (for example switching to a different speaker, or altering the emphasis or emotion of their speech) without recording a whole new database.bestätigt 24.09.2026
  • parametric TTS, where all the information required to generate the data is stored in the parameters of the model, and the contents and characteristics of the speech can be controlled via the inputs to the model.bestätigt 24.09.2026
  • Existing parametric models typically generate audio signals by passing their outputs through signal processing algorithms known as vocoders.bestätigt 24.09.2026
  • Researchers usually avoid modelling raw audio because it ticks so quickly: typically 16,000 samples per second or more, with important structure at many time-scales.bestätigt 24.09.2026
  • MOS are a standard measure for subjective sound quality tests, and were obtained in blind tests with human subjects (from over 500 ratings on 100 test sentences).bestätigt 24.09.2026

Vocoder im Glossar08 Sprachsynthese

Alle Quellen

Tippen Sie los.

↑↓ auswählenEnter öffnenDie Suche läuft im Browser. Nichts wird übertragen.