Dokumentationerreichbar
OpenAI, Voice activity detection (VAD)
developers.openai.com (externe Seite)
Die Seite hinter dem Schalter, der entscheidet, ob eine Aufnahme überhaupt beim Modell ankommt. Zwei Betriebsarten stehen dort nebeneinander: eine Lautstärkeschwelle, die lauteren Ton verlangt und in lauter Umgebung besser fährt, und eine zweite, die stattdessen anhand der gesprochenen Wörter entscheidet, ob jemand zu Ende geredet hat. Beleg dafür, dass die Lautstärkeprüfung vor der Erkennung ein dokumentierter Griff der Anbieter ist, und dafür, wo sie aufhört.
geprüft 24.09.2026
Worauf sich diese Seite beruft, wörtlich, abgerufen am 07.08.2026:
A higher threshold will require louder audio to activate the model, and thus might perform better in noisy environments.
bestätigt 24.09.2026It uses periods of silence to automatically chunk the audio.
bestätigt 24.09.2026Semantic VAD is a new mode that uses a semantic classifier to detect when the user has finished speaking, based on the words they have uttered.
bestätigt 24.09.2026With this mode, the model is less likely to interrupt the user during a speech-to-speech conversation, or chunk a transcript before the user is done speaking.
bestätigt 24.09.2026
Voice Activity Detection im Glossar02 Sprache und Audio03 Wenn das Gespräch nicht in Runden läuft