Dokumentationerreichbar
OpenAI, Using realtime models
developers.openai.com (externe Seite)
Die Anleitung zum Anweisen eines Sprachmodells, das rohen Ton verarbeitet. Sie belegt nebenbei, was ein solches Modell auseinanderhalten kann: Stille, Hintergrundgeräusch, Wartemusik, Fernsehton und ein Gespräch, das jemand anderem gilt. Dieselbe Seite führt die Ansage vor einer Wartezeit als eigenes Bauteil mit Namen und benennt ihre Kehrseite: Schlecht gemacht erhöht sie die gefühlte Wartezeit. Die Anweisungen stehen dort für zwei Modellfassungen nebeneinander und lauten verschieden; zitiert ist hier die neuere.
geprüft 24.09.2026
Worauf sich diese Seite beruft, wörtlich, abgerufen am 05.08.2026:
If the latest audio is silence, background noise, hold music, TV audio, side conversation, or speech not addressed to you
bestätigt 24.09.2026Treat audio as unclear if it is ambiguous, noisy, silent, unintelligible, partially cut off, or if you are unsure of the exact words the user said.
bestätigt 24.09.2026Preambles are short spoken updates that keep a voice agent feeling responsive while it reasons, looks something up, or calls a tool.
bestätigt 24.09.2026Before calling tools in the commentary channel, briefly tell the user what you are doing.
bestätigt 24.09.2026Used well, they reassure the user that the assistant is working. Used poorly, they become filler and increase perceived latency.
bestätigt 24.09.2026gpt-realtime-2 generates preambles by default.
bestätigt 24.09.2026you are about to call a tool that may take noticeable time
bestätigt 24.09.2026
Preamble im Glossar02 Sprache und Audio05 Wissen anbinden06 Tools anbinden