> ## Documentation Index
> Fetch the complete documentation index at: https://docs.baseten.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Speech and audio

> Choose a path for transcription, speech generation, speaker diarization, and audio understanding on Baseten.

export const AudioDiagrams = () => {
  const duration = 5.021416666666667;
  const speakers = [{
    id: "SPEAKER_00",
    start: 0.25,
    end: 2.1959583333333335,
    text: "Could we move our meeting to Friday morning?"
  }, {
    id: "SPEAKER_01",
    start: 2.5459583333333335,
    end: 4.671416666666667,
    text: "Friday works. Let's meet at ten."
  }];
  const speakersAt = time => speakers.filter(s => time >= s.start && time < s.end).map(s => s.id);
  const waveform = [2, 2, 2, 2, 2, 5, 23, 6, 16, 17, 15, 15, 14, 14, 42, 32, 10, 11, 11, 13, 13, 8, 3, 2, 25, 15, 6, 5, 34, 30, 19, 14, 14, 11, 10, 11, 12, 9, 8, 12, 3, 2, 2, 2, 2, 2, 2, 2, 2, 4, 8, 38, 40, 24, 18, 19, 16, 11, 10, 16, 15, 11, 2, 18, 14, 5, 2, 2, 2, 4, 28, 32, 15, 18, 9, 12, 12, 10, 12, 7, 2, 11, 7, 15, 18, 13, 4, 2, 2, 2, 2, 2, 2, 2, 2, 2];
  const styles = `
  .audio-viz { --av-bg:#fff; --av-panel:#f6f8f5; --av-line:#d9e1d9; --av-ink:#183326; --av-muted:#53665a; --av-green:#146b40; --av-green-bg:#e5f5e9; --av-pink:#91366f; --av-pink-bg:#f8eaf3; color:var(--av-ink); border:1px solid var(--av-line); border-radius:12px; margin:24px 0; overflow:hidden; background:var(--av-bg); font-size:14px; line-height:1.5; }
  .dark .audio-viz { --av-bg:#071b10; --av-panel:#0f2518; --av-line:#334b3c; --av-ink:#e2eee5; --av-muted:#afc1b3; --av-green:#73dfa1; --av-green-bg:#163c28; --av-pink:#efb0d7; --av-pink-bg:#3e2336; }
  .audio-viz p { margin:0; }
  .audio-viz .av-mono { font-family:ui-monospace,SFMono-Regular,Menlo,monospace; font-size:12px; }
  .audio-viz .av-muted { color:var(--av-muted); }
  .audio-viz .av-kicker { font-size:11px; font-weight:600; letter-spacing:.1em; text-transform:uppercase; color:var(--av-muted); }
  .audio-viz .av-header { padding:20px 24px 16px; display:flex; justify-content:space-between; align-items:center; gap:12px; border-bottom:1px solid var(--av-line); }
  .audio-viz .av-header > .av-mono { white-space:nowrap; flex-shrink:0; }
  .audio-viz .av-title { font-size:16px; font-weight:600; margin:0 0 3px; }
  .audio-viz .av-body { padding:20px 24px; }
  .audio-viz .av-wave { position:relative; margin-top:12px; padding:8px 0; background:var(--av-panel); border:1px solid var(--av-line); border-radius:6px; }
  .audio-viz svg { display:block; width:100%; height:70px; }
  .audio-viz .av-ticks { position:relative; height:18px; margin-top:5px; color:var(--av-muted); }
  .audio-viz .av-ticks span { position:absolute; transform:translateX(-50%); white-space:nowrap; }
  .audio-viz .av-cursor { position:absolute; top:0; bottom:0; border-left:2px solid var(--av-ink); pointer-events:none; }
  .audio-viz .av-player { width:100%; height:36px; margin-top:16px; color-scheme:light; }
  .dark .audio-viz .av-player { color-scheme:dark; }
  .audio-viz .av-output { margin-top:20px; border-top:1px solid var(--av-line); padding-top:18px; }
  .audio-viz .av-track { display:grid; gap:6px; margin-top:14px; }
  .audio-viz .av-lane { position:relative; height:28px; background:var(--av-panel); border-radius:4px; border:1px solid var(--av-line); }
  .audio-viz .av-segment { position:absolute; height:100%; min-width:3px; border-radius:3px; background:var(--av-green-bg); border:1px solid var(--av-green); }
  .audio-viz .av-segment.pink { background:var(--av-pink-bg); border-color:var(--av-pink); }
  .audio-viz .av-speaker { color:var(--av-green); }
  .audio-viz .av-speaker.pink { color:var(--av-pink); }
  .audio-viz .av-turn { display:grid; grid-template-columns:100px 1fr; gap:12px; padding:12px; margin:12px -12px 0; border-radius:6px; border:1px solid transparent; }
  .audio-viz .av-turn[data-active=true] { border-color:var(--av-line); background:var(--av-panel); }
  .audio-viz .av-footer { padding:14px 24px; background:var(--av-panel); border-top:1px solid var(--av-line); font-size:12px; color:var(--av-muted); }
  .audio-viz .av-footer a { color:inherit; text-decoration:underline; }
  @media (max-width:540px) {
    .audio-viz .av-header, .audio-viz .av-body { padding:16px; }
    .audio-viz .av-footer { padding:14px 16px; }
    .audio-viz .av-turn { grid-template-columns:88px 1fr; gap:8px; }
    .audio-viz .av-mono { font-size:11px; }
  }
  @media (prefers-reduced-motion:reduce) { .audio-viz .av-cursor { display:none; } }
`;
  const [time, setTime] = React.useState(0);
  const [started, setStarted] = React.useState(false);
  const [failed, setFailed] = React.useState(false);
  const active = started && !failed ? speakersAt(time) : [];
  return <figure className="audio-viz" aria-label="Speaker diarization">
      <style>{styles}</style>
      <div className="av-header">
        <div>
          <div className="av-title">Speaker diarization</div>
          <p className="av-muted">
            Play the conversation to follow each speaker's turn.
          </p>
        </div>
        <span className="av-mono av-muted">{duration.toFixed(2)} s</span>
      </div>
      <div className="av-body">
        <div className="av-kicker">Audio in</div>
        <div className="av-wave">
          <svg viewBox="0 0 600 90" preserveAspectRatio="none" aria-hidden="true">
            {waveform.map((height, index) => <line key={index} x1={6 + index * 6.18} x2={6 + index * 6.18} y1={45 - height * 0.8} y2={45 + height * 0.8} stroke={speakersAt((index + 0.5) / waveform.length * duration).includes("SPEAKER_01") ? "var(--av-pink)" : "var(--av-green)"} strokeWidth="2.4" strokeLinecap="round" />)}
          </svg>
          {started && <div className="av-cursor" style={{
    left: `${Math.min(time / duration, 1) * 100}%`
  }} />}
        </div>
        <div className="av-ticks av-mono" aria-hidden="true">
          {[0, 1, 2, 3, 4, duration].map(tick => <span key={tick} style={{
    left: `${tick / duration * 100}%`,
    transform: tick === 0 ? "none" : tick === duration ? "translateX(-100%)" : undefined
  }}>
              {Number(tick.toFixed(2))} s
            </span>)}
        </div>
        <audio className="av-player" controls preload="metadata" aria-label="Play the generated conversation" onPlay={() => setStarted(true)} onTimeUpdate={event => setTime(event.currentTarget.currentTime)} onSeeked={event => {
    setStarted(true);
    setTime(event.currentTarget.currentTime);
  }} onEnded={() => setTime(duration)} onLoadedData={() => setFailed(false)} onError={() => setFailed(true)}>
          <source src="/assets/audio/kokoro-meeting-13887dbb.wav" type="audio/wav" onError={() => setFailed(true)} />
        </audio>
        {failed && <p role="status">
            Audio couldn't load. The transcript and diagrams are still available
            below.
          </p>}
        <div className="av-output">
          <div className="av-kicker">Speaker intervals out</div>
          {speakers.map((speaker, index) => <div className="av-track" key={speaker.id}>
              <span className={`av-mono av-speaker ${index ? "pink" : ""}`}>
                {speaker.id}
              </span>
              <div className="av-lane" aria-hidden="true">
                <div className={`av-segment ${index ? "pink" : ""}`} style={{
    left: `${speaker.start / duration * 100}%`,
    width: `${(speaker.end - speaker.start) / duration * 100}%`
  }} />
              </div>
            </div>)}
          {speakers.map((speaker, index) => <div className="av-turn" key={speaker.id} data-active={active.includes(speaker.id)}>
              <div>
                <span className={`av-mono av-speaker ${index ? "pink" : ""}`}>
                  {speaker.id}
                </span>
                <br />
                <span className="av-mono av-muted">
                  {speaker.start.toFixed(2)}–{speaker.end.toFixed(2)} s
                </span>
              </div>
              <p data-audio-transcript="true">{speaker.text}</p>
            </div>)}
        </div>
      </div>
      <figcaption className="av-footer">
        Audio generated with{" "}
        <a href="/examples/text-to-speech">Kokoro on Baseten</a>.
      </figcaption>
    </figure>;
};

Deploy speech-to-text and text-to-speech models on dedicated infrastructure, or send audio to audio-capable Model APIs. Choose batch or real-time transcription, speech generation, or audio understanding for your application.

Start with an optimized deployment from the [Model Library](https://www.baseten.co/library/). Each model card describes its input format, supported languages, hardware, and benchmarks. For custom preprocessing or inference code, define a [Truss model package](/development/model/overview) and deploy it with the [Baseten CLI](/reference/cli/baseten/model#push).

## Transcription

Convert speech to text with automatic speech recognition (ASR). Choose a client workflow that matches how audio reaches your application:

* **Batch transcription:** Send a complete audio file to [Whisper](https://www.baseten.co/library/whisper/) or [Qwen3-ASR](https://www.baseten.co/library/qwen-3-asr-1-7b/). Start with the [batch quickstart](/inference/audio/quickstart#transcribe-a-batch).
* **Real-time transcription:** Keep a WebSocket connection open and send audio as it arrives. Start with the [real-time quickstart](/inference/audio/quickstart#deploy-a-real-time-model) to deploy Qwen3-ASR and receive partial and final transcripts.

Real-time models have model-specific protocols. Qwen3-ASR uses JSON messages containing base64 audio; the real-time Whisper deployment accepts binary audio frames. Follow the [transcription WebSocket reference](/reference/inference-api/predict-endpoints/streaming-transcription-api) for the model you deploy.

## Speech generation

Convert text to speech (TTS) and return audio to your application. [Qwen3-TTS Streaming](https://www.baseten.co/library/qwen3-tts-12hz/) accepts text over WebSocket and returns audio incrementally. Its model card also describes voice cloning from reference audio.

Listen to speech samples in the [upstream Qwen3-TTS demo](https://huggingface.co/spaces/Qwen/Qwen3-TTS). Use the Baseten model card to select a deployment with the voice controls your application needs.

To write your own Python inference code, follow [Generate speech with Kokoro](/examples/text-to-speech). This example deploys a speech model with the Baseten CLI and returns generated audio.

## Speaker diarization

Associate time intervals in a recording with speaker labels. Use diarization when you need to separate speakers in a meeting, call, or interview. Labels such as `SPEAKER_00` distinguish speakers within the recording; they don't establish a person's identity.

For batch processing of complete recordings, use [MOSS Transcribe Diarize](/examples/models/transcription/moss-transcribe-diarize) to generate transcripts with speaker labels and timestamps. For standalone diarization, browse the [Model Library](https://www.baseten.co/library/) and review the [pyannote performance article](https://www.baseten.co/blog/how-baseten-makes-pyannotes-diarization-models-96x-faster/).

<AudioDiagrams />

## Audio understanding

Ask a model to summarize, reason about, or answer questions about a recording. [Audio input with Model APIs](/inference/model-apis/audio) sends audio and text in a request to a hosted model. Use this path when you need a model's response to the audio, with no dedicated deployment to manage.

## Voice agents

Connect real-time transcription to an LLM, then send the LLM's response to a speech-generation model. Use [Model APIs](/inference/model-apis/overview) for the LLM stage, or deploy an LLM on [Dedicated Inference](/deployment/concepts).

Review the selected model's [reasoning controls](/inference/model-apis/reasoning#control-reasoning-depth) when evaluating response latency. Measure the full turn from the end of the user's speech to the start of response playback, alongside the [speech latency measurements](/inference/audio/performance#latency-and-quality-measurements). For framework support, see the LiveKit entry in [Integrations](/inference/integrations).

## Custom speech models

Fine-tune a speech model with Training Jobs using the ML Cookbook:

* **[Qwen3-ASR](https://github.com/basetenlabs/ml-cookbook/tree/main/examples/qwen3-asr-transformers):** Prepare audio/transcript pairs, train a checkpoint, and deploy it for transcription.
* **[Whisper](https://github.com/basetenlabs/ml-cookbook/tree/main/examples/whisper-transformers):** Fine-tune Whisper with Transformers and serve the checkpoint with the recipe's inference example.
* **[Qwen3-TTS](https://github.com/basetenlabs/ml-cookbook/tree/main/examples/qwen3-tts-transformers):** Prepare speech data, fine-tune voice generation, and serve the checkpoint with the recipe's Truss implementation.

For checkpoint deployment paths, see [Serve your trained model](/training/deployment#speech-checkpoints).

## Next steps

* [Transcribe speech](/inference/audio/quickstart) with the batch or real-time client and a supplied recording.
* [Measure voice inference performance](/inference/audio/performance) to choose concurrency and evaluate deployment cost.
