Skip to main content
Microsoft’s VibeVoice-ASR speech-to-text model, returning JSON segments with speaker labels and timestamps through an OpenAI-compatible API.

Setup

Sign in to Baseten with Truss, then install the OpenAI SDK.
Sign in to Baseten
Install the OpenAI SDK
This preset serves VibeVoice-ASR on a single H100 through vLLM with an OpenAI-compatible chat completions endpoint, tuned for low-latency transcription with speaker labels and timestamps.

Hardware

H100

Engine

vLLM 0.14.1

Context

32K

Concurrency

32

Write the config

Create and move into the project directory:
Then create a file named config.yaml and paste the following:
config.yaml
This config runs the vllm/vllm-openai:v0.14.1 image with Microsoft’s VibeVoice plugin patches applied at startup, serving weights pre-mounted at /models/vibevoice-asr so cold starts skip the 9.2 GB Hugging Face download. The server runs in eager mode with a 32k context and up to 16 concurrent sequences, exposing the model as vibevoice on the chat completions endpoint.

Flags

The start_command passes these flags to the engine. Each one controls a runtime or serving behavior:

Deploy

Push the config to Baseten:
You should see output similar to:
truss push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.

Call the model

Your deployment serves an OpenAI-compatible chat completions API at /v1/chat/completions that accepts audio inputs. Send audio as an audio_url content item on a chat message. The model returns the transcription as the assistant message content.
main.py

Next steps

Call your model

Endpoint anatomy, authentication, and sync versus async inference

Autoscaling

Scale replicas with traffic, including scale to zero