Setup
Install the Baseten CLI and sign in, then install thewebsockets library.
Install and sign in to BasetenFor other platforms or a specific version, see the Baseten CLI install reference.
- macOS or Linux
- Windows
Terminal
Install websockets
uvx truss login --browser and deploy with uvx truss push.
Pick the model you want to deploy. Each tab is a self-contained recipe.
- Mini 4B
- Mini 3B Batch
mistralai/Voxtral-Mini-4B-Realtime-2602 is a 4B-parameter encoder-decoder model.This preset serves Voxtral Mini Realtime on an RTX PRO 6000, tuned for low-latency streaming transcription.Then create a file named You should see output similar to:
Hardware
RTX_PRO_6000
Engine
vLLM (0.22.0 custom build)
Context
10K
Concurrency
48
Write the config
Create and move into the project directory:config.yaml and paste the following:config.yaml
Flags
Thestart_command passes these flags to the engine. Each one controls a runtime or serving behavior:Deploy
Push the config to Baseten with the Baseten CLI, or with the Truss CLI if you prefer it:baseten model push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.Call the model
This preset exposes a WebSocket streaming endpoint at/v1/realtime for low-latency, incremental transcription. See the streaming transcription API reference for the message protocol, Python client example, and supported audio formats.Next steps
Call your model
Endpoint anatomy, authentication, and sync versus async inference
Autoscaling
Scale replicas with traffic, including scale to zero