/v1/audio/transcriptions gets diarized output without changing its request shape.
Setup
Install the Baseten CLI and sign in, then install the OpenAI SDK.Install and sign in to BasetenFor other platforms or a specific version, see the Baseten CLI install reference.
- macOS or Linux
- Windows
Terminal
Install the OpenAI SDK
uvx truss login --browser and deploy with uvx truss push.
This preset serves MOSS-Transcribe-Diarize on one H100 with SGLang Omni pinned to v0.1.2, optimized for single-pass transcription that returns speaker labels and timestamps alongside the transcript.
Hardware
H100
Engine
SGLang (dev build)
Context
32K
Concurrency
128
Write the config
Create and move into the project directory:config.yaml and paste the following:
config.yaml
/v1/audio/transcriptions gets diarized output without changing its request shape.
Flags
Thestart_command passes these flags to the engine. Each one controls a runtime or serving behavior:
Deploy
Push the config to Baseten with the Baseten CLI, or with the Truss CLI if you prefer it:baseten model push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.
Call the model
Your deployment serves an OpenAI-compatible chat completions API at/v1/audio/transcriptions that accepts audio inputs.
Send audio as an audio_url content item on a chat message. The model returns the transcription as the assistant message content.
- Python
- cURL
main.py
Next steps
Call your model
Endpoint anatomy, authentication, and sync versus async inference
Autoscaling
Scale replicas with traffic, including scale to zero