Skip to main content
Alibaba’s Qwen3-ASR is a compact 1.7B speech-to-text model with multilingual transcription support.

Setup

Install the Baseten CLI and sign in, then install the OpenAI SDK.
Install and sign in to Baseten
Terminal
For other platforms or a specific version, see the Baseten CLI install reference.
Install the OpenAI SDK
Prefer not to install? Sign in with uvx truss login --browser and deploy with uvx truss push. This preset serves Qwen3-ASR on a single RTX PRO 6000 through vLLM, tuned for fast multilingual transcription.

Hardware

RTX_PRO_6000

Engine

vLLM (0.22.0-cu129 build)

Concurrency

256

Write the config

Create and move into the project directory:
Then create a file named config.yaml and paste the following:
config.yaml

Flags

The start_command passes these flags to the engine. Each one controls a runtime or serving behavior:

Deploy

Push the config to Baseten with the Baseten CLI, or with the Truss CLI if you prefer it:
You should see output similar to:
baseten model push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.

Call the model

Your deployment serves an OpenAI-compatible chat completions API at /v1/chat/completions that accepts audio inputs. Send audio as an audio_url content item on a chat message. The model returns the transcription as the assistant message content.
main.py

Next steps

Call your model

Endpoint anatomy, authentication, and sync versus async inference

Autoscaling

Scale replicas with traffic, including scale to zero