Setup
Install the Baseten CLI and sign in, then install the OpenAI SDK.Install and sign in to BasetenFor other platforms or a specific version, see the Baseten CLI install reference.
- macOS or Linux
- Windows
Terminal
Install the OpenAI SDK
uvx truss login --browser and deploy with uvx truss push.
This preset serves Mellum2 12B A2.5B Instruct on a single H100 through vLLM, optimized for low-latency code generation.
Hardware
H100
Engine
vLLM 0.23.0
Context
128K
Concurrency
128
Write the config
Create and move into the project directory:config.yaml and paste the following:
config.yaml
vllm/vllm-openai:v0.23.0 image, the release that adds MellumForCausalLM support, and streams weights from JetBrains/Mellum2-12B-A2.5B-Instruct with the Run:ai streamer. The Hermes tool-call parser enables OpenAI-compatible function calling, and a concurrency ceiling of 128 keeps the deployment throughput-friendly for coding assistants.
Flags
Thestart_command passes these flags to the engine. Each one controls a runtime or serving behavior:
Deploy
Push the config to Baseten with the Baseten CLI, or with the Truss CLI if you prefer it:baseten model push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.
Call the model
Your deployment serves an OpenAI-compatible API. Now call your deployment to run inference:- Python
- cURL
main.py
tools array. The server returns structured tool_calls on the response:
Next steps
Call your model
Endpoint anatomy, authentication, and sync versus async inference
Autoscaling
Scale replicas with traffic, including scale to zero