Skip to main content
JetBrains’ Mellum2 open MoE code model (12B total, 2.5B active) with a 131k-token context window and tool calling.

Setup

Install the Baseten CLI and sign in, then install the OpenAI SDK.
Install and sign in to Baseten
Terminal
For other platforms or a specific version, see the Baseten CLI install reference.
Install the OpenAI SDK
Prefer not to install? Sign in with uvx truss login --browser and deploy with uvx truss push. This preset serves Mellum2 12B A2.5B Instruct on a single H100 through vLLM, optimized for low-latency code generation.

Hardware

H100

Engine

vLLM 0.23.0

Context

128K

Concurrency

128

Write the config

Create and move into the project directory:
Then create a file named config.yaml and paste the following:
config.yaml
This config runs the official vllm/vllm-openai:v0.23.0 image, the release that adds MellumForCausalLM support, and streams weights from JetBrains/Mellum2-12B-A2.5B-Instruct with the Run:ai streamer. The Hermes tool-call parser enables OpenAI-compatible function calling, and a concurrency ceiling of 128 keeps the deployment throughput-friendly for coding assistants.

Flags

The start_command passes these flags to the engine. Each one controls a runtime or serving behavior:

Deploy

Push the config to Baseten with the Baseten CLI, or with the Truss CLI if you prefer it:
You should see output similar to:
baseten model push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.

Call the model

Your deployment serves an OpenAI-compatible API. Now call your deployment to run inference:
main.py
To let the model call tools, pass a tools array. The server returns structured tool_calls on the response:

Next steps

Call your model

Endpoint anatomy, authentication, and sync versus async inference

Autoscaling

Scale replicas with traffic, including scale to zero