Setup
Install the Baseten CLI and sign in, then install the OpenAI SDK.Install and sign in to BasetenFor other platforms or a specific version, see the Baseten CLI install reference.
- macOS or Linux
- Windows
Terminal
Install the OpenAI SDK
uvx truss login --browser and deploy with uvx truss push.
This preset serves GLM-5.3 Flash on H100:8 from its native FP8 checkpoint, with a multi-token prediction head proposing five speculative tokens per step to speed up decoding.
Hardware
H100 × 8
Engine
vLLM (glm53-flash build)
Context
1M
Concurrency
16
Write the config
Create and move into the project directory:config.yaml and paste the following:
config.yaml
/app/checkpoint/model and serves the OpenAI-compatible API on port 8000 with tensor parallel size 8 across the eight H100 GPUs. The engine caps each request at 1,048,576 tokens and runs 16 concurrent sequences, keeping weights in FP8 while the KV cache stays in BF16, because the current implementation does not support an FP8 KV cache on Hopper. Requests that omit reasoning_effort inherit the server default of high, and each prompt accepts one image and one video alongside its text.
Flags
Thestart_command passes these flags to the engine. Each one controls a runtime or serving behavior:
Deploy
Push the config to Baseten with the Baseten CLI, or with the Truss CLI if you prefer it:baseten model push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.
Call the model
Your deployment serves an OpenAI-compatible API. Now call your deployment to run inference:- Python
- cURL
main.py
reasoning_content field on the response. Read it alongside the final answer:
tools array. The server returns structured tool_calls on the response:
Next steps
Call your model
Endpoint anatomy, authentication, and sync versus async inference
Autoscaling
Scale replicas with traffic, including scale to zero