Serving GLM-5.2 on Baseten: >280 TPS and <0.8s TTFT
Melvin Vivas · X video post · 2026-06-23 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
Baseten hosts the open model GLM-5.2 and claims more than 280 tokens per second with under 0.8s time to first token. The creator points it out as a hosting option for GLM 5.2.
Key points
- GLM-5.2 is available on Baseten's model library
- Claimed throughput: more than 280 tokens per second
- Claimed latency: under 0.8s time to first token (TTFT)
- TPS and TTFT are the key metrics when choosing an inference provider
Resources mentioned
- GLM-5.2 | Model library (Baseten) · website · baseten.co · check price
Baseten model library page for trying and deploying GLM-5.2, including its vision support.
Also in: GLM-5.2 Vision on Baseten: Turning Images into Code (Melvin Vivas on X · notes)
Try this
- Try GLM-5.2 on Baseten
More in LLMOps, Deployment & Monitoring
- Serve GLM-5.2 NVFP4 with vLLM on NVIDIA Blackwell
- Self-Hosting GLM 5.2 with Modal Auto Endpoints
- Where to Access GLM 5.2: Inference Providers and Gateways
- 2.3x Faster Ideogram 4 in ComfyUI with INT8 and SageAttention
- Model routing: Rayline picks the best model per task
- Running GLM-5.2 locally on a 256GB Mac with Unsloth