Serve GLM-5.2 NVFP4 with vLLM on NVIDIA Blackwell
Melvin Vivas · X post · 2026-06-27 · Open on X
Topics: LLMOps, Deployment & Monitoring, Fine-tuning & Model Customization · Level: advanced
Summary
vLLM now supports NVIDIA's official NVFP4-quantized checkpoint of GLM-5.2. On Blackwell GPUs it uses less memory than FP8 while matching FP8 accuracy on reasoning, coding and long-context benchmarks. You can serve it with one command.
Key points
- NVIDIA released an official NVFP4 (4-bit floating point) checkpoint of GLM-5.2.
- It targets Blackwell GPUs and uses less memory than FP8.
- It is reported to match FP8 accuracy on reasoning, coding and long-context benchmarks.
- Serve command:
vllm serve nvidia/GLM-5.2-NVFP4.
Resources mentioned
- vLLM · repo · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
Open-source, high-throughput LLM inference and serving engine with continuous batching and efficient KV-cache management (PagedAttention).
Also in: Ollama vs vLLM: From Local AI Demo to Production Inference Serving (Bashiri Smith on Facebook · notes), Gemma 4 Gets Up to 3x Faster with MTP Drafters (Melvin Vivas on X · notes) - nvidia/GLM-5.2-NVFP4 · tool · huggingface.co · free
NVIDIA's NVFP4-quantized release of the open GLM-5.2 model, with its model card on Hugging Face.
Also in: Model card: NVIDIA's NVFP4-quantized GLM-5.2 on Hugging Face (Melvin Vivas on X · notes)
Try this
- Serve the model with `vllm serve nvidia/GLM-5.2-NVFP4` on Blackwell hardware.
More in LLMOps, Deployment & Monitoring
- Frontier Model Costs: Why Top Models Can Sink a Solo Dev's MRR
- GLM 5.2 Hits 446 tok/s on Fireworks AI
- Fastest GLM-5.2 Provider: Fireworks AI at 343 tok/s
- Self-Hosting GLM 5.2 with Modal Auto Endpoints
- Where to Access GLM 5.2: Inference Providers and Gateways
- Serving GLM-5.2 on Baseten: >280 TPS and <0.8s TTFT