Self-Hosting Qwen3.8 27B FP8 on an H100 for $6.50/hr
Melvin Vivas · X post · 2026-08-19 · Open on X
Topics: LLMOps, Deployment & Monitoring · Level: intermediate
Summary
You can run your own copy of Qwen3.8 27B in FP8 on Baseten. It needs a single H100 instance, which costs about $6.50 per hour, a useful number for estimating self-hosting costs.
Key points
- Qwen3.8 27B quantized to FP8 fits on one H100 GPU.
- An H100 instance on Baseten costs about $6.50/hr.
- FP8 quantization lowers memory needs, so a 27B model can run on a single GPU.
Resources mentioned
- Baseten · tool · x.com · paid
Model inference platform offering dedicated GPU deployments for serving open models.
Also in: Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference (Melvin Vivas on X · notes), Coworker Desktop App Works with Any OpenAI-Compatible Endpoint (Melvin Vivas on X · notes), Running GLM 5.3 Flash on Baseten with the Pi Coding Agent (Melvin Vivas on X · notes), GLM 5.3 Flash Now Available on Baseten (Melvin Vivas on X · notes) and 6 more - Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more
More in LLMOps, Deployment & Monitoring
- Running DeepSeek V4 Pro 0813 on Baseten with the Pi Agent
- Red Hat's Quantized Qwen3.8-2.4T-A95B Variant (NVFP4/FP8)
- One-Click Deploy Link for Qwen3.8 27B on Baseten
- Hosted Qwen3.8 27B Costs More Than GPT 5.6 Luna, So Run It Locally
- Running Qwen3.8 27B Locally with Hermes
- AIBackends: Python Library for Local-Model AI Workflows