AI Engineer Study Library

Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference

Melvin Vivas · X post · 2026-08-29 · Open on X

Topics: LLMOps, Deployment & Monitoring, Evaluation (Evals) & Testing · Level: intermediate

Summary

The creator serves Qwen3.8 27B in FP8 on a dedicated NVIDIA H100 through Baseten's dedicated inference. He gives the price as $6.50 per hour ($0.10833 per minute), which is a useful reference point for what it costs to self-host a mid-size open model. He also offers to evaluate the model with other people's agents.

Key points

Resources mentioned

Try this

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring