Red Hat's Quantized Qwen3.8-2.4T-A95B Variant (NVFP4/FP8)
Melvin Vivas · X post · 2026-08-20 · Open on X
Topics: LLMOps, Deployment & Monitoring, Industry Trends & Job Market · Level: advanced
Summary
Red Hat AI published a quantized version of Qwen3.8 on Hugging Face called Qwen3.8-2.4T-A95B-NVFP4-FP8. The name indicates a 2.4T-parameter MoE model with about 95B active parameters, quantized to NVFP4/FP8 for more efficient serving.
Key points
- Red Hat AI released its own quantized Qwen3.8 variant.
- Model name: Qwen3.8-2.4T-A95B-NVFP4-FP8.
- Naming: 2.4T total parameters, A95B = about 95B active (MoE), NVFP4/FP8 = quantization formats.
- The model card is on Hugging Face under RedHatAI.
Resources mentioned
- RedHatAI/Qwen3.8-2.4T-A95B-NVFP4-FP8 · Hugging Face · tool · huggingface.co · free
Hugging Face model card for Red Hat AI's NVFP4/FP8 quantized Qwen3.8 2.4T-A95B model.
Try this
- Read the model card to see the quantization details and serving requirements.
More in LLMOps, Deployment & Monitoring
- Serving Ornith-1.5-35B-A3B at 128k Context on an RTX 3090 with llama.cpp
- Running Ornith-1.5-35B-A3B (Q4_K_M) at 128k Context on an RTX 3090 with llama.cpp
- Running DeepSeek V4 Pro 0813 on Baseten with the Pi Agent
- One-Click Deploy Link for Qwen3.8 27B on Baseten
- Self-Hosting Qwen3.8 27B FP8 on an H100 for $6.50/hr
- Hosted Qwen3.8 27B Costs More Than GPT 5.6 Luna, So Run It Locally