Qwen3.8 Model Family Now on Novita AI: Flash, 27B and 2.4T-A95B
Melvin Vivas · X video post · 2026-08-28 · 0:25 · 348 views · Open on X
Topics: LLM Fundamentals, Industry Trends & Job Market · Level: intermediate
Summary
Melvin Vivas shares a quoted post saying three Qwen3.8 models (Qwen3.8-Flash, Qwen3.8-27B and Qwen3.8-2.4T-A95B) are now available on the Novita AI inference platform. The models are aimed at coding, agentic workflows, research and long-context tasks. The post is a model-release announcement, not a tutorial. Its main use is knowing which new model options exist and how their architectures differ.
Key points
- Three Qwen3.8 models are now available on Novita AI: Qwen3.8-Flash, Qwen3.8-27B and Qwen3.8-2.4T-A95B.
- Stated target uses: coding, agentic workflows, research and long-context tasks.
- Qwen3.8-Flash is a 125B-parameter Mixture-of-Experts (MoE) model with only 6B active parameters, and it accepts text, image and video input.
- Qwen3.8-27B is a dense vision-language model. The rest of its description is cut off in the caption.
- Going by the usual Qwen naming, the name Qwen3.8-2.4T-A95B suggests an MoE model with about 2.4T total parameters and about 95B active per token. The post itself doesn't spell this out.
- MoE vs. dense: an MoE model with few active parameters (like Flash) can be cheaper and faster per token than its total size suggests. A dense model uses all of its parameters for every token.
Resources mentioned
- Qwen3.8-Flash · tool · qwen.ai · check price
A 125B-parameter MoE model with 6B active parameters that accepts text, image and video input. - Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more - Qwen3.8-2.4T-A95B · tool · huggingface.co · free
The largest Qwen3.8 model; by its name, about 2.4T total parameters with about 95B active. - Novita · tool · novita.ai · check price
A cloud platform where you can call open models like the Qwen3.8 family through an API.
Also in: Free GLM-5.2 via Hugging Face Inference Providers in coding agents (Melvin Vivas on X · notes)
More in LLM Fundamentals
- Finding Models to Run Locally on Hugging Face
- Free Nemotron 3.5 Lightning on OpenRouter Supports Thinking
- Creator's Top 3 Closed Models: Fable 5, GPT 5.6 Sol, Grok 4.6
- Qwen3.8-27B Free on Groq at ~450 Tokens per Second
- Colab Notebook: Entity Extraction with GLiNER2.5 in aibackends
- Cohere Parse Beats Frontier LLMs at Receipt Parsing