Qwen3.8-27B Free on Groq at ~450 Tokens per Second
Melvin Vivas · X post · 2026-08-28 · Open on X
Topics: LLM Fundamentals, LLMOps, Deployment & Monitoring · Level: beginner
Summary
Qwen3.8-27B is available for free on GroqCloud and runs at about 450 tokens per second. The post points out that Groq (the inference provider) is not the same as Grok (xAI's model). You can access it through console.groq.com with the model ID qwen/qwen3.8-27b.
Key points
- Qwen3.8-27B is free on Groq.
- Speed is about 450 tokens per second.
- Groq (fast inference provider) is not Grok (xAI's model).
- Model ID to use: qwen/qwen3.8-27b.
- Get access through the GroqCloud console at console.groq.com.
Resources mentioned
- GroqCloud Console · tool · console.groq.com · free
Groq's developer console for fast LLM inference APIs. - Groq (@GroqLLC) on X · person · x.com · free
Groq's official X account. - Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more
Try this
- Try qwen/qwen3.8-27b for free in the GroqCloud console.
More in LLM Fundamentals
- Free Nemotron 3.5 Lightning on OpenRouter Supports Thinking
- Creator's Top 3 Closed Models: Fable 5, GPT 5.6 Sol, Grok 4.6
- Qwen3.8 Model Family Now on Novita AI: Flash, 27B and 2.4T-A95B
- Colab Notebook: Entity Extraction with GLiNER2.5 in aibackends
- Cohere Parse Beats Frontier LLMs at Receipt Parsing
- You don't need frontier LLMs for everything: use SLMs