Qwen 3.8 27B Matches GPT 5.6 Luna (max) on Artificial Analysis
Melvin Vivas · X post · 2026-08-18 · Open on X
Topics: LLM Fundamentals, Industry Trends & Job Market · Level: beginner
Summary
On Artificial Analysis's intelligence comparison, Qwen 3.8 27B, a fairly small open-weight model, scores about the same as GPT 5.6 Luna at max reasoning. The post shows how to use independent benchmark sites to compare models.
Key points
- Qwen 3.8 27B scores on par with GPT 5.6 Luna (max) on Artificial Analysis's intelligence comparison.
- A 27B open-weight model now competes with a frontier proprietary model on this index.
- Artificial Analysis compares models on intelligence, performance (speed) and price.
- Check independent leaderboards like this one when choosing a model, rather than relying only on vendor claims.
Resources mentioned
- Artificial Analysis – Comparison of AI Models across Intelligence, Performance, and Price · website · artificialanalysis.ai · free
Benchmark site comparing AI models on an intelligence index, speed and price, here filtered to the GPT models in Codex.
Also in: Comparing GPT Models in Codex by Intelligence and Cost per Task (Melvin Vivas on X · notes) - Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more - GPT 5.6 Luna · tool · openai.com · paid
The OpenAI model used inside Codex for the demo. The transcript gives the variant name as 'Soul', which is unclear.
Also in: Set Codex subagent model and reasoning to save usage limits (Melvin Vivas on X · notes), Use GPT-5.6 Luna in Codex for Terminal Tasks (Melvin Vivas on X · notes), Match Reasoning Effort to Task Length in Codex (Astra/Sol) (Melvin Vivas on X · notes), Run Coworker desktop agents cheaply with GPT-5.6 Luna on OpenRouter (Melvin Vivas on X · notes) and 24 more
Try this
- Use Artificial Analysis to compare models on intelligence, speed and price before choosing one.
More in LLM Fundamentals
- Running Hermes on DeepSeek V4 Pro via Baseten Model APIs
- Find Discounted Models with OpenRouter's New Filter
- Cost-Aware Model Routing in Codex: Default to Luna, Escalate to Sol
- Run LFM2.5-2.6B locally with llama-cpp-python in Colab
- GLM-5.3 free weekend trial announcement
- Swapping AI SDK for pi-ai as the LLM provider layer