OpenRouter Pareto Code: cost-optimized coding router
Melvin Vivas · X post · 2026-05-10 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Dev Tools & Productivity, LLM Fundamentals · Level: intermediate
Summary
OpenRouter launched Pareto Code, a free, experimental router for coding requests. You set min_coding_score in your request, and it sends the request to the cheapest code-capable model that meets that score, using Artificial Analysis rankings. The creator calls it a game-changer for cutting coding costs.
Key points
- Pareto Code is a free, experimental coding router from OpenRouter.
- Set the
min_coding_scoreparameter in your request to define your quality bar. - It routes to the cheapest code-capable model that meets that bar.
- Models are ranked using Artificial Analysis benchmarks.
- The Pareto frontier of cost vs. quality updates in real time as new models arrive.
Resources mentioned
- OpenRouter · tool · openrouter.ai · free
A single OpenAI-compatible API that routes requests to hundreds of models from many providers, with one bill and model fallbacks.
Also in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), The Jev model is now on OpenRouter (Melvin Vivas on X · notes), Kev-4B model, an alternative to Jev, now available on OpenRouter (Melvin Vivas on X · notes), Space Bunny Alpha: stealth 1M-context flash model on OpenRouter (Melvin Vivas on X · notes) and 42 more - OpenRouter Pareto Code · tool · openrouter.ai · free
Free, experimental router that picks the cheapest coding model meeting your min_coding_score. - Artificial Analysis · website · artificialanalysis.ai · free
Independent benchmarking site that compares AI models and providers, including text-to-speech, on quality, speed and price.
Also in: ElevenLabs Eleven v4 & v4 Turbo: Emotive, Multilingual Text-to-Speech Models (Melvin Vivas on X · notes), Be Skeptical of Model Leaderboards: Muse Spark vs Astra (Melvin Vivas on X · notes), Pipette: Open-Source Benchmarking for On-Device Models (Melvin Vivas on X · notes), GLM 5.2 Hits 446 tok/s on Fireworks AI (Melvin Vivas on X · notes) and 1 more
Try this
- Try setting min_coding_score in OpenRouter requests to lower coding costs.
More in LLMOps, Deployment & Monitoring
- Unsloth MTP GGUFs make Qwen3.6 run 1.4x faster locally
- Cheap TTS serving: Qwen3-TTS on vLLM-Omni at $3 per 1M characters
- LM Studio MLX v1.8.1: Vision Model Batching and Better Caching
- Speeding Up Gemma 4 Inference: MTP (3x) vs. DFlash Speculative Decoding (6x)
- Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code
- Runpod Flash Reaches GA: Deploy AI Workloads from Python