Pipette: Open-Source Benchmarking for On-Device Models
Melvin Vivas · X post · 2026-08-24 · Open on X
Topics: Evaluation (Evals) & Testing, Industry Trends & Job Market · Level: intermediate
Summary
The creator points to Pipette, a newly released open-source suite for evaluating on-device models, built in partnership with Artificial Analysis. Most benchmarks measure models served in the cloud. Pipette instead targets how capable and how fast models are when they run on-device.
Key points
- Pipette is an open-source model evaluation suite for on-device intelligence
- It was released in partnership with Artificial Analysis
- Most benchmarking platforms measure the capabilities and speed of foundation models served in the cloud
- Pipette fills the gap by benchmarking models that run locally on devices
Resources mentioned
- Pipette · tool · liquid.ai · free
An open-source suite for evaluating the capability and speed of on-device models. - Artificial Analysis · website · artificialanalysis.ai · free
Independent benchmarking site that compares AI models and providers, including text-to-speech, on quality, speed and price.
Also in: ElevenLabs Eleven v4 & v4 Turbo: Emotive, Multilingual Text-to-Speech Models (Melvin Vivas on X · notes), Be Skeptical of Model Leaderboards: Muse Spark vs Astra (Melvin Vivas on X · notes), GLM 5.2 Hits 446 tok/s on Fireworks AI (Melvin Vivas on X · notes), GLM-5.2 on Fireworks: Top Open-Weights Model on GDPval-AA (Melvin Vivas on X · notes) and 1 more
More in Evaluation (Evals) & Testing
- Is OpenAI Evals Still Maintained?
- Running LLM evals with promptfoo on local llama.cpp models
- The 8 Layers of Evaluating Production RAG and Agent Systems
- Idea: A Benchmark for Codex Usage-Limit Consumption
- Why Evaluation Is the Skill That Sets Top AI Engineers Apart
- AI-Written Code Still Needs Human Testing