Jev + WebMCP Solves 100% of Benchmark Tasks at About 112× Lower Cost
Melvin Vivas · X post · 2026-09-18 · Open on X
Topics: AI Agents, Tool Use & MCP, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
A quoted benchmark claims that Jev with Mercury 2.5, a fast and low-cost LLM, used WebMCP to solve 100% of tasks at about 112× lower model cost than GPT-6 Astra using computer use with code execution. The point is that structured tool interfaces like WebMCP can let cheap models beat expensive computer-use agents.
Key points
- WebMCP exposes website actions as structured tools instead of making the agent operate the screen.
- Jev + Mercury 2.5 + WebMCP reportedly solved 100% of the benchmark tasks.
- Model cost was about 112× lower than GPT-6 Astra using computer use with code execution.
- Lesson: a good tool interface can matter more than model size for web agents.
Resources mentioned
- WebMCP · tool · developer.chrome.com · free
A proposed web standard that lets websites expose tools and context to AI agents through the Model Context Protocol in the browser.
Also in: WebMCP: Exposing Existing App Features as Agent Tools (Melvin Vivas on X · notes), Docs7: Serving Agent-Readable Documentation to Coding Agents via MCP (Melvin Vivas on X · notes) - Jev · tool · typesafe.ai · paid
A decision model that makes turn-level forecasts on AI agent conversations using only structural signals (turns, tool calls, workflow stages, timing).
Also in: Generating a Repo Promo Video with a Claude Skill on Sonnet 5.5 vs Opus 5.5 (Melvin Vivas on X · notes), GLiDE by Fastino Labs: A Post-Trainable Reasoning Decision Model (Melvin Vivas on X · notes), Using Jev as a Reranker to Augment RAG Retrieval (Melvin Vivas on X · notes), Using Jev (TypeSafe AI) as a Decision Model for Email Classification (Melvin Vivas on X · notes) and 12 more - Mercury 2.5 · tool · inceptionlabs.ai · paid
A very fast language model, reported at about 1,100 tokens/sec with better agentic performance than Mercury 2.
Also in: Mercury 2.5 runs at 1,100 tokens/sec (Melvin Vivas on X · notes) - GPT-6 Astra · tool · openai.com · paid
The model announced in the quoted launch post, pitched as the developer's most capable model for work, coding, science and cybersecurity, and able to operate a computer.
Also in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), Pi Agent Council: Ask Multiple LLMs in Parallel and Compare Their Advice (Melvin Vivas on X · notes), Use GPT-6.1 Sol by Default, Save Astra for Emergencies (Melvin Vivas on X · notes), Dots in ChatGPT: always-on AI agents that you hand responsibilities to (Melvin Vivas on X · notes) and 49 more
Try this
- Compare a cheap LLM using structured WebMCP tools against a frontier model using computer use on the same web tasks, measuring both success rate and cost.
More in AI Agents, Tool Use & MCP
- Coworker: Open-Source Desktop App for AI Agents
- Melvin Vivas's refreshed website: AI agents, local models, workflows
- Running Gemma4-E2B tool calling locally on an iPhone
- Agent Monitor: Visualize Codex and Claude Code Subagents, Traces and Costs
- Intent Classification and Routing in a Multi-Agent Desktop App
- Coworker: Open-Source Desktop App for AI Agents, with Jev