How Anthropic's data team automated 95% of analytics queries with Claude
Melvin Vivas · X post · 2026-06-04 · Open on X
Topics: Evaluation (Evals) & Testing, AI Agents, Tool Use & MCP, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
The creator shares Anthropic's announcement that its data team automated 95% of business analytics queries with Claude, calling it the way to automate. The linked blog post explains how they used evals, ablations and online validation to make the automation reliable.
Key points
- Anthropic's data team reports that Claude automates 95% of its business analytics queries.
- The blog post covers their evaluation approach (evals) for the analytics agent.
- They used ablations to work out which parts of the system add value.
- Online validation was used to check quality in production after offline evals.
- Pattern: pair agent automation with offline evals plus online monitoring before trusting it.
Resources mentioned
- Anthropic blog post: automating business analytics queries with Claude · article · claude.com · free
How Anthropic's data team automated 95% of analytics queries with Claude, including evals, ablations and online validation. - Claude Opus 5.5 · tool · claude.ai · paid · open in a browser to verify · recommended by both Bashiri Smith & Melvin Vivas
Anthropic's AI assistant, used throughout the guide to tailor resumes, add live roles to the tracker, match connections to target companies and find hiring managers.
Also in: Generating a Repo Promo Video with a Claude Skill on Sonnet 5.5 vs Opus 5.5 (Melvin Vivas on X · notes), Comparing Coding Models on the Same Task in Devin iOS (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), AI Engineer Roadmap Overview: From ML Foundations to RAG, Agents & Ops (Bashiri Smith on Facebook · notes) and 53 more - ClaudeDevs (@ClaudeDevs) on X · person · x.com · free
Anthropic's developer account on X, which shared the analytics automation post.
Also in: Claude Code Agents Can Now Talk to Each Other Natively (Melvin Vivas on X · notes), Follow @ClaudeDevs for Claude Developer Updates (Melvin Vivas on X · notes)
Try this
- Read Anthropic's blog post on how they evaluate and validate their analytics automation.
- Build a text-to-SQL analytics agent with Claude and measure it with an eval set, ablations and online validation.
More in Evaluation (Evals) & Testing
- OCR Testing a Small Vision Model with LLM-Made Ground Truth
- Code Arena Fullstack Benchmark: Kimi K3 Ranks #1
- Test AI-Generated Apps: They Are Not Bug-Free
- End-to-end test harnesses are key to automated AI coding
- Read Benchmark Numbers, Not Highlights: the Muse Spark Lesson
- Gemma 4 on the Arena Leaderboard