OCR Testing a Small Vision Model with LLM-Made Ground Truth
Melvin Vivas · X post · 2026-08-16 · Open on X
Topics: Evaluation (Evals) & Testing, LLM Fundamentals · Level: intermediate
Summary
Melvin Vivas tests the OCR quality of Liquid AI's small LFM2.5-VL-3B vision-language model. He has a stronger model (GPT 5.6 Luna) write the ground-truth transcriptions, then has Codex check the small model's outputs against them. This is a cheap way to evaluate a small local model when you don't have hand-labeled data.
Key points
- Model under test: LFM2.5-VL-3B, a 3B-parameter vision-language model, used for OCR.
- Ground truth comes from a stronger frontier model (GPT 5.6 Luna) instead of manual labeling.
- A coding agent (Codex) compares the small model's outputs with the ground truth.
- Pattern: strong model labels, small model predicts, agent or LLM checks the results.
- Ground truth made by an LLM can contain errors, so spot-check a sample by hand.
Resources mentioned
- LFM2.5-VL-3B · tool · huggingface.co · free
Liquid AI's lightweight vision-language model for screen and document understanding, grounding and tool calling.
Also in: AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers (Melvin Vivas on X · notes), AIBackends Adds Support for LFM2.5-VL-3B (Melvin Vivas on X · notes), Using LFM2.5-VL-3B's Vision Capabilities (Liquid AI Guide) (Melvin Vivas on X · notes), Quick test of Liquid AI's LFM2.5-VL-3B vision model (Melvin Vivas on X · notes) and 1 more - GPT 5.6 Luna · tool · openai.com · paid
The OpenAI model used inside Codex for the demo. The transcript gives the variant name as 'Soul', which is unclear.
Also in: Set Codex subagent model and reasoning to save usage limits (Melvin Vivas on X · notes), Use GPT-5.6 Luna in Codex for Terminal Tasks (Melvin Vivas on X · notes), Match Reasoning Effort to Task Length in Codex (Astra/Sol) (Melvin Vivas on X · notes), Run Coworker desktop agents cheaply with GPT-5.6 Luna on OpenRouter (Melvin Vivas on X · notes) and 24 more - OpenAI Codex · tool · openai.com · paid
OpenAI's coding agent. In the diagram it writes code, fixes review findings and drives the build loop. The creator also used it to make this video.
Also in: An agent bot that installs and drives Codex on its own (Melvin Vivas on X · notes), Asking a Coder bot to install Codex (Melvin Vivas on X · notes), Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Codex Cloud Environments Must Be Saved & Published Before Use (Melvin Vivas on X · notes) and 240 more
Try this
- Build an OCR evaluation harness: a frontier model creates the ground truth, a small local VLM makes predictions, and an agent or script scores the differences.
More in Evaluation (Evals) & Testing
- AI-Written Code Still Needs Human Testing
- LoopsBench: Testing Coding Agents on Long-Horizon Tasks
- Agent Loops Need Real Evals and Feedback
- Code Arena Fullstack Benchmark: Kimi K3 Ranks #1
- Test AI-Generated Apps: They Are Not Bug-Free
- How Anthropic's data team automated 95% of analytics queries with Claude