Cohere's North Micro Vision: a small open-source vision model for documents
Melvin Vivas · X post · 2026-08-13 · Open on X
Topics: LLM Fundamentals, Industry Trends & Job Market · Level: beginner
Summary
Cohere released North Micro Vision, its smallest vision-language model so far. It is built for document understanding and is open-source under Apache 2.0. The weights are on Hugging Face, so you can try it locally.
Key points
- North Micro Vision is Cohere's smallest vision-language model (VLM) so far.
- It is aimed at document understanding tasks.
- It is released open-source under the Apache 2.0 license, which allows commercial use.
- The model weights can be downloaded from Hugging Face.
Resources mentioned
- Cohere · person · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
AI company that builds language, embedding and reranking models, and now the Transcribe speech-recognition model.
Also in: Basic RAG Pipeline in 60 Seconds: From Documents to Grounded Answers (Bashiri Smith on Facebook · notes), Cohere Parse Beats Frontier LLMs at Receipt Parsing (Melvin Vivas on X · notes), Cohere Parse: Pricing vs Parse Bench Score (Melvin Vivas on X · notes), Local Audio Transcription with Cohere Transcribe on WebGPU (Melvin Vivas on X · notes) and 1 more - North Micro Vision · tool · huggingface.co · free
Cohere's smallest open-source vision-language model, built for document understanding. - Hugging Face · website · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
Platform for hosting and finding ML models, datasets and papers. The quoted post says LocateAnything was trending there.
Also in: Using an ML agent to train an open-source TTS model on your voice (Melvin Vivas on X · notes), Deploy Open-Source Models with Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes) and 38 more
Try this
- Download the North Micro Vision weights from Hugging Face and try them on document-understanding tasks.
More in LLM Fundamentals
- Swapping AI SDK for pi-ai as the LLM provider layer
- Open Models Aren't Always Local Models
- Using LFM2.5-VL-3B's Vision Capabilities (Liquid AI Guide)
- Quick test of Liquid AI's LFM2.5-VL-3B vision model
- LFM2.5-VL-3B: a lightweight VLM for screens, documents and tool calls
- NVIDIA Nemotron 3.5 Lightning: Fast Open MoE Model for Agents