Local Audio Transcription with Cohere Transcribe on WebGPU
Melvin Vivas · X post · 2026-04-04 · Open on X
Topics: LLM Fundamentals, LLMOps, Deployment & Monitoring · Level: beginner
Summary
The creator shows Cohere's speech-to-text model running locally in the browser through WebGPU. He links the model card on Hugging Face and a Hugging Face Space where anyone can try it.
Key points
- Cohere released a speech-to-text model: cohere-transcribe-03-2026.
- With WebGPU, it runs entirely in the browser on your own machine, with no server.
- A Hugging Face Space lets you try it with no setup.
- Running transcription in the browser keeps your audio private.
Resources mentioned
- CohereLabs/cohere-transcribe-03-2026 · tool · huggingface.co · free
Model card on Hugging Face for Cohere's audio transcription (speech-to-text) model. - Cohere Transcribe WebGPU (Hugging Face Space) · tool · huggingface.co · free
Browser demo by CohereLabs that runs the transcription model locally using WebGPU. - Cohere · person · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
AI company that builds language, embedding and reranking models, and now the Transcribe speech-recognition model.
Also in: Basic RAG Pipeline in 60 Seconds: From Documents to Grounded Answers (Bashiri Smith on Facebook · notes), Cohere Parse Beats Frontier LLMs at Receipt Parsing (Melvin Vivas on X · notes), Cohere Parse: Pricing vs Parse Bench Score (Melvin Vivas on X · notes), Cohere's North Micro Vision: a small open-source vision model for documents (Melvin Vivas on X · notes) and 1 more - Hugging Face · website · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
Platform for hosting and finding ML models, datasets and papers. The quoted post says LocateAnything was trending there.
Also in: Using an ML agent to train an open-source TTS model on your voice (Melvin Vivas on X · notes), Deploy Open-Source Models with Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes) and 38 more
Try this
- Open the Cohere Transcribe WebGPU Space and transcribe some audio locally in your browser.
- Build a private, in-browser transcription app on WebGPU using Cohere Transcribe.
More in LLM Fundamentals
- Gemma 4 E4B: Local Image Understanding at 131K Context in 6GB VRAM
- A Visual Guide to Gemma 4 Architecture (Maarten Grootendorst)
- Gemma 4 Architecture: MoE, Encoders and Per-Layer Embeddings
- Gemma 4 26B for OCR in LM Studio
- Gemma 4: Google's Apache 2.0 Open-Weight Models for Local Hardware
- Qwen 3.6 Plus Preview Is Free for a Limited Time on OpenRouter