AuK: Open-Source Model for Generating and Editing Speech
Melvin Vivas · X video post · 2026-09-18 · 1:53 · 277 views · Open on X
Topics: Industry Trends & Job Market, AI Dev Tools & Productivity · Level: beginner
Summary
Melvin Vivas shares a demo of AuK, a new open-source foundation model that both generates and edits speech. It takes natural-language instructions plus reference audio through a single interface, and the quoted post calls it "Nano Banana for audio." The video is a reel of audio samples with no spoken explanation. It is mostly a model announcement, not a tutorial.
Key points
- AuK is described as an open-source foundation model for both speech generation and speech editing.
- Input: natural-language instructions plus a reference audio clip, all through one interface.
- Zero-shot TTS: it clones a voice from reference audio with no training on that speaker.
- Instruction-controlled generation: you control style or delivery with text instructions.
- Content editing: you can change words in existing speech. The demo plays original and edited clips in pairs (0:39-1:15).
- Whisper conversion: it turns normal speech into whispered speech. This is a voice style, not OpenAI's Whisper model.
- The quoted post calls it 'Nano Banana for audio', meaning instruction-based editing like Google's image model, but for speech.
Resources mentioned
- AuK · tool · github.com · free
Open-source foundation model for unified speech generation and editing (zero-shot TTS, instruction control, content editing, whisper conversion).
Also in: AuK: Tencent Hunyuan's Unified Audio Generation and Editing Model (TTS Demo) (Melvin Vivas on X · notes) - Nano Banana (Gemini 2.5 Flash Image) · tool · developers.googleblog.com · free
Google's instruction-based image generation and editing model, used here as a comparison point for AuK.
More in Industry Trends & Job Market
- Qwen-Image-2.1 Announced as an Open-Weights Image Model
- Qwen3.8-LiveTranslate: real-time simultaneous interpretation model
- Open models now dominate token volume on Vercel AI Gateway
- Qwen3.8-Omni-Flash: Qwen's agentic omni-modal model
- TypeSafe AI's Jev and System One models: the launch article
- Liquid AI's LFM2-Longevity models for aging-data analysis