Running MiniMax H3 Locally on a Single RTX 3090 (Demo)
Melvin Vivas · X video post · 2026-08-06 · 0:15 · 514 views · Open on X
Topics: Industry Trends & Job Market, LLMOps, Deployment & Monitoring · Level: beginner
Summary
This is a 15-second demo clip with no speech. Melvin Vivas shows output he generated locally with a model he calls "Minimax H3" on one consumer NVIDIA RTX 3090 GPU, and says the model is "promising". It's a data point that this model can run on consumer hardware. It doesn't explain the setup, settings or performance numbers.
Key points
- The model, which the creator calls "Minimax H3", was run locally instead of through a cloud API.
- The hardware was one NVIDIA RTX 3090, a consumer GPU with 24 GB of VRAM.
- The creator's verdict is only "This model is promising!" No benchmarks or quality metrics are given.
- The post doesn't say which inference stack, quantization or settings were used, or how long generation took.
- The clip has no speech, so the only information is the caption and the visual demo.
Resources mentioned
- MiniMax H3 · tool · huggingface.co · free
MiniMax video generation model that turns a starting image and an audio track into a lip-synced video.
Also in: MiniMax H3 Image + Audio-to-Video Lip-Sync Demo (with Irodori-TTS v3) (Melvin Vivas on X · notes), Generating AI Video Locally with MiniMax H3 in ComfyUI on an RTX 3090 (Melvin Vivas on X · notes) - NVIDIA GeForce RTX 3090 · tool · nvidia.com · paid
Consumer GPU with 24 GB of VRAM, used here to run a long-context LLM locally.
Also in: Portable Computer Now Runs AI Agents and Models Locally on NVIDIA RTX PCs (Melvin Vivas on X · notes), Long-Context Local LLM on an RTX 3090: .env Config and ~64 tok/s Benchmark (Melvin Vivas on X · notes), Running Ornith-1.5-35B-A3B (Q4_K_M) at 128k Context on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Generating AI Video Locally with MiniMax H3 in ComfyUI on an RTX 3090 (Melvin Vivas on X · notes) and 2 more - Melvin Vivas · person · x.com · free
A creator who shares experiments with AI agents and developer tools on X.
Also in: How LLMs Work: A Motion-Graphics Explainer Made in One Shot with Claude Opus 5.5 (Melvin Vivas on X · notes), Concept Demo: An Agent Monitor Built on OpenAI Codex Traces (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent (Melvin Vivas on X · notes)
Try this
- Try running the Minimax H3 model locally on a 24 GB consumer GPU and write down VRAM use, generation time and output quality.
More in Industry Trends & Job Market
- Grok 4.6 release: better than Grok 4.5 at the same price
- Meta Muse Glimmer-30B: Open Weights Model That Beats Qwen3.6 37B on Agentic Tasks
- Liquid AI Documentation and the Future of Local AI
- Qwen3.8-27B announced to run locally on 17GB RAM/VRAM
- Qwen3.8-Max announced with open weights for Max and 27B
- Grok Can Now Analyze Videos