2.3x Faster Ideogram 4 in ComfyUI with INT8 and SageAttention
Melvin Vivas · X post · 2026-06-20 · Open on X
Topics: LLMOps, Deployment & Monitoring · Level: advanced
Summary
A quoted benchmark shows how to speed up Ideogram 4.0 image generation in ComfyUI: switch from FP8 weights to INT8, and from PyTorch SDP attention to SageAttention. On an RTX 3090 this made generation about 2.3x faster, and it costs nothing.
Key points
- Drop PyTorch SDP attention and FP8 weights; use INT8 weights with SageAttention instead.
- Benchmark setup: RTX 3090, same JSON prompt, 48 steps.
- INT8 + SageAttention: 78s.
- INT8 + PyTorch SDP: 104s.
- FP8 + SageAttention: 155s. FP8 + PyTorch SDP was the slowest (the number is cut off).
- The overall speedup is about 2.3x, at no cost.
Resources mentioned
- ComfyUI · tool · x.com · free
The official X account for ComfyUI, the node-based tool for generative AI workflows, which announces things like ComfyUI Agent.
Also in: ComfyUI Agent: early access announcement (Melvin Vivas on X · notes), Comfy Router: One API for Image, Video, 3D and Audio Model Providers (Melvin Vivas on X · notes), Generating AI Video Locally with MiniMax H3 in ComfyUI on an RTX 3090 (Melvin Vivas on X · notes), Running MiniMax M3 Locally in ComfyUI on an RTX 3090 (Melvin Vivas on X · notes) and 23 more - Ideogram 4.0 · tool · ideogram.ai · free
Ideogram's text-to-image model, released with open weights that you can download, fine-tune and run yourself.
Also in: Ideogram 4.0 Released as an Open-Weights Image Model (Melvin Vivas on X · notes) - SageAttention · tool · github.com · free
An optimized attention kernel that speeds up inference. - PyTorch · tool · pytorch.org · free · recommended by both Bashiri Smith & Melvin Vivas
Deep-learning framework. Its built-in scaled dot-product attention (SDP) was the baseline in this benchmark.
Also in: Step-by-Step Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes), Wannabe vs $200K+ AI Engineer: RAG, Agents, Context and Fine-Tuning Mistakes (Bashiri Smith on Facebook · notes), AI DevBox v1.3.0: a GPU-ready Docker image with coding-agent CLIs (Melvin Vivas on X · notes), Docker image with coding agents pre-installed on a CUDA + PyTorch base (Melvin Vivas on X · notes) and 5 more
Try this
- In ComfyUI, switch Ideogram 4 from FP8 to INT8 weights.
- Use SageAttention instead of PyTorch SDP attention.
- Benchmark on your own GPU with a fixed prompt and step count.
- Benchmark different weight-precision and attention-kernel combinations on your own GPU for a diffusion model.
More in LLMOps, Deployment & Monitoring
- Self-Hosting GLM 5.2 with Modal Auto Endpoints
- Where to Access GLM 5.2: Inference Providers and Gateways
- Serving GLM-5.2 on Baseten: >280 TPS and <0.8s TTFT
- Model routing: Rayline picks the best model per task
- Running GLM-5.2 locally on a 256GB Mac with Unsloth
- Tracking LLM Spend, Tokens & Guardrails with OpenRouter's Activity Explorer