Topic 9 of 16 in the learning path
LLMOps, Deployment & Monitoring
Shipping to production: serving, gateways, observability, latency and cost.
Reels and posts (114)
- Deploy Open-Source Models with Hugging Face Inference Endpoints · Melvin Vivas, X: The creator recommends Hugging Face Inference Endpoints as the easiest current way to deploy open-source models.
- Deploying Qwen3.8 27B on Hugging Face Inference Endpoints · Melvin Vivas, X: The creator deployed the open-weights Qwen3.8 27B model on a dedicated Hugging Face Inference Endpoint.
- Using Hugging Face credits: Jobs, Inference Endpoints and Open Models · Melvin Vivas, X: The creator thanks Victor Mustar of Hugging Face for a credit giveaway and plans to use the credits to deploy Qwen3.8 27B on an Inference Endpoint.
- FireRouter: cost-aware model routing between open models and Claude Opus · Melvin Vivas, X: Fireworks AI released FireRouter, its first router model.
- Run Laya Decision models locally with Unsloth on 4GB RAM · Melvin Vivas, X: Unsloth now lets you run Laya Decision models, described as an alternative to Jev, on your own machine with as little as 4GB of RAM.
- AI Backends: A Production AI Workflow Engineering Site (Link Share) · Melvin Vivas, X · 0:15: A 15-second X post with no speech.
- n8n vs Apache Camel for enterprise AI workflows · Melvin Vivas, X: The creator argues that n8n was not built for enterprise use, so he uses Apache Camel for AI workflow integration instead.
- Deploying AI workflow integrations with Apache Camel in Docker · Melvin Vivas, X: The creator is testing Apache Camel for deployable AI workflow integrations.
- Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M · Melvin Vivas, X · 2:16: This short demo shows how Liquid AI's LFM 2.5-Encoder-350M routes prompts zero-shot.
- AIBackends v0.8.1 adds GLiNER2.5-Decide local classification · Melvin Vivas, X: Release v0.8.1 of the creator's open-source library AIBackends adds support for Fastino's GLiNER2.5-Decide model.
- DigitalOcean Serverless Inference Now Serves OpenAI GPT-6 Models · Melvin Vivas, X: DigitalOcean now offers OpenAI models through its Serverless Inference product.
- llama.cpp / Llama-macOS v0.5.0 release · Melvin Vivas, X: The creator shares a v0.5.0 release from ggml-org, the team behind llama.cpp, for running LLMs locally.
- Comfy Router: One API for Image, Video, 3D and Audio Model Providers · Melvin Vivas, X · 0:47: Melvin Vivas shares the announcement of Comfy Router.
- How claude.ai was made 3x faster using Claude itself · Melvin Vivas, X: Melvin Vivas notices that Opus 5.5 feels faster in chat.
- How Cursor cut agent token costs by 7% without losing quality · Melvin Vivas, X: Melvin Vivas quotes Cursor's announcement that it cut token costs by 7% with no drop in agent quality.
- GPT-6 Prompt Caching: Why It Stretches Your Usage Limits · Melvin Vivas, X: Melvin Vivas shares OpenAI's announcement of better prompt caching for GPT-6.
- Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs · Melvin Vivas, X: Unsloth released GGUF quantizations of Qwen-Image-2.1, so the 7B image model runs locally on 12GB of VRAM.
- Run llama.cpp GGUF Checkpoints in Hugging Face Transformers · Melvin Vivas, X: Hugging Face Transformers can now load the same GGUF quantized checkpoints used by llama.cpp.
- Run GGUF models directly in Hugging Face Transformers · Melvin Vivas, X: Hugging Face Transformers can now run GGUF models directly.
- Devin Fusion: Multi-Model Routing to Cut Agentic Coding Costs · Melvin Vivas, X: Melvin Vivas shares Cognition's Devin Fusion, a multi-model routing approach for agentic coding.
- Fly.io Sprites Get a Price Cut · Melvin Vivas, X: A short news post saying Sprites on Fly.io are now cheaper.
- Building a Local Model Server with ONNX Support · Melvin Vivas, X: The creator is building a tool to serve local models and is adding ONNX support for classification models.
- Jev model added to the AIBackends API via Vercel AI Gateway · Melvin Vivas, X: The creator added support for TypeSafe AI's Jev model to his open-source AIBackends API server, routed through Vercel AI Gateway.
- LiteRT: Google's on-device AI runtime · Melvin Vivas, X: A short thank-you to Google for LiteRT, its runtime for running ML models on devices (the successor to TensorFlow Lite).
- Adding Vercel AI Gateway as a provider in AIBackends with Devin · Melvin Vivas, X: The creator used the Devin coding agent to add Vercel AI Gateway (aigw) as a new provider for the Jev model in his open-source AIBackends API.
- Jev was free on Vercel AI Gateway until Sept 25 · Melvin Vivas, X: The creator shares Vercel's announcement that TypeSafe AI's Jev model was free on Vercel AI Gateway until September 25.
- Benchmarking LLM Endpoints with NVIDIA Dynamo AIPerf · Melvin Vivas, X: The creator recommends NVIDIA Dynamo AIPerf for inference engineering.
- Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash · Melvin Vivas, X · 0:26: This short post reacts to a quoted announcement of Inco Splash, an open-source inference engine built for Qwen models on Apple silicon.
- Hugging Face Cache Deduplication with Xet in huggingface_hub v1.32 · Melvin Vivas, X: A quoted post explains that huggingface_hub v1.32 stores identical Xet-backed files only once, even when several repos share them.
- Agent Monitor: see traces, tokens and costs of your coding agents · Melvin Vivas, X: The creator shares Agent Monitor, his open-source tool for seeing what coding agents do behind the scenes.
- Agent Monitor: Per-Model Usage Stats for Codex and Claude Code Agents · Melvin Vivas, X: Melvin Vivas added per-model usage stats to Agent Monitor, his open-source dashboard for Codex and Claude Code subagents.
- Agent Monitor: Track Token Usage and Cost for Codex and Claude Code Agents · Melvin Vivas, X · 0:01: Melvin Vivas shows Agent Monitor, his open-source dashboard for watching coding-agent sessions.
- Agent Monitor: Dashboard for Codex Subagents, Tokens and Cost · Melvin Vivas, X: Melvin Vivas released Agent Monitor on GitHub.
- Concept Demo: An Agent Monitor Built on OpenAI Codex Traces · Melvin Vivas, X · 0:04: This is a 4-second, silent video in which Melvin Vivas shows an "Agent Monitor" dashboard built from OpenAI Codex execution traces and asks viewers whether they would use it.
- Running Local Models on NVIDIA DGX Sparks (Alex Ellis) · Melvin Vivas, X: Melvin Vivas recommends Alex Ellis's article on why and how his team bought four NVIDIA DGX Sparks to run local models.
- Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s · Melvin Vivas, X · 0:12: A short post sharing a working local-inference setup for Qwen3.8-27B quantized in the EXL3 format on one 24 GB RTX 3090.
- llama.cpp v0.4.1 release announcement · Melvin Vivas, X: Short repost announcing a new llama.cpp release, v0.4.1, the popular engine for running LLMs locally.
- DeepSeek V4.1 Flash Off-Peak Pricing as a Cheap Fallback Model · Melvin Vivas, X: The creator calls DeepSeek V4.1 Flash good and very cheap, and suggests it as a fallback when other usage limits run out.
- How inference engines work: the full life of an LLM request · Melvin Vivas, X: The creator bookmarks a quoted post that shares slides from a talk on how LLM inference engines work.
- SGLang v0.5.19 release: new models and beam search · Melvin Vivas, X: Announces SGLang v0.5.19, an open-source LLM serving and inference engine.
- llama.cpp's built-in web UI for testing local models · Melvin Vivas, X: The default web UI that comes with llama.cpp is now good enough for testing local inference.
- Local Speech-to-Text with LFM2.5-Audio-1.5B and llama.cpp · Melvin Vivas, X · 0:16: This is a 16-second demo with no speech.
- Daytona Offers GPU Sandboxes (H100, RTX 4090/5090) · Melvin Vivas, X: Daytona now has GPU sandboxes.
- Long-Context Local LLM on an RTX 3090: .env Config and ~64 tok/s Benchmark · Melvin Vivas, X · 0:12: Melvin Vivas shares a short post saying a local LLM setup from @MiaAI_lab runs on a single 24 GB RTX 3090.
- Docker Template Bundling Coding Agents for GPU Cloud Hosting · Melvin Vivas, X: Melvin Vivas is using Codex to build a Docker image for GPU hosts like Runpod or QuickPod.
- Why Claude Fable 5.1 Costs Less: Cheaper Cache Reads · Melvin Vivas, X: According to Cursor Bench, Claude Fable 5.1 is cheaper to run than Fable 5.
- Ollama vs vLLM: From Local AI Demo to Production Inference Serving · Bashiri Smith, Facebook · 0:45: The video explains why developers often use Ollama while companies use vLLM.
- Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context · Melvin Vivas, X: The creator shares the .env settings he used to run Qwen3.8-27B in EXL3 quantization on a single 24GB RTX 3090 with a 220K-token context, using a DFlash2 draft model for speculativ
- Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090 · Melvin Vivas, X · 0:12: Melvin Vivas shares his results running Mia's (@MiaAI_lab) experimental Qwen3.8-27B EXL3 release on one 24GB RTX 3090.
- Running Qwen3.8 27B Locally on an RTX 3090 with llama.cpp and the Pi Harness · Melvin Vivas, X · 0:23: A 23-second demo with no speech.
- Low-Cost Agent Run: DeepSeek V4 Flash via OpenRouter in ohmypi · Melvin Vivas, X: A cost data point: a coding agent session in ohmypi using DeepSeek V4 Flash through OpenRouter cost about $0.30.
- Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference · Melvin Vivas, X: The creator serves Qwen3.8 27B in FP8 on a dedicated NVIDIA H100 through Baseten's dedicated inference.
- Running GLM 5.3 Flash on Baseten with the Pi Coding Agent · Melvin Vivas, X: The creator runs the GLM 5.3 Flash model through Baseten's inference platform and uses it inside the Pi coding agent.
- AIBackends v0.7.0: Prompt Routing with Liquid AI's LFM2.5 Encoder · Melvin Vivas, X: Melvin Vivas released version 0.7.0 of his AIBackends Python library, which adds prompt routing with Liquid AI's LFM2.5-Encoder-350M.
- GLM 5.3 Flash Now Available on Baseten · Melvin Vivas, X: An announcement that the GLM 5.3 Flash model can now be used through Baseten's model library, a hosted service for running models.
- Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M · Melvin Vivas, X · 2:16: This demo shows how a small encoder model, Liquid AI's LFM 2.5-Encoder-350M, can route prompts zero-shot.
- Novita AI Spot GPU Instances: Cheap GPU Compute · Melvin Vivas, X: The creator says Novita AI's spot GPU instances are the cheapest he has seen.
- Why Codex Usage Drains Faster: Prompt Cache Hit Rate · Melvin Vivas, X: Melvin shares an OpenAI Codex update from Tibo (@thsottiaux): some users drained their Codex rate limits faster because the prompt-cache hit rate dropped that week.
- GPT-5.6 Sol Price Cuts on Vercel AI Gateway (70% Off Input) · Melvin Vivas, X: Vercel AI Gateway cut prices again for GPT-5.6 Sol.
- How to Reduce Latency in a Production AI Agent (Interview Answer) · Bashiri Smith, Facebook · 1:22: This reel uses a mock interview to show how to answer "How do you reduce latency in a production AI agent?" The answer starts by clarifying the agent's architecture.
- Move Non-Coding Work to Local Models (Hermes + llama.cpp) · Melvin Vivas, X: The creator argues for moving some AI work to local models so you depend less on cloud providers and their changing limits.
- Serving Ornith-1.5-35B-A3B at 128k Context on an RTX 3090 with llama.cpp · Melvin Vivas, X: Gives an exact llama-server command for running the Q4_K_M quantization of Ornith-1.5-35B-A3B with a 128k context window, with every layer on a single RTX 3090.
- Running Ornith-1.5-35B-A3B (Q4_K_M) at 128k Context on an RTX 3090 with llama.cpp · Melvin Vivas, X · 1:51: This is a short local-inference benchmark.
- Running DeepSeek V4 Pro 0813 on Baseten with the Pi Agent · Melvin Vivas, X: The creator runs the new DeepSeek V4 Pro 0813 model through Baseten's inference platform and uses it with Pi as the agent harness.
- Red Hat's Quantized Qwen3.8-2.4T-A95B Variant (NVFP4/FP8) · Melvin Vivas, X: Red Hat AI published a quantized version of Qwen3.8 on Hugging Face called Qwen3.8-2.4T-A95B-NVFP4-FP8.
- One-Click Deploy Link for Qwen3.8 27B on Baseten · Melvin Vivas, X: A follow-up post sharing the Baseten deploy page for Qwen3.8 27B, so you can spin up your own dedicated deployment of the model.
- Self-Hosting Qwen3.8 27B FP8 on an H100 for $6.50/hr · Melvin Vivas, X: You can run your own copy of Qwen3.8 27B in FP8 on Baseten.
- Hosted Qwen3.8 27B Costs More Than GPT 5.6 Luna, So Run It Locally · Melvin Vivas, X: The creator found that Qwen3.8 27B costs more on OpenRouter than OpenAI's GPT 5.6 Luna, and decided to run the open-weight model on his own machine instead.
- Running Qwen3.8 27B Locally with Hermes · Melvin Vivas, X: The creator runs Qwen3.8 27B locally with Hermes and calls the model really good.
- AIBackends: Python Library for Local-Model AI Workflows · Melvin Vivas, X: The creator shares his open-source project AIBackends.
- AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers · Melvin Vivas, X: Release notes for version 0.4.0 of AIBackends, the creator's open-source Python package.
- AIBackends Adds Support for LFM2.5-VL-3B · Melvin Vivas, X: The creator says he has finished adding support for Liquid AI's LFM2.5-VL-3B vision model to AIBackends.
- Running LFM2.5-2.6B Q4_K_M Locally with llama.cpp and Pi · Melvin Vivas, X: The creator runs Liquid AI's LFM2.5-2.6B at Q4_K_M quantization through llama.cpp's llama-server and connects it to the Pi agent.
- Run Qwen3.8-27B locally with llama.cpp (llama-server) · Melvin Vivas, X: A complete llama-server command for running Unsloth's Q4_K_M GGUF build of Qwen3.8-27B locally from Hugging Face.
- Unsloth releases Qwen3.8-27B GGUF quantizations · Melvin Vivas, X: Announces that Unsloth has published GGUF builds of Qwen3.8-27B, so you can run it locally with llama.cpp-compatible tools.
- Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp · Melvin Vivas, X: Shares a guide to running Muse Glimmer 30B on one RTX 3090 with llama.cpp under WSL Ubuntu on Windows 11.
- Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent · Melvin Vivas, X: Melvin Vivas tests Meta's Muse Glimmer 30B open model with the Hermes agent.
- Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent · Melvin Vivas, X: This post is a step-by-step guide to serving Unsloth's quantized Muse Glimmer 30B GGUF model on one RTX 3090 (24 GB) using llama.cpp built with CUDA under WSL Ubuntu on Windows 11.
- Running Liquid AI LFM2.5-2.6B Locally with llama-server · Melvin Vivas, X: This post shares a llama.cpp `llama-server` command for serving Liquid AI's LFM2.5-2.6B from its Q8_0 GGUF.
- Adding LFM2.5-2.6B support to AIBackends with Cursor · Melvin Vivas, X: The creator announces he is adding support for Liquid AI's small LFM2.5-2.6B model to his site AIBackends, building it with Cursor and the Fable model.
- Zero-Shot Prompt Routing by Task Complexity with LFM2.5-Encoder · Melvin Vivas, X · 2:16: Melvin Vivas demos zero-shot prompt routing with the LFM2.5-Encoder bidirectional encoders.
- Run LFM2.5-2.6B locally with llama.cpp · Melvin Vivas, X: The creator quotes his own post that runs Liquid AI's LFM2.5-2.6B GGUF with llama-server.
- Serving LFM2.5-2.6B with llama-server: full command · Melvin Vivas, X: A full llama-server command for running Liquid AI's LFM2.5-2.6B GGUF (Q8_0) on a GPU.
- Personal AI Computer to Run DeepSeek V4-Flash Locally · Melvin Vivas, X: Autonomous is building a desktop 'Personal AI Computer' that runs DeepSeek V4-Flash on-premises: private and with no per-token cost.
- Running DeepSeek V4 Flash Locally: RAM Needs for 4-bit and 3-bit Quants · Melvin Vivas, X: Quoting Unsloth, the creator jokes about the hardware needed to run DeepSeek V4 Flash 0731 locally.
- aibackends 0.3.0: Model Caching Speeds Up PII and OCR Inference · Melvin Vivas, X: The creator's Python library aibackends now caches models.
- Monitor GPU Usage With nvtop Instead of nvidia-smi · Melvin Vivas, X: The creator recommends nvtop, an interactive, htop-style GPU monitor, instead of running nvidia-smi to check GPU usage.
- Cost-Saving Model Fallback Chain: Grok → Codex → OpenRouter DeepSeek · Melvin Vivas, X: Melvin describes a cost-saving fallback chain for his Hermes agent.
- Running a Personal AI Agent for $0.69/Day with DeepSeek V4 Flash · Melvin Vivas, X: A real cost figure: one day of running DeepSeek V4 Flash through OpenRouter as the main agent model in Hermes cost $0.692.
- Hosting a Personal AI Agent: Local Mac Mini vs Cloud VM Privacy Trade-off · Melvin Vivas, X: Melvin Vivas asks for cheaper alternatives to a Mac Mini for running a personal AI agent (Hermes).
- Frontier Model Costs: Why Top Models Can Sink a Solo Dev's MRR · Melvin Vivas, X: Quoting a post about a team's large 24-hour bill from using Fable 5, the creator warns that solo developers who rely on the most expensive frontier models can easily end up with ne
- GLM 5.2 Hits 446 tok/s on Fireworks AI · Melvin Vivas, X: Fireworks AI reports serving GLM 5.2 at 446 tokens per second, as measured by Artificial Analysis.
- Fastest GLM-5.2 Provider: Fireworks AI at 343 tok/s · Melvin Vivas, X: Says Fireworks AI is currently the fastest GLM-5.2 provider at about 343 tokens/sec.
- Serve GLM-5.2 NVFP4 with vLLM on NVIDIA Blackwell · Melvin Vivas, X: vLLM now supports NVIDIA's official NVFP4-quantized checkpoint of GLM-5.2.
- Self-Hosting GLM 5.2 with Modal Auto Endpoints · Melvin Vivas, X: Says you can serve the open-weight GLM 5.2 model on your own infrastructure using Modal's new Auto Endpoints.
- Where to Access GLM 5.2: Inference Providers and Gateways · Melvin Vivas, X: Lists the providers that serve the GLM 5.2 model: Z.ai's own coding plan, Together AI, Baseten, Fireworks AI, Vercel AI Gateway and OpenRouter.
- Serving GLM-5.2 on Baseten: >280 TPS and <0.8s TTFT · Melvin Vivas, X: Baseten hosts the open model GLM-5.2 and claims more than 280 tokens per second with under 0.8s time to first token.
- 2.3x Faster Ideogram 4 in ComfyUI with INT8 and SageAttention · Melvin Vivas, X: A quoted benchmark shows how to speed up Ideogram 4.0 image generation in ComfyUI: switch from FP8 weights to INT8, and from PyTorch SDP attention to SageAttention.
- Model routing: Rayline picks the best model per task · Melvin Vivas, X: This post introduces model routing: sending each task to whichever LLM handles it best.
- Running GLM-5.2 locally on a 256GB Mac with Unsloth · Melvin Vivas, X: Thanks to Unsloth's quantized builds, GLM-5.2 can now run locally on a Mac with 256GB of unified memory.
- Tracking LLM Spend, Tokens & Guardrails with OpenRouter's Activity Explorer · Melvin Vivas, X · 2:15: Melvin Vivas shares OpenRouter's launch video for its new Activity Explorer, a real-time dashboard for AI usage across a team.
- Bonsai Image Model Running In-Browser with WebGPU · Melvin Vivas, X: Melvin Vivas shares a Hugging Face Space where the Bonsai image generation model runs entirely in the browser using WebGPU.
- Running Qwen3.6-27B Fully in the Browser With WebGPU and wllama · Melvin Vivas, X: Xuan Son Nguyen (@ngxson) of Hugging Face showed Qwen3.6-27B running entirely in a web browser on WebGPU, with no cloud.
- Unsloth MTP GGUFs make Qwen3.6 run 1.4x faster locally · Melvin Vivas, X: Unsloth released experimental GGUFs of Qwen3.6 that use Multi-Token Prediction (MTP).
- Cheap TTS serving: Qwen3-TTS on vLLM-Omni at $3 per 1M characters · Melvin Vivas, X: A provider reports serving the open Qwen3-TTS model on vLLM-Omni for $3 per 1M characters.
- LM Studio MLX v1.8.1: Vision Model Batching and Better Caching · Melvin Vivas, X: LM Studio's latest MLX engine update adds batching for vision models in beta and improves caching for faster inference.
- OpenRouter Pareto Code: cost-optimized coding router · Melvin Vivas, X: OpenRouter launched Pareto Code, a free, experimental router for coding requests.
- Speeding Up Gemma 4 Inference: MTP (3x) vs. DFlash Speculative Decoding (6x) · Melvin Vivas, X · 0:05: Melvin Vivas shares a quick inference-speed update: Gemma 4 ran about 3x faster with Multi-Token Prediction (MTP), which now ships natively in Gemma 4, and up to 6x faster with DFl
- Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code · Melvin Vivas, X · 0:24: This is a short demo with no narration.
- Runpod Flash Reaches GA: Deploy AI Workloads from Python · Melvin Vivas, X · 0:50: Melvin Vivas reshares Runpod's announcement that Flash is now generally available (GA).
- Z.ai's Lessons from Serving GLM-5 for Coding Agents at Scale · Melvin Vivas, X: The creator shares Z.ai's blog post "Scaling Pain" about debugging GLM-5 at scale.
- Running AI Locally: Rebuilding AIBackends as a Python Library with Open Models · Melvin Vivas, X · 0:07: Melvin Vivas shows a short demo of AIBackends, his project for running AI tasks locally.
- Run Gemma 4 Locally with llama.cpp in Two Commands · Melvin Vivas, X: The post gives a quick way to run Gemma 4 locally with llama.cpp on macOS.
- LM Studio's LM Link: Use a Local Model on Another Machine · Melvin Vivas, X: LM Link, a feature of LM Studio, connects one LM Studio instance to another.
Watch, free (4)
- Large Language Model Operations (LLMOps) Explained (IBM Technology) · video · youtube.com · free
An IBM explainer video on running and operating LLM applications (LLMOps). The exact title was shown on screen only.
Mentioned in: AI Engineer Roadmap Before 2027: Fundamentals, RAG, Agents, Ops, Evals (Bashiri Smith on Facebook · notes) - LLMOps (DeepLearning.AI) · course · learn.deeplearning.ai · free
A DeepLearning.AI course on putting LLM applications into production. The exact course was shown on screen only. Free with a DeepLearning.AI account during its platform beta; certificates are paid.
Mentioned in: AI Engineer Roadmap Before 2027: Fundamentals, RAG, Agents, Ops, Evals (Bashiri Smith on Facebook · notes) - Made With ML · course · madewithml.com · free
End-to-end MLOps curriculum; the guide calls it the best free resource in the category.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Original LFM 2.5-Encoder prompt routing demo video (YouTube) · video · youtube.com · free
The full original YouTube demo of zero-shot prompt routing with LFM 2.5-Encoder. The caption link appears truncated and looked broken when checked.
Mentioned in: Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M (Melvin Vivas on X · notes)
Read and use (148)
- OpenRouter · tool · openrouter.ai · free
A single OpenAI-compatible API that routes requests to hundreds of models from many providers, with one bill and model fallbacks.
Mentioned in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), The Jev model is now on OpenRouter (Melvin Vivas on X · notes), Kev-4B model, an alternative to Jev, now available on OpenRouter (Melvin Vivas on X · notes), Space Bunny Alpha: stealth 1M-context flash model on OpenRouter (Melvin Vivas on X · notes) and 43 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Mentioned in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 29 more - AIBackends · repo · aibackends.com · free
Open-source API server runtime for common AI use cases that supports many models and providers (Ollama, LM Studio, OpenRouter, OpenAI, Anthropic).
Mentioned in: Building a production website with Opus 5.5 and TanStack (Melvin Vivas on X · notes), CamelFlow: open-source visual viewer for Apache Camel routes (Melvin Vivas on X · notes), AI Backends: A Production AI Workflow Engineering Site (Link Share) (Melvin Vivas on X · notes), Demo: Claude Opus 5.5 Generating a Motion-Graphics Video for AIBackends (Melvin Vivas on X · notes) and 17 more - Ollama · tool · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
Open-source tool for downloading and running LLMs on your own machine with minimal setup.
Mentioned in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), Adding Vercel AI Gateway as a provider in AIBackends with Devin (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), AIBackends: An API Layer Between Your App and AI Models (Now with Jev) (Melvin Vivas on X · notes) and 13 more - Vercel · tool · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
Deploy web apps and frontends.
Mentioned in: OpenAI DevDay 2026 Recap: Dots, Agents API, Codex Cloud & Marketplace (Melvin Vivas on X · notes), 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes), Jev model added to the AIBackends API via Vercel AI Gateway (Melvin Vivas on X · notes), Open models now dominate token volume on Vercel AI Gateway (Melvin Vivas on X · notes) and 8 more
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Baseten · tool · x.com · paid
Model inference platform offering dedicated GPU deployments for serving open models.
Mentioned in: Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference (Melvin Vivas on X · notes), Coworker Desktop App Works with Any OpenAI-Compatible Endpoint (Melvin Vivas on X · notes), Running GLM 5.3 Flash on Baseten with the Pi Coding Agent (Melvin Vivas on X · notes), GLM 5.3 Flash Now Available on Baseten (Melvin Vivas on X · notes) and 7 more - Docker · tool · docs.docker.com · free · recommended by both Bashiri Smith & Melvin Vivas
Containerize everything you ship.
Mentioned in: Deploying AI workflow integrations with Apache Camel in Docker (Melvin Vivas on X · notes), Devin's cloud Ubuntu sandbox ships with Docker pre-installed (Melvin Vivas on X · notes), 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes), 8-Week Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes) and 6 more
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Runpod · tool · x.com · paid
GPU cloud platform for running, training and serving AI models; Flash deploys workloads to it.
Mentioned in: AI DevBox v1.3.0: a GPU-ready Docker image with coding-agent CLIs (Melvin Vivas on X · notes), Why LLM data agents need solid data foundations (Runpod) (Melvin Vivas on X · notes), Docker image with coding agents pre-installed on a CUDA + PyTorch base (Melvin Vivas on X · notes), GPU devbox Docker image with coding agents on Runpod (Melvin Vivas on X · notes) and 6 more - Fireworks AI · tool · x.com · free
Fireworks AI's official X account, which posts news about inference and serving open models.
Mentioned in: FireRouter: cost-aware model routing between open models and Claude Opus (Melvin Vivas on X · notes), GLM 5.2 Hits 446 tok/s on Fireworks AI (Melvin Vivas on X · notes), GLM 5.2 Speed vs Opus 4.8 and GPT 5.5 (Melvin Vivas on X · notes), GLM 5.2 on Fireworks Makes Opus 4.8 and GPT 5.5 Feel Slow (Melvin Vivas on X · notes) and 5 more - donvito/agent-monitor · repo · github.com · free
Open-source dashboard that visualizes Codex and Claude Code subagents, with a traces view, a token breakdown and estimated cost.
Mentioned in: Monitor Codex and Claude Code Subagents with agent-monitor (Melvin Vivas on X · notes), Agent Monitor: see traces, tokens and costs of your coding agents (Melvin Vivas on X · notes), Agent Monitor: Visualize Codex and Claude Code Subagents, Traces and Costs (Melvin Vivas on X · notes), Agent Monitor: Visualize Coding-Agent Subagents, Traces and Costs (Melvin Vivas on X · notes) and 3 more - NVIDIA GeForce RTX 3090 · tool · nvidia.com · paid
Consumer GPU with 24 GB of VRAM, used here to run a long-context LLM locally.
Mentioned in: Portable Computer Now Runs AI Agents and Models Locally on NVIDIA RTX PCs (Melvin Vivas on X · notes), Long-Context Local LLM on an RTX 3090: .env Config and ~64 tok/s Benchmark (Melvin Vivas on X · notes), Running Ornith-1.5-35B-A3B (Q4_K_M) at 128k Context on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Generating AI Video Locally with MiniMax H3 in ComfyUI on an RTX 3090 (Melvin Vivas on X · notes) and 3 more - FastAPI · tool · fastapi.tiangolo.com · free
Wrap your models as real inference APIs.
Mentioned in: Step-by-Step Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes), 8-Week Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes), Software Engineer to AI Engineer: Job Boards, Stack, Projects and Learning Sites (Bashiri Smith on Facebook · notes), AI Engineer Roadmap for 2026 in 60 Seconds (Bashiri Smith on Facebook · notes) and 1 more
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - LFM2.5-Encoder-350M · tool · huggingface.co · free
A 350M-parameter encoder model from Liquid AI that does zero-shot prompt classification and routing against categories you supply at runtime.
Mentioned in: Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M (Melvin Vivas on X · notes), AIBackends v0.7.0: Prompt Routing with Liquid AI's LFM2.5 Encoder (Melvin Vivas on X · notes), Model Routing Fine-Tuned on Your Agent Harness Traces (Melvin Vivas on X · notes), Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M (Melvin Vivas on X · notes) and 1 more - Langfuse · tool · langfuse.com · free
Listed under platforms that combine tracing and evaluation.
Mentioned in: Step-by-Step Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes), 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes), Taking a RAG App to Production: Evals, Guardrails, Cost and Tracing (Bashiri Smith on Facebook · notes)
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook - LangSmith · tool · docs.smith.langchain.com · free
Tracing, logging and evals; also used to track cost and latency.
Mentioned in: Step-by-Step Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes), 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes)
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook; The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - MiaAI_lab (@MiaAI_lab on X) · person · x.com · free
X account that published the experimental Qwen3.8-27B EXL3 release for serving long contexts on 24GB GPUs.
Mentioned in: Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s (Melvin Vivas on X · notes), Long-Context Local LLM on an RTX 3090: .env Config and ~64 tok/s Benchmark (Melvin Vivas on X · notes), Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context (Melvin Vivas on X · notes), Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090 (Melvin Vivas on X · notes) - MLX · repo · github.com · free
Apple's open-source array and ML framework for efficient inference on Apple Silicon.
Mentioned in: LM Studio MLX v1.8.1: Vision Model Batching and Better Caching (Melvin Vivas on X · notes), Gemma 4 Gets Up to 3x Faster with MTP Drafters (Melvin Vivas on X · notes), 1-bit Bonsai 8B Runs On-Device on iPhone at 40+ tok/s (Melvin Vivas on X · notes), Qwen 3.5 2B Runs On-Device on iPhone with MLX (Melvin Vivas on X · notes) - Qwen3.8-27B-GGUF · tool · huggingface.co · free
Unsloth's GGUF quantized versions of the Qwen3.8 27B model, ready for llama.cpp and other GGUF runtimes.
Mentioned in: Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Running Qwen3.8 27B Locally on an RTX 3090 with llama.cpp and the Pi Harness (Melvin Vivas on X · notes), Run Qwen3.8-27B locally with llama.cpp (llama-server) (Melvin Vivas on X · notes), Unsloth releases Qwen3.8-27B GGUF quantizations (Melvin Vivas on X · notes) - Together AI · tool · x.com · paid
Inference platform for hosting and calling open-weights models.
Mentioned in: Fine-tune and Deploy Qwen3.8 27B on Together AI (Melvin Vivas on X · notes), Where to Access GLM 5.2: Inference Providers and Gateways (Melvin Vivas on X · notes), Where to Access GLM-5.2: Provider Roundup (Melvin Vivas on X · notes), Free GLM-5.2 via Hugging Face Inference Providers in coding agents (Melvin Vivas on X · notes) - unsloth/Muse-Glimmer-30B-GGUF · tool · huggingface.co · free
Unsloth's GGUF quantizations of Meta's Muse Glimmer 30B open-weights model, used here with the UD-Q4_K_XL quant.
Mentioned in: Three Local GGUF Models That Fit on an RTX 3090 (24GB) (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent (Melvin Vivas on X · notes) - DVC (Data Version Control) · tool · dvc.org · free · open in a browser to verify
Version your data like code.
Mentioned in: Five Core AI Engineering Topic Areas: A Study Roadmap Checklist (Bashiri Smith on Facebook · notes), 17-Step AI Engineer Roadmap: From Basic RAG to Agents, Evals, LLMOps & Governance (Bashiri Smith on Facebook · notes)
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - ExLlamaV3 (EXL3) · repo · github.com · free
turboderp's inference library and the EXL3 quantization format it uses to run quantized LLMs on consumer NVIDIA GPUs.
Mentioned in: Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s (Melvin Vivas on X · notes), Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context (Melvin Vivas on X · notes), Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090 (Melvin Vivas on X · notes) - Hugging Face Inference Endpoints · tool · endpoints.huggingface.co · paid
Managed service for deploying models from the Hugging Face Hub on dedicated infrastructure.
Mentioned in: Deploy Open-Source Models with Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes) - LiteLLM · tool · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
An open-source LLM gateway/proxy that gives one API for many model providers, with cost tracking and pricing features.
Mentioned in: 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes), Using Codex to Fix a LiteLLM Pricing-Markup Bug in Docker (Melvin Vivas on X · notes), Codex Finds a Bug in a LiteLLM Feature (Melvin Vivas on X · notes) - LiteRT · tool · ai.google.dev · free
Google's on-device runtime (formerly TensorFlow Lite) for running ML models and LLMs on mobile and edge devices.
Mentioned in: Gemma 4 Runs Locally On-Device in the Antigravity SDK (Melvin Vivas on X · notes), On-Device AI: Running Gemma 4 E2B Offline on an iPhone with LiteRT (Melvin Vivas on X · notes), LiteRT: Google's on-device AI runtime (Melvin Vivas on X · notes) - nvtop · tool · github.com · free
An open-source, top-style monitor for GPU usage in the terminal.
Mentioned in: GPU devbox Docker image with coding agents on Runpod (Melvin Vivas on X · notes), GPU-Ready AI Devbox Docker Image on Runpod (PyTorch 2.8 + CUDA 12.8) (Melvin Vivas on X · notes), Monitor GPU Usage With nvtop Instead of nvidia-smi (Melvin Vivas on X · notes) - QuickPod · tool · console.quickpod.io · paid
A cloud service for renting GPU pods to host and run your own models.
Mentioned in: Reusable GPU devbox: PyTorch/CUDA plus six coding agents (Melvin Vivas on X · notes), Docker Template Bundling Coding Agents for GPU Cloud Hosting (Melvin Vivas on X · notes), Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code (Melvin Vivas on X · notes) - Replicate · tool · replicate.com · check price
A cloud platform for running open and hosted AI models through an API; you pay for what you use.
Mentioned in: Replicate's Guide to Making Videos with Seedance 2.0 (Melvin Vivas on X · notes), Seedance 2.0 Video Generation Model Now on Replicate (Melvin Vivas on X · notes), LTX-2.3 Video Generation Model Now Available on Replicate (Melvin Vivas on X · notes) - Unsloth LLM Tutorials (Unsloth Documentation) · docs · unsloth.ai · free
Unsloth's guides for running and fine-tuning open LLMs locally, including the Qwen3.8-Next guide.
Mentioned in: Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Run Qwen3.8-Flash-Next (125B MoE) Locally with Unsloth GGUFs (Melvin Vivas on X · notes), Run Qwen3.6-27B Locally in 18GB RAM with Unsloth GGUFs (Melvin Vivas on X · notes) - vLLM · repo · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
Open-source, high-throughput LLM inference and serving engine with continuous batching and efficient KV-cache management (PagedAttention).
Mentioned in: Ollama vs vLLM: From Local AI Demo to Production Inference Serving (Bashiri Smith on Facebook · notes), Serve GLM-5.2 NVFP4 with vLLM on NVIDIA Blackwell (Melvin Vivas on X · notes), Gemma 4 Gets Up to 3x Faster with MTP Drafters (Melvin Vivas on X · notes) - @aivandroid · person · x.com · free
The developer who maintains the geocine/llama-swap fork.
Mentioned in: Serving LFM2.5-2.6B with llama-server: full command (Melvin Vivas on X · notes), Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code (Melvin Vivas on X · notes) - Cloudflare · tool · x.com · check price
Cloud platform listed as a supported BYOS sandbox provider.
Mentioned in: OpenAI Agents API with Bring Your Own Sandbox as a backend for a 'software factory' (Melvin Vivas on X · notes), Using Codex as a Sysadmin to Find a Cloudflare DNS Issue (Melvin Vivas on X · notes) - dflash2 · repo · github.com · free
Open-source speculative decoding method that uses a block-diffusion drafter to speed up LLM inference, with support for Gemma 4.
Mentioned in: Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090 (Melvin Vivas on X · notes), Speeding Up Gemma 4 Inference: MTP (3x) vs. DFlash Speculative Decoding (6x) (Melvin Vivas on X · notes) - ggml-org/Llama-macOS · repo · github.com · free
A ggml-org GitHub repo described as 'A cosy home for your LLMs': a macOS app for running local models.
Mentioned in: llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) - GitHub Actions · tool · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
GitHub's CI/CD automation platform, used to run the self-healing docs workflow.
Mentioned in: 5 Weekend AI Engineering Projects: Cost Routing, Caching, Evals & Observability (Bashiri Smith on Facebook · notes), Use Claude Code /loop to Watch and Auto-Fix Failing GitHub Actions (Melvin Vivas on X · notes) - GLM-5.2 | Model library (Baseten) · website · baseten.co · check price
Baseten model library page for trying and deploying GLM-5.2, including its vision support.
Mentioned in: GLM-5.2 Vision on Baseten: Turning Images into Code (Melvin Vivas on X · notes), Serving GLM-5.2 on Baseten: >280 TPS and <0.8s TTFT (Melvin Vivas on X · notes) - LLMOps: Managing Large Language Models in Production · book · oreilly.com · paid · open in a browser to verify
Abi Aryan's O'Reilly book on running LLM applications in production (paid). The guide linked an unofficial PDF copy; this is the publisher's page.
Mentioned in: 5 Books to Move from Software Engineer to AI/ML Engineer (Bashiri Smith on Facebook · notes)
In a shared PDF: 5 books to upgrade from software engineer into Ai/ML engineer - shared in this reel on Facebook - Mac Mini · tool · apple.com · paid
Apple desktop computer popular as always-on local hardware for personal AI agents.
Mentioned in: Hosting a Personal AI Agent: Local Mac Mini vs Cloud VM Privacy Trade-off (Melvin Vivas on X · notes), Running Local Models for Agents: Tool Use, Context and Quantization (Melvin Vivas on X · notes) - MLflow · tool · mlflow.org · free
Experiment tracking and a model registry.
Mentioned in: Five Core AI Engineering Topic Areas: A Study Roadmap Checklist (Bashiri Smith on Facebook · notes)
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Modal · tool · x.com · check price
Serverless cloud platform for running and serving AI models and GPU workloads.
Mentioned in: OpenAI Agents API with Bring Your Own Sandbox as a backend for a 'software factory' (Melvin Vivas on X · notes), Self-Hosting GLM 5.2 with Modal Auto Endpoints (Melvin Vivas on X · notes) - Novita · tool · novita.ai · check price
A cloud platform where you can call open models like the Qwen3.8 family through an API.
Mentioned in: Qwen3.8 Model Family Now on Novita AI: Flash, 27B and 2.4T-A95B (Melvin Vivas on X · notes), Free GLM-5.2 via Hugging Face Inference Providers in coding agents (Melvin Vivas on X · notes) - NVIDIA DGX Spark · tool · nvidia.com · paid
NVIDIA's compact desktop AI computer for running and fine-tuning models locally.
Mentioned in: Running Local Models on NVIDIA DGX Sparks (Alex Ellis) (Melvin Vivas on X · notes), NVIDIA buying Hugging Face: what it could mean for open models (Melvin Vivas on X · notes) - NVIDIA RTX 3090 · tool · nvidia.com · paid
A consumer GPU with 24GB of VRAM, often used to run LLMs locally.
Mentioned in: Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context (Melvin Vivas on X · notes), Running Hermes Agent with Local Models on an RTX 3090 (Melvin Vivas on X · notes) - NVIDIA T4 GPU · tool · nvidia.com · free
The data-center GPU (about 15 GB of VRAM) that Colab offers on its free tier.
Mentioned in: Run Notebooks on a Free GPU with Google Colab (T4, 15GB VRAM) (Melvin Vivas on X · notes), Run Local Models on a Free GPU with Google Colab (T4) (Melvin Vivas on X · notes) - nvidia-smi · tool · developer.nvidia.com · free
An NVIDIA command-line tool that shows the attached GPU, its VRAM and how much is in use.
Mentioned in: Run Local Models on a Free GPU with Google Colab (T4) (Melvin Vivas on X · notes), Monitor GPU Usage With nvtop Instead of nvidia-smi (Melvin Vivas on X · notes) - Terraform · tool · terraform.io · free · recommended by both Bashiri Smith & Melvin Vivas
Infrastructure-as-code tool for creating and managing cloud infrastructure.
Mentioned in: Step-by-Step Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes), Cautionary Tale: Claude Code Wiped a Production Database via Terraform (Melvin Vivas on X · notes) - WebGPU · tool · github.com · free
A browser API for GPU compute that lets machine-learning models run fast on the user's device.
Mentioned in: Running Qwen3.6-27B Fully in the Browser With WebGPU and wllama (Melvin Vivas on X · notes), Real-Time Speech Transcription in the Browser with Voxtral and WebGPU (Melvin Vivas on X · notes) - AIBackends (PyPI 0.4.0) · tool · pypi.org · free
The creator's Python package for running local AI model backends.
Mentioned in: AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers (Melvin Vivas on X · notes) - aibackends 0.7.0 (PyPI) · tool · pypi.org · free
PyPI page for the AIBackends Python package, version 0.7.0, which adds prompt routing.
Mentioned in: AIBackends v0.7.0: Prompt Routing with Liquid AI's LFM2.5 Encoder (Melvin Vivas on X · notes) - Alex Ellis · article · blog.alexellis.io · free
Alex Ellis explains why his team moved local LLM work from RTX 3090s to four NVIDIA DGX Sparks.
Mentioned in: Running Local Models on NVIDIA DGX Sparks (Alex Ellis) (Melvin Vivas on X · notes) - Arize Phoenix · tool · github.com · free
Listed under platforms that combine tracing and evaluation.
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook - autonomous-ai/autonomous-computer · repo · github.com · free
Planned open-source build of a personal AI computer for running DeepSeek locally.
Mentioned in: Personal AI Computer to Run DeepSeek V4-Flash Locally (Melvin Vivas on X · notes) - AWS · tool · aws.amazon.com · free
Amazon Web Services cloud platform for deploying and scaling applications, with a free tier.
Mentioned in: 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes) - AWS Application Load Balancer (ALB) · tool · aws.amazon.com · paid
AWS's managed load balancer for HTTP and HTTPS traffic.
Mentioned in: Using Codex as a Sysadmin to Find a Cloudflare DNS Issue (Melvin Vivas on X · notes) - Baseten Deploy (model library) · tool · app.baseten.co · paid
Baseten's list of models you can deploy.
Mentioned in: One-Click Deploy Link for Qwen3.8 27B on Baseten (Melvin Vivas on X · notes) - Baseten Deploy: Qwen3.8 27B · tool · app.baseten.co · paid
Baseten page for deploying Qwen3.8 27B as your own dedicated endpoint.
Mentioned in: One-Click Deploy Link for Qwen3.8 27B on Baseten (Melvin Vivas on X · notes) - Better prompt caching for GPT-6 · article · openai.com · free
OpenAI's announcement of improved prompt caching for GPT-6 models.
Mentioned in: GPT-6 Prompt Caching: Why It Stretches Your Usage Limits (Melvin Vivas on X · notes) - Bonsai Image WebGPU (Hugging Face Space by webml-community) · tool · huggingface.co · free
Browser demo that runs the Bonsai image generation model locally with WebGPU.
Mentioned in: Bonsai Image Model Running In-Browser with WebGPU (Melvin Vivas on X · notes) - Braintrust · tool · braintrust.dev · free
Listed under platforms that combine tracing and evaluation.
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook - ChatGPT sites · tool · help.openai.com · paid
OpenAI's hosting feature that publishes web apps to *.chatgpt.site subdomains.
Mentioned in: Building and Deploying a Flappy Bird Clone with GPT-6 Astra in Codex (Melvin Vivas on X · notes) - Comfy Router · tool · comfy.org · paid
ComfyUI's single API that routes requests to frontier image, video, 3D and audio models from different providers.
Mentioned in: Comfy Router: One API for Image, Video, 3D and Audio Model Providers (Melvin Vivas on X · notes) - Darkbloom · tool · darkbloom.dev · paid
Eigen Labs' inference service that provides the free model capacity on OpenRouter.
Mentioned in: Free gpt-oss-20b and Gemma 4 26B on OpenRouter (Melvin Vivas on X · notes) - DeepInfra · tool · deepinfra.com · check price
An inference provider for open models.
Mentioned in: Free GLM-5.2 via Hugging Face Inference Providers in coding agents (Melvin Vivas on X · notes) - DigitalOcean Droplets · tool · digitalocean.com · paid
DigitalOcean's virtual machines, used here as cloud dev environments for Codex.
Mentioned in: Run Codex in the Cloud: Spin Up DigitalOcean Droplets with the Codex Plugin (Melvin Vivas on X · notes) - DigitalOcean Model Library · website · digitalocean.com · check price
DigitalOcean's catalog for comparing and using large language models available through its inference service.
Mentioned in: DigitalOcean Serverless Inference Now Serves OpenAI GPT-6 Models (Melvin Vivas on X · notes) - DigitalOcean Serverless Inference · tool · digitalocean.com · paid
DigitalOcean's managed serverless API for calling hosted LLMs.
Mentioned in: DigitalOcean Serverless Inference Now Serves OpenAI GPT-6 Models (Melvin Vivas on X · notes) - Eigen Labs · person · x.com · check price
Company whose Darkbloom infrastructure serves the free model capacity on OpenRouter.
Mentioned in: Free gpt-oss-20b and Gemma 4 26B on OpenRouter (Melvin Vivas on X · notes) - Evidently AI · tool · evidentlyai.com · free
Drift detection and monitoring.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - FireRouter · article · fireworks.ai · free
Fireworks AI's router model that sends each request to either open models or Claude Opus, depending on the task and cache.
Mentioned in: FireRouter: cost-aware model routing between open models and Claude Opus (Melvin Vivas on X · notes) - Fireworks AI blog: GLM 5.2 fast inference · article · fireworks.ai · free
Fireworks AI blog post on serving GLM 5.2 at high token throughput.
Mentioned in: GLM 5.2 Hits 446 tok/s on Fireworks AI (Melvin Vivas on X · notes) - Fly.io (@flydotio) on X · website · x.com · free
The official X account of Fly.io, a cloud platform for deploying apps, which posts news about products like Sprites.
Mentioned in: Fly.io Sprites Get a Price Cut (Melvin Vivas on X · notes) - Friendli · tool · friendli.ai · paid
An inference provider on OpenRouter (the transcript says 'Friendly').
Mentioned in: OpenRouter MCP: Pick, Price, and Test LLMs from Inside Your Coding Agent (Melvin Vivas on X · notes) - geocine/llama-swap (GitHub) · repo · github.com · free
A fork of llama-swap that reliably swaps models for any local server compatible with the OpenAI or Anthropic API, such as llama.cpp or vLLM.
Mentioned in: Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code (Melvin Vivas on X · notes) - ggml · repo · github.com · free
Tensor library behind llama.cpp and the GGUF format, with Metal kernels for Apple GPUs.
Mentioned in: Run GGUF models directly in Hugging Face Transformers (Melvin Vivas on X · notes) - ggml-org kernels on Hugging Face · repo · huggingface.co · free · open in a browser to verify
ggml kernels on Hugging Face that power fast local inference on Mac in Transformers.
Mentioned in: Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes) - GLM-5.3 on AWS Marketplace · website · aws.amazon.com · paid
AWS Marketplace listing for the GLM-5.3 model (the link looked broken when checked).
Mentioned in: GLM-5.3 now available on AWS Marketplace (Melvin Vivas on X · notes) - GLM-5.3 | Baseten Model library · website · baseten.co · check price
Baseten's page for the full GLM-5.3 model.
Mentioned in: GLM 5.3 Flash Now Available on Baseten (Melvin Vivas on X · notes) - GLM-5.3-Flash | Baseten Model library · website · baseten.co · check price
Baseten's page for running the GLM-5.3-Flash model.
Mentioned in: GLM 5.3 Flash Now Available on Baseten (Melvin Vivas on X · notes) - Google Cloud Agent Platform · tool · docs.cloud.google.com · paid
Google Cloud platform where you can deploy a self-hosted GLM 5.2.
Mentioned in: GLM 5.2 open model now deployable on Google Cloud (Melvin Vivas on X · notes) - Grafana · tool · grafana.com · free
An open-source dashboarding and monitoring platform for visualizing metrics, logs and traces.
Mentioned in: 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes) - Groq (@GroqLLC) on X · person · x.com · free
Groq's official X account.
Mentioned in: Qwen3.8-27B Free on Groq at ~450 Tokens per Second (Melvin Vivas on X · notes) - Hostinger · tool · x.com · paid
Web hosting provider that offers VMs, which the creator says are marketed for running Hermes.
Mentioned in: Hosting a Personal AI Agent: Local Mac Mini vs Cloud VM Privacy Trade-off (Melvin Vivas on X · notes) - How inference engines actually work (talk slides) · article · x.com · free
Talk slides that follow an LLM request through an inference engine: caching, batching, paged attention, chunked prefill, sampling and agentic loops.
Mentioned in: How inference engines work: the full life of an LLM request (Melvin Vivas on X · notes) - How we made claude.ai faster (claude.dev blog) · article · claude.dev · free
A blog post on how Claude was used to measure, debug and improve claude.ai performance, with prompts and methods included.
Mentioned in: How claude.ai was made 3x faster using Claude itself (Melvin Vivas on X · notes) - Hugging Face Blog: Transformers + llama.cpp quants · article · huggingface.co · free
Hugging Face blog post announcing that llama.cpp GGUF quantized checkpoints can run in Transformers.
Mentioned in: Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes) - Hugging Face Inference Providers · tool · huggingface.co · check price
A unified API for calling open models hosted by many inference providers.
Mentioned in: Free GLM-5.2 via Hugging Face Inference Providers in coding agents (Melvin Vivas on X · notes) - Hugging Face Jobs · tool · huggingface.co · paid
Hugging Face service for running compute jobs on HF infrastructure, and agents can launch them.
Mentioned in: Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes) - Inco Splash · tool · github.com · free
Open-source LLM inference engine built for Qwen models on Apple silicon, aimed at fast local decoding.
Mentioned in: Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes) - Infron · tool · infron.ai · free
A model-hosting platform that is currently offering Qwen3.8-27B for free.
Mentioned in: Qwen3.8-27B Is Free on Infron: 256K-Context Multimodal Model (Melvin Vivas on X · notes) - Jev on Vercel AI Gateway · website · vercel.com · check price
Vercel AI Gateway's model page for TypeSafe AI's Jev model.
Mentioned in: Jev was free on Vercel AI Gateway until Sept 25 (Melvin Vivas on X · notes) - LFM2.5-2.6B GGUF on Hugging Face (Liquid AI) · tool · huggingface.co · free · open in a browser to verify
Liquid AI's official GGUF files for the LFM2.5-2.6B small model.
Mentioned in: Running Liquid AI LFM2.5-2.6B Locally with llama-server (Melvin Vivas on X · notes) - LFM2.5-Encoder-230M · tool · huggingface.co · free
A small bidirectional encoder model from the LFM2.5 family. It stays fast at long context on CPU and can be used for zero-shot classification and prompt routing.
Mentioned in: Zero-Shot Prompt Routing by Task Complexity with LFM2.5-Encoder (Melvin Vivas on X · notes) - Liquid4All Cookbook – audio-transcription-cli example · repo · github.com · free
Open-source example code for a fully local audio-to-text CLI that serves LFM2.5-Audio-1.5B with llama.cpp.
Mentioned in: Local Speech-to-Text with LFM2.5-Audio-1.5B and llama.cpp (Melvin Vivas on X · notes) - llama-cpp-python · tool · github.com · free
Python bindings for llama.cpp for running quantized GGUF models locally.
Mentioned in: Run LFM2.5-2.6B locally with llama-cpp-python in Colab (Melvin Vivas on X · notes) - llama-swap (original by mostlygeek) · repo · github.com · free
The original open-source llama-swap proxy for swapping models on demand on local LLM servers.
Mentioned in: Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code (Melvin Vivas on X · notes) - LLM Engineer's Handbook · book · pauliusztin.ai · paid
A book by Paul Iusztin and Maxime Labonne on building production LLM systems from start to finish, including RAG, fine-tuning and LLMOps.
Mentioned in: 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes) - LM Link (LM Studio) · tool · lmstudio.ai · check price
An LM Studio feature for using your local models remotely from another LM Studio instance.
Mentioned in: LM Studio's LM Link: Use a Local Model on Another Machine (Melvin Vivas on X · notes) - LongCat-Image-Edit · tool · github.com · free
Image-editing diffusion model served through SGLang-Diffusion.
Mentioned in: SGLang v0.5.19 release: new models and beam search (Melvin Vivas on X · notes) - M5 Max MacBook Pro · other · apple.com · paid
Apple laptop with M5 Max chip, used as the hardware for the local inference benchmark.
Mentioned in: Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes) - Modal Auto Endpoints · tool · modal.com · paid
Modal feature for spinning up inference endpoints for models you choose.
Mentioned in: Self-Hosting GLM 5.2 with Modal Auto Endpoints (Melvin Vivas on X · notes) - Muse Glimmer 30B GGUF on Hugging Face (Unsloth) · tool · huggingface.co · free · open in a browser to verify
Unsloth's GGUF quantizations of Muse Glimmer 30B for running locally.
Mentioned in: Muse Glimmer 30B: Unsloth GGUF Release and Run Guide (Melvin Vivas on X · notes) - Namespace · tool · x.com · check price
Cloud infrastructure provider whose offering includes on-demand macOS instances.
Mentioned in: Namespace macOS Instances as Computers for AI Agents (Melvin Vivas on X · notes) - Novita AI (@novita_labs) on X · tool · x.com · check price
Novita AI's X account; the company is a cloud provider offering GPU instances and model APIs.
Mentioned in: Novita AI Spot GPU Instances: Cheap GPU Compute (Melvin Vivas on X · notes) - NVIDIA blog: Dynamo AIPerf · article · nvda.ws · free
NVIDIA's blog post explaining how to load-test LLM endpoints with AIPerf.
Mentioned in: Benchmarking LLM Endpoints with NVIDIA Dynamo AIPerf (Melvin Vivas on X · notes) - NVIDIA Dynamo · repo · github.com · free
Open-source framework for distributed, disaggregated LLM inference serving.
Mentioned in: Jensen Huang on NVIDIA's Long-Term Commitment to Open Nemotron Models (Melvin Vivas on X · notes) - NVIDIA Dynamo AIPerf · tool · github.com · free
NVIDIA's benchmarking tool for measuring TTFT, ITL, latency and throughput of LLM endpoints under realistic load.
Mentioned in: Benchmarking LLM Endpoints with NVIDIA Dynamo AIPerf (Melvin Vivas on X · notes) - NVIDIA GB200 NVL72 (NVLink 72) · tool · nvidia.com · paid
NVIDIA's rack-scale system with 72 GPUs connected by NVLink.
Mentioned in: Jensen Huang on NVIDIA's Long-Term Commitment to Open Nemotron Models (Melvin Vivas on X · notes) - Ollama (@ollama on X) · tool · x.com · free
A tool for running open LLMs locally on your own machine.
Mentioned in: Coworker: Free Desktop AI Agent App with Cloud or Local Models (Melvin Vivas on X · notes) - ONNX · tool · onnx.ai · free
Open Neural Network Exchange, an open format and runtime ecosystem for portable ML model deployment.
Mentioned in: Building a Local Model Server with ONNX Support (Melvin Vivas on X · notes) - OpenAI Pricing · website · openai.com · free
OpenAI model pricing page (now redirects to ChatGPT pricing).
Mentioned in: Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) (Melvin Vivas on melvinvivas.com · notes) - OpenAI Ultrafast · tool · openai.com · paid
High-speed inference mode running at about 300 tokens/s.
Mentioned in: OpenAI DevDay recap: Dots, GPT-6.1 Sol, Codex Cloud, Agents API (Melvin Vivas on X · notes) - OpenRouter Activity Explorer · tool · openrouter.ai · check price
OpenRouter's real-time usage dashboard with Overview, Trends, Explore and Guardrails tabs for analyzing LLM spend, tokens, caching and security events.
Mentioned in: Tracking LLM Spend, Tokens & Guardrails with OpenRouter's Activity Explorer (Melvin Vivas on X · notes) - OpenRouter API endpoint · docs · openrouter.ai · free
The base URL of OpenRouter's API, used as base_url in the Codex provider config. It is an API endpoint, not a web page.
Mentioned in: Running Codex with DeepSeek V4 Flash through OpenRouter (Melvin Vivas on X · notes) - OpenRouter Discord · community · discord.com · free
OpenRouter's community server, where members get early access to and test new features.
Mentioned in: OpenRouter Adds Video Generation: One API for Veo, Seedance, Wan and Sora (Melvin Vivas on X · notes) - OpenRouter Pareto Code · tool · openrouter.ai · free
Free, experimental router that picks the cheapest coding model meeting your min_coding_score.
Mentioned in: OpenRouter Pareto Code: cost-optimized coding router (Melvin Vivas on X · notes) - OpenTelemetry · tool · opentelemetry.io · free
An open-source observability standard and toolkit for collecting traces, metrics and logs.
Mentioned in: 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes) - Qwen3-TTS · tool · github.com · free
Open text-to-speech model from the Qwen family.
Mentioned in: Cheap TTS serving: Qwen3-TTS on vLLM-Omni at $3 per 1M characters (Melvin Vivas on X · notes) - Qwen3.8-27B-DFlash2-EXL3-5.0bpw · tool · huggingface.co · free
A DFlash2 draft model for Qwen3.8-27B, used for speculative decoding.
Mentioned in: Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context (Melvin Vivas on X · notes) - Railway · tool · railway.com · free
Deploy apps and APIs quickly.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Rayline (@RaylineAI) on X · tool · x.com · check price
A model router that sends each task to the best-suited model.
Mentioned in: Model routing: Rayline picks the best model per task (Melvin Vivas on X · notes) - RedHatAI/Qwen3.8-2.4T-A95B-NVFP4-FP8 · Hugging Face · tool · huggingface.co · free
Hugging Face model card for Red Hat AI's NVFP4/FP8 quantized Qwen3.8 2.4T-A95B model.
Mentioned in: Red Hat's Quantized Qwen3.8-2.4T-A95B Variant (NVFP4/FP8) (Melvin Vivas on X · notes) - Release aibackends v0.7.0 — LFM2.5 Prompt Routing (donvito/aibackends) · repo · github.com · free
GitHub release notes for AIBackends v0.7.0 with LFM2.5 prompt routing.
Mentioned in: AIBackends v0.7.0: Prompt Routing with Liquid AI's LFM2.5 Encoder (Melvin Vivas on X · notes) - Releases · donvito/aibackends · repo · github.com · free
Release notes for AIBackends, the creator's open-source Python library for running local AI backends.
Mentioned in: AIBackends v0.8.1 adds GLiNER2.5-Decide local classification (Melvin Vivas on X · notes) - Render · tool · render.com · free
Simple app and service hosting.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Runpod Blog: Flash is GA · article · runpod.io · free
Runpod's blog post announcing that Flash is generally available.
Mentioned in: Runpod Flash Reaches GA: Deploy AI Workloads from Python (Melvin Vivas on X · notes) - Runpod PyTorch 2.8 + CUDA 12.8 template · tool · runpod.io · check price
Official Runpod pod template with PyTorch 2.8 and CUDA 12.8.
Mentioned in: GPU devbox Docker image with coding agents on Runpod (Melvin Vivas on X · notes) - runpod/flash (GitHub) · repo · github.com · free
Runpod's open-source Python SDK and framework for defining infrastructure and deploying multimodal and distributed AI inference from the terminal.
Mentioned in: Runpod Flash Reaches GA: Deploy AI Workloads from Python (Melvin Vivas on X · notes) - SageAttention · tool · github.com · free
An optimized attention kernel that speeds up inference.
Mentioned in: 2.3x Faster Ideogram 4 in ComfyUI with INT8 and SageAttention (Melvin Vivas on X · notes) - Scaling Pain of Coding Agent Serving (Z.ai blog) · article · z.ai · free
Z.ai's write-up of the lessons it learned debugging GLM-5 while serving coding agents at scale.
Mentioned in: Z.ai's Lessons from Serving GLM-5 for Coding Agents at Scale (Melvin Vivas on X · notes) - SGLang · repo · github.com · free
Open-source, high-performance serving engine for running LLMs and multimodal models.
Mentioned in: SGLang v0.5.19 release: new models and beam search (Melvin Vivas on X · notes) - SGLang-Diffusion · tool · docs.sglang.io · free
SGLang's add-on for serving diffusion and image-generation models.
Mentioned in: SGLang v0.5.19 release: new models and beam search (Melvin Vivas on X · notes) - TensorRT-LLM · repo · github.com · free
NVIDIA's open-source library for optimized LLM inference on NVIDIA GPUs.
Mentioned in: Jensen Huang on NVIDIA's Long-Term Commitment to Open Nemotron Models (Melvin Vivas on X · notes) - Transformers.js · tool · huggingface.co · free
Hugging Face's JavaScript library for running transformer models directly in the browser or in Node.js.
Mentioned in: Real-Time Speech Transcription in the Browser with Voxtral and WebGPU (Melvin Vivas on X · notes) - Unsloth Desktop · tool · unsloth.ai · free
Unsloth's desktop app for running local models and serving them through compatible APIs.
Mentioned in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes) - Unsloth Documentation – DeepSeek V4 guide · docs · unsloth.ai · free
Unsloth's guide to running DeepSeek V4 models locally with quantized GGUFs.
Mentioned in: Running DeepSeek V4 Flash Locally: RAM Needs for 4-bit and 3-bit Quants (Melvin Vivas on X · notes) - Unsloth Documentation: Decision Laya guide · docs · unsloth.ai · free
Unsloth's guide to running and serving Laya Decision models locally.
Mentioned in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes) - Unsloth Qwen-Image-2.1-GGUF (Hugging Face) · repo · huggingface.co · free · open in a browser to verify
Unsloth's GGUF quantizations of the Qwen-Image-2.1 model, for running it locally.
Mentioned in: Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes) - Unsloth Qwen3.8-Flash-Next-GGUF (Hugging Face) · tool · huggingface.co · free
Quantized GGUF weights of Qwen3.8-Flash-Next published by Unsloth on Hugging Face.
Mentioned in: Run Qwen3.8-Flash-Next (125B MoE) Locally with Unsloth GGUFs (Melvin Vivas on X · notes) - unsloth/DeepSeek-V4-Flash-0731-GGUF (Hugging Face) · repo · huggingface.co · free · open in a browser to verify
Hugging Face repository with the GGUF quantized weights of DeepSeek V4 Flash 0731.
Mentioned in: Running DeepSeek V4 Flash Locally: RAM Needs for 4-bit and 3-bit Quants (Melvin Vivas on X · notes) - Vercel changelog: GPT-5.6 Sol is now 50% off at a lower price · article · vercel.com · free
Vercel's changelog post announcing the GPT-5.6 Sol price cuts on AI Gateway.
Mentioned in: GPT-5.6 Sol Price Cuts on Vercel AI Gateway (70% Off Input) (Melvin Vivas on X · notes) - vLLM (@vllm_project) · tool · x.com · free
An open-source, high-throughput engine for serving LLMs, with tool-call and reasoning parsers.
Mentioned in: Serve LFM2.5-2.6B with vLLM and connect it to Hermes (Melvin Vivas on X · notes) - vLLM-Omni · tool · github.com · free
Version of the vLLM inference engine for serving omni-modal models such as TTS.
Mentioned in: Cheap TTS serving: Qwen3-TTS on vLLM-Omni at $3 per 1M characters (Melvin Vivas on X · notes) - Voxtral WebGPU (demo) · tool · huggingface.co · free
A browser demo that transcribes speech in real time on your own device using Voxtral-Mini-4B on WebGPU.
Mentioned in: Real-Time Speech Transcription in the Browser with Voxtral and WebGPU (Melvin Vivas on X · notes) - Weights & Biases · tool · wandb.ai · free
Experiment tracking and dashboards.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Weights & Biases Weave · tool · github.com · free
Listed under platforms that combine tracing and evaluation.
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook - wllama · repo · github.com · free
A WebAssembly/WebGPU binding of llama.cpp for running GGUF models in the browser.
Mentioned in: Running Qwen3.6-27B Fully in the Browser With WebGPU and wllama (Melvin Vivas on X · notes) - Xet (Hugging Face storage backend) · tool · huggingface.co · free
Hugging Face's chunk-based storage backend for large files on the Hub, which makes deduplication possible.
Mentioned in: Hugging Face Cache Deduplication with Xet in huggingface_hub v1.32 (Melvin Vivas on X · notes) - Xuan Son Nguyen (@ngxson) · person · x.com · free
Hugging Face engineer who builds wllama and works on llama.cpp.
Mentioned in: Running Qwen3.6-27B Fully in the Browser With WebGPU and wllama (Melvin Vivas on X · notes)
Build
- Have an agent launch Hugging Face Jobs to run batch workloads automatically. (from Using Hugging Face credits: Jobs, Inference Endpoints and Open Models)
- Build a simple router that sends easy prompts to a cheap open model and hard ones to a frontier model, then compare cost and accuracy. (from FireRouter: cost-aware model routing between open models and Claude Opus)
- Serve Laya locally through Unsloth Desktop and point an existing Jev-based app at the compatible API. (from Run Laya Decision models locally with Unsloth on 4GB RAM)
- Package an AI workflow as an Apache Camel service in a Docker container and deploy it to AWS, Google Cloud or Azure (from Deploying AI workflow integrations with Apache Camel in Docker)
- A device-assistant router that sends simple function calls (weather, timers) to a small model and complex multi-step agentic tasks (trip planning and booking) to a large model. (from Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M)
- A topic router that sends prompts to specialist agents (for example a soccer agent) added at runtime. (from Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M)
- A coding or math prompt router that sends each prompt to the right language-specific or domain-specific handler. (from Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M)
- Build a local router that classifies incoming requests by intent and sentiment, and checks them against a policy, before sending them to the right handler. (from AIBackends v0.8.1 adds GLiNER2.5-Decide local classification)
- Build a router that sends simple classification jobs to GPT-6 Luna and complex reasoning jobs to GPT-6 Sol. (from DigitalOcean Serverless Inference Now Serves OpenAI GPT-6 Models)
- Use an LLM coding assistant to measure and fix performance bottlenecks in one of your web apps. (from How claude.ai was made 3x faster using Claude itself)
- Build a small coding agent, then measure the token savings from selective tool loading and prompt caching against a quality baseline. (from How Cursor cut agent token costs by 7% without losing quality)
- Refactor a local-inference backend so it loads GGUF models through Transformers with ggml kernels. (from Run llama.cpp GGUF Checkpoints in Hugging Face Transformers)
- Build a local model server that hosts both LLMs and ONNX classification models. (from Building a Local Model Server with ONNX Support)
- Add a new model gateway provider (e.g. Vercel AI Gateway) to a multi-provider AI backend API. (from Adding Vercel AI Gateway as a provider in AIBackends with Devin)
- Build an observability dashboard for your own agent's subagents that tracks tokens and estimated cost. (from Agent Monitor: Dashboard for Codex Subagents, Tokens and Cost)
- Build an agent monitor dashboard that reads OpenAI Codex session traces and shows each step, tool call and result on a timeline. (from Concept Demo: An Agent Monitor Built on OpenAI Codex Traces)
- Run a 27B model locally with a 262K context on one consumer GPU, then compare tokens/s with and without MTP drafting and at different KV-cache quantization settings. (from Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s)
- Build a fully local, private audio-to-text transcription CLI using LFM2.5-Audio-1.5B and llama.cpp. (from Local Speech-to-Text with LFM2.5-Audio-1.5B and llama.cpp)
- Benchmark local LLM throughput on your own GPU at different context sizes and KV-cache quantization settings, with and without MTP speculative decoding. (from Long-Context Local LLM on an RTX 3090: .env Config and ~64 tok/s Benchmark)
- Build a Docker image for a GPU cloud that comes with PyTorch, CUDA, llama.cpp and coding agents already installed. (from Docker Template Bundling Coding Agents for GPU Cloud Hosting)
- Build a CRM app with SQLite and a draggable kanban board as a smoke test for coding models (from Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context)
- Self-host a long-context (262K) 27B model on one consumer GPU and benchmark tokens/s with different KV cache quantization and speculative decoding settings. (from Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090)
- Build a fully local, free AI coding assistant: Qwen3.8 27B (Unsloth Q4_K_M GGUF) served by llama.cpp on a single RTX 3090 and driven by the Pi harness. (from Running Qwen3.8 27B Locally on an RTX 3090 with llama.cpp and the Pi Harness)
- Build a cost-aware router that sends simple prompts to a small or free model and hard prompts to a frontier model. (from AIBackends v0.7.0: Prompt Routing with Liquid AI's LFM2.5 Encoder)
- Build a device-assistant router that sends simple function calls (weather, timers) to a small model and multi-step agentic tasks (trip planning and booking) to a larger model. (from Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M)
- Build a code router that sends each prompt to a handler or model for the right programming language. (from Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M)
- Add a new domain agent (for example a soccer agent) by adding its category to the router at runtime. (from Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M)
- Build a local personal assistant: an agent framework connected to a llama.cpp server for everyday non-coding tasks. (from Move Non-Coding Work to Local Models (Hermes + llama.cpp))
- Benchmark the tokens/sec of a local quantized MoE model at different context and output lengths. (from Serving Ornith-1.5-35B-A3B at 128k Context on an RTX 3090 with llama.cpp)
- Benchmark a local MoE model on your own GPU across different context sizes (16k, 64k, 128k) and output lengths, then chart how tok/s changes. (from Running Ornith-1.5-35B-A3B (Q4_K_M) at 128k Context on an RTX 3090 with llama.cpp)
- Build a local web-to-markdown agent using a small quantized model with llama.cpp and Pi. (from Running LFM2.5-2.6B Q4_K_M Locally with llama.cpp and Pi)
- Connect the local llama-server endpoint to an agent harness and test tool use (from Run Qwen3.8-27B locally with llama.cpp (llama-server))
- Run a fully local coding agent on Muse Glimmer 30B with a single consumer GPU. (from Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp)
- Build a fully local agent stack: llama.cpp serving Muse Glimmer 30B on a 24 GB GPU, driven by Hermes, OpenClaw or Pi. (from Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent)
- Run a fully local coding/agent setup: llama.cpp serving a 30B GGUF model on a 24 GB GPU, used as the backend for Hermes Agent, OpenClaw or Pi. (from Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent)
- A device-assistant router that sends simple requests (weather, timers) to a small model and trip planning/booking to a complex agent. (from Zero-Shot Prompt Routing by Task Complexity with LFM2.5-Encoder)
- A domain-agent router, e.g., a soccer agent added as a runtime category to handle World Cup questions. (from Zero-Shot Prompt Routing by Task Complexity with LFM2.5-Encoder)
- A coding router that sends prompts to the agent for the right programming language, or a math/topic classifier router. (from Zero-Shot Prompt Routing by Task Complexity with LFM2.5-Encoder)
- A cost-saving gateway that catches off-topic or low-value requests and sends them to a cheap model or declines them. (from Zero-Shot Prompt Routing by Task Complexity with LFM2.5-Encoder)
- Build a local on-prem machine to run an open model privately. (from Personal AI Computer to Run DeepSeek V4-Flash Locally)
- Build a model router that tracks provider usage limits and fails over automatically: Grok → Codex → OpenRouter. (from Cost-Saving Model Fallback Chain: Grok → Codex → OpenRouter DeepSeek)
- Set up a personal AI assistant agent that runs on a cheap model through OpenRouter, and track its daily cost. (from Running a Personal AI Agent for $0.69/Day with DeepSeek V4 Flash)
- Deploy an open-weight LLM (GLM 5.2) on Modal and compare its cost and latency against a hosted API. (from Self-Hosting GLM 5.2 with Modal Auto Endpoints)
- Benchmark different weight-precision and attention-kernel combinations on your own GPU for a diffusion model. (from 2.3x Faster Ideogram 4 in ComfyUI with INT8 and SageAttention)
- Build a simple router that classifies each task and sends it to the best-suited model. (from Model routing: Rayline picks the best model per task)
- Build a fully local in-browser chat app with wllama and a quantized GGUF model. (from Running Qwen3.6-27B Fully in the Browser With WebGPU and wllama)
- Benchmark Gemma 4 inference three ways: no speculation, native MTP, and DFlash. Measure tokens per second and check that outputs stay the same. (from Speeding Up Gemma 4 Inference: MTP (3x) vs. DFlash Speculative Decoding (6x))
- Self-host an open-weight coding model on a rented GPU and use it as the backend for Claude Code. Compare its quality, speed and cost with hosted models. (from Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code)
- Build a Python library that wraps local AI tasks such as PII redaction and image processing behind one simple API, using open models like Gemma, Qwen and gliner-PII. (from Running AI Locally: Rebuilding AIBackends as a Python Library with Open Models)
- Build a privacy filter that redacts PII locally with gliner-PII before forwarding prompts to a hosted LLM. (from Running AI Locally: Rebuilding AIBackends as a Python Library with Open Models)
- Build an app on a laptop that sends inference requests to a local LLM on a separate GPU PC through LM Link (from LM Studio's LM Link: Use a Local Model on Another Machine)