Topic 3 of 16 in the learning path
LLM Fundamentals
How large language models work: transformers, tokens, model APIs, choosing models.
Reels and posts (149)
- GLiDE by Fastino Labs: A Post-Trainable Reasoning Decision Model · Melvin Vivas, X · 0:38: Melvin Vivas shares Fastino Labs' launch of GLiDE, a "thinking decision model" for decisions that need a lot of reasoning.
- Use GPT-6.1 Sol by Default, Save Astra for Emergencies · Melvin Vivas, X: A model-selection tip: GPT-6.1 Sol is capable enough for everyday use, so save Astra for emergencies only.
- Use Sonnet 5.5 instead of Opus 5.5 for faster video-clipping tasks · Melvin Vivas, X: The creator found that Claude Sonnet 5.5 can handle his 'create clips from a long video' task and runs much faster than Opus 5.5, so he switched models for that task.
- Guide: Building with Claude Sonnet 5.5 · Melvin Vivas, X: The creator shares the official claude.dev blog guide on building with Claude Sonnet 5.5.
- Running LLMs Locally Without an Expensive Rig · Melvin Vivas, X: A beginner path to running LLMs locally on ordinary hardware.
- The Jev model is now on OpenRouter · Melvin Vivas, X: A short news post saying a model called Jev can now be used through OpenRouter.
- How LLMs Work: A Motion-Graphics Explainer Made in One Shot with Claude Opus 5.5 · Melvin Vivas, X · 1:00: Melvin Vivas shares a 1-minute motion-graphics video explaining how large language models work.
- Kev-4B model, an alternative to Jev, now available on OpenRouter · Melvin Vivas, X: The creator shares that Kev-4B, a small 4B-parameter model he calls an alternative to Jev, can now be used through OpenRouter.
- GLiNER decision model demos from Fastino Labs · Melvin Vivas, X: The creator points to demos from Fastino Labs built on their GLiNER-based decision models.
- Intent classification for support using GLiNER2.5-Decide notebook · Melvin Vivas, X: Fastino's GLiNER2.5-Decide runs on CPU or GPU and can classify the intent of customer support messages.
- NeoHorse-Jev-4B: Open 4B Model for Structured Decisions · Melvin Vivas, X: NeoHorse-Jev-4B is a new open-weight 4B model, presented as an alternative to Jev.
- Getting started with GLiNER2.5-Decide in a Colab notebook · Melvin Vivas, X: Melvin Vivas shares a Google Colab notebook for trying Fastino Labs' GLiNER2.5-Decide.
- GLiNER2.5-Decide: A 340M Encoder Model for Deterministic Classification · Melvin Vivas, X: Shares the release of GLiNER2.5-Decide, a 340M-parameter open-weight model built on an encoder.
- GPT-6 Sol vs Opus 5.5 in a Livestream Comparison · Melvin Vivas, X: An informal comparison from a livestream in which Opus 5.5 gave better results than GPT-6 Sol.
- Opus 5.5 vs GPT-6 Sol for Writing: Latent Space AINews Test · Melvin Vivas, X: Melvin Vivas says Opus 5.5 is great at writing and quotes a Latent Space post.
- Grok 4.7 Launch: Official xAI Announcement · Melvin Vivas, X: The post links to xAI's official 'Introducing Grok 4.7' page as the place to learn everything about the new model.
- Running Gemma 4 Models Offline on an iPhone · Melvin Vivas, X: A demo of Google's Gemma 4 models chatting fully on-device on an iPhone with no internet.
- Possible Open-Weights Version of Jev on Hugging Face · Melvin Vivas, X: A quoted post asks whether a model published by 'convaiinnovations' on Hugging Face (the name looks like 'laya') is an open-weights version of Jev.
- Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores · Melvin Vivas, X: Ternary Bonsai 2 27B has been released.
- Jev model from TypeSafe AI for cheap support-ticket triage · Melvin Vivas, X: The creator recommends Jev, a model from TypeSafe AI, for support-ticket triage.
- Union Alpha: Free Stealth Model on OpenRouter · Melvin Vivas, X: The creator shares a new free stealth model, Union Alpha, available on OpenRouter.
- Model Routing for Coding: Estimate Task Difficulty, Then Pick a Model · Melvin Vivas, X: Melvin Vivas suggests using the newly announced Jev model for coding by first estimating how hard a task is and then sending it to a suitable model.
- Spark-X2.5-4B: a small local model with a 1M-token context window · Melvin Vivas, X: The creator points to Spark-X2.5-4B, a trending 4B-parameter model built to run locally.
- Free Qwen-3.8 27B Model via Infron · Melvin Vivas, X: A short pointer saying the Qwen-3.8 27B model is available for free through Infron.
- Qwen3.8-27B Is Free on Infron: 256K-Context Multimodal Model · Melvin Vivas, X: A quoted post announces that Qwen3.8-27B is free to use on the Infron platform.
- DeepSeek 4.1 Flash speed on the official API: about 325 tokens/s · Melvin Vivas, X: The creator reports that DeepSeek 4.1 Flash generates about 325 tokens per second through DeepSeek's official API.
- Read the GPT-6 Astra launch article to learn what the model can do · Melvin Vivas, X: After using GPT-6 Astra for a few days, the creator went back to read OpenAI's long launch article to understand what the model can really do.
- When a higher reasoning level is worth the extra cost · Melvin Vivas, X: Running GPT-6 Astra at a higher reasoning level costs more per task, but it makes sense when your requirements are clear.
- Gemma 4 as a Strong Small Model for Local and On-Device Use · Melvin Vivas, X: A short post saying that Google's Gemma 4 is still one of the strongest small models you can run locally and on-device.
- Running Qwen3.8 27B Locally with llama.cpp for Writing · Melvin Vivas, X: The creator says he likes how Qwen3.8 27B writes and that he runs it locally with llama.cpp.
- Three Local GGUF Models That Fit on an RTX 3090 (24GB) · Melvin Vivas, X: The creator lists three quantized GGUF models he has run on one RTX 3090 with 24GB of VRAM.
- Comparing GPT Models in Codex by Intelligence and Cost per Task · Melvin Vivas, X: The creator shares an Artificial Analysis comparison of every GPT model available in Codex.
- Open-weight Nemotron models for finance and healthcare · Melvin Vivas, X: Fastino and NVIDIA released two open-weight models for specific industries, built on Nemotron 3.5 Lightning: one for finance and one for healthcare.
- Qwen3.8 27B Quantization Benchmark: 4-Bit Is Enough for Agentic Coding · Melvin Vivas, X: Quesma benchmarked quantized versions of Qwen3.8 27B on the agentic coding benchmark Terminal-Bench 2.1.
- Finding Models to Run Locally on Hugging Face · Melvin Vivas, X: Melvin Vivas says his weekend hobby is browsing Hugging Face for models he can run locally.
- Free Nemotron 3.5 Lightning on OpenRouter Supports Thinking · Melvin Vivas, X: Nemotron 3.5 Lightning is available for free on OpenRouter and supports thinking (reasoning) mode.
- Creator's Top 3 Closed Models: Fable 5, GPT 5.6 Sol, Grok 4.6 · Melvin Vivas, X: Melvin Vivas ranks his top three closed models from his own experience.
- Qwen3.8 Model Family Now on Novita AI: Flash, 27B and 2.4T-A95B · Melvin Vivas, X · 0:25: Melvin Vivas shares a quoted post saying three Qwen3.8 models (Qwen3.8-Flash, Qwen3.8-27B and Qwen3.8-2.4T-A95B) are now available on the Novita AI inference platform.
- Qwen3.8-27B Free on Groq at ~450 Tokens per Second · Melvin Vivas, X: Qwen3.8-27B is available for free on GroqCloud and runs at about 450 tokens per second.
- Colab Notebook: Entity Extraction with GLiNER2.5 in aibackends · Melvin Vivas, X: The creator shares a free Google Colab notebook that shows how to use Fastino Labs' GLiNER2.5 model for information extraction through his aibackends library.
- Cohere Parse Beats Frontier LLMs at Receipt Parsing · Melvin Vivas, X: The creator tested Cohere Parse on a hard receipt with many sub-items.
- You don't need frontier LLMs for everything: use SLMs · Melvin Vivas, X: The creator recommends a talk by Rachel Nabors arguing that you don't need the biggest, most advanced (frontier) LLMs for every job.
- Run Qwen3.8-Flash-Next (125B MoE) Locally with Unsloth GGUFs · Melvin Vivas, X: Shares news that Qwen3.8-Flash-Next, a 125B mixture-of-experts model, can now run locally using Unsloth's GGUF quantizations.
- GPT 5.6 Sol (medium) in ChatGPT for planning tasks · Melvin Vivas, X: The creator recommends GPT 5.6 Sol at medium reasoning effort in ChatGPT for planning.
- GLM-5.2 Vision on Baseten: Turning Images into Code · Melvin Vivas, X: Melvin Vivas shares the news that GLM-5.2 now supports vision (image input) in production, available only on Baseten.
- Try the Free Stealth Model Ox Alpha on OpenRouter · Melvin Vivas, X: Ox Alpha, an unreleased 'stealth' model, is free to use on OpenRouter for now, and people say it is good.
- GPT Image 2 Adds Transparent Background Support in the OpenAI API · Melvin Vivas, X · 0:28: This is a short announcement.
- Ornith-1.5: Open-Source LLM Family (9B Dense to 397B MoE) · Melvin Vivas, X: Ornith-1.5 is a new family of open-source LLMs in three sizes: 9B dense, 35B MoE and 397B MoE.
- Running Hermes on DeepSeek V4 Pro via Baseten Model APIs · Melvin Vivas, X: The creator used leftover Baseten Model API credits to run his Hermes setup on DeepSeek V4 Pro 0813, and says he likes it.
- Find Discounted Models with OpenRouter's New Filter · Melvin Vivas, X: OpenRouter now has a filter that lists all discounted models.
- Cost-Aware Model Routing in Codex: Default to Luna, Escalate to Sol · Melvin Vivas, X: Compares two GPT-5.6 variants in Codex: Luna scores about 52 on intelligence at about $0.05 per task, and Sol scores about 61 at about $1.30 per task.
- Qwen 3.8 27B Matches GPT 5.6 Luna (max) on Artificial Analysis · Melvin Vivas, X: On Artificial Analysis's intelligence comparison, Qwen 3.8 27B, a fairly small open-weight model, scores about the same as GPT 5.6 Luna at max reasoning.
- Run LFM2.5-2.6B locally with llama-cpp-python in Colab · Melvin Vivas, X: The creator shares a Colab notebook that runs the small LFM2.5-2.6B model with llama-cpp-python.
- GLM-5.3 free weekend trial announcement · Melvin Vivas, X: A heads-up about a free weekend trial of the GLM-5.3 model that started at 00:00 Beijing Time (UTC+8) on August 16.
- Swapping AI SDK for pi-ai as the LLM provider layer · Melvin Vivas, X: The creator replaced the AI SDK in his own AI library, which connects to multiple AI providers, with pi-ai from Pi (@pidotdev).
- Open Models Aren't Always Local Models · Melvin Vivas, X: The post separates 'open' (downloadable weights) from 'local' (small enough to run on your own computer).
- Using LFM2.5-VL-3B's Vision Capabilities (Liquid AI Guide) · Melvin Vivas, X: Points to Liquid AI's guide for its small vision-language model, LFM2.5-VL-3B.
- Cohere's North Micro Vision: a small open-source vision model for documents · Melvin Vivas, X: Cohere released North Micro Vision, its smallest vision-language model so far.
- Quick test of Liquid AI's LFM2.5-VL-3B vision model · Melvin Vivas, X: The creator shares a quick test of LFM2.5-VL-3B, a new lightweight vision-language model from Liquid AI.
- LFM2.5-VL-3B: a lightweight VLM for screens, documents and tool calls · Melvin Vivas, X: Liquid AI released LFM2.5-VL-3B, a lightweight vision-language model.
- NVIDIA Nemotron 3.5 Lightning: Fast Open MoE Model for Agents · Melvin Vivas, X: The creator reacts to NVIDIA's launch of Nemotron 3.5 Lightning.
- Claude Sonnet 5 Pricing Made Permanent ($2/$10 per M Tokens) · Melvin Vivas, X: The creator jokes that everyone is using Opus 4.8 while quoting Anthropic's announcement.
- Running Liquid AI Models On-Device with the Apollo iPhone App · Melvin Vivas, X: The creator tries Liquid AI's models directly on his iPhone through the Liquid Apollo app.
- Muse Glimmer 30B: Unsloth GGUF Release and Run Guide · Melvin Vivas, X: Meta released Muse Glimmer, a 30B open model under the Apache 2.0 license.
- Transferable KV Cache: Reusing One Model's Cache in Another · Melvin Vivas, X: Melvin highlights NVIDIA research, shared by Avi Chawla, showing that a KV cache can be moved from one model to another.
- DeepSeek V4 Flash at 90% Off on Nous Portal · Melvin Vivas, X: A heads-up that DeepSeek V4 Flash (0731) has a 90% discount on Nous Portal for about one more week.
- Qwen3.8-Max on the Frontend Code Arena cost-performance frontier · Melvin Vivas, X: The creator shares a quoted Arena post saying Qwen3.8-Max changed the cost-performance Pareto frontier in Frontend Code Arena.
- Running a 28.9M-Parameter LLM on an $8 ESP32 Microcontroller · Melvin Vivas, X: A developer used a trick borrowed from Google's Gemma models to fit a 28.9M-parameter language model on an $8 ESP32 microcontroller.
- GPT 5.6 Luna: Good, Fast and Cheap · Melvin Vivas, X: The creator says GPT 5.6 Luna is good, fast and low-cost.
- Luna Model Gets a Price Cut · Melvin Vivas, X: The creator agrees that Luna is a great low-cost model and says a recent price cut makes it even better value.
- GPT 5.6 Luna Is Enough for ChatGPT Chrome Extension Tasks · Melvin Vivas, X: The creator reports that the cheaper GPT 5.6 Luna model is good enough for tasks in the ChatGPT Chrome extension.
- DeepSeek V4 Flash 0731 Available on OpenRouter · Melvin Vivas, X: The new DeepSeek V4 Flash 0731 release is now available on OpenRouter.
- 1-bit Kimi K3 GGUF Running Locally vs Claude Opus 5 and GPT 5.6 · Melvin Vivas, X: Unsloth released a 1-bit quantized GGUF of Kimi K3 and compared it with Claude Opus 5 and GPT 5.6 on the same creative coding prompt.
- Opus 5 vs Fable 5: Cost per Task in Cursor Benchmarks · Melvin Vivas, X: Based on Cursor benchmarks, the creator says Claude Opus 5 costs about 50% less than Fable 5, uses fewer tokens per task and has a lower cost per task.
- Opus 5 Beats Fable 5 on Cost per Task and Tokens · Melvin Vivas, X: The creator points to a cost-per-task comparison showing Claude Opus 5 is cheaper and uses fewer tokens than Fable 5.
- Use NVIDIA Nemotron 3 Ultra Free on OpenRouter · Melvin Vivas, X: NVIDIA's Nemotron 3 Ultra model can be used for free through OpenRouter.
- Try NVIDIA Nemotron 3 for Free on OpenRouter · Melvin Vivas, X: A short tip that NVIDIA's Nemotron 3 model can be used for free on OpenRouter.
- Hy3 (Tencent Hunyuan) Free on Nous Portal for a Limited Week · Melvin Vivas, X · 0:24: This short post says that Tencent Hunyuan's Hy3 model is free on Nous Research's Nous Portal for one more week.
- Gemini 3.6 Flash and 3.5 Flash-Lite Released for Agents · Melvin Vivas, X: Google released new Gemini models aimed at faster, cheaper agents.
- Kimi K3 ranks #1 on Design Arena for frontend building · Melvin Vivas, X: The creator calls Kimi K3 the new 'design king' because it ranked #1 on Design Arena.
- Grok 4.5 model card released · Melvin Vivas, X: The post shares that the Model Card for Grok 4.5 has been released.
- Comparing Frontier Model API Prices: Grok 4.5, GPT 5.6, Opus 4.8, Fable 5 · Melvin Vivas, X: Lists API prices per million tokens for four frontier models.
- GLM 5.2 Speed vs Opus 4.8 and GPT 5.5 · Melvin Vivas, X: The creator reacts to how fast GLM 5.2 is, quoting his own comparison post.
- GLM 5.2 on Fireworks Makes Opus 4.8 and GPT 5.5 Feel Slow · Melvin Vivas, X: The creator compared GLM 5.2 served by Fireworks AI with Opus 4.8 and GPT 5.5 and found GLM 5.2 far faster.
- Compare GLM-5.2 API Providers on Artificial Analysis · Melvin Vivas, X: Gives Artificial Analysis as the source for GLM-5.2 provider benchmarks.
- NVIDIA's Quantized GLM-5.2 Released on Hugging Face (MIT License) · Melvin Vivas, X: NVIDIA released a quantized GLM-5.2 checkpoint on Hugging Face under the MIT License.
- Model card: NVIDIA's NVFP4-quantized GLM-5.2 on Hugging Face · Melvin Vivas, X: The post links to the Hugging Face model card for nvidia/GLM-5.2-NVFP4, a version of GLM-5.2 published by NVIDIA in the NVFP4 low-precision format.
- Ornith-1.0: open-source LLM family for agentic coding · Melvin Vivas, X: Ornith-1.0 is a new family of open-source LLMs built for agentic coding.
- Sakana Fugu model now available on OpenRouter · Melvin Vivas, X: A short news post saying Sakana's Fugu model is now on OpenRouter.
- How GPT Works: From Token Embeddings to Multi-Head Attention · Melvin Vivas, X · 15:11: This video walks through how a decoder-only GPT is put together, using a Galton board as the analogy for predicting the next token.
- GLM-5.2 on Fireworks: Top Open-Weights Model on GDPval-AA · Melvin Vivas, X: Points to Fireworks as a provider for GLM-5.2.
- Where to Access GLM-5.2: Provider Roundup · Melvin Vivas, X: Lists the providers that offered GLM-5.2 when the post was written: Z.ai's coding plan, Together AI, Baseten, Fireworks, Vercel AI Gateway and OpenRouter.
- Google Interactions API GA: one endpoint for inference and agents · Melvin Vivas, X: Google's Interactions API is now generally available.
- Google Interactions API is GA: recommended for Gemini · Melvin Vivas, X: Google's Interactions API is now generally available.
- GLM 5.2 open model now deployable on Google Cloud · Melvin Vivas, X: Z.ai's open GLM 5.2 model, a strong coding model for long-running agents, can now be self-hosted on Google Cloud's Agent Platform.
- GLM 5.2 as a Default Open Model for Claude Code · Melvin Vivas, X: The creator says GLM 5.2 is now his default open-weight model.
- Use the Right Z.ai Endpoint for the GLM 5.2 Coding Plan · Melvin Vivas, X: This post shares a tip from Ivan Fioravanti: GLM coding plan users should call the coding-optimized endpoint (api.z.ai/api/coding/paas/v4), not the general one (api.z.ai/api/paas/v
- GLM 5.2 Leads Open-Source Models on Datacurve's Deep SWE Leaderboard · Melvin Vivas, X: The creator reports that Z.ai's GLM 5.2 is the top open-source model on the Deep SWE leaderboard from Datacurve.
- GLM 5.2 vs Opus 4.8 for Landing Page Design: 6x Cheaper · Melvin Vivas, X: A side-by-side test built a landing page with GLM 5.2 and with Claude Opus 4.8.
- Free GLM-5.2 via Hugging Face Inference Providers in coding agents · Melvin Vivas, X: For a limited time, Hugging Face made GLM-5.2 free through its Inference Providers API across several providers.
- Gemini 3.5 Live Translate: Real-Time Speech Translation via the Live API · Melvin Vivas, X · 1:40: This is an announcement and demo of Gemini 3.5 Live Translate, which is now available in the Gemini Live API and Google AI Studio.
- GLM-5.2 released: open weights, 1M context, two reasoning levels · Melvin Vivas, X: Z.ai released GLM-5.2, an open-weights frontier model.
- Running Local Models for Agents: Tool Use, Context and Quantization · Melvin Vivas, X · 1:45: Melvin Vivas explains why leaderboard scores don't tell you whether a self-hosted model can work inside an agent workflow like OpenClaw.
- GLM-5.2 by Z.ai: open-weights model with a 1M-token context · Melvin Vivas, X: Z.ai released GLM-5.2, a frontier-level model with open weights.
- Gemini 3.5 Live Translate: Real-Time Speech Translation via the Gemini Live API · Melvin Vivas, X · 1:40: Melvin Vivas says he used Google's live translation to follow a wedding ceremony in Switzerland.
- Free gpt-oss-20b and Gemma 4 26B on OpenRouter · Melvin Vivas, X: OpenRouter added free capacity for two open-weight models, gpt-oss-20b and Gemma 4 26B.
- Open-weight model updates: Kimi 2.7 and GLM 5.2 · Melvin Vivas, X: The creator asks what's new in open models and names Kimi 2.7 and GLM 5.2 as recent releases.
- Claude support for Apple's Foundation Models framework · Melvin Vivas, X: The creator reacts to news that Claude now supports Apple's Foundation Models framework.
- DiffusionGemma: Google's Experimental Diffusion-Based Text Model · Melvin Vivas, X: Melvin Vivas shares Google's announcement of DiffusionGemma, an experimental open model under the Apache 2.0 license.
- Run Gemma 4 12B locally with LM Studio · Melvin Vivas, X: The creator recommends LM Studio as the easiest way to run Google's Gemma 4 12B.
- Run Gemma 4 12B on 8GB RAM with Unsloth Dynamic GGUFs · Melvin Vivas, X: Unsloth's Dynamic GGUF quantizations let Gemma 4 12B run locally on only 8GB of RAM.
- Gemma 4 12B: encoder-free multimodal model for laptops · Melvin Vivas, X: A reaction to Google's announcement of Gemma 4 12B, a unified, encoder-free multimodal model made to run on laptops.
- Claude Opus 4.8 System Card (official PDF) · Melvin Vivas, X: Links to Anthropic's official system card for Claude Opus 4.8, which the creator calls everything you need to know about the new model.
- Claude Opus 4.8 on OpenRouter: Pricing and Fast Mode · Melvin Vivas, X: Claude Opus 4.8 is now on OpenRouter at the same price as Opus 4.7.
- Claude Opus 4.8 Launch: What's New · Melvin Vivas, X: Shares Anthropic's announcement of Claude Opus 4.8.
- DeepSeek makes its DeepSeek-V4-Pro discount permanent · Melvin Vivas, X: A reaction to DeepSeek announcing that its discounted API pricing for DeepSeek-V4-Pro is now permanent.
- Qwen3.7-Max: Qwen's flagship model for agents · Melvin Vivas, X: A reaction to Alibaba's Qwen3.7-Max launch.
- MiniCPM-V 4.6 1.3B: Small Open-Source Vision/OCR Model for Edge Devices · Melvin Vivas, X: The creator shares the open-source release of MiniCPM-V 4.6 (1.3B parameters), the smallest model in the MiniCPM-V family.
- Enable Gemma 4 MTP Speculative Decoding on iPhone · Melvin Vivas, X: Explains how to get faster Gemma 4 inference on an iPhone: turn on speculative decoding in the app's settings.
- Gemma 4 Gets Up to 3x Faster with MTP Drafters · Melvin Vivas, X: Gemma 4 now has Multi-Token Prediction (MTP) drafters for speculative decoding.
- Hy-MT1.5-1.8B-1.25bit: A 440MB Offline Phone Translation Model · Melvin Vivas, X: An open-sourced 1.8B-parameter translation model, quantized to 1.25 bits so it takes only 440MB, runs fully offline on a phone.
- Run Qwen3.6-27B Locally in 18GB RAM with Unsloth GGUFs · Melvin Vivas, X: Qwen3.6-27B can run locally in about 18GB of RAM using Unsloth Dynamic GGUF quantizations.
- Qwen3.6-27B: Dense Open Model for Local Agentic Coding · Melvin Vivas, X: The creator calls Qwen3.6-27B a must-try for coding locally.
- Gemma 4 Now Available Through the Gemini API (Dev Use Only) · Melvin Vivas, X: Google's open Gemma 4 models can now be called through the Gemini API.
- Run Gemma 4 Offline on an iPhone with the Locally AI App · Melvin Vivas, X · 0:43: This short demo shows how to run Google's open Gemma 4 model on an iPhone for free with the Locally AI app.
- OpenRouter Adds Video Generation: One API for Veo, Seedance, Wan and Sora · Melvin Vivas, X · 3:27: Melvin Vivas shares OpenRouter's launch video for video generation.
- Running Gemma 4 E2B Video Understanding Locally in WSL · Melvin Vivas, X: The creator reports getting video understanding working locally with Google's small Gemma 4 E2B model on Windows Subsystem for Linux (WSL).
- Using Gemma 4 via the Gemini API and Google AI Studio · Melvin Vivas, X: Gemma 4 is now available through the official Gemini API and Google AI Studio.
- Run Gemma 4 on Your Phone with Google AI Edge Gallery · Melvin Vivas, X: Shows that you can run Google's open Gemma 4 model directly on an iPhone or Android phone using the Google AI Edge Gallery app.
- Gemma 4 E4B: Local Image Understanding at 131K Context in 6GB VRAM · Melvin Vivas, X: The creator reports running Google's Gemma 4 E4B model locally for image understanding.
- A Visual Guide to Gemma 4 Architecture (Maarten Grootendorst) · Melvin Vivas, X: The creator recommends Maarten Grootendorst's illustrated guide to how the Gemma 4 models are built.
- Gemma 4 Architecture: MoE, Encoders and Per-Layer Embeddings · Melvin Vivas, X: The creator shares a quoted post about 'A Visual Guide to Gemma 4', which explains the architecture of Google's new open models.
- Local Audio Transcription with Cohere Transcribe on WebGPU · Melvin Vivas, X: The creator shows Cohere's speech-to-text model running locally in the browser through WebGPU.
- Gemma 4 26B for OCR in LM Studio · Melvin Vivas, X: The creator found Gemma 4 26B very good at OCR (reading text from images) when he ran it locally in LM Studio.
- Gemma 4: Google's Apache 2.0 Open-Weight Models for Local Hardware · Melvin Vivas, X: This post shares Google's Gemma 4 launch.
- Qwen 3.6 Plus Preview Is Free for a Limited Time on OpenRouter · Melvin Vivas, X: Alibaba's Qwen 3.6 Plus Preview is free on OpenRouter for a limited time.
- 1-bit Bonsai 8B Runs On-Device on iPhone at 40+ tok/s · Melvin Vivas, X: A demo shows PrismML's 1-bit Bonsai 8B running locally on an iPhone 17 Pro at more than 40 tokens per second, which the original poster calls a first for a dense 8B model on iPhone
- MiniMax 2.7 One-Shots a Linear Clone at 95% Lower Cost · Melvin Vivas, X: Sherry Jiang used MiniMax 2.7 to build a Linear clone in one shot in about 10 minutes.
- Qwen3.5 0.8B Does Real-Time Local Video Captioning · Melvin Vivas, X: A roughly 1GB Qwen3.5 0.8B model captions video in real time on a Mac Studio M2 Ultra.
- Qwen3.5 0.8B Real-Time Video Captioning on Mac Studio · Melvin Vivas, X: This reshares the same demo: Qwen3.5 0.8B (about 1GB) captions video locally on a Mac Studio M2 Ultra at under 1s per frame.
- How to Get a GLM-5-Turbo Rate Limit Increase for OpenClaw · Melvin Vivas, X: OpenClaw power users who need a higher GLM-5-Turbo rate limit can DM Lou (@louszbd) with their User ID.
- GLM-5-Turbo: A Fast GLM-5 Variant for Agents Like OpenClaw · Melvin Vivas, X: Z.ai released GLM-5-Turbo, a faster version of GLM-5 tuned for agent environments such as OpenClaw.
- Real-Time Speech Transcription in the Browser with Voxtral and WebGPU · Melvin Vivas, X · 0:39: A short demo of real-time speech-to-text running entirely in the browser.
- NVIDIA build.nvidia.com: Try Hosted NIM Model APIs · Melvin Vivas, X: A short post pointing to build.nvidia.com, NVIDIA's catalog where you can try hosted models through NVIDIA NIM APIs.
- Qwen 3.5 2B Runs On-Device on iPhone with MLX · Melvin Vivas, X: Alibaba's Qwen 3.5 runs locally on an iPhone 17 Pro as a 2B model at 6-bit quantization, using MLX optimized for Apple Silicon.
- Qwen 3.5 One-Shots a Python Stock Analyzer App with yfinance · Melvin Vivas, X: The creator reports that Qwen 3.5 built a working Python stock analyzer app with the yfinance library from a single prompt.
- GPT-5.3-Codex Now Available in OpenAI's Responses API · Melvin Vivas, X: OpenAI made GPT-5.3-Codex available to all developers through the Responses API.
- Local Qwen3.5-35B-A3B on a 24GB GPU Builds a Full Game From One Spec · Melvin Vivas, X: The creator shares a quoted demo in which Qwen3.5-35B-A3B, running locally on a 24GB VRAM GPU, generated a complete game (10 files, 3,483 lines of code) from a single detailed spec
- Qwen3.5 Multimodal Model Now on NVIDIA AI Endpoints · Melvin Vivas, X: Alibaba's Qwen3.5 is now available through NVIDIA AI endpoints.
Watch, free (5)
- Hugging Face LLM Course · course · huggingface.co · free
Hands-on course on transformers, tokenizers and fine-tuning.
Mentioned in: How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes), AI Engineer Roadmap for 2026 in 60 Seconds (Bashiri Smith on Facebook · notes)
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - [1hr Talk] Intro to Large Language Models (Andrej Karpathy) · video · youtube.com · free
A video by Andrej Karpathy that introduces how large language models work.
Mentioned in: 6 Free Videos to Move from Software Engineer to AI Engineer (Bashiri Smith on Facebook · notes), How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes) - Building Systems with the ChatGPT API · course · deeplearning.ai · free
A DeepLearning.AI course on building with LLM APIs. Free with a DeepLearning.AI account during its platform beta; certificates are paid.
Mentioned in: How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes) - OpenAI with Python: A Step-by-Step Guide for Beginners (George Shakan) · video · youtube.com · free
A video tutorial on starting to build with the OpenAI API (creator and title not shown).
Mentioned in: How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes) - X post by DonvitoAI: animation of how an LLM works · video · x.com · free
The linked X post with the short animation that shows how an LLM works.
Mentioned in: Claude Opus 4.8 Generates an Animation of How an LLM Works in One Shot (Melvin Vivas on X · notes)
Read and use (218)
- Claude Opus 5.5 · tool · claude.ai · paid · open in a browser to verify · recommended by both Bashiri Smith & Melvin Vivas
Anthropic's AI assistant, used throughout the guide to tailor resumes, add live roles to the tracker, match connections to target companies and find hiring managers.
Mentioned in: Generating a Repo Promo Video with a Claude Skill on Sonnet 5.5 vs Opus 5.5 (Melvin Vivas on X · notes), Comparing Coding Models on the Same Task in Devin iOS (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), AI Engineer Roadmap Overview: From ML Foundations to RAG, Agents & Ops (Bashiri Smith on Facebook · notes) and 54 more
In a shared PDF: The Interview Engine: The 7-Step System to Land AI Engineering Interviews - shared in this reel on Facebook - GPT-6 Astra · tool · openai.com · paid
The model announced in the quoted launch post, pitched as the developer's most capable model for work, coding, science and cybersecurity, and able to operate a computer.
Mentioned in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), Pi Agent Council: Ask Multiple LLMs in Parallel and Compare Their Advice (Melvin Vivas on X · notes), Use GPT-6.1 Sol by Default, Save Astra for Emergencies (Melvin Vivas on X · notes), Dots in ChatGPT: always-on AI agents that you hand responsibilities to (Melvin Vivas on X · notes) and 50 more - Hugging Face · website · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
Platform for hosting and finding ML models, datasets and papers. The quoted post says LocateAnything was trending there.
Mentioned in: Using an ML agent to train an open-source TTS model on your voice (Melvin Vivas on X · notes), Deploy Open-Source Models with Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes) and 39 more - GLM 5.2 · tool · huggingface.co · free
A GLM-family LLM (the transcript says 'GLM-5-2') that the demo ranked as a strong, affordable model for coding and design.
Mentioned in: Qwen3.8-Max on the Frontend Code Arena cost-performance frontier (Melvin Vivas on X · notes), Models That Work With the Hermes Agent for Personal Productivity (Melvin Vivas on X · notes), Open Model GLM 5.2 Helped Mitigate OpenAI-Caused Cyberattack (Melvin Vivas on X · notes), AI News Roundup: Claude Fable 5, Scientist AI, ZCode, NVIDIA RL (Melvin Vivas on X · notes) and 28 more - Gemma 4 · tool · ai.google.dev · free
Google's family of open-weight models in several sizes, built to run on devices and offline, with multimodal and agentic abilities, and open to fine-tuning.
Mentioned in: Gemma 4 Runs Locally On-Device in the Antigravity SDK (Melvin Vivas on X · notes), On-Device AI: Running Gemma 4 E2B Offline on an iPhone with LiteRT (Melvin Vivas on X · notes), Running Gemma 4 Models Offline on an iPhone (Melvin Vivas on X · notes), Fine-tuning Gemma4-E2B on your own tweet style with Unsloth (Melvin Vivas on X · notes) and 26 more - GPT 5.6 Luna · tool · openai.com · paid
The OpenAI model used inside Codex for the demo. The transcript gives the variant name as 'Soul', which is unclear.
Mentioned in: Set Codex subagent model and reasoning to save usage limits (Melvin Vivas on X · notes), Use GPT-5.6 Luna in Codex for Terminal Tasks (Melvin Vivas on X · notes), Match Reasoning Effort to Task Length in Codex (Astra/Sol) (Melvin Vivas on X · notes), Run Coworker desktop agents cheaply with GPT-5.6 Luna on OpenRouter (Melvin Vivas on X · notes) and 25 more - Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Mentioned in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 25 more - Luna · tool · openai.com · paid
The other model/agent the classifier routes tasks to; the post doesn't describe it further.
Mentioned in: GPT-6.1 Sol May Beat Luna for Subagents (Melvin Vivas on X · notes), Picking models for orchestrator and subagent roles in Codex (Melvin Vivas on X · notes), GPT-6 Sol Ultra Subagents Use Up Limits Fast (Melvin Vivas on X · notes), GPT-6 Sol vs Opus 5.5 in a Livestream Comparison (Melvin Vivas on X · notes) and 15 more - GPT 5.6 Sol · tool · openai.com · paid
A GPT-family model available through an API, whose API and credit pricing was cut by over 20% for three months.
Mentioned in: Creator's Top 3 Closed Models: Fable 5, GPT 5.6 Sol, Grok 4.6 (Melvin Vivas on X · notes), Coworker v0.3.1: Telegram Streaming Sync, Built Fast with Codex (Melvin Vivas on X · notes), Coworker: Open-Source Grok Bot Clone Built on Pi and CopilotKit (Melvin Vivas on X · notes), GPT 5.6 Sol (medium) in ChatGPT for planning tasks (Melvin Vivas on X · notes) and 13 more - Jev · tool · typesafe.ai · paid
A decision model that makes turn-level forecasts on AI agent conversations using only structural signals (turns, tool calls, workflow stages, timing).
Mentioned in: Generating a Repo Promo Video with a Claude Skill on Sonnet 5.5 vs Opus 5.5 (Melvin Vivas on X · notes), GLiDE by Fastino Labs: A Post-Trainable Reasoning Decision Model (Melvin Vivas on X · notes), Using Jev as a Reranker to Augment RAG Retrieval (Melvin Vivas on X · notes), Using Jev (TypeSafe AI) as a Decision Model for Email Classification (Melvin Vivas on X · notes) and 13 more - Fable · tool · anthropic.com · paid
Named as what the creator used with Devin for this build; the post gives no details about what it is.
Mentioned in: Demo: One-Shotting a Flappy Bird iPhone App with Devin and Fable 5.1 (Melvin Vivas on X · notes), Subagents with mixed models: a strong planner and a fast executor (Melvin Vivas on X · notes), Cognition's SWE-2 Coding Model Now in Devin (Melvin Vivas on X · notes), GPT-6 Astra Access Across Plans vs Fable on Claude Max (Melvin Vivas on X · notes) and 11 more - GPT-6.1 Sol · tool · openai.com · paid
A new OpenAI model announced at DevDay 2026 (name as written in the machine transcript), offered with Fast and Ultra fast speed tiers.
Mentioned in: JevDev: Open-Source UI Tool for Experimenting with Jev (Typesafe.ai) (Melvin Vivas on X · notes), Comparing Coding Models on the Same Task in Devin iOS (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), Pi Agent Council: Ask Multiple LLMs in Parallel and Compare Their Advice (Melvin Vivas on X · notes) and 10 more - Grok 4.5 · tool · x.ai · paid
An LLM said to be trained in partnership with SpaceXAI, pitched as a general model beyond software engineering.
Mentioned in: Grok 4.5 Works Well as the Model Behind Hermes Agent (Melvin Vivas on X · notes), Models That Work With the Hermes Agent for Personal Productivity (Melvin Vivas on X · notes), Agent-Made Video in 10 Minutes: Hermes Agent + Grok 4.5 + Hyperframes (Melvin Vivas on X · notes), Personal Assistant Agent on Hermes: Morning Briefings and Inbox Triage (Melvin Vivas on X · notes) and 10 more - LFM2.5-2.6B · tool · huggingface.co · free
Liquid AI's small language model, which can be paired with the LFM2.5-VL-3B vision model.
Mentioned in: Coworker: Open-Source Work Agent That Runs on Small Local Models (Melvin Vivas on X · notes), Coworker with Liquid AI LFM2.5-2.6B via LM Studio on a Mac (Melvin Vivas on X · notes), Zero-Cost Coworker Setup: OpenRouter Free Models + Local LFM2.5 (Melvin Vivas on X · notes), Coworker: A Subscription-Free AI Agent on Local LFM2.5-2.6B (Melvin Vivas on X · notes) and 10 more - Fable 5 · tool · anthropic.com · paid
A frontier model used as the comparison point in the Claude Opus 5 announcement; the post gives no other details about it.
Mentioned in: Claude Opus 5.5 in Devin: #1 on FrontierCode 1.1 (Melvin Vivas on X · notes), Creator's Top 3 Closed Models: Fable 5, GPT 5.6 Sol, Grok 4.6 (Melvin Vivas on X · notes), Opus 5 vs Fable 5: Cost per Task in Cursor Benchmarks (Melvin Vivas on X · notes), Opus 5 Beats Fable 5 on Cost per Task and Tokens (Melvin Vivas on X · notes) and 8 more - Qwen 3.5 · tool · qwen.ai · free
Alibaba Qwen open model family with small on-device variants and toggleable reasoning.
Mentioned in: Workshop AI: Building Apps with Cloud and Local Agents (GLM 5, Qwen 3.5) (Melvin Vivas on X · notes), Qwen3.5 0.8B Does Real-Time Local Video Captioning (Melvin Vivas on X · notes), Qwen3.5 0.8B Real-Time Video Captioning on Mac Studio (Melvin Vivas on X · notes), Qwen 3.5 Is the Local Model That Holds Up in Claude Code (Melvin Vivas on X · notes) and 8 more - OpenAI API · tool · platform.openai.com · paid · recommended by both Bashiri Smith & Melvin Vivas
OpenAI API documentation, including function calling and structured outputs.
Mentioned in: How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes), Livestream: Building a React Native ChatGPT App with Cursor and OpenAI (Melvin Vivas on X · notes), GPT-Live-1: OpenAI's Full-Duplex Voice Model for Voice Agents in the API (Melvin Vivas on X · notes), The 3 Levels of AI Engineering: LLM Apps → Production → Agentic Systems (Bashiri Smith on Facebook · notes) and 5 more
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - DeepSeek · tool · deepseek.com · free
Open-weight large language models from DeepSeek, known for strong reasoning and coding at low cost.
Mentioned in: Open models now dominate token volume on Vercel AI Gateway (Melvin Vivas on X · notes), jina-ocr-v1: Turning PDFs, Scans and Tables into Markdown (Melvin Vivas on X · notes), Bolt.new Adds Open Models (GLM, DeepSeek, Kimi) via Bolt Forge (Melvin Vivas on X · notes), OpenCode Go: Open Coding Models (DeepSeek, Qwen, Kimi, GLM) for $10/mo (Melvin Vivas on X · notes) and 5 more - Qwen 3.6 35B (MTP) · tool · huggingface.co · free
Qwen 3.6 mixture-of-experts model (35B total, about 3B active parameters) with multi-token prediction for local inference.
Mentioned in: Use a Local Model for Confidential Data with Your Agent (Melvin Vivas on X · notes), Running Qwen 3.6 35B locally on an RTX 3090 for agent tool calling (Melvin Vivas on X · notes), Models That Work With the Hermes Agent for Personal Productivity (Melvin Vivas on X · notes), Run Hermes Agent locally with Qwen 3.6 35B MTP in LM Studio (Melvin Vivas on X · notes) and 5 more - DeepSeek V4 Flash · tool · huggingface.co · free
DeepSeek model used as the comparison baseline in the tool-calling benchmark.
Mentioned in: Low-Cost Agent Run: DeepSeek V4 Flash via OpenRouter in ohmypi (Melvin Vivas on X · notes), Running Codex with DeepSeek V4 Flash through OpenRouter (Melvin Vivas on X · notes), LFM2.5-2.6B Matches DeepSeek-V4-Flash on Tool Calling; LEAP Fine-Tuning (Melvin Vivas on X · notes), DeepSeek V4 Flash at 90% Off on Nous Portal (Melvin Vivas on X · notes) and 4 more - Claude Opus 4.7 · tool · anthropic.com · paid
Anthropic's top Opus model, aimed at long-running agentic tasks that need precise instruction following.
Mentioned in: Shipping Features Fast in Cursor with Multiple AI Models (Melvin Vivas on X · notes), Composer 2.5 as the Default Coding Model in Cursor (Melvin Vivas on X · notes), Which AI coding model to use for which task (Melvin Vivas on X · notes), Claude Opus 4.7 Fast Mode in Cursor: Speed vs. Cost Trade-off (Melvin Vivas on X · notes) and 3 more - Claude Sonnet 5.5 · tool · anthropic.com · paid
Anthropic's newly released mid-tier Claude model, reported to be strong at agentic coding.
Mentioned in: Generating a Repo Promo Video with a Claude Skill on Sonnet 5.5 vs Opus 5.5 (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), GPT-6.1 and Sonnet 5.5 Released the Same Week (Melvin Vivas on X · notes), Use Sonnet 5.5 instead of Opus 5.5 for faster video-clipping tasks (Melvin Vivas on X · notes) and 3 more - GPT 5.5 · tool · openai.com · paid
OpenAI models the creator used as the coding model inside Cursor.
Mentioned in: Models That Work With the Hermes Agent for Personal Productivity (Melvin Vivas on X · notes), Running GPT-5.5 via Codex as Hermes's Main Model (Melvin Vivas on X · notes), Conductor Walkthrough: Running Parallel Coding Agents in Isolated Git Worktrees (Melvin Vivas on X · notes), Composer 2.5 as the Default Coding Model in Cursor (Melvin Vivas on X · notes) and 3 more - GPT-5.4 · tool · openai.com · paid
OpenAI frontier model (Thinking and Pro variants) for reasoning, coding and agents.
Mentioned in: Opinion: GPT-5.4 in Cursor 3 Replaces Opus for Coding (Melvin Vivas on X · notes), Splitting AI Tools: Cursor 3 + GPT-5.4 for Code, Claude for Learning (Melvin Vivas on X · notes), Creator's Coding Setup: Cursor 3 with GPT 5.4 (Melvin Vivas on X · notes), Open-Source GLM-5.1 Beats GPT-5.4 on SWE-Bench Pro (Melvin Vivas on X · notes) and 3 more - Cohere · person · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
AI company that builds language, embedding and reranking models, and now the Transcribe speech-recognition model.
Mentioned in: Basic RAG Pipeline in 60 Seconds: From Documents to Grounded Answers (Bashiri Smith on Facebook · notes), Cohere Parse Beats Frontier LLMs at Receipt Parsing (Melvin Vivas on X · notes), Cohere Parse: Pricing vs Parse Bench Score (Melvin Vivas on X · notes), Cohere's North Micro Vision: a small open-source vision model for documents (Melvin Vivas on X · notes) and 2 more - DeepSeek V4.1 Flash · tool · api-docs.deepseek.com · paid
A fast DeepSeek language model available through DeepSeek's official API.
Mentioned in: Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), DeepSeek 4.1 Flash speed on the official API: about 325 tokens/s (Melvin Vivas on X · notes), One-Shot Agent Management UI with DeepSeek Harness for $0.19 (Melvin Vivas on X · notes), DeepSeek Harness: Open-Source, Browser-Based Agent for DeepSeek V4.1 (Melvin Vivas on X · notes) and 2 more - Kimi K3 · tool · huggingface.co · free
Moonshot AI's large multimodal LLM with a 1M-token context window and Kimi Delta Attention.
Mentioned in: Multi-Teacher On-Policy Distillation (MOPD) in 2026 (Melvin Vivas on X · notes), Qwen3.8-Max on the Frontend Code Arena cost-performance frontier (Melvin Vivas on X · notes), 1-bit Kimi K3 GGUF Running Locally vs Claude Opus 5 and GPT 5.6 (Melvin Vivas on X · notes), Code Arena Fullstack Benchmark: Kimi K3 Ranks #1 (Melvin Vivas on X · notes) and 2 more - LFM2.5-VL-3B · tool · huggingface.co · free
Liquid AI's lightweight vision-language model for screen and document understanding, grounding and tool calling.
Mentioned in: OCR Testing a Small Vision Model with LLM-Made Ground Truth (Melvin Vivas on X · notes), AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers (Melvin Vivas on X · notes), AIBackends Adds Support for LFM2.5-VL-3B (Melvin Vivas on X · notes), Using LFM2.5-VL-3B's Vision Capabilities (Liquid AI Guide) (Melvin Vivas on X · notes) and 2 more - Qwen3.6-27B · tool · huggingface.co · free
Dense open-weight Qwen model. The MTP GGUF version runs at about 140 tokens/s.
Mentioned in: Meta Muse Glimmer-30B: Open Weights Model That Beats Qwen3.6 37B on Agentic Tasks (Melvin Vivas on X · notes), Running Qwen3.6-27B Fully in the Browser With WebGPU and wllama (Melvin Vivas on X · notes), Unsloth MTP GGUFs make Qwen3.6 run 1.4x faster locally (Melvin Vivas on X · notes), Qwen 3.6 27B: An Open Local Model with Benchmarks Near Claude Opus 4.5 (Melvin Vivas on X · notes) and 2 more - Andrej Karpathy · person · karpathy.ai · free · recommended by both Bashiri Smith & Melvin Vivas
AI researcher and educator known for clear, free lessons on neural networks and LLMs.
Mentioned in: Karpathy's Tip: Ask LLMs to Explain in ASD-STE100 (Melvin Vivas on X · notes), Grok Team Bots: Shared AI Teammates in Slack and Grok (Melvin Vivas on X · notes), 6 Free Videos to Move from Software Engineer to AI Engineer (Bashiri Smith on Facebook · notes), Claude Opus 5.5 Released: Fable 5.1-Level Performance at 40% Lower Cost (Melvin Vivas on X · notes) and 1 more - Claude Fable 5.1 · tool · anthropic.com · paid
An Anthropic Claude model used as the performance reference for Opus 5.5.
Mentioned in: Claude Opus 5.5 Released: Fable 5.1-Level Performance at 40% Lower Cost (Melvin Vivas on X · notes), GPT-6 Astra Tops Vending-Bench, Beating Claude Fable 5.1 (Melvin Vivas on X · notes), Claude Fable 5.1 Runs a 38-Hour Unattended ML Task (Melvin Vivas on X · notes), Claude Fable 5.1 Scores 73.4% on CursorBench 3.2 (Melvin Vivas on X · notes) and 1 more - Claude Opus 4.6 · tool · anthropic.com · paid
Anthropic's Claude Opus model, version 4.6, which Claude Code was running in this comparison.
Mentioned in: GLM-5.1 vs Claude Code (Opus 4.6): One-Shot Three.js Racing Game Eval (Melvin Vivas on X · notes), GLM-5.1 Reportedly Beats Claude Opus 4.6 on Cybersecurity (Melvin Vivas on X · notes), MiniMax 2.7 One-Shots a Linear Clone at 95% Lower Cost (Melvin Vivas on X · notes), When one coding model gets stuck, switch to another (Melvin Vivas on X · notes) and 1 more - GLM-5.1 · tool · huggingface.co · free
Open-source large language model from Z.ai (Zhipu), built for coding and long-running agent tasks.
Mentioned in: GLM-5.1 vs Claude Code (Opus 4.6): One-Shot Three.js Racing Game Eval (Melvin Vivas on X · notes), GLM-5.1 Reportedly Beats Claude Opus 4.6 on Cybersecurity (Melvin Vivas on X · notes), Open-Source GLM-5.1 Beats GPT-5.4 on SWE-Bench Pro (Melvin Vivas on X · notes), GLM-5.1: Open-Source Model for Long-Running Coding Agents (Melvin Vivas on X · notes) and 1 more - Qwen3.8 · tool · qwen.ai · free
Upcoming large Qwen model (2.4T parameters) that Qwen says will be released with open weights.
Mentioned in: SGLang v0.5.19 release: new models and beam search (Melvin Vivas on X · notes), Qwen3.8-Max on the Frontend Code Arena cost-performance frontier (Melvin Vivas on X · notes), Qwen3.8-Max announced with open weights for Max and 27B (Melvin Vivas on X · notes), Qwen3.8 (2.4T params) Announced as Upcoming Open-Weight Model (Melvin Vivas on X · notes) and 1 more - Gemini · tool · gemini.google.com · free
Google's multimodal AI model family. In this prototype it interprets voice, pointer position and on-screen content, and writes code to act on what the user wants.
Mentioned in: Multimodal AI Pointer: Combining Voice, Mouse Pointing and Vision with Gemini (Melvin Vivas on X · notes), Why Claude Works Well as a Brainstorming Partner (Melvin Vivas on X · notes), Workshop AI: Building Apps with Cloud and Local Agents (GLM 5, Qwen 3.5) (Melvin Vivas on X · notes), Building a Nano Banana 2 Image-Gen App with Memex Managed AI Connectors (Melvin Vivas on X · notes) - Gemini Flash · tool · deepmind.google · free
Google's fast, low-cost Gemini model, suited to auxiliary tasks like web browsing and vision.
Mentioned in: Testing Gemini 3.8 Flash in Cursor with a CRM Smoke Test (Melvin Vivas on X · notes), Gemini 3.8 Flash Available in Cursor CLI (Melvin Vivas on X · notes), Models That Work With the Hermes Agent for Personal Productivity (Melvin Vivas on X · notes), Cut Hermes Agent Token Costs with Gemini Flash as an Auxiliary Model (Melvin Vivas on X · notes) - GLiNER2.5-Decide · tool · huggingface.co · free
A 340M-parameter open-weight encoder model that answers user-defined typed questions and rules for fast, deterministic classification.
Mentioned in: AIBackends v0.8.1 adds GLiNER2.5-Decide local classification (Melvin Vivas on X · notes), Intent classification for support using GLiNER2.5-Decide notebook (Melvin Vivas on X · notes), Getting started with GLiNER2.5-Decide in a Colab notebook (Melvin Vivas on X · notes), GLiNER2.5-Decide: A 340M Encoder Model for Deterministic Classification (Melvin Vivas on X · notes) - GLM-5.3 · tool · huggingface.co · free
A large language model in the GLM family (Z.ai / Zhipu), offered free for one weekend.
Mentioned in: GLM 5.3 Open Weights Release Delayed for Framework Support (Melvin Vivas on X · notes), Paperscrolling: Browse Trending AI Research Papers Like a Social Feed (alphaXiv) (Melvin Vivas on X · notes), GLM-5.3 now available on AWS Marketplace (Melvin Vivas on X · notes), GLM-5.3 free weekend trial announcement (Melvin Vivas on X · notes) - GPT-6 Luna · tool · community.openai.com · paid
An OpenAI model aimed at fast, high-volume classification and routing.
Mentioned in: DigitalOcean Serverless Inference Now Serves OpenAI GPT-6 Models (Melvin Vivas on X · notes), Codex Orchestrator v0.2.1: GPT-6 Sol orchestrator with Luna subagents (Melvin Vivas on X · notes), DeepSWE results: GPT-6 Sol slightly below GPT-5.6 Sol, but cheaper (Melvin Vivas on X · notes), OpenAI Releases GPT-6 Sol and GPT-6 Luna (Melvin Vivas on X · notes) - Grok 4.6 · tool · x.ai · paid
xAI's latest frontier LLM, announced as an improvement over Grok 4.5 at the same price.
Mentioned in: Creator's Top 3 Closed Models: Fable 5, GPT 5.6 Sol, Grok 4.6 (Melvin Vivas on X · notes), Hermes (Grok 4.6) vs ChatGPT Work on a text-to-PDF task (Melvin Vivas on X · notes), Plan with a Strong Model, Execute with a Cheaper One (Melvin Vivas on X · notes), Grok 4.6 release: better than Grok 4.5 at the same price (Melvin Vivas on X · notes) - Hugging Face Transformers · tool · github.com · free
Open-source library for loading and running pretrained models, which now supports GGUF directly.
Mentioned in: Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), Run GGUF models directly in Hugging Face Transformers (Melvin Vivas on X · notes), AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers (Melvin Vivas on X · notes), Gemma 4 Gets Up to 3x Faster with MTP Drafters (Melvin Vivas on X · notes) - Kimi · tool · github.com · free
Open-weight large language models from Moonshot AI, aimed at agent and coding tasks.
Mentioned in: Bolt.new Adds Open Models (GLM, DeepSeek, Kimi) via Bolt Forge (Melvin Vivas on X · notes), OpenCode Go: Open Coding Models (DeepSeek, Qwen, Kimi, GLM) for $10/mo (Melvin Vivas on X · notes), Orchestrator + Subagents in opencode with Open Models (Melvin Vivas on X · notes), Why Chinese AI Models Are Mostly Open Source (Melvin Vivas on X · notes) - Muse Spark 1.3 · tool · research.meta.ai · free
Meta's model release, available for free in OpenCode at the time of the post.
Mentioned in: Try Muse Spark 1.3 for Free in OpenCode (Melvin Vivas on X · notes), Be Skeptical of Model Leaderboards: Muse Spark vs Astra (Melvin Vivas on X · notes), Muse Spark 1.3 Is Free to Try in OpenCode (Melvin Vivas on X · notes), Installing Meta's Muse Code CLI (Needs a Paid Plan) (Melvin Vivas on X · notes) - Nous Portal · tool · portal.nousresearch.com · paid
Nous Research's model-access portal, where you can sign up and use hosted models such as Hy3.
Mentioned in: Alternative Coding Plans for Open Models: OpenCode Go, Ollama Cloud, Nous (Melvin Vivas on X · notes), DeepSeek V4 Flash at 90% Off on Nous Portal (Melvin Vivas on X · notes), Connecting Hermes Agent to the Buzz Agent-First Chat App via the Native Gateway (Melvin Vivas on X · notes), Hy3 (Tencent Hunyuan) Free on Nous Portal for a Limited Week (Melvin Vivas on X · notes) - Qwen3.5-2B · tool · huggingface.co · free
Small open-weight Qwen language model, used here as the base for fine-tuning.
Mentioned in: LoRA Fine-Tune Qwen3.5-2B on Your Tweets with Unsloth Studio (Melvin Vivas on X · notes), Post-training Qwen3.5-2B on your own X posts with Unsloth Studio (Melvin Vivas on X · notes), Fine-Tuned Qwen3.5-2B LoRA Model That Writes X Posts in Your Style (Melvin Vivas on X · notes), First LoRA Run on Qwen3.5-2B with Codex as Training Companion (Melvin Vivas on X · notes) - DeepSeek API · tool · platform.deepseek.com · paid
DeepSeek's official API platform for buying credits and managing API keys.
Mentioned in: DeepSeek 4.1 Flash speed on the official API: about 325 tokens/s (Melvin Vivas on X · notes), Getting Started with DeepSeek Harness and the DeepSeek API Platform (Melvin Vivas on X · notes), DeepSeek-V4-Flash Official API Launches in Public Beta (Melvin Vivas on X · notes) - DeepSeek V4 · tool · huggingface.co · free
A DeepSeek LLM, accessed through OpenRouter as the final fallback.
Mentioned in: Multi-Teacher On-Policy Distillation (MOPD) in 2026 (Melvin Vivas on X · notes), Models That Work With the Hermes Agent for Personal Productivity (Melvin Vivas on X · notes), Cost-Saving Model Fallback Chain: Grok → Codex → OpenRouter DeepSeek (Melvin Vivas on X · notes) - DeepSeek V4 Flash 0731 · tool · huggingface.co · free
DeepSeek's open-weight V4 Flash model, released as quantized GGUFs for local use.
Mentioned in: DeepSeek V4 Flash 0731 Available on OpenRouter (Melvin Vivas on X · notes), Running DeepSeek V4 Flash Locally: RAM Needs for 4-bit and 3-bit Quants (Melvin Vivas on X · notes), DeepSeek V4 Flash 0731 Agentic Benchmarks and Use with Hermes Agent (Melvin Vivas on X · notes) - Gemini API · docs · ai.google.dev · free
Google's developer API and docs for calling Gemini and Gemma models.
Mentioned in: Gemini 3.5 Live Translate: Real-Time Speech Translation via the Live API (Melvin Vivas on X · notes), Gemma 4 Now Available Through the Gemini API (Dev Use Only) (Melvin Vivas on X · notes), Using Gemma 4 via the Gemini API and Google AI Studio (Melvin Vivas on X · notes) - Gemma 4 12B · tool · blog.google · free · open in a browser to verify
Google's open-weight, encoder-free multimodal model under Apache 2.0 for on-device use.
Mentioned in: Run Gemma 4 12B locally with LM Studio (Melvin Vivas on X · notes), Run Gemma 4 12B on 8GB RAM with Unsloth Dynamic GGUFs (Melvin Vivas on X · notes), Gemma 4 12B: encoder-free multimodal model for laptops (Melvin Vivas on X · notes) - Gemma 4 26B · tool · ai.google.dev · free
Mid-sized multimodal open model in Google's Gemma 4 family.
Mentioned in: Free gpt-oss-20b and Gemma 4 26B on OpenRouter (Melvin Vivas on X · notes), Local Prompt-to-Image Pipeline: ComfyUI + LM Studio Gemma + Z-Image Turbo (Melvin Vivas on X · notes), Gemma 4 26B for OCR in LM Studio (Melvin Vivas on X · notes) - GLiNER2.5 · tool · fastino.ai · free
Fastino Labs' models for extraction and classification that run on CPU, with long context and joint information extraction.
Mentioned in: Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails (Melvin Vivas on X · notes), Colab Notebook: Entity Extraction with GLiNER2.5 in aibackends (Melvin Vivas on X · notes), AIBackends v0.6.0: Run GLiNER2.5 Extraction on CPU (Melvin Vivas on X · notes) - GLM · tool · github.com · free
Chinese open-weight LLM family from Zhipu AI.
Mentioned in: OpenCode Go: Open Coding Models (DeepSeek, Qwen, Kimi, GLM) for $10/mo (Melvin Vivas on X · notes), Orchestrator + Subagents in opencode with Open Models (Melvin Vivas on X · notes), Why Chinese AI Models Are Mostly Open Source (Melvin Vivas on X · notes) - NVIDIA Nemotron 3.5 · tool · openrouter.ai · free
NVIDIA Nemotron model with thinking support, available free on OpenRouter.
Mentioned in: Coworker: Free Desktop AI Agent App with Cloud or Local Models (Melvin Vivas on X · notes), Free Nemotron 3.5 Lightning on OpenRouter Supports Thinking (Melvin Vivas on X · notes), Using Free OpenRouter Models (Nemotron 3.5) With Coworker (Melvin Vivas on X · notes) - Ornith-1.5-35B-A3B · tool · huggingface.co · free
An open-weight mixture-of-experts language model (35B total, about 3B active parameters), run here as a Q4_K_M GGUF quantization.
Mentioned in: Running a Local Ornith Model as the Backend for a Hermes Agent (Melvin Vivas on X · notes), Serving Ornith-1.5-35B-A3B at 128k Context on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Running Ornith-1.5-35B-A3B (Q4_K_M) at 128k Context on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes) - Qwen · tool · qwen.ai · free
Alibaba's family of open-weight LLMs.
Mentioned in: OpenCode Go: Open Coding Models (DeepSeek, Qwen, Kimi, GLM) for $10/mo (Melvin Vivas on X · notes), Orchestrator + Subagents in opencode with Open Models (Melvin Vivas on X · notes), Why Chinese AI Models Are Mostly Open Source (Melvin Vivas on X · notes) - A Visual Guide to Gemma 4 · article · newsletter.maartengrootendorst.com · free
Illustrated walkthrough of the Gemma 4 architecture: MoE, vision encoder, per-layer embeddings and audio encoder.
Mentioned in: A Visual Guide to Gemma 4 Architecture (Maarten Grootendorst) (Melvin Vivas on X · notes), Gemma 4 Architecture: MoE, Encoders and Per-Layer Embeddings (Melvin Vivas on X · notes) - Artificial Analysis – Comparison of AI Models across Intelligence, Performance, and Price · website · artificialanalysis.ai · free
Benchmark site comparing AI models on an intelligence index, speed and price, here filtered to the GPT models in Codex.
Mentioned in: Comparing GPT Models in Codex by Intelligence and Cost per Task (Melvin Vivas on X · notes), Qwen 3.8 27B Matches GPT 5.6 Luna (max) on Artificial Analysis (Melvin Vivas on X · notes) - DeepSeek V4 Pro 0813 · tool · baseten.co · paid
DeepSeek large language model, served here through Baseten Model APIs.
Mentioned in: Running DeepSeek V4 Pro 0813 on Baseten with the Pi Agent (Melvin Vivas on X · notes), Running Hermes on DeepSeek V4 Pro via Baseten Model APIs (Melvin Vivas on X · notes) - Gemini 3.5 Flash · tool · blog.google · free · open in a browser to verify
Google's fast Gemini model, used as the agent model throughout the demo.
Mentioned in: Gemini 3.5 Flash Released (Melvin Vivas on X · notes), Google Antigravity 2.0: Multi-Agent Desktop App Walkthrough (Subagents, Scheduled Tasks) (Melvin Vivas on X · notes) - Gemini 3.5 Live Translate · tool · blog.google · paid · open in a browser to verify
Google's real-time speech translation model that detects language switches automatically and supports more than 70 languages.
Mentioned in: Gemini 3.5 Live Translate: Real-Time Speech Translation via the Live API (Melvin Vivas on X · notes), Gemini 3.5 Live Translate: Real-Time Speech Translation via the Gemini Live API (Melvin Vivas on X · notes) - Gemini Live API · docs · ai.google.dev · check price
Google's API for real-time, streaming audio sessions with Gemini models, used here for live translation.
Mentioned in: Gemini 3.5 Live Translate: Real-Time Speech Translation via the Live API (Melvin Vivas on X · notes), Gemini 3.5 Live Translate: Real-Time Speech Translation via the Gemini Live API (Melvin Vivas on X · notes) - GLiDE · tool · fastino.ai · check price
Fastino Labs' decision model, which uses adaptive thinking (fast probabilities first, reasoning only when uncertain).
Mentioned in: GLiDE: Fastino's Decision Model with Adaptive Thinking (Melvin Vivas on X · notes), GLiDE by Fastino Labs: A Post-Trainable Reasoning Decision Model (Melvin Vivas on X · notes) - GLiNER2.5 Extraction Colab Notebook · repo · github.com · free
An example notebook that runs GLiNER2.5 information extraction with AIBackends in Colab.
Mentioned in: Colab Notebook: Entity Extraction with GLiNER2.5 in aibackends (Melvin Vivas on X · notes), AIBackends v0.6.0: Run GLiNER2.5 Extraction on CPU (Melvin Vivas on X · notes) - GLM 5.3 Flash · tool · huggingface.co · free
A fast model from Z.ai's GLM family, used here for agentic coding.
Mentioned in: Coworker Desktop App Works with Any OpenAI-Compatible Endpoint (Melvin Vivas on X · notes), Running GLM 5.3 Flash on Baseten with the Pi Coding Agent (Melvin Vivas on X · notes) - GLM-5 · tool · github.com · free
An open-weight large language model from Zhipu AI (Z.ai) that can run locally.
Mentioned in: Z.ai's Lessons from Serving GLM-5 for Coding Agents at Scale (Melvin Vivas on X · notes), Workshop AI: Building Apps with Cloud and Local Agents (GLM 5, Qwen 3.5) (Melvin Vivas on X · notes) - GLM-5-Turbo docs (Z.ai) · docs · docs.z.ai · free
Z.ai's API guide for the GLM-5-Turbo model.
Mentioned in: How to Get a GLM-5-Turbo Rate Limit Increase for OpenClaw (Melvin Vivas on X · notes), GLM-5-Turbo: A Fast GLM-5 Variant for Agents Like OpenClaw (Melvin Vivas on X · notes) - Google Gemma (@googlegemma) · person · x.com · free
Official X account for Google's Gemma open models, with release news and examples.
Mentioned in: Running Gemma 4 Models Offline on an iPhone (Melvin Vivas on X · notes), Running Gemma4-E2B tool calling locally on an iPhone (Melvin Vivas on X · notes) - gpt-3.5-turbo · tool · platform.openai.com · paid
OpenAI chat model; a cheaper alternative to text-davinci-003.
Mentioned in: Natural Language to API Calls with LangChain APIChain and OpenAI (Melvin Vivas on melvinvivas.com · notes), Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) (Melvin Vivas on melvinvivas.com · notes) - GPT-Live · tool · openai.com · free
OpenAI model behind ChatGPT Voice that can speak, listen and coordinate work at the same time.
Mentioned in: ChatGPT Voice on Desktop: Directing Codex and Agents by Voice (Melvin Vivas on X · notes), AI News Roundup: Grok 4.5, GPT-5.6 Launch, GPT-Live Voice Models (Melvin Vivas on X · notes) - GPT-Realtime-Translate · tool · developers.openai.com · paid
OpenAI's realtime model that translates speech live across 70 languages while the speaker is talking.
Mentioned in: GPT-Realtime-2 & Realtime-Translate: Reasoning Voice Agents and Live Translation (Melvin Vivas on X · notes), GPT-Realtime-2 & GPT-Realtime-Translate: Reasoning Voice Agents and Live Translation (Melvin Vivas on X · notes) - Grok 4.7 · tool · x.ai · paid
xAI's large language model, available through the xAI API or an X Premium subscription.
Mentioned in: One-shot Demo: Building a Flyable App with Grok 4.7 in Cursor (Melvin Vivas on X · notes), Ask Your Agent to Configure Itself: Adding Image Generation (Melvin Vivas on X · notes) - LiquidAI/LFM2.5-2.6B-GGUF · tool · huggingface.co · free
The Hugging Face repo with GGUF quantized weights of Liquid AI's LFM2.5-2.6B model.
Mentioned in: Run LFM2.5-2.6B locally with llama.cpp (Melvin Vivas on X · notes), Serving LFM2.5-2.6B with llama-server: full command (Melvin Vivas on X · notes) - Mercury 2.5 · tool · inceptionlabs.ai · paid
A very fast language model, reported at about 1,100 tokens/sec with better agentic performance than Mercury 2.
Mentioned in: Jev + WebMCP Solves 100% of Benchmark Tasks at About 112× Lower Cost (Melvin Vivas on X · notes), Mercury 2.5 runs at 1,100 tokens/sec (Melvin Vivas on X · notes) - MiniMax (official) · person · x.com · free
MiniMax's official X account; MiniMax is the AI lab behind the M3 and H3 models.
Mentioned in: Running MiniMax M3 Locally in ComfyUI on an RTX 3090 (Melvin Vivas on X · notes), MiniMax 2.7 One-Shots a Linear Clone at 95% Lower Cost (Melvin Vivas on X · notes) - Muse Glimmer-30B · tool · developer.meta.com · free
Open-weight 30B model aimed at agentic work, tool use and long-horizon reasoning.
Mentioned in: Fine-Tune Muse Glimmer 30B with LoRA or Full-Parameter on Fireworks (Melvin Vivas on X · notes), Meta Muse Glimmer-30B: Open Weights Model That Beats Qwen3.6 37B on Agentic Tasks (Melvin Vivas on X · notes) - Nano Banana 2.0 · tool · blog.google · free
Google's Gemini image generation model, described as having Pro-level performance at Flash speed.
Mentioned in: Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Building a Nano Banana 2 Image-Gen App with Memex Managed AI Connectors (Melvin Vivas on X · notes) - NVIDIA Nemotron · tool · nvidia.com · free
NVIDIA's family of open foundation LLMs, including Nemotron 3 Super and the upcoming Ultra.
Mentioned in: NVIDIA Agent Toolkit: open Nemotron models for domain-specific agents (Melvin Vivas on X · notes), Jensen Huang on NVIDIA's Long-Term Commitment to Open Nemotron Models (Melvin Vivas on X · notes) - NVIDIA Nemotron 3.5 Lightning · tool · huggingface.co · free
Open 30B MoE model (3B active) from NVIDIA, built for fast, high-volume agent tasks.
Mentioned in: Open-weight Nemotron models for finance and healthcare (Melvin Vivas on X · notes), NVIDIA Nemotron 3.5 Lightning: Fast Open MoE Model for Agents (Melvin Vivas on X · notes) - nvidia/GLM-5.2-NVFP4 · tool · huggingface.co · free
NVIDIA's NVFP4-quantized release of the open GLM-5.2 model, with its model card on Hugging Face.
Mentioned in: Serve GLM-5.2 NVFP4 with vLLM on NVIDIA Blackwell (Melvin Vivas on X · notes), Model card: NVIDIA's NVFP4-quantized GLM-5.2 on Hugging Face (Melvin Vivas on X · notes) - OpenAI Responses API · docs · platform.openai.com · check price
OpenAI's API for building model-powered apps and agents with tool use.
Mentioned in: OpenAI DevDay 2026 Recap: Dots, Agents API, Codex Cloud & Marketplace (Melvin Vivas on X · notes), GPT-5.3-Codex Now Available in OpenAI's Responses API (Melvin Vivas on X · notes) - OpenRouter Documentation · docs · openrouter.ai · free
OpenRouter's docs explaining video generation parameters, the async request flow and the video models route.
Mentioned in: OpenRouter MCP: Pick, Price, and Test LLMs from Inside Your Coding Agent (Melvin Vivas on X · notes), OpenRouter Adds Video Generation: One API for Veo, Seedance, Wan and Sora (Melvin Vivas on X · notes) - Qwen 3.5 35B · tool · huggingface.co · free
An open-weight Mixture-of-Experts LLM from Alibaba's Qwen team that can run locally on a 24GB GPU.
Mentioned in: Running Qwen 3.5 35B Locally in LM Studio for Vision and Prompts (Melvin Vivas on X · notes), Local Qwen3.5-35B-A3B on a 24GB GPU Builds a Full Game From One Spec (Melvin Vivas on X · notes) - Qwen3.5-LiveTranslate · tool · qwen.ai · check price
Alibaba Qwen's real-time speech translation model: understands 60 languages, speaks 29, with low latency and hot-word customization.
Mentioned in: Qwen 3.5 Live Translate: Real-Time Speech Translation Demo (Melvin Vivas on X · notes), Qwen3.5-LiveTranslate: Real-Time Audio-Visual Interpretation Model (Melvin Vivas on X · notes) - Qwen3.6-Plus · tool · qwen.ai · free
Alibaba's Qwen model, now available in OpenCode Go and recommended for coding.
Mentioned in: Qwen3.6-Plus and Qwen3.5-Plus Coding Models Now Available in OpenCode (Melvin Vivas on X · notes), Qwen 3.6 Plus Preview Is Free for a Limited Time on OpenRouter (Melvin Vivas on X · notes) - Qwen3.8-27B EXL3 · tool · huggingface.co · free
Experimental EXL3-quantized build of the Qwen3.8-27B model that can run with 200K+ context on a single 24GB GPU.
Mentioned in: Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s (Melvin Vivas on X · notes), Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090 (Melvin Vivas on X · notes) - AI SDK (Vercel) · tool · ai-sdk.dev · free
Vercel's TypeScript toolkit for calling LLM providers and building AI apps.
Mentioned in: Swapping AI SDK for pi-ai as the LLM provider layer (Melvin Vivas on X · notes) - AINews: Claude Opus 5.5, the new default (Latent Space) · article · latent.space · free
Latent Space AINews issue explaining why Claude Opus 5.5 became the default model after a side-by-side test with GPT-6 Sol.
Mentioned in: Opus 5.5 vs GPT-6 Sol for Writing: Latent Space AINews Test (Melvin Vivas on X · notes) - Anthropic · tool · anthropic.com · check price
An AI model provider (Claude models) whose models are available through managed connectors.
Mentioned in: Building a Nano Banana 2 Image-Gen App with Memex Managed AI Connectors (Melvin Vivas on X · notes) - Anthropic API · tool · anthropic.com · paid
Anthropic's API for calling Claude models from your own applications.
Mentioned in: The 3 Levels of AI Engineering: LLM Apps → Production → Agentic Systems (Bashiri Smith on Facebook · notes) - Anthropic API docs · docs · docs.anthropic.com · free
Claude API reference. The guide also points to it for safety guidance.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Apple Foundation Models framework · docs · developer.apple.com · free
Apple's framework that lets apps on Apple platforms use language models.
Mentioned in: Claude support for Apple's Foundation Models framework (Melvin Vivas on X · notes) - Attention Is All You Need · paper · arxiv.org · free
The 2017 Google paper that introduced the Transformer and scaled dot-product / multi-head attention.
Mentioned in: How GPT Works: From Token Embeddings to Multi-Head Attention (Melvin Vivas on X · notes) - Avi Chawla · person · x.com · free
AI educator who posts explanations of ML and LLM research; author of the quoted thread.
Mentioned in: Transferable KV Cache: Reusing One Model's Cache in Another (Melvin Vivas on X · notes) - Bonsai 8B (PrismML) · tool · prismml.com · free
1-bit quantized dense 8B language model built for on-device inference.
Mentioned in: 1-bit Bonsai 8B Runs On-Device on iPhone at 40+ tok/s (Melvin Vivas on X · notes) - Building with Claude Sonnet 5.5 (claude.dev Blog) · article · claude.dev · free
Official claude.dev blog guide on building applications with Claude Sonnet 5.5.
Mentioned in: Guide: Building with Claude Sonnet 5.5 (Melvin Vivas on X · notes) - Claude Fable 5 / Mythos 5 · tool · anthropic.com · paid
Anthropic Claude models that were redeployed, per the creator.
Mentioned in: AI News Roundup: Claude Fable 5, Scientist AI, ZCode, NVIDIA RL (Melvin Vivas on X · notes) - Claude for Foundation Models (Claude blog) · article · claude.com · free
Anthropic's announcement of Claude support for Apple's Foundation Models framework.
Mentioned in: Claude support for Apple's Foundation Models framework (Melvin Vivas on X · notes) - Claude Opus · tool · anthropic.com · paid
Anthropic's frontier model, used by FireRouter for the harder tasks.
Mentioned in: FireRouter: cost-aware model routing between open models and Claude Opus (Melvin Vivas on X · notes) - Claude Platform · tool · platform.claude.com · paid
Anthropic's developer platform for building with Claude models, where Managed Agents is available in public beta.
Mentioned in: Claude Managed Agents: Anthropic's Hosted Harness for Building and Deploying Agents (Melvin Vivas on X · notes) - Claude Sonnet 5 · tool · anthropic.com · paid
Anthropic model priced at $2/M input and $10/M output tokens.
Mentioned in: Claude Sonnet 5 Pricing Made Permanent ($2/$10 per M Tokens) (Melvin Vivas on X · notes) - Cohere Parse (Hugging Face Space) · tool · huggingface.co · free
Cohere's document-parsing model, with a demo you can try on Hugging Face Spaces.
Mentioned in: Cohere Parse Beats Frontier LLMs at Receipt Parsing (Melvin Vivas on X · notes) - Cohere Transcribe WebGPU (Hugging Face Space) · tool · huggingface.co · free
Browser demo by CohereLabs that runs the transcription model locally using WebGPU.
Mentioned in: Local Audio Transcription with Cohere Transcribe on WebGPU (Melvin Vivas on X · notes) - CohereLabs/cohere-transcribe-03-2026 · tool · huggingface.co · free
Model card on Hugging Face for Cohere's audio transcription (speech-to-text) model.
Mentioned in: Local Audio Transcription with Cohere Transcribe on WebGPU (Melvin Vivas on X · notes) - Compare AI Models: Pricing, Context & Benchmarks | OpenRouter (discounted filter) · website · openrouter.ai · free
OpenRouter model list filtered to discounted models, with pricing and context details.
Mentioned in: Find Discounted Models with OpenRouter's New Filter (Melvin Vivas on X · notes) - convaiinnovations / laya (Hugging Face) · website · huggingface.co · free
A Hugging Face model repo said to possibly be an open-weights version of Jev. The link was truncated and broken when checked.
Mentioned in: Possible Open-Weights Version of Jev on Hugging Face (Melvin Vivas on X · notes) - DealignAI models on Hugging Face · website · huggingface.co · free
DealignAI's Hugging Face page listing abliterated models you can run locally.
Mentioned in: Abliterated Local Models from DealignAI on Hugging Face (Melvin Vivas on X · notes) - DiffusionGemma · tool · ai.google.dev · free
An experimental open Gemma model that uses diffusion to generate blocks of text in parallel.
Mentioned in: DiffusionGemma: Google's Experimental Diffusion-Based Text Model (Melvin Vivas on X · notes) - Documentation | Claude Platform · docs · platform.claude.com · free
Main documentation hub for building with Claude on the Claude Platform.
Mentioned in: Anthropic's Prompting Best Practices Guide for Claude Opus 4.8 (Melvin Vivas on X · notes) - dots3.note · tool · huggingface.co · free
Open model newly supported in SGLang.
Mentioned in: SGLang v0.5.19 release: new models and beam search (Melvin Vivas on X · notes) - Exploring Language Models (Maarten Grootendorst's Substack) · newsletter · newsletter.maartengrootendorst.com · free
Maarten Grootendorst's newsletter of visual guides to LLM concepts and architectures.
Mentioned in: A Visual Guide to Gemma 4 Architecture (Maarten Grootendorst) (Melvin Vivas on X · notes) - Fastino-Nemotron-3.5-Lightning-Finance · tool · huggingface.co · free
Open-weight finance model from Fastino, built on NVIDIA Nemotron 3.5 Lightning.
Mentioned in: Open-weight Nemotron models for finance and healthcare (Melvin Vivas on X · notes) - Fastino-Nemotron-3.5-Lightning-Healthcare · tool · huggingface.co · free
Open-weight healthcare model from Fastino, built on NVIDIA Nemotron 3.5 Lightning.
Mentioned in: Open-weight Nemotron models for finance and healthcare (Melvin Vivas on X · notes) - Gemini 3 · tool · deepmind.google · check price
Google's flagship model family. The video says its research was reused to build Gemma.
Mentioned in: Gemma 4: Google's Open Models for On-Device, Offline AI (Melvin Vivas on X · notes) - Gemini 3.1 Pro · tool · deepmind.google · check price
Google's Gemini model that powers Gemini-SQL2.
Mentioned in: Gemini-SQL2: text-to-SQL with Gemini 3.1 Pro, SOTA on BIRD (Melvin Vivas on X · notes) - Gemini 3.5 Flash-Lite · tool · deepmind.google · check price
Fast, low-cost Google model for everyday tasks.
Mentioned in: Gemini 3.6 Flash and 3.5 Flash-Lite Released for Agents (Melvin Vivas on X · notes) - Gemini 3.6 Flash · tool · blog.google · check price · open in a browser to verify
Google model that is more token-efficient than 3.5 Flash at the same cost.
Mentioned in: Gemini 3.6 Flash and 3.5 Flash-Lite Released for Agents (Melvin Vivas on X · notes) - ggml-org/gemma-4-26B-A4B-it-GGUF · tool · huggingface.co · free
Quantized GGUF build of the Gemma 4 26B-A4B instruction-tuned model, made for llama.cpp.
Mentioned in: Run Gemma 4 Locally with llama.cpp in Two Commands (Melvin Vivas on X · notes) - ggml-org/Qwen3.8-27B-GGUF · tool · huggingface.co · free
A GGUF-quantized build of Qwen3.8 27B for llama.cpp-style local inference.
Mentioned in: Three Local GGUF Models That Fit on an RTX 3090 (24GB) (Melvin Vivas on X · notes) - GLiNER · repo · github.com · free
An open-source family of lightweight, generalist models for named entity recognition and extraction.
Mentioned in: GLiNER decision model demos from Fastino Labs (Melvin Vivas on X · notes) - GLM (Zhipu AI / Z.ai) · tool · huggingface.co · free
A family of open-weight large language models from Zhipu AI (Z.ai), strong at coding and agent tasks.
Mentioned in: Bolt.new Adds Open Models (GLM, DeepSeek, Kimi) via Bolt Forge (Melvin Vivas on X · notes) - GLM 5.2 launch blog (Z.ai) · article · z.ai · free
Z.ai's launch post with the details of the open GLM 5.2 coding model.
Mentioned in: GLM 5.2 open model now deployable on Google Cloud (Melvin Vivas on X · notes) - GLM 5.3 / GLM 5.3 Flash · tool · cursor.com · paid
Open-weight model family; the Max version is the top open-weight scorer on CursorBench 4.0.
Mentioned in: GLM 5.3 and GLM 5.3 Flash now in Cursor (Melvin Vivas on X · notes) - GLM-5.2 (max): API Provider Performance Benchmarking & Price Analysis | Artificial Analysis · website · artificialanalysis.ai · free
Artificial Analysis page comparing GLM-5.2 API providers by output speed, latency and price.
Mentioned in: Compare GLM-5.2 API Providers on Artificial Analysis (Melvin Vivas on X · notes) - GLM-5.2 on Fireworks AI · tool · fireworks.ai · check price
Fireworks-hosted inference endpoint for the GLM-5.2 open-weights model.
Mentioned in: GLM-5.2 on Fireworks: Top Open-Weights Model on GDPval-AA (Melvin Vivas on X · notes) - GLM-5V-Turbo · tool · docs.z.ai · check price
A multimodal vision coding model in the GLM family that generates code from images, videos, design drafts and document layouts.
Mentioned in: GLM-5V-Turbo: A Vision Coding Model for Multimodal Inputs (Melvin Vivas on X · notes) - Google Interactions API · docs · ai.google.dev · free
Google's main API for Gemini models and agents, using a single /interactions endpoint.
Mentioned in: Google Interactions API GA: one endpoint for inference and agents (Melvin Vivas on X · notes) - google-genai SDK · repo · github.com · free
Google Gen AI Python SDK used to call Gemini and Gemma models (generate_content, function calling, images).
Mentioned in: Using Gemma 4 via the Gemini API and Google AI Studio (Melvin Vivas on X · notes) - GPT Image 2 (gpt-image-2) · tool · community.openai.com · paid
OpenAI's image generation model, available through the API, which now supports transparent and semi-transparent backgrounds.
Mentioned in: GPT Image 2 Adds Transparent Background Support in the OpenAI API (Melvin Vivas on X · notes) - GPT-5 Mini · tool · developers.openai.com · paid
An OpenAI model used as the LLM-based retrieval baseline.
Mentioned in: Using Jev as a Reranker to Augment RAG Retrieval (Melvin Vivas on X · notes) - GPT-5.3-Codex model documentation · docs · developers.openai.com · free
OpenAI developer docs for the GPT-5.3-Codex model.
Mentioned in: GPT-5.3-Codex Now Available in OpenAI's Responses API (Melvin Vivas on X · notes) - GPT-5.6 Sol - API Pricing & Benchmarks (OpenRouter) · website · openrouter.ai · check price
OpenRouter page with pricing and benchmarks for OpenAI's GPT-5.6 Sol model.
Mentioned in: Use GPT-5.6 Sol on OpenRouter at 50% off when Codex limits run out (Melvin Vivas on X · notes) - gpt-oss-20b · tool · huggingface.co · free
OpenAI's 20B-parameter open-weight language model.
Mentioned in: Free gpt-oss-20b and Gemma 4 26B on OpenRouter (Melvin Vivas on X · notes) - Granite 4.2 · tool · research.ibm.com · free
IBM's open Granite language model.
Mentioned in: SGLang v0.5.19 release: new models and beam search (Melvin Vivas on X · notes) - Grok 4.5 Model Card · pdf · media.x.ai · free
The official model card for Grok 4.5, describing its capabilities and evaluations.
Mentioned in: Grok 4.5 model card released (Melvin Vivas on X · notes) - GroqCloud Console · tool · console.groq.com · free
Groq's developer console for fast LLM inference APIs.
Mentioned in: Qwen3.8-27B Free on Groq at ~450 Tokens per Second (Melvin Vivas on X · notes) - Hy-MT1.5-1.8B-1.25bit · tool · huggingface.co · free
An open-source 1.25-bit translation model, 440MB, covering 33 languages for offline use on phones.
Mentioned in: Hy-MT1.5-1.8B-1.25bit: A 440MB Offline Phone Translation Model (Melvin Vivas on X · notes) - Hy3 (Tencent Hunyuan) · tool · github.com · free
Tencent Hunyuan's LLM, built for low-cost agent use with strong coding, tool calling, reasoning and long-context skills.
Mentioned in: Hy3 (Tencent Hunyuan) Free on Nous Portal for a Limited Week (Melvin Vivas on X · notes) - Improving Language Understanding by Generative Pre-Training (original GPT paper) · paper · cdn.openai.com · free
OpenAI's 2018 paper introducing GPT, a decoder-only Transformer pretrained generatively and then fine-tuned.
Mentioned in: How GPT Works: From Token Embeddings to Multi-Head Attention (Melvin Vivas on X · notes) - Infron (@InfronAI) on X · tool · x.com · free
Infron's account on X. The post says it offers free access to Qwen-3.8 27B.
Mentioned in: Free Qwen-3.8 27B Model via Infron (Melvin Vivas on X · notes) - Interactions API: our primary interface for Gemini models and agents · article · blog.google · free
Google blog post announcing that the Interactions API is generally available and what it can do.
Mentioned in: Google Interactions API is GA: recommended for Gemini (Melvin Vivas on X · notes) - Introducing Grok 4.7 (xAI) · article · x.ai · free
xAI's official announcement page for the Grok 4.7 model.
Mentioned in: Grok 4.7 Launch: Official xAI Announcement (Melvin Vivas on X · notes) - Kev-4B · tool · openrouter.ai · paid
A small (4B) language model, now available to call through OpenRouter.
Mentioned in: Kev-4B model, an alternative to Jev, now available on OpenRouter (Melvin Vivas on X · notes) - Kimi 2.7 · tool · huggingface.co · free
Moonshot AI's Kimi model release, named as a recent open-model update.
Mentioned in: Open-weight model updates: Kimi 2.7 and GLM 5.2 (Melvin Vivas on X · notes) - Kimi K2.5 · tool · huggingface.co · free
An open-source large language model from Moonshot AI.
Mentioned in: Cursor Composer 2 Is Built on Open-Source Kimi K2.5 (Melvin Vivas on X · notes) - Kimi K2.7 · tool · cursor.com · paid
Open-weight LLM from Moonshot AI, now available in Cursor.
Mentioned in: Open models in Cursor: Kimi K2.7 vs GLM 5.2 (Melvin Vivas on X · notes) - Kimi K3 (Moonshot AI) · tool · kimi.ai · check price
A Moonshot AI large language model, used here as the Hermes agent's model.
Mentioned in: Connecting Hermes Agent to the Buzz Agent-First Chat App via the Native Gateway (Melvin Vivas on X · notes) - Laya Decision models · tool · laya.convaiinnovations.com · free
A model family pitched as an alternative to Jev that can run locally on low-memory hardware.
Mentioned in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes) - LFM2-24B-A2B · tool · huggingface.co · free
Liquid AI's Mixture-of-Experts model with 24B total and about 2B active parameters.
Mentioned in: Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework (Melvin Vivas on X · notes) - LFM2.5-350M (Liquid AI) · tool · huggingface.co · free
Tiny 350M-parameter language model from Liquid AI, used as the base model for fine-tuning.
Mentioned in: Fine-Tune LFM2.5-350M with GRPO in TRL for Structured Outputs (Melvin Vivas on X · notes) - LFM2.5-8B-A1B · tool · liquid.ai · free
Liquid AI's Mixture-of-Experts model with 8B total and about 1B active parameters.
Mentioned in: Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework (Melvin Vivas on X · notes) - LFM2.5-Audio-1.5B · tool · huggingface.co · free
Liquid AI's 1.5B-parameter audio model from the LFM2.5 family, used here for speech-to-text.
Mentioned in: Local Speech-to-Text with LFM2.5-Audio-1.5B and llama.cpp (Melvin Vivas on X · notes) - Ling-3.0-flash / Ling-3.0-tiny · tool · huggingface.co · free
Ling 3.0 family of open language models in flash and tiny sizes.
Mentioned in: SGLang v0.5.19 release: new models and beam search (Melvin Vivas on X · notes) - Liquid AI Docs – Vision Capabilities guide · docs · docs.liquid.ai · free
Guide to using LFM2.5-VL-3B for image prompts, OCR, layout annotation, grounding and tool calling.
Mentioned in: Using LFM2.5-VL-3B's Vision Capabilities (Liquid AI Guide) (Melvin Vivas on X · notes) - Liquid Apollo · tool · liquid.ai · free
Liquid AI's mobile app for running its models on-device.
Mentioned in: Running Liquid AI Models On-Device with the Apollo iPhone App (Melvin Vivas on X · notes) - Lou (@louszbd) on X · person · x.com · free
X account that is offering GLM-5-Turbo rate limit increases.
Mentioned in: How to Get a GLM-5-Turbo Rate Limit Increase for OpenClaw (Melvin Vivas on X · notes) - Maarten Grootendorst · person · x.com · free
Author of visual guides to LLMs and co-author of Hands-On Large Language Models; worth following for clear explanations of model architectures.
Mentioned in: A Visual Guide to Gemma 4 Architecture (Maarten Grootendorst) (Melvin Vivas on X · notes) - Meta Muse Glimmer 30B · tool · huggingface.co · free
A 30B open model from Meta that can be fine-tuned and run locally.
Mentioned in: Fine-Tune Muse Glimmer 30B for Free with Unsloth (incl. GRPO) (Melvin Vivas on X · notes) - Meta Muse Spark · tool · free
Meta Superintelligence Labs' first model, which now powers Meta AI.
Mentioned in: Read Benchmark Numbers, Not Highlights: the Muse Spark Lesson (Melvin Vivas on X · notes) - MiniCPM-V 4.6 · tool · github.com · free
Open-source 1.3B vision-language model from OpenBMB, built for OCR and image understanding on edge devices.
Mentioned in: MiniCPM-V 4.6 1.3B: Small Open-Source Vision/OCR Model for Edge Devices (Melvin Vivas on X · notes) - MiniMax M2.5 · tool · huggingface.co · free
An open-weight large language model from MiniMax, optimized for lightweight, high-frequency agent workflows.
Mentioned in: MiniMax M2.5: first open-weight model in Notion Custom Agents (Melvin Vivas on X · notes) - Mistral AI · website · mistral.ai · check price
The AI company that builds the Voxtral speech models.
Mentioned in: Real-Time Speech Transcription in the Browser with Voxtral and WebGPU (Melvin Vivas on X · notes) - Muse Spark 1.2 (Meta) · tool · research.meta.ai · check price
Meta's multimodal model for visual-to-code, perception-to-action and audio-visual understanding tasks.
Mentioned in: Meta Muse Spark 1.2: Multimodal Model Release Announcement (Melvin Vivas on X · notes) - nanoGPT (Karpathy) · repo · github.com · free
A minimal, readable repo for training GPT.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Nemotron 3 · tool · research.nvidia.com · free
NVIDIA's Nemotron 3 language model family.
Mentioned in: Try NVIDIA Nemotron 3 for Free on OpenRouter (Melvin Vivas on X · notes) - NeoHorse-Jev-4B (ModelScope collection) · tool · modelscope.ai · free
Open-weight 4B model that turns application states into structured decisions and probabilities.
Mentioned in: NeoHorse-Jev-4B: Open 4B Model for Structured Decisions (Melvin Vivas on X · notes) - North Micro Vision · tool · huggingface.co · free
Cohere's smallest open-source vision-language model, built for document understanding.
Mentioned in: Cohere's North Micro Vision: a small open-source vision model for documents (Melvin Vivas on X · notes) - NVIDIA Nemotron 3 Ultra · tool · developer.nvidia.com · free
NVIDIA's Nemotron 3 Ultra open model.
Mentioned in: Use NVIDIA Nemotron 3 Ultra Free on OpenRouter (Melvin Vivas on X · notes) - NVIDIA NIM APIs (build.nvidia.com) · website · build.nvidia.com · check price
NVIDIA's catalog of hosted model endpoints you can try and call through NIM APIs.
Mentioned in: NVIDIA build.nvidia.com: Try Hosted NIM Model APIs (Melvin Vivas on X · notes) - OpenAI API (announcement) · article · openai.com · free
OpenAI's blog post introducing public API access to its models.
Mentioned in: Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) (Melvin Vivas on melvinvivas.com · notes) - OpenAI GPT-3 Models (docs) · docs · developers.openai.com · free
OpenAI docs page on the GPT-3 model family; the link now appears broken.
Mentioned in: Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) (Melvin Vivas on melvinvivas.com · notes) - OpenAI Image API · tool · developers.openai.com · paid
OpenAI's image generation and editing API.
Mentioned in: Building and Deploying a Full App with the Codex App, MCP Servers and Skills (Melvin Vivas on X · notes) - OpenAI Models (docs) · docs · platform.openai.com · free
OpenAI documentation listing the available models and what they can do.
Mentioned in: Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) (Melvin Vivas on melvinvivas.com · notes) - OpenAI, Claude & Gemini API Tutorial in Python (Machine Learning Plus) · article · machinelearningplus.com · free
A reference guide for working with LLM APIs from providers other than OpenAI (the exact source isn't shown).
Mentioned in: How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes) - OpenRouter Python SDK · tool · openrouter.ai · free
OpenRouter's Python SDK for calling models from your application code.
Mentioned in: OpenRouter MCP: Pick, Price, and Test LLMs from Inside Your Coding Agent (Melvin Vivas on X · notes) - Opus 5.5 · tool · anthropic.com · paid
Anthropic Claude model, one of the three models asked in the council demo.
Mentioned in: Pi Agent Council: Ask Multiple LLMs in Parallel and Compare Their Advice (Melvin Vivas on X · notes) - Ornith-1.0 · tool · github.com · free
A family of open-source LLMs for agentic coding, from 9B dense up to 397B MoE.
Mentioned in: Ornith-1.0: open-source LLM family for agentic coding (Melvin Vivas on X · notes) - Ornith-1.5 · tool · ornith.ai · free
Open-source LLM family (9B dense, 35B MoE, 397B MoE) trained with self-improving strategies.
Mentioned in: Ornith-1.5: Open-Source LLM Family (9B Dense to 397B MoE) (Melvin Vivas on X · notes) - ornith-ai/Ornith-1.5-35B-A3B-GGUF · tool · huggingface.co · free
A GGUF build of Ornith 1.5, a 35B mixture-of-experts model with about 3B active parameters.
Mentioned in: Three Local GGUF Models That Fit on an RTX 3090 (24GB) (Melvin Vivas on X · notes) - Ox Alpha (OpenRouter) · tool · openrouter.ai · free
Stealth model that is free to use through OpenRouter's API.
Mentioned in: Try the Free Stealth Model Ox Alpha on OpenRouter (Melvin Vivas on X · notes) - pi-ai · tool · github.com · free · open in a browser to verify
A library from Pi that gives one interface to many LLM providers.
Mentioned in: Swapping AI SDK for pi-ai as the LLM provider layer (Melvin Vivas on X · notes) - Quesma blog: Qwen3.8 27B quantizations benchmarked · article · quesma.com · free
A benchmark of Qwen3.8 27B quantization levels on the Terminal-Bench 2.1 agentic coding benchmark.
Mentioned in: Qwen3.8 27B Quantization Benchmark: 4-Bit Is Enough for Agentic Coding (Melvin Vivas on X · notes) - Qwen 27B Opus-distilled model · tool · huggingface.co · free
An open-weight Qwen model of about 27B parameters, distilled from Claude Opus outputs for coding.
Mentioned in: Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code (Melvin Vivas on X · notes) - Qwen 3.5 Omni · tool · qwen.ai · free
Alibaba's omni-modal Qwen model that accepts audio and video input and can generate code from spoken, visually grounded instructions.
Mentioned in: Audio-Visual Vibe Coding with Qwen 3.5 Omni: Spoken Specs to Web App (Melvin Vivas on X · notes) - Qwen 3.6 · tool · qwen.ai · free
An open-weight LLM family from Alibaba's Qwen team that can run locally.
Mentioned in: Use local Qwen 3.6 to write JSON prompts for Ideogram 4 (Melvin Vivas on X · notes) - Qwen 3.7 Plus · tool · qwen.ai · check price
Alibaba's Qwen model family, latest preview release.
Mentioned in: Qwen 3.7 Plus Preview Ranks #16 in the Vision Arena (Melvin Vivas on X · notes) - Qwen 4B VL · tool · huggingface.co · free
Qwen's small 4B vision-language model, used here for OCR.
Mentioned in: aibackends 0.3.0: Model Caching Speeds Up PII and OCR Inference (Melvin Vivas on X · notes) - Qwen models (Alibaba) · tool · github.com · free
Alibaba's family of open models (LLM and multimodal) that can run locally.
Mentioned in: Running AI Locally: Rebuilding AIBackends as a Python Library with Open Models (Melvin Vivas on X · notes) - Qwen VL · tool · github.com · free
Alibaba's Qwen family of vision-language models, which can be used for OCR and document understanding.
Mentioned in: Mistral OCR 4 launch, and the missing Qwen VL comparison (Melvin Vivas on X · notes) - Qwen2.5 · tool · github.com · free
Alibaba's family of open-weight LLMs in many sizes, often used as fine-tuning bases.
Mentioned in: Using Codex to Generate Synthetic Training Pairs for Fine-tuning Qwen2.5 (Melvin Vivas on X · notes) - Qwen2.5-0.5B · tool · huggingface.co · free
Tiny open-weight Qwen model, named as the base for Kev-0.5B.
Mentioned in: Kev-0.5B: A Tiny Open-Source Decision Model to Train on a MacBook (Melvin Vivas on X · notes) - Qwen3.5 4B · tool · huggingface.co · free
Open-weight base model from Alibaba's Qwen family.
Mentioned in: NuExtract3: Open 4B VLM for OCR and Structured JSON Extraction (Melvin Vivas on X · notes) - Qwen3.5-0.8B · tool · huggingface.co · free
Small Qwen model used as the comparison baseline.
Mentioned in: MiniCPM-V 4.6 1.3B: Small Open-Source Vision/OCR Model for Edge Devices (Melvin Vivas on X · notes) - Qwen3.5-27B-GGUF (Unsloth) · tool · huggingface.co · free
Unsloth's GGUF quantizations of Qwen3.5-27B, made for llama.cpp.
Mentioned in: Run Qwen3.5-27B GGUF with llama-server for Claude Code on an RTX 3090 (Melvin Vivas on X · notes) - Qwen3.5-Plus · tool · alibabacloud.com · paid
An earlier Qwen Plus model, also available in OpenCode Go.
Mentioned in: Qwen3.6-Plus and Qwen3.5-Plus Coding Models Now Available in OpenCode (Melvin Vivas on X · notes) - Qwen3.7-Max · tool · qwen.ai · check price
Alibaba Qwen's flagship model, built for coding and productivity agents.
Mentioned in: Qwen3.7-Max: Qwen's flagship model for agents (Melvin Vivas on X · notes) - Qwen3.8 27B (FP8) · tool · huggingface.co · free
Open-weight Qwen language model, 27B parameters, served in FP8 precision.
Mentioned in: Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference (Melvin Vivas on X · notes) - Qwen3.8-2.4T-A95B · tool · huggingface.co · free
The largest Qwen3.8 model; by its name, about 2.4T total parameters with about 95B active.
Mentioned in: Qwen3.8 Model Family Now on Novita AI: Flash, 27B and 2.4T-A95B (Melvin Vivas on X · notes) - Qwen3.8-Flash · tool · qwen.ai · check price
A 125B-parameter MoE model with 6B active parameters that accepts text, image and video input.
Mentioned in: Qwen3.8 Model Family Now on Novita AI: Flash, 27B and 2.4T-A95B (Melvin Vivas on X · notes) - Qwen3.8-Flash-Next · tool · huggingface.co · free
Open 125B MoE model from Qwen that is tuned to run fast on CPU RAM or unified memory.
Mentioned in: Run Qwen3.8-Flash-Next (125B MoE) Locally with Unsloth GGUFs (Melvin Vivas on X · notes) - r/LocalLLaMA · community · reddit.com · free
Reddit community for running and discussing local and open LLMs.
Mentioned in: Jensen Huang on NVIDIA's Long-Term Commitment to Open Nemotron Models (Melvin Vivas on X · notes) - Rachel Nabors · person · x.com · free
Speaker whose talk explains why you don't need frontier LLMs for every task.
Mentioned in: You don't need frontier LLMs for everything: use SLMs (Melvin Vivas on X · notes) - Run Gemma with the Gemini API | Google AI for Developers · docs · ai.google.dev · free
Official guide to calling Gemma models through the Gemini API.
Mentioned in: Gemma 4 Now Available Through the Gemini API (Dev Use Only) (Melvin Vivas on X · notes) - Running a 28.9M Parameter LLM on an $8 Microcontroller (Hackster.io) · article · hackster.io · free
Hackster.io news article about fitting a 28.9M-parameter LLM onto an ESP32 using a trick from Gemma.
Mentioned in: Running a 28.9M-Parameter LLM on an $8 ESP32 Microcontroller (Melvin Vivas on X · notes) - Sebastian Raschka · person · sebastianraschka.com · free
An ML researcher and author known for clear breakdowns of LLM research and architectures (Ahead of AI newsletter, Build a Large Language Model (From Scratch)).
Mentioned in: 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes) - Space Bunny Alpha (OpenRouter) · tool · openrouter.ai · check price
A stealth flash model on OpenRouter with adjustable reasoning, a 1M-token context and multimodal input.
Mentioned in: Space Bunny Alpha: stealth 1M-context flash model on OpenRouter (Melvin Vivas on X · notes) - Spark-X2.5-4B · tool · github.com · free
A 4B-parameter open model for local use with a 1M-token context window, multilingual support and tool-use/agent abilities.
Mentioned in: Spark-X2.5-4B: a small local model with a 1M-token context window (Melvin Vivas on X · notes) - Spark2.5 · tool · check price
Model newly supported in SGLang.
Mentioned in: SGLang v0.5.19 release: new models and beam search (Melvin Vivas on X · notes) - SWE-1.7 · tool · cognition.com · paid
Cognition's coding model, said to be near frontier quality at low cost and 1000 tok/s.
Mentioned in: Cognition Releases SWE-1.7: Near-Frontier Coding Model at Low Cost (Melvin Vivas on X · notes) - System Card: Claude Opus 4.8 · pdf · cdn.sanity.io · free
Anthropic's 244-page system card for Claude Opus 4.8 (May 28, 2026).
Mentioned in: Claude Opus 4.8 System Card (official PDF) (Melvin Vivas on X · notes) - Ternary Bonsai 2 27B · tool · prismml.com · free
A ternary-quantized 27B model, 9x smaller than full precision, that keeps 98.2% of benchmark performance.
Mentioned in: Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) - text-davinci-003 · tool · platform.openai.com · paid
OpenAI GPT-3 completion model; the default LLM in llama_index at the time (now deprecated).
Mentioned in: Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) (Melvin Vivas on melvinvivas.com · notes) - Union Alpha on OpenRouter · tool · openrouter.ai · free
Free multimodal stealth model with 256K context and tool calling, for research, coding and agents.
Mentioned in: Union Alpha: Free Stealth Model on OpenRouter (Melvin Vivas on X · notes) - Unsloth Gemma 4 12B IT GGUF (Hugging Face) · tool · huggingface.co · free · open in a browser to verify
Quantized Dynamic GGUF weights of Gemma 4 12B instruct for low-RAM local inference.
Mentioned in: Run Gemma 4 12B on 8GB RAM with Unsloth Dynamic GGUFs (Melvin Vivas on X · notes) - Unsloth Qwen3.6-27B GGUF (Hugging Face) · tool · huggingface.co · free · open in a browser to verify
Quantized GGUF weights of Qwen3.6-27B by Unsloth for running locally.
Mentioned in: Run Qwen3.6-27B Locally in 18GB RAM with Unsloth GGUFs (Melvin Vivas on X · notes) - Voxtral-Mini-4B (Voxtral Realtime) · tool · huggingface.co · free
Mistral AI's streaming speech-recognition (ASR) model with a custom causal audio encoder, supporting 13 languages.
Mentioned in: Real-Time Speech Transcription in the Browser with Voxtral and WebGPU (Melvin Vivas on X · notes) - Z.ai coding plan API endpoint · docs · api.z.ai · paid
Coding-optimized OpenAI-compatible endpoint for the GLM coding plan; the full path is https://api.z.ai/api/coding/paas/v4.
Mentioned in: Use the Right Z.ai Endpoint for the GLM 5.2 Coding Plan (Melvin Vivas on X · notes) - Z.ai general API endpoint (paas/v4) · docs · api.z.ai · paid · open in a browser to verify
General-purpose Z.ai API endpoint, which should not be used with the coding plan.
Mentioned in: Use the Right Z.ai Endpoint for the GLM 5.2 Coding Plan (Melvin Vivas on X · notes)
Build
- Build a tool that automatically creates short clips from a long video. (from Use Sonnet 5.5 instead of Opus 5.5 for faster video-clipping tasks)
- Set up a local LLM on your own laptop with Ollama or LM Studio and compare several small GGUF models (from Running LLMs Locally Without an Expensive Rig)
- Prompt a frontier model such as Claude Opus 5.5 to create a one-shot motion-graphics explainer of an AI concept, such as tokenization or attention. (from How LLMs Work: A Motion-Graphics Explainer Made in One Shot with Claude Opus 5.5)
- Build a demo that uses a decision model for something beyond text parsing. (from GLiNER decision model demos from Fastino Labs)
- Build a customer support intent classifier with GLiNER2.5-Decide that runs on CPU. (from Intent classification for support using GLiNER2.5-Decide notebook)
- Use GLiNER2.5-Decide to build a fast rule-based classifier from your own typed questions and rules. (from Getting started with GLiNER2.5-Decide in a Colab notebook)
- Build an offline chat app on iPhone running Gemma 4 locally (from Running Gemma 4 Models Offline on an iPhone)
- Build a support-ticket triage pipeline using the Jev model. (from Jev model from TypeSafe AI for cheap support-ticket triage)
- Build a coding-task router that scores how hard a task is and sends it to a cheap, fast model or a frontier model. (from Model Routing for Coding: Estimate Task Difficulty, Then Pick a Model)
- Run Qwen3.8 27B at Q4_K_M locally and test it as the backend for a coding agent. (from Qwen3.8 27B Quantization Benchmark: 4-Bit Is Enough for Agentic Coding)
- Build a receipt-parsing pipeline that extracts line items and sub-items, and compare Cohere Parse with frontier LLMs on accuracy and cost. (from Cohere Parse Beats Frontier LLMs at Receipt Parsing)
- Build a tool that turns a design screenshot into code using GLM-5.2 vision. (from GLM-5.2 Vision on Baseten: Turning Images into Code)
- A custom sticker generator that shows how the stickers would look on a laptop or tote bag (from GPT Image 2 Adds Transparent Background Support in the OpenAI API)
- A game where users generate their own sprites in-game, including semi-transparent ones like ghosts (from GPT Image 2 Adds Transparent Background Support in the OpenAI API)
- A tool that generates transparent product images and marketing assets for website mockups (from GPT Image 2 Adds Transparent Background Support in the OpenAI API)
- Benchmark different quantizations (e.g., Q4_K_M vs Q8_0) of a small model on a Colab T4 and compare latency and throughput. (from Run LFM2.5-2.6B locally with llama-cpp-python in Colab)
- Build a local tool-calling agent on top of a small quantized model with llama-cpp-python. (from Run LFM2.5-2.6B locally with llama-cpp-python in Colab)
- Build a document OCR and layout-extraction pipeline with a small local vision model. (from Using LFM2.5-VL-3B's Vision Capabilities (Liquid AI Guide))
- Combine LFM2.5-VL-3B (vision) with LFM2.5-2.6B (text) into a small local multimodal assistant. (from LFM2.5-VL-3B: a lightweight VLM for screens, documents and tool calls)
- Run a tiny language model on a cheap microcontroller such as an ESP32. (from Running a 28.9M-Parameter LLM on an $8 ESP32 Microcontroller)
- Give the same creative coding prompt to several models, including a local quantized one, and compare the results. (from 1-bit Kimi K3 GGUF Running Locally vs Claude Opus 5 and GPT 5.6)
- Train a small character-level GPT on your own text (for example your video scripts) so it writes in your style. Use the video's settings: batch 4, block size 8, 32-dim embeddings, ASCII vocabulary. (from How GPT Works: From Token Embeddings to Multi-Head Attention)
- Build a landing page with two models and compare cost and quality (from GLM 5.2 vs Opus 4.8 for Landing Page Design: 6x Cheaper)
- A live translation app for presentations: the speaker shares a URL or QR code, and attendees join on their phones to hear the talk in their chosen language (one Live API session per language). (from Gemini 3.5 Live Translate: Real-Time Speech Translation via the Live API)
- A translator for multilingual meetings that follows speakers automatically when they switch languages. (from Gemini 3.5 Live Translate: Real-Time Speech Translation via the Live API)
- A multilingual live-presentation app: the speaker shares mic audio, attendees join by URL or QR code, and each listens in their own language through Gemini Live API sessions. (from Gemini 3.5 Live Translate: Real-Time Speech Translation via the Gemini Live API)
- A real-time translation companion for live events such as ceremonies or meetings, where speakers switch between languages. (from Gemini 3.5 Live Translate: Real-Time Speech Translation via the Gemini Live API)
- Fine-tune Gemma 4 12B on your own dataset with Unsloth Studio and run it locally. (from Run Gemma 4 12B on 8GB RAM with Unsloth Dynamic GGUFs)
- Benchmark Gemma 4 speed with and without MTP drafters on your own hardware. (from Gemma 4 Gets Up to 3x Faster with MTP Drafters)
- Build an offline mobile translation app on top of Hy-MT1.5-1.8B-1.25bit. (from Hy-MT1.5-1.8B-1.25bit: A 440MB Offline Phone Translation Model)
- Run Gemma 4 E2B locally (for example in WSL) and test it on your own video clips. (from Running Gemma 4 E2B Video Understanding Locally in WSL)
- Build a small tool-calling assistant on Gemma 4 using system instructions and function calling through the google-genai SDK. (from Using Gemma 4 via the Gemini API and Google AI Studio)
- Build a private, in-browser transcription app on WebGPU using Cohere Transcribe. (from Local Audio Transcription with Cohere Transcribe on WebGPU)
- Build a local document OCR pipeline with Gemma 4 26B running in LM Studio. (from Gemma 4 26B for OCR in LM Studio)
- Build a Linear-style issue tracker clone with an AI model and compare cost and quality across models. (from MiniMax 2.7 One-Shots a Linear Clone at 95% Lower Cost)
- Build a local real-time video captioning app with a small vision-language model like Qwen3.5 0.8B. (from Qwen3.5 0.8B Does Real-Time Local Video Captioning)
- Build a streaming video description tool with a small local vision-language model. (from Qwen3.5 0.8B Real-Time Video Captioning on Mac Studio)
- Build a privacy-preserving speech transcription web app that runs Voxtral fully on-device with Transformers.js and WebGPU. (from Real-Time Speech Transcription in the Browser with Voxtral and WebGPU)
- Python stock analyzer app using the yfinance library, generated by an LLM (from Qwen 3.5 One-Shots a Python Stock Analyzer App with yfinance)
- Build a coding agent powered by GPT-5.3-Codex via the Responses API. (from GPT-5.3-Codex Now Available in OpenAI's Responses API)
- Have a local LLM build a complete browser space-shooter game from one spec prompt, with enemy types, particles, procedural audio, power-ups and boss fights. (from Local Qwen3.5-35B-A3B on a 24GB GPU Builds a Full Game From One Spec)