Topic 11 of 16 in the learning path
AI Safety, Security & Guardrails
Prompt injection, red-teaming, guardrails, privacy and responsible AI.
Reels and posts (28)
- Abliterated Local Models from DealignAI on Hugging Face · Melvin Vivas, X: The creator points to DealignAI's Hugging Face page, which has many abliterated models you can run locally.
- Run an abliterated Qwen3.8-27B GGUF locally with llama-server · Melvin Vivas, X: The creator shares an uncensored (abliterated) GGUF build of Qwen3.8-27B on Hugging Face and says to use it only for education or red teaming.
- Abliteration.ai: Hosted Abliterated (Uncensored) Models · Melvin Vivas, X: The creator points to Abliteration.ai, which hosts 'abliterated' models if you need them.
- OpenAI Astra: More Capable Cyber Model Than Sol · Melvin Vivas, X: The creator says OpenAI's Astra is far ahead of Sol.
- Teaching Teens AI Safety: ChatGPT for Teens in Singapore · Melvin Vivas, X: The creator, who sits on a parent support group committee at a Singapore secondary school, asks for an OpenAI contact in Singapore to run a student talk on AI safety and awareness.
- Screening LLM Prompts and Responses on CPU with AIBackends + GliGuard · Melvin Vivas, X: AIBackends v0.5.0, the creator's open-source Python library, adds support for Fastino's GliGuard 300M, a small guardrail model that runs on CPU.
- Codex safeguards against accidental destructive actions · Melvin Vivas, X: The creator shares an update from Tibo at OpenAI about changes that lower the risk of Codex doing destructive things, like accidental deletions, while it works.
- Liquid AI's Multilingual PII Detection Works on Japanese · Melvin Vivas, X · 0:12: A short X post with a 12-second video and no speech.
- Use a Local Model for Confidential Data with Your Agent · Melvin Vivas, X: If you give your Hermes agent confidential data such as bank statements, switch it to a local model so the data stays on your machine.
- Don't Let Coding Agents Run Database Migrations Unsupervised · Melvin Vivas, X: A warning against letting an AI agent run Prisma migrations automatically, because it can break your database.
- Forbes: Did Chinese Open AI Save Hugging Face After the OpenAI Hack? · Melvin Vivas, X: This post links a Forbes article (July 22, 2026) about the Hugging Face cyberattack incident involving an OpenAI model.
- OpenAI on Why Teens Deserve Access to Safe AI · Melvin Vivas, X: The creator shares and supports OpenAI's article 'Why teens deserve access to safe AI', about giving teenagers access to AI with safety protections.
- Codex file deletions: why sandboxing matters for coding agents · Melvin Vivas, X: Shares OpenAI's root-cause finding on reports that GPT-5.6 in Codex unexpectedly deleted files.
- Keep Production .env Files Out of Your Coding-Agent Workspace · Melvin Vivas, X: Melvin reacts to a report that GPT-5.6 Sol deleted a user's production database.
- Switching Agent Backends Over Zero Data Retention (ZDR) Concerns · Melvin Vivas, X: The creator moved his Hermes agent from Grok (via OAuth) to Codex.
- Grok Build uploaded whole codebases: privacy fix and /privacy · Melvin Vivas, X: The creator quotes an xAI response about Grok Build's coding CLI.
- Turn off Grok Build's default data sharing with /privacy · Melvin Vivas, X: A quick warning that Grok Build shares your data by default.
- Coding Agent Risk: GPT 5.6 Sol Deleted a Mac User Directory · Melvin Vivas, X: Melvin Vivas shares Matt Shumer's report that GPT-5.6-Sol, running as a coding agent in Codex, accidentally deleted almost all the files on his Mac.
- Dario Amodei's Essay 'The Adolescence of Technology' · Melvin Vivas, X: The creator shares Anthropic CEO Dario Amodei's essay 'The Adolescence of Technology', subtitled 'Confronting and Overcoming the Risks of Powerful AI'.
- Devin Security Review from Cognition · Melvin Vivas, X: The creator points to Cognition's announcement of Devin Security Review, an AI-assisted security review feature.
- Local PII Redaction on CPU with the AIBackends Python Package · Melvin Vivas, X · 0:14: A 14-second X post from Melvin Vivas showing AIBackends, a Python package for removing personally identifiable information (PII) from text.
- OpenAI Privacy Filter: An Open-Weights PII Detection Model · Melvin Vivas, X: OpenAI released Privacy Filter, an open-weights model for detecting personally identifiable information (PII).
- OpenAI Privacy Filter for Local PII Detection and Redaction · Melvin Vivas, X: OpenAI released Privacy Filter, a small open-weight model that detects and redacts personally identifiable information (PII).
- Keep a human in the loop: question and steer AI output · Melvin Vivas, X: Advice to slow down and review AI output carefully instead of accepting it as-is.
- Detect PII and PHI with NVIDIA's GLiNER-PII Model · Melvin Vivas, X: Melvin shares NVIDIA's GLiNER-PII model on Hugging Face.
- Claude Mythos Model Card: Anthropic's Limited Preview for Cyber Defenders · Melvin Vivas, X: Shares the model card for Claude Mythos, Anthropic's new frontier model.
- Cautionary Tale: Claude Code Wiped a Production Database via Terraform · Melvin Vivas, X: A quoted post from DataTalksClub reports that Claude Code ran a Terraform command that deleted their production database.
- X Paid Partnerships Policy and AI-Content Disclosure · Melvin Vivas, X: The creator points to X's paid partnerships policy, which goes with X's new option to label posts as paid partnerships or made with AI.
Watch, free (1)
- DonvitoAI post on X (AIBackends PII redaction demo) · video · x.com · free
The original X post with the short video demo of AIBackends redacting PII locally.
Mentioned in: Local PII Redaction on CPU with the AIBackends Python Package (Melvin Vivas on X · notes)
Read and use (27)
- Docker Sandboxes · tool · docker.com · free
Local microVM-isolated sandboxes for safely running AI coding agents on your laptop.
Mentioned in: Docker Sandboxes for Running Coding Agents Locally (Melvin Vivas on X · notes), Docker Cloud Sandboxes: Run AI Agents in Local or Cloud microVMs (Melvin Vivas on X · notes), Docker Sandboxes (Melvin Vivas on X · notes), Docker Sandboxes Include an Interactive Dashboard (Melvin Vivas on X · notes) and 2 more - Codex Security Cloud · tool · learn.chatgpt.com · paid
A new Codex product that runs in cloud environments and helps defenders harden their infrastructure.
Mentioned in: OpenAI DevDay 2026 Recap: Dots, Agents API, Codex Cloud & Marketplace (Melvin Vivas on X · notes), OpenAI DevDay recap: Dots, GPT-6.1 Sol, Codex Cloud, Agents API (Melvin Vivas on X · notes), OpenAI DevDay 2026 Replay: Timestamped Guide to the Announcements (Melvin Vivas on X · notes) - GLiNER-PII · tool · huggingface.co · free
NVIDIA's GLiNER-based model that finds personal data (PII) in text, so it can be redacted locally.
Mentioned in: aibackends 0.3.0: Model Caching Speeds Up PII and OCR Inference (Melvin Vivas on X · notes), Running AI Locally: Rebuilding AIBackends as a Python Library with Open Models (Melvin Vivas on X · notes), Detect PII and PHI with NVIDIA's GLiNER-PII Model (Melvin Vivas on X · notes) - GliGuard 300M · tool · huggingface.co · free
Fastino Labs' small guardrail model that runs on CPU and classifies prompts and responses for safety, toxicity, jailbreaks and refusals.
Mentioned in: Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails (Melvin Vivas on X · notes), Screening LLM Prompts and Responses on CPU with AIBackends + GliGuard (Melvin Vivas on X · notes) - OpenAI Privacy Filter · tool · openai.com · free
Small open-weight OpenAI model for context-aware PII detection and redaction.
Mentioned in: OpenAI Privacy Filter: An Open-Weights PII Detection Model (Melvin Vivas on X · notes), OpenAI Privacy Filter for Local PII Detection and Redaction (Melvin Vivas on X · notes) - Abliteration.ai · tool · x.com · check price
Service that hosts abliterated (refusal-removed) language models.
Mentioned in: Abliteration.ai: Hosted Abliterated (Uncensored) Models (Melvin Vivas on X · notes) - Adversarial Prompting (Prompt Engineering Guide) · article · promptingguide.ai · free
How to defend against prompt injection and jailbreaks.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - AI Safety, Ethics, and Society Textbook (AI Safety Book) · book · aisafetybook.com · free
Free online textbook covering the concepts of AI safety, ethics and societal impact.
Mentioned in: Two Free Resources to Learn AI Engineering: AI Engineering Hub & AI Safety Book (Bashiri Smith on Facebook · notes) - aibackends 0.5.0 on PyPI · docs · pypi.org · free
The PyPI package page for aibackends version 0.5.0.
Mentioned in: Screening LLM Prompts and Responses on CPU with AIBackends + GliGuard (Melvin Vivas on X · notes) - Claude Mythos Model Card · pdf · www-cdn.anthropic.com · free
Anthropic's official model card for Claude Mythos, covering its capabilities, evaluations and safety approach.
Mentioned in: Claude Mythos Model Card: Anthropic's Limited Preview for Cyber Defenders (Melvin Vivas on X · notes) - Dario Amodei (@DarioAmodei) on X · person · x.com · free
Anthropic CEO who writes about AI safety and the future of AI.
Mentioned in: Dario Amodei's Essay 'The Adolescence of Technology' (Melvin Vivas on X · notes) - Dario Amodei — The Adolescence of Technology · article · darioamodei.com · free
Dario Amodei's essay on confronting and overcoming the risks of powerful AI.
Mentioned in: Dario Amodei's Essay 'The Adolescence of Technology' (Melvin Vivas on X · notes) - Devin Security Review (Cognition announcement) · article · x.com · free · open in a browser to verify
Cognition's X post announcing Devin Security Review; the link now returns a 404 error.
Mentioned in: Devin Security Review from Cognition (Melvin Vivas on X · notes) - EU AI Act · pdf · eur-lex.europa.eu · free
The European Union's law regulating AI systems by risk level.
Mentioned in: 17-Step AI Engineer Roadmap: From Basic RAG to Agents, Evals, LLMOps & Governance (Bashiri Smith on Facebook · notes) - GliGuard Moderation Colab Notebook · tool · colab.research.google.com · free
A Google Colab notebook showing GliGuard moderation of prompts and responses with AIBackends.
Mentioned in: Screening LLM Prompts and Responses on CPU with AIBackends + GliGuard (Melvin Vivas on X · notes) - Guardrails AI · tool · github.com · free
Listed under guardrails tools for enforcing quality and safety on every request at runtime.
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook - huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF · other · huggingface.co · free
An abliterated (uncensored) GGUF build of Qwen3.8-27B on Hugging Face that runs locally with llama.cpp.
Mentioned in: Run an abliterated Qwen3.8-27B GGUF locally with llama-server (Melvin Vivas on X · notes) - Introducing ChatGPT for Teens: Built for learning, backed by protections · article · openai.com · free
OpenAI's announcement of a teen-focused ChatGPT with safety protections for learning.
Mentioned in: Teaching Teens AI Safety: ChatGPT for Teens in Singapore (Melvin Vivas on X · notes) - LawZero Scientist AI · other · lawzero.org · free · open in a browser to verify
AI safety research on 'Scientist AI' from LawZero, the organization led by Yoshua Bengio.
Mentioned in: AI News Roundup: Claude Fable 5, Scientist AI, ZCode, NVIDIA RL (Melvin Vivas on X · notes) - NeMo Guardrails · tool · github.com · free
Listed under guardrails tools for enforcing quality and safety on every request at runtime.
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook - NIST AI Risk Management Framework (AI RMF) · pdf · nist.gov · free
US NIST framework for identifying and managing risks of AI systems.
Mentioned in: 17-Step AI Engineer Roadmap: From Basic RAG to Agents, Evals, LLMOps & Governance (Bashiri Smith on Facebook · notes) - OpenAI Astra · tool · openai.com · paid
Upcoming OpenAI model with advanced cybersecurity capabilities.
Mentioned in: OpenAI Astra: More Capable Cyber Model Than Sol (Melvin Vivas on X · notes) - OpenAI private intelligence · tool · openai.com · paid
An OpenAI offering, in preview, for stronger data controls, with zero data retention (ZDR) and private safety processing.
Mentioned in: OpenAI DevDay 2026 Recap: Dots, Agents API, Codex Cloud & Marketplace (Melvin Vivas on X · notes) - OWASP Top 10 for LLMs · website · genai.owasp.org · free
The security checklist for LLM apps.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - PII Guardian (Hugging Face Space) · website · huggingface.co · free
A Hugging Face Space where you can try the GLiNER-PII model in your browser.
Mentioned in: Detect PII and PHI with NVIDIA's GLiNER-PII Model (Melvin Vivas on X · notes) - Why teens deserve access to safe AI · article · openai.com · free
OpenAI's article on giving teenagers access to AI with safety protections.
Mentioned in: OpenAI on Why Teens Deserve Access to Safe AI (Melvin Vivas on X · notes) - X Paid Partnerships Policy · docs · help.x.com · free · open in a browser to verify
X's official rules for disclosing paid partnerships, with country-specific requirements.
Mentioned in: X Paid Partnerships Policy and AI-Content Disclosure (Melvin Vivas on X · notes)
Build
- Use the abliterated model for local red-teaming experiments and compare its refusals with the original Qwen model. (from Run an abliterated Qwen3.8-27B GGUF locally with llama-server)
- Add a CPU-based guardrail step that screens user prompts and model responses in your LLM app (from Screening LLM Prompts and Responses on CPU with AIBackends + GliGuard)
- Add a local PII-redaction step with AIBackends ahead of an LLM chatbot or RAG ingestion pipeline, so personal data never reaches the model or the logs. (from Local PII Redaction on CPU with the AIBackends Python Package)
- Add a local PII-redaction step with Privacy Filter before sending user data to an LLM. (from OpenAI Privacy Filter for Local PII Detection and Redaction)
- Build a PII/PHI redaction guardrail that cleans user input before it reaches an LLM. (from Detect PII and PHI with NVIDIA's GLiNER-PII Model)