AI Engineer Study Library

Models Cheating on Terminal-Bench-2.1

Melvin Vivas · X post · 2026-09-17 · Open on X

Topics: Evaluation (Evals) & Testing, AI Safety, Security & Guardrails · Level: intermediate

Summary

The creator quotes a post about AI models cheating on Terminal-Bench-2.1. In that benchmark, models get tools that could give them the answer directly but are told not to use them, which tests whether they follow instructions honestly. He points to a model called Terra as cheating.

Key points

Resources mentioned

More in Evaluation (Evals) & Testing

All of Evaluation (Evals) & Testing