AI-Written Code Still Needs Human Testing
Melvin Vivas · X post · 2026-08-21 · Open on X
Topics: Evaluation (Evals) & Testing, AI Dev Tools & Productivity · Level: beginner
Summary
Software development isn't just prompting. A hidden cost of AI coding is the time developers spend checking that the generated code actually works. Browser automation helps, but people still need to test the user experience by hand.
Key points
- Plan time for verifying AI-generated code. It is a real, often overlooked part of the work.
- Automated browser testing is useful but doesn't replace human testing.
- Test the user experience by hand to see whether the product actually works for users.
Try this
- Add a manual UX testing step after AI-generated features, on top of automated browser tests.
More in Evaluation (Evals) & Testing
- Pipette: Open-Source Benchmarking for On-Device Models
- Idea: A Benchmark for Codex Usage-Limit Consumption
- Why Evaluation Is the Skill That Sets Top AI Engineers Apart
- LoopsBench: Testing Coding Agents on Long-Horizon Tasks
- Agent Loops Need Real Evals and Feedback
- OCR Testing a Small Vision Model with LLM-Made Ground Truth