Most AI 'QA agents' re-audit the whole repo every run, re-report the same findings until you stop reading, call flaky and stale tests alike 'failures', and sign off with an LGTM. And once an agent is writing the code, a loop where it also tests its own work has no independent gate left in it.
What I builtA QA agent with a stored baseline, so a repeat run is a delta — NEW, REGRESSED, STILL_OPEN, RESOLVED — with regressions ranked first and every open finding carrying its age. It has no Edit tool by design, quarantines flaky tests only with an expiry date attached, requires a citation before it will write a failure off as a stale expectation, and closes on one of four verdicts where 'blocked' is a legitimate answer. Every line a finding cites is hashed, so the next run knows which evidence the code moved and re-verifies exactly that.
Built withZero-dependency Claude Code plugin, 10 slash commands behind one front door (/verdict:run picks baseline, delta, or a scoped review itself), the same doctrine as five skills for Cursor, Codex, and any other coding agent, a Python MCP server exposing 10 read-only tools, a fact harness that measures counts and diffs so the model contributes judgment only, an exit-code release gate for CI, a local tier that runs unattended nights on an 8B model with zero Claude tokens, a 24-technique test-design catalog, 5 report templates, and six hooks — write-scope, Bash-scope, and a Stop hook that blocks a run which tried to hand-write its own state. MIT.
TestedShips with its own eval: seven scored fixtures — seeded defects, a delta run against an authored history, a TypeScript twin, root-cause with a decoy, a spec review with no code, AI-generated slop, and an adversarial repo whose suite prints ALL TESTS PASSED while exiting 1 — plus a key it did not write: 40 SWE-bench Verified issues, the maintainers' fix located in 36 of 37 scored runs. Published results include the misses and the control: plain Claude Code located those defects just as well in a twelfth of the time, and the repo says so next to the 8/8s. 1,338 tests and a scorer-regression corpus run in plain CI with no model.
ResultPublic and MIT, 73 tagged releases in its first month. Runs unattended nightly against a production codebase — its first fully unattended run refused to execute the suite because a live .env sat in the checkout, said so, and still delivered a delta report, then found a leaked API token untracked in the repo. Pointed at strangers' repositories, it found a rotate_file in boltons that, at keep=1, deletes the file it rotates and a FilePerms bug that had lived 4,595 days, both reported upstream. Turned on itself, it filed the defect that made its own anti-fabrication check imitation-proof: every run now signs the run history with a hash chain a copied state cannot reproduce.