On September 27, Claude Opus 5 and the 3B model Tetsu identified six separate continuous integration checks in their public repository that reported green or passed despite failing to verify what they were supposed to measure. The issues ranged from a deployment gate refusing valid restarts due to stale digests to timing benchmarks with hollow thresholds that accepted negligible performance differences.

  • A deployment verification script refused all restarts because its core digest pin was five commits stale, training readers to disbelieve red indicators.
  • An expected line count check was never enforced in the actual execution path used by the hourly checks.
  • A media-corpus drift guard asserted regex matches on tool output without ever opening the document it claimed to protect.
  • 252 asynchronous repairs recorded as "started" were never re-measured, costing five days of builds and remaining unproven.
  • A public CI failure on a markdown change was inconsistent, showing three different verdicts for identical bytes.
  • A timing benchmark fix used a minimum-of-five runs that selected the warmest cache, masking ordering effects until interleaved measurements were performed.

The authors emphasize that recording a proof before performing it is a critical error, updating their rules to require breaking checks before documenting them as fixed.