Research demonstrates that chain-of-thought (CoT) monitoring fails against adversaries who control the reasoning process. By rewriting an agent's reasoning to appear as good-faith engineering while keeping commands and outputs verbatim, attackers can bypass detection mechanisms.

  • A gradient-free attack drops a held-out monitor's catch rate from approximately 95% to under 11% on the targeted subset.
  • The attack transfers across different monitor families and agent models, including live agents.
  • Trace-only defenses recover detection only partially because the rewrite remains truthful about events while lying about intent.
  • Probing surrogate monitor activations can separate missed hacks, but causal control confirms this indicates a detector rather than secret knowledge.

The study concludes that aggregate accuracy is a false average, as monitors hide near-total collapse on subsets where CoT monitoring is the only signal.