A peer-to-peer ledger with a fail-closed ethics gate composed of a distilled bag-of-words classifier, a semantic seat, and a panel of small open models (qwen2.5:7b, llama3.2:3b, gemma2:2b) reveals that judges sharing features exhibit correlated errors rather than independent ones.

  • Among 1,226 labelled violations, the distilled judge and a same-features naive Bayes model wrongly admitted the same transaction 4.3% of the time, approximately 7.7 times higher than the 0.56% expected if their errors were independent.
  • The evaluation shows the distilled judge wrongly admits 3.4% of labelled violations (2.6 to 4.3) compared to 23.9% for the same-features model forced to decide, while incurring a 34.8% abstention rate.
  • The dataset of 3,444 labels was generated by a panel of three small open models with Fleiss’ kappa of 0.857 over 374 voted rows, with no human labelling involved.

The findings indicate that diversity of implementation does not guarantee independence of error for judges sharing features and training data, challenging the assumption of independence in majority-of-graders designs.