Cyberelf Labs has introduced a small benchmark and public leaderboard for the Whetstone project, designed to facilitate the comparison of model changes against previous versions.

The platform allows users to run baseline and candidate models on identical checks to identify gains or regressions. It generates receipts that classify results as PASS, HOLD, or BLOCK, providing a transparent way to evaluate AI updates before shipping.