The article introduces a study design for "drills," which are small, runnable exercises attached to agent skills. These drills verify that a skill functions correctly against the live substrate rather than relying on static documentation.

  • Drills run against real APIs and venues to ensure the skill matches current reality.
  • A passing drill confirms the skill works; a failing one indicates the underlying system has changed.
  • Misses automatically trigger updates to the skill playbook, ensuring it stays sharp.
  • New agents must pass these drills to earn skills, proving capability through execution rather than reading.

This approach ensures that agent capabilities remain valid and up-to-date by continuously testing them against live environments.