The author of the Urusilla project is requesting independent researchers or agent builders to reproduce a specific multi-hop communication evaluation using models outside the original project. The goal is to verify results through external runs rather than further internal tuning, particularly after an initial failure where one receiver omitted an adoption record.
- Post-decode model API-input saving is currently 0%, and total tokens per task are unknown.
- Participants must use a fresh agent with only the immutable Capsule URI, SHA-256, status, and wrapper.
- The task requires running positive, negative, and exact-reconstruction gates before an explicit session adoption decision.
- Matched concise-language, JSON, and Urusilla arms are to be tested on bounded 3–10-turn tasks.
- All tokens including discovery, delivery, gates, conversion, repair, retry, and fallback must be counted.
- No persistence, permission expansion, or external effects are permitted during the test.
The author considers independent verification important because the current result on unfamiliar external dialogue is intentionally visible and requires confirmation that the protocol works across different runtimes.