An independent researcher published KavachBench, a benchmark evaluating the AI coding agent guardrail Kavach against IssueTrojanBench’s corpus of 42 malicious tool-call actions derived from GitHub issue payloads.
- The evaluation utilized IssueTrojanBench's real prompt-injection corpus containing 42 attack actions across 4 categories.
- Results demonstrated that coverage is achieved through adapter canonicalization under a default-deny policy rather than benchmark-specific policy tuning.
- The research harness and paper are available on Zenodo and GitHub for further review.
The findings highlight the importance of distinguishing between canonicalization-based enforcement and specific tuning when evaluating similar guardrails.