Anthropic has released new evaluation tooling for Claude Code, introducing a `claude-api` plugin with `build_eval` and `hill-climb` commands designed to help developers build evaluations, check graders, and improve applications against them.
- The tool automatically discovers issues in conversation traces, such as problems with human handoff, formatting, and voice agents.
- It allows users to create evaluators for specific failure modes, though the reviewer noted the initial scope was too broad.
- The workflow includes commands to check graders and iteratively improve application performance based on evaluation results.
The author of the review suggests holding off on using the tool until it better supports data exploration before committing to an evaluator, noting that the current interface for reviewing labels is cumbersome.