An audit of ChatGPT, Claude, and Gemini reveals that API evaluation metrics do not reliably transfer to deployed chatbot interfaces. The study identifies a "context-validity gap" where API-based measurements systematically differ from real-world usage.
- API evaluations score 3.4 percentage points higher in accuracy than interface evaluations.
- Test-retest agreement is 2.1 percentage points higher for API access compared to interface access.
- For ChatGPT, the performance drop from switching to an interface exceeds the difference between GPT 5.3 and GPT 5.4.
- Adjusting system prompts, sampling parameters, and reasoning settings via API controls does not reliably eliminate this performance gap.
These findings complicate the use of API evaluations as proxies for deployed systems, indicating that benchmark scores may overstate actual user-facing capabilities.