Benchmark · agentic
Berkeley Function-Calling Leaderboard
Berkeley's function/tool-calling accuracy leaderboard.
| Date | Model | Score | Source |
|---|---|---|---|
| 2026-08-13 | LFM2.5-VL-3B | 32.5% | Liquid AI releases LFM2.5-VL-3B, a 3B on-device vision-language model with tool calling |
| 2026-08-12 | LFM2.5-VL-3B | 32.5% | LFM2.5-VL-3B improves screen understanding and function calling for edge deployment |
| 2026-08-08 | Pokee-Isaac 28B | 70.94% | Pokee AI releases Pokee-Isaac 28B, a 10M-token context model for in-boundary deployment |
| 2025-08-06 | GLM-4.5 | 90.6% | Zhipu AI releases GLM-4.5 models matching Claude and DeepSeek |
| 2025-08-06 | GLM-4.5-Air | 76.4% | Zhipu AI releases GLM-4.5 models matching Claude and DeepSeek |
| 2024-12-17 | Falcon3-10B-Instruct | 86.3% | Falcon3 family releases five open models with improved science, math, and code capabilities |