Researchers present TokenProbe, a framework addressing the high token consumption of Chain-of-Thought reasoning by exploiting the non-uniform value of tokens within reasoning traces. The method distinguishes core tokens carrying decisive content from redundant filler using normalized log probability signals.
- TokenProbe utilizes token-level log probability to identify and compress low-confidence, exploratory tokens.
- It introduces an efficient GRPO objective that selectively removes redundant tokens to improve the accuracy-token efficiency trade-off.
- Empirical results show a 76% reduction in token usage compared to the baseline while maintaining reasoning quality.
- Under matched reasoning-length budgets, the approach outperforms strong baselines such as Gemini-3.1-Pro.
This work demonstrates that selectively compressing redundant tokens yields Pareto improvements, allowing models to achieve higher efficiency without sacrificing performance.