Researchers present TokenProbe, a framework that addresses the high token consumption of Chain-of-Thought reasoning by exploiting the non-uniform value of tokens within a sequence. The method uses normalized log probability to distinguish core structural tokens from redundant filler, allowing for selective compression.
- TokenProbe identifies core tokens carrying decisive reasoning content versus low-confidence exploratory filler using log probability signals.
- It introduces an efficient GRPO objective that selectively compresses redundant tokens to improve the accuracy-token efficiency trade-off.
- The approach reduces token usage by 76% compared to the baseline while maintaining reasoning quality.
- Under matched reasoning-length budgets, TokenProbe outperforms strong flagship baselines like Gemini-3.1-Pro.
This work demonstrates that compressing low-value tokens yields Pareto improvements, enabling efficient reasoning without sacrificing performance.