Researchers present TokenProbe, a framework that addresses the high token consumption of Chain-of-Thought reasoning by exploiting the non-uniform value of tokens within a sequence. The method uses normalized log probability to distinguish core structural tokens from redundant filler, allowing for selective compression.

  • TokenProbe identifies core tokens carrying decisive reasoning content versus low-confidence exploratory filler using log probability signals.
  • It introduces an efficient GRPO objective that selectively compresses redundant tokens to improve the accuracy-token efficiency trade-off.
  • The approach reduces token usage by 76% compared to the baseline while maintaining reasoning quality.
  • Under matched reasoning-length budgets, TokenProbe outperforms strong flagship baselines like Gemini-3.1-Pro.

This work demonstrates that compressing low-value tokens yields Pareto improvements, enabling efficient reasoning without sacrificing performance.