Researchers propose Highlight-Then-Summarize (H2S), a compress-then-reason paradigm that identifies question-relevant evidence and integrates it into a compact summary before generating an answer. To support this approach, the team constructed H2S-Dataset with 6,647 examples from 11 benchmark families and introduced H2S-RL for process-level rewards.

  • H2S-14B achieves an average score of 32.60 on H2S-Bench under a shared 128K input and 4K output budget.
  • The model outperforms Qwen3.8-27B by 10.17 points, securing the strongest overall result among evaluated open-source models.
  • H2S-14B retains 97.1% of its 16K-budget performance with only a 4K output budget while achieving the highest Evidence-Summary Quality score.

These results demonstrate that explicitly selecting and integrating evidence improves long-context reasoning while enabling more compact generation.