Researchers propose Highlight-Then-Summarize (H2S), a compress-then-reason paradigm that identifies question-relevant evidence and integrates it into a compact summary before generating an answer. The team constructs H2S-Dataset with 6,647 examples and introduces H2S-RL to provide process-level rewards for evidence selection.
- H2S-14B achieves an average score of 32.60 on the seven-task H2S-Bench suite under a shared 128K input and 4K output budget.
- The model outperforms Qwen3.8-27B by 10.17 points, obtaining the strongest overall result among evaluated open-source models.
- H2S-14B retains 97.1% of its 16K-budget performance with only a 4K output budget while achieving the highest Evidence-Summary Quality score.
These results demonstrate that explicitly selecting and integrating evidence improves long-context reasoning while enabling more compact generation.