Researchers introduce ReFold, a training-free rendering layer that preserves interaction history while compressing the model's rendered context for long-horizon LLM agents. It removes inter-turn redundancy by replacing displayed content with stubs and folding finished turns into notes, using chunked rewriting to maintain prefix caches.

  • Removes content already shown in earlier turns and folds agent-reported finished turns into one-line notes.
  • Uses chunked rendering to rewrite the cached prefix once every few steps instead of at every step.
  • Ensures all removals are strictly reversible, allowing restoration from history if a removal is incorrect.
  • Reduces token consumption by up to 2.5x and halves KV-cache memory per session without degrading task success rates.
  • Avoids up to 92% of forced compactions under capped context budgets and reduces request queuing delays by up to 100%.

ReFold operates as a plug-and-play solution across standard ReAct-style harnesses, accelerating inference by up to 1.7x while cutting inference costs by up to 3.4x.