Researchers introduce ReFold, a training-free rendering layer that compresses the model's rendered context in long-horizon LLM agents without discarding the underlying interaction history. It removes inter-turn redundancy by replacing displayed content with stubs and folding finished turns into notes, using chunked rendering to rewrite the cached prefix periodically.

  • ReFold eliminates two types of inter-turn redundancy: previously displayed content and agent-reported finished turns.
  • The method uses chunked rendering to update the cached prefix every few steps rather than at every step.
  • All removals are strictly reversible, allowing restoration from history if a removal is incorrect.
  • Evaluations across five benchmarks and two frontier LLMs show up to 2.5x reduction in token consumption and halved KV-cache memory.
  • Under capped context budgets, ReFold avoids up to 92% of forced compactions.
  • In concurrent serving workloads, it reduces request queuing delays by up to 100%, accelerates inference by up to 1.7x, and cuts costs by up to 3.4x.

ReFold is designed as a plug-and-play solution for standard ReAct-style harnesses, addressing the limitations of existing predictive context management methods that introduce runtime overhead and invalidate prefix caches.