The vLLM project released version 0.29.0, which includes a core bugfix regarding the dense prefix cache.

  • The update applies the dense prefix cache default setting specifically to hybrid models.