Turboderp has released significant updates to ExLlamaV3, introducing CPU offload capabilities for Mixture of Experts (MoE) experts. The update also includes support for the newly released GLM-5.3-Flash and Qwen3.8-Flash models.

  • Implementation of CPU offload specifically for MoE experts.
  • Support for Qwen-3.8-Flash-Next with ngram disk offload capabilities.
  • Integration of GLM-5.3-Flash via the exl3 format.
  • Introduction of a new self-calibrated optimization technique.
  • Various other optimizations and improvements to the engine.

These updates expand the hardware flexibility and model compatibility of ExLlamaV3, allowing users to leverage CPU resources for specific model components.