Turboderp has released significant updates to ExLlamaV3, introducing CPU offload capabilities for Mixture of Experts (MoE) experts. The update also includes support for the newly released GLM-5.3-Flash and Qwen3.8-Flash models.
- Implementation of CPU offload specifically for MoE experts.
- Support for Qwen-3.8-Flash-Next with ngram disk offload capabilities.
- Integration of GLM-5.3-Flash via the exl3 format.
- Introduction of a new self-calibrated optimization technique.
- Various other optimizations and improvements to the engine.
These updates expand the hardware flexibility and model compatibility of ExLlamaV3, allowing users to leverage CPU resources for specific model components.