The ISTA Deep Algorithms and Systems Lab has released quantized GGUF versions of the sparse mixture-of-experts model Qwen3.8-Flash-Next, alongside an experimental capability-targeted build with half its experts removed.
- The release includes four quantized GGUFs ranging from 2.40 to 3.50 bpw (66.4 to 83.6 GB) using GSQ and RCO techniques.
- At 3.50 bpw, the model matches the BF16 base on every benchmark evaluated, including GPQA-Diamond and LiveCodeBench v6.
- The expert-pruned Coder build removes 256 of 512 routed experts per layer, resulting in an average bitwidth of 1.89 bits per parameter.
- This pruning reduces the resident working set to 29.6 GB, allowing the 176.9B-parameter model to fit within a single 32 GB accelerator.
The quantized models maintain high performance relative to their size, while the pruned Coder build retains 91.3% of SWE-bench Verified and 98.7% of LiveCodeBench v6 scores compared to the BF16 baseline.