Qwen has released the open weights for its Qwen 3.8 27B model. The checkpoint is available on Hugging Face.
Qwen releases open weights for Qwen 3.8 27B
ISTA releases GSQ-RCO GGUFs for Qwen3.8-Flash-Next and 50% expert-pruned Coder build
The ISTA Deep Algorithms and Systems Lab has released quantized GGUF versions of the sparse mixture-of-experts model Qwen3.8-Flash-Next, alongside an experimental capability-targeted build with half its experts removed.
Quantization-Aware Healing recovers 4-bit LLMs faster than QAT
The authors introduce Quantization-Aware Healing (QAH), a pipeline that distills 4-bit students directly from uncompressed models to recover performance lost during structural compression and quantization. Applied to GPT-OSS 120B, this method produces Hypernova-60B, which matches or beats the bfloat16 source on 7 of 9 benchmarks while using roughly 4 times less weight memory.
Quantization-Aware Healing recovers 4-bit LLMs faster than QAT
The authors introduce Quantization-Aware Healing (QAH), a method to recover reasoning and coding capabilities in structurally compressed, 4-bit large language models. Unlike standard quantization-aware training which re-fits the compressed model, QAH distills the 4-bit student directly from the original uncompressed model.
PrismML releases Bonsai 27B, compressing Qwen3.6-27B to 1-bit and ternary weights
PrismML has released Bonsai 27B, a low-bit quantization of the Qwen3.6-27B model that enables running large multimodal models on laptops and phones without retraining.
Qwen3.5 122B A10B at UD-Q2_K_XL fits in 64GB RAM and improves internal knowledge
A user demonstrates that Qwen3.5 122B A10B quantized with UD-Q2_K_XL can fit into 64GB of system RAM, replacing the previous best option for that constraint.