A test comparing Q8 and IQ3 XXS turbo4 quantized versions of Qwen3.6 27B shows that Q8 excels in API safety and input sanitization, while IQ3 XXS turbo4 performs better in thread management and modular code design. The model recommends merging both approaches: using Q8 for initial launch protection and IQ3 XXS for atomic writes and thread lifecycle, forming a combined Phase 1 foundation.
Qwen3.6 27B Quantization Performance Test Results
Best Settings for 48GB VRAM with Qwen 3.6 27B
A user shares optimized settings for running Qwen 3.6 27B with Q8_0 quantization on an RTX 4090 and RTX 3090 setup using llama.cpp. The configuration includes tensor split, 999 layers on GPU, 250k context, speculative decoding, and unified KV cache, achieving 75-100t/s throughput with vision and MTP support.
ISTA releases GSQ-RCO GGUFs for Qwen3.8-Flash-Next and 50% expert-pruned Coder build
The ISTA Deep Algorithms and Systems Lab has released quantized GGUF versions of the sparse mixture-of-experts model Qwen3.8-Flash-Next, alongside an experimental capability-targeted build with half its experts removed.
qwen4exp adds recurrent state rollback support for MTP speculative decoding
The qwen4exp project has implemented recurrent state rollback support to enable effective Multi-Token Prediction (MTP) speculative decoding. This change allows the target state to move back by the number of rejected draft tokens, preventing unnecessary serialization of the entire recurrent state to host memory.
Quantization-Aware Healing recovers 4-bit LLMs faster than QAT
The authors introduce Quantization-Aware Healing (QAH), a pipeline that distills 4-bit students directly from uncompressed models to recover performance lost during structural compression and quantization. Applied to GPT-OSS 120B, this method produces Hypernova-60B, which matches or beats the bfloat16 source on 7 of 9 benchmarks while using roughly 4 times less weight memory.
Quantization-Aware Healing recovers 4-bit LLMs faster than QAT
The authors introduce Quantization-Aware Healing (QAH), a method to recover reasoning and coding capabilities in structurally compressed, 4-bit large language models. Unlike standard quantization-aware training which re-fits the compressed model, QAH distills the 4-bit student directly from the original uncompressed model.