NVIDIA has released the Qwen3.6-27B-NVFP4 model on Hugging Face.
The checkpoint is now available for download.
NVIDIA has released the Qwen3.6-27B-NVFP4 model on Hugging Face.
The checkpoint is now available for download.
The authors introduce Quantization-Aware Healing (QAH), a pipeline that distills 4-bit students directly from uncompressed models to recover performance lost during structural compression and quantization. Applied to GPT-OSS 120B, this method produces Hypernova-60B, which matches or beats the bfloat16 source on 7 of 9 benchmarks while using roughly 4 times less weight memory.
The authors introduce Quantization-Aware Healing (QAH), a method to recover reasoning and coding capabilities in structurally compressed, 4-bit large language models. Unlike standard quantization-aware training which re-fits the compressed model, QAH distills the 4-bit student directly from the original uncompressed model.
NVIDIA has demonstrated serving the Qwen 3.8 2.4T parameter model with configurable reasoning capabilities on its GB300 NVL72 hardware platform.
Qwen has released the open weights for its Qwen 3.8 27B model. The checkpoint is available on Hugging Face.
Alibaba has released the open weights for Qwen3.8-2.4T-A95B (also known as Qwen3.8-Max), its largest open-weight model, designed to bring near-frontier capabilities to the open ecosystem.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy