Topic · Inference efficiency
lab Hugging Face Blog · 29d ago GPQA Diamond · 67.4% · 39 views

Multiverse Computing's QAH makes 4-bit GPT-OSS outperform its full-precision original

Multiverse Computing introduces Quantization-Aware Healing (QAH), a method that distills compressed, quantized large language models directly from their original pre-compression teachers. Applied to a GPT-OSS 120B model compressed to 60B parameters and quantized to MXFP4, the approach produces a model that beats its own full-precision bfloat16 version on 7 of 9 benchmarks.