The article introduces Bonsai 27B, which is described as the first model of its class capable of running on a mobile phone.
No further details or specific features are provided in the source text.
The article introduces Bonsai 27B, which is described as the first model of its class capable of running on a mobile phone.
No further details or specific features are provided in the source text.
The article introduces Quantization-Aware Healing, a method for creating a compressed 4-bit model that achieves performance superior to its full-precision original.
Meta has released MobileMoE, a collection of on-device Mixture-of-Experts (MoE) language models designed to optimize the quality-efficiency trade-off for mobile devices. The family includes three model scales with sub-billion active parameters and weight footprints under 3 GB in INT4 format.
The authors introduce Quantization-Aware Healing (QAH), a pipeline that distills 4-bit students directly from uncompressed models to recover performance lost during structural compression and quantization. Applied to GPT-OSS 120B, this method produces Hypernova-60B, which matches or beats the bfloat16 source on 7 of 9 benchmarks while using roughly 4 times less weight memory.
The authors introduce Quantization-Aware Healing (QAH), a method to recover reasoning and coding capabilities in structurally compressed, 4-bit large language models. Unlike standard quantization-aware training which re-fits the compressed model, QAH distills the 4-bit student directly from the original uncompressed model.
Joakimpalm-Zen has updated Xyntetik Runner to train LoRA adapters directly using the same quantized GGUF weights it serves for inference, eliminating the need for a separate FP16 training copy or runtime.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy