Qwen has released the open weights for its Qwen 3.8 27B model. The checkpoint is available on Hugging Face.
Qwen releases open weights for Qwen 3.8 27B
PrismML releases Bonsai 27B, compressing Qwen3.6-27B to 1-bit and ternary weights
PrismML has released Bonsai 27B, a low-bit quantization of the Qwen3.6-27B model that enables running large multimodal models on laptops and phones without retraining.
Qwen3.5 122B A10B at UD-Q2_K_XL fits in 64GB RAM and improves internal knowledge
A user demonstrates that Qwen3.5 122B A10B quantized with UD-Q2_K_XL can fit into 64GB of system RAM, replacing the previous best option for that constraint.
Ternary Qwen3.6 27B runs at 60 tk/s on dual RTX 3090s
A user reports successfully running the ternary-quantized Qwen3.6 27B model on two NVIDIA RTX 3090 GPUs. The setup achieves a throughput of 60 tokens per second with stable tool calling capabilities.
PrismML compresses Qwen 3.6-27B to <4GB for iPhone
PrismML has developed a compressed version of Alibaba's open-source Qwen 3.6 model, enabling the 27-billion-parameter large language model to run on an iPhone 17 Pro. The startup reduced the model's size from approximately 54 gigabytes to less than 4 gigabytes using a mathematical technique derived from Caltech research.
NVIDIA releases Qwen3.6-27B-NVFP4 model
NVIDIA has released the Qwen3.6-27B-NVFP4 model on Hugging Face.