User Achilles1089 has uploaded an uncensored "fable and opus 4.8" hybrid model to Hugging Face, claiming the numbers demonstrate its effectiveness.
The model is.
User Achilles1089 has uploaded an uncensored "fable and opus 4.8" hybrid model to Hugging Face, claiming the numbers demonstrate its effectiveness.
The model is.
Alibaba's Qwen team has announced Qwen 3.8 Max, a new 2.4T-parameter flagship model focused on coding, long-horizon agentic work, and multimodal reasoning. The company confirmed that open-weight versions of both Qwen 3.8 Max and the smaller Qwen 3.8-27B will be released next week.
Alibaba announced the open-weight release of Qwen3.8, a 2.4-trillion-parameter model, while also open-sourcing its Zhenwu AI chip software stack to reduce reliance on Nvidia's CUDA ecosystem.
LLM providers currently subsidize expensive API usage to build ecosystems, planning to raise prices later. As subsidies dwindle, users may face steep price increases—like $2k per month—making access costly and threatening widespread adoption, especially for individuals relying on affordable hardware to run models.
A user shares optimized settings for running Qwen 3.6 27B with Q8_0 quantization on an RTX 4090 and RTX 3090 setup using llama.cpp. The configuration includes tensor split, 999 layers on GPU, 250k context, speculative decoding, and unified KV cache, achieving 75-100t/s throughput with vision and MTP support.
Recent AI model releases show that high-intelligence, low-cost models are increasingly dominated by open-weight models like DeepSeek, Qwen, GLM, Kimi, and MiniMax. For most real-world applications, the performance gap between frontier closed models and strong open models is shrinking faster than cost differences, making open models competitive in terms of both capability and price.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy