STAGE is a pipeline that generates text-to-JSON training data by using LLMs to synthesize reports and JSON schemas, validated against underlying spreadsheets. Evaluations on STAGE-Eval show it improves Qwen3-4B exact match from 31.37% to 74.27% and value accuracy from 45.46% to 90.69%.
STAGE: Source-Grounded Data Generation for Text-to-JSON
The economics of AI are starting to favor open models
Recent AI model releases show that high-intelligence, low-cost models are increasingly dominated by open-weight models like DeepSeek, Qwen, GLM, Kimi, and MiniMax. For most real-world applications, the performance gap between frontier closed models and strong open models is shrinking faster than cost differences, making open models competitive in terms of both capability and price.
Alibaba announces Qwen 3.8 Max, a 2.4T-parameter model with open weights coming next week
Alibaba's Qwen team has announced Qwen 3.8 Max, a new 2.4T-parameter flagship model focused on coding, long-horizon agentic work, and multimodal reasoning. The company confirmed that open-weight versions of both Qwen 3.8 Max and the smaller Qwen 3.8-27B will be released next week.
Qwen3.8 open-weight release, Kimi Code CLI, Netflix LLM stack, Alibaba chip software
Alibaba announced the open-weight release of Qwen3.8, a 2.4-trillion-parameter model, while also open-sourcing its Zhenwu AI chip software stack to reduce reliance on Nvidia's CUDA ecosystem.
Achilles1089 releases uncensored Qwen Fable hybrid model
User Achilles1089 has uploaded an uncensored "fable and opus 4.8" hybrid model to Hugging Face, claiming the numbers demonstrate its effectiveness.
AGC-Bench introduces unified benchmark and AGC-Judge to measure artificial general creativity
Researchers introduce AGC-Bench, a unified benchmark for artificial general creativity constructed from 3,101 screened papers and covering 78 datasets across domains like brainstorming and STEM. To address bias in automated evaluation, the team fine-tunes Qwen3-30B on bias-corrected ratings to create AGC-Judge, an open-weight model that robustly scores new creativity benchmarks.