A setup using four 5060 Ti GPUs (totaling $1800) achieves 55 tokens per second with Qwen3.6-27B-FP8, supporting 262K context length and bfloat16 KV cache. The configuration uses P2P and FlashInfer, with benchmark results showing 55.67 output token throughput and 65.25% speculative decoding acceptance rate.
$1800 GPU cost runs Qwen3.6-27B with 262K context and 55 tok/s
qwen4exp adds recurrent state rollback support for MTP speculative decoding
The qwen4exp project has implemented recurrent state rollback support to enable effective Multi-Token Prediction (MTP) speculative decoding. This change allows the target state to move back by the number of rejected draft tokens, preventing unnecessary serialization of the entire recurrent state to host memory.
Qwen3.6 35B-A3B generates flight simulator in single prompt
A Reddit user demonstrated Qwen3.6 35B-A3B generating a complete, relaxing flight simulator with mountains, clouds, and endless procedural terrain from a single prompt.
Qwen 27B local performance on consumer hardware
A user reports that Qwen 27B, quantized to q6kxl and running with multi-token prediction on a system with 4090 and 3090 GPUs, achieves decode speeds of 50-90 tokens/s and pre-fill speeds of 1500-2200 token/s. The model reliably interfaces with various APIs and generates functional code for single-page apps, LaTeX docs, parsers, and crawlers.
LLMs Benchmarked for Web Vulnerability Detection
A study evaluates six LLMs on detecting real-world web vulnerabilities in WordPress plugins, finding detection rates vary by model and prompt design. Claude Opus 4.6 achieved the highest detection rate at 63%, while Qwen 3.5 only reached 35%, and no model consistently identified all baseline vulnerabilities across iterations.
Benchmark Evaluation of Small Language Models for Arabic NLP
A benchmark of 240 Arabic test items across eight domains and ten skills assesses twelve small language models in zero-shot settings. Gemma 3 (12B) achieved the highest overall score (4.548/5), followed by Aya and C4AI Command Arabic, with performance linked more to Arabic alignment and instruction-following than model size. Common failure modes include prompt leakage, hallucination, and weak task adherence.