A setup using four 5060 Ti GPUs (totaling $1800) achieves 55 tokens per second with Qwen3.6-27B-FP8, supporting 262K context length and bfloat16 KV cache. The configuration uses P2P and FlashInfer, with benchmark results showing 55.67 output token throughput and 65.25% speculative decoding acceptance rate.
$1800 GPU cost runs Qwen3.6-27B with 262K context and 55 tok/s
Qwen3.6 35B-A3B generates flight simulator in single prompt
A Reddit user demonstrated Qwen3.6 35B-A3B generating a complete, relaxing flight simulator with mountains, clouds, and endless procedural terrain from a single prompt.
Qwen 27B local performance on consumer hardware
A user reports that Qwen 27B, quantized to q6kxl and running with multi-token prediction on a system with 4090 and 3090 GPUs, achieves decode speeds of 50-90 tokens/s and pre-fill speeds of 1500-2200 token/s. The model reliably interfaces with various APIs and generates functional code for single-page apps, LaTeX docs, parsers, and crawlers.
LLMs Benchmarked for Web Vulnerability Detection
A study evaluates six LLMs on detecting real-world web vulnerabilities in WordPress plugins, finding detection rates vary by model and prompt design. Claude Opus 4.6 achieved the highest detection rate at 63%, while Qwen 3.5 only reached 35%, and no model consistently identified all baseline vulnerabilities across iterations.
Benchmark Evaluation of Small Language Models for Arabic NLP
A benchmark of 240 Arabic test items across eight domains and ten skills assesses twelve small language models in zero-shot settings. Gemma 3 (12B) achieved the highest overall score (4.548/5), followed by Aya and C4AI Command Arabic, with performance linked more to Arabic alignment and instruction-following than model size. Common failure modes include prompt leakage, hallucination, and weak task adherence.
Benchmark Evaluation of Small Language Models for Arabic NLP
A benchmark of 240 Arabic test items across eight domains and ten skills assesses twelve small language models in zero-shot settings. Gemma 3 (12B) achieved the highest overall score (4.548/5), followed by Aya and C4AI Command Arabic, with performance linked more to Arabic alignment and instruction-following than model size. Common failure modes include prompt leakage, hallucination, and weak task adherence.