The hexagon backend updates flash attention to use head-parallel partitioning for row-split multicore execution, addressing a previous limitation where cores read the full KV cache. By partitioning by KV heads when divisible by core count, each core reads only its specific shard of the KV cache, restoring memory bandwidth benefits.
- Qwen3-0.6B inference speed increases from 6977 to 11026 tokens/s (+58%) at 4c row-split.
- llama-3.2-3B performance rises from 3717 to 5522 tokens/s (+49%).
- Qwen3.5-4B sees a minor 4% gain, while Gemma-4 MoE shows no change due to MoE FFN dominance.
- The feature is controlled by GGML_HEXAGON_FA_HEAD_SPLIT and falls back to token-block splitting when head counts are not divisible by core count.
This optimization significantly accelerates prefill throughput for specific model configurations on Hexagon hardware.