The llama.cpp project released build b11302, which addresses a critical concurrency issue identified by ThreadSanitizer in the CI pipeline. The update resolves a data race within the sparse attention mechanism where multiple CPU threads were incorrectly writing to the same memory elements.

  • Dead indexer slots are now assigned unique scatter rows (n_kv + slot) instead of sharing a single sentinel row.
  • Live slots address disjoint cells, ensuring that scatter indices for a token remain unique across all selection paths.
  • The fix eliminates overlaps between invisible pools selected by top_k and the tail cells of tokens.

This change ensures thread-safe operation of the sparse indexer mask construction, preventing undefined behavior caused by concurrent writes to shared memory locations.