The llama.cpp server now permits RANK pooling batch splitting for causal LLM rerankers like Qwen3 and Qwen3-VL, resolving previous limitations that broke long-document and multimodal reranking.

  • Previously, the server rejected RANK-pooling inputs larger than n_ubatch, forcing all tokens into a single physical batch.
  • The fix exposes llama_get_causal_attn and llama_model_is_causal to dynamically check attention types instead of hardcoding architecture checks.
  • can_split() now allows chunked prefill for RANK pooling when the context is causal, removing duplicated logic in the graph builder.

This change enables efficient processing of long documents and multimodal inputs by leveraging chunked prefill capabilities inherent to causal models.