Researchers introduce Queen, a 4-billion-parameter chess-language model that plays at the level of a typical Grandmaster while explaining its moves and plans. The framework combines an encoder-decoder architecture with an iterative distillation algorithm to enable domain-specific reasoning.
- Integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, trained via a question-answering curriculum.
- Uses a natural-language analog of the Bellman update to analyze positions after candidate moves and consolidate explanations.
- Gains over 900 Elo points (1782 to 2697) over seven iterations, surpassing frontier models in playing strength and puzzle accuracy.
- Explanations are fluent and approach GPT-5.6-Sol (high) in coherence despite having three orders of magnitude fewer parameters.
The authors suggest this architecture provides a recipe for applying language models to domains with silent expert encoders, such as robotics and computer use.