loo5x has released llama-modes v0.4.0, a fork of llama.cpp that enables loading a single GGUF model once to perform multiple structured inference modes without swapping models or adding classifiers.
- BOOLEAN mode directly scores Yes/No answers.
- CHOICE mode scores arbitrary supplied candidates, including multi-token labels.
- SCALE mode evaluates ordinal or interval scales, returning discrete distributions and derived statistics like median or expected value.
- The project includes a Windows CUDA release, local React demo, and API documentation.
This approach allows developers to reuse the same model weights in memory for different inference primitives, exposing structured scoring directly from ordinary local language models.