The author released Dev-4B, a small local model based on Qwen3-4B-Instruct that answers typed questions about documents with calibrated confidence. It uses a built-in router to decide whether to provide a quick answer or escalate to step-by-step reasoning only when the initial prediction is likely wrong.

  • Evaluated on 7,100 frozen test questions, Dev-4B achieved 81.8% accuracy in 1.37 seconds per question.
  • This outperforms always-quick answers (76.3%, 0.16 s/q) and always-reasoning approaches (78.0%, 10.2 s/q).
  • The model uses a LoRA adapter switched on for questions and a tiny router, with MLX 8-bit builds available.

The tool is designed for local calibration and offline routing experiments, particularly for Mac indie developers.