Nanbeige4.2-3B is a compact general agentic model with 3B non-embedding parameters, pretrained from scratch on 28T tokens using a Looped Transformer architecture that reuses layer stacks to increase capacity without adding parameters.
The model utilizes mixed-mode RLHF over Think and Non-Think responses, length-controlled reasoning RL, and agentic RL with outcome and process rewards to stabilize training. It demonstrates strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining competitive reasoning capabilities in mathematics, coding, and science.
Extensive evaluations show that Nanbeige4.2-3B outperforms larger models, including Qwen3.5-9B and Gemma4-12B, across diverse agentic benchmarks. Its performance with OpenClaw supports its use as a compact local personal assistant.