Tmax presents the strongest open RL recipe for terminal agents, achieving 27% on Terminal-Bench 2.0 with only 9B parameters. It uses a novel data taxonomy to generate over 2.5x more terminal environments than prior datasets, enabling efficient training with a simple, outcome-only recipe. The dataset, models, and code are open-sourced at https://github.com/hamishivi/tmax.
Tmax: A Simple RL Recipe for Terminal Agents
Benchmarks
| Benchmark | Model | Score |
|---|---|---|
| Terminal-Bench 2 | Tmax | 27% |
H Company releases Holo4 open-weight computer-use models for desktop, web, and mobile
H Company has released Holo4, a family of generalist computer-use vision-language models designed to click, type, write code, and call tools across desktop, web, Android, and API environments. The release includes two model sizes: Holo4 27B (dense) and Holo4 35B-A3B (Mixture of Experts with 3B active parameters), both supporting a 256K context window.
IQuestLab releases IQuest-Q1, a 320B MoE model for agentic coding
IQuest has released IQuest-Q1, a Mixture-of-Experts (MoE) model designed for agentic coding, reasoning, and multi-step tool use. The model comprises approximately 320 billion total parameters, with an estimated 15 billion parameters activated per token. It is available on Hugging Face.
HealthClaw: open-source agent with self-evolving memory for longitudinal personal health management
Researchers developed HealthClaw, an open-source agent architecture designed to update support as a person's routines, preferences, measurements, and risks change over time. It separates shared safety rules and medical knowledge from private longitudinal memory containing profile facts, reusable procedures, and episodic traces.
DiaLLM investigates robustness-generation gap in English dialect adaptation
The DiaLLM study addresses the disconnect between understanding and producing dialectal English by continually pretraining three open-weight language model families on the International Corpus of English. It applies implicit and explicit post-training paradigms combined with three alignment strategies to compare Australian, Indian, and Northern British English.
DiaLLM investigates robustness-generation gap in English dialect adaptation
The DiaLLM study reveals that dialectal robustness and generation are dissociated in large language models, as benchmarks shaped by continual pretraining do not capture how alignment reshapes actual output. The authors continually pretrain three open-weight model families on the International Corpus of English and apply implicit and explicit post-training paradigms combined with three alignment strategies for Australian, Indian, and Northern British English.