Janas-LLM is a new C-based, GPL-3.0 inference engine designed to run Mixture-of-Experts models larger than RAM on ordinary Linux machines without GPUs by streaming experts from the SSD.
The developer is requesting community testing to tune performance across diverse hardware configurations, specifically targeting AMD CPUs, Intel processors without AVX-VNNI, systems with 16 GB or less of memory, and SATA SSDs. The provided script automates building the engine, downloading Qwen models, verifying bit-identical arithmetic results, measuring speed on an idle machine, and reporting findings via GitHub issues.
This effort aims to improve the engine's self-tuning capabilities for threads, GPU usage, and expert loading strategies across a wider variety of hardware than the developer's single test machine.