A user has released Juud_engine, an experimental MIT-licensed fork of Strata v0.1.39 designed for local inference of the Qwen3.8-Flash-Next model on consumer hardware like the RTX 4090. The project introduces two opt-in CPU-side decode changes: shortened worker spin after CPU expert batches and skipped CPU activation quantization for tokens routed entirely to the GPU.
- Benchmarks on an i7-13700 and RTX 4090 showed median paired decode throughput improvements of 3.1–9.1% and latency gains of 0.5–5.6% across various context lengths.
- A separate control run with different cache settings reported higher throughput gains of 9.6–15.6%, though output consistency varied between test configurations.
- The repository includes source diffs, build instructions, model hashes, and a verification script for independent replication.
The author invites feedback on the benchmark methodology and requests independent replications to validate the performance claims.