The author has released AiiStream Q3.6, an open-source technique that allows the Qwen3.6-35B-A3B model to run on Apple Silicon devices with only 8 to 16 GB of RAM. By keeping routed experts on SSD and using early loading predictions, the method reduces MLX memory usage to approximately 1.7 GB while maintaining byte-identical outputs to the standard mlx_lm implementation.

Key performance metrics on a 16 GB M5 MacBook Air include:

  • 15.0 tokens/s in controlled benchmarks, which is 72.6% faster than direct SSD streaming.
  • Approximately 11.8 tokens/s with other applications running.
  • A time-to-first-token of around 2.3 seconds for short prompts.

The technique is designed to work on 8 GB Macs as well, though performance figures for that configuration remain estimates based on the current release.