The llama.cpp project has added fast attention vector (fa-vec) tuning records for Apple M3 Max, M5, and M5 Pro GPUs to improve Metal performance.

  • Tunings for the Apple M5 were generated on an M5 machine using f16 and q8_0 data types.
  • Records for the Apple M5 Pro cover F16, Q4_0, and Q8_0 formats across 20 GPU cores.
  • Support for the M3 Max includes F16 and Q8_0 configurations on a MacBook Pro with 64GB in low power mode.

These updates extend Metal acceleration capabilities to newer Apple silicon architectures.