The llama.cpp project has added an A8 Q8_0 non-MoE dp4a binary kernel for OpenCL.

  • The change is tracked as pull request #29439 in the ggml-org/llama.cpp repository.
  • It introduces support for a specific quantization format (A8 Q8_0) using the dp4a instruction set within the OpenCL backend.