The llama.cpp project has released build b10269, which includes a fix for the dflash wo_a reshape issue during model loading.
This release provides binaries for macOS (Apple Silicon and Intel), iOS, Linux (Ubuntu x64, arm64, s390x with CPU, Vulkan, ROCm 7.2, OpenVINO, and SYCL backends), Android (arm64), Windows (CPU, OpenCL Adreno, CUDA 12/13, Vulkan, OpenVINO, SYCL, HIP), and openEuler (x86 and aarch64 with ACL Graph).
The update is available for download across all supported platforms to ensure correct model loading behavior.