llama.cpp b10728 रिलीज़ Flash Attention K,V shared memory fp16 tiles के लिए CUDA XOR swizzle का कार्यान्वयन पेश करती है। यह अपडेट DGX Spark हार्डवेयर पर Flash Attention में shared memory race condition को भी हल करता है।
- CUDA के लिए fp16 tiles के साथ XOR swizzle flash attention लागू करता है।
- 32-bit shared pointer की जगह 64-bit generic pointer के उपयोग को ठीक करता है।
- swizzle के लिए टेस्ट केस जोड़ता है और सिर्फ swizzled पथ के लिए synchronization को गेट करता है।
- nbatch_2%32==0 वाले non-power-of-2 आकारों के लिए swizzling की अनुमति देता है।
- Flash Attention swizzle ldmatrix तर्क को helper functions में पुनर्गठित करता है।
रिलीज़ CPU, CUDA, Vulkan, ROCm, OpenVINO और SYCL बैकएंड्स के लिए macOS, Linux, Windows, Android और openEuler के लिए बाइनरी प्रदान करती है।