The llama.cpp project released build b11282, which includes a fix to define the `CUDA_ARCH` macro for MUSA device passes. Previously, the MUSA vendor header failed to define this variable, causing architecture tests in shared ggml-cuda sources to evaluate to zero and resulting in empty kernel bodies.
- The fix reports the newest architecture similar to the HIP backend while explicitly excluding NVIDIA-only features incompatible with MUSA.
- `CUDA_ARCH` is now defined only for device passes to align with CUB's detection logic and nvcc behavior.
- Redundant `defined(CUDA_ARCH)` checks in architecture comparisons were removed since the macro is undefined in host passes for CUDA and MUSA.
This change ensures that kernel bodies gated on architecture, such as the q8_0 -> f16 dequantization kernel, compile correctly for MUSA devices instead of expanding to empty bodies.