The llama.cpp project has released build b11298, which introduces support for the dflash format. This update enables both conversion of models to this format and subsequent feature extraction capabilities.

  • The model conversion tool has been updated to handle dflash inputs.
  • Feature extraction functionality is now available for dflash models.
  • A fix was applied to the 'cont' component during development.
  • Binaries are provided for macOS, Linux, Windows, Android, and openEuler across CPU, GPU (CUDA, Vulkan, ROCm, OpenCL), and NPU backends.

This release allows users to utilize dflash models within the llama.cpp ecosystem.