The llama.cpp project has added support for the embeddingGemma2 model, which handles text, vision, and audio inputs. This update is implemented via pull request #30054.
llama.cpp adds support for embeddingGemma2 multimodal embeddings
llama.cpp v0.6.0 adds extended batch API, GLM-5.3-Flash support, and decision model endpoints
llama.cpp v0.6.0 introduces the new llama_batch_ext extended batch API for mixed token and embedding inputs, along with support for the GLM-5.3-Flash 320B hybrid model and the Clef decision model.
llama.cpp b11240 adds typed multimodal input support to /v1/embeddings
The llama.cpp server now supports typed content (vision, audio, video) input for the /v1/embeddings endpoint. This update enables multimodal embedding requests by accepting OpenAI-style wrapped content arrays, allowing text and image_url parts to be processed together.
Liquid AI releases LFM2.5-VL-3B-DSpark speculative drafter for up to 3.13x faster decoding
Liquid AI has released LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The drafter adds approximately 280M parameters and accelerates token generation without altering the model's output distribution.
Liquid AI releases LFM2.5-VL-DSpark to accelerate vision-language model inference
Liquid AI has released the LFM2.5-VL-DSpark draft model, a speculative decoding drafter designed to accelerate inference for the LFM2.5-VL-3B vision-language model. The drafter captures hidden states from fixed layers of the target model to generate candidate tokens, adding only 280M parameters (8.9% overhead) while maintaining identical dimensionality across modalities.
Liquid AI releases DSpark draft model for LFM2.5-VL-3B
Liquid AI has released an experimental DSpark draft model for its vision-language model LFM2.5-VL-3B, enabling faster decoding on edge devices and GPUs without changing output quality.