The llama.cpp server now supports typed content (vision, audio, video) input for the /v1/embeddings endpoint. This update enables multimodal embedding requests by accepting OpenAI-style wrapped content arrays, allowing text and image_url parts to be processed together.

  • Accepts OpenAI-style wrapped content array format for multimodal embeddings.
  • Text parts are concatenated while image_url parts are decoded via handle_media and spliced with process_mtmd_prompt.
  • Legacy formats (plain string, token arrays, mixed arrays) continue to work unchanged.
  • Bare content arrays are rejected with a migration message.
  • Disables KV prefix reuse for stateless embedding/rerank tasks to prevent incorrect cache sharing across requests.

This change allows users to generate embeddings for multimodal inputs using the standard OpenAI API format while ensuring that repeated inputs do not incorrectly share cached key-value pairs.