A discussion on Hugging Face proposes replacing traditional text retrieval in RAG systems with LLM-native latent representations. The idea involves encoding documents into internal model states during indexing and retrieving those states for inference, potentially bypassing the cost of re-processing retrieved text.

  • Documents are converted to LLM-native latent representations for storage instead of embeddings.
  • Queries retrieve these latent states to inject into Transformer layers, avoiding tokenization and prefill costs.
  • Key challenges include compressing representations without losing critical details like names or conditions.
  • Technical hurdles involve injecting memory into appropriate layers and storing data on NVMe/SSD for efficient loading.
  • The author seeks existing research on using internal LLM states as persistent external RAG memory.

The proposal aims to reduce RAG prefill latency by leveraging the model's own semantic processing capabilities rather than discarding embeddings after retrieval.