A researcher is seeking an arXiv cs.CL endorsement for a paper introducing Simple Retrieval-Reconstruction Memory (SRRM), a lightweight benchmark designed to evaluate long-context memory in Small Language Models.
Unlike existing retrieval-oriented benchmarks, SRRM jointly evaluates targeted retrieval, global information reconstruction, and a new metric called Memory Preservation Ratio (MPR). The authors tested the Qwen2.5 model family (0.5B, 1.5B, and 3B) under context loads of 10–100 factual statements.
The results showed that the Qwen2.5-1.5B model consistently outperformed the larger 3B model across retrieval accuracy, reconstruction performance, and MPR, suggesting that parameter count alone does not necessarily improve long-context memory preservation.