An independent researcher has published a literature review analyzing whether model compression techniques like quantization, pruning, and distillation actually reduce measured energy consumption or merely lower FLOPs.

  • The review draws on a corpus of 53 works from 2019 to 2026.
  • Software energy estimators and hardware counters (RAPL, NVML) disagree by margins comparable to the savings claimed by compression methods.
  • FLOPs correlate weakly with measured energy once memory bandwidth, batch size, and backend kernel realization are accounted for.
  • Compression rankings on one hardware platform do not reliably transfer to another.

The author shares the review and associated data tables on Zenodo and GitHub, while seeking a cs.LG endorsement to post the work on arXiv.