Inference efficiency
lab Microsoft Research Blog · 6h ago · 14 views

Microsoft Research shows offloading physical AI inference improves robot performance and battery life

Microsoft Research challenges the assumption that physical AI inference must run exclusively on onboard robot GPUs, demonstrating that offloading to edge or cloud infrastructure significantly enhances mobile manipulation workloads. Their systematic study of robotics workloads reveals that relying on onboard compute limits performance, battery life, and scalability.

arxiv arXiv cs.LG · 22h ago · 12 views

CliffCompaction reduces long-horizon coding agent costs by up to 50% while maintaining performance

Researchers introduce CliffCompaction, an autocompaction technique designed for agents handling millions of tokens, which cuts costs by up to 50% under bounded context. The method maintains or improves performance on Terminal-Bench and achieves state-of-the-art results on KernelBench by keeping information faithful through truncation rather than rephrasing.