Tencent introduces KuaFu, a unified behavior-compression layer that condenses raw user history into compact representations for conversational agents and recommenders. The system uses a two-axis projector to compress each behavior item into 2-4 tokens of width 128-256, addressing bottlenecks in processing long sequences for a billion users.

KuaFu matches or exceeds uncompressed single-task production models across four profiling tasks while raising per-GPU throughput by 37%-350% and saving 190 GPUs. On public benchmarks, it outperforms prior compressors at the same ratio, with a 4B model surpassing an 8B counterpart on RecBench.

After running on Tencent's advertising and recommendation platform for ten months, KuaFu lifted overall GMV by 1.37%.