Summary
Nvidia researchers have developed a cross-model KV cache transfer technique using simple linear math, which can replace costly AI model handoffs and reduce compute costs and latency in multi-LLM workflows by up to 25 times faster than recomputing.