A Lifecycle-Aware KV Cache System for Agent Inference
Agent inference extends a single model call into a long-lived workflow of planning, tool execution, observation integration, validation, and retries, turning the KV cache from a temporary per-request tensor into execution state spanning multiple turns.