The streaming detection piece is what I'd want to understand better. How are you catching that the model is committing to a tool call before generation finishes? Pattern-matching on partial token sequences, or something at the structured output layer that gives you earlier signal?
On the caching claim, I want to make sure I'm reading this right.. Mutating at the head of a long context should bust prefix cache for everything after the insertion point. The win would come if the stable prefix before the insertion is long enough that most of the cache value lives there anyway, and tool invocations tend to cluster toward the end of a conversation. Is that the workload shape this is designed around?
On the caching claim, I want to make sure I'm reading this right.. Mutating at the head of a long context should bust prefix cache for everything after the insertion point. The win would come if the stable prefix before the insertion is long enough that most of the cache value lives there anyway, and tool invocations tend to cluster toward the end of a conversation. Is that the workload shape this is designed around?