Text agents can afford to think. A voice agent cannot, three seconds of dead air on a video call is read as a dropped connection, not as deliberation.We handle this with two mechanisms. First, tools are declared with an expected duration, and anything slow enough triggers a spoken acknowledgement, a short line generated from the tool's name rather than a canned filler, while the call resolves in the background.Second, a tool result is injected mid-turn rather than starting a new one, so the model receives it as a continuation of the utterance it is already speaking and the answer arrives as one sentence instead of two.