An explicit limit on how long each step may take, set so the steps add up to less than what the user will wait for.
You have done this if
You gave retrieval 800 ms and the model 20 seconds so the whole request stayed under the 30-second gateway limit.
Say it in a review
Each hop has a timeout and they sum to less than the user-facing deadline.
On the AI Application map API Gateway, Retrieval, Model