Quality changing over time without any code change, because the data, the users' questions or the model behind an API changed.
You have done this if
Answer quality slipped after the provider updated the model version behind the same name.
Say it in a review
We pin model versions and track eval scores weekly to catch drift in the data or the model.
On the AI Application map Observability, Model