Chat Completions API now allows you to pass an opaque session ID and the router attempts to keep the same eligible model across conversation turns, expiring after 30 minutes of inactivity.
Fixes the “why did its tone change mid-conversation?” complaint on routed chat workloads and improves prompt-cache reuse
Read more, including code examples: https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/model-router?tabs=foundry-responses#keep-chat-completions-requests-on-the-same-model-preview