8. Switching Models Mid-Conversation
Switch to a different model during a conversation to compare answers or escalate a difficult turn, without losing history.
OUPI lets you change the model at any point during a conversation. When you switch, the full conversation history is preserved — the new model reads everything that came before and continues seamlessly. This is powerful for two workflows: comparing answers (ask the same question, then switch models to see a different perspective) and escalating (start with a balanced model for routine turns, then bring in a frontier model only for the turn that demands deeper reasoning). Because only the hard turn uses the expensive model, you keep costs under control while still getting top-quality output when it matters.
To decide which model to switch to, think in families rather than names — names change often, but the families stay stable. Balanced models (like Claude Sonnet 5, GPT-5.6 Terra, or Gemini 3.8 Flash) handle most professional work well and are your sensible default. Frontier models (like Claude Opus 5, GPT-6 Astra, or Gemini 3 Pro) are for hard reasoning, long multi-document synthesis, or complex code. Fast models (like Claude Haiku 4.5 or GPT-5.6 Luna) suit short, repetitive, or low-stakes tasks. Specialists — reasoning models, code models, or live-web models — shine when the task matches their strength exactly.
A common cost-saving pattern: run most of the conversation on a balanced model, then switch to a frontier model only for the one difficult turn (e.g., a tricky legal analysis or a multi-step calculation). Frontier models can cost 10–30× more per token than fast ones, so reserving them for the hard part makes a real difference to your credit balance.
Keep context length in mind when you switch. Your conversation history counts toward the new model's context window. Recent frontier and balanced models can read up to about a million tokens, but lighter or older models have much smaller windows. If your conversation is already long and you switch to a fast model, it may not be able to read the full history. For very long threads, consider using a knowledge base so the model reads only what is relevant, rather than the entire conversation.
If you don't have a strong reason to pick a specific model, leave automatic model selection on — OUPI will route each request to a suitable model based on the task and length. Switch manually only when you want a specific provider, need to handle a very long document, or know the task is especially hard.
Open a conversation in OUPI's chat. Send a moderately complex question using a balanced model. Then switch to a frontier model and ask the same question again. Compare the two answers side by side — notice how the history is preserved and the frontier model can reference earlier turns.
You can switch models at any point in a conversation without losing history. Use this to compare answers across models or to escalate a single difficult turn to a more powerful model. Think in families (frontier, balanced, fast, specialist) to pick the right target. Watch context length when switching to lighter models on long threads. Use automatic model selection when you have no specific reason to choose, and switch manually when you do.