← Blog
Field note

LLM Model Migration: GPT-6 Sol & Claude Opus 5.5 Without Regressions

Thien Nguyen · Oct 6, 2026

If your AI app runs on GPT-5.1, GPT-5.3-Codex or GPT-5.4-Nano, OpenAI shuts those models down in the API on 1 April 2027. Older GPT-5 and o3 snapshots go sooner, on 11 December 2026. Claude Sonnet 4.5 retires on 30 November 2026. To migrate without regressions, don't just swap the model ID. Freeze a golden test set from real traffic, fix the breaking API changes, re-tune prompts and reasoning effort against that set, shadow the new model on live traffic, then roll out behind a fallback. For most production apps that takes two to six weeks, and the evals usually take longer than the code changes.

This playbook covers the process. Deadlines, parameters and prices come from the OpenAI, Anthropic and Google deprecation, model and pricing pages as checked on 7 October 2026. For token prices across every architecture, see our AI app development cost guide. For an eval plan that will satisfy an auditor, see LLM testing for regulated industries.

Which models are being retired, and when?

These are the API deadlines most likely to hit a production app in the next six months:

Shutdown dateVendorWhat goesRecommended replacementAnnounced
31 Oct 2026 (read-only), 30 Nov 2026 (shutdown)OpenAIEvals platform (dashboard and API)OpenAI recommends Promptfoo3 Jun 2026
30 Nov 2026OpenAIReusable prompts (v1/prompts) and Agent BuilderPrompts in your own code; Agents SDK or ChatGPT workspace agents3 Jun 2026
30 Nov 2026Anthropicclaude-sonnet-4-5-20250929claude-sonnet-5-530 Sep 2026
11 Dec 2026OpenAIOriginal GPT-5 snapshots (gpt-5, -mini, -nano, -pro) and o3 / o3-proGPT-5.6 Sol, Terra or Luna11 Jun 2026
6 Jan 2027OpenAItts-1, tts-1-hd, gpt-4o-mini-tts snapshotsgpt-realtime-2.1-mini1 Oct 2026
1 Apr 2027OpenAIgpt-5.1, gpt-5.3-codex, gpt-5.4-nanogpt-6-sol (Nano → gpt-6-luna)1 Oct 2026
7 May 2027Googlegemini-3.1-flash-litegemini-3.5-flash-liteSee Gemini deprecations page

Keep three more dates on your radar. Claude Haiku 4.5 is still active, but Anthropic only commits to keeping it until "not sooner than" 15 October 2026, so a deprecation notice could arrive at any time. Anthropic promises at least 60 days' notice before retiring a publicly released model. Gemini 3.8 Flash doesn't retire, but its price doubles on 1 January 2027. And if your regression suite lives in OpenAI's Evals platform, export it this month. It becomes read-only on 31 October.

Timeline of LLM API retirements from October 2026 to May 2027: OpenAI Evals read-only 31 Oct, Claude Sonnet 4.5 and Agent Builder 30 Nov, GPT-5 and o3 snapshots 11 Dec, TTS 6 Jan, GPT-5.1, GPT-5.3-Codex and GPT-5.4-Nano 1 Apr 2027, gemini-3.1-flash-lite 7 May 2027, with replacements
Figure 1: Timeline of LLM API retirements from October 2026 to May 2027: OpenAI Evals read-only 31 Oct, Claude Sonnet 4.5 and Agent Builder 30 Nov, GPT-5 and o3 snapshots 11 Dec, TTS 6 Jan, GPT-5.1, GPT-5.3-Codex and GPT-5.4-Nano 1 Apr 2027, gemini-3.1-flash-lite 7 May 2027, with replacements

What breaks when you move to GPT-6 Sol or Claude Opus 5.5?

Both vendors made request settings that used to be harmless into hard errors. Changing the model string alone will fail in production. These are the changes their migration guides list:

AreaGPT-6 Sol / GPT-6.1 Sol / GPT-6 LunaClaude Opus 5.5
Sampling parametersRemove temperature, top_p and top_logprobs when reasoning effort is not none (and logprobs on Chat Completions)Any non-default temperature, top_p or top_k returns a 400 error
Reasoning / thinkingEffort levels none to max, default medium. GPT-6.1 Sol has no none; use low. Replace minimal with low and compareAdaptive thinking is always on. thinking: disabled and manual budgets are rejected. Effort is the only control, and its default is medium (Opus 5 defaulted to high)
Tool callingUse the Responses API for tools on GPT-6.1 Sol. On Chat Completions, GPT-6 Sol and Luna call functions only at effort noneForced tool choice (any or a named tool) is rejected. Use auto with strict tools or structured outputs
Response shapeReasoning settings change output length and latencyResponses can start with thinking blocks, so code that reads content[0].text breaks. Thinking blocks must be passed back unmodified in tool loops
Prefill and cachingReplace prompt_cache_retention with prompt_cache_options.ttl set to "30m"Assistant prefill is rejected; use structured outputs or system instructions
Safety and refusalsCheck refusals on your own edge casesNew stop_reason: "refusal" categories, with optional server-side fallback to another model
Regional processingFast mode isn't available with EU data residencyUS-only inference costs 1.1x

Two of these cause silent regressions rather than errors. On Opus 5.5, text that the model writes between tool calls now arrives in thinking blocks, which are empty by default. An app that streamed those progress updates to users goes quiet, and nothing throws an exception. On GPT-6 Sol and Opus 5.5, an unset effort parameter means medium, which can change answer depth and latency without any code change.

Monitoring dashboards for an AI application
Monitoring dashboards for an AI application

How do the new models compare on cost and latency?

List prices per million tokens (input / output) from the vendors' pricing pages on 7 October 2026. The per-1,000 columns reuse the assumptions from our cost guide: 6,000 input and 500 output tokens for a RAG answer, and 60,000 input and 3,000 output for an agent task. There is no caching, and reasoning tokens are not included.

MigrationPrice before → afterPer 1,000 RAG answersPer 1,000 agent tasksWatch out for
GPT-5.1 → GPT-6 Sol$1.25 / $10 → $2 / $10$12.50 → $17.00$105 → $150Input costs 60% more. Cached input is $0.20 (Sol) or $0.10 (6.1 Sol)
GPT-5.3-Codex → GPT-6 Sol$1.75 / $14 → $2 / $10$17.50 → $17.00$147 → $150Roughly flat. Output-heavy jobs get cheaper
GPT-5.4-Nano → GPT-6 Luna$0.20 / $1.25 → $0.10 / $0.50$1.83 → $0.85$15.75 → $7.50About half the price. Check quality on your hardest cases
Claude Sonnet 4.5 → Sonnet 5.5$3 / $15 → $2 / $10$25.50 → ~$22.10$225 → ~$195The newer tokenizer produces ~30% more tokens for the same text, so savings are ~13%, not 33%
Claude Opus 5 → Opus 5.5$5 / $25 → $4 / $20$42.50 → $34.00$375 → $300Anthropic reports 40% lower cost on typical workloads at default settings, partly from the lower default effort

Three cost traps sit outside the list price. First, Anthropic bills thinking tokens as output tokens, and on Opus 5.5 every request thinks, because thinking can't be turned off. Check reasoning-token usage on the OpenAI side too. A workload that previously ran without thinking can produce more output tokens even at a lower price. Second, GPT-6 Sol and GPT-6.1 Sol charge $4 input and $15 output above 272K tokens of context. Claude 4.6 and later models charge the standard rate across the full 1M window. Third, batch processing costs half on both vendors, so move evals and back-office jobs there.

No vendor publishes latency guarantees for these models. Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5, and it offers a fast mode at $8 / $40. OpenAI's GPT-6.1 Sol has no none effort level, which puts a floor under its response time. Measure p50 and p95 latency on your own prompts at the effort you plan to ship. Don't trust a launch chart.

Why do model upgrades cause regressions?

Newer models are better on average and worse on some of your cases. The usual causes are:

  • Prompts tuned to the old model. Workarounds for old weaknesses, such as "always answer in JSON" or "think step by step", can confuse a model that already does those things.
  • Changed defaults. Effort, verbosity and refusal behaviour shift between versions. A request that sets nothing now behaves differently.
  • Output format drift. Small changes in markdown, field order or citation style break parsers and UI.
  • Tool-use behaviour. Models differ in how readily they call tools, how many calls they make and how they handle tool errors. Agents feel this most.
  • Token counts. A new tokenizer changes context-window headroom, truncation points and cost.
  • Latency budgets. More thinking can push p95 latency past a timeout that never fired before.

Without an eval set, you find these through customer complaints. Avoiding that is the whole point of the next section.

What does an eval-driven migration plan look like?

Six steps, in this order. Two to six weeks is realistic for one production feature, depending mainly on whether you already have evals.

  1. Inventory. List every model ID, endpoint, SDK version and parameter in use. Include background jobs and scripts. Anthropic's Console can export usage broken down by API key and model, and your OpenAI usage dashboard shows which models you call.
  2. Build the golden set. Sample 200–500 real requests across your main intents, plus every past incident and known edge case. Record the current model's outputs as the baseline and label what "correct" means for each case.
  3. Fix the breaking changes. Work through the table above. Read content blocks by type, remove sampling parameters, choose an effort level explicitly and move tool calls to the supported API.
  4. Tune against the set. Run the golden set on the candidate model at two or three effort levels. Remove old prompt workarounds one at a time and keep only changes that improve the score.
  5. Shadow live traffic. Send a copy of production requests to the new model without showing users the result. Compare quality scores, refusals, tool calls, cost and latency for at least a week.
  6. Canary, then cut over. Move 5%, then 25%, then 100% of traffic. Keep the old model as a fallback until its shutdown date, and keep the eval suite in CI for the next migration.
Six-step eval-driven LLM migration plan: inventory, golden set of 200 to 500 real cases, fix breaking changes, tune, shadow traffic for a week, canary 5%, 25%, 100%, plus example regression pass rules
Figure 2: Six-step eval-driven LLM migration plan: inventory, golden set of 200 to 500 real cases, fix breaking changes, tune, shadow traffic for a week, canary 5%, 25%, 100%, plus example regression pass rules

What should the regression test cover?

Use one row per risk and set the threshold before you see the results. The thresholds below are examples. Set your own from your baseline.

TestWhat it measuresExample pass rule
Task accuracyCorrect answer or action on the golden set, scored by rules, human labels or a calibrated LLM graderNo worse than baseline minus 1 point overall, and no drop on any critical intent
Format validityJSON or schema parses, required fields present100% on structured outputs
Tool-call correctnessRight tool, valid arguments, no forbidden calls, call count per taskZero forbidden calls; median call count within 20% of baseline
RefusalsOver-refusal on legitimate requests, under-refusal on red-team promptsNo new refusals on the golden set; all red-team cases blocked
GroundingAnswers supported by retrieved sources, with citations that resolveNo unsupported claims on the regulated-content subset
Latencyp50 and p95 end to end, including toolsp95 under your product timeout with margin
CostTokens and dollars per task, including reasoning tokensWithin the re-forecast budget

If your industry is regulated, keep the run results, model versions and thresholds as evidence. A model change is exactly the kind of event auditors ask about. Our regulated-industry eval plan maps these tests to HIPAA, PCI DSS and the EU AI Act, and AI agent evaluation and observability covers the tooling.

How do you roll out the new model safely in production?

  • Put the model behind a config switch, not a code change, so rollback takes seconds.
  • Route by task, not by vendor. Easy tasks can move to cheaper models such as GPT-6 Luna or Claude Haiku 4.5. Keep the expensive model for steps where your evals show it matters.
  • Configure fallback deliberately. If a refusal or outage moves a conversation to another model, test that path. Anthropic's guide warns that the fallback model will usually run without Opus 5.5's thinking blocks.
  • Watch online metrics for two weeks after cutover: user corrections, thumbs-down, escalations, tool errors, cost per task and p95 latency.
  • A/B test prompt changes with statistics, not intuition. A difference of two points on 50 samples is usually noise.

What is the model-migration checklist?

  • Every model ID, SDK version and parameter in use is listed, including batch jobs.
  • Each model has a shutdown date and an owner in your tracker.
  • Evals are exported from any platform that is shutting down, such as OpenAI Evals by 31 October 2026.
  • A golden set of 200–500 real cases, with edge cases and past incidents, has a recorded baseline.
  • Sampling parameters, prefills, forced tool choice and disabled thinking are removed where they are rejected.
  • Response parsing reads blocks by type, and thinking blocks are passed back unmodified.
  • Effort or reasoning level is set explicitly and chosen from an eval sweep.
  • max_tokens is re-budgeted for thinking plus answer.
  • Pass thresholds are written down before the candidate runs.
  • Shadow traffic has run for at least a week with quality, cost and latency compared.
  • Canary steps and a one-switch rollback are tested.
  • Token cost is re-forecast, including tokenizer changes, long-context tiers and 2027 price changes.
  • The eval suite stays in CI for the next migration.

How can BeevR help with a model migration?

A migration is a well-bounded piece of work: inventory, evals, code changes and rollout. That makes it suitable for a fixed price. Our AI agent development work is priced by phase from $10K, and our published packages run from the $4K Pitch Demo to the $38K Flagship Sprint, which includes a full test suite, automated deploy and rollback, and full observability. Our open-source agent framework, Kite, ships prompt A/B testing with statistical confidence intervals on real traffic. It exists because "the new prompt feels better" is not a release criterion. If you're unsure whether your agent should be rebuilt rather than migrated, our framework guide and agent cost guide are good places to start.

FAQ

When does GPT-5.1 stop working in the API?

On 1 April 2027. OpenAI announced on 1 October 2026 that gpt-5.1, gpt-5.3-codex and gpt-5.4-nano will be shut down that day, with gpt-6-sol as the replacement for the first two and gpt-6-luna for Nano. Requests to a shut-down model fail.

Is GPT-5.5 being deprecated in the API?

Not as of 7 October 2026. OpenAI's deprecations page lists no shutdown date for gpt-5.5, which is priced at $5 input and $30 output per million tokens. Changes to which models appear in ChatGPT are separate from API deprecations.

Should we move to GPT-6.1 Sol or Claude Opus 5.5?

Run both on your golden set and decide on the evidence. GPT-6.1 Sol costs $2 / $10 with a 1.05M context. Opus 5.5 costs $4 / $20 with a 1M context and always-on thinking. Price per token matters less than cost per successful task, and that depends on your prompts, tools and effort level.

Can we skip evals if the new model is "better"?

Benchmarks average over tasks that aren't yours. Anthropic itself notes that benchmark margins have become a less reliable guide to real-world differences. A 200-case golden set takes days to build and is the only way to know whether your app got better or worse.

How long does an LLM migration take?

Changing the model ID takes minutes. A safe migration of one production feature typically takes two to six weeks. Most of that is building the eval set, tuning prompts and effort, and running shadow traffic long enough to trust the numbers.

BeevR is a senior, founder-led AI studio in Hanoi, Vietnam. We work at a fixed price per phase, you own the code and IP from day one, and we build AI for regulated industries. If you have a model deadline coming up, tell us which models you run. We'll map the breaking changes and the eval work before you commit to anything.