If your AI app runs on GPT-5.1, GPT-5.3-Codex or GPT-5.4-Nano, OpenAI shuts those models down in the API on 1 April 2027. Older GPT-5 and o3 snapshots go sooner, on 11 December 2026. Claude Sonnet 4.5 retires on 30 November 2026. To migrate without regressions, don't just swap the model ID. Freeze a golden test set from real traffic, fix the breaking API changes, re-tune prompts and reasoning effort against that set, shadow the new model on live traffic, then roll out behind a fallback. For most production apps that takes two to six weeks, and the evals usually take longer than the code changes.
This playbook covers the process. Deadlines, parameters and prices come from the OpenAI, Anthropic and Google deprecation, model and pricing pages as checked on 7 October 2026. For token prices across every architecture, see our AI app development cost guide. For an eval plan that will satisfy an auditor, see LLM testing for regulated industries.
These are the API deadlines most likely to hit a production app in the next six months:
| Shutdown date | Vendor | What goes | Recommended replacement | Announced |
|---|---|---|---|---|
| 31 Oct 2026 (read-only), 30 Nov 2026 (shutdown) | OpenAI | Evals platform (dashboard and API) | OpenAI recommends Promptfoo | 3 Jun 2026 |
| 30 Nov 2026 | OpenAI | Reusable prompts (v1/prompts) and Agent Builder | Prompts in your own code; Agents SDK or ChatGPT workspace agents | 3 Jun 2026 |
| 30 Nov 2026 | Anthropic | claude-sonnet-4-5-20250929 | claude-sonnet-5-5 | 30 Sep 2026 |
| 11 Dec 2026 | OpenAI | Original GPT-5 snapshots (gpt-5, -mini, -nano, -pro) and o3 / o3-pro | GPT-5.6 Sol, Terra or Luna | 11 Jun 2026 |
| 6 Jan 2027 | OpenAI | tts-1, tts-1-hd, gpt-4o-mini-tts snapshots | gpt-realtime-2.1-mini | 1 Oct 2026 |
| 1 Apr 2027 | OpenAI | gpt-5.1, gpt-5.3-codex, gpt-5.4-nano | gpt-6-sol (Nano → gpt-6-luna) | 1 Oct 2026 |
| 7 May 2027 | gemini-3.1-flash-lite | gemini-3.5-flash-lite | See Gemini deprecations page |
Keep three more dates on your radar. Claude Haiku 4.5 is still active, but Anthropic only commits to keeping it until "not sooner than" 15 October 2026, so a deprecation notice could arrive at any time. Anthropic promises at least 60 days' notice before retiring a publicly released model. Gemini 3.8 Flash doesn't retire, but its price doubles on 1 January 2027. And if your regression suite lives in OpenAI's Evals platform, export it this month. It becomes read-only on 31 October.

Both vendors made request settings that used to be harmless into hard errors. Changing the model string alone will fail in production. These are the changes their migration guides list:
| Area | GPT-6 Sol / GPT-6.1 Sol / GPT-6 Luna | Claude Opus 5.5 |
|---|---|---|
| Sampling parameters | Remove temperature, top_p and top_logprobs when reasoning effort is not none (and logprobs on Chat Completions) | Any non-default temperature, top_p or top_k returns a 400 error |
| Reasoning / thinking | Effort levels none to max, default medium. GPT-6.1 Sol has no none; use low. Replace minimal with low and compare | Adaptive thinking is always on. thinking: disabled and manual budgets are rejected. Effort is the only control, and its default is medium (Opus 5 defaulted to high) |
| Tool calling | Use the Responses API for tools on GPT-6.1 Sol. On Chat Completions, GPT-6 Sol and Luna call functions only at effort none | Forced tool choice (any or a named tool) is rejected. Use auto with strict tools or structured outputs |
| Response shape | Reasoning settings change output length and latency | Responses can start with thinking blocks, so code that reads content[0].text breaks. Thinking blocks must be passed back unmodified in tool loops |
| Prefill and caching | Replace prompt_cache_retention with prompt_cache_options.ttl set to "30m" | Assistant prefill is rejected; use structured outputs or system instructions |
| Safety and refusals | Check refusals on your own edge cases | New stop_reason: "refusal" categories, with optional server-side fallback to another model |
| Regional processing | Fast mode isn't available with EU data residency | US-only inference costs 1.1x |
Two of these cause silent regressions rather than errors. On Opus 5.5, text that the model writes between tool calls now arrives in thinking blocks, which are empty by default. An app that streamed those progress updates to users goes quiet, and nothing throws an exception. On GPT-6 Sol and Opus 5.5, an unset effort parameter means medium, which can change answer depth and latency without any code change.

List prices per million tokens (input / output) from the vendors' pricing pages on 7 October 2026. The per-1,000 columns reuse the assumptions from our cost guide: 6,000 input and 500 output tokens for a RAG answer, and 60,000 input and 3,000 output for an agent task. There is no caching, and reasoning tokens are not included.
| Migration | Price before → after | Per 1,000 RAG answers | Per 1,000 agent tasks | Watch out for |
|---|---|---|---|---|
| GPT-5.1 → GPT-6 Sol | $1.25 / $10 → $2 / $10 | $12.50 → $17.00 | $105 → $150 | Input costs 60% more. Cached input is $0.20 (Sol) or $0.10 (6.1 Sol) |
| GPT-5.3-Codex → GPT-6 Sol | $1.75 / $14 → $2 / $10 | $17.50 → $17.00 | $147 → $150 | Roughly flat. Output-heavy jobs get cheaper |
| GPT-5.4-Nano → GPT-6 Luna | $0.20 / $1.25 → $0.10 / $0.50 | $1.83 → $0.85 | $15.75 → $7.50 | About half the price. Check quality on your hardest cases |
| Claude Sonnet 4.5 → Sonnet 5.5 | $3 / $15 → $2 / $10 | $25.50 → ~$22.10 | $225 → ~$195 | The newer tokenizer produces ~30% more tokens for the same text, so savings are ~13%, not 33% |
| Claude Opus 5 → Opus 5.5 | $5 / $25 → $4 / $20 | $42.50 → $34.00 | $375 → $300 | Anthropic reports 40% lower cost on typical workloads at default settings, partly from the lower default effort |
Three cost traps sit outside the list price. First, Anthropic bills thinking tokens as output tokens, and on Opus 5.5 every request thinks, because thinking can't be turned off. Check reasoning-token usage on the OpenAI side too. A workload that previously ran without thinking can produce more output tokens even at a lower price. Second, GPT-6 Sol and GPT-6.1 Sol charge $4 input and $15 output above 272K tokens of context. Claude 4.6 and later models charge the standard rate across the full 1M window. Third, batch processing costs half on both vendors, so move evals and back-office jobs there.
No vendor publishes latency guarantees for these models. Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5, and it offers a fast mode at $8 / $40. OpenAI's GPT-6.1 Sol has no none effort level, which puts a floor under its response time. Measure p50 and p95 latency on your own prompts at the effort you plan to ship. Don't trust a launch chart.
Newer models are better on average and worse on some of your cases. The usual causes are:
Without an eval set, you find these through customer complaints. Avoiding that is the whole point of the next section.
Six steps, in this order. Two to six weeks is realistic for one production feature, depending mainly on whether you already have evals.

Use one row per risk and set the threshold before you see the results. The thresholds below are examples. Set your own from your baseline.
| Test | What it measures | Example pass rule |
|---|---|---|
| Task accuracy | Correct answer or action on the golden set, scored by rules, human labels or a calibrated LLM grader | No worse than baseline minus 1 point overall, and no drop on any critical intent |
| Format validity | JSON or schema parses, required fields present | 100% on structured outputs |
| Tool-call correctness | Right tool, valid arguments, no forbidden calls, call count per task | Zero forbidden calls; median call count within 20% of baseline |
| Refusals | Over-refusal on legitimate requests, under-refusal on red-team prompts | No new refusals on the golden set; all red-team cases blocked |
| Grounding | Answers supported by retrieved sources, with citations that resolve | No unsupported claims on the regulated-content subset |
| Latency | p50 and p95 end to end, including tools | p95 under your product timeout with margin |
| Cost | Tokens and dollars per task, including reasoning tokens | Within the re-forecast budget |
If your industry is regulated, keep the run results, model versions and thresholds as evidence. A model change is exactly the kind of event auditors ask about. Our regulated-industry eval plan maps these tests to HIPAA, PCI DSS and the EU AI Act, and AI agent evaluation and observability covers the tooling.
max_tokens is re-budgeted for thinking plus answer.A migration is a well-bounded piece of work: inventory, evals, code changes and rollout. That makes it suitable for a fixed price. Our AI agent development work is priced by phase from $10K, and our published packages run from the $4K Pitch Demo to the $38K Flagship Sprint, which includes a full test suite, automated deploy and rollback, and full observability. Our open-source agent framework, Kite, ships prompt A/B testing with statistical confidence intervals on real traffic. It exists because "the new prompt feels better" is not a release criterion. If you're unsure whether your agent should be rebuilt rather than migrated, our framework guide and agent cost guide are good places to start.
On 1 April 2027. OpenAI announced on 1 October 2026 that gpt-5.1, gpt-5.3-codex and gpt-5.4-nano will be shut down that day, with gpt-6-sol as the replacement for the first two and gpt-6-luna for Nano. Requests to a shut-down model fail.
Not as of 7 October 2026. OpenAI's deprecations page lists no shutdown date for gpt-5.5, which is priced at $5 input and $30 output per million tokens. Changes to which models appear in ChatGPT are separate from API deprecations.
Run both on your golden set and decide on the evidence. GPT-6.1 Sol costs $2 / $10 with a 1.05M context. Opus 5.5 costs $4 / $20 with a 1M context and always-on thinking. Price per token matters less than cost per successful task, and that depends on your prompts, tools and effort level.
Benchmarks average over tasks that aren't yours. Anthropic itself notes that benchmark margins have become a less reliable guide to real-world differences. A 200-case golden set takes days to build and is the only way to know whether your app got better or worse.
Changing the model ID takes minutes. A safe migration of one production feature typically takes two to six weeks. Most of that is building the eval set, tuning prompts and effort, and running shadow traffic long enough to trust the numbers.
BeevR is a senior, founder-led AI studio in Hanoi, Vietnam. We work at a fixed price per phase, you own the code and IP from day one, and we build AI for regulated industries. If you have a model deadline coming up, tell us which models you run. We'll map the breaking changes and the eval work before you commit to anything.