If you are starting an OpenAI integration in late 2026, three things in most tutorials are now wrong. OpenAI recommends the Responses API over Chat Completions for new projects. Structured Outputs on Responses uses text.format, not response_format. And the whisper-1 family is on a published removal date. This piece gives the current surface, the error codes the docs actually list, the caching rules that decide your bill, and the retention settings you need before sending regulated data.
Responses, not Chat Completions, for anything new
The migration guide states it plainly: "While Chat Completions remains supported, Responses is recommended for all new projects." The differences that affect your code:
- Chat Completions works on a Messages array; Responses works on Items.
- The
nparameter (parallelchoices) is gone. - Responses are stored by default. Pass
store: falseto turn that off, and note that this is a data-governance decision, not a performance one. - Conversation state can be carried with
previous_response_idinstead of resending the whole history yourself.
OpenAI publishes two numbers for the switch, both from its own internal testing: a 3% improvement on SWE-bench with the same prompt and setup, and a 40% to 80% improvement in cache utilisation compared with Chat Completions. Treat those as vendor-measured, not independently verified.
The Structured Outputs shape changed, and the constraint list is longer than people expect
On the Responses API the schema goes in text.format:
{
"model": "gpt-5.6-terra",
"input": "Extract the invoice fields.",
"text": {
"format": {
"type": "json_schema",
"name": "invoice",
"strict": true,
"schema": {
"type": "object",
"properties": {
"number": { "type": "string" },
"total": { "type": "number" },
"currency": { "type": "string", "enum": ["TRY", "USD", "EUR"] },
"due_date": { "type": ["string", "null"], "format": "date" }
},
"required": ["number", "total", "currency", "due_date"],
"additionalProperties": false
}
}
}
}
Three rules in that snippet are non-negotiable and cause most of the 400s. The root must be an object and cannot be an anyOf (so a top-level discriminated union will not validate). Every field must be listed in required; optionality is expressed as a nullable union, which is why due_date above is ["string", "null"]. And every object needs additionalProperties: false.
Unsupported JSON Schema keywords, per the guide: allOf, not, dependentRequired, dependentSchemas, if, then, else. Fine-tuned models drop more, including pattern, format, minimum, maximum and the array length constraints.
There are hard numeric limits too: at most 5,000 object properties, 10 levels of nesting, 1,000 enum values per schema, and a combined string length of 120,000 characters across property names, definition names and enum values (see OpenAI, Structured Outputs). Auto-generating a schema from a wide database table will hit these.
One useful behaviour that is easy to exploit: output key order follows schema order. If you want the model to reason before committing to a classification, put a shortrationalefield before thelabelfield in the schema. Put it after, and the model has already chosen.
Prompt caching is the single biggest lever on cost, and it is prefix-based
Caching is on by default and keys off the shared prefix of the rendered request. The documentation is explicit that "cache reuse requires the entire rendered prefix to match." The minimum cacheable prompt is 1,024 tokens for GPT-5.6 and later, and varies by request settings for earlier models. Reused tokens are billed at the cached-input rate, discounted up to 90%.
What silently breaks the cache: changing model, tools, parallel_tool_calls, text.format or reasoning.effort (see OpenAI, Prompt caching). The tools entry is the one that catches teams out, because it includes tool ordering. A tool registry that iterates a hash map will reorder itself between deploys and quietly cost you the discount. Sort your tool list deterministically.
Measure it rather than assume it. The usage object reports input_tokens_details.cached_tokens and input_tokens_details.cache_write_tokens. If cached_tokens is zero on a workload with a long fixed system prompt, your prefix is not stable.
Retention is configurable: prompt_cache_options.ttl on GPT-5.6 and later (currently the single value "30m"), and prompt_cache_retention on earlier models with in_memory or 24h. Cache-write pricing is not an additive fee; a token is billed as uncached, cached, or cache-write.
Rate limits: read the headers, not the dashboard
Limits are enforced at organisation and project level, not per user, and the response headers tell you exactly where you stand:
x-ratelimit-limit-requests
x-ratelimit-limit-tokens
x-ratelimit-remaining-requests
x-ratelimit-remaining-tokens
x-ratelimit-reset-requests
x-ratelimit-reset-tokens
x-ratelimit-limit-project-tokens
x-ratelimit-remaining-project-tokens
x-ratelimit-reset-project-tokens
Retry-After
Reset values come back as durations (6m0s), not timestamps. Log x-ratelimit-remaining-tokens on every call; watching it approach zero is far more actionable than reacting to 429s.
Tiers are unlocked by cumulative spend: Tier 1 at $5 paid, Tier 2 at $50, Tier 3 at $100, Tier 4 at $250, Tier 5 at $1,000, with monthly usage ceilings from $100 up to $200,000. Units include RPM, RPD, TPM, TPD, images per minute, and audio minutes per minute for some streaming models. One limit that is easy to miss: vector store file uploads share 300 requests per minute per vector store (see OpenAI, Rate limits).
The error codes in current docs are not the ones you memorised
| HTTP | code | What to actually do |
|---|---|---|
| 429 | slow_down | Ramp traffic gradually; retry with backoff |
| 429 | credit_balance_exhausted | Retrying will not help. Alert billing. |
| 429 | project_spend_limit_exceeded | A project cap, not a rate limit. Raise the cap. |
| 400 | invalid_request_error with error.param = service_tier | Your org is not entitled to that tier |
| 403 | — | Unsupported country or region |
| 503 | server_is_overloaded | Transient; honour Retry-After |
The distinction that matters for your retry policy: Retry-After appears on transient 429s and on 503s. The docs are explicit that its presence "does not mean that quota, billing, or other errors that require user action can be resolved by retrying." A retry loop that treats all 429s alike will burn through your backoff budget on a billing problem.
Batch API: half price if you can wait a day
The Batch API gives 50% lower cost, a separate and significantly higher rate limit pool, and a 24-hour turnaround. Supported endpoints include /v1/responses, /v1/chat/completions, /v1/embeddings, /v1/moderations and the image endpoints.
Practical constraints: input is JSONL with a unique custom_id per line, one input file can target only a single model, the upload limit is 200 MB, and requests with stream: true are rejected.
This is the right tool for backfills, nightly classification, embedding generation and evaluation runs. It is the wrong tool for anything a user is waiting on. Splitting your workload into interactive and batch paths is usually the largest single cost reduction available, larger than switching to a cheaper model.
Data retention: what you must configure before sending regulated data
Two facts from the data controls page. First, since 1 March 2023 data sent to the API is not used to train OpenAI models unless you explicitly opt in. Second, by default all API usage produces abuse-monitoring logs retained for up to 30 days.
Zero Data Retention removes content from those logs and, on /v1/responses and /v1/chat/completions, makes store behave as false even when your request sets it to true. ZDR requires prior approval from OpenAI; it is not a toggle you flip alone.
Data residency is set per project, with regional endpoints such as eu.api.openai.com. Two caveats worth writing into your design document: regional processing for eligible models released after 5 March 2026 carries a 10% surcharge, and residency does not cover system data, which the docs list as account data, metadata, usage statistics and your structured output schema. If your JSON schema itself encodes sensitive business logic, that is not covered.
Check the deprecation page before you pin a model
The deprecation notice period is at least six months for generally available models, at least three months for specialised variants, and can be as short as two weeks for preview models. Currently published dates worth knowing:
whisper-1,gpt-4o-transcribe,gpt-4o-mini-transcribeandgpt-4o-transcribe-diarizewere deprecated on 26 August 2026 and are removed from the API on 26 February 2027. The replacements named in the docs aregpt-live-transcribeandgpt-transcribe.- Reusable prompts (
v1/prompts) shut down on 30 November 2026 (see OpenAI, Deprecations). OpenAI's current advice is to "store production prompts in your application code instead of creating reusable prompt objects." - Self-serve fine-tuning stops accepting new jobs on 6 January 2027.
Pin an explicit model snapshot in production, and put the deprecation page in your quarterly review. A model string in a config file is a dependency with an expiry date.
Next step
Before writing more integration code, do one measurement: log usage.input_tokens_details.cached_tokens for a day of real traffic. If it is near zero, fix your prompt prefix and tool ordering before optimising anything else. Then split whatever does not need a live response onto the Batch API. Those two changes typically move the bill more than a model swap does.
For voice workloads, the transcription side has its own set of limits and a migration deadline; we covered them in the speech-to-text guide. If you want an outside look at an LLM integration before it goes to production, the contact page is the place to start, and AI integrations describes how we approach this work.
Verified on 25 September 2026 from OpenAI's own documentation, retrieved as markdown by appending .md to the docs URL: migrate-to-responses, rate-limits, error-codes, structured-outputs, function-calling, prompt-caching, batch, your-data and the deprecations page, plus the pricing page. Model names, prices and deprecation dates move quickly; confirm against the live docs before you commit a model string to production.