When ChatGPT or Codex fails, incident titles are usually a variant of “elevated errors”. For 15 of the 106 incidents on OpenAI’s page since July, the company went further and published a write-up with a stated root cause. We read all of them. Ten of the 14 distinct write-ups blame a change or action inside OpenAI itself, such as a configuration update, a code change or an automated cleanup. Three blame an unnamed infrastructure provider’s maintenance, and one cites a capacity shortfall. Anthropic’s status page rarely states a cause, and the two engineering postmortems we found concern response quality, not outages.
Key takeaways
- OpenAI linked a write-up from 15 of the 106 incidents in its feed (14%), which cover 14 distinct events because two incidents on 25 July share one write-up.
- By our reading of OpenAI’s wording, 10 of the 14 were triggered by something inside OpenAI: 5 configuration or deployment changes, 3 code or software changes and 2 automation or security actions. 3 were an infrastructure provider’s maintenance and 1 was a capacity shortfall.
- None of the 14 names a provider. “Infrastructure provider” appears 11 times and “cloud provider” once; Azure, AWS, Google Cloud and Microsoft appear nowhere.
- On Anthropic’s page, only 3 of 50 incidents name a cause in an update (a Windows update, an unnamed “upstream cloud provider” and a Google Play issue). Its two engineering postmortems we found cover model-quality bugs, not availability.
Last reviewed 6 October 2026. We read OpenAI’s public RSS feed (106 incidents, 9 July to 6 October; we saved a copy of it), opened every incident page that links a write-up, and read the 15 write-ups in full. We also read Anthropic’s public incident feed (its 50 most recent incidents, 25 July to 5 October) and the engineering index and two postmortem pages on anthropic.com. Times in the write-ups are Pacific (PDT is UTC minus 7); we keep OpenAI’s wording and label them. Our grouping of causes is our own reading, not OpenAI’s.
How many OpenAI incidents come with a write-up?
OpenAI’s status page marks an incident with a write-up when one exists, and the write-up lives at status.openai.com/incidents/<id>/write-up. We read OpenAI’s incident feed (retrieved 2026-10-06), which lists 106 incidents, and checked each incident page for a write-up link. Fifteen have one. The oldest is dated 14 July, so every write-up falls in about the last twelve weeks, and 91 incidents in the same feed have no write-up we could find. That is not necessarily a gap, since we have not checked how long or severe those 91 were, and we cannot tell how OpenAI decides which ones get a write-up.
One recent case is worth watching. On 29 September, OpenAI marked an incident titled “Elevated errors across ChatGPT, Codex, and the API including the Agents API” resolved after about 5 hours 20 minutes (17:52 to 23:14 UTC), with the note that “the detailed Root Cause Analysis (RCA) will be published in the next 5 business days”. Counting from that day, the window runs to about 6 October. When we read the incident page on 6 October at about 16:40 UTC there was no write-up link yet, so we have not included it and cannot say what it will conclude.
What does each write-up say caused the incident?
Each row below is our short paraphrase of the “Root Cause” and “Summary” sections of OpenAI’s own write-up, linked in the first column. The window is the impact period OpenAI gives, in Pacific time as written, with the duration we calculated from it. Where a write-up describes two separate periods, we show both.
| Date (2026) | Product | Window OpenAI gives (Pacific) | Stated cause, in brief | Our grouping |
|---|---|---|---|---|
| 25 Sep | Codex (sign in with ChatGPT) | 3:33 to 4:47 p.m., about 74 min | A credential leak-detection system wrongly flagged internal service credentials as leaked; they were revoked by a manual operation that bypassed safeguards. | Automated cleanup or security action |
| 3 Sep | ChatGPT and Codex | 7:43 to about 8:20 a.m., about 37 min | The routing system misread a configuration update and stopped sending requests to available services; returning traffic then overloaded some of them. | Configuration or deployment change |
| 2 Sep | ChatGPT Work Mode | 4:44 to 5:10 p.m., about 26 min | A configuration update meant to enable another service removed access that existing Work Mode services needed. | Configuration or deployment change |
| 2 Sep | New account creation | 10:49 to 11:31 a.m., about 42 min | An automated cleanup removed access a sign-up service still needed, because its usage records did not capture all internal activity. | Automated cleanup or security action |
| 31 Aug | ChatGPT conversations (Free, Go, signed-out) | 5:20 to 5:51 p.m., about 31 min | A code change to a supporting feature introduced a compatibility issue; it was rolled back at 5:46 p.m. | Code or software change |
| 31 Aug | ChatGPT Work | 7:40 a.m. to 12:53 p.m., about 313 min | A code change raised demand while background activity added network connections, exceeding network capacity; retries and memory pressure added to it. | Code or software change |
| 19 Aug | Sign-in and ChatGPT | 4:50 to 5:19 p.m. PT, about 29 min | A deployment applied two incompatible configurations to the authentication systems. | Configuration or deployment change |
| 30 Jul | ChatGPT conversations | 6:20 to 7:47 a.m., about 87 min | Available capacity was insufficient for the processing some conversations needed. | Capacity shortfall |
| 21 and 27 Jul | Image generation in ChatGPT | 21 Jul 4:52 to 7:47 a.m. (residual errors to 7:50 p.m.); 27 Jul 6:20 a.m. to 1:22 p.m., about 422 min | A software update for image generation increased memory use and cut capacity; new capacity hit the same behaviour, causing the second period. | Code or software change |
| 25 Jul | ChatGPT, API Platform, Codex | 1:59 to 2:50 and 4:24 to 4:58 a.m., about 85 min in two periods | A configuration update with routing changes was incompatible with an internal service version; a deployment-system bug then re-applied it automatically. | Configuration or deployment change |
| 23 Jul | ChatGPT, API Platform, Codex, file features | 7:44 a.m. to 2:26 p.m., about 402 min | An infrastructure provider’s network maintenance expanded in scope and removed routes; its automated DDoS protections then dropped traffic. | Infrastructure provider’s maintenance |
| 21 Jul | Logins and sign-ups (ChatGPT, API Platform, Codex) | 11:18 to 11:40 a.m. and 3:00 to 3:10 p.m., about 32 min in two periods | Unintended maintenance by an infrastructure provider reduced a database’s backup capacity; the provider then restarted it twice, once by faulty automation. | Infrastructure provider’s maintenance |
| 19 Jul | ChatGPT and Codex | 7:08 to 8:05 a.m., about 57 min | Cloud maintenance made a regional database replica unavailable; remaining capacity was too small and failover did not move enough traffic. | Infrastructure provider’s maintenance |
| 14 Jul | ChatGPT on the web | 4:39 to 5:19 p.m., about 40 min | A web traffic-routing configuration change sent far more traffic than intended to a small subset of capacity. | Configuration or deployment change |
The windows are mostly short. In 11 of the 14 write-ups the longest single impact period is under 90 minutes, and the median of the longest single period in each write-up is 46.5 minutes (the table shows totals where there are two periods). Three run longer: 31 August ChatGPT Work (about 5 hours), 23 July (about 6.7 hours, the provider-maintenance incident that also reached the API Platform and file features) and the second image-generation period on 27 July (about 7 hours). The first image-generation period, on 21 July, lasted about 175 minutes, with residual errors for a subset of requests until 7:50 p.m. Durations are approximate, taken from OpenAI’s “approximately” times, and they describe the impact period OpenAI chose to report, which may differ from how long any given user was affected.
What the causes have in common
Ten of the 14 write-ups name a trigger that starts inside OpenAI. Five are configuration or deployment changes: a routing configuration on 3 September, a Work Mode access change on 2 September, two incompatible authentication configurations on 19 August, an incompatible production configuration on 25 July and a traffic-routing change on 14 July. Three are code or software changes (31 August twice, and the image-generation update). Two are automated or security actions that removed access a service still needed (25 September and the 2 September account-creation cleanup).
Three more start with an infrastructure provider’s maintenance, and one is an unattributed capacity shortfall on 30 July. Even in the provider cases, OpenAI’s own text lists its own contributing factors. On 19 July it says existing failover “did not automatically redirect enough traffic” away from the affected region, and the improvements it lists include automatic failover and cross-region rerouting. On 23 July they include regional failover and traffic steering. We read that as OpenAI treating those incidents as partly a resilience problem on its side, though that is our interpretation, not a stated conclusion.
The classification has limits. A write-up often names several things: the 31 August ChatGPT Work incident was triggered by a code change but was resolved partly when “our infrastructure provider added network capacity”, and the 3 September incident was prolonged by overload during recovery. We grouped by the first trigger OpenAI states. A reader who weighted the trigger differently would move one or two incidents between groups, but we would not expect them to move the overall pattern from “mostly internal changes”.
An outage the news blamed on a cloud provider
We looked at the 3 September incident in an earlier article, where several outlets pointed to an Azure outage as the shared cause of ChatGPT, Claude and Gemini problems: News said one Azure outage took down ChatGPT, Claude and Gemini together. We checked the three companies’ own pages. OpenAI’s write-up for that day, in the table above, describes its own routing system misreading a configuration update and does not mention Azure, Microsoft or any outside vendor. That does not prove no provider was involved anywhere, and the write-up is OpenAI’s account of its own incident only. It does show what OpenAI itself says about that day, and it is the more specific source.
What OpenAI says it will change
Every write-up ends with improvements, and the same ideas recur. Several promise safer configuration rollouts: on 3 September, staged rollouts and an evaluation of health-check-based automatic rollback, and on 14 July, limits on how much production traffic a staged rollout can shift and confirmation prompts and impact previews for risky routing changes. Several promise stronger checks before a change ships: on 25 July, validating configuration against the service versions that must support it, and on 19 August, warnings for incompatible or incomplete configuration. A few promise resilience work so a failing part does not stop the product, such as on 31 August keeping conversations available when a nonessential feature fails and on 21 July letting new logins work through a database outage. We did not count how many of the promised items have been delivered, and the write-ups do not report that.
One item is specific. On 25 July the write-up says a bug in the deployment system let an automated deployment re-apply a configuration that had already been rolled back, and that the bug “was subsequently fixed”. It also describes two separate periods of impact, 51 and 34 minutes, which fits the two incidents in the feed for that day.
What does Anthropic publish?
Anthropic’s public status page runs on Atlassian Statuspage, and we read its feed (retrieved 2026-10-06; 50 incidents, 25 July to 5 October). Only three incidents state a cause in an update. On 10 September, “Degraded functionality for Claude Cowork on Windows” said that a Windows update released on 8 September had left Cowork unable to run local commands, that Microsoft had developed a fix, and the closing update linked Microsoft’s 14 September update. We covered that in our article on the Windows incident. On 28 August, “Elevated errors on Claude Code and Claude Cowork” said Anthropic had identified “an issue with an upstream cloud provider”, without naming it. On 16 September, “Issues with Google Play subscriptions” said Google had confirmed an issue with Google Play. In the other 47 incidents we found no stated cause: we searched every update for cause wording and read the matches, and some say only that a root cause was identified without saying what it was.
That does not mean Anthropic publishes no postmortems. On its Engineering page (retrieved 2026-10-06) we found two, and both are about response quality, not availability. “An update on recent Claude Code quality reports” (23 April 2026) traces complaints to three changes: a default reasoning-effort change on 4 March (reverted 7 April), a caching bug from 26 March that cleared older reasoning on every turn (fixed 10 April), and a system-prompt line limiting verbosity added on 16 April (reverted 20 April). It states the API was not affected. “A postmortem of three recent issues” (17 September 2025) describes three infrastructure bugs that intermittently degraded responses between August and early September 2025: a context-window routing error, output corruption from a TPU misconfiguration and an approximate top-k compiler bug.
So the two companies publish different things. OpenAI’s write-ups are short, structured and about availability incidents, with a summary, impact, root cause, resolution and prevention list for each. Anthropic’s two postmortems are long engineering essays about quality regressions, published months apart, while its status updates for outages mostly carry no cause. We checked the status feed and the Engineering index only. We did not search Anthropic’s other channels, so we cannot say there are no others. The two postmortems date from September 2025 and April 2026, so they do not cover the same period as OpenAI’s July to September 2026 write-ups. We also cannot say which approach serves users better, since they answer different questions.
What this means when something breaks
- Check what the write-up says was unaffected. The 25 September Codex write-up says access through customer-provided API keys was unaffected while sign-in with ChatGPT failed, and several others say the API Platform or core inference was not affected. If your work can run through an API key, that may be a workaround, but only the write-up for that incident can tell you.
- Do not assume a provider outage. In OpenAI’s own account, as we read it, most of these incidents began with a change inside OpenAI. When several services fail at once, read each company’s own incident page before accepting one shared explanation. Our earlier look at a related question is in our comparison of ChatGPT’s and Claude’s September incident counts.
- Treat the status page’s window as OpenAI’s account. Our measurement of the monitoring stage shows how much of an incident’s displayed length can come after the company says a fix is in.
What we could not verify
- We cannot confirm any write-up’s account independently. Each is OpenAI’s own description, and we have no way to check, for example, that a configuration was the trigger.
- We do not know how OpenAI decides which incidents get a write-up, so the 14% coverage figure describes what was published, not what was warranted.
- Our five-way grouping is our own reading of OpenAI’s wording. Several write-ups name more than one factor, and we grouped by the first trigger stated.
- The durations are calculated from the approximate times OpenAI gives, in Pacific time as written, and cover the impact period OpenAI reports, not individual users’ experience.
- On 19 July the write-up says “cloud infrastructure maintenance” and only later names “the cloud provider”, so we infer the maintenance was a provider’s; and the three provider cases may or may not be the same provider.
- We did not verify that any promised improvement was delivered, and the 29 September incident’s promised write-up had not appeared when we looked.
- On the Anthropic side we read the status feed and the Engineering index only, and its feed holds its 50 most recent incidents, not a full history.
How we researched this, and the data
On 6 October 2026 we fetched OpenAI’s RSS feed, which lists each incident with its start and resolution time, and opened every incident page to see whether it links a write-up. For each of the 15 linked write-ups we read the Summary, Impact, Root Cause, Resolution and Prevention sections, and for the two 25 July incidents, which share one write-up, we counted it once. We then read Anthropic’s JSON incident feed for cause wording in every update, and fetched the two Anthropic engineering pages. The times and the 106-incident total are from OpenAI’s own feed (retrieved 2026-10-06), and the Anthropic counts are from its incident feed (retrieved 2026-10-06). The group counts were tallied by hand from the table above, and the vendor-name and phrase counts were done by a script over the write-up text. Our related lifecycle data for both providers’ incidents is in this CSV.
Frequently asked questions
OpenAI gave its own explanation for several September incidents. On 3 September it said a routing configuration change prevented requests from reaching its services, on 2 September that a configuration update removed access Work Mode services needed and that an automated cleanup removed access an account-creation service needed, and on 25 September that a credential leak-detection system wrongly flagged internal Codex credentials. Each is OpenAI’s own account.
Not for most incidents on its status page. In the feed we read, 15 of 106 listed incidents (14%, including some non-outage notices) link a write-up, covering 14 distinct events. We could not tell how OpenAI decides which incidents get one. One recent incident, on 29 September, had a write-up promised within 5 business days that we had not found by 6 October.
None name a provider. Three of the 14 describe maintenance by an unnamed infrastructure or cloud provider as the trigger (the write-ups do not say whether it is the same one), and the phrase “infrastructure provider” appears 11 times across four write-ups, but Azure, AWS, Google Cloud and Microsoft are not mentioned. Ten of the 14 describe a trigger inside OpenAI.
On its status page only 3 of 50 incidents state a cause. Its Engineering page has two postmortems we found, from 23 April 2026 and 17 September 2025, both about degraded response quality and not availability. We did not search other Anthropic channels.
