OpenAI’s write-up for its 29 September 2026 outage says a feature rollout caused routine settings requests to trigger repeated background checks, which opened connections across internal services until shared network capacity ran out. The write-up was linked from the incident page by 7 October, about eight days after the incident. Its start and end times roughly agree with the status page’s, but its mitigation time does not: it says most customer-facing impact was mitigated within about 30 minutes, while the status page kept saying “investigating” for about another hour and a half, and its “mitigation applied” update came about 100 minutes after the write-up’s mitigation time.
Key takeaways
- Cause, per OpenAI’s write-up (our paraphrase): a feature rollout made settings requests trigger repeated background checks, and failed checks generated more attempts, exhausting shared network connection capacity.
- OpenAI says impact started at about 10:30 a.m. PDT (17:30 UTC), 22 minutes before the status page’s first post at 17:52 UTC, and that most customer-facing impact was mitigated by about 11:00 a.m. PDT (18:00 UTC).
- The status page said “investigating” until 19:08 UTC and posted “mitigation applied” at 19:39 UTC, then four identical monitoring updates until “resolved” at 23:14 UTC. The write-up says some services had residual degradation until about 4:07 p.m. PDT (23:07 UTC).
- At least twice before, OpenAI described an internal change followed by shared capacity running short: 2 June 2026 (a close match) and 31 August 2026 (a looser one).
Last reviewed 7 October 2026. We read OpenAI’s write-up and the incident’s updates from its JSON feed (both retrieved 2026-10-07), plus two earlier write-ups. The write-up gives times in Pacific time (PDT is UTC minus 7) and calls them approximate; the status page’s times are exact UTC timestamps. Where the two disagree we say so and do not pick one. The grouping and comparisons are our own reading.
What does OpenAI’s write-up say happened?
The write-up, retrieved 2026-10-07, says that on 29 September at approximately 10:30 a.m. PDT some users experienced elevated errors in ChatGPT, Codex and the API, including the Agents API. Users saw failed requests, timeouts or difficulty signing in to ChatGPT and Codex, some API requests failed, and some conversations and tasks could not start, resume or complete.
The root cause section says a feature rollout caused routine settings requests to trigger repeated background checks. Those checks opened new connections across multiple internal services, and failed checks generated additional attempts. That exhausted shared network connection capacity and stopped otherwise healthy services from reliably communicating with each other. Engineers disabled the rollout, and the write-up says most customer-facing impact was mitigated within about 30 minutes, at about 11:00 a.m. PDT. To deal with residual errors they routed traffic around unhealthy infrastructure, reduced repeated requests and increased database cache capacity, and the remaining degradation was resolved by about 4:07 p.m. PDT.
The prevention section has two items: OpenAI says it disabled the triggering feature and removed the code that initiated repeated checks during settings requests, and that it deployed a fix for expensive database queries and is expanding network capacity and reviewing other queries to reduce recovery bottlenecks. It does not name the feature, say how many users were affected, or say whether the rollout was staged. Unlike several of its other write-ups, it does not list changes to rollout or validation practice.
How do the write-up’s times compare with the status page’s?
The incident’s updates in OpenAI’s JSON feed (retrieved 2026-10-07) give exact UTC times. We placed the write-up’s approximate Pacific times beside them.
| UTC | Pacific (PDT) | Source | What it says |
|---|---|---|---|
| about 17:30 | about 10:30 a.m. | Write-up | Impact begins. |
| 17:52 | 10:52 a.m. | Status page | First post: investigating elevated errors in ChatGPT, Codex and the API. |
| about 18:00 | about 11:00 a.m. | Write-up | Most customer-facing impact mitigated. The write-up also says engineers disabled the rollout, but gives no time for that step. |
| 18:24 | 11:24 a.m. | Status page | Investigating; the Agents API is now named. |
| 19:08 | 12:08 p.m. | Status page | “We’re still investigating.” |
| 19:39 | 12:39 p.m. | Status page | Mitigation applied, monitoring recovery. Repeated at 21:09, 22:03 and 22:47. |
| about 23:07 | about 4:07 p.m. | Write-up | Remaining degradation resolved. |
| 23:14 | 4:14 p.m. | Status page | Resolved; a root cause analysis promised within 5 business days. |
| 6 Oct, 08:26 | 6 Oct, 1:26 a.m. | Status page | A second “resolved” update: “All impacted services have now fully recovered.” |
Two gaps and one match. OpenAI’s write-up puts the start about 22 minutes before the status page’s first post; that lag looks ordinary, since the 25 September Codex write-up shows 25 minutes (start 3:33 p.m. PDT, first status post at 22:58 UTC). A separate support-chat incident was posted at 17:49 UTC. The mitigation gap is the unusual one: the write-up’s time, about 18:00, is about 100 minutes (99 by the clock times, which are approximate) before the status page’s “mitigation applied” update at 19:39, and the page still said “still investigating” at 19:08, about 70 minutes after the write-up’s mitigation time. The match is the end: the status page’s “resolved” at 23:14 UTC is within a few minutes of the write-up’s end of residual degradation, about 23:07. For comparison, on 25 September the write-up’s recovery time (about 4:47 p.m. PDT, 23:47 UTC) and the status page’s monitoring update (23:45 UTC) were two minutes apart, so a gap this large is not how every incident looks.
These are two accounts and not one account with an error. The write-up’s times are approximate, and “most customer-facing impact” is not the same as “all impact”. The status page’s “mitigation applied” is a stock sentence OpenAI uses across incidents, and it may describe the later work on residual errors rather than the first step, which is our guess and not something either source says. We cannot tell which account better matches what users saw, and we do not have independent measurements of error rates.
What was the three and a half hours of “monitoring”?
The status page posted its first “monitoring” update at 19:39 UTC and resolved at 23:14, a stage of about 215 minutes (3 hours 35 minutes by the clock times), with the same sentence posted four times. In our measurement of status-page stages we noted that we could not tell, from the feed alone, whether a long stage reflects real risk or a slow close-out. For this incident the write-up gives part of the answer: it says some services continued to experience residual degradation until about 4:07 p.m. PDT, which is inside that stage. So by OpenAI’s own account, the monitoring stage here was not just waiting.
The write-up does not say which services were degraded after the first 30 minutes, how badly, or for how many users, so we cannot say how much of the 215 minutes involved real impact. We also do not know whether the status page’s “monitoring” updates were meant to cover the residual degradation or only the first mitigation. A reader who relied on the status page alone would have seen only an unchanging “monitoring” message for the whole afternoon.
Has OpenAI described this kind of failure before?
In part, at least twice in the write-ups we read, though our systematic review starts in July and earlier ones may exist. The 2 June write-up describes an internal service repeatedly opening connections to a shared backend, much as the September one does. The 31 August write-up is a looser match: a code change increased demand and background activity added network connections, which exceeded available capacity. Each trigger is different.
| Date (2026) | Trigger, per OpenAI | What it says was exhausted or degraded | Write-up |
|---|---|---|---|
| 2 June (PDT) | A configuration rollout caused an internal service to repeatedly establish connections to a shared backend dependency. | Capacity within shared infrastructure; Responses API latency, Codex 429 errors, ChatGPT login failures. | June write-up |
| 31 August | A code change increased demand on ChatGPT Work systems while background activity added network connections. | Available network capacity; Work tasks failed or were slow for about 5 hours. | August write-up |
| 29 September | A feature rollout made settings requests trigger repeated background checks that opened new connections. | Shared network connection capacity; failures across ChatGPT, Codex and the API. | September write-up |
The three differ in what was changed (configuration, code, a feature), and in what the write-ups promise afterwards. The June one lists stronger validation, canary deployments and rollout controls. The August one lists reduced network demand, better retry behaviour and more capacity. The September one lists removing the triggering code and expanding network capacity. In our review of the write-ups linked in OpenAI’s feed since July, 10 of 14 distinct write-ups named an internal trigger. With this one, which we group as a configuration or deployment change (a judgement call, since the write-up also says the code was removed), it is 11 of 15 in our reading. We would not conclude that OpenAI has a recurring root cause: each write-up is OpenAI’s own account, and three write-ups, one of them a loose match and one outside our July to October review, are too few to establish a pattern. They are consistent with a family of failures, shared capacity running short after an internal change, which OpenAI has described more than once.
What this means if you build on OpenAI’s API
- The status page’s “mitigation applied” may not mean you are fully recovered. For this incident OpenAI’s write-up says residual errors lasted until about 4:07 p.m. PDT, long after both the write-up’s mitigation and the status page’s. If you run batch jobs or agents, watch your own error rate before restarting them.
- Check the write-up for what was and was not affected. This one says some Agents API requests failed. Others of OpenAI’s write-ups say the API was unaffected, so the answer is incident-specific.
- Don’t assume the explanation arrives with the incident. OpenAI promised an analysis within 5 business days. We did not find it at about 16:40 UTC on 6 October, the fifth business day, and it was linked by 7 October. During the incident itself, only the status page was available.
What other pages said
The one aggregator we checked, UptimeRobot’s page for 29 September (retrieved 2026-10-07), repeated the status page’s facts and gave no cause. It lists two OpenAI incidents that day, one for ChatGPT, the APIs and related services starting at 17:52 UTC and one for the Help Center and support chat starting at 17:49 UTC, and gives no cause for either. OpenAI’s own feed has the support-chat incident too, from 17:49 to 23:18 UTC. The write-up does not mention support chat, so we do not know whether they shared a cause. We did not check user-report sites such as Downdetector, so we have no user-side timeline to test the two accounts against.
What we could not verify
- We cannot independently confirm the write-up’s cause or any of its times. Both are OpenAI’s account.
- The write-up’s times are approximate. A start of 17:30 UTC may be a few minutes either way, so the 22-minute gap before the first status post is approximate.
- We do not know why the status page’s mitigation update came about 100 minutes after the write-up’s mitigation time, or what the 6 October “resolved” update was for. The feed does not say.
- We do not know exactly when the write-up was published. It was not linked from the incident page at about 16:40 UTC on 6 October and was linked when we looked on 7 October. The write-up also mentions database queries and cache capacity in its recovery and prevention sections without saying how they relate to the rollout.
- We did not measure user-side error rates, so we cannot say how many users were affected in either the first 30 minutes or the residual period.
- The comparison with the 2 June and 31 August incidents rests on OpenAI’s descriptions only, and the three write-ups describe similar mechanisms, not the same one.
How we researched this, and the data
On 6 and 7 October 2026 we read the incident page and fetched the write-up from OpenAI’s status site. We extracted the Summary, Impact, Root Cause, Resolution and Prevention sections, and pulled the incident’s updates with their UTC timestamps from OpenAI’s incident feed (retrieved 2026-10-07). We converted the write-up’s Pacific times to UTC by adding seven hours, and calculated the gaps in minutes by script. We then read the two earlier write-ups in full and checked the aggregator page for 29 September. Our earlier articles on OpenAI’s September incident count and on the write-ups published since July give the wider context.
Frequently asked questions
OpenAI’s write-up says a feature rollout caused routine settings requests to trigger repeated background checks. The checks opened new connections across internal services, and failed checks produced more attempts, which exhausted shared network connection capacity. It says engineers disabled the rollout. That is OpenAI’s own account.
The status page shows 17:52 to 23:14 UTC, about 5 hours 20 minutes. The write-up puts the start at about 10:30 a.m. PDT (17:30 UTC), says most impact was mitigated within about 30 minutes, and says residual degradation on some services lasted until about 4:07 p.m. PDT (23:07 UTC).
They give times in different time zones and at different precision, and describe overlapping but not identical events. The write-up separates the first 30 minutes from the residual period. The status page shows a first post at 17:52 UTC, a mitigation update at 19:39 UTC and a resolved update at 23:14 UTC. The start and end roughly agree, the mitigation times differ by about 100 minutes, and neither source explains why.
At least twice in the write-ups we read. The 2 June 2026 write-up describes a similar mechanism (a configuration rollout that made an internal service repeatedly open connections to shared infrastructure), and the 31 August one is a looser match (a code change and background activity that exceeded network capacity). Each is OpenAI’s account of its own incident.
