OpenAI Outage Puts Codex Reliability on Trial

OpenAI resolved two July 25 incidents across APIs, ChatGPT and Codex, then kept a later ChatGPT conversation issue under monitoring.

By Arkolith Newsroom3 min read
an empty logo-free server operations room during a cloud-service incident response.

OpenAI logged three separate service incidents on July 25, including two resolved elevated-error events that affected APIs, ChatGPT and Codex, followed by a later ChatGPT conversation issue that was still under monitoring on Sunday morning in Europe.

The incident matters because Codex and ChatGPT Work are no longer side experiments. They are production workflow tools, and a failure window now interrupts code review, agent tasks, workspace search and ordinary chat at the same time.

What OpenAI reported

OpenAI's OpenAI status history lists two July 25 elevated-error incidents that affected 12 API components, 15 ChatGPT components and four Codex components. One began at 9:17 a.m. on July 25 and was marked fully recovered at 11:08 a.m. The OpenAI broad elevated-error incident was identified at 11:35 a.m. and marked fully recovered at 11:57 a.m.

The later incident was narrower but still live at publication time. OpenAI's current OpenAI system status said it was experiencing issues with elevated errors affecting ChatGPT conversations, with mitigation implemented and recovery under monitoring. The OpenAI ChatGPT conversation incident detail says the impact period began around 1:00 p.m. PT on July 25 and could prevent some users from loading or continuing conversations.

Photo: an empty logo-free server operations room during a cloud-service incident response

Why Codex changes the outage math

A chatbot outage is disruptive. An agent outage is operationally different because users may have delegated a multi-step task, a code review, a workspace search or a file-dependent workflow before the service began failing.

That distinction showed up in the conversation around the outage. In a Tibo Sottiaux X post, OpenAI's Tibo Sottiaux said usage limits had been reset for all Codex and ChatGPT Work users after an almost global outage during the night. The post is evidence of the operator response and reset, not a substitute for the status-page record of affected services.

The reset also leaves a commercial question. If agent products are sold as work infrastructure, users will judge them less like experimental chat tools and more like cloud services: by recovery time, incident frequency, transparency and whether failed work can be resumed without hidden cost.

The countercase

The official record does not prove that all users were affected, and OpenAI cautions that availability metrics are aggregated across tiers, models and error types. Individual customer experience can vary by subscription tier and feature.

It also matters that the two broad API, ChatGPT and Codex incidents were resolved on July 25. The live issue at publication time concerned ChatGPT conversations, not a fresh official Codex outage. Reports of stream disconnects and failed tasks on X are useful discovery signals, but they should not be treated as a confirmed Codex incident unless OpenAI records or verifies them.

What to watch next

The follow-through is whether OpenAI closes the conversation incident cleanly and whether the July cluster becomes a reliability footnote or a buyer objection. The same operational standard is why OpenAI's earlier Hugging Face security incident and medical-records rollout are not only feature stories. They are tests of whether high-trust AI products can explain boundaries when something breaks.

For enterprise and developer teams, the watch items are plain: incident timelines, component-level status detail, retry semantics for interrupted agent work, billing or limit adjustments after failures, and whether Codex tasks can resume from a known state instead of restarting from scratch.

This article is informational only and is not investment advice.

#OpenAI#Codex#ChatGPT#AI infrastructure#Reliability