I work at OpenAI and I was the Incident Commander for yesterday's outage.
We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.
Ha. It's a role/title for the lifetime of the incident -- it's useful to have someone to keep things moving, keep track of workstreams, and to know who the decision-maker is, especially for bigger incidents. I'm just a SWE who works on infrastructure.
Note: Incident Command is almost certain an allusion (or implementation) of the Incident Command System [1]
It is a common system in all kinds of emergency response scenarios, including local emergency services (fire/police/ambulance) and it scales all the way to massive disasters.
It is especially useful to clarify command structures when multiple response entities need to coordinate. That is true even within organizations like public companies, where the reporting structures may be distinct.
Actually, nevermind... this looks more like Anthropic and xAI coincided due to shared xAI infra (6:23am and 6:30am), and OpenAI's issue was more likely then a coincidence (7:43am).
We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.