---
title: "Incident Report: September 30, 2026: Domain Routing Disruption"
description: "Railway experienced a platform-wide disruption to domain routing on September 30, 2026. For roughly five minutes, between about 07:35 and 07:40 UTC, new connections to domains hosted on Railway returned 404 errors."
date: 2026-09-30T22:42:12.172Z
authors: ["Angelo Saraceno"]
category: "Engineering"
url: https://blog.railway.com/p/incident-report-sept-30-2026-routing-disruption
---

# Incident Report: September 30, 2026: Domain Routing Disruption

On September 30, 2026, domains hosted on Railway returned 404 not found errors for new connections for a five minute period for all customers.

While rolling out a change to our distributed routing control plane, the new deployment didn’t wait for the schema updates needed to support new functionality. Our distributed routing control plane is the system that tells our edge how to route user requests to the workloads behind every domain.

Because of this mismatch, route lookups failed, and any attempt to reach a domain served by Railway-hosted workloads returned an error. During the incident period, our edge treated each failed lookup as if "this domain doesn't exist" instead of falling back to the last route it knew.

After five minutes of the system automatically reconciling to the intended working state - it restored most traffic; then routing recovered on its own as soon as the schema update finished.

## Impact

On September 30, 2026, between roughly 07:35 and 07:40 UTC, new connections to domains hosted on Railway (which include custom domains and `*.up.railway.app` domains) returned HTTP 404 in all regions.

Public access and some private networking lookups between services failed during this window. User service uptime was not impacted, only access as workloads kept running throughout we are able to confirm that no data was lost.

## Incident Timeline

*All times are UTC on September 30, 2026.*

- **07:29**: An engineer merges a change to our routing control plane after testing it in our staging environment. The change includes a database schema update along with a new version of the routing service.
- **\~07:35**: Our deployment system builds and rolls out the new routing service in every region. Unfortunately, a race in the configuration ran the schema update pipeline in parallel with the deployment pipeline. Route lookups immediately begin to fail, and internal alerting engages the on-call response team.
- **\~07:40**: Our CI system applies the schema update, the routing service retries its lookups, and domains immediately begin routing normally again.
- **07:44**: Engineers identify the cause: the new routing service came online before the schema it depended on.
- **07:47**: We confirm routing has fully recovered. No rollback is needed.
- **07:48**: A public incident is posted to our status page.
- **08:17**: The incident is marked resolved.
- **After**: The engineering team pauses other work and focuses entirely on delivering remediations to prevent this issue from recurring.

The full incident is available on our [Status Page](https://status.railway.com/incident/IFBXJTGG).

## What Happened?

To roll out a new feature that adds safeguards to hosted workloads, we needed a database schema update to enable the new functionality.

The distributed control plane was introduced recently to keep running in case of a major cloud outage. We run this control plane all over the world, and roll out pipelines in parallel. An error in the pipelines configuration also can the schema migrations in parallel. This opened us up to a race case, particularly apparent with larger migrations.

Existing connections weren’t disrupted but new visits to user pages would error for the period in which the schema was mismatched.

Concretely:

1. It was possible for new versions of our network plane to be deployed faster than migrations.
2. The change in question was a change to the domain fallback behavior; which led to the widespread impact.

We’ll explain how these two factors combined to the observed failure.

### 1) A schema update and a binary rollout in the wrong order

We are building an authentication layer that lets customers put a login in front of their services. With all features that we develop, we ship features disabled after extensive testing in our development and staging environments. However, it required a modified data type in the database that stores routing data, along with a new version of the routing service that reads the new table format.

Those two pieces deploy through separate pipelines.

The schema update runs as a migration job within one system, and the routing service is built by another. When the routing service was part of a larger monolith (aka the non-distributed control plane), a single pipeline enforced that order. After we split it out, into the distributed control plane to add additional resiliency, the entire system was parrallelized.

On September 30, the new routing service build won the race. As it rolled out within every region the API started querying columns that didn't exist yet. Because those queries touched the modified database table, every route lookup returned an error.

Once the migration finished, API calls returned 200 again, which led to the quick recovery.

That raises a question: why did all new connections return the error that they did?

### 2) Joint impact and blast radius

For every request we process, our edge asks the routing service which service the request for that domain should go to. The routing service answers with the matching routes: sometimes one, sometimes many, and none if the domain doesn't exist.

During the incident, the lookup itself failed, and the edge handled that failure the same way it handles a domain with no routes: by returning a 404. Because the fallback behavior had been modified, the edge couldn’t fall back to the last route it had seen for that domain.

This routing service also resolves private network DNS, so private networking lookups between services failed in the same window.

Once the migration fully completed, synchronizing with the deployment, the next retry succeeded and routing resumed. This is when all customers reported recovery.

## Preventative Measures

Immediately, the network engineering team identified and triaged components of the system, including the following:

- **Forcing serial ordering of the migration rollout**
  - For this change and others, we perform dry runs of the migrations and have built in cautions. That said, we have acted on additional opportunities to unify the deployment pipeline to make it so their deploy order and blast radius get checked before merge. We had these checks in place when this system was a monolith, and we have tested and ensured the correct behavior within the system.
- **Blocking binary’s blue/green rollout on schema validation**
  - This one is simple: rollouts are independently validate the schema version they are authored with, and the service refuses to go live until the matching database schema is in place.
- **Shard based rollouts per region:**
  - We maintain shard based rollouts for the majority of other systems at the company. The distributed control plane is a new component, and we will be adding shard based rollouts to it to isolate impact going forward.

## Our Incident Procedure

Every major outage triggers a work stoppage on the affected system and requires that we ship guardrails that eliminate the entire class of issue. The team then performs a root cause analysis and reviews all of our operating procedures. For our networking systems, this applies doubly given how critical these services are.

After this outage, we tested the new procedure and confirmed that it mitigates the root cause. Our engineering team takes events like these seriously, and we spare no precaution in making the platform safer and more reliable for you and your customers.


---

Open this post in a browser: https://blog.railway.com/p/incident-report-sept-30-2026-routing-disruption
