Infra Log

Need a Logo?\ View Our Brand Assets

Infra Log

A note on incidents: incidents are internal events for our infrastructure and engineering teams. Incidents often correspond to degraded service on our platform, but not always. This log aims for 100% fidelity to internal incidents, and is a superset both of our status page events and of customer-impacting events on the platform. It includes events reported to subsets of customers on their personal status pages, as well as events without any status page impact.


# September 2: Sprites API partial outage from SQLite database evictions (22:54UTC)

Sprites API stores per-org state in lightweight SQLite databases, each managed by a local (to the organization) control plane node. The nodes periodically write per-database heartbeats to object storage to prove they’re alive. If a node misses a heartbeat, it assumes something is wrong and evicts all of its databases, which then attempt to restore from their Litestream replicas on other nearby nodes.

For about 30 minutes, multiple control plane nodes started failing their heartbeat writes. The failures hit nodes across several regions but were primarily concentrated in Europe. Each failed heartbeat triggered a bulk eviction, displacing hundreds of databases to other failing nodes. Databases continued to hop between nodes until heartbeats resumed.

We’re still not sure exactly what caused the heartbeat writes to start failing. Our leading hypothesis is intermittent connectivity loss between our nodes and Tigris (our object storage provider), amplified by a bug in the retry logic that made subsequent attempts always fail.

September 2: Sidekiq periodic scheduler not running after Redis timeout

# September 2: Sidekiq periodic scheduler not running after Redis timeout (07:15UTC)

Sidekiq uses a leader election to decide which worker process enqueues periodic (aka cron) jobs. The newly elected leader hit a transient network issue communicating with Redis at exactly the wrong moment: the process still became the elected scheduler, but it never completed initialization of its periodic job manager. That left scheduled background work not being enqueued for roughly five hours. The missed jobs included billing usage aggregation, certificate renewal checks, and rechecks for newly requested certificates that had not passed their initial domain validation.

Once we detected the problem, we restarted the wedged scheduler process to restore periodic job execution, and backfilled the missed billing usage.

We’re adding direct monitoring for missed periodic jobs so it doesn’t take us five hours to notice, and improving the reliability of the network between the Sidekiq workers and Redis.


# August 31: Fly-proxy HTTP/2 routing broke due to missing Consul service rows (15:20UTC)

A manual cleanup/config change to our Consul service registration data briefly removed the HTTP/2 routing entries for fly-proxy on a subset of hosts. During that window, some HTTP/2-based traffic to Machines on those hosts could fail. We mitigated the issue by restarting the service-registration sync agent on affected hosts/regions to republish the missing entries, restoring normal routing.


# August 30: Stuck background job queue blocked Sprite deletions (22:10UTC)

An Oban background job queue used by Sprites got stuck when repeated attempts to place it on a Sprites API Machine failed. This caused Sprite deletions and some other background operations to fail. The queue recovered automatically about an hour later, and Sprite deletion requests began succeeding again.

Investigation traced the placement failures to a Sprites API Machine with a corrupted ext4 volume. The Machine had been migrated between hosts several days earlier, which led us to investigate the migration path.

When we migrate a volume, we use Linux’s dm-clone to allow the destination Machine start before the entire volume has been copied: reads for uncopied blocks come from the source host, while writes and already-copied blocks use the destination disk. After the copy completes, the running Machine continues using the clone device until its next restart.

Although the clone and its underlying destination ultimately access the same storage, Linux treats them as separate block devices and maintains an independent page cache for each. A nightly volume check read filesystem metadata directly from the destination device and cached it there. The running Machine continued updating that metadata through the clone, but those writes did not invalidate the destination device’s cached copy. When the Machine was eventually restarted, it read a mixture of stale metadata and current disk contents, corrupting the filesystem.

We now ensure the kernel has invalidated the underlying device’s cache before allowing a Machine to start. The volume check and other maintenance commands have also been updated to read through the clone device when it exists, preventing the incoherent cache from being created in the first place.


# August 26: WireGuard gateways broke after schema migration mismatch (18:10UTC)

We ran a routine deploy of wggwd, the service that sets up customer WireGuard peers, which shouldn’t have impacted anything as the latest code was already deployed (or so we thought). Soon after the deploy we noticed new peer additions failing with a SQL error (wggwd runs a sqlite database on each server to cache peers).

This was tracked down to a new feature whose SQL migration had been merged but not deployed; wggwd was erroring on a select * that picked up more columns than the code expected. However, we were confused by the error, since it was not possible for the wggwd service to both have restarted to run migrations, and not restarted as to fail on the new schema. It turns out this was caused by our alerting, that runs a wggwd check command on an interval; the check command was incorrectly set up to run migrations when opening the database connection.

We closed out this incident by restarting wggwd everywhere to pick up the new code, fixing the check command so it doesn’t run migrations, and ensuring deployments actually restart the server process.


# August 24: Metrics backlog (09:59UTC)

A surge of hot & fresh time-series blew out our storage/query cardinality and caused metrics to appear delayed by 30-60 minutes. We spun up some short term metrics capacity to absorb the load and caught up on the backlog without dropping anything along the way. Once caught up, the metrics cluster has been ticking away happily, albeit at its presumably anomalous new volume. (We’ll get to that soon.)


# August 20: Development node advertised Fly public DNS anycast IPs and broke it (20:08UTC)

An internal development server was misconfigured to advertise our public DNS’s IP address, but the node’s DNS service had been failing. As a result, some recursive resolvers routed queries to that location and experienced brief DNS resolution errors. We resolved this by removing that server from the anycast/DNS serving range so it stopped receiving DNS traffic.

Of course, the most obvious answer here going forward is that this kind of misconfiguration should not happen again, and we need to have more aggressive alerts about misconfigured DNS. Another route of improvement is to make our two DNS anycast IPs actually redundant to each other by announcing them from different subsets of PoPs.

August 20: Some API endpoints briefly rejected flyctl requests

# August 20: Some API endpoints briefly rejected flyctl requests (14:46UTC)

We deployed a change to one of our API backends to reject authentication using legacy OAuth “f01 ” tokens. Due to a bug, this also ended up rejecting the token “bundles” used by flyctl - these bundles include one Macaroon token per org, as well as an user-specific OAuth token which uses the same “f01” prefix as the legacy tokens.

We rolled out an update to the token parsing/rejection logic to allow mixed-token bundles (while still rejecting legacy tokens), and flyctl behavior returned to normal; we also added a code test to help catch similar issues in the future.

August 20: MPGv1 clusters in ORD-0 failed to restart

# August 20: MPGv1 clusters in ORD-0 failed to restart (07:19UTC)

While moving the underlying orchestration for our legacy Managed Postgres v1 service in ORD, some of the control-plane Machines didn’t auto-start (and briefly re-stopped), which prevented a number of ORD-0 Postgres clusters from restarting cleanly. Most clusters recovered quickly once the orchestration layer caught up, but a smaller set remained degraded because replica Machines could not be recreated, including failures returning “insufficient resources to create new machine with existing volume.” We resolved the incident by letting the control plane recover and then manually rescheduling/repairing the remaining affected replicas until all health checks were passing again.


# August 14: Petsem primary crash loop from corrupted volume (19:26UTC)

The primary node of Petsem (our secrets database) suffered from disk corruption during a routine deployment. Petsem can run in a degraded state when the primary is unavailable, so while creating new apps and updating secrets failed during the impact period, it was still possible to perform read actions, such as deploying new machines.

We restored service by promoting an up-to-date replica to become the new primary, repointing the other replicas to follow it, and then provisioning a new replica to return the cluster to full redundancy.


# August 13: Stuck nftables reload blocked flyd startup causing API errors (17:28UTC)

While deploying an nftables change to our worker hosts, a simultaneous flyd (our orchestrator for Fly Machines) deployment happened which saw a number of hosts stuck trying to restart flyd. Because of this, API requests targeted at machines on those hosts were temporarily failing.

flyd was stuck restarting because its systemd service was waiting for nftables to reload, due to Requires=nftables.service and After=nftables.service clauses. These clauses are necessary to ensure that we do not accidentally expose customer VMs to the Internet in case of nftables issues. However, in this case, nftables was taking an abnormally long time to reload, which in turn blocked flyd from restarting in the interim (note: this is due to systemd‘s service startup ordering; in this case, it is not strictly necessary to queue flyd like this since the nftables service was never being fully restarted, only reloaded, but it is how systemd works).

We traced the slow nftables reload down to a suboptimal save/restore script we use for the purpose of machines’ network policies – they are implemented as nftables rules and must be restored on reload. The script to reload these rules was looping through every single machine on a host calling nft list chain repeatedly, which added up to a long time on the more busy hosts. It would not have been a problem if flyd wasn’t being deployed at the same time, either. This incident eventually resolved itself once nftables finished reloading, but we’ll optimize how the reload script works so that this does not happen when we need to reload nftables again. We would like to also optimize the way we use nftables for network policies to reduce the sizes of nftables on worker hosts.


# August 10: Egress IP issues in SJC (06:57UTC)

We saw intermittent outbound connectivity failures for traffic using static egress IP addresses in SJC. The issue turned out to be caused by an unhealthy WireGuard tunnel between a worker host and its egress gateway. This is, in fact, a known bug internally: sometimes our WireGuard tunnels just lose connectivity to one of the peers. Usually this requires a reset of the WireGuard tunnel (fun fact: if you do it wrong, resetting the tunnel might also deadlock the rest of the kernel’s network stack), which we usually will do when we receive an alert for this. Unfortunately, on this day, that specific alert was caught in the middle of another alert storm, and was not noticed and handled for a significant amount of time. We were able to resolve this quickly once we realized this fact, but this incident did tell us a few things we need to work on:

  1. The alert and incident handling process. This is something we now have an entire team for, and we are hopeful that ignoring alerts inadvertently during an alert storm will be a thing of the past soon;

  2. Reliability of our internal network mesh (WireGuard) itself. This bug is already fixed upstream, so any new hosts we have provisioned recently will not see the same issue. For existing hosts, kernel upgrades will be applied as they are gradually rebooted. Of course, this is a somewhat slow process and unless we decided to reboot every existing host at once, we’ll likely continue to see these issues from time to time. Fortunately, we have been working on the ability to reroute traffic over the WireGuard mesh when this kind of connectivity issues happen, including but not limited to this WireGuard bug. Once that work is completed, hopefully these occasional issues will cause much less disruption from users’ perspective.


# August 8: Global Corrosion lag caused Machines API app-not-found errors (23:52UTC)

A subset of very large service-definition updates from machines with large numbers of services could not be applied to our global service-discovery database (Corrosion) within the configured timeout on some nodes. These updates were repeatedly retried, slowing down or blocking the application of other updates. Since the Machines API (via Flaps) reads app metadata from global Corrosion, some newly created apps and recent updates were not visible everywhere, resulting in elevated app-not-found errors.

We mitigated the issue by adjusting the timeout and Corrosion’s retry behavior so that failed updates do not retry indefinitely and block other updates. We also rolled out a fix to chunk updates when a machine has a large number of services, preventing such updates from causing similar propagation issues in the future.


# August 4: Certificate issuance delays from Sidekiq overload (14:45UTC)

A surge of certificate-check and renewal work overwhelmed our background job system, causing significant delays issuing and renewing TLS certificates. This was caused by a few reasons, but the main one being an internal DNS resolver dedicated to the certificate renewal path: it is configured with no caching and aggressively recurses to upstream whenever possible in order to minimize delays in observing DNS updates for customers’ domains. That, however, was initially configured at a time where we were operating at a much smaller scale than we are today, and on this specific day, a spike of DNS renewal tasks combined with possibly slow lookups for some domains completely paralyzed the TLS certificate issuance queue. In the process of investigating this issue, we also discovered that the TLS certificate renewal task has been trying to renew certificates for deleted apps (with no success, of course), which did not help with the excessive queuing delays seen on this day.

We have put in several mitigations for this. Firstly, the DNS resolver used for TLS certificate renewal is now allowed to do some minimal caching to act as a buffer against temporary spikes in TLS renewal tasks and domain resolution failures. This does mean that updated DNS configuration could take a little while to propagate to our certificate issuer, but since this is on the order of seconds, it should not be noticeable under normal operation, and it provided enough protection for us to resolve the Sidekiq queuing issue. We also fixed several issues in our codebase that did not filter out deleted apps correctly, so that we are less likely to experience these spikes in the first place.

This is not the first time we’ve had Sidekiq-related incidents this year. We have plans to improve reliability around certificate issuance – our earlier work to move certificates out of Vault due to numerous incidents being an example – and another part of it is to decouple our certificate issuing pipeline from the rest of our API implementation, a Ruby on Rails app. The hope is that in this process we will make it both more scalable and easier to understand and debug.


# August 3: BGP misconfiguration dropped ingress traffic (14:57UTC)

A BGP misconfiguration while provisioning new edge capacity caused most traffic from Europe (and some US) endpoints to be dropped. The misconfiguration has been fixed and we have implemented safeguards against this kind of issue in the future.


# August 1: IAD MPGv2 Patroni etcd lease-expiration storm (16:40UTC)

MPG v2 in the IAD region saw a burst of Patroni health-check failures and timeouts when the region’s shared Patroni etcd cluster became overloaded. Existing databases were unaffected because Patroni entered failsafe mode, but new clusters could not be created. Investigation pointed to an intrinsic etcd problem: adding a new user bumps etcd’s Auth Revision, invalidating existing JWT tokens and forcing clients to reconnect. On reconnect, a thundering herd of clients overloaded etcd and brought it down. We mitigated by scaling up the Patroni etcd machines in IAD and lowering the bcrypt cost, and we’re exploring longer-term fixes.


# July 20: tkdb primary host outage caused token validation failures (07:07UTC)

The host running the primary node of tkdb, our token validation service, went offline due to maintenance and did not manage to reboot successfully due to a separate hardware issue. This caused token validation to time out, which in turn led to widespread failures in API operations that depend on authentication (Machines API, flyctl, etc.) and elevated 500s from platform services. API and dashboard functionality recovered over a few minutes once we restored the host to full service.

The root cause of this incident is, of course, the host failure, but a huge contributing factor is that tkdb is single-primary, and, unlike petsem (which powers app secrets and TLS certs), read replicas of tkdb do not function if the primary node is entirely offline. This is because, as a token validation service, the read replicas must be able to receive notifications about revocation to uphold security guarantees. We are currently looking into ways we can relax this requirement, keeping replicas up during primary outages, while avoiding impact on security properties provided by tokens.


# July 19: Missing app logs because we forgot a `systemctl restart` (14:32UTC)

Some customers reported that their app logs appeared to be missing. Initially we suspected that this was related to log ingestion (we use vector on each physical host), but further investigation revealed that… we forgot to restart the query service after adding a new storage node to the logging cluster, so it only queried the previous set of nodes. Because each app’s log stream is written to one storage node, apps whose logs ended up stored on the new node would return empty or incomplete results. Restarting the log query service restored normal log visibility.


# July 16: Vault throws one last tantrum on its way out the door (11:28UTC)

Recently we have been migrating our certificate storage away from Vault to our Petsem codebase, to reduce the number of distinct cluster-shaped things we need to think about. We reached the 100% mark on the rollout a few days before this incident, so this body of work was all but done.

This is, of course, when Vault decided to break in some manner. Since we were still validating things and hadn’t rolled out our code sans feature flags yet, there remained calls to Vault in both our GraphQL API and the Fly Proxy, particularly as a fallback for unknown certificates. Though these weren’t in the hot path of healthy requests, Vault hanging caused some issues across the services, and caused TLS handshakes to fail for otherwise functional Fly Apps for some number of minutes.

Unlike most writeups here, we don’t have (or, need) a firm root cause on exactly what broke and where. As soon as this kicked off we made the call to rip out the Vault functionality, since that was up next anyway. Without Vault in play, we got to blissfully ignore whatever went wrong with it. And we didn’t have much of a reason to dig into exactly which Vault/Petsem interaction in the proxy misbehaved, as the code that housed it was razed.


# July 14: SJC worker hosts locked up and rebooted (14:53UTC)

While debugging network packet loss in SJC, a tc qdisc (queuing discipline) configuration update was (unnecessarily) pushed out to all of SJC workers, which caused a number of them to lock up. We promptly started rebooting them and nursing them back to health.

Honestly, we are still not sure why this happened in the first place. The tc qdisc configuration was, after further validation, later very carefully deployed to all of our hosts in all regions, and none experienced a similar problem whatsoever. qdisc is not supposed to just randomly lock up entire systems: at most, it should cause network issues that we can easily recover from. So, it seems likely that we must have hit some weird kernel or NIC driver bug, but so far we do not know exactly which one. One thing is for sure though: we’ll have to be much more careful about even the most unassuming changes related to NICs.


# July 9: A couple of deployment issues (15:11UTC)

This day saw 2 separate instances of deployment-related issues centered around IAD happening around the same time.

The first one was related to Depot builders, where flyctl deploy would wait seemingly indefinitely for a Depot builder to become available. Initially, we believed that this was related to a single IAD host under immense I/O pressure, and was only affecting a subset of Depot builders with backing volumes on that host. Failed hosts with Depot volumes can sometimes prevent existing Depot builders from being reused until recreated via an explicit flyctl --recreate-builder flag. However, after we resolved the host issue (which took a considerable amount of time), it became clear that the Depot issue was not limited to the one host, but was impacting builds for any IAD builders. This is when an internal incident was declared and a status page posted.

It turns out that this was another case of capacity-related Depot issue. Because machines used for Depot builders are pretty large, when we’re under capacity constraints, Depot builders may not be able to successfully start with an existing volume, because that host has no spare capacity to run the builder machine. Usually, machines can auto-migrate on start when capacity error happens, but this is explicitly not enabled for machines and volumes as large as builders. Under these circumstances, no new Depot builder should land in IAD, but existing ones will continue to be attempted and fail to start. The fix here, before adding more capacity, was to purge existing Depot builders from IAD in order to have them recreated in other North American regions as needed.

Separately, our Docker registry, which runs as a Fly App managed by us, started throwing a lot of 5xx errors, while the Depot issue above was happening. This accounted for another bulk of deployment issues happening on this day, which unfortunately compounded with Depot-related issues above. Initially we thought the host DNS resolvers were acting up, since the registry app is throwing a lot of DNS-related errors. After ruling that out, we realized that even pinging the host (fdaa::3, where the DNS resolver also lives) from these machines was showing latency up to 100ms – definitely too much for basically loopback traffic! At first, we thought this is due to CPU starvation since the machines were all at 100% utilization. Scaling up just the CPU did bring down the latency, but it was still at several 10s of milliseconds. Some further investigation made us realize that these machines are running a lot of traffic over their TAP interfaces, and we do know that our TAP interfaces, for one reason or another, starts to struggle a bit under higher utilization (> 1Gbps). Creating more machines for the registry to spread the traffic out immediately solved the issue.

We came out of this incident with some plans for improvement:

  1. We should know about widespread deployment issues way sooner than we did on this day. We did not put up a statuspage because we assumed that the issue was related to a downed host, which already had its separate host issue, but it was in fact unrelated. We now have better monitoring for Depot issues related to machine and volume placement, and we will hopefully be paged much, much sooner should something like this happen again.

  2. Even when there is capacity pressure, ideally, Depot builders should still be available, just placed in different regions, for new or existing users. We are working with Depot to add support for placing builders using [geo region aliases](/content/docs/machines/guides-examples/machine-placement/#geographic-groups-and-aliases ""/index.html), which allows much more flexible machine placement, and we are planning to integrate this with our API and flyctl so that local capacity issues no longer automatically translate to much wider deployment failures.

  3. For registry, besides similar alerting, figuring out the performance bottleneck on our TAP interfaces is on our plan as well. Fixing this will improve experience across all machines, with the caveat that we’ll likely still place constraints to ensure fairness between machines on a host.


# July 3: Power supply failure in ORD (00:07UTC)

For redundancy, servers generally have two power supplies, connected to two independent power feeds: if one power supply fails, or one power feed goes down, the server can keep operating on the other power supply. During normal operation, the load is shared between both supplies.

Our servers are, of course, no exception. However, when one power feed of an ORD datacenter went down, we found our servers’ CPUs heavily throttled, to the point where they were unable to do any useful work. Later discussing with the provider, we found a misconfigured setting: the server would throttle the CPUs when one power supply failed or lost power.

Now, this isn’t a “oh duh, it should obviously not be like that” situation: this is a safety feature designed to not overload one power feed. Single-power feed operation, in some cases, can just be a fallback for the server to shut down safely without the ability to support normal tasks, even though this is not true in our case – we expect it to provide better uptime at full performance. The throttling setting makes sense as a safe default, and we should be explicitly opting out of the safety feature when we are sure single-power feed operation is safe in our case.

As a result of this incident, we are now working with our providers to ensure all of our servers’ settings are configured best for how our power feeds are set up.


# July 2: Certificate issuance outage (21:40UTC)

We issue certificates using Let’s Encrypt, who had a networking hiccup when failing over datacenters for maintenance. This caused some requests to fail, making our issuance jobs bail out and re-queue themselves. Apart from a handful of lucky hostnames who squeaked through, certificate issuance was out of action for a few hours. We renew certificates well before they expire, so no existing traffic was affected. Any apps that were in the process of setting up their hostname would have seen a delay in getting their first certificate. Once the upstream incident was resolved, all queued issuances and renewals were processed.


# July 1: Runaway cleanup job filled Redis (05:44UTC)

Our GraphQL API runs its background jobs on Sidekiq. When an app is deleted, we enqueue a cleanup job to tear down any associated resources, and that job in turn enqueues a job for every Machine the app ever had - even ones that had already been deleted.

One app had a very large Machine history. Its cleanup job loaded that entire history into memory at once to enqueue the per-Machine jobs, which pushed the worker’s memory high enough that the supervisor recycled it - killing the process and restarting the job from scratch on another worker. The loop repeated and quickly piled up tens of millions of per-Machine jobs. That filled the disk on the Redis instance backing Sidekiq; once Redis could no longer write its append-only file, job processing stalled and customers saw API errors for around 40 minutes. Already-running Machines were unaffected and kept serving traffic.

We mitigated by extending the Redis volume and restarting Redis, and shipped a change to make the cleanup job safer - it now iterates in smaller batches without pulling everything into memory and skips already-deleted Machines. Once Redis recovered, the backlog drained and error rates returned to normal.


# June 28: VictoriaMetrics ingestion delays and backlog (08:00UTC)

We had an extended period where our hosted metrics pipeline fell behind, causing Grafana dashboards and alerts (including fly-metrics.net) to show missing or delayed data. The underlying issue was uneven load distribution into our metrics ingestion “aggregator” layer, which led to CPU starvation on a subset of ingestion hosts and large metric queue backlogs that then took significant time to drain. We mitigated by rebalancing ingestion traffic (making the load balancer aware of backend load), tuning queue/throughput settings, and temporarily adding processing capacity to speed up backfill; metrics ingestion resumed and the remaining lag gradually cleared.


# June 25: Corrosion migration caused elevated CPU and routing failures in BOM and NRT (12:06UTC)

An ongoing Corrosion database migration on some hosts in BOM and NRT caused SQLite query planner statistics to become stale. As a result, the SQLite query planner chose full table scans for a particular fly-proxy query instead of using the appropriate indexes. This led to high CPU utilization and increased SQLite lock contention within fly-proxy. The resulting long-running read transactions prevented SQLite WAL truncation, further compounding the issue.

Users experienced HTTP 502 responses and connection resets for traffic routed through the affected regions, and requests to some Machine API endpoints timed out.

The issue was mitigated by stopping the migration and fly-proxy on the affected hosts, truncating the WAL and running ANALYZE on the affected tables to refresh SQLite’s query planner statistics. Once the statistics were refreshed, the query planner resumed using the correct indexes, and fly-proxy was restarted on the affected hosts. The migration script has also been updated to run ANALYZE after migrating each table to prevent this issue from recurring.


# June 23: Corrosion OOMs and wedged systemd (13:45UTC)

Corrosion, our internal state propagation system, was OOM-killed on a small number of hosts. This is usually not a big deal (save for the part where we need to figure out why it exceeded the generous memory limit we gave it), since the systemd unit is configured to restart the service automatically, and most other services on a host fail gracefully when Corrosion is down temporarily. This time though, the systemd units were stuck activating with seemingly no progress made, which triggered us to open an internal incident to investigate, and we also quickly increased Corrosion’s memory limit on all hosts to prevent any further OOM’s.

This turned out to be related to some fly-proxy changes we made recently. 2 facts that directly caused this:

  1. fly-proxy‘s systemd unit contains a After=corrosion.service;

  2. fly-proxy now waits for much longer on shutdown for unfinished connections.

When Corrosion got killed this time, it was in the middle of a fly-proxy rollout, which means that the fly-proxy unit was still deactivating on a lot of hosts due to the new shutdown logic. systemd enforces a reverse order based on After= constraints on shutdown, which means that before the last fly-proxy process stopped, a new corrosion.service cannot start successfully.

Since the proxy does not necessarily require a running Corrosion process nowadays thanks to our lazy-loading work, especially not on shutdown, the correct fix here is to simply remove that After= dependency, which we promptly did. The other loose end in this incident is why Corrosion OOM’d in the first place. Our conclusion here is that it is also related to the fly-proxy rollout: fly-proxy depends on something called Corrosion “updates”, which are generated in Corrosion and sent through a Corrosion-internal mpsc channel. The channel became blocked while the proxy restarted, which pushed Corrosion’s memory usage over limit, causing the OOM.

Because Corrosion updates are designed not to be fully consistent, and fly-proxy can deal with missed updates gracefully, we shipped changes to start dropping updates under this type of pressure. This should prevent this incident from repeating in the future.


# June 17: Singapore network outage (03:18UTC)

One of our upstream providers in Singapore region experienced an unexpected failure on a core router. This caused intermittent but complete loss of connectivity for a couple hours on June 18th; the issue reappeared on June 22nd, after which the affected device was replaced.


# June 15: Machines API Outage (15:07UTC)

Our token service [tkdb](/content/blog/operationalizing-macaroons/ ""/index.html) became unavailable, which broke our macaroon token verification path. This caused a broad control-plane outage for approximately an hour: dashboard login/SSO and many Machines API operations failed, while existing app traffic continued to route normally.

On this day, tkdb needed to be migrated between two internal hosts. Many Machines API operations rely on tkdb for macaroon minting and verification, and tkdb is itself a Fly App that relies on the Machines API.

Migrating tkdb is something we have a known, written procedure for. The general idea is to:

  1. [Cordon](/content/docs/machines/api/machines-resource/#route-requests-away-from-or-back-to-a-machine ""/index.html) the tkdb primary

  2. Stand up a new Machine, let it catch up as a replica, then promote it to primary

  3. Point all replicas at the new primary

During this window, minting new macaroons will be unavailable, but verifications will still work and the Machines API will remain functional.

A while back, a fly-proxy change inadvertently shifted the semantics of Machine cordoning. The new cordoning flow meant that cordoned Machines would not ever be load balanced to, but they were still eligible for requests that specifically name them. That is, requests using [fly-replay](/content/docs/networking/dynamic-request-routing/#how-fly-replay-works ""/index.html) to an instance, or the [fly-force-instance-id](/content/docs/networking/dynamic-request-routing/#the-fly-force-instance-id-header ""/index.html) header, would now each arrive at a cordoned Machine.

Because of that, when we cordoned the primary tkdb Machine, it would still receive replays from replicas. Additionally, stopping the primary wouldn’t work, as these same replays would auto-start the targeted Machine. The right decision here would have been to restore and uncordon this primary, and pause the migration. The less-right decision, which is what we did, was to destroy the old primary before promoting the new one. In theory this is roughly equivalent, and is a safe way to migrate the broadly similar petsem cluster, which is why it seemed reasonable.

The (overlooked) quirk for tkdb is that in order to guarantee a bounded lag on token revocation, replicas that fall too far behind the primary will throw their toys and proceed to replay all requests to the primary, rather than use a stale revocation list. In practice this meant that we were suddenly left with a cluster of replicas that would not function, which took out the Machines API for customers and for ourselves. Without the Machines API, we were unable to easily bring up the new primary Machine required to resolve this.

From here we had to resort to spinning up the required Machine by hand with flyd on a host, which has its own set of internal hurdles to overcome. The main responder in this incident was based in Europe, which additionally meant that their Machines API operations would prefer flaps, and thus tkdb, in Europe rather than the new primary in North America. These factors weren’t a problem in the initial incident, but they did complicate and slow the response for getting everything back up and running.

The technical root cause here is the cordoning flaw in fly-proxy, and the organizational root cause is that our documentation and runbooks for tkdb were not adequate to prevent this from happening. Both of these things are being fixed: In fly-proxy, cordoning will be fixed to prevent all proxy-routed traffic, and our internal documentation is being reviewed and improved for tkdb, and more broadly where other services have the same gap.


# June 12: Elevated Sprites error rates in SIN (02:27UTC)

Some Sprites API and dashboard requests hitting a SIN edge returned 500s for a subset of orgs with Sprites in the SIN Region. This was caused by a mismatch in code versions running in SIN and SYD regions, as well as the specific way the API handles some organization requests.

In addition to serving general API requests, each API node acts as the ‘manager’ for a subset of organizations in its region. Certain organization request types for some actions will be replayed through the network to that org’s manager node. The ultimate source of truth for this management data lives in object storage. If an API node fails, another node in the region steps in to manage the Sprites the dead node was responsible for. If both nodes in a region fail, API nodes in the next closest region will take up management. So on, so forth.

Earlier in the day one of the two API nodes in SIN failed due to corruption on its local cache volume. The failover process worked as designed, with the other SIN node taking over management of all orgs on the dead node. We brought up a replacement node in SIN and that started re-syncing with the cluster.

While the replacement node was still syncing, the second healthy SIN node restarted. Since it was the only healthy node in the region, this triggered the next closest region (SYD) to take over management of all SIN orgs. Again this succeeded without issue.

However when the two SIN nodes came up, they were unable to replay requests to the nodes in SYD. After some investigation we identified an earlier update had failed to update the SIN nodes to the latest code version. This caused an incompatibility, with the shape of request the SIN nodes were sending not matching what the SYD nodes were expecting. After identifying the issue we pushed out a manual version update to the new SIN nodes and normal operation resumed.

This only impacted requests hitting a SIN region API node, for Sprites that had their management re-homed to the SYD nodes. Requests hitting any other region’s edge for those same Sprites continued working as expected.


# June 11: Everyone lives on NULL island (06:03UTC)

Newly created WireGuard peers and Depot builders were placed in AMS (Amsterdam) regardless of where the user actually was. We had removed some datacenter coordinate metadata from Consul, believing it to be unused – but part of our control plane still sourced region coordinates from it, and without that data it located every region at (0,0). The logic picks the region nearest the user, but with distances all tied at zero it fell through to the alphabetically first region, AMS. Existing peers and builders were unaffected; only newly placed ones landed in the wrong region.

We mitigated this by sourcing region coordinates from our database instead of Consul, then cleaned up the Depot builders that had been pinned to AMS. We’ve also hardened the placement logic to fail loudly if coordinates go missing in the future.

June 11: ORD transit loss destabilized Managed Postgres clusters

# June 11: ORD transit loss destabilized Managed Postgres clusters (01:18UTC)

We hit a period of severe packet loss and intermittent connectivity on network transit paths into/out of ORD, which made some Managed Postgres nodes unable to reliably reach Patroni’s DCS (Kubernetes). That triggered leadership churn (primaries demoting / failovers) and left some replicas unable to participate cleanly, causing intermittent connection errors for Managed Postgres clusters in ORD. Service stabilized once reachability improved, and we also rescheduled affected replicas away from the worst-impacted hosts to keep clusters steady during ongoing network flaps.


# June 10: Two GRU edges OOM (19:35UTC)

Two edge nodes in the GRU region couldn’t keep up with exporting their metrics, and the backlog caused the hosts to run out of memory. The immediate impact lasted about 17 minutes, at which point we rebooted the affected nodes and returned them to the routing pool. Some traffic was disrupted, though this was not a complete outage in the region. We’ve tidied up our metrics pipeline a little since, paring back some high-cardinality metrics that contributed to the heavy load.


# June 4: Stale 6PN mappings wreaking havoc (09:18UTC)

This is another case where, as we were working towards improving the platform, we ended up with multiple ways of doing one thing, some of which are considered legacy and should eventually be removed, but the removal was never completed. An unexpected interaction between the old and new systems then wreaked havoc.

In this case, the system in question is [6PN](/content/docs/networking/private-networking/ ""/index.html), our private network powered by Wireguard that connects all of your Machines. When this system was designed, each Machine’s private 6PN address was bound to the host where it was created. This made routing simple to implement, but also started to cause issues when we migrated Machines between different hosts. The reason is that some apps depended on a static 6PN address per Machine: even our own legacy unmanaged Postgres offering depended on it, despite the fact that these addresses were never meant to be stable.

At some point, we finally decided that this is not sustainable and we should, instead, meet the expectation of a majority of apps: that is, to keep 6PN addresses stable. The first iteration of this work was a simple DNAT, where machines still get new 6PN addresses, but an eBPF program rewrites packets targeting a machine’s old 6PN address(es) to the new one. The price we pay is that, because technically the 6PN address still changes, we need to keep track of every single 6PN address a Machine has ever had. This is all stored in Corrosion, which bloated its storage, not to mention the map we needed to synchronize into the eBPF program.

This was changed roughly a year ago. Instead of keeping track of all old 6PN addresses, we simply made it so that Machines do not get reassigned a new 6PN on a new host if it is migrated. All Machines created after this change retain their initial 6PN after migration. Of course, this alone would break routing, because that depended on a per-host fixed 6PN prefix. Some sort of NAT is still needed, but now we only need to keep track of a Machine’s current host and its initial 6PN.

…which brings us to today. We have two types of “stable 6PN” Machines: some before the change above, and some after. The intention was that when an old-style 6PN Machine gets migrated, it will become a new-style stable 6PN Machine with all the new-style plumbing. We’d delete unneeded Corrosion entries in this case and slowly drain them away as they’re moved around. At some point this year, we realized that the Corrosion subscriptions used for old-style 6PN DNAT were creating a lot of load on Corrosion. As a result, we shipped a change to only apply 6PN DNAT rules once when the service responsible for this is started, since we did not except any new Machines to be created with old-style 6PN anymore. However, there was an oversight: in some cases, the existence of old-style DNAT rules actually overrides new-style stable 6PN’s rewriting logic. So, when a Machine gets migrated to use new-style stable 6PN, it is possible that some peers might still be rewriting its address to a host-specific address that no longer exists.

This exact scenario started happening first for our multi-tenant Consul clusters (used for unmanaged Postgres and LiteFS), and then for some customer Machines as a spike of rebalancing migrations happened for various reasons. A considerable amount of time was spent on triaging the issue because it was an unexpected failure mode. We did not expect that old-style stable 6PN would interact with new stable 6PN in this way, especially not several weeks after the last round of changes were deployed.

We mitigated this problem by adding code to delete old-style 6PN DNAT entries when new-style stable 6PN rules are set up. This, unfortunately, briefly caused another bug where the daemon responsible for this became too slow to catch up with Corrosion (we need Corrosion to decide whether a Machine has been migrated and thus needs rules for stable 6PN), which caused issues with Managed Postgres in LAX for a little while. This was then fixed up by making the cleanup code opportunistic and non-blocking for the main processing path.

We see a few directions as the next steps to preventing this from happening again:

  1. Old-style 6PN DNAT mappings should really not exist anymore. We need to migrate all the remaining machines that still use it to new-style stable 6PN addresses.

  2. The reason why the Corrosion subscription and its processing code became slow was partially due to the query’s inefficiency; we’re working on addressing that too.

  3. We need a way to gracefully recover from such an event; the daemon should not just miss updates.

  4. Finally, we should be able to “fill in” missing stable 6PN NAT rules even if the Corrosion subscription happened to miss some updates. The subscription can still be used for updates, but not as the single point of failure.


# May 30: Deploys blocked by billing error (02:18UTC)

For a few hours, deploys for some organizations were failing with a “We require your billing information” error, despite having just added payment methods or credits to their organizations. This was due to a mis-ordered deployment of a new Corrosion schema.

For some context: organization information is managed by our central GraphQL API backed by a local database in iad; when an organization is updated, for instance when the billing information is updated, the GraphQL API pushes the changes to the global Corrosion cluster so it can be read by the Machines API. When new information needs to be stored in Corrosion, we need to deploy two changes: a global change to the Corrosion (sqlite) schema, and a change to the GraphQL API to push the new data to the global cluster.

Earlier in the day, we had prepared a change to push some new organization data to Corrosion. This is usually a safe change, however this time the GraphQL API was deployed prior to the global schema being updated. This caused all organization updates to fail to be propagated to Corrosion, thus causing the Machines API to not know about the updated billing status of organizations. To resolve this incident, we quickly reverted the change to the GraphQL API and backfilled the missing data in Corrosion.

We are looking into ways to alert on repeated sync failures, as well as failing GraphQL API deployments if the Corrosion schema is out of date.


# May 28: West coast edge proxies overloaded (21:08UTC)

This incident requires some background which will become important later:

The incident started with us noticing flappiness in our US west coast regions, primarily in SJC at the beginning. Our logs and metrics indicated that the lazy loader latency was high, on the order of 500 ms to several seconds. This means that many new requests will need to wait that long or even longer to be served. On the other hand, proxy’s CPU usage was not especially high, and neither was the inbound connection rate. We’ve seen this kind of issue before: it usually is indicative of inefficient sqlite queries, certain apps with excessively large state stored in Corrosion, or general host performance issues. At this point, we happened to have spotted one app with extremely large state in Corrosion, and quickly “concluded” that it must be contributing to the issue, so we put in a temporary mitigation and deployed the proxy in SJC.

It momentarily seemed to improve the situation, but latency quickly shot through the roof again after the new proxy processes warmed up. We began doubting whether it is inefficient sqlite queries, which we ruled out, or whether there was lock contention simply due to our recent growth resulting in increased connection rates. This is also the point where we noticed Airtime reporting increased bandwidth in SJC, but it was below what we have concluded before was the ceiling of what a single edge server could handle. In either case, our edge capacity in SJC was also underprovisioned due to a couple of servers being out of production, so we decided to first shift Anycast traffic to LAX and see if it handles the load better.

Again, initially it seemed to help, but after a while LAX started struggling as well (side note: at certain points we also attempted to shift traffic out of the west coast entirely, which was why edges in other regions may have been momentarily affected). We finally decided to adjust down the bandwidth limit of Airtime, even though we were pretty sure our edges could take the level of traffic seen throughout this incident. It did bring softirq CPU usage and host load average down, but the proxy was still struggling with slow lazy loader queries. We bounced the proxy, which seemed to clear up the lazy loader issues as well. This marks the end of the first acute phase of this incident.

It would have been nice if this was the actual end of the incident. It was not, and it was primarily due to 2 other issues:

  1. Airtime, the system we used to limit impact of traffic spikes, works entirely within one single process and does not propagate its knowledge outside. This would not have been a problem (we initiated a hard-kill of all pending-shutdown proxy processes when we bounced them), if not for:

  2. Due to a bug with how our proxy deployment script interacts with systemd, we have somehow left multiple instances of the proxy running indefinitely on some of the affected nodes (TLDR: systemctl kill does not actually transition a unit to a stopped state; combined with Restart=always it simply causes the process to restart);

The combination of these two means that any limit we set in Airtime could, at any point, become effectively doubled if some heavy connections landed on a different proxy instance, causing the same issue to repeat after the initial phase was resolved. It is also worth noting that the fact that we needed to bounce proxy processes after tuning Airtime is itself contributing to this issue: that revealed that there are issues with queuing behavior around the lazy loader. Specifically, it seems that it is possible to end up with effectively infinite queues waiting on the sqlite connections when lazy loader itself is slow (due to softirq contending with userspace for CPU under high load, for example), which will not resolve unless the process itself is bounced (and in turn, that revealed the other issues causing recurrence of the incident).

In summary, this incident was caused by a combination of factors:

  1. Our edge capacity is underprovisioned in some regions; they have not caught up with our recent growth in user base.

  2. Airtime’s tuning no longer matches reality, either due to a shift in traffic patterns or other non-bandwidth scaling issues in the proxy.

  3. A bug caused multiple active proxy instances to coexist without code to handle shared state.

  4. The lazy loader exhibits runaway queuing behavior at high load.

We’re working hard to address each and every one of these issues. As a starter, we are going to provision significantly more edge capacity in the coming weeks/months. We have addressed the bug that caused multiple proxy instances to coexist, and changed Airtime so that, for now, it applies a much stricter limit when it is not the expected active proxy instance. We have fixed load-shedding behavior in the lazy loader so that there is a more reasonable upper bound on the maximum latency serving requests. Other work is currently under way:

  1. We believe that the reason why proxy seems to run into lazy loader-related performance issues much earlier now, compared to before, is due to our single coarse-grained lock on the proxy’s in-memory state is no longer scaling well as we grow. We have observed high queuing delays not in sqlite queries, but simply in trying to insert data into the in-memory service catalog. We’re planning to shard the catalog and move to finer-grained locking, assisted with testing such as Antithesis to ensure migration to this does not cause more outages.

  2. We are going to rework Airtime so that it reacts better to overall system load instead of just the proxy. This will hopefully serve as a backstop when we somehow end up with multiple proxy processes running, or when any non-proxy processes on the same host consume any of the bandwidth headroom.

  3. We’re looking into better monitoring for when the proxy is not under its expected configuration.

May 28: DNS cache was broken for CNAME'd domains

# May 28: DNS cache was broken for CNAME'd domains (09:59UTC)

Some customers saw persistent DNS resolution failures for certain external hostnames that only cleared when we restarted corro-dns, our recursive DNS resolver. It turns out that the domains they were trying to resolve had intermittent failures upstream. The weird thing is that by itself should not cause persistent problems: even though corro-dns does cache DNS responses, it only caches failures for a very brief moment and will retry pretty quickly if one resolution failed. The cache should eventually be populated with a valid response, and if more upstream errors happen, corro-dns is allowed to serve an expired cache in that case.

It turns out that this cache logic failed to take into account cases where a domain A is CNAME‘d onto domain B, and only domain B failed to resolve. In that case, corro-dns ended up with a cached CNAME entry for A -> B, but without any corresponding entry for B. A subsequent request for domain A will hit the cache for the CNAME, but corro-dns will not spawn a new query for domain B since it thinks we’ve already hit the cache. It then returns only the CNAME record to the client, and most clients will not spawn another query either and will just report to the user that no A or AAAA records are returned. This situation will not clear itself until the TTL of the CNAME record expires, which in this case was very long.

We mitigated this issue for now by skipping cache when any unexpected failure happens while resolving a domain. The root cause, however, is that corro-dns caches full DNS responses and not individual DNS records, and does not “fill in” additional records when only a CNAME can be cached. Our plan is to refactor this layer of caching to prevent similar bugs in the future.


# May 27: App creation timeouts from petsem-certs disk full (03:43UTC)

New app creation requests (including flyctl apps create) were failing with 504 timeouts because they weren’t able to create certificates in petsem-certs, our new certificate store (which we’re in the process of provisioning and migrating to from Vault). While petsem-certs is not yet operational, we are writing certificates to the store, which (disappointingly) caused it to run out of disk space. Existing apps and Machines were unaffected - only app creation timed out. We restored app creation by expanding the storage.

Since we’re dual-writing to both petsem-certs and Vault, and the Fly Proxy is reading from Vault, we didn’t expect the loss of petsem-certs to have any impact and it hadn’t yet been hooked up to our monitoring. We also had excessive retries on requests to the store, which caused the issue to present as a timeout rather than a failure, so we’ve tuned that as well.


# May 20: SYD egress IP networking broken on new workers (04:46UTC)

Some newly provisioned hosts in our Sydney (SYD) region failed to be configured properly for [egress IP](/content/docs/networking/egress-ips/ ""/index.html) connectivity. As a result, a number of Machines using egress IPs in the region were unable to access the network. During the incident, we immediately migrated the affected Machines to known-good hosts.

Recently, we moved configuration for some infra components (including the VXLAN interface backing egress IPs) to a new, more scalable system. The rollout appeared to be successful, but an interaction with a legacy deployment method caused the configuration service to not be restarted correctly - so VXLAN worked on existing hosts, but would not be provisioned on new hosts.

Our egress IP monitoring was set up in a world where egress IPs were machine-scoped rather than app-scoped (see this forum post for more context). As such, a couple monitoring Machines were set up in each region, and not every host was being monitored - as that would require one IP address for every host. After this incident, we ported egress IP monitoring to app-scoped IPs with a Machine running on every host.


# May 19: Fly dashboard outage from broken GraphQL API deploy (23:13UTC)

A deploy of Fly’s GraphQL API service (which backs much of flyctl and parts of the dashboard) reduced the number of healthy instances enough that the HAProxy layer in front of it began timing out its /status health checks and returning fast 503s, which showed up as intermittent dashboard/API failures and elevated deploy failure rates. We recovered by removing the broken instances and cloning a known-good Machine to restore capacity until HAProxy backends were stable and green again.

Afterward we found that a local fly deploy of the GraphQL API service could produce a broken image because it skipped a CI-only step that fetches a supporting binary (resulting in an empty placeholder being copied into the image), and we merged a change to prevent that failure mode.

May 19: IAD observability disrupted by NATS misconfiguration

# May 19: IAD observability disrupted by NATS misconfiguration (19:11UTC)

A server being decommissioned began advertising a bad NATS config (specifically, an empty connection URL), which caused the logs/metrics exporters of various hosts in the IAD region to crash. During the incident, many Machines in IAD had missing metrics, and some customers may have also seen gaps or delays in log delivery. We mitigated by removing the decommissioned server from the NATS cluster and restarting the affected metrics exporters across IAD to restore normal telemetry flow, and are discussing options to remove the reliance on NATS from the logs/metrics pipelines.

May 19: Proxy and Corrosion in SIN weren’t on the same page

# May 19: Proxy and Corrosion in SIN weren’t on the same page (11:13UTC)

During a rollout of the Fly Proxy, a new Corrosion query started throwing errors on a subset of hosts in Singapore. This query relied on a new column in our Corrosion schema, which had been rolled out globally the day prior. It turns out these hosts had received the new schema but hadn’t successfully reloaded it.

Once the new proxy came up, it failed to load apps from Corrosion and couldn’t serve any traffic. This made machines on these hosts unavailable, and caused a wave of Managed Postgres (MPG) healthcheck failures in the region.

During the incident this was fixed by forcing a reload of the Corrosion schema on these hosts, after which traffic returned to normal and all MPG cluster alerts resolved.

We made two changes to prevent this happening in the future. First, we didn’t notice this during the schema rollout as Corrosion didn’t return an error for a failed reload. Corrosion now returns an error code when this happens, so we can revisit those hosts after a rollout. Second, this is the sort of thing we should catch in the proxy’s bluegreen deployment. This error wasn’t hit until after the proxy marked itself healthy, though, so it had already taken over as the primary. Now the proxy prepares all SQL queries against Corrosion during its startup sequence, so the new proxy won’t successfully come up if any of these fail.


# May 16: FRA managed Postgres control-plane outage (12:23UTC)

Managed Postgres clusters in the FRA region became intermittently unreachable after the regional Kubernetes control plane (FKS) got overloaded/stuck and the Kubernetes API began timing out. Because Patroni uses Kubernetes for coordination in this setup, those API failures prevented clusters from reliably determining primary/replica state, causing widespread connection failures and flapping health. We recovered by reducing resource pressure, defragmenting the affected etcd instances, restoring the control plane’s ability to reconcile, and then repairing clusters one-by-one.


# May 15: Bad TLS cert update broke Consul (16:12UTC)

A configuration automation run accidentally overwrote Consul TLS certificates with invalid ones, which caused Consul lookups to fail across parts of the fleet. We have spent a lot of time in the past couple of years to remove Consul as a key dependency, and as such most aspects of our platform were not directly impacted: all running Machines were unaffected and fly-proxy routing remained functional. The impact was concentrated on the few control-plane operations that still rely on Consul: mainly fly ssh console and OIDC tokens. We restored the correct certificates and restarted Consul agents to pick up the fixed TLS configuration, after which errors and alerts subsided.


# May 12: Usage ingestion blocked by stuck Oban jobs (16:48UTC)

Usage ingestion fell behind and then stopped making progress when several background jobs hung indefinitely, eventually consuming all available worker concurrency for the ingestion queue. This was triggered by a bug in an Elixir decimal dependency where converting certain values (like 0.0) to an integer could loop forever, causing specific volume-usage receipt processing jobs to never finish. We fixed this by updating the dependency to a version containing the upstream bugfix, after which ingestion resumed and the backlog drained; once the queue cleared, the delayed ~18 hours of usage data was backfilled and reflected normally.


# May 11: Kernel upgrade caused machine `stdout` to become wedged by Cloud Hypervisor (14:30UTC)

This is probably one of the more interesting / confusing / complicated bugs we have had since the revival of Infra Log. It began on this day with us receiving reports about an outage of the Upstash Redis extension, reflected on their status page as well. Upstash Redis, when used as a Fly extension, runs on Fly Machines, just like other customers. The only difference is that, for various reasons, we run their Machines using Cloud Hypervisor rather than Firecracker. This has never caused problems before, and initially we were pretty certain this is an issue on Upstash side. As we worked with them to investigate, though, we got something confusing: the report that these Machines are stuck writing to stdout!

The stdout of a process running inside the machine is a pipe, one end acting as stdout and the other end connected to init, where it splice()s the pipe into a vsock connected to the hypervisor. A process on the host, called firefly, collects these logs from vsocks. For the stdout (write) end of the pipe to be stuck, one of these steps must have gone wrong: either firefly on host is failing to collect logs fast enough, or init inside the Machine is failing to splice() them into the vsock.

We were quickly able to rule out an issue with firefly. This means something went wrong when copying logs into the vsock, but that is extremely weird: the init process does use Rust async with Tokio, but at the end of the day, it is just repeatedly calling splice() to move pages around. We suspected all sorts of wakeup issues on the Rust side, because the way splice() works does not play nicely with Rust async and we had to use another pair of internal pipes as a workaround (we’ll not get into the details here, but the TL;DR is one cannot tell which side of the splice() caused an EWOULDBLOCK or EAGAIN; see something like tokio-splice2 if you are curious).

That was not what was happening. The Rust side was behaving just fine (* for some definition of “fine”, but no weird wakeup issues). Instead, the vsock itself seemed to be getting stuck. The only change(s) recently in this path was a batch of kernel upgrades due to recent vulnerabilities. We found this issue on Cloud Hypervisor’s GitHub repo, which points to a kernel change incompatible with older Cloud Hypervisor builds. The change was also backported to LTS kernels, which we upgraded recently, but without the corresponding Cloud Hypervisor fixes. Because the majority of Machines do not use Cloud Hypervisor (except GPU machines, and specific organizations), this did not show up in our testing. And because the change only affects large transfers, it only really triggers when a large write is performed on stdout, and by extension, the vsock. That is also a gap in our testing: not many apps emit logs as often as some of our biggest customers.

We ended up resolving the incident by upgrading Cloud Hypervisor to a version with the fixes. To mitigate similar issues in the future, we also introduced safeguards in init that will not allow a vsock or host-side issue to block the stdout pipe indefinitely. Unfortunately, the splice() + Rust async issue combined with the lack of a way to determine a Linux pipe’s remaining capacity reliably means we had to resort to a simple timeout on the vsock write side for this, which will still introduce some latency to stdout when the vsock side is wedged. It will not be an infinite delay, though, and it will also emit a corresponding warning for visibility when a timeout happens.


# May 6: Machines API hitting failed hosts in SIN (00:01UTC)

Some ListMachines API calls returned 500s—primarily for apps/orgs that had Machines in the sin region. This was caused by these API queries hitting failed hosts in SIN. For some context, when an organization has machines located in regions faraway from the one that handled the Machines API request, we replay / forward that request to their respective regions for more up-to-date information. At the time of the incident, one host in SIN appeared up (accepting TCP connections) but responded every connection attempt with a reset. fly-proxy, our load-balancer component, had an independent bug that prevented it from treating these requests as retryable. Cordoning that host mitigated the incident, and the fly-proxy bug has also been fixed since.


# May 5: Petsem primary host lost networking (IAD) (13:13UTC)

A worker host running the primary instance of Petsem, our secrets storage service, lost network connectivity after a NIC driver stall. We restored service about 30 minutes after the start of impact by reloading the host’s NIC driver.

During the incident, requests to set/update secrets and create apps failed globally. Some other platform functionality was also affected because an internal Redis used for rate limiting was on the same host. However, existing apps/machines continued running, and existing secrets continued being accessible from our read replicas.


# April 28: Machines API bug caused `fly deploy` to create duplicate Machines (23:41UTC)

On April 28th, some fly deploy runs incorrectly received an empty Machines list from the Machines API for apps that had already been deployed. When that happened, flyctl created new Machines instead of updating the existing ones, resulting in duplicates.

This was tracked down to new affinity behavior in flaps (our Machines API service). This will likely make its way to a Fresh Produce near you sometime soon, but the abstract is that after some operations, API requests are replayed to the same flaps instance for a brief duration (while state propagates through Corrosion, our distributed database).

Another place flaps uses fly-replay is when fanning out to list Machines from multiple regions, where it used the replay itself as a signal to strictly return Machines local to its region. When this received a new affinity replay, it returned the response for a fanout replay instead (that is, it did not list Machines from other regions). So, if all your Machines were in yyz, but you had affinity with flaps in ord, flyctl would be given an empty list of Machines.

At 00:28 UTC on April 29th, we mitigated the issue by disabling app affinity in the API. Since then the bug has been fixed, but duplicate Machines created during the incident will persist until removed manually. Customers who ran fly deploy on flyctl v0.4.41 or v0.4.42 between April 28th and April 29th can run fly scale show to review their apps’ Machine counts. If more Machines appear than are wanted, they can be removed individually with fly machines destroy <id> --force.


# April 27: MPG provisioning failures from revoked org token (17:10UTC)

A revoked token caused new Managed Postgres clusters to fail when bootstrapping the default database and user during a 6-hour window. We mitigated the problem by rotating the token and redeploying the secret, and added alerts to detect these types of failures more quickly in the future.


# April 23: GitHub integration management callbacks returned 500s on Fly.io secondary nodes (14:59UTC)

Some dashboard flows for adding/updating GitHub integrations intermittently returned 500 errors (notably the GitHub app callback endpoint after opening “Manage GitHub Integration” link) when those requests were served by a secondary instances (no database connections, used for static routes). Existing GitHub integrations and deployment activity weren’t affected, but users trying to manage integrations could hit errors (sometimes succeeding after a refresh). We deployed a fix to ensure these callback paths are handled correctly regardless of which instance receives the request, and followed up by patching a few related endpoints found via Sentry.

April 23: Extension provider polling overloaded Postgres

# April 23: Extension provider polling overloaded Postgres (11:12UTC)

An extension provider increased how often they polled one of our private API endpoints. The endpoint ran an expensive Postgres query, and at the higher rate it saturated CPU on the database backing our dashboard and GraphQL API. This caused intermittent 500s on the dashboard and GraphQL API endpoints for about 40 minutes. The provider reverted the polling frequency change and traffic dropped back to normal.

The query was checking whether an organization had a registered extension with the provider, but it was scanning far more rows than it needed to. We rewrote it to short-circuit on the first match. We are also adding rate limiting on this endpoint to stop a similar spike from saturating the database again.


# April 20: Duplicate Wireguard Mesh IPs Wreaking Havoc (14:53UTC)

Some background: at Fly.io, we run a fleet of bare metal servers hosting your workload, be it Machines or Sprites, all connected over a Wireguard mesh. When we provision new servers, something has to set up Wireguard such that it is reachable by the rest of the fleet. This is done by something we call flywire. It generates a Wireguard public / private key pair, sends the public part to our Consul cluster to be read by other nodes, and picks a IP in a private /8 range.

You have probably read about our last incident where some nodes just lost connectivity over this Wireguard mesh. This incident began as what looked like a recurrence of that one: some nodes not being able to talk to others. Only that this time, resetting the wg0 interface did not do anything to fix the issue. On one of the affected edge servers, we also noticed that NATS (used to propagate app load information, logs, etc.) is using an abnormally high amount of CPU. This actually gave us some clue, since its logs kept complaining about some of its peers do not report the expected regions (they should be in sin but report as fra, for example).

We went to check on those nodes in fra as well. Turns out, they have the same IPs as the problematic nodes in sin! In fact, after a quick sweep of our entire fleet using a script, we found a couple more pairs / triples of servers with this exact same problem. They were all provisioned recently, and we were also lucky that many of them were not yet set up to accept new Machines. Duplicate IPs are problematic, because other nodes may end up selecting one but not the other as the “active” peer, causing partial connectivity. Most of our platform components also assume that Wireguard IPs are unique. We quickly took all of them out of production to investigate.

It turned out that there were two bugs in the provisioning process that caused this:

  1. When generating new Wireguard key pairs and IPs, we acquire a Consul lock on the respective resources, but the duration of the lock only covers generating the IP. We do check duplicate IPs at this stage, but by the time we write the IP into Consul, the lock would have been released already. Any parallel writers could cause a classic TOCTOU condition.

  2. In some cases though, nodes all get the first IP available in the /8 range. That is too unlikely to be explained away by pure chance. Rather, the bug here is that our code to generate the next IP by checking consul ignored errors emitted by Consul and just defaulted to the first IP in that case.

As an immediate measure, we have reset the wg0 IP addresses of all these servers and added an alert when we detect duplicates. We are also going to fix the two bugs in our provisioning script to avoid this in the future.


# April 17: Vault outage broke TLS certificate lookups (13:04UTC)

A failure in our certificate store (Vault) caused fly-proxy to time out or fail while resolving some TLS certificates, leading to intermittent TLS handshake errors for affected apps. The same issue also affected a subset of MPG clusters.

The issue started after a migration left the Vault cluster in a bad state, where the Raft leader was stopped before transferring leadership to a different node. As Vault (and other Raft-based services such as Consul) need to load the entirety of the database into memory at process start, a cold boot of the leader would have taken hours; so we restored service by rebuilding the cluster.

This is not the first time Vault has caused fly-proxy outages. Longer term, we have plans to migrate to something more resilient to this specific failure mode.


# April 14: WireGuard wg0 one-way host connectivity (10:51UTC)

Over the past few weeks, we observed individual pairs of hosts fail to send traffic over our global WireGuard mesh. Specifically, the tunnel between the hosts would appear up and handshaking correctly, but packets would only flow in one direction (or sometimes none at all). Neither WireGuard configuration nor firewall rules were able to explain this behaviour. This caused a few issues: fly-proxy on affected hosts wouldn’t be able to talk to each other, breaking load balancing in some cases; a more severe problem is with static egress IPs, since the return path depends on edge nodes being able to forward packets back to workers – if one edge node happens to lose connectivity with a worker in this way, some packets might be silently dropped depending on which node upstream flow hashing decides to forward the packets to.

Eventually, we tracked this issue down to a regression in the 5.15 stable kernel tree. We attempted to resolve this problem by removing and re-add the peer, but that caused Netlink in the kernel to hang, as described in this LKML thread. Fortunately, we later realized that even though resetting one single peer would hang, restarting the entire Wireguard interface (by downing the interface and re-initializing it) does not. This causes much less disruption to customer workloads on affected hosts, and we quickly fixed up all that we could find.

To close up the incident, we added an alert for any WireGuard peers stuck in this way, and scheduled a kernel upgrade to a later version in the future.


# April 12: High edge CPU usage resulting in high latency in ORD (18:43UTC)

ORD edge nodes became CPU-saturated, which made traffic entering through the ORD region intermittently slow (and in some cases time out). Profiling on the affected edges showed fly-proxy spending an unexpectedly large amount of time in pthread_mutex_{lock,unlock} calls. This is weird, because fly-proxy itself does not, in fact, use pthread mutexes – it uses locks from the parking_lot crate, which is based on futex system calls directly. Eventually, we traced the likely cause of this lock contention to SQLite, which is used to directly access Corrosion’s local database to load app metadata used for routing. We reduced the number of SQLite connections fly-proxy opens on the ORD edges, which immediately dropped CPU usage and brought lookup latencies back to normal.

We are, however, still unsure exactly which lock in SQLite caused the contention: initially, we suspected the per-connection lock SQLite uses to prevent concurrent access, but our Rust side code (based on rusqlite) has explicitly marked connections as !Sync and therefore they are never shared between threads in the first place. Our current hypothesis is that this is due to rusqlite’s use of the flag SQLITE_ENABLE_MEMORY_MANAGEMENT, which puts a mutex on SQLite’s per-process page cache. However, we are still unable to definitely confirm that this is the case due to the lack of stack traces through SQLite during the incident (and that we have not managed to reproduce the issue at all). We have enabled more instrumentation in our code, which will hopefully give us more complete stack trace profiles should this happen again.


# April 10: NRT Machines API errors thrashed Managed Postgres (18:37UTC)

Managed Postgres clusters in NRT intermittently went unavailable when the regional Machines API began timing out and occasionally returning truncated 502 responses, which caused Kubernetes (via our virtual-kubelet) and the Postgres operator to repeatedly reschedule and recreate Machines. That feedback loop produced extra operator/pgbouncer Machines and kept some pods stuck “not initialized,” breaking routing for affected clusters for several minutes at a time. We stabilized the region by shifting load off unhealthy workers, bringing additional NRT capacity online, and restarting the affected control components, after which cluster health checks returned to normal.


# April 8: SYD host I/O saturation (23:14UTC)

Some workers in our SYD region became saturated on disk I/O, which in turn caused several managed Postgres clusters to go unhealthy (with some temporarily offline) and led to slower/less reliable machine operations on affected hosts. This was mainly caused by a large amount of machines suspending at once, and a lack of concurrency limits / queuing on this operation. We addressed this by limiting how much I/O is allocated to writing machines’ memory snapshots to disk on suspend, and adding limits to the number of suspending machines allowed at once.


# April 6: Web Sidekiq backlog from stuck usage jobs (19:32UTC)

This was a weird one – our Sidekiq instance, the one that powers background jobs of our GraphQL API and dashboard, repeatedly got stuck processing long-running billing and usage sync jobs. That led to large job backlogs and delayed processing. We first attempted to mitigate by scaling up and restarting stuck workers, which worked for a while, but eventually everything ground to a halt again. After chasing down a few false leads, we eventually tracked this down to two major causes:

(1) when Sidekiq workers use too much memory, Sidekiq sends SIGUSR2 to kill the process, which stops the process from accepting new work and wait for any existing work to complete before exitting; (2) however, our database connections for those billing / usage jobs sometimes got stuck without any proper timeout. When their worker processes are killed with SIGUSR2, they get into a state where they neither exit nor make any further progress; eventually, we are left with no workers that can process jobs.

We hardened our setup by adding proper timeouts to both Sidekiq worker shutdown and the database connection pool. We also separated billing / usage jobs into their own queue to avoid blocking other tasks that need to be processed.


# April 5: Private networking outage in Sydney (again) (10:37UTC)

Private networking between our Sydney region and other regions started failing again as an upstream provider was filtering UDP traffic again. This was resolved promptly by moving traffic away from the affected provider.


# March 29: Sprite creation errors in SJC and AMS (14:36UTC)

Sprites are provisioned from pools, which are our regional buffers for spinning up the underlying apps and machines. This was a brief incident where the pool monitors, responsible for replenishing these pools, weren’t running in a couple of regions. Once the monitors were up again, pools refilled and sprite creation returned to normal. Ideally we catch this before the pools are empty, so we’re adjusting our alerting to make sure that’s the case.


# March 27: Sprites API errors (18:00UTC)

Some Sprites API endpoints saw elevated errors for a subset of organizations when the per-org database was resuming from an idle state. Only requests needing fresh sprite data were affected, whereas already-running Sprites continued to work normally, including new connections to them. Our orchestrator normally catches anything like this and tries again, but this was a novel failure mode it didn’t recognize as retryable.

March 27: IAD CPU crunch

# March 27: IAD CPU crunch (17:56UTC)

IAD had a spike of machine start failures due to increased CPU usage across machines in the region. We had some dormant hosts in IAD sitting aside for boring things like system upgrades, so we brought some of these back online to fill the gap while spinning other resources up.


# March 26: ORD machine creates bogged down (15:18UTC)

A solid chunk of machine creation requests in ORD were timing out. We tracked it down to one ORD server that had become a very frequent placement target but was taking extremely long to complete flyd machine create operations. In poking at it, we noticed this host’s bolt store on disk was huge, which bogged down flyd enough that flaps, our Machines API frontend, timed out the request before it completed. This was mitigated by pulling the host as a placement option, and then compacting its event store before putting it back in the pool.

March 26: FRA region outage

# March 26: FRA region outage (12:35UTC)

Both redundant fibre links to our primary provider’s rack in FRA dropped offline. This caused most apps hosted in the FRA region to become inaccessible. The links came back up after about 40 minutes, but some Managed Postgres clusters needed additional time to catch up.

Honestly, we (and our provider) are unsure why this happened; it was likely human error of some sort. We don’t expect this to happen again.


# March 24: GraphQL timeouts (21:04UTC)

We called an incident due to elevated errors through our GraphQL API, where requests were timing out when talking to our primary database. All errors were coming from a single machine in IAD, where the physical host for that machine seemed to be misbehaving and couldn’t hold a connection, so it earned itself a reboot. The host and its machines were behaving once it came back up, and everything seems fine, but this is a server we’re going to squint at with suspicion if it misbehaves in the future.


# March 23: Errors viewing logs in Grafana (15:02UTC)

The logs panel on fly-metrics.net started throwing up an error due to an internal mTLS certificate expiring. This prevented customers from viewing logs in Grafana only, and both fly logs and the dashboard log viewer continued to work. We fixed this, then briefly broke it again, then fixed it.


# March 20: DFW capacity, again (07:20UTC)

The capacity we provisioned in DFW but two days earlier was slurped up in short order, so we were back to machines on some hosts being unable to start due to resource gates. With the popularity of the region, we found our placement would concentrate new machines onto the few best hosts at any given moment, which would itself create new bursts of start failures. In response to this incident we once again brought new hosts online, alongside improving our placement logic to better handle this case.


# March 19: Metrics outage (06:26UTC)

Metrics graphs were impacted when the storage backing our metrics pipeline filled up, which stalled ingestion. Typically a metrics outage like this will backfill the missing data on recovery, but a configuration mistake this time meant we briefly accepted metrics without forwarding them on, resulting in roughly an hour of permanently lost metrics. We still have some open investigations on this one, such as why the storage filled up (it shouldn’t have), and why our alarms didn’t fire (they should have).


# March 18: SJC disappeared (briefly) (13:57UTC)

A firmware bug on a switch upstream of us caused traffic to our sjc servers to be dropped. This was fixed within a few minutes. For a moment, most instances in sjc were unreachable: apps saw connection failures, and some Managed Postgres clusters were knocked into unhealthy states. The upstream issue was resolved immediately, so this was largely a monitoring incident on our end.

March 18: DFW capacity errors

# March 18: DFW capacity errors (09:21UTC)

Another region hit its capacity ceiling after higher-than-normal growth. We were able to spin up more physical hosts over the following hours, during which machine start errors in DFW settled back down.


# March 17: A bunch of wedged Sprites (11:30UTC)

A large number of sprites became difficult or impossible to wake after migrations left them in a failed state due to capacity limits in their region. This showed up as 502/503 responses from sprites that should have been started on-demand. From the perspective of the Fly Proxy this was strictly correct: failed should be a terminal state when a machine fails to launch. This incident showed that this state was mistakenly set for machines that failed to start after being migrated, which is very much not the same thing. Our quick fix was rolling out a proxy change that tries to start failed machines anyway, and once that put out the fire we focused on cleaning up our migrations to handle capacity issues better across the board.


# March 16: Tight capacity in ORD and SIN (09:12UTC)

Demand in many of our regions is growing a lot this year. So much so that our ORD and SIN regions hit capacity faster than we could bring new hardware online. The main impact from a constrained region is that existing machines may fail to start if the physical host they’re on is above our safety margin for resource usage. Deploys also see some effect, when a valid host to place a new machine can’t be found. In this case, the incident was largely resolved by provisioning new ORD hardware and rebalancing some workloads across the region.


# March 14: Sprites API didn't like numbers (13:52UTC)

We called an incident due to increased errors for certain customers through the Sprites API. It turned out to be a newly-introduced bug that only affected organizations with slugs that had leading numbers, which was promptly fixed. Outside of these organizations there was no impact to Sprites.


# March 11: GraphQL overloaded (10:26UTC)

Our GraphQL API became overloaded when requests backed up after our secret storage service was temporarily unavailable. Working through the increased backlog of requests led to timeouts and failing health checks for the dashboard and GraphQL endpoints (the Machines API itself remained available). We mitigated the issue about 25 minutes after impact began by scaling up the GraphQL API and increasing its request concurrency.


# March 7: Sydney WireGuard outages from upstream UDP filtering (14:39UTC)

WireGuard connectivity in our Sydney region intermittently failed because an upstream network backbone was dropping/filtering UDP traffic. This caused 6PN private networking to fail between Sydney and other regions, as well as some internal services to become degraded, until the upstream issue was corrected; service returned to normal within minutes once the filtering stopped.


# March 6: App-scoped IPs deleted by mistake (04:11UTC)

Our egress IPs are assigned in regional blocks. A while ago, when we deprecated some regions, these blocks and assignments were migrated to the new destinations. On this day, an automated egress IP reseed reintroduced stale “unassigned” egress IP records for a few deprecated regions; these were then synchronized into regional Corrosion clusters, which makes use of our fly-force-region header through the Fly Proxy, which itself rewrites deprecated regions to their migration destinations. This inadvertently deleted the assigned app-scoped egress IP entries which were migrated to the new regions. Affected apps temporarily lost the expected app-scoped egress IP, until we manually re-synced the assigned egress IPs. We then removed the stale deprecated-region records and added safeguards to keep deprecated regions out of the reseed path and to harden internal regional-routing requests so this can’t recur.


# March 5: BGP route leak sent North America traffic to Singapore (19:20UTC)

A BGP routing issue caused a significant amount of North American traffic to be routed to Fly edges in Singapore (sin), increasing latency and causing some disconnects. Impact stopped once we fixed up routing.


# March 3: GraphQL mutations failing (20:15UTC)

Some GraphQL mutations were unavailable after a Redis node encountered an underlying disk/filesystem failure. We restored service by bringing up a fresh Redis node and switching over ~20 minutes after the start of the incident. We’re still investigating why the original node’s storage failed.

March 3: Cost Explorer errors from internal timeouts

# March 3: Cost Explorer errors from internal timeouts (10:37UTC)

The Dashboard Cost Explorer intermittently loaded very slowly and sometimes failed with an error because requests to our billing provider’s spend-breakdown endpoint timed out. The page was flaky, especially for fresh queries with different date ranges or per-app breakdowns. We shipped UI changes to reduce reliance on the problematic endpoint and show a summary view for larger organizations.


# March 2: Petsem got overwhelmed (21:15UTC)

A slow database query in Petsem caused secret lookups to start timing out, preventing some machines from starting during the incident. Our theories were pointing toward a machine on one of our hosts that was trying very hard to initialize a volume with a missing block device, triggering a large number of secret lookups. We recreated the missing block device, and also flipped some feature flags around to reduce load on Petsem. Fortunately, one of these actions worked and Petsem recovered, marking the end of the public incident.

This failure mode was triggered because this app had a lot of volume encryption keys stored in Petsem – even though it only had a handful of existent volumes, we don’t delete the encryption keys when volumes are deleted, as we still need them to decrypt volume snapshots. Usually the impact of a slow read query is limited, as we have multiple read replicas and can easily scale them up. However, changes to routing (as part of the regionalization project) had caused most of North America (including SJC, where the problematic volume was located) to route directly to the write primary, instead of a nearby read replica (which was seeing almost no load throughout the incident). This meant, not only was Petsem unable to accept writes, but it also couldn’t accept reads.

As a response to this incident, we’re working to migrate Petsem to a new database schema that is compatible with the database indexes needed to make these read queries consistently fast.

March 2: Our certs vault went down and took some proxies with it

# March 2: Our certs vault went down and took some proxies with it (20:09UTC)

A partial update to Vault, which we use to store TLS certificates, left it in a state where requests would hang for several minutes. In the face of most Vault issues, the Fly Proxy will continue with the cached certs it has on hand, limiting impact only to those hostnames that are new or have not been served recently. Annoyingly, Vault wasn’t down down, it was just slow. As our proxies encountered hostnames with uncached certificates, a growing pile of requests stalled on the hung-but-not-failing Vault lookup. Eventually there were enough requests in this state on some proxies that we simply ran out of connection slots, causing widespread connection errors. Service recovered after a few minutes once Vault finished its update.


The infra log took a long sabbatical, but we’re bringing it back with a new format.

Instead of weekly roundups, each post now covers a single incident, published roughly seven days after it occurs, give or take.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)

It’s looking quiet. A little… too quiet.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)


#

A note on incidents: incidents are internal events for our infrastructure team. Incidents often correspond to degraded service on our platform, but not always. This log aims for 100% fidelity to internal incidents, and is a superset both of our status page events and of customer-impacting events on the platform. It includes events reported to subsets of customers on their personal status pages, as well as events without any status page impact.

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)

Anchor February 16 Outage postmortem

Anchor Narrative

At approximately 04:45EST on February 16, we experienced a total outage of our IAD region. For logistical reasons having to do with the centrality of Ashburn, Virginia to the proper function of the entire Internet, IAD is our primary API region, so this outage took our API down with it. During the four hours on Sunday morning that IAD was down, our users could neither deploy nor modify Fly Machines or applications in any region of our fleet, and Fly Machines hosted in IAD would not have been available to users.

We record lots of incidents in this log, but severe, sustained outages are uncommon. When they do happen, they tend to be relatively gnarly combinations of distributed systems issues and operations fallibility. In other words, they tend to be interesting to write about.

Not this one. This outage was simple: our IAD upstream provider had a core switch (a switch upstream of our racks, deeply embedded in their regional network) that faulted out, and it didn’t have a redundancy. That made recovery an hours-long project rather than a minutes-long project. The switch that failed wasn’t a top-of-rack device in one of our racks, but rather a transit switch.

We operate a global fleet with hardware in over 35 regions. While sustained platform-wide are uncommon, regional network cuts are less uncommon. There are parts of the world that are difficult to run stable networks in relative to London or New York, especially in Latin America. We would not normally write a detailed postmortem of an outage that cut off GRU or SCL. The Fly.io platform is resilient to whole-region outages in most of the world.

IAD is different; it hosts our APIs and much of our API backing stores. If our IAD data centers are cut off completely, our API won’t function.

We operate in several providers in IAD, with multiple upstreams. For example, our HashiCorp cluster — the giant servers managing our Consul and Vault deployments — is hosted at Equinix. Some of our worker servers, on which customer workloads run, are also hosted in Equinix. But the bulk of our servers, including the ones that happen to be hosting our Fly APIs (which are Fly Apps themselves) are not at Equinix (the hardware product we take advantage of at Equinix is both nosebleed expensive and has been sunset).

Much of the operator/engineer time we spent during this outage was aimed at bringing our APIs up in a “backup” region (EWR, in Secaucus, to be specific). To their credit, our team got us there, just as our upstream provider managed to restore connectivity. But we want to be clear that being able to bring our API up in a region other than IAD is not a business goal of ours. As a general rule we accept the risk that if IAD falls off the Internet, people won’t be able to deploy Fly Machines.

This sounds weird (even to some of our own engineers). But maintaining an “IAD-resilient” API is not a costless choice. Keeping our core infrastructure constantly ready to bring up in EWR wouldn’t just be expensive and distracting, but would also add further complexity to our infrastructure. If you scroll back (and back, and back) on this log, you’ll see that overall, it’s complexity, and not backhoes in Virginia, that are our real adversary.

Anchor Incident Timeline

Anchor Forward Looking Statements

Throughout this postmortem, we’ve been at pains to be clear that API resilience from a sustained failure in IAD is not part of our service model. We’ve made a strategic choice to simplify our platform by allowing our API to depend on the availability of the IAD region. This means that if a truly epic bolt of lightning strikes Ashburn, Virginia, we might experience a sustained API outage. Existing Fly Machines outside of IAD will continue to run (that kind of resilience is part of our service model) but deployments and modifications will not function.

We’re impressed with and grateful to our infra and platform engineering teams for coming very close to the finish line of migrating our API out of IAD, on the fly, in the span of just a couple of hours. This postmortem provides some detail about that work out of appreciation for the engineering acumen required to pull that off. But we’re generally waving engineers off of the work of making that kind of migration easier or more reliable.

Instead, all of our effort in response to this is about bulletproofing our connectivity in IAD.

There are two broad things we’re doing to prevent future network cuts in our busiest regions, starting with IAD. The first is working with our upstream provider on network engineering, and the second is diversifying our upstream connectivity. Both efforts are active and ongoing; we should see payoffs (expect them in this infra-log) in the coming weeks.

With our upstream, we’re inventorying and auditing the entire network between our metal and the cross connects to transit providers. We were surprised to learn, that Sunday morning, that our upstream had a non-redundant transit switch in their architecture. We’re going hunting for more of them, and we’ll work in partnership with them to ensure those weaknesses are rectified.

In running postmortems with staff from that upstream, we’ve discovered some important process and communications issues that we’ve been able to resolve. Some of what we’ve uncovered is the result of Fly.io growing organically alongside our upstream partner, which has resulted in a less-than-optimal physical architecture for our hardware and theirs. Some of it is also expectations management; much of the rest of their server footprint is operating in CDN configurations, where the cost of a sustained region outage is less-than-optimal page load times. And some of it is the communications process running between our organizations, with staff engineers on our side talking directly to staff on their side, which has the benefit of warding off downtime (“can we move this server?” “no!”) but also deferring maintance (“can we move this ser—” “no!” “—ver to make the power space we need to install a redundant switch here?”).

Because our strategy depends on high availability for IAD, we’re also investing in additional cross-connects (and network paths to them) from other providers. If this was simply a matter of striking up contracts with transit providers, it would be done by now. Unfortunately for us, as just related, we’ve grown organically and messily in the two data centers our upstream operates in, and server numbering is going to complicate transit diversity for us. We’re in the beginning stages of planning out the automated provisioning and renumbering that will make this possible. It’s not a huge lift, but it’s more of a lift than you’d expect.

Every major incident we deal with surfaces internal process issues we can improve. Here, we’re not especially psyched about how coordinated our immediate response to the outage was: it took us something like 20 minutes to have a coherently communicated understanding of what was happening, from the time the initial pages hit our team. This was complicated by the fact that some of our alerting infrastructure has IAD dependencies. While we don’t plan to make our whole API migrateable, our alerting infrastructure needs to function (better) no matter what.

This infra-log is a product of our engineering organization, and does a pretty good job of surfacing strong work from our infra, platform, and fullstack engineers. In this specific incident, we’d like to go out of our way to call out Matthew Cunningham, our VP of Finance. Matthew is excellent for any number of reasons, but here in particular he’s taken the lead on driving forward work and investments now underway with our upstream to ensure an incident like this is unlikely to recur. You’d never know to read his commentary about cabinets, cross connects, and network inventories that he was a fleece-vest finance person. Thanks, Matthew!


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)

turns out that Saleem, in fixing bugs in our guest-resident SSH server hallpass, broke bug compatibility with flyctl.

Anchor This Week In Engineering

We have an excuse every week for not updating this, don’t we? There’s two blog posts coming this week about infra work, but also we’re doing investigative work with our upstream on the outages they’ve experienced; that, and a lot of on-call afterhours stuff, and we’re going to lay off the infra team this week too.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)

Anchor This Week In Engineering

Extending last week’s “new rule of the infra-log”, of none of the incidents in a week have anything really directly to do with us, the infra-log author is still not going to hassle the team for updates. Also: we’ve lit the blog back up, and some of what the infras would tell me really belong on the blog, not here.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)

Yep. We got nothing.

The new rule of the infra-log: if there are zero incidents of any sort, the author of the infra-log doesn’t get to pester engineering for updates. I’m sure something will happen somewhere before next week (the author

said, fate-temptingly) and we’ll have some engineering updates for the

next post.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)

Anchor This Week In Engineering

Ben A. is working on observability/telemetry for Fly Machine migrations. A Fly Machine is brought into being through a recorded log of finite state machine steps stored in a BoltDB; “migration” is just another FSM, as far as our orchestrator is concerned. The migration FSM is somewhat complicated. We need better visibility into how each FSM step can fail. A reasonable way to think about this is that we’re trying to do for migrations, which are a fully asynchronous distributed system, what we do with oTel tracing elsewhere in our platform: get clear sightlines without scraping logs, which is what we’ve been stuck with ‘til now. See, Ben, I had no trouble writing something interesting out of that.

Dusty created a unified closed loop system for tracking hardware issues across all our various providers, in order to stop being our single source of “which providers are working on which issues” truth for us. Those kinds of issues are now bridged into our Slack and recorded in a single issue database. I could tell a story about how this will resolve hardware issues and bring capacity online more quickly for users, but really, this is just making life better for our infra team.

Will and Tim are working on stable 6PN addressing for Fly Machines. This is motivated by a terrible mistake we made several years ago when we launched Fly Postgres and had it auto-configure with IPv6 address literals; those addresses embed hardware addresses in them, which cause obvious problems when we need to migrate a database from one host to another. Will also did a deep dive into IO scheduling and performance; there are things we can be doing to manage IOPS load between different applications, and also things we can do with our hypervisors to increase IO performance (Firecracker, our mainstay hypervisor, is single-process-per-VM, which means that CPU, disk, and network can all cause contention with each other; there are fixes for this, but they depend on kernel updates we don’t uniformly deploy).

Somtochi spent the week on a support rotation, working with Lillian (welcome back Lillian!) to address gnarly customer issues. We all do rotations in support, except for me; I cop out by saying I’m too busy writing this bulletin.

Peter is moving stuff forward on regionalizing fly-proxy. The core idea here is to decreate the distributed systems failure blast radius, by relaxing the design constraint that every proxy region has fine-grained knowledge of every Fly Machine on our platform. This is straightforward to do for HTTP services, which have a command/control system that allows us to bounce requests around. It’s a real problem for raw TCP services, though: once we make a TCP connection, we can’t naively retry it elsewhere. Peter has a bananas solution for this that (I am not making this up) uses the TCP URG pointer as a signaling mechanism. We’ll write more about this if it doesn’t end up dooming us all.

Kaz continues to work on the AMD-Vi reboot process (there’s an AMD virtualization hardware bug that pops up randomly once an affected host gets past a threshold uptime, so we’re doing orderly maintenance-window reboots of upcoming potential victims). This week: lots of customer notification work.

Steve pitched in to Fly Kubernetes (FKS), building a backup system that stores etcd data in object storage. Steve was also in the middle of the Consul outage we documented this week: that turns out to have been a hardware issue, a machine burning out all its NVMe drives within an hour, and then a scramble to replace it.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact.)

Anchor This Week In Engineering

Lots of little stuff.

Somtochi spent the week hunting a Corrosion bug that caused periodic stalls, and, long story short, it was an IOPS contention issue with a Corrosion server deployed alongside a particularly demanding customer application. OK, moving on.

JP worked out a runbook for a flyd database capacity issue. Recall that every worker server in our fleet is the sole source of truth for the workloads running on it (we run a “decentralized” orchestrator); that source of truth is a BoltDB kept by flyd, our primary orchestrator service. That database records every transition that happens on every Fly Machine on the server (basic operations like starting or stopping a Fly Machine might incur many such transitions under the hood), and, over time, that database accumulates garbage, gradually slowing flyd down. We now have alerting for flyd database operations that exceed a time threshold, and a runbook for compacting the database when it does. Boring, but good. Moving on.

Will bit off all his fingernails keeping tabs on the rollout of our new CPU scheduling. We’re writing a blog post about this. Will’s also working on a long-term infra white whale of ours, which is providing stable private IPv6 addresses for Fly Machines even as they’re migrated to new hardware; this is difficult because our IPv6 addresses encode hardware addresses as part of our routing discipline. There’s some neat engineering happening here.

Kaz has been leading our (P)reventative (H)ost (M)aintenance (P)roject (P.H.M.P.), which addresses a problem we have on large AMD server hardware: because of a firmware bug, those machines can lock up if their uptime exceeds a (very large) threshold, which some of our worker servers do, because we go way out of our way not to reboot them. We’ve been working through a hit list of servers that we’ve succeeded “too” well on, deliberately performing maintenance reboots on them, which involves migrating workloads off them first, because modern servers take forever to reboot, which is why we try to avoid rebooting them. Anyways, that’s what Kaz’s life is now: rebooting lots of servers.

Peter is working on regionalizing fly-proxy, which means altering our Anycast routing system so that every region in the world doesn’t need to maintain state for individual workloads in other regions, a change that addresses two of our most severe outages. But Peter was also mostly on winter vacation during this time period. We’ll check up on him again next week.

Dusty gave me an update that had something like 9 animated emojis in it, about addressing flyd reliability issues. The long and the short of it: he did a study of all recurrent logged errors from flyd and root-caused each of them in turn, regardless of whether metrics and active health checks were nominal for the impacted flyd instances. He found a bunch of Fly Machines in inconsistent states, a lot of machines with dedicated swap block devices that were unhappy, and some edge servers that were only partially-provisoned (we have a lot of edge servers, and our system generally routes around janky edges). Also: he found a bug in our metrics reporting code that had our workers continuing to report metrics for dead Fly Machines, which produced some happy graphs of of metrics volumes once it got deployed.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; this week, unlike ordinary weeks, all 3 rows of the chart are “fresh.”)

Anchor This Week In Engineering

Your author has spent this week writing a blog post and is thus derelict in their duty to interview our infra team about the fun stuff they’ve been working on, and does not have the chutzpah to reach out to the team at 9:00PM. We’ll maybe cut a special interim update tomorrow, but we don’t like leaving you hanging on the incident reports.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; this week, unlike ordinary weeks, all 3 rows of the chart are “fresh.”)

That’s it! Normally, even though major incidents are pretty rare, there’s a smattering of little things to post here; “incidents” we flagged internally that didn’t merit a status page update, or that had limited impact. Not this week, though!

Anchor These Weeks In Engineering

The lack of incidents to write about isn’t really just happenstance. We locked the platform down in anticipation of the holidays; major changes, such as to our state propagation system, our Anycast routers, or Fly Machine scheduling, were all frozen. We don’t want to get paged on New Years Eve any more than you do.

So! Not many major things to report this last month! Engineering updates should pick up next week.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Engineering

We’re back, baby! (But give us a break over the next week or so, for obvious reasons).

Ben and Simon worked out that we were double-storing container images for customer apps. [Recall that we exploit containerd](/content/blog/docker-without-docker/ ""/index.html) to “stage” the block devices we boot VMs onto with customer app images; once as the products of the containerd snapshotting plugin (so, as LVM snapshots, stored in a large LVM thin pool) and once as a blob in the containerd content store (as temporary storage while containerd makes the LVM snapshot). The blob store is on our root storage device. This is bad. containerd GC’s the content store; this is good. But flyd, our orchestrator, labels content in containerd in such a way that GC doesn’t work. This is bad. Ben and Simon are fixing that, which is good.

First and foremost, Dusty provisioned all our standby hardware; we now have zero unprovisioned servers. If we own it, it’s ready to go into prod. But more ambitiously, Dusty is leading efforts on what we’re calling the “hardware resiliency” project, which is what it sounds like. Most of the docket for this project right now is about I/O performance, and heavily concentrated on our volume backups (because that’s a major I/O load that we actually control, unlike your apps, which we do not). Another completed item on that checklist that rhymes with a prior outage is a fleetwide audit of all our certificates and their expirations, which is now done.

JP and Senyo released Pilot, our new init. Why don’t I just let Annie explain this one?

Tom‘s white whale right now is sourcing new “burst capacity” for us. We have long-term stable hardware and hosting, but it’s not super fast to bring online, and every once in awhile we get spikes of demand in particular regions. We have providers we can “burst” into while we bring long-term capacity online for those regions, but it’s pricey and a little precarious, so we’re building a deeper back bench of providers to burst into. Tom also re-did all our BGP4 configuration management, so we now have an effective CI/CD process fo BGP4 announcement changes, as well as a full-documented configuration.

Peter is doing some of the most important reliability work in the company getting fly-proxy load balancing working with regionalized Corrosion. Regionalized Corrosion, again, is the effort to take most of the state we track right now and make it local (within a region, like IAD or SYD) rather than global. This breaks a huge assumption we’ve made about what information our proxies have to work with! We now have regionalized load balancing up and running behind a feature flag (your apps: not balanced regionally yet), which is huge. Another huge lift: getting fly-replay, which is like our signature feature, working in the new world order where a given fly-proxy actually doesn’t know what all the specific instances of an app are; this is tricky because fly-replay lets you direct requests to specific instances by name.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).


#

Anchor Trailing 2 Weeks Incidents

The infra-log took last week off for Thanksgiving; the trailing two weeks in this update are two “fresh” weeks worth of incidents, though basically nothing happened in the first of those weeks.

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor November 25 Outage Postmortem

Anchor Narrative

At approximately 15:00EST on November 25, we experienced a fleetwide severe orchestration outage. What this means is that for the duration of the incident, both deployments of new applications and changes to existing applications were disrupted; during the acute phase of the outage Fly Machines could not be updated; for the back half of the outage, our API was unavailable. Service was restored completely at 02:30EST.

This was a compound outage with two distinct causes and impacts. The first approximately mirrored the October 22nd orchestration outage and involved a distributed systems failure in Corrosion, our state sharing system. The second was an API limit problem, which combined with an error in a customer app had the effect of denying service to our API. The two outages overlapped chronologically but were resolved serially, extending the duration of the incident.

We’re going to explain the outage, provide a timeline of what happened, and then get into some of what we’re doing to keep anything like it from happening again.

Orchestration is the process by which software from our customers gets translated to virtual machines running on Fly.io’s hardware. When you deploy an app on Fly.io, or when your Github Actions CI kicks off an update after you merge to your main branch, you communicate with our orchestration APIs to package your code as a container, ship it to our container registry, and arrange to have that container unpacked as a virtual machine on one or more of our worker servers.

At the heart of our orchestration scheme is a state-sharing system called Corrosion. Corrosion is a globally-synchronized SQLite database that records the state of every Fly Machine on our platform. Corrosion uses CRDT semantics (via the cr-sqlite crate) to handle SWIM-gossipped updates from worker servers around the world; a reasonable first approximation holds that every edge and worker server in our fleet runs a copy of Corrosion and, through gossip updates, synchronizes its own copy of the (rather large) global state for the Fly.io platform.

The proximate cause of this Corrosion incident is straightforward. About 5 minutes before the incident began, a developer deployed a schema change to Corrosion, fleet-wide.

The change added a nullable (and usually-null) column to the table in Corrosion that tracks all configured services on all Fly Machines (that is to say: if you edit your fly.toml to light up 8443/tcp on your app, this table records that configuration on every Fly Machine started under that app). Surprisingly to the developer, the CRDT semantics on the impacted table meant that Corrosion backfilled every row in the table with the default null value. The table involved is the largest tracked by Corrosion, and this generated an explosion of updates.

As with the previous Corrosion outage, because this is a large-scale distributed system, Corrosion quickly drove tens of gigabytes of traffic, saturating switch links at our upstream.

This outage is prolonged by a belief that the root cause is an inconsistent set of schemas on different instances of Corrosion.

The incident begins (and is alarmed and declared and status-paged) promptly after the schema change is deployed. The deployment is immediately halted, and investigation begins. Corrosion is driving enough traffic in some regions to impact networking, and cr-sqlite‘s CRDT code is consuming enough CPU and memory on many hosts throw Corrosion into a restart loop. Now the deployment is allowed to complete, to rule out inconsistency as a driver of the update storm. The deployment doesn’t worsen the outage, but does take time, and doesn’t improve the situation.

As with the October 22nd outage, the Corrosion problem is resolved when the decision is made to re-seed the database from an external source of truth. This time, the schema change complicates the process: a backup snapshot of the Corrosion database from prior to the schema change is needed, and downloading and uncompressing it adds time to the resolution.

As with the previous outage, once the snapshot is in place, re-seeding Corrosion takes approximately 20 minutes and resolves the Corrosion half of the outage.

At the same time this is happening, a corner-case interaction between a malfunctioning customer app and our API is choking out our API server.

The customer’s app runs untrusted code on behalf of users (this is a core feature of our platform). It does so by creating a new Fly Machine for each run, loading code on to it, running it to completion, and then destroying the Fly Machine. This works, but is not how the platform is meant to be used; rather, our orchestrator assumes users will ahead-of-time create pools of Fly Machines (dynamically resizing them as needed), starting and stopping them to handle incoming workloads; a stop of an existing Fly Machine resets it to its original state. Start and stop are much, much faster than create and destroy.

The customer’s app is suddenly popular, and begins creating dozens of Fly Machines every second, at a rate steadily increasing throughout the outage. This exercises a code path not expected to be run in a tight loop and missing a rate limit. In our central Rails API server, which is implicated in create requests (but not starts and stops), this has the effect of jamming the process up with expensive SQL queries.

A different team is investigating and working on resolving this incident alongside the previous one. The team attempts to scale up to accommodate the load, first at the database layer, and then with larger Rails app servers; dysfunction in the Rails API makes the latter difficult and time-consuming, and ultimately neither scale-up resolves the problem: paradoxically, as we create additional capacity for create requests, the lack of backpressure amplifies the number of incoming create requests we receive.

30 minutes before the end of the outage, we reach the customer, who disables their scheduling application. The API outage promptly resolves.

Anchor Incident Timeline

This timeline makes reference to the Corrosion outage as “Incident 1”, and the API flood as “Incident 2”.

Anchor Forward-Looking Statements

A significant fraction of this outage rhymes with our previous orchestration outage, and much of what we’re working on in response to that outage applies here as well.

The most significant thing we’re doing to minimize the impact of outages like these in the future is to reduce global state. Currently, every physical in our fleet has a high-fidelity record of every individual Fly Machine running on the platform. This is a consequence of the original architecture of Fly.io, and it’s a simplifying assumption (“anywhere we need it, we can get any data we want”) that we’ve taken advantage of over the years.

Because of the increased scale we’re working at, we’ve reached a fork in the road. We can continue running into corner-cases and bottlenecks as we scale and manage high-fidelity global state, and develop the muscles to handle those, or we can break the simplifying assumption and do the engineering required to retrofit “regionality” into our state. As was the case late this summer with fly-proxy, we’re choosing the latter: running multiple regional high-fidelity Corrosion clusters, with Fly-Machine-by-Fly-Machine detail about what’s running where, and a single global low-fidelity cluster with pointers down to the regional clusters.

The payoff for regionalized state is a reduced blast radius for distributed systems failures. There’s effectively nothing that can happen with Corrosion today that doesn’t echo across our whole fleet. In a regionalized configuration, most problems can be contained to a single region, where they’ll be more straightforward to resolve and have drastically less impact on the platform. Corrosion, an open-source project we maintain, is already capable of running this way; the work is in how it’s integrated, particular to our routing layer.

This work has been ongoing for over a month, but it’s a big lift and we can’t rush it. So: we’re cringing this overlap this outage has with our last one, but it’s not for lack of staffing and effort on the long-term fix.

Two immediately evident pieces of low-hanging fruit that we have already picked in the last week:

First, as we said last time, responding to Corrosion issues by efficiently re-seeding state continues to be an effective and relatively fast fix to these issues. Re-seeding was complicated this time by the schema change that precipitated the event. We’ve begun creating processes to simplify and speed up re-seeding under worst-case circumstances. Additionally, some of the delay in kicking off the re-seeding, as with last time, resulted from a cost/benefit calculation: re-seeding requires resynchronizing our API and flyd servers across our fleet with Corrosion, which isn’t automated; the hope was that Corrosion would converge and reach acceptable P95 performance levels soon enough not to need to do that work. We’re building tooling to minimize that work in the future, so that doesn’t need to be part of the calculation.

We’ve also added comprehensive “circuit-breaker” limits to Corrosion. For already-deployed apps, even a total Corrosion outage shouldn’t break routing; Corrosion synchronizes a SQLite database, and our routing layer can simply read that database, whether or not Corrosion is running. But during the acute phase of this outage, Corrosion wasn’t just not running effectively; it was also consuming host and (especially) network resources. Corrosion now has internal rate limits on gossip traffic, and our hosts have external limits, in the network stack and OS scheduler, to stop runaway processes; this is being rolled out now.

Second, the back half of this outage was due to a pathological condition we hit because we lacked a rate limit in an expensive API operation we didn’t expect users to drive in a tight loop. Obviously, that’s a bug, one we’ve fixed. But there was a process issue here as well: we identified the “pathological” app (it wasn’t doing anything malicious, and in fact was trying to do something we built the platform to do! it was just using the wrong API call to do it), but then engaged in heroics to try to scale up to meet the demand. Without backpressure, this doesn’t do anything.

So we’re also building out a process runbook for handling/quarantining applications that have spun out, one that incident response teams don’t have to think hard about when in the middle of high-priority incidents.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Engineering

Despite a quiet week incident-wise, the whole team was unusually interrupt-driven this week; a consequence of catching a bunch of stuff before it could actually become an incident.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Engineering

Somtochi is working on bringing up regional Corrosion clusters. Recall that Corrosion is our state-sync service; think of it as a replacement for Consul, driven by SWIM gossip rather than Raft consensus, and with a SQLite interface (it essentially gossip-syncs a big SQLite database of everything running). In the wake of the Anycast outage from a few months back, we’ve been working on splitting Corrosion into a much smaller global cluster than we currently run (that is: gossiping less state into the global cluster) and then supporting it with regional clusters. A good first approximation of what we’re talking about: the global cluster knows every Fly App running in every region on the fleet, but the regional clusters know the specific Fly Machines for those Fly Apps running on the worker physicals in their region. Anyways, that’s what Somtochi is working on; this week, that mostly involved teaching our Corrosion-backed internal DNS service how to fetch information for machines in another regional cluster.

JP and Jerome diagnosed and fixed a gnarly volume migration bug that temporarily broke Jerome’s Fediverse server. If a Fly Volume is extended while that volume is in the process of being migrated (meaning that behind the scenes, dm-clone is still “hydrating” the volume over a temporary iSCSI connection from the origin worker physical), the underlying volume operation could apply to the wrong block device (the temporary clone device, not the final device). This was a missing step in the flyd FSM for restarting Fly Machines, now fixed.

Will upgraded VictoriaMetrics. The one incident we had last week was from an aborted partial attempt to upgrade Vicki. Well, we succeeded this week. In the process of investigating that outage and completing the upgrade this week, Will spotted a perf issue in upstream Vicki that degraded cache performance in Vicki clusters with large numbers of tenants (like we operate), and wrote an upstream PR for it.

Steve was on support rotation. Engineers across the team all do time, a couple days at a time, as technical escalation for our support team. Our support team is great, but being directly exposed to customers is as helpful for product engineering as it is for the support team. Steve and Peter are also hip-deep in working out plans to reboot large numbers of worker physicals, which is a fun problem we’ll be writing about in the weeks to come. Nothing dramatic is going on, we just need a process to reliably schedule reboots and maintenance windows.

Peter spent the week rolling out lazy-loading Corrosion state in fly-proxy. Currently, all the state we hold about every app running on the fleet is kept in-memory in fly-proxy (the component that picks up your HTTP requests from the Internet and relays them to your Fly Machines). As part of the work we’re doing to make Corrosion more resilient (along with regional clusters), we’re changing this, so that we load state for apps only when they’re actually requested. By way of example: the author of this log had one bourbon too many back in 2022 and booted up “paulgra-ham.com” on Fly.io, which is an app that has never once been requested since. Ever since that moment, fly-proxy has assiduously kept abreast of the current state of “paulgra-ham.com”, every minute of every day on every edge and worker in the fleet. This is dumb, and makes fly-proxy brittle. So we’re not doing it anymore.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Engineering

We apologize for the delay this week. We’re a US company and the US was eventful! Also, there wasn’t much incident stuff to write about. We pledge to be more timely in the weeks to come.

Somtochi is back to doing surgery on Corrosion. It now exposes a lighter-weight update interface that streams the primary keys of updated nodes over an HTTP connection, rather than repeatedly applying queries. Corrosion also favors newer updates over older ones during sync, which speeds time-to-recovery when bringing nodes online and dealing with large volumes of updates.

Akshit has been working on static egress IPs. Some of our customers run Fly Machines that interface with remote APIs, and some of those APIs have IP filters. Normally, Fly Machines aren’t promised any particular egress IP address for outgoing connections, but we’re rolling out a feature that assigns a routable static egress IP. Akshit also wrote a runbook for diagnosing issues with egress IPs and its integration with nftables and our internal routing system.

Dusty built a custom iPXE installation process for bringing up new hardware on the fleet. Our hardware providers rack and plug in our servers, and PXE pulls a custom initrd and kernel down from our own infrastructure, eliminating an old process where we effectively had to uninstall an operating system configuration our hardware was shipped to us with, making it faster to roll out new hardware, and hardening our installation process.

In response to capacity issues in some regions (particularly in Europe), Kaz rolled out default per-organization capacity limits. These kinds of circuit-breaker limits are par for the course in public clouds, but we’re relatively new and had been getting away with not having them. We’re happy to let you scale up pretty much arbitrarily! But it’s to everyone’s benefit if we default to some kind of cap, because our DX makes it really easy to scale to lots of Fly Machines without thinking about it. Most capacity issues we’ve had over the last year have taken the form of “someone decided to spontaneously create 10,000 Fly Machines in one very specific region”, so this should be an impactful change.

We run a relatively large (for our age) hardware fleet, and we generate a lot of logs. We have a relatively large (for our age) logging cluster that absorbs all those logs. Well, now we absorb 80% less log traffic, because Tom spent a week using Vector (or, rather, holding it better) to parse and drop dumb and duplicative stuff.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor October 22 Orchestration Outage Postmortem

Anchor Narrative

At 14:00 EST on October 22, we experienced a fleetwide severe orchestration outage. What this means is that for the duration of the incident, both deployments of new applications and changes to existing applications were disrupted; during the most acute stage of the outage, lasting roughly an hour and 40 minutes, that disruption was almost total (Fly Machines could not be updated), and for roughly another 2 hours new application deployments did not function (but changes to existing applications did). Service was restored completely at 21:15 EST.

This outage proceeded through several phases. The earliest acute phase was the worst of it, and subsequent phases restored various functions of the platform, so that towards the end of the outage it was largely functional for most customers. At the same time, up until the end of of the outage, Fly.io’s orchestration layer was significantly disrupted. That makes this the longest significant outage we’ve recorded, not just on this infra-log but in the history of the company.

Orchestration is the process by which software from our customers gets translated to virtual machines running on Fly.io’s hardware. When you deploy an app on Fly.io, or when your Github Actions CI kicks off an update after you merge to your main branch, you communicate with our orchestration APIs to package your code as a container, ship it to our container registry, and arrange to have that container unpacked as a virtual machine on one or more of our worker servers.

You can broadly split our orchestration system into three pieces:

  1. flyd, our distributed process supervisor; flyd understands how to download a container, transform it into a block device with a Linux filesystem, boot up a hypervisor on that block device, connect it to the network, and keep track of the state of that hypervisor,

  2. the state sharing system; an individual flyd instance knows only about the Fly Machines running on its own host, by design, and it publishes events to a logically separate state sharing system so that other parts of our platform (most notably our Anycast routers and our API) know what’s running where, and

  3. our APIs, which allow customers to create, start, and stop new Fly Machines; our APIs are comprised of the Fly Machines API, which interacts directly with flyd to start and stop machines, and our GraphQL API, which is used to deploy new applications and manage existing applications.

The outage we experienced broke (2), our state sharing system, but had ripple effects that disrupted (1) and (3).

The outage was a cascading failure with a simple cause: a long-lived CA certificate for a deprecated state-sharing orchestration component expired. Our state sharing system is made up of two major parts:

  1. consul, or “State-Sharing Classic”, manages the configuration of our server components, registers available services on Fly Apps, and manages health checks for individual Fly Machines. consul is a Raft cluster of database servers that take updates from “agent” processes running on all our physical servers. consul used to be the heart of all our state-sharing, but was superseded 18 months ago, by

  2. corrosion, or “New State-Sharing”, tracks the state of every Fly Machine, and every available service, and service health. corrosion is a SWIM-gossip cluster that replicates a SQLite database across our fleet.

We began replacing consul with corrosion because of scaling issues as our fleet grew. It’s the nature of our service that every physical server needs, at least in theory, information about every app deployment, in order to route requests; this is what enables an edge in Sydney to handle requests for an app that’s only deployed in Frankfurt. consul can be deployed with regional Raft clusters, but not in a way that shares information automatically between those clusters. Since 2020, we’ve instead operated it in a single flat global cluster. Rather than do a lot of fussy in-house consul-specific engineering to make regional clusters work, we built our own state sharing system, wrapped around the dynamics of our orchestrator. This project is mostly complete.

What we have not completed is a complete severance of consul from the flyd component of our orchestrator. flyd still updates consul when Fly Machine events (like a start, stop, or create) occur. Those consul updates are slow, because consul doesn’t want to scale the way we’re holding it. But that doesn’t normally matter, because our “live” state-sharing updates come from corrosion, which normally has p95 update times around 1000ms. Still, some of these consul operations do need to complete, especially for Fly Machine creates.

consul runs over mTLS secure connections; that means everything that talks to it needs a CA certificate (to validate the consul server certificate) and a client certificate (to prove that it’s authorized to talk to consul).

At around 14:00EST on the day of the outage, consul’s CA certificate expired. Every component of our fleet which had any dependence on consul stopped working immediately. Fly Machine creates (but not starts and stops) depend on consul, as does some of our telemetry and internal fleet deployment capability. flyctl deploy stopped working.

To resolve this problem, we need to re-key the entire fleet; a new CA certificate, new server certificates, and new client certificates. Complicating matters: our internal deployment system (fcm/fsh) relies on consul to track available physical servers. All told, it takes us about 45 minutes to restore enough connectivity to deploys, and another 45 minutes to completely rekey the fleet.

At this point, basic Fly Machines API operations are completing. But there’s another problem: vault, our old secret storage system, is still used for managing disk encryption secrets, and by our API. The fleet rekeying has broken connectivity to vault. Complicating matters further, vault has extraordinarily high uptime, and so when its configuration is updated and the service is bounced, it doesn’t come back cleanly. For about 90 minutes, a team of infra engineers works to diagnose the problem, which turns out to be a different set of certificates that have expired; we’re able to perform X.509 surgery to restore them without rekeying another cluster.

The biggest problem in the outage (in terms of difficulty, if not raw impact) now emerges. During the window in which consul was completely offline, flyd has been queueing state updates and retrying them on an exponential backoff timer. These updates can’t complete until consul is back online, but all of them are events that corrosion consumes. They pile up, rapidly and dramatically.

By the time consul is restored, corrosion is driving 150gB/s of traffic, saturating switch links with our upstream. The data it’s trying to ship is mostly worthless, but it doesn’t know that. It’s a distributed system based on gossip, so it’s not simple to filter out the garbage.

For 6.5 hours, through the acute phase (in which deploys of existing apps aren’t functioning) and subacute phase (in which deploys of new apps aren’t functioning reliably), this will be the major problem we contend with. We need corrosion in order to inform our Anycast proxies of which Fly Machines are available to route traffic to. During the subacute phase of the outage, routing to existing Fly Machines continues functions, but changes to Machines take forever to propagate: at the beginning of the subacute phase, as much as 30 minutes; by the end, P99 latencies of several minutes — still far too slow for real-time Anycast routing.

Ultimately, the decision is made to restore the corrosion cluster from a snapshot (we made snapshots daily), and fill in the gaps (“reseeding” the cluster) from source-of-truth data. This process begins at 18:00 EST and completes by 18:30 EST, at which time P99 latencies for corrosion are back under 2000ms.

At this point, orchestration is almost fully functional. We have one remaining problem: deploys for apps that involve creating new volumes (which includes most new apps) fail, because our GraphQL API server needs to talk to consul to complete them (and only them), and it’s disconnected due to the rekeying.

Now that corrosion is stabilized, we’re able to safely redeploy the API server. The deployment hits a snag, which results in HTTP 500 errors from the API for about 20 minutes, at which point we’ve successfully redeployed, restoring the API.

Minutes later, with no known disruption or instability in the platform, the incident is put into “Monitoring” mode.

Anchor Incident Timeline

Anchor Forward-Looking Statements

The simplest thing to observe here is that it shouldn’t have been possible for us to approach the expiration time of a load-bearing internal certificate without warnings and escalations. Ordinarily we think about situations like this in terms of proximal causes (“we were missing a critical piece of alerting”) and some root cause; here, it’s more useful to look at two distinct root causes.

The first cause of this incident was that our infra team was overscheduled on project work. For the past year, we’ve pursued a blended ops/development strategy, with increasing responsibility inside the infra team for platform component development. If you follow the infra log, you’ve seen a lot of that work, especially with the Fly Machine volume migration work, which was largely completed by infra team members. We have developed a tendency to think about reliability work in terms of big projects with reliability payoffs. That makes sense, but needs to be balanced with nuts-and-bolts ops work. The “fix” for this problem will be denominated in big-ticket infrastructure projects we defer into next year to make room for old-fashioned systems and network ops work.

The second cause of this incident is a system architecture decision.

We shipped the first iteration of the Fly.io platform on a single global Consul cluster. That made sense at the time. It took years for us to scale to a point where Consul became problematic. When we approached that point, we had a decision to make:

The latter decision was sensible: we could make Consul scale, but making it fast enough for real-time routing to Fly Machines that start in under 200ms was challenging; a new, gossip-based system, taking advantage of architectural features that eliminated the need for distributed consensus, would make it much easier to address that challenge.

Unfortunately, we chose a half-measure. We replaced Consul with corrosion, but we retained Consul dependencies inside of our orchestration system, using Consul as a kind of backstop for corrosion, and keeping data in Consul relatively fresh so that old components could continue using it. Consul inevitably became a dusty old corner in our architecture, and so nobody was up nights worrying about managing it. Thus, our longest-ever outage.

The moral of the story is, no more half-measures. Work has already begun to completely sever Consul from our orchestration (we’ll still be using it for what it’s good at, which is managing configuration information across our fleet; Consul is great, it’s just not meant to be run, in its default configuration, for a half-million applications globally).

Finally, you may notice from the timeline that it took an odd amount of time to pull the trigger on restoring and reseeding the corrosion cluster, especially since once we did so, the process was completed in just 30 minutes. Restoring corrosion was straightforward because we have a tested runbook for doing so. But that runbook doesn’t have higher-level process information about when to restore corrosion, and what service impact to expect when doing so. If we’d had that information ready, we could have decided to perform the restore much earlier, shaving potentially 4 hours off the disruption.

,


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Stuff got done, but to generate these updates, the author of the infra-log needs to go interview infra people 1:1, and infra is heads-down responding to the incident from this week to foreclose on something like it happening again; some of that work is the same as the work we’re doing responding to the August Anycast routing incident (also a state explosion problem, also addressible by regionalizing state propagation and distributing aggregates globally instead of fine-grained updates), some of it isn’t; we’ll write more about it next week.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Let’s make up for some lost time.

Peter has deployed the first stage of the “sharding” of fly-proxy, the engine of our Anycast request routing system. Recall from our September 1 Anycast outage that one major identified problem was that we run a global, flat, unsegmented topology of proxies; as a result, a control-plane outage is as likely to disrupt the entire fleet as it is to disrupt a single proxy. We’re pursuing two strategies to address that: regional segmentation, which limits the propagation of control-plane updates (in potentially somewhat the same fashion as an OSPF area does) and sharding of instances. Sharding here means that, within a single region and on a single edge physical server, we run multiple instances of the proxy.

The first stage of making that happen is to add a layer of indirection between the kernel network stack and our proxy; that layer, the fly-proxy-acceptor, picks up incoming TCP connections from the kernel, and then routes them to particular instances of the “full” proxy using a Unix domain socket and file descriptor passing. This allows us to add and remove proxy instances without reconfiguring or contending for the same network ports. In the early stages of deployment, both the proxy-acceptor and the proxy itself listen for TCP connections (meaning the acceptor can blow up, and we’ll continue to handle connections, though nothing has blown up yet).

Unix file descriptor passing is textbook Unix systems programming, literally, you can find it in the W. Richard Stevens books, but it’s surprisingly tricky to get right; for instance, connect and accept completion are separate events, and we have to be fastidious about which instances we route file descriptors to (the bug where you let two different proxies see the same request file descriptor is very problematic).

Peter, Dov, and Pavel have been in a protracted disagreement with systemd. From a few weeks back: Dov added systemd watchdog support to fly-proxy. Recall that the diagnosis of the September 1 outage involved us noticing that the entire proxy event loop had locked up (it was a mutex deadlock, that’s what happens in a deadlock). It shouldn’t have been possible for the proxy to lock up without us noticing, and now it can’t.

Anyways, Dov read the systemd source code as it relates to watchdogs to make sure that when the proxy entered a shutting down state, the watchdog would be disabled. Things seemed fine, but then alerts began firing every time we did a deploy; the watchdog was tripping while the proxy was doing its orderly shutdown. Peter discovered a bug in systemd: it assumes that signal handling and watchdog logic share a thread.

In our case they don’t, which created a race condition that triggered watchdogs right after the systemd unit went into stopping state, which caused systemd to re-enable the watchdog. We stopped preempting the watchdog task and let it run until proxy’s bitter end.

There was more. In some cases it can take greater than 10 seconds (our watchdog length) for the fly-proxy to exit, after our tokio::main is complete. Boom, watchdog kill. “Ok, fine, you got us!” we said to systemd, and simply disabled the watchdog at runtime when the watchdog task was preempted. This, finally, worked, and proxies would no longer get watchdog killed when shutting down.

Except that sometimes they did? Turns out that our few older-distro hosts (remember: we have up-to-date kernels everywhere, but not up-to-date distros; systemd is the one big problem with that) use a pretty old version of systemd. That systemd does not support disabling the watchdog at runtime. Peter landed what we hope is the final blow this week; instead of disabling the watchdog at runtime, he set it to a very large non-zero value. You may read further adventures of Peter, Dov, and Pavel in their battles with systemd next week.

Speaking of distro updates, Steve continues our steady march towards getting our whole fleet on a recent distro. He’s picked up where Ben left off a few weeks ago, testing and re-testing and re-re-testing our provisioning to ensure that swapping distros out from under our running workloads doesn’t confuse our orchestration; we now have something approaching unit/integration testing for our OS provisioning process.

Tom spent the week spiking alternative log infrastructure to replace ElasticSearch, with which we are now at our wit’s end. We’re generally pretty reliable at log ingestion with ES, but experience sporadic ES outages with log retrieval. What we’ve come to learn as a business is that our customers are less sanguine about log disruption than we are; what sometimes feels to us like secondary infrastructure reads as core platform health to them. That being the case, we can’t keep limping with the ES architecture we booted up in 2021.

Finally: a couple weeks ago, Daniel had an on-call shift, and was, like everyone working an on-call shift here, triggered by alerts about storage capacity issues; everybody on-call sees at least a couple of these. You check to make sure the host isn’t actually running out of space, clear the alert, and go back to sleep. Unless you’re Daniel.

Daniel has had it in for the way we track available volume storage since back when he shipped GPU support for Fly Machines. There are two big problems with the way we’ve been doing this: the first is, going back to 2021 when we first shipped volumes, the system of record for available storage has been the RDS database backing our GraphQL API; that’s a design that predates flyd and our move away from centralized resource scheduling. The second big problem is that flyd itself has erroneous logic for querying available storage in our LVM configuration (it pulls disk usage from the wrong LVM object, causing it to misreport available space.

The result of this situation is that we’ve been managing available storage, and, worse, storage resource scheduling (deciding which physical server to boot up new Fly Machines on) manually — and, not just manually, but largely in response to alerts, some of which are arriving in the middle of the night.

Daniel fixed the flyd resource calculation and surfaced it to our Fly Machines API service, starting in Sao Paolo, where our API storage tracking went from reporting an average 95% storage utilization across all our physicals to an average 5%. The change has since been rolled out fleetwide, and, in addition to reducing alert load, has drastically improved Fly Machine scheduling. In every region we now have significantly more headroom and, just as importantly, more diversity in our options for deploying new Machines.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

It’s coming! But the infra log author was late getting to this update and doesn’t want to put all the infra people on the spot, so we’re getting the update about our one incident this week up first.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Peter worked on restructuring the connection handling code in fly-proxy, the engine for our Anycast layer, to support process-based sharding of proxy instances. This is work responding to the September 1 Anycast outage; the proximate cause of that outage was a Rust concurrency bug, which we’ve now audited for, but the root cause was the fact that a single concurrency bug could create a fleetwide outage to begin with. Process-based sharding runs multiple instances of fly-proxy on every edge, spreading customer load across them, not for performance (the single-process fly-proxy is probably marginally more performant) but to reduce the blast radius of any given bug in the proxy.

Kaz is rolling out size-aware Fly Machine limits. Obviously (it may have been more obvious to you than to us), you can’t expose something like the Fly Machines API without some kind of circuit-breaker-style limits on the resources a single user can request). Our current limits are coarse-grained: N concurrent Fly Machines, regardless of size. Clearly, these limits should be expressed in terms of underlying resources — a shared-1x is a tiny fraction of a performance-16x. Getting this working has required us to rationalize and internally document the relationships between these scaling parameters. Most of our users will never notice this (especially if we do it well), but it should make it less likely that you’ll hit a limit and have to ask support to remove it.

Somtochi has continued working on Corrosion, the SWIM-gossip SQLite state-sharing system the proxy uses to route traffic. The net effect of Corrosion is a synchronized SQLite database that is currently available across our whole fleet of edges, workers, and gateway servers. We’re refactoring this architecture to reduce the number of machines that will keep Corrosion replicas, and allowing them to subscribe and track changes to Corrosion databases stored elsewhere (this allows us to deploy more edges, by reducing the compute and storage requirements for those hosts).

Will did a bunch of reliability and ergonomic work on fsh, our internal deployment tool; better DX for fsh means more reliable deploys of new code means fewer incidents. fsh now integrates with PagerDuty to abort deploys automatically if incidents occur during a deployment; it will also fail fast on errors (this is an issue on fleet-scale rollouts, where it can be hard to spot errors across hundreds of servers being updated); it now directly supports staged deploys (something we were hacking together with shell scripts previously) and stepwise concurrency (slow-start style, running a single deploy, then 2 on success, and so on).

Dusty is continuing his capacity planning work by integrating our business intelligence tools with our capacity dashboard, so we can reflect dollar costs and revenue into our technical capacity planning.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Will did a deep dive into I/O scheduling in our LHR clusters, after Tigris nudged us about performance/reliability issues in their FoundationDB cluster running in that region. Using metrics, system configuration, and statekeeping data, Will isolated Tigris’s workloads to a concentrated cluster of SSDs with a particular make/model, which we now know to have an iffy performance envelope for the kinds of work we do. The bigger problem wasn’t so much the drives as it was the scheduling we did: because Tigris created their series of volumes for this region in rapid succession, they all got scheduled on a small subset of our storage capacity in the region. Worse, a consequence of our scheduling algorithm was that the periodic snapshot backups of these volumes were all scheduled to fire in tandem, concentrating a large amount of I/O activity on a small number of drives (the graph of what was happening looked like a stable EKG). Scheduling improvements are a theme of the infra work that we’re doing right now, but this issue in particular surfaced an unexpected (and straightforward to fix) issue: we needed to be adding jitter to the timing of our backup jobs.

Ben is hip-deep in a fleetwide upgrade of our worker OS distributions. This is a tricky and annoying problem. Most of what runs on a worker physical for us is software we build and ship ourselves, in some cases several times a week. Beyond that, we have an established runbook for upgrading our OS kernels; we don’t run the distro-version-native kernel version anywhere (we have fussy eBPF dependencies, among other things). But the distro itself, which in particular sticks us with a specific version of systemd, is a gigantic pain to upgrade; we have consistent OS kernels across the fleet, but not consistent distro versions. That’s changing, but it’s a complicated process, involving workload migrations and, in some cases, reprovisioning servers, which surfaces fun bugs like “the semi-random identifiers we create for Fly Machines are influenced by the provenance of the worker physical on which it was created, meaning a reprovisioned host can cause Machine ID collisions”. That bug hasn’t happened, because Ben is auditing the stack for problems like that.

Dusty has been the capacity czar for the past couple months. As you can see from this week’s BOS outage, this is an important issue for us. We’ve integrated our scheduler state with system metrics and used that to create new capacity threshold numbers for the fleet, which now informs our provisioning; a lot of the guesswork has been taken out of where we’re shipping and racking new servers. We’ve shifted some capacity (ORD gave some servers to EWR, for instance), and reallocated a backlog of servers to different regions. We now have a runbook for provisioning new capacity in existing regions that uses a lightweight version of our machine migration (for apps without storage) to rebalance as capacity is added.

JP has been working on improving Fly Machine scheduling across the fleet, which was also implicated in an incident this week. We now have stricter placement logic that ensures multiple Fly Machines for the same app created concurrently are distributed across worker physicals. Our Machines API now handles some of the retry logic that our CLI, flyctl, was using before; unlike flyctl, which is open-source and doesn’t have APIs that can specifically place a Fly Machine on a specific worker physical, our API gateways do have visiblity into the physicals in a region, and we now take advantage of that.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

In the interests of getting this infra log update up in a timely fashion and also giving the infra log writer a break, we’re going to talk about this week in infra engineering… next week.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Simon is deep — perhaps approaching crush depth — on the volume storage problem of making Fly Machines create faster. The most expensive step in the process of bringing up a Fly Machine is preparing its root filesystem. Today, we rely on containerd for this. When a worker brings up a Fly Machine, it makes sure we’ve pulled the most recent version of the app’s container from our container registry into a local containerd, and then [“checks out” the container image](/content/blog/docker-without-docker/ ""/index.html) into a local block device. If we can replace this process with something faster, we can narrow the gap between starting a stopped Fly Machine (already ludicrously fast) and creating one (this can take double-digit seconds today). The general direction we’re exploring is pre-preparing root filesystems and storing them in fast object storage, and serving those images to newly-created Fly Machines over nbd. Essentially, this puts the [work we did on inter-host migration](/content/blog/machine-migrations/ ""/index.html) to work making the API faster.

JP made all our Machines API servers restart more gracefully, by replacing direct socket creation with systemd socket activation. Prior to this change, a redeployment of flaps, the Machines API server, would bounce hundreds of flaps instances across our fleet, causing dozens of API calls to fail as service ports briefly stopped listening. That’s ~fine, in that flyctl, our CLI server, knows to retry these calls, but obviously it still sucks. Delegating the task of maintaining the socket to systemd eliminates this problem, and improves our API reliability metrics.

Ben and JP began the process of getting [Machine migrations](/content/blog/machine-migrations/ ""/index.html) onto [flyd](/content/blog/carving-the-scheduler-out-of-our-orchestrator/ ""/index.html)’s v2 FSM implementation. What’s a v2 FSM? We’re glad you asked! Recall: flyd, our scheduler, is essentially a BoltDB event log of finite state machine steps; “start a existing Fly Machine” is an FSM, as is “create a Fly Machine” or “migrate it to another host”. v1 FSMs (which are all of our in-production flyd FSMs) are pretty minimal. v2 FSMs add observability and introspection to current steps, and tree structures of parent-child relationships to chain FSMs into power combos; they also have a callback API for failure recovery. This is all stuff inter-host migration can make good use of; with observable, coordinated, tracked migration as a first-class API on flyd, we can get more aggressive about using Machine migration to rebalance the fleet and to quickly and automatically react to hardware incidents.

We have a couple customers that make gigantic Machine create requests — many thousands at once. To make these kind of transactions more performant, we parallelized them in flyctl. But these parallel requests are evaluated by our scheduler in isolation, which has resulted in suboptimal placement of Machines (the two most common cases being multiple Machines for the same app scheduled unnecessarily on the same worker, and unbalanced regions with some lightly loaded and some heavily loaded workers; “Katamari scheduling”). Kaz and JP made fixes both to flyctl and to our scheduler backend to resolve this; in particular, we now explicitly manage a “placement ID”, tracked statefully across scheduling requests using memdb, that allows users of our APIs to spread workloads across our hardware.

In related news, Dusty has been working on improved capacity planning. The most distinctive thing about the Fly Machines API, as orchestration APIs go, is that we’re explicit about the possibility of Machine creation operations (new reservations of compute capacity) failing; the most obvious reason a Machine create might fail is that it’s directed to a region that’s at capacity. What we have learned over several years of providing this API is that customers are not as thrilled with the computer science elegance of this “check if it fails and try again elsewhere” service model as we are. So we’ve been moving heaven and, uh, Dusty to make sure this condition happens as rarely as possible. Dusty’s big project over the last week: integrating our existing host metrics (the capacity metrics you’d think about by default, like CPU and disk utilization, IOPS, etc) with Corrosion, our state tracking database. Exported host metrics are a high-fidelity view into what our hosts actually see, while Corrosion is a distilled view into what we are trying to schedule on hosts. We’ve now got Corrosion reflected into Grafana, which has enabled Dusty to build out a bunch of new capacity planning dashboards.

Dusty also moved half of our AMS region to new hardware; half the region to go!

Peter worked a support rotation. We schedule product engineers to multi-day tours of duty alongside our support engineers, which means watching incoming support requests and pitching in to help solve problems. Peter reports his most memorable support interaction was doing Postgres surgery for a customer who had blown out their WAL file by enabling archive_mode, which preserves WAL segments, without setting archive_command, giving Postgres no place to send the segments.

Tom continued his top-secret work that we can’t write about, except that to say this week it involved risk-based CPU priorities and Machine CPU utilization tracking.

Now, deep breath:

Anchor September 1 Routing Layer Outage Postmortem

(A less formal version of this postmortem was posted on our community site the day after the incident.)

Anchor Narrative

At 3:30PM EST on September 1, we experienced a fleetwide near-total request-routing outage. What this means is that for the duration of the incident, which was acute for roughly 40 minutes and chronically recurring for roughly another hour, apps hosted on Fly.io couldn’t receive requests from the Internet. This is a big deal; our most significant outage since the week we started the infra-log (in which we experienced roughly the same WireGuard mesh outage, which also totally disrupted request routing, twice in a single week). We record lots of incidents in this log, but very few of them disable the entire platform. This one did.

Request routing is the mechanism by which we accept connections from the Internet, figure out what protocol they’re using, match them to customer applications, find the nearest worker physical running that application, and shuttle the request over to that physical so customer code can handle it. Our request routing layer is broadly comprised of these four components:

  1. Anycast routing, which allows us to publish BGP4 updates to our upstreams in all our regions to attract traffic to the closest region.

  2. fly-proxy, our Anycast request router, in its “edge” configuration. In this configuration, fly-proxy works a lot like an application-layer version of an IP router: connections come in, the proxy consults a routing table, and forwards the request.

  3. That same fly-proxy code in its “backhaul” configuration, which cooperates with the edge proxy to bring up transports (usually HTTP/2) to efficiently relay requests from edges to customer VMs.

  4. Corrosion, our state propagation system. When a Fly Machine associate with a routable app starts or stops on a worker, flyd publishes an update to Corrosion, which is gossiped across our fleet; fly-proxy subscribes to Corrosion and uses the updates to build a routing table in parallel across all the thousands of proxy instances across our fleet.

During the September 1 outage, practically every instance of fly-proxy running across our fleet became nonresponsive.

Generally, platform/infrastructure components at Fly.io are designed to cleanly survive restarts, so that as a last resort during an incident we can attempt to restore service by doing a fleetwide bounce of some particular service. Bouncing fly-proxy is not that big a deal. We did that here, it restarted cleanly, and service was restored. Briefly. The fleet quickly locked back up again.

Our infra team continued applying the defibrillator paddles to fly-proxy while the proxy team diagnosed what was happening.

The critical clue, identified about 50 minutes into the incident, was that proxyctl, our internal CLI for managing the proxy, was hanging on stuck fly-proxy instances. There’s not a lot of mechanism in between proxyctl and the proxy core; if proxyctl isn’t working, fly-proxy is locked, not just slowly processing some database backlog or grinding through events. The team immediately and correctly guessed the proxy was deadlocked.

fly-proxy is written in Rust. If you’re a Rust programmer, the following code pattern may or may not be familiar to you, and you may taste pennies in your mouth seeing it:

Wrap text Copy to clipboard

    // RWLock self.load
    if let (Some(Load::Local(load))) = (&self.load.read().get(...)) {
        // do a bunch of stuff with `load`
    } else {
        self.init_for(...);
    }

An RWLock is a lock that can be taken multiple times concurrently for readers, but only exclusively during any attempt to write. An if let in Rust is an if-statement that succeeds if a pattern matches; here, if self.load.read().get() returns Some instance, rather than None; this is a Rust error checking idiom. In the success case, the result is available inside the success arm of the if let as load. The else arm fires if self.load.read().get() returns None.

The way this if let statement looks, it would appear that the lifetime of the read lock taken in attempting the success case is only the length of the success arm of the if statement, and that the lock is dropped if the else arm triggers. But that is not what happens in Rust. Rather: if let is syntactic sugar for this code:

Wrap text Copy to clipboard

    match &self.load.read().get() {
        Some(load) => { /* do a bunch of stuff with `load` */ },
        _ => {
            self.init_for(...);
        },
    }

It is clearer, in this de-sugared code block, that the read() lock taken spans the whole conditional, not just the success arm.

Unfortunately for us, buried a funcall below init_for() is an attempt to take a write lock. Deadlock.

This is a code pattern our team was already aware of, and this code was reviewed by two veteran Rust programmers on the team, before it was deployed, but neither spotted the bug, most likely because the conflicting write lock wasn’t lexically apparent in the PR diff.

The PR that introduced this bug had been deployed to production several days earlier. It introduced “virtual services”, which decouple request routing from Fly Apps. Conventionally-routed services on Fly.io are tied to apps; the fly.toml configuration for these apps “advertise” services connected to the app, which ultimately end up pushed through Corrosion into the proxy’s routing table. Virtual services enable flexible query patterns that match subsets of Fly Machines, by metadata labels, to specific URL paths. We’re generally psyched about virtual services, and they’re important for [FKS, the Fly Kubernetes Service](/content/blog/fks/ ""/index.html).

The deadlock code path occurs when a specific configuration of virtual service is received in a Corrosion update. That corner case had not occurred in our staging testing, or on production for several days after deployment, but on September 1 a customer testing out the feature managed to trigger it. When that happened, a poisonous configuration update was rapidly gossiped across our fleet, deadlocking every fly-proxy that saw it. Bouncing fly-proxy broke it out of the deadlock, but only long enough for it to witness another Corrosion subscription update poisoning its service catalog again. Distributed systems. Concurrency. They’re not as easy as computer science classes tell you they are.

Because we had a strong intuition this was a deadlock bug, and because it’s easy for us to isolate recent deployed changes to fly-proxy, and because this particular if let``RWLock bug is actually a known Rust foot-gun, we worked out what was happening reasonably quickly. We disabled the API endpoints that enabled users to create virtual services, and rolled out a proxy code fix, restoring services shortly thereafter.

Complicating the diagnosis of this incident was a corner case we ran into a with sysctl change we had made across the fleet. To improve graceful restarts of the proxy, we had applied tcp_migrate_req, which migrates requests across sockets in the same REUSEPORT group. Under certain circumstances with our code, this created a condition where the “backhaul” proxy stopped receiving incoming connection requests. This condition impacted only a very small fraction (roughly 10 servers total) of our physical fleet, and was easily resolved fleetwide by disabling the sysctl; it did slow down our diagnosis of the “real” problem, however.

Anchor Incident Timeline

Anchor Forward-Looking Statements

The most obvious issue to address here is the pattern of concurrency bug we experienced in the proxy codebase. Rust’s library design is intended to make it difficult to compile code with deadlocks, at least without those deadlocks being fairly obvious. This is a gap in those safety ergonomics, but an easy one to spot. In addition to code review guidelines (it is unlikely that another if let concurrency bug is going to make it through code review again soon), we’ve deployed semgrep across all our repositories; this is a straightforward thing to semgrep for.

The deeper problem is the fragility of fly-proxy. The “original sin” of this design is that it operates in a global, flat, unsegmented topology. This simplifies request routing and makes it easy to build the routing features our customers want, but it also increases the blast radius of certain kinds of bugs, particularly anything driven by state updates. We’re exploring multiple avenues of “sharding” fly-proxy across multiple instances, so that edges run parallel deployments of the proxy. Reducing the impact of an outage from all customers to 1/N customers would have simplified recovery and minimized the disruption caused by this incident, with potentially minimum added complexity.

One issue we ran into during this outage was internal dependencies on request routing; one basic isolation method we’re exploring is sharding off internal services, such as those needed to run alerting and observability.

fly-proxy is software, and all software is buggy. Most fly-proxy bugs don’t create fleetwide disasters; at worst, they panic the proxy, causing it to quickly restart and resume service alongside its redundant peers in a region. This outage was severe both because the proxy misbehavior was correlated, and because the proxy hung rather than panicking. A straightforward next step is to watchdog the proxy, so that our systems notice N-second periods during which the proxy is totally unresponsive. Given the proxy architecture and our experience managing this outage, watchdog restarts could function like a pacemaker, restoring some nominal level of service automatically while humans root-cause the problem. This was our first sustained correlated proxy lock-up, which is why we hadn’t already done that.

This incident sucked a lot. We regret the inconvenience it caused. Expect to read more about improvements to request routing resilience in infra-log updates to come.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

This was a holiday week during which we experienced a significant reliability incident that has the infra-log busily copyediting an in-progress postmortem, so the infra-log is giving itself (and the team) a break. More of the continuing adventures of Peter and Somtochi in the weeks to come. Thanks!


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Dusty put Fly.io-authored Ansible provisioning for our hardware upstream providers into motion, getting us close to capping off a project to streamline the provisioning of new physicals. Additionally, he set up IPMI sensor alerting across our fleet, so we can do early detection of hardware issues (our fleet is now of the size where hardware failures, while rare as a fraction of the fleet itself, aren’t totally out of the ordinary); we now have better early alerting for server physical issues, which is important because with the completion of the migration project we’re in a much better position to preemptively migrate workloads.

Somtochi is back into Corrosion, responding to incidents from a couple weeks ago. Changes include queue caps (one incident was caused by an unbounded queue of changes from nodes) and fixing a bug that was causing Corrosion to request way more data than it needs when syncing up a new node with the fleet. She also set up sampled otel telemetry for fly-proxy (fly-proxy telemetry is tricky because of the enormous volume of requests we handle).

Tom did important top-secret work that we are not in a position to share but will one day be very fun to talk about. Read between the lines. Also, the previous week, which Peter monopolized with the petertron4000, Tom did a bunch of Postgres monitoring work, because we’re gearing up to get a lot more serious about managing Postgres.

Peter mitigated a longstanding compatibility issue with us and Cloudflare; either our HTTP/2 implementation (which is just, like, Rust hyperium or something) or theirs is doing something broken; when we have problems we now automatically downgrade to HTTP/1.1 for their source IPs.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

August 15th was spicy; with the exception of the Elixir API issue, the problems were contained regionally or to a small minority of specific hosts, but it was an infra-intensive day.

Anchor This Week In Infra Engineering

Several infra people were out on vacation this week, and the rest of the team did interesting work that deserves a showcase, but we’re giving this week in infra engineering over to Peter; everybody else will get their due next week.

One of the features we offer to applications hosted on Fly.io are a dense set of Prometheus performance metrics, along with Grafana dashboards to render them. One of the things our users can monitor with these metrics is TLS handshake latency. Why do we expose TLS latency metrics? We don’t know. It was easy to do.

Anyways, a consequence of exposing TLS latency metrics is that customers can, if they’re watchful, notice spikes in TLS latency. And so someone did, and reported to us anomalously slow TLS handshakes originated from AWS, which set off several weeks of engineering effort to isolate the problem.

Two things complicated the analysis. First, while we could identify slower-than-usual latencies, we couldn’t (yet) isolate extreme latency. Second, packet-level analytics didn’t show any weird intra-handshake packet latency originating “from us”; the TCP 3WH completed quickly, our TLS ServerHello messages rapidly followed after ClientHello, etc.

What we did notice were specific clients that appeared to have large penalties on “their side” inside the TLS handshake; a delay of up to half a second between the TCP ACK for our ServerHello and the ChangeCipherSpec message that would move the handshake forward. Reverse DNS for all these events traced back to AWS. This was interesting evidence of something, but doesn’t dispose of the question of whether there are events causing our own Anycast proxies to lag.

To better catch and distinguish these cases, what we needed was a more fine-grained breakdown of the handshake. A global tcpdump would produce a huge pcap file on each of our edges within within seconds, and even if we could do that, it would be painful to filter out connections with slow TLS handshakes from those files.

Instead, Peter ended up (1) making fly-proxy log all slow handshakes; and (2) writing a tool called petertron4000 (real infra-log heads may remember its progenitor, petertron3000) to temporarily save all packets in a fixed-size ring buffer, ingest the slow-handshake log events from systemd-journald, and (3) write durable logs of the flows corresponding to those slow handshakes.

With petertron4000 in hand, he was able to spot check a random selection of edges. The outcome:

Peter was able to reproduce the problem directly on EC2: running 8 concurrent curl requests on a t2.medium reliably triggers it. Accompanying the TLS lag: high CPU usage on the instance. Further corroborating it: we can get the same thing to happen curl-ing Google.

This was an investigation set in motion by an ugly looking graph in the metrics we expose, and not from any “organically observed” performance issues; a customer saw a spike and wanted to know what it was. We’re pretty sure we know now: during the times of those spikes, they were getting requests from overloaded EC2 machines, which slow-played a TLS handshake and made our metrics look gross. :shaking-fist-emoji:.

That said: by dragnetting slow TLS handshakes, we still uncovered a bunch of pessimal routing issues, and we’re on a quest this summer to shake all of those out. The big question we’re working on now: what’s the best way to roll something like the petertron4000 (the petertron4500) so that it can run semipermanently and alert for us? Fun engineering problem! We definitely don’t want to be doing high-volume low-level packet capture work indefinitely on all our edges, right? Stay tuned.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Akshit began a project to diversify our edge providers and edge routing. Recall that our production infra is broadly divided into edge hosts (that receive traffic from the Internet, terminate TLS, and route it) and worker hosts (that actually host Fly Machine VMs). We have more flexibility on which providers and datacenters we can run edge hosts on, because they don’t require chonky servers. Akshit is working out the processes and network engineering (like per-region-provider IP blocks) required for us to take better advantage of available host and network inventory for edges. Ideally, we’ll wrap this project up with same-region backup routing (via different providers) in our key regions.

Peter spent the week sick. He wants you to know he feels better now.

Steve is working on rehabilitating RAID10 hosts. This is a beast of an issue that has been taunting us since late 2022: we took delivery of a bunch of extremely chonky worker servers that would handle our workloads just fine for a period, and then lock up in unkillable LVM threads. We solved those problems for customers by migrating workloads off those machines, and now Steve is doing the storage subsystem brain surgery required to find out if we can bring them back into service.

Somtochi has moved from Pet Semetary (our Vault replacement, which she got deployed fleetwide) to Corrosion2 (which she drastically improved the performance and resiliency of) to fly-proxy, the engine for our Anycast network. She’s picking up where Peter left off, with backup routing, by extending it to raw TCP connections and not just HTTP (reminder: if it speaks TCP on any port, you can run it here.)

Dusty is working with one of our upstream hardware providers to get us end-to-end control over machine provisioning, rather than having them hand off physicals with BMC connections for us to provision. Faster, tighter physical host provisioning means we can bring up capacity in regions more quickly; we’ve been O.K. there to date, but that leads us to

John is working on “infinite capacity” burst provisioning processes, which is a shorthand for “you can ask us for 1000 16G Fly Machines” (it has happened) “and we will just automatically spin up the underlying hardware capacity needed to satisfy that request”. We’re a ways off on this but expect it to be a theme of our updates (if it pans out). Again, this is primarily of interest to people who expect to have sudden or sporadic needs for large numbers of Fly Machines in very specific places.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Akshit rolled out opt-in granular bandwidth billing. The new bandwidth billing scheme saves most of our customers money (especially if you make good use of our private networks, for instance by running highly utilized Postgres clusters), but, because it can end up costing a bit extra for users that don’t use private networks are are deploying in expensive regions (most notably India), it’s opt-in for existing customers. This work involved working through bugs with our upstream billing partner; Akshit has our sympathies. Akshit was also part of the response to the July 29 GQL outage, which meant they spent a chunk of this week reworking parts of our GQL server CI/CD system.

Steve got cross-connects deployed between Fly.io and Oracle Cloud, in order to accelerate object storage for our partners; objects stored off-network should no longer traverse the public Internet.

Andres and Matt improved synthetic monitoring (we built a new synthetic monitoring system a few weeks back), notably by creating and deploying new reference apps for us to measure. We have improved visibility into behavior we weren’t directly alerting on, like obviously-broken routing (think Asia->Europe->Asia). Synthetics surfaces some fly-proxy bugs, which got fixed. We flirted with making synthetics a customer-visible feature and decided we hadn’t worked out the privacy issues yet.

Ben, Dusty, Steve and John continued migrating workloads from old servers to newly provisioned ones; this involved building out more migration tooling, fixing bugs in migration tooling, and wrestling with particularly persnickity physical servers in Asia. We are asked to relay the following: “Servers go out. Servers come in. We are thus ever trapped in samsara”. Ok then.

Peter rolled out fallback routing fleetwide. He writes it up better than I do, as usual. In addition to metrics-based fallback routing, we now have rule-based routing that takes known backbone topology issues into account. Peter also resolved the LVM2 metadata issues we talked about a few weeks ago, and is deep into debugging (very) sporadic TLS handshake time delays.

Kaz updated and simplified the public status page, which now does a better job of answering the most important question (is the problem my app, or something going wrong at Fly.io?).


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

The Infra Log author had the flu this (mercifully uneventful) week, so let’s just do this week and the next together in one block.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

One thing you’re starting to see now is that Fly Machine migration and host draining is ironed out enough to be a casual solution to problems that would not have been casually resolved a year ago; “bring up new capacity, move the noisy workloads there” is a no-escalation runbook for us now. See the last 10 infra-logs for some of the effort that it took us to get to this point.

Akshit shipped a new egress billing model, which applies only to new customers for now. Under the new scheme, we segregate egress bandwidth by region in invoices, and private networking (between Fly Machines in different regions) is now cheaper. Our product team shipped a new billing system last month, and billing improvements are likely to be a continuing theme of our work.

Andres continued improving our internal synthetics monitoring systems.

Ben fixed several Fly Machine migration bugs: migration RPCs were breaking the configuration for static assets (we can serve static file HTTP directly off our proxies without bouncing HTTP requests off your Fly Machines, if you ask us to); we had a coordination bug in one of the FSMs our orchestrator flyd uses to migrate volumes; and high-availability Postgres cluster migrations were made less tricky (we do these by hand currently, for reasons we’re blogging about this week).

Matt shipped alert-critic, a chat-ops service that monitors our busiest alerting channels and tracks first-responder satisfaction with those alerts, in order to generate reports that spotlight problematic alerts that are either poorly reviewed or that don’t end up needing responses at all.

Peter generated network telemetry data to inform a fleetwide rollout of fly-proxy fallback routing, which routes requests through our overlay network during periods of network instability, at the application layer, automatically. This was deployed in Singapore last week, and is deployed more widely this week.

Tom overhauled our alerting layer for internal server health check alerts; we have hundreds of these, and they currently route directly from our health check system to our on-call system (and thus, ultimately, PagerDuty). We’ve scaled past our alert system’s ability to reliably alert (for very ambitious values of “reliably”). The new alert system routes through Vector, like the rest of our logs and telemetry, and fires alerts from Grafana; both these systems are used for customer workloads and were built to scale, unlike our internal server health check system.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Will shipped bottomless storage volumes backed by Tigris. This is big! Last fall, Matthew Ingwersen announced log-structured virtual disks that cache blocks while writing them to object storage for durability — the net effect is a “bottomless volume” that is continuously in snapshotted state. The tradeoff for this is, you had to write them to off-network object storage, like S3, which adds an order of magnitude latency to uncached blocks. Tigris is S3-flavored object storage that is both directly attached to Fly.io and also localized to the regions we operate in, which drastically improves performance. It’s early days yet, and this feature is experimental, but we’d like to get this tuned well enough to be a sane default choice for general-purpose storage.

Andres shipped a first cut of a new synthetic monitoring system (“synthetics” is the cool-kid way of saying “actually making requests and seeing if they complete”, as opposed to watching metrics). We had some synthetic monitoring, but now we have substantially more, broken out into regions, particularly for the APIs reachable from flyctl, our CLI.

Akshit and Steve worked on internal bandwidth tracking, in part to support the egress pricing work Akshit talked about a few weeks back. Steve’s work gives us improved visibility for our own internal traffic between all pairs of servers, regions, and data centers.

John worked on our continuing theme of migrating from and decommissioning older hardware, and, in the process, resolved a gnarly problem with LVM2 metadata stores running near capacity. LVM2 is the userland correspondant to devicemapper, the kernel’s block storage framework; if you think of LVM2 and devicemapper together as an implementation of a software RAID controller, you’re not far off. LVM2 virtualizes block storage devices on top of physical devices, and reserves space on each physical to track metadata about which sectors are being used where; if space runs out, all hell breaks loose, and extending metadata space is tricky to do, but is much less tricky now. This is one of these random backend infra engineering problems that make migrations tricky (to balance workloads between servers and migrate off old servers, you sometimes want to migrate jobs to places where there’s LVM2 metadata pressure) which, once solved, makes it much easier for us to migrate jobs without ceremony. Maybe you have to have dealt with LVM2 PV metadata issues for them to be as interesting to you as they are to us. We’ll shut up now.

Dusty is on a top-secret mission to increase the speed of OCI image pulls from containerd. Recall: you deploy, and push a Docker image to our registry. Then, a worker server, running containerd, pulls that image from the registry into its own local storage, and converts it to a block storage device we can boot a VM on. That containerd image pull is the dominant factor in how long it takes to create a Fly Machine, and we’re like create to be asymptotically as fast as start (which is so fast you can start a Fly Machine to handle an incoming HTTP request on the fly).

Peter shipped fallback routing in fly-proxy, and we can’t write it up any better than he did, so go follow that link.

Tom did a bunch of anti-abuse stuff we’re not allowed to talk about. In lieu of a fun writeup of the anti-abuse stuff Tom did, we’ve instead been asked to describe the on-call drama that kept him busy for much of the week:


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

The 4th of July hit on a Thursday this year, making this an extended holiday weekend for a big chunk of our team.

The big stories this week are mostly the same as last week. We continued deploying and ironing out bugs in Corrosion record compaction, we migrated off a bunch of old physical servers and continued building out migration tooling to make it even easier to drain workloads from arbitrary servers, and we improved incident alerting for customers in our UI and in flyctl.

The most important work that happened in this abbreviated week was all internal process stuff: we roadmapped out the next 12 weeks of infra work for networking, block storage, observability, hardware provisioning, and Corrosion. Lots of new projects hitting, which we’ll be talking about in upcoming posts.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Somtochi rolled out a major change to the way we track distributed state with Corrosion. Because Corrosion is a distributed system (based principally on SWIM gossip) and no distributed system is without sin, we have to carefully watch the amount of data it consumes; updates are relatively easy to propagate, but eliminating space for old, overridden data is difficult; this is the “compaction” problem. Somtochi and Jerome worked out a straightforward scheme for doing compaction, but it required adding an index to a table that had been growing without bound for many months, and would potentially trigger multi-minute startup lags everywhere Corrosion needed to get reinstalled. Instead of doing that, we “re-seeded” Corrosion, taking a known-good dataset from one of our nodes, compacted, and then using it as the basis for new Corrosion databases. This was rolled out on many hundreds of hosts without event, and on a small number of edge servers (which have much slower disks) with some events, which you just read about above.

Akshit worked on improving the metrics we’re using for bandwidth billing, putting us in a position to true up bandwidth accounting by more carefully tracking inter-region (like, Virginia to Frankfurt) traffic, especially for users with app clusters where only some of the apps have public addresses. You’ll hear more about this from us! This is just the infra side of that work.

After Peter wrote a brief postmortem of an incident from last week, Ben Ang worked out a system to more carefully track deployments of internal components, especially when those deployments happen piecemeal as opposed to full-system redeploys. Since the first question you ask when you’re handling an incident is “what changed”, anything that gives us quicker answers also gives us shorter incidents.

Dusty, John, Simon, and Peter all worked on draining old servers, migrating Fly Machines to newer, faster hardware. This is all we’ve been talking about here for the last month or so, and it’s happening at scale now.

Andres got tipped off by an I/O performance complaint on a Mumbai worker and ended up tracking down a small network of crypto miners. The hosting business; how do you not love it? Andres did other stuff this week, too, but this was the only one that was fun to write about.

Will wrapped up his NATS log-shipping work. We’ll let him tell the story.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Intra-region host migrations are unblocked again! This is huge for us.

Peter worked with our upstream providers to eliminate pathological AS-path routes impacted by recent APAC undersea cable cuts. This work started with us noticing relatively high packet loss in Asia regions, and resulted in us drastically reducing timeouts in our own telemetry and tooling, and network quality for users. A very big win that we’re looking to compound with better monitoring and tooling. He also figured out a configuration bug that was causing Fly Machines not to use BBR congestion control on private networking traffic, which is now fixed.

Dusty and Matt got all our multi-node Postgres clusters in condition to migrate (recall: multi-node Postgres clusters had been problematic for us, because they were configured to use literal IPv6 addresses for their peer configurations, and migration breaks those addresses, which embed routing information).

In addition to spending 30 working hours getting a single email announcement (about migrations) out to customers, John shipped our 6PN address forwarding tooling, along with Ben, out to the fleet, making it possible to migrate clusters that refer to literal IPv6 addresses. Dusty, Peter, John and Matt began draining hosts, moving the Machines running on them to most stable, modern, resilient systems on better upstreams, and lining us up to decom the much older machines. Ben drained an old server live during our internal Town Hall meeting. It was an emotional moment.

Still a bunch of people out this week! It’s summer (for most of us)!


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Can’t complain too much. There may come a day when we are large enough not to experience transient failures somewhere in the world, but that day is not this day. Two things we’re working aggressively on:

Anchor This Week In Infra Engineering

Bunch of people out this week! It’s summer (for most of us)!

Andres shipped a long-overdue feature for flyctl: if you run a flyctl command that involves some physical host on our platform (most commonly: the worker server your Machine is on), we’ll warn you if we’re currently dealing with an issue on that host. We’ve had these notices in the UI for a bit, and Andres recently shipped email alerts for any host drama that impacts your Fly Machines, but we suspect this might be the more important reporting channel, since so many of our users are CLI users.

Ben integrated some work from Saleem on our ProdSec team that, during a Fly Machine migration, makes the original Machine’s 6PN address still appear to work for other members of the same network. Recall: our 6PN private network feature works under the hood by embedding routing information into IPv6 addresses; moving a Machine from one physical worker to another breaks that routing. This is only a problem for a small subset of apps that embed literal IPv6 addresses in their configurations. Saleem’s work applies network address translation during and after migrations; Ben’s work links this capability into Corrosion, our global state sharing system, to keep everyone’s Machine updated.

Peter is working on stalking cluster apps people have deployed that use statically-configured 6PN addresses, and thus need the mitigation Ben is working on. He’s doing that by detecting connections that originate prior to DNS lookups, and tracking them in SQLite databases, using a tool we call petertron3000.

Akshit and Ben did a bunch of work this week updating and improving metrics, for internal vs. edge traffic, [FlyCast](/content/docs/networking/private-networking/#flycast-private-fly-proxy-services ""/index.html) traffic, gateways, and flyd. Ben also caught and fixed some flyd migration bugs.

Kaz did a bunch of bug fixing and ops work in the background, but this week we’ll call out the stuff he’s been doing with customer comms, in particular this Machine Create success rate metric on our public status page, which is now much more accurate.

Simon did some rocket surgery on flyd to ensure that applications that are migrated with multiple deployed instances are migrated serially rather than concurrently, to eliminate corner cases in distributed applications.

Steve spent some time talking to Oracle about cross connects, because we have users and partners that want especially fast and reliable connectivity to Oracle OCI. So that’ll happen.

Steve also spent a bunch of time this week refactoring parts of fcm, our bespoke, Bourne Shell based physical host provisioning tool, so that it can be run from arbitrary production hosts rather than the specially designated host that it runs from now. I mean, it can’t be, not yet, but we’re… steps closer to that? We don’t know why he did this work. Sometimes people just get nerd sniped. This page is all about transparency, and Steve is this week’s designated Victim Of Transparency.

Will is working with Shaun on our platform team on a volumes project so awesome that we don’t want to spoil it yet. (Similarly: Somtochi is still working on the huge Corrosion project she was working on last week which is also such a big deal you won’t hear about it until it ships or fails).


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

This week’s series of small regional incidents kept the infra team hopping.

Apart from incident response, this week’s work looked a lot like the last week’s. Rather than break it out by person, we’ll just document the themes:

Physical host migration remains the biggest ticket item for the infra team. We’re pushing forward on decommissioning old Equinix data center deployments and moving them to newer, more resilient, more cost-effective hardware. The big obstacle we’re facing right now remains applications that may (sometimes surreptitiously) be saving and reusing literal 6PN IPv6 private networking addresses, rather than DNS names. Because 6PN addresses are bound to specific physical hardware, these apps may break when migrated, which isn’t acceptable. We’re doing lots of things, from careful manual migration of apps (like Fly Postgres) where we control the cluster, to alerting and eBPF-based fixes. We knocked out another dozen or two old physicals this week.

Better host status alerting is a big deal for us. We’re going to keep seeing regional and host-local outages, which is just the nature of running a large fleet of physical servers. We’re now doing email alerting for customers on impacted hosts, and have PRs in for displaying alerts whenever a user touches an app impacted by a host alert with flyctl to continue closing that loop.

Corrosion scalability and reliability work continues; Somtochi has some design changes that could further minimize the amount of state we have to share, which we’ll talk about more when they pay off. Corrosion is a super important service in our infrastructure (it’s the basis for our request routing) and reliability improvements since infra took it over have been a big win.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

Anchor This Week In Infra Engineering

Short week. Couple people out sick.

Kaz worked on getting Fly Machine creation success rates onto our status page, which you should see soon. The two most important things you can know about the Fly Machines API: “create” and “start” are two different operations (“start” is the fast one; you can pre-“create” a bunch of stopped machines and start them whenever you need them), and “create” can fail; for instance, you can ask for more resources than are available in the region you target. [Read more about that here](/content/blog/carving-the-scheduler-out-of-our-orchestrator/ ""/index.html). We (well, Kaz, but we agree with him) want the success rates for this operation to be visible to customers.

Dusty and Simon spent the week heads-down on Postgres cluster migration. Read last week’s bulletin for more on that. We’re getting somewhere, but we’re not done until we can push a button and safely clear all the Machines of a physical server without having to worry too much about it.

Will won his next boss battle with NATS. We’ve successfully upgraded the whole fleet to current NATS (recall: the last attempt drove a terabit-scale message storm), on a custom branch with some of his fixes from last week. Metrics are down up to 90% across the board (a good thing) and problems we’ve been having with connection stability after network outages (inevitable at our scale!) seem to have resolved. Will’s writing a Fresh Produce release about this and we won’t steal any more of his thunder here.

Matt spent the week making log monitoring more resilient. “Logs” here mean “the platform feature we offer that ships logs off physical servers and to customers using NATS”. What Matt’s doing is, we run a Machine on every physical server in our fleet, the “debug app”, and it checks various things and freaks out and generates alerts when things go wrong. One more thing “debug” does now is track our server inventory, and make sure we’re getting NATS logs from all of them. In other words, another constantly-running, all-points end-to-end test of log shipping, from the vantage point of our customers.

Tom is doing topological work on Corrosion. As we keep saying, we have “edge” servers and “worker” servers; the “edges” are much, much smaller than the “workers”, and we don’t want to tax them too much, so they can just do their thing terminating TLS and routing traffic. But that routing function depends on Corrosion, our gossip-based state tracking system, and Corrosion is expensive. One answer, which Tom is pursuing, is for (most) edges not to run it at all, but instead to be remote clients of it on other machines.

Dave (and Matt and Will and Simon) did a bunch of hiring work, including revamping our challenges and updating our internal processes for reviewing them. We should be much more responsive to infra candidates (we already were within tolerances, but we’re raising the bar for ourselves).


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

So, yeah, pretty easy week, as far as the infra team is concerned.

Anchor This Week In Infra Engineering

The big news for the past several weeks has been intra-region Fly Machine migration: minimal downtown migration of workloads, including large volumes, from one physical worker server to another. We hit a snag here: Fly Postgres wasn’t designed originally to be migrated, so many instances of it are intolerant to being moved and booted up on new 6PN IPv6 addresses. A bunch of work is happening to resolve this; we’ll certainly be migrating Fly Postgres instances in the near future, it’s just a question of “how”.

Simon is designing migration tooling work to make large-scale migration and host draining work for everything we can safely migrate, including single-node Postgres instances ( do not run single-node Postgres instances in production! — but thanks for being easy to migrate).

Ben A. worked on migrating and draining workloads to balance workers. Fly Machines bill customers primarily for the time the Machine is actually running. When they’re stopped, a Machine is a commitment to some amount of resources on its associated worker, and a promise that we will start that Machine within some n-hundred millisecond time budget. This commitment/promise dance is drastically simpler and less expensive for us to honor now that we can migrate stopped Machines.

Now that Pet Semetary is up and running, Somtochi has switched up and is working with Sage on Corrosion, our Rust-based gossip statekeeping system. Corrosion is a (large) SQLite database managed by SWIM gossip updates. The work this week is primarily testing and bugfixing, but they may have figured out a way to reduce the size of our database by a factor of 3, which we’ll certainly write about next week if it pans out.

Dusty and Sage have been adding more edge hosts to keep up with capacity. Dusty also began trial migrations of Fly Postgres, using an IP mapping hack by Saleem, and built some internal dashboards to assist in the sort of manual host rebalancing work that Ben A. was doing this week.

Akshit cleaned up some log messages. Yawn. But also he graduated university! Congratulations to Akshit.

John rolled out a fleetwide fix for an interaction between Corrosion and our eBPF UDP forwarding path. You can run a fully-functional DNS server as a Fly Machine, because our Anycast network handles UDP as well as TCP. We do this by transparently encapsulating and routing UDP in the kernel using XDP and TC BPF. The routing logic for this scheme is written into BPF maps by a process (udpcatalogd) that subscribes to Corrosion. We decommissioned a large physical worker in AMS, which generated a big Corrosion update, which tickled a bug in a particular SQL query pattern only udpcatalogd uses, which caused ghost services for that AMS ex-worker to get stuck in our routing maps. There were a bunch of fixes for this, but the immediate thing that cleared the problem operationally involved… turning udpcatalogd off for a moment and then back on, fleetwide. Thank, John! John also did a bunch of retrospective work on our learnings abouting Fly Postgres clusters. He’s also taken up residency on our public community site. Go say ‘hi’ to him (or complain about something that infra is involved in, and he’ll apparently show up.)

Will has been heads-down in NATS land for the past 2 weeks, after an attempted NATS upgrade briefly melted our network a bit over a week ago. Will has been chatting with the Synadia people about our topology, and, in the meantime, found two scaling issues in the NATS core code that drive excess system chatter in our current topology. He’s prepared a couple upstream PRs.

Steve spent some time this week building new features for Drift, our Elixir/Phoenix internal server hardware inventory tool. Drift now tracks server lifecycle (for things like decommissioning), and, where our upstreams support it, automatically adding new servers to our inventory.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

This was an easier week than last week. The middle outage, migrating Postgres clusters, was very noticeable to impacted customers — but also quick root-caused. The other two incidents were limited in scope. Unless you’re carefully watching [Fly Metrics](/content/docs/metrics-and-logs/metrics/ ""/index.html). Are you using [Fly Metrics](/content/docs/metrics-and-logs/metrics/ ""/index.html)? An incident that broke them is a weird place to pitch them, we know, but they’re pretty neat and you get them for free. Hold our feet to the fire on them being reliable!

Anchor This Week In Infra Engineering

Dusty provisioned new hardware capacity in San Jose, Singapore, Warsaw, Sydney, Atlanta, and Seattle.

Will had a conversation with engineers at Synadia (last week’s NATS outage hit right during an all-hands meeting for them!) and got some advice on reconfiguring our internal NATS topology, shifting most of our hosts to “leaf” nodes and minimizing the number of “clustering” notes we have per region; this should trade an imperceptible amount of latency (which doesn’t matter with our NATS use case) for drastically reduced chatter. Thanks, Synadians!

Akshit finished an upgrade to [Firecracker](/content/blog/sandboxing-and-workload-isolation/#firecracker ""/index.html) 1.7 across our fleet. 1.7 does asynchronous block I/O with io_uring. We’ve noticed, since we rolled out Cloud Hypervisor for our GPU workloads (ask us about the security work we had to do here!) that Cloud Hypervisor was doing a better job handling busy disks than the version of Firecracker we were running. We’re optimistic that the new version will close the gap.

Steve finished up the provisioning tooling for the fou-tunnel-and-SNAT monstrosity that we talked about last week: giving Fly Machines static IP addresses, for people who talk to IP-restricted 3rd party APIs.

Tom and Ben A (ask us how many Bens work here!) completed the migration and draining of workloads from the cursed “edge worker” machines we mentioned last week. Edge-workers are no more. In the process, Tom debugging a bunch of draining tooling issues (being good at this is a big deal, because we’d like to be able to drain a sus server anywhere in the world at the drop of a hat), and Ben wrote up internal playbooks for draining hosts. Requiescat, Edge Workers, 2020-2024.

Simon continued low-level work on Machine/Volume migration, which is the platform kernel of the host draining stuff Tom and Ben were doing. This week’s work focused on large volume migration. Recall that our migration system causes the “source” physical server to temporarily serve as an ad-hoc SAN for the “target” physical, allowing us to “move” a Machine from one physical to another in seconds while the actual volume block clone happens in the background; Simon’s instrumentation work may have shaved about ~10s off this process (about 1/3rd of the total time).

Andres got host alerting (notifying users of hardware issues with the specific hosts they’re using, both on their personal status page and directly via email) integrated with our internal support admin tool.

Somtochi rolled out the first iteration of Pet Semetary to our [flyd orchestrator](/content/blog/carving-the-scheduler-out-of-our-orchestrator/ ""/index.html). We now have two (hardware-isolated) secret stores: Hashicorp Vault, and our internal Pet Semetary. The big thing here is, if we can’t read secrets, we can’t boot machines; now, if Vault has an availability issue, we “fall back” to Pet Semetary. Requeiscat, Vault-related Outages, 2020-2024.


#

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact; see last week’s bulletin for details about the top row of incidents).

A pretty straightforward week. The most painful incident was the Vault “outage”, in part because it happened on the eve of us cutting over to Pet Semetary, our Vault replacement; in our new post-petsem world, it’ll take an outage of both Vault and PetSem to disrupt deploys. The other two incidents were more limited in scope.

Anchor This Week In Infra Engineering

Dusty built out telemetry and monitoring for Fly Machine migration, in preparation for a regional migration of some Machines to a new upstream provider.

In addition to doing a cubic heckload of [routine hiring work](/content/docs/hiring/ ""/index.html) (do these updates sound fun? [we’re hiring!](/content/jobs/infra-engineer/ ""/index.html)), Matt and Tom revised one of our technical work sample tests, eliminating an inadvertent cheat code some candidates had discovered; a comprehensively broken environment we ask candidates to diagnose had a way to straightforwardly dump out the changes we had made to break it. Respect to those candidates for figuring that out, and helping us level up the challenge a bit.

Steve has had a fun week. He’s working on shipping (you heard it here first) static IP address assignments for individual Fly Machines — this means Fly Machines can make direct requests to the Internet (for instance, to internal on-prem APIs) with predictable IP addresses. The original plan was to run an IGP across our fleet, but Steve worked out a combination of fou tunnels and SNAT that keeps our routing discipline static while allowing address to float. It’s a neat trick.

Steve would also like us to tell you that he rebooted dev-pkt-dc10-9b7e.

Ben built out tooling for host draining. Last week we talked about Simon’s work shipping inter-server volume migrations. Now that we can straightforwardly move workloads between physicals, storage and all, we can rebuild the “drain” feature we had when with Hashicorp Nomad back in 2020 (before we had storage), which means that when servers get janky (inevitable at our scale), or things need to be rebalanced, we can straightforwardly move all the Fly Machines to new physical homes, with minimal downtime. There’s a lot of corner cases to this (for instance: not all the volumes on a physical are necessarily attached to Machines), so this is a tooling-intensive problem.

Andres and Kaz re-established telemetry, metrics, and alerting on our Rails API, after an incident last week - it didn’t directly impact deploys, but would have made incidents involving API server problems, which are not unheard of, harder to detect and more difficult to resolve.

Kaz worked on fly-proxy-initiated Fly Machine migration. True fact: you can start a Fly Machine with an HTTP request; if a request is routed to a Fly Machine in stopped state, it’ll start. Kaz is working towards automatic migration of Machines from hosts that overloaded (i.e., exceeding our internal utilization thresholds): instead of starting on a busy machine, we can initiate a migration to a less-loaded machine. Recall that the core idea of our migration system is temporary SAN-style connections: a Machine can boot up on a new physical long before its entire volume has been copied over. Automatic migration isn’t happening yet, but it’s getting closer.

Akshit worked on cloud-hypervisor integration with our flyctl developer experience. cloud-hypervisor is like Firecracker except Intel ships it instead of AWS (they are both memory-safe Rust KVM hypervisors with minimal footprints; they even share a bunch of crates). We use cloud-hypervisor for [GPU machines](/content/gpu ""/index.html) because it supports VFIO IOMMU device passthrough (ask us about the security work we did here, please). Operating cloud-hypervisor is similar enough to Firecracker that it’s almost a drop-in, but we’re still smoothing out the differences so they feel indistinguishable to users.

Tom and John are decommissioning our old, cursed “edge workers”. We run mainly two kinds of servers: edges that take traffic from the Internet and feed them into our proxy network, and workers that run actual Fly Machines. For historical reasons (those being: the founders made annoying decisions) we have on one of our upstreams a bunch of dual-role machines. Not for long. You may not like it, but this is what peak performance looks like:

Wrap text Copy to clipboard

root@edge-nac-fra1-558f: ~
$ danger-host-self-destruct-i-want-pain
!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
!!!!!DANGER!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!DANGER!!!!!
!!!!!DANGER!!!!!  _____          _   _  _____ ______ _____   !!!!!DANGER!!!!!
!!!!!DANGER!!!!! |  __ \   /\   | \ | |/ ____|  ____|  __ \  !!!!!DANGER!!!!!
!!!!!DANGER!!!!! | |  | | /  \  |  \| | |  __| |__  | |__) | !!!!!DANGER!!!!!
!!!!!DANGER!!!!! | |  | |/ /\ \ | . ` | | |_ |  __| |  _  /  !!!!!DANGER!!!!!
!!!!!DANGER!!!!! | |__| / ____ \| |\  | |__| | |____| | \ \  !!!!!DANGER!!!!!
!!!!!DANGER!!!!! |_____/_/    \_\_| \_|\_____|______|_|  \_\ !!!!!DANGER!!!!!
!!!!!DANGER!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!DANGER!!!!!
!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!

This script will TOTALLY DECOMMISSION and DESTROY this host and REMOVE IT
PERMANENTLY from the Fly.io fleet.
To proceed, enter the hostname: edge-nac-fra1-558f

Correct, this host is edge-nac-fra1-558f.

To proceed, repeat verbatim "Yes, IRREVERSIBLY decommission"
-> Yes, IRREVERSIBLY decommission
This is your LAST CHANCE. Press ENTER to run away to safety. Press '4' to begin.

Migration is a theme of this bulletin; like we said last week, it has been kind of our “white whale”.

We have not forgotten last week’s promise to publish Matt’s incident handling process documents, but Matt wants to clean them up a bit more. We’ll keep mentioning it in updates until Matt lets us release them.

This is a small fraction of our infra team! These are just highlights; things that stuck out to us at the end of the week.


#

This is a new thing we’re doing to surface the work our infra team does. We’re trying to accomplish two things here: 100% fidelity reporting of internal incidents, regardless of how impactful they are, and a weekly highlights reel of project work by infra team members. We’ll be posting these once a week, and bear with us while we work out the format and tone.

Anchor Trailing 2 Weeks Incidents

(Larger boxes are longer, darker boxes are higher-impact).

This was a difficult time interval, dominated by a pair of first-of-their-kind outages in the control plane for our global WireGuard mesh, which subjected us to several days of involuntary chaos testing, followed by a surprisingly long upstream power loss in one of our regions. “Incidents” for infra engineering occur somewhat routinely; these were atypically impactful to customers.

Anchor This Week In Infra Engineering

Somtochi completed an initial integration between flyd, our orchestrator, and Pet Semetary, our internal replacement for Hashicorp Vault. Fly Machines now read from both secret stores when they’re scheduled. This is the first phase of real-world deployment for Pet Semetary. Because Vault relies on a centralized Raft cluster with global client connections, and because secrets reads have to work in order to schedule Fly Machines, it has historically been a source of instability (though not within the last few months, after we drastically increased the resources we allocate to it). Pet Semetary has a much simpler data model, relying on LiteFS for leader/replica distribution, and is easier to operate. Somtochi’s work makes deployments significantly more resilient.

Simon got Fly Machine inter-server volume migrations working reliably, the payoff of a months-long project that is one of the “white whales” of our platform engineering. Volumes attached to Fly Machines are locally-attached NVMe storage; Fly Machines without Volumes can be trivially moved from one server to another, but Volumes historically could not be without an uptime-sapping snapshot restore. The new migration system exploits dm-clone, which effectively creates temporary SAN connections between our physical servers to allow Fly Machines to boot on physical while reading from a Volume on another physical while the Volume is cloned. Simon’s work allows us to drain workloads from sus physical machines, and to rebalance workloads within regions.

Andres built new internal tooling for host-specific customer alerts. At the scale we’re operating at, host failures are increasingly common; more hosts, more surface area for cosmic rays to hit. These issues generally impact only the small subset of customers deployed on that hardware, so we report them out in “personal status pages”. But we’re a CLI-first platform, and many of our customers don’t use our Dashboard. Andres has rolled out preemptive email notification, so customers get direct notification.

Dusty beefed up metrics and alerting around “stopped” Fly Machines. The premise of the Machines platform is that Machines are reasonably fast to create, but ultra-fast to start and stop: you can create a pool of Machines and keep them on standby, ready to start to field specific requests. Making this work reliably requires us to carefully monitor physical host capacity, so that we’re always ready to boot up a stopped Fly Machine. This is capacity planning issue unique to our platform.

Will continued our ongoing project to move all of Fly Metrics off of special-purpose hosts on OVH, which hosts have been flappy over the years, and onto Fly Machines running on our own platform. Metrics consumes an eye-popping amount of storage, and Will spent the week adding storage nodes to our new Fly Machine metrics cluster.

Matt capped off the week by, appropriately enough, fleshing out our incident response and review process documentation. We could say more, but what we’ll probably do instead is just make them public next week.