Disaster recovery initial setup
This page captures the initial disaster-recovery considerations that should be validated before go-live. Recurring failover and failback operations belong in the Administration Guide.
Detection of need
We recommend deep observability provision around Terraform Enterprise both in terms of live metrics consumption(opens in new tab) and processing the diagnostics API(opens in new tab). Further, we strongly recommend connecting alerts to these metrics in case 200 responses are not experienced.
These provisions afford the Terraform Enterprise platform owner the ability to keep a close eye on the health of the platform. It is always better to be able to notify your user base of an outage situation than experience them notifying you.
On receiving a system health alert, the platform team can understand the level of outage and notify senior management accordingly, but we suggest that the decision to failover requires human confirmation.