Zanda Health

Zanda Knowledge Base

Reliability and Availability

Zanda stays available through redundant infrastructure, constant monitoring, and a fast, practiced incident response when something does go wrong.

Zanda stays online through infrastructure built for failure, constant monitoring, and a practiced response when something does go wrong. This article explains how those pieces fit together, and how Zanda tracks its uptime.


How Zanda engineers infrastructure resilience

Zanda has engineered its infrastructure for resilience on top of Amazon Web Services (AWS): running across three AWS regions, with redundant systems and automatic failover built into both databases and servers. That infrastructure is distributed across multiple availability zones within each region, so a failure in one physical location doesn’t take the whole system down with it. Failover happens within a customer’s own region, not across regions, so data stays within its agreed storage region at all times, and data submitted just before a database failover is preserved through it. See Data Protection, Backups, and Disaster Recovery for how that regional replication works in detail.

Zanda handles an individual data-center failure automatically, without manual intervention. Zanda has experienced individual data-center disruptions, each handled through automatic failover, but has never experienced a whole-region outage - that would be an extremely rare event, a different scale of failure that would require restoration by AWS itself, alongside the disaster recovery process described in the article linked above.

Zanda has also built capacity to adjust automatically. When demand increases, the system scales up within roughly 30 seconds to a couple of minutes, and Zanda anticipates predictable peaks, such as the start of the business day in major time zones, ahead of time rather than only reacting to them. If an update needs to be rolled back, that typically completes in about 15 minutes, and Zanda tests infrastructure failover regularly, documenting it as part of the annual ISO 27001 audit.


How uptime is measured

Zanda sets its availability objective at 99.9%, a target it has exceeded consistently since 2010.

Zanda measures availability against every customer HTTP request, not a sample: a timeout or a 500-level server error counts against the measurement, so the figure reflects what customers actually experienced rather than an idealized best case. Zanda tracks it both regionally and globally, and reviews and reports on it internally every month.

Most maintenance happens without any service interruption at all - across a full year, the brief outages that planned maintenance does require add up to a matter of minutes. On the rare occasions planned maintenance does require a brief outage, that maintenance is scheduled outside business hours in the affected region, customers get at least a week’s notice by email, and the specific timing, converted to the customer’s own time zone, is published on the status page beforehand. The resulting downtime is excluded from the uptime calculation, since it is planned rather than a failure.


Proactive monitoring

Zanda runs automatic monitoring systems and extensive logging continuously, capturing errors and performance data without recording customer content.

Zanda alerts engineering and support teams immediately when a threshold is crossed, covering error rates, slow database queries, queue delays across messages, emails, and background processing, and the health of third-party integrations. Beyond real-time alerting, it also reviews logs weekly, to catch issues that are too infrequent to trip an automatic threshold but still worth investigating, and tests backups weekly too, investigating any failure promptly - see Data Protection, Backups, and Disaster Recovery for how that restore testing fits into disaster recovery.


Responding to incidents

An incident can start two ways: monitoring detects an anomaly, or a customer reports a problem. Either is enough to activate the incident response process - it does not wait for both. Zanda staffs on-call engineers across time zones, so someone is available to respond as soon as an issue is flagged, and issues affecting many customers get the highest-urgency response, with typical incident recovery work measured in minutes rather than hours. A catastrophic, region-level event is a different scale of recovery, covered separately in Data Protection, Backups, and Disaster Recovery.

Once an incident is active, Zanda centralizes communication and coordination in one place internally: engineering investigates and works the fix while support keeps customers updated directly, and the public status page is updated every 10 to 15 minutes for as long as the incident is active, so customers checking in have current information rather than a stale notice. The status page identifies which region is affected when an incident is regional, and incident information is also available through the support chat. See Zanda System Status Page for how to access it and subscribe to proactive notifications.

After an incident is resolved, Zanda runs a retrospective that documents the root cause and the improvements that follow from it. Zanda also has its incident management process itself reviewed as part of the ISO 27001 audit, checking the process, not just individual incidents.


Frequently Asked Questions

Does a single data center going down affect my access to Zanda?

No, not on its own. Infrastructure is spread across multiple availability zones within a region specifically so that one data center failing does not interrupt service. A failure affecting an entire AWS region is handled differently, through AWS’s own restoration process and disaster recovery.

Does planned maintenance count against the 99.9% availability objective?

No. Most maintenance causes no interruption at all. The rare planned maintenance that does require a brief outage is scheduled outside business hours and excluded from the uptime calculation, since it is a planned event rather than an unplanned failure.

How quickly does Zanda notice a problem?

Monitoring triggers alerts immediately when a threshold is crossed, so many issues are caught before a customer would notice them. Less frequent issues that do not cross an automatic threshold are still caught through weekly log reviews.

How much notice do I get before planned maintenance that needs downtime?

At least a week, by email. The specific timing, converted to your own time zone, is also published on the status page beforehand. Most maintenance causes no interruption at all, so this only applies to the rare occasions a brief outage is needed.

Does my data ever move to a different region during a failover?

No. Zanda runs infrastructure across three AWS regions, but a database failover stays within the region a customer’s data is stored in - it does not move data to a different region.

Can I get notified automatically if there is an incident?

Yes. See Zanda System Status Page for how to subscribe to proactive notifications, or check incident details through the support chat at any time.

Related articles

Was this article helpful?