Tech Citadels Solutions logoTECH CITADELSSOLUTIONS

Service Commitments

What we commit to, and what we do not

These are the defaults a managed engagement runs to. They cover the things that depend on us rather than on your infrastructure, which is why you will not find an uptime percentage on this page.

Why there is no uptime percentage here

Your platform runs in your cloud account or on your own hardware. Availability turns on your region, your network, your change windows and your capacity, none of which we own. A single number published to every visitor would be either meaningless or a promise that could not be kept, so the availability target is set per environment in the agreement, alongside the error budget and the recovery objectives. How that works is described under Reliability.

For context, what the common targets actually cost to reach, on a 30 day month:

TargetDowntime allowedWhat it takes
99.5%3h 36mA single well-run instance
99.9%43mRedundancy with tested failover
99.95%22mMultiple availability zones, automated failover
99.99%4m 19sMultiple regions, and the infrastructure bill that comes with it

The last column is the point. No amount of on-call discipline delivers four nines on a single instance, so a target is only worth agreeing next to the architecture that can carry it.

Response

These depend on our on-call rotation rather than on your environment, so they are the same for every managed engagement. Acknowledged means a named engineer is awake, has picked it up and has started work, not that an automated ticket was opened.

SeverityMeansAcknowledged withinUpdates
Sev1Service down, or data at risk15 minutes, day or nightEvery 30 minutes until resolved
Sev2Major function degraded, no workaround1 hour, day or nightEvery 2 hours
Sev3Minor issue, a workaround existsNext business dayOn resolution

Practice

WhatCommitmentWhy it is worded that way
Blameless postmortemWithin 5 business days of a Sev1 or Sev2You get the same document we do, with the actions and their owners named.
Restore drillQuarterly, with the measured time reportedA backup that has never been restored is not a backup. The recovery objective is measured rather than asserted.
Security patchingWindow proposed within 72 hours for an actively exploited vulnerability, 7 days for other critical onesProposed rather than applied, because the change window is yours to approve. We will tell you what we think the risk of waiting is.
Spend and capacity reviewMonthly, per workloadHeadroom against real growth and instances sized against real usage, so the bill is explainable rather than a surprise at renewal.

What we do not commit to

Worth saying plainly, because a commitment page that promises everything is worth nothing.

A resolution time

How long a fix takes depends on the fault, on vendors upstream of us, and on you approving a change. Any firm publishing a guaranteed resolution time is either padding it heavily or planning to miss it.

Uptime for infrastructure we do not run

Your cloud account, your network and your hardware are yours. If a region degrades or a change lands outside a window we agreed, that is not something we can be accountable for, and pretending otherwise would make every other number here less credible.

Availability of third-party model providers

When a commercial model provider has an outage, routing fails over to the alternatives configured for you, which is a large part of why the routing layer exists. What we do not do is promise you a provider we do not operate will be up.

How these become binding

What is on this page is the standard we hold ourselves to and report against. The agreement signed for an engagement fixes the availability target and recovery objectives for your environment, and sets the remedy if a commitment is missed. Ask for that in writing before you sign, from us or from anyone else.

These apply to managed engagements with an ongoing retainer. A fixed-fee build without managed operations is scoped in its own statement of work.

Questions about any of this go to [email protected].