← All posts

WAN SLAs and MTTR: what really counts

99.9% sounds good and still allows 8.8 hours of downtime a year: how to spot a solid WAN SLA — measurement method, committed metrics, MTTR and penalties.

Minimalistische Wanduhr neben Netzwerk-Kabelkanal mit orangen Patchkabeln – WAN-SLAs und MTTR: Was wirklich zählt

When companies buy site connectivity, one single number often decides in the end: 99.9 percent. It features prominently in the offer, it sounds reassuring, and it says almost nothing. Because what you really buy with a line is a promise — the promise of how quickly things continue when it fails. That promise lives in the SLA, and that is where a close look pays off.

99.9% is not 99.9%

First, the sober arithmetic of what availability classes actually permit:

  • 99% — up to 3.7 days of outage per year
  • 99.9% — up to 8.8 hours per year
  • 99.99% — up to 53 minutes per year
  • 99.999% — up to 5.3 minutes per year

For a sales office, 99.9% may be perfectly fine. For a production line or a logistics hub, 8.8 hours per year are a different matter — especially since they may occur in one piece.

More important than the number, however, is the measurement method, because two providers with the same percentage can promise entirely different things. Does planned maintenance count as an outage or not? Does a partial outage at half the bandwidth count as available? Is availability measured per site, or averaged across the whole network so that one permanently degraded site disappears in the statistics? And do the figures refer to a month or a year? Comparing percentages alone means comparing incomparables.

MTTR: the underestimated number

The second figure that decides everyday reality is time to repair — usually stated as MTTR (mean time to repair). It determines how long a site is actually offline when something happens. Here too, the differences sit in the details:

  • Guaranteed or target? An "aspired" value without consequences is a statement of intent, not a commitment.
  • When does the clock start? At ticket creation — or only once the provider has confirmed the fault? Hours can pass in between.
  • Response is not restoration. A two-hour response time helps little if restoration takes another eight afterwards. Both values belong in the contract separately.
  • Service hours. Fault clearance "within eight hours" during business hours Monday to Friday means: the Friday-evening fault gets handled on Monday. Critical sites need 24/7 fault clearance — and that is either written into the contract or it does not exist.

For redundantly connected sites the focus shifts: what counts there is less the repair time of a single line and more the question of whether failover kicks in automatically and fast enough to go unnoticed.

Quality is more than availability

A line can be available and still unusable. The application experience is decided by latency, jitter and packet loss — and precisely these values are missing from many SLAs, or appear only as non-binding guidance. For orientation: for voice, ITU-T G.114 considers an end-to-end latency of up to roughly 150 milliseconds acceptable; jitter targets are usually below 30 milliseconds and packet loss below one percent. Where things start to hurt in practice depends on your applications — what matters is that the values sit in the contract as committed metrics, not as "best effort".

And: commitments without your own measurements remain theory. How to build baselines and turn measurement series into solid fault tickets is what we covered in our post on WAN monitoring — SLA management and monitoring are two halves of the same job.

Penalties: steering, not compensation

Service credits never make up for the business damage of an outage — a few percent of a monthly fee bears no relation to a warehouse at standstill. Their real value is different: they make the commitment measurable and give the provider an economic reason to take faults seriously. For that to work, three things must be clear: the trigger threshold (from exactly which point does the credit apply), the amount, and the mechanism — is it credited automatically, or does it require a claim within a deadline that nobody files in day-to-day business?

Then there is the small print: how far does the force majeure clause reach? Who is liable on the last mile if a third-party operator sits there? Such gaps only surface during an incident — when it is too late.

Several providers, one promise

After an SD-WAN migration, the provider landscape is deliberately diverse: different carriers per site and per line increase resilience. The flip side: many contracts, many SLA definitions, many fault processes. This is exactly where responsibility gaps appear — every provider measures its own network, and nobody feels responsible for the path in between.

That is why SLA management is part of the managed service at KAEMI: we consolidate the providers' contracts and fault-clearance processes, verify independently with our own measurements and drive faults with the right provider — including escalation when clearance stalls. For our customers, that means one contact and end-to-end responsibility. What that looks like across more than 250 sites with multiple providers: see our case study from industry .

The checklist for your next contract

  1. Choose the availability class per site by criticality, not across the board
  2. Clarify the measurement method: maintenance windows, partial outages, per site or averaged, month or year
  3. Agree MTTR as a guaranteed value, not a target
  4. Define the start of the repair clock (ticket creation, not confirmation)
  5. Have response and restoration times stated separately
  6. Check service hours: 24/7 fault clearance for critical sites
  7. Anchor latency, jitter and packet loss as committed metrics
  8. Service credits with a clear threshold and automatic crediting
  9. Check the last mile and force majeure in the small print
  10. Build your own measurement data to hold commitments to account

An SLA is not an annex to the connectivity contract — it is the actual product you are buying. Those who understand measurement methods, repair times and penalties, and verify them with their own data, negotiate on equal terms and face no surprises during an incident. And if you would rather not shoulder that effort yourself: that is exactly what managed SD-WAN is for — connectivity, provider and SLA management from a single source.

Want to connect your sites with performance and resilience?

KAEMI handles design, rollout and 24/7 management of your SD-WAN — including carrier management and redundancy.