← All posts

WAN monitoring: from green lights to real transparency

Reachable doesn't mean good: why device monitoring alone is no longer enough, which metrics matter and how measurement data makes SLAs enforceable.

Helles Network Operations Center mit Monitoring-Dashboards – WAN-Monitoring: Vom grünen Lämpchen zur echten Transparenz

The rollout is done, every site has been migrated, and the dashboard shows nothing but green. This is exactly where the part begins that decides the value of a WAN project: everyday operation. Whether a network is good only shows month after month — and you can only judge it if you measure it.

The uncomfortable truth in many networks: users notice problems before IT does. The video call stutters, the ERP feels sluggish, and at some point someone picks up the phone. Whoever starts looking only then is searching under time pressure and without data.

Green lights don't tell you much

Classic monitoring checks reachability: the device answers ping, the interface is up, utilization stays below the threshold. All of that can be true — and the video conference is still unusable. As little as one percent packet loss visibly disrupts real-time applications while the availability statistics keep reporting 100 percent. Reachability is the minimum requirement, not proof of quality.

On top of that, latency and jitter decide the application experience, yet they don't appear in pure device monitoring at all. A line can be "up" and still have been taking a detour for days because a routing path changed.

Most of your traffic runs through other people's networks

The second blind spot is structural. With SD-WAN and SASE/SSE , the majority of traffic runs across the internet: through provider networks, across exchange points, into cloud platforms. But SNMP and device monitoring end at your own network edge — exactly where the interesting problems start.

Whether a provider currently has an issue at its peering , whether a cloud service is reached over an unfavourable path, whether the bottleneck sits in your own overlay or in the provider's underlay: only those who measure the entire path can see it, not just their own devices.

Three levels of transparency

In practice, a three-level setup has proven itself — levels that build on each other rather than replacing one another:

  • Level 1 – devices: SNMP, ping, interface counters. Answers the question of whether your own hardware is running. Limitation: it only sees your own devices.
  • Level 2 – traffic flows: Flow data (NetFlow/IPFIX) shows who talks to whom, which applications occupy the bandwidth and where congestion builds up. Limitation: visibility ends at the network edge.
  • Level 3 – application experience: Synthetic tests and measurements from the user's perspective check the complete path into the application ( observability , digital experience monitoring). Answers the question that matters in the end: does the performance arrive at the user — and if not, where is it stuck?

Without level 1 the foundation is missing; without level 3 the user experience remains guesswork. Anyone setting up today plans all three from the start.

The metrics that matter

Four values belong in continuous collection per line and per critical application: latency, jitter, packet loss and availability — complemented by the response times of your most important applications. More important than any single value is the baseline: what is normal at this site, on this route? Only with that reference does a measurement become a statement.

Alerts belong on degradation, not just on total outage. And thresholds need judgement: too sensitive creates alert fatigue, too blunt means users are once again faster than your monitoring.

Measurement data is SLA evidence

With carriers, documentation is what counts in the end. Your own independent measurements turn "it feels slow" into a solid ticket with a time window, a route and values. That speeds up fault clearance because the provider receives verifiable data — and if committed values are missed persistently, the measurement series become the basis for service credits and contract discussions.

Without your own measurements, all you have is the provider's statistics. But the provider measures its own network, not the path your data actually takes.

Monitoring is a process, not a tool

The most frequently overlooked point: tools alone change nothing. Dashboards nobody looks at are decoration. What matters is the process behind them — who assesses the alerts, who correlates underlay and overlay, who reports the fault to the right provider and escalates when clearance stalls?

That is why monitoring is an integral part of the managed service at KAEMI: we monitor lines and overlay across all sites, correlate anomalies and manage fault clearance with the providers — including escalation. Through our own autonomous system (AS 51834) we see routing and peering anomalies directly in the routing data, often before they are reported as incidents. What that looks like across more than 250 sites with multiple providers: see our case study from industry .

The pragmatic way in

Nobody has to start with a mega project. A realistic roadmap:

  1. Take stock: what is measured today — and who actually reads it?
  2. Close the basics: capture all WAN routers and lines cleanly in device monitoring
  3. Activate flow data (NetFlow/IPFIX) for the WAN routes
  4. Establish baselines: latency, jitter, packet loss, availability per line
  5. Set up synthetic tests for your five most important applications
  6. Configure alerting for degradation instead of outage
  7. Clarify responsibility: who reacts, who escalates, who reports?
  8. Establish SLA reporting: compare measurements against contractual commitments

A WAN nobody measures cannot be steered — only repaired once it is already stuck. Those who treat transparency as part of the network from day one, rather than an accessory, spot degradation before their users do, talk to providers on equal terms and make capacity decisions based on data instead of gut feeling. That is exactly what managed SD-WAN is for: connectivity and the transparency over it from a single source.

Want to connect your sites with performance and resilience?

KAEMI handles design, rollout and 24/7 management of your SD-WAN — including carrier management and redundancy.