22586 - SCTE Broadband - Sep2026 Complete v1

FROM THE INDUSTRY

Tell me how your company manages observability oversight. Our security intelligence comes from three key sources. First, our Internet Mapper platform analyses global internet intelligence, including carrier data and synthetic transactions, to build a real- time view of the internet’s overall health. Second, our customers provide valuable telemetry. During incidents such as a Cloudflare outage, customer data enables us to identify problems extremely quickly, allowing organisations to detect outages around 90% faster than the industry average. Finally, we enrich this with intelligence from our network of carrier and service provider relationships, providing additional context to pinpoint the source and impact of an issue.

large scale knowledge of things that are going on. And then we infer and we understand how those could possibly affect our customers as well as ourselves. so that’s the first major piece. What kind of customers do you work with? Our customers are Fortune 1K, primarily in independent industries, oil and gas, healthcare, financial services, internationally. Over the last year we’ve built a self-service version of our product for smaller customers. We work with big internet exchanges within the EU, companies who actually use our products with their own customers to help them have more resilient networks. Not only do we sell directly to enterprises, or in this case, with our with our new self- service platform to smaller organisations, but we also have a large number of carriers, internet exchanges, MSPs, especially in the EU, that are always looking at that. And that kind of goes across from the EU into Asia PAC, which has been a big arena for us. The world is an increasingly unstable place. Bad state actors, hostile neighbours, hysterical media amplifying attacks out of all proportion. How much of your work is of malicious nature?* The silver lining for all of us is that, statistically speaking, purely malicious activities still account for a smaller portion of overall downtime. However, that slice of the pie is growing rapidly and becoming far more severe. We are seeing a surge in large-scale malicious campaigns across multiple vectors. As you noted regarding state actors, current geopolitical tensions have fueled a rise in targeted infrastructure attacks, supply chain compromises, and sophisticated ransomware. Core internet infrastructure remains a massive target— with submarine cable systems primarily being attacked at a software level, there remains physically caused outages as a results of ship anchors dragging and ripping the cables apart.

are still non-malicious in nature. They are most often the result of human error, flawed software updates, or complex cloud misconfigurations. While it’s somewhat reassuring that we aren’t always under active attack, the sheer volume of these non-malicious outages is unacceptably high, and the industry’s Mean Time to Resolution (MTTR) is still far too long. At WanAware, our goal is to drastically reduce that impact. We operate on the assumption that outages—whether from a hostile state actor or a simple bad deployment—are inevitable. The critical challenge is compression: how do we take a catastrophic two-day or week- long outage and resolve it in a matter of minutes? Achieving that comes down to two fundamental components. First, you must be able to rapidly isolate the root cause. Second, doing so requires a deep, automated understanding of the environment’s dependencies, interdependencies, and the exact context of how a customer utilises their infrastructure and services. How do you do that? WanAware takes a fundamentally different approach to observability, which is why we’ve seen significant improvements in Mean Time to Repair (MTTR). Most platforms rely on analysing huge volumes of logs and metrics to infer the source and impact of an issue, but 90–95% of those alerts are false positives, creating alert fatigue and slowing incident response.

How has this process developed over time?

We’ve seen a significant increase in outages over the past 18 months, largely because IT environments have become far more complex. Twenty years ago, most applications ran in one or two data centres connected by private networks. Today, organisations rely on cloud services, SaaS platforms, AI applications and multiple third-party providers, creating far more dependencies across the internet. That means a problem with a major provider can have a cascading impact across thousands of organisations, making outages more frequent, harder to diagnose and far more disruptive.

We start by building a complete understanding of a customer’s

Goodness me. Are we all doomed?

infrastructure. Our platform automatically ingests and enriches asset data, mapping the dependencies between systems to create a real-time relationship graph. That context allows us to identify the true root cause of an issue rather than simply reporting symptoms. Our globally distributed monitoring network also extends visibility beyond a customer’s own infrastructure. We can identify outages caused by intermediate providers between networks—issues that individual carriers may not even detect. By combining infrastructure context with global telemetry, we pinpoint the source and impact of incidents in seconds, dramatically reducing troubleshooting time and helping organisations restore services much faster.

Certainly not. We just have to be smart about it and we have to think about things a little bit differently. This is obviously a reactive initiative as opposed to proactive. Tell me about your preventative measures. We provide both preventative and reactive measures that happen as close to the event as possible. Those are two of the major things that we do from a preventative perspective. First, we have

However, despite the frightening headlines, the vast majority of day-to-day outages

* Want to know more? Turn to page XX for our Long Read on subsea cable sabotage.

Volume 48 No.23 SEPTEMBER 2026

79

Made with FlippingBook - Online magazine maker