Thought Leaders

From Uptime to Experience: The AI-Driven Shift in Modern Observability

mm
Add Unite.AI to your preferred sources on Google

In 2001, IBM wrote an autonomous IT manifesto. The Vision of Autonomic Computing broke down self-management into four pillars: self-optimization, self-healing, self-configuration, and self-protection. I was at Microsoft when IBM laid out this IT vision. We reacted by proposing technology ideas such as the autonomous data center, but ultimately, it was a dream that was way ahead of its time. There was no practical way to bring that vision to reality.

In my Microsoft days, I was on the team behind Clippy. Even though the animated paperclip assistant was infamously intrusive, the idea behind it was sound: computers should actively help humans do their jobs. We just didn’t have the computational horsepower and AI to make it possible. 25 years later, we finally do.

From Service Level to Experience Level

The concept of observability didn’t originate in IT. In 1960, Hungarian-American engineer and mathematician Rudolf E. Kálmán coined the term “observability” to describe how well a system can be measured by its outputs. Then, in 2013, Twitter adopted the term in a series of blog posts, effectively saying that old-fashioned monitoring, through all the commercial off-the-shelf tools they had available, was designed for a different era of technology and didn’t work in microservice-scale architectures.

Think of it like a doctor examining a patient. They can check their pulse, take their blood pressure, and observe other external features to indirectly assess the internal health of the patient. In IT, we need to do the same. When there’s a flicker in the patient’s pulse, we need to know whether that means there are issues in the kidney or the liver. At the scale and complexity of operations that Twitter was dealing with even 20 years ago (the company was only serving 100 million users with real-time tweets and feeds), observability required a different tooling and monitoring approach.

Today’s systems have become even larger and more complex, with dependencies on content delivery networks, caching and distributing bitmaps, fonts, JavaScript files, and so on all around the world. Truly understanding the performance of real-world applications is no easy feat.

When IT gets paged at 4 o’clock in the morning, someone has to get out of bed and figure out whether the issue is due to a bad sector on a hard drive or a bad actor trying to penetrate and cause havoc in the infrastructure. It doesn’t really matter which one it is: at the end of the day, their job is to keep all the systems running. Thankfully, to assess application health today, we can take in all the available telemetry: every network device, every application, thousands of out-of-the-box integrations, ticket flow through JIRA or Atlassian, and so many other signals.

This is where Experience Level Objectives (XLOs) come in. You’ve likely heard of Service Level Agreements (SLAs) and Service Level Objectives (SLOs), but XLOs take the next step by measuring whether your customers and employees are getting the level of experience that they want. It’s about quality, not just uptime. From a technical perspective, the only way to achieve XLOs is to have visibility from the NIC to the end-user device.

Last October, AWS US-EAST-1 went down. Catchpoint detected the issue 16 minutes before Amazon publicly acknowledged it. Customers with that visibility were able to react before their users could feel the outage’s consequences.

The promise of observability is like being Smokey Bear: detect where there’s smoke before there’s a fire. Done right, observability lets you douse a prairie blaze before it becomes a conflagration that takes out the Palisades in California. Smokey is the early warning system that can detect those small whiffs of smoke regardless of where they’re coming from: an AWS issue, an Oracle issue, a GCP issue, a Microsoft Azure issue, or something awry in your infrastructure.

AI Scales Security Systems

No human operator can keep tabs on today’s infrastructure systems. The only way to monitor systems at scale, ingesting petabytes of log data and trillions of metrics per day, is to use AI.

For example, say you want to track the read/write performance on a disk or the input/output or packet buffer overruns within your networking environment. You can use a dynamic threshold to define what normal looks like, or a deterministic way to look at time series data over the past week, month, year, or whatever timeframe you want, and establish normal performance thresholds. Once you have this statistical analysis, you can set levels for two standard deviations from the mean, so that when something happens outside of that range, you get an alert that performance is potentially abnormal.

Highly complex systems, however, can receive thousands of alerts per day. Dashboards start flashing, and people start getting paged. Sifting through all these alerts is not a good use of humans’ time. Indeed, Vectra estimates that organizations receive an average of 2,992 security alerts per day, 63% of which go unaddressed.

AI tools can cut down these thousands of alerts per day to just a few dozen. I remember one case where a single issue on a single NIC on a single machine caused 2,000 downstream alerts. Thanks to AI, the client was able to perform alert correlation and get to a much faster root cause analysis, which in turn concluded that one issue at this one point in time was causing the company’s whole dashboard to turn red.

AI Is Making IT Exciting Again

I took some time off after Cisco acquired Splunk in 2023. Over the next two years, I watched as my friends and former colleagues founded companies to use AI in ways that were not possible even five years ago. (Remember that if ChatGPT were a human child, it would be a three-year-old).

IT teams need help detecting smoke before the alarm goes off, not more dashboards to stare at. They. In a way, this is the same problem that IBM, Twitter, and even Microsoft with Clippy have all been trying to address.

This is the reason I decided to jump back in. The technology has finally reached a point where we can deliver on the original promise of observability and autonomous IT.

Garth Fort is Chief Product Officer at LogicMonitor, where he leads global product strategy and execution for the company’s AI-powered observability platform, LM Envision. A seasoned technology executive, Garth brings more than 20 years of experience driving product innovation, business growth, and cloud transformation across some of the world’s most respected enterprise software companies.