Files
nexus/sreweekly/markdown/477/06-observability-2025-navigating-costs-complexity-and-the-rise-of-ai.md
2026-09-12 17:23:01 +08:00

21 KiB
Raw Blame History

Observability 2025: Navigating Costs, Complexity, and The Rise of AI

简介

In this piece, I’ll delve into four macro challenges facing observability today, explore strategies that are emerging across the industry to address them, and offer my perspective on the trajectory of this crucial domain in the year to come.

正文

Observability 2025: Navigating Costs, Complexity, and The Rise of AI

In 2024, I had the privilege of engaging with top minds in Software, Platform, and Site Reliability Engineering at some of the most innovative companies in the world. A recurring set of observability challenges emerged, that are repeatedly hindering organisations.

Observability is no longer a luxury; it's a survival skill in today's hyper-connected world. It's the X-ray vision that allows organisations to peer deep into their software stacks, understand their inner workings, and proactively anticipate and resolve issues before they impact customers.

In this piece, I'll delve into four macro challenges facing observability today, explore strategies that are emerging across the industry to address them, and offer my perspective on the trajectory of this crucial domain in the year to come.

Four Observability Challenges Facing Organizations in 2024

The Soaring Cost of Observability

Turning Data into Actionable Insights

Something I wrote about back in September. 

Impact on Productivity

Something I originally explored back in March. 

Data Observability

cascading failures when real-time data applications encounter disruptions. This limited insight hinders their ability to proactively detect and resolve data flow and data integrity issues before they impact downstream systems and ultimately, the end-user experience.

During my discussions, I also encountered other areas of concern, such as the management of Generative AI applications. However, a recurring theme emerged: a lack of clear understanding of the specific operational challenges faced in these domains. Many individuals lacked firsthand experience with the intricacies of these technologies, relying primarily on hypothetical scenarios rather than practical, real-world challenges.

To truly grasp the operational realities, it's essential to gain hands-on experience. By actively engaging with these technologies and encountering the challenges firsthand, organisations can develop a deeper understanding of their operational needs and implement effective solutions.

How The Industry Is Responding

Before delving deeper into the observability challenges we've outlined, it's crucial to acknowledge the rapidly converging landscape of vendors. This convergence is driven by a confluence of factors, including aggressive acquisitions and the organic expansion of capabilities across the industry.

The rise of AI, particularly LLMs and AI Agents, has significantly lowered the barrier for vendors to integrate sophisticated analytical capabilities. This democratization of AI empowers vendors to offer innovative solutions like predictive maintenance and automated root cause analysis, blurring traditional product boundaries and intensifying competition. This dynamic landscape is evidenced by the emergence of new players and the reinvention of existing ones, all vying for dominance in this highly competitive and evolving market.

Controlling Observability Costs

Data Strategy: Effective cost containment in observability hinges on a robust data strategy. Begin by clearly defining your objectives:

Based on these objectives, meticulously define the essential metrics, logs, and traces required to gain the necessary insights. Avoid collecting unnecessary data, as this inflates costs, clutters systems, and hinders efficient data utilization by end-users. This focused approach ensures that resources are allocated effectively and that valuable insights are not obscured by irrelevant information.

On the ground it is also important to focus on eliminating situations where inaccurate or incomplete data can generate misleading insights and squander valuable resources. This is particularly critical when leveraging AI for observability data analysis, as flawed data inevitably leads to flawed outcomes. Implementing rigorous data validation and cleaning processes to ensure the reliability and integrity of your observability data is critical in this regard.

For organizations building applications on Kubernetes, eBPF technology offers a compelling solution to the challenge of comprehensive observability. A key advantage of eBPF lies in its ability to bypass a significant non-technical hurdle: the reliance on developers to manually instrument code. By leveraging eBPF, organizations can dynamically instrument applications without code modifications, reducing developer burden and minimizing performance overhead compared to traditional methods.

      An evaluation I was involved in 2024 identified [Odigos](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fodigos%2Eio%2F&urlhash=QGCT&trk=article-ssr-frontend-pulse_little-text-block) as a leading provider in this space. Their platform excels in seamless integration with existing tools and platforms. Odigos further distinguishes itself with its deep dive into data observability, particularly for event-driven applications and database environments. They continue to demonstrate leadership through ongoing innovation and recently secured significant backing from a funding round led by Venture Guides, with participation from Salesforce Ventures, Mango Capital, and Firestreak Ventures.
      Other noteworthy options include Beyla (now an open-source project under Grafana) and Pixie. For a more comprehensive list, including the use of eBPF for network and security management, refer to [https://ebpf.io/applications/](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Febpf%2Eio%2Fapplications%2F&urlhash=fJHU&trk=article-ssr-frontend-pulse_little-text-block).
    

      Telemetry Pipelines: These platforms facilitate edge-based data processing, including transformation and filtering, thereby minimizing the volume of data transmitted to backend observability systems. Established players in this space include [Chronosphere](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fchronosphere%2Eio%2F&urlhash=fEqx&trk=article-ssr-frontend-pulse_little-text-block) (following their acquisition of Calyptia), [Cribl](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fcribl%2Eio%2F&urlhash=QgPt&trk=article-ssr-frontend-pulse_little-text-block), [Edge Delta](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fedgedelta%2Ecom%2F&urlhash=2EIT&trk=article-ssr-frontend-pulse_little-text-block), [Mezmo](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fwww%2Emezmo%2Ecom%2F&urlhash=hC1l&trk=article-ssr-frontend-pulse_little-text-block), [ObserveIQ,](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fobserveiq%2Ecom%2F&urlhash=fuNr&trk=article-ssr-frontend-pulse_little-text-block) and [Middleware](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fmiddleware%2Eio%2F&urlhash=liyz&trk=article-ssr-frontend-pulse_little-text-block). If your organisation already leverages telemetry pipelines for other purposes, consider leveraging those existing investments. This can lead to significant cost savings through shared infrastructure, expertise, and potential licensing discounts.
    

      For organizations exclusively focused on Cloud Native Observability and leveraging OpenTelemetry, [ControlTheory](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fwww%2Econtroltheory%2Ecom%2F&urlhash=hqlJ&trk=article-ssr-frontend-pulse_little-text-block) is emerging as a noteworthy contender worth tracking. Founded by industry veterans with decades of experience, ControlTheory is developing a platform that emphasizes policy-driven telemetry management. This approach aims to provide valuable intelligence to significantly reduce the ongoing cost of ownership for organizations embracing OpenTelemetry.
    

      Evaluate Alternatives Backends:  Some of the vendors I mentioned earlier and platforms from [Dash0](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fwww%2Edash0%2Ecom%2F&urlhash=2R33&trk=article-ssr-frontend-pulse_little-text-block) and [Grafana](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fgrafana%2Ecom%2F%3Fsrc%3Dggl-s%26mdm%3Dcpc%26cnt%3D118483912276%26camp%3Db-grafana-exac-emea%26trm%3Dgrafana&urlhash=eN3f&trk=article-ssr-frontend-pulse_little-text-block) can offer attractive pricing models relative to incumbents. However, it's crucial to carefully consider the potential costs associated with migration.

Actionable Data & Productivity

These two go hand in hand and amidst these challenges, a new force is emerging: Artificial Intelligence (AI). LLMs are poised to revolutionize observability by:

      Pioneers in this space include [Resolve.ai](https://www.linkedin.com/redir/redirect?url=http%3A%2F%2FResolve%2Eai&urlhash=Fbk4&trk=article-ssr-frontend-pulse_little-text-block) and [Flip.ai](https://www.linkedin.com/redir/redirect?url=http%3A%2F%2FFlip%2Eai&urlhash=-k4R&trk=article-ssr-frontend-pulse_little-text-block). Following quickly on their heels are companies like [Cleric](https://www.linkedin.com/redir/redirect?url=http%3A%2F%2Fcleric%2Eio%2F&urlhash=jgIf&trk=article-ssr-frontend-pulse_little-text-block), [Deductive](https://www.linkedin.com/redir/redirect?url=http%3A%2F%2Fdeductive%2Eai%2F&urlhash=0j4F&trk=article-ssr-frontend-pulse_little-text-block), [Traversal](https://www.linkedin.com/redir/redirect?url=http%3A%2F%2Ftraversal%2Ecom%2F&urlhash=mIKV&trk=article-ssr-frontend-pulse_little-text-block). A number of established vendors have also experienced a resurgence including [Logz.io](https://www.linkedin.com/redir/redirect?url=http%3A%2F%2FLogz%2Eio&urlhash=TJMZ&trk=article-ssr-frontend-pulse_little-text-block) and [Science Logic](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fsciencelogic%2Ecom%2F&urlhash=msYz&trk=article-ssr-frontend-pulse_little-text-block), leveraging LLMs in this context. 

While LLMs are undoubtedly making significant strides, I have yet to observe concrete evidence demonstrating their ability to definitively establish causal relationships beyond mere correlations. The prevalent terms they employ, such as 'probable explanations' or 'likely causes', suggest a cautious approach to attributing causality.

Undeniably, vendors leveraging LLMs have achieved remarkable advancements in their capabilities. However, true autonomy necessitates a profound understanding of cause-and-effect relationships with exceptional precision.

      Vendors explicitly claiming to offer Causal Reasoning capabilities within this domain include [Causely](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fwww%2Ecausely%2Eai%2F&urlhash=7TSd&trk=article-ssr-frontend-pulse_little-text-block), [Dynatrace](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fwww%2Edynatrace%2Ecom%2F&urlhash=qgzT&trk=article-ssr-frontend-pulse_little-text-block), and [Senser](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fsenser%2Etech%2F&urlhash=OSQ8&trk=article-ssr-frontend-pulse_little-text-block). I delve deeper into this critical topic in my article, [Fools Gold Or Future Fixer: Can AI-powered Causality Crack the RCA Code For Cloud Native Applications?](https://www.linkedin.com/pulse/fools-gold-future-fixer-can-ai-powered-causality-crack-mallaband-xz1xf/?trackingId=lO3ATt1mTf67rQr8IMKTtQ%3D%3D&trk=article-ssr-frontend-pulse_little-text-block)" where I meticulously examine the requirements for genuine causal inference and outline crucial evaluation criteria. When evaluating potential solutions in this area, carefully consider these requirements. Build a shortlist of promising vendors and conduct thorough Proof-of-Concept (POC) evaluations to validate their claims and assess their suitability for your specific needs.

The democratization of AI access will undoubtedly accelerate the evolution of existing and new offerings in this space. We can expect a surge in innovation as vendors strive to capitalize on the power of LLMs. A key question that intrigues me is: Which vendors will be the first to effectively combine the strengths of LLMs with robust causal reasoning capabilities?

This convergence of technologies holds immense potential. LLMs excel at processing and generating human-like text, while causal reasoning algorithms are crucial for understanding and predicting cause-and-effect relationships within complex systems. The vendor that successfully integrates these two powerful technologies will likely gain a significant competitive advantage.

      [CausalLens](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fcausalens%2Ecom%2F&urlhash=pKNx&trk=article-ssr-frontend-pulse_little-text-block), a leader in causal AI for data science, serves as a compelling example of how these technologies can be integrated in other domains. While currently focused on data science applications, the core principles of causal inference are highly relevant to the observability space. We can expect to see innovative solutions emerge that leverage LLMs and causal reasoning techniques to enhance observability platforms.
    

      Since first publishing this article a company called [Nofire.ai](https://www.linkedin.com/redir/redirect?url=http%3A%2F%2FNofire%2Eai&urlhash=Y5f3&trk=article-ssr-frontend-pulse_little-text-block) left comments claiming to offer this capability already.  

The Rise of AI and Data Observability

      Data Delay & Missing Data - In my June article, ['Real-time Data & Modern UXs: The Power and the Peril When Things Go Wrong,'](https://www.linkedin.com/pulse/real-time-data-modern-uxs-power-peril-when-things-go-wrong-mallaband-audmf/?trackingId=Wj56Uc3FQzqq0bau6jg1Gg%3D%3D&trk=article-ssr-frontend-pulse_little-text-block) I explored how various industries leverage real-time data to deliver exceptional user experiences. However, I also highlighted the critical risks associated with data lags and missing data, particularly in these dynamic environments.
    

      My 2024 engagement with customers revealed a concerning trend: many organizations operate with limited visibility into the intricate flow of data across their systems. This blind spot is particularly acute in environments characterized by asynchronous communication, such as those involving services that interact through messaging platforms/services, as well as interactions between service, distributed databases caches... This lack of understanding leaves organisations vulnerable to [cascading failures](https://www.linkedin.com/pulse/beyond-blast-radius-demystifying-mitigating-cascading-mallaband-ee0ie/?trackingId=Xnv8y0IfQUqRUMNagjx5wg%3D%3D&trk=article-ssr-frontend-pulse_little-text-block) when real-time data applications encounter disruptions.

Fortunately, the OpenTelemetry community has made significant progress in developing instrumentation to support the discovery and monitoring of service relationships. For organizations seeking to avoid vendor lock-in and ensure future flexibility, adopting OpenTelemetry should be the default approach. This includes exploring the potential benefits of leveraging eBPF for enhanced system-level observability.

Furthermore, organizations should prioritize the exploration of intelligent telemetry analysis techniques. These techniques are crucial for streamlining incident management and proactively identifying early warning signs of potential issues within complex service dependency chains.

      Data Integrity - Addressing data integrity challenges is crucial for modern organizations. Several vendors, including [Monte Carlo Data](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fwww%2Emontecarlodata%2Ecom%2F&urlhash=f4gM&trk=article-ssr-frontend-pulse_little-text-block), [Acceldata](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fwww%2Eacceldata%2Eio%2F&urlhash=8F3k&trk=article-ssr-frontend-pulse_little-text-block), [Bigeye](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fwww%2Ebigeye%2Ecom%2F&urlhash=hf88&trk=article-ssr-frontend-pulse_little-text-block), and [DataKitchen,](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fwww%2Ebigeye%2Ecom%2F&urlhash=hf88&trk=article-ssr-frontend-pulse_little-text-block) are some of the players developing innovative solutions in this space. While I haven't delved deeply into the specifics of each vendor's offerings, I recognize the critical importance of these solutions based on my interactions with customers. 

Data integrity issues can have significant consequences, impacting everything from operational efficiency and customer experience to regulatory compliance. By leveraging solutions from these and other emerging vendors, organizations can proactively identify and address data quality problems, ensuring the accuracy and reliability of their data assets.

Conclusion

The observability landscape is evolving at a breakneck pace. The emergence of AI, the increasing complexity of distributed systems, and the growing emphasis on data-driven decision-making are all driving rapid innovation in this space.

      As we move forward, it will be fascinating to observe how the competitive landscape unfolds. Which vendors will emerge as leaders in this dynamic market? How will the integration of AI and machine learning revolutionize observability practices? And who will ultimately claim the coveted "[O11Y Awards](https://www.linkedin.com/posts/john-hayes-o11y_o11ys-observability-devops-activity-7279984955391184896-K9ZB?utm_source=share&utm_medium=member_desktop&trk=article-ssr-frontend-pulse_little-text-block)" in 2025? I cannot wait to hear from the infamous 
  
    [John Hayes](https://uk.linkedin.com/in/john-hayes-o11y?trk=article-ssr-frontend-pulse_little-mention)

and Observability 360 in December regarding this.

The answers to these questions remain to be seen, but one thing is certain: the journey towards a truly comprehensive and insightful understanding of our complex systems has only just begun.

BTW I do realise that a lot of vendors are missing from this article so if you would like to have a say on this whether you work on the end user side or for a vendor please share your thoughts in the comments.

Just published a new Newsletter which features this and 11 other articles in my Observability 2025 series https://www.linkedin.com/pulse/unlock-future-digital-operations-your-essential-guide-mallaband-ej8ke/?trackingId=W7dd9sStQPaFlgdmajOeWw%3D%3D

In case you missed it the next installment has arrived https://www.linkedin.com/posts/andrew-mallaband-88b1b7_observability-sre-aiops-activity-7333092140253659137-9tBa?utm_source=share&utm_medium=member_ios&rcm=ACoAAAAHeysBfS7vSo-aICN2qukOww4KbZOM3wc

https://www.linkedin.com/pulse/ai-sre-tooling-navigating-hype-reality-nascent-market-mallaband-ptc2e?utm_source=share&utm_medium=member_ios&utm_campaign=share_via

Nice write up Andrew Mallaband Instana should definitely belong to the list here. It supports Pearlian do-calculus-based causalAI for root cause identification for over a year. More information can be found here: https://www.ibm.com/products/blog/probable-root-cause-accelerating-incident-remediation-with-causal-ai https://www.ibm.com/new/announcements/introducing-probable-root-cause-enhancing-instanas-observability

Next article in the series coming on Wednesday https://www.linkedin.com/posts/andrew-mallaband-88b1b7_observability-ai-devops-activity-7297248923503521792-8flF?utm_source=share&utm_medium=member_desktop&rcm=ACoAAAAHeysBfS7vSo-aICN2qukOww4KbZOM3wc