August 4, 2026

Observable AI: Why Most Investments Stall

By EverOps

Closing the AI Visibility Gap: A Fireside Chat with EverOps' CEO, Stephen Koza, and Delivery Director, Joseph Angeja

Ask an engineering team whether their AI tools have made them more productive, and the honest answer usually comes back as "maybe." Engineers say maybe, support says maybe, and leadership stays uncertain because nobody has the objective data to say otherwise. The licenses are paid for, and the pilots are running, yet the question the board keeps asking still goes unanswered.

EverOps CEO Stephen Koza and Delivery Director Joseph Angeja regularly see this gap play out in client engineering teams. Angeja spends his days inside the incident channels and planning meetings where AI tools either earn their keep or quietly get abandoned, and his read on why the gap persists points to something leadership rarely measures.

Their answer throughout it all keeps coming back to the idea that the problem organizations think they have with AI usually turns out to be a problem with what they can see. Throughout their insightful 30-minute conversation, the two leaders share their top priorities and takeaways worth carrying into your next planning conversation.

Pressure Creates Movement Before It Creates Strategy

There are two figures that helped frame the discussion. Last year, a Gartner survey of 506 CIOs found that 72% of organizations are breaking or losing money on AI. Similarly, McKinsey's 2025 AI survey found that roughly 6% of high performers generate meaningful EBIT impact. What separates those groups has nothing to do with the technology itself, which Angeja was direct about in explaining what happens instead when board pressure lands on an engineering org.

"Pressure creates movement before it creates strategy."

Teams start piloting tools, individual departments adopt their own solutions, and six months later leadership finds dozens of AI initiatives running in parallel with no consistent way to measure any of them. The first month of a rollout might look encouraging, but then a wide gap opens between the engineers who have built the tools into their daily work and those who use them only occasionally. Everybody has an opinion about whether productivity improved, and nobody has the data to settle it.

Angeja framed the underlying discipline gap in a line worth repeating to any executive team debating its next AI purchase.

"Organizations underestimate how much operational discipline is required to turn AI from a tool into a capability."

Koza made the same point from a different angle, drawing on a pattern he has watched play out for years. 

"Most problems are not technology problems. They're something else masquerading as a technology problem."

In this case, movement without measurement is just motion, and most organizations can't yet tell the difference from the inside.

What Observability Has to Answer Once AI Enters the Picture

Observability earned its place watching infrastructure health, covering servers, cloud resources, and application performance. The arrival of AI across the organization adds a second set of questions that the same practice now has to answer. 

  • Who is actually using these tools? 
  • Which workflows improved? 
  • What does the combined spend look like across every AI platform in play, and does any of it correlate to an outcome the business cares about?

"Historically, observability told us whether our infrastructure was healthy. Today, it also needs to tell us whether our AI investments are healthy."

That expansion pushes the scope of observability out of systems and into business processes and workflows, which asks something new of teams whose instrumentation was built for uptime. It also creates a compounding problem for anyone still fighting tool sprawl, because the data needed to evaluate AI adoption ends up scattered across the same disconnected systems. EMA research found that only 46% of IT organizations consider themselves fully successful with the observability tools they already own, and layering AI measurement on top of that foundation widens the gap rather than helping close it.

For teams that are struggling with the same thing and want a structured read on where they stand, we suggest starting with our Observability Maturity Assessment, which surfaces blind spots before they become blockers.

Where AI Already Earns Its Keep in Operations

One of the most valuable stretches of the conversation covered a workflow Angeja's team built with a fintech client, offering the clearest picture of what a well-built observability foundation makes possible for AI.

Incident management has a familiar shape. An alert fires, an engineer opens dashboards, reviews logs, checks recent changes against open pull requests, creates a ticket, and only then starts troubleshooting. Every step costs time, and the work of assembling context from four or five systems arrives at the exact moment it is most expensive.

Connecting the observability platforms, the service management layer, and the project tracking system into one source is what collapsed that sequence. By the time an engineer joined the incident channel, a summary was already waiting, including what changed, the likely root cause, the affected services, and recommended next steps. The engineer still made every call. The investigative legwork was completed before they arrived, and the mean time to resolution decreased accordingly.

Angeja was careful to name the prerequisite that makes any of it work. The workflow depends on trustworthy platforms and data, which means the foundation has to be genuinely in place before the automation sits on top of it. A fragmented data layer gives an AI workflow nothing reliable to reason across, no matter how capable the model is.

That same principle scales up to how leaders consume operational data. Angeja described leaders managing 10 dashboards across dozens of tools and saw the interface change. 

"AI becomes the interface between leadership and that operational data." 

A leader asks which initiatives are producing measurable value or where the overspend is, and the answer assembles itself from every underlying source rather than from five separate dashboards. He puts the timeline at 12 to 24 months for this to become a standard operating model, with the furthest-along teams already there.

The Metrics That Answer the Board's Question

Asked what he would put in front of an executive team to show whether AI investment is working, Angeja named three categories worth building before the next board conversation.

Adoption comes first, covering who is using which tools and how consistently. Koza flagged a useful trap here, pointing to the company that stood up public leaderboards for token consumption and scrapped them once it became clear how easily they were gamed. Reaching the top of that leaderboard mostly required wasting money on tokens, which tells leadership virtually nothing about value. Usage only matters when it connects to an outcome.

Productivity is next, and where that connection is made, through mean time to resolution, deployment frequency, engineering throughput, and the extent to which repetitive work has genuinely been automated rather than merely assisted. Angeja mentioned one organization that mandated shorter sprint cycles while holding delivery volume steady, which is the kind of result that actually survives scrutiny in a board deck. Financial impact closes the set by tying measurable value back to actual spend, including the downstream effects on cloud costs and incident-driven downtime.

Sentiment sits outside all three. Positive sentiment is good news on its own terms, and it answers a different question than ROI does. High adoption alongside flat operational metrics is a very different conversation from high adoption alongside improving ones, and only the data tells you which one you are having.

Start With What You Already Have

For leaders who suspect they are behind, the wrong move is buying more tools, which is how the sprawl started. The first productive step is an honest read of the current state: 

  • What is deployed today 
  • What is being measured
  • Where the gaps in visibility and governance lie
  • Which business outcomes actually need improvement. 

Once that baseline exists, sequencing the work gets much easier.

Koza summarized it with a line one of their clients uses. 

“You don't want to be a hammer chasing a nail.”

In other words, starting with the tool and hunting for a problem it solves is how organizations end up with dozens of initiatives and no measurable return.

What EverOps’ assessments consistently surface, Angeja noted, is duplicate tooling across departments, a lack of adoption measurement, and the absence of operational workflows or governance. That pattern leads to the conclusion that gives this whole conversation its point.

"What organizations often discover is they don't have an AI problem. They have a visibility problem."

The fix in every one of these cases is to see clearly before spending more.

Take the Next Step

An assessment tends to surface more in an hour than a quarter of internal debate does, since it's built to look at the same blind spots Angeja described. Starting there, with the Observability Maturity Assessment or the parallel AI Opportunity Assessment, will help give you an honest baseline before you touch another dashboard or another dollar.

Access the Full Webinar Now 

For the full conversation between Koza and Angeja, including how agentic workflows are being built in practice and where operational tooling is headed over the next two years, watch the on-demand recording of Observable AI: How Engineering Leaders Make AI Investments Pay Off now: https://www.geteverops.com/observable-ai-webinar-replay 

Alongside this impactful conversation, we recommend downloading our recent report, The Infrastructure Imperative, which draws on research from Gartner, McKinsey, CNCF, and Google DORA to explain why infrastructure maturity shapes AI outcomes as much as AI itself does.

Frequently Asked Questions

What was the Observable AI: How Engineering Leaders Make AI Investments Pay Off webinar about?

The webinar was a fireside chat between EverOps CEO Stephen Koza and Senior Delivery Director Joseph Angeja on why AI investments stall inside engineering organizations and how observability gives leaders the visibility to prove whether AI is actually paying off.

How can I watch the webinar?

The full conversation is available as an on-demand recording that you can sign up for and access here: https://www.geteverops.com/observable-ai-webinar-replay. The fireside chat between Joseph Angeja and Stephen Koza covers additional ground beyond what is in this article, including how agentic workflows are built in practice and how to make them sustainable over the long term. 

What is observability, and why does it matter for AI adoption?

Observability is the practice of measuring whether infrastructure, applications, and now AI investments are performing the way an organization expects. As AI tools spread across an organization, observability becomes the layer that shows who is using them, which workflows actually improved, and whether the spend is producing a measurable return.

What is the Observability Maturity Assessment?

The Observability Maturity Assessment was created by the EverOps team and is designed to give engineering leaders an honest read on their current state, covering what's deployed, what's being measured, and where the gaps in visibility and governance lie before an organization adds another tool or AI initiative.

How is the AI Opportunity Assessment different from the Observability Maturity Assessment?

The Observability Maturity Assessment focuses on the health and coverage of an organization's existing monitoring and data foundation. The AI Opportunity Assessment focuses specifically on where AI can create measurable value once that foundation is in place, making the two a natural pair for organizations early in AI adoption.

What services does EverOps provide beyond observability?

EverOps works with high-growth engineering teams as a cloud services partner, helping them ship faster, reduce risk, and optimize cloud spend, with observability and AI adoption support built into that broader practice.

How can my organization get started with EverOps?

The Observability Maturity Assessment or the AI Opportunity Assessment are the recommended starting points for most organizations. For additional questions about which fits best or what the best starting point is, you can also reach out to our team today here: https://www.everops.com/talk