From Blind Spots to Clear Insights: Monitoring AI Agent Conversations
A Deep Dive into the Observability Stack that Tamed our Multi-Cloud AI Agents

Building an AI demo is easy. Building a production-grade AI platform that manages costs, latency, and accuracy across thousands of conversations is a different beast entirely.
In our journey to build a robust AI agent platform, we quickly discovered that standard logging approaches, the kind that served us well for traditional web apps, simply weren't sufficient. When a user asks a question, it might trigger a chain reaction: a retrieval step (RAG), a tool execution, and a summarisation call, all routed through different providers like AWS Bedrock, Azure AI, or GCP Vertex.
In this blog, you will find a transparent look at the specific engineering challenges we faced, the open-source stack we built (using Grafana, Loki, and Prometheus), and the structured data strategy that turned our blind spots into clear insights.
The Challenge: Why AI Agents Break Traditional Monitoring
In a standard microservice architecture, you monitor HTTP 500s and average latency. If the server is up and responding fast, you’re green. In the world of AI Agents, "green" dashboards can hide massive problems.
We identified four distinct challenges that required a new approach:
- Cost Unpredictability: A standard API endpoint costs roughly the same per hit. An LLM call can range from a fraction of a cent to a dollar depending on the model (GPT-4o vs. Mini) and token depth.
- The "Black Box" Workflow: A single user request isn't just one database query; it is a multi-step conversation involving tools and reasoning loops.
- Provider Chaos: We use a multi-cloud strategy (AWS, Azure, GCP). Debugging errors across three different provider schemas is a nightmare without unification.
- Performance Nuance: "Total Duration" is a bad metric for streaming. A 10-second response is fine if the first token arrives in 200ms (TTFT). If the user waits 10 seconds for a blank screen, that's a churn event.
We realised we needed to move from asking "Is the service healthy?" to asking "What is the model thinking, and how much did that thought cost?"