You answered successfully, but not correctly - observing AI hijinks
Submitted to Codemotion Milan 2026 on . Still waiting to hear back.
- Level
- Intermediate
- Length
- Standard Talk (30 min)
- Track
- Cloud and DevOps
Abstract, as submitted
An AI-powered feature can hit every SLO and still hand the user a wrong, hallucinated, off-tone, or just not very good answer. Token counts, latency, and HTTP 200s won’t tell you that.
This talk walks through what to instrument inside an AI-using app to catch and correct quality problems, not just performance and cost ones. I’ll demonstrate how to instrument a typical enterprise app and how to monitor it, covering the standard signals (tokens, cost, tool calls, model/version), the less standard ones (retrieval quality for RAG, tool-call failure rate, user feedback as a first-class telemetry attribute), and what dashboards and alerts actually catch a degraded model in production.
This is the text that went into the CFP form, kept as submitted - not a later rewrite of it. Talks get retitled and reworked between submission and stage, so what was actually delivered may differ.