You answered successfully, but not correctly - observing AI hijinks

Submitted to Infobip Shift Zadar 2026 on . Rejected .

Level
Intermediate
Track
Cloud Development

Abstract, as submitted

An AI-powered feature can hit every SLO and still hand the user a wrong, hallucinated, off-tone, or just not very good answer. Token counts, latency, and HTTP 200s won’t tell you that.

This talk walks through what to instrument inside an AI-using app to catch and correct quality problems, not just performance and cost ones. I’ll demonstrate how to instrument a typical enterprise app and how to monitor it, covering the standard signals (tokens, cost, tool calls, model/version), the less standard ones (retrieval quality for RAG, tool-call failure rate, user feedback as a first-class telemetry attribute), and what dashboards and alerts actually catch a degraded model in production.

This is the text that went into the CFP form, kept as submitted - not a later rewrite of it. Talks get retitled and reworked between submission and stage, so what was actually delivered may differ.