Anomaly detection

MCPs for Observability Stacks

MCPs for Observability Stacks

Diana Todea, Head of Developer Relations at VictoriaMetrics, demonstrates how Model-Controller-Pair (MCP) servers enhance observability stacks for faster, smarter troubleshooting. She highlights the use of AI-driven anomaly detection, natural language querying, and customizable "skills" to investigate production issues, manage metrics, and generate alerts, showcasing practical applications with two interconnected MCP servers.

Learned Execution Graphs for Anomaly Detection & Drift in APIs — Ritvik Pandya, JP Morgan Chase

Learned Execution Graphs for Anomaly Detection & Drift in APIs — Ritvik Pandya, JP Morgan Chase

Ritvik Pandya from JP Morgan introduces 'learn execution graphs'—short-lived DAGs representing API request processing—to detect anomalies and drift in high-throughput systems. This approach localizes performance issues to exact nodes, identifies skipped or reordered steps, and differentiates between transient anomalies and fundamental system drift (structural, volume, or behavioral). It leverages per-client baselines and OpenTelemetry data to reduce mean time to discovery (MTTD) from hours to seconds, enhancing system reliability and observability.

AI, AIOps & Agentic AI in Data Storage Observability

AI, AIOps & Agentic AI in Data Storage Observability

Prabira Acharya explains that managing data storage without observability is like driving a car without a dashboard. The talk outlines the seven pillars of storage observability, the critical role of AI in analyzing vast amounts of data for anomaly detection and predictive analytics, and the evolution toward agentic AIOps for creating self-healing and self-managing storage infrastructures.

AI Agents: Transforming Anomaly Detection & Resolution

AI Agents: Transforming Anomaly Detection & Resolution

Martin Keen explores how agentic AI can significantly reduce IT downtime and Mean Time To Repair (MTTR) by moving beyond naive data dumps and embracing context-aware analysis. The key lies in using topology-aware correlation to curate relevant data for an AI agent, which can then systematically identify the root cause, provide explainable insights, and generate actionable remediation steps, ultimately augmenting human SREs rather than replacing them.