Energy · 2025
AI Productionization & Observability at E.ON
AI productionization for 8M+ energy customers under KRITIS. Six monitoring tools consolidated into one platform. 2M+ monthly LLM interactions at 92% accuracy.
Contents

AI productionization for 8M+ energy customers with no MLOps baseline. Six monitoring tools, no shared view when the grid broke.
Production AI at enterprise scale. One observability platform across IT, OT, and grid with agentic diagnostics.
- LLM interactions at 92% accuracy
- 2M+ monthly
- Model deployment
- 45% faster
- Monitoring tools consolidated
- 6 → 1
The shift
- Grid observability
Engagement
- Client
- E.ON
- Role
- AI-native PO | Technical Manager, Observability Platform
- Timeframe
- 2025
- Industry
- Energy
E.ON runs 1.6 million km of energy networks for 47 million customers across 17 countries. When something breaks, people lose power. My engagement had two remits: get AI into production for 8M+ energy customers, and fix the observability foundation underneath it. The second turned out to be a precondition for the first. AI-assisted operations reason over telemetry, and E.ON's telemetry was split across six monitoring tools.
Six tools across IT, OT and grid operations, none with the full picture. Decentralised renewables were making the grid more volatile by the quarter, and alert fatigue was burning out on-call teams. Ripping the six out would have thrown away years of scripts, integrations and habits wrapped around them, so we consolidated underneath them instead. OpenTelemetry became the single, vendor-neutral standard for traces, metrics and logs, with New Relic as the platform layer: APM, infrastructure monitoring, Kubernetes auto-discovery via eBPF, AI-assisted root cause analysis.
The part I pushed hardest was semantic conventions: getting teams to name things consistently. Automated correlation only works on labels that agree, and the AI diagnostics we wanted to build later depended on it.
For critical infrastructure, getting AI into production is mostly governance and delivery work. The MLOps pipelines we built replaced manual release steps with automated testing, and model deployment time fell 45% while production incidents fell 60%. The customer-service LLM deployment now handles 2M+ monthly interactions at 92% accuracy. All of it ships inside the regulatory envelope: BSI C5 controls, NIS2, KRITIS/IT-SiG 2.0, with DSFA documentation for every AI-driven use of customer data.
On that baseline we piloted the agentic layer: MCP-connected LLMs against New Relic in AWS and Azure, doing automated diagnostics and remediation with human-in-the-loop guardrails. A working group from observability, operations and data science checked that the collected metrics mapped to defined SLOs, so the pilot had to prove its value in numbers the operators already trusted.
Pathpoint mapped customer journeys, order flows and backend services to shared KPIs, and grid operators got real-time state estimation and congestion detection. An incident used to arrive as "this service is down". Now it arrives as "this is affecting X customers in their billing flow", and the prioritisation arguments got shorter.
Self-service dashboards and templates took observability to thousands of users without queueing on a central team. A reusable onboarding playbook covering collectors, alerts, SLOs and dashboards meant new teams reached production-grade observability in days.
- Production AI serving 8M+ energy customers: 2M+ monthly LLM interactions at 92% accuracy
- 45% faster model deployment and 60% fewer production incidents through MLOps pipelines with automated testing
- Six monitoring tools consolidated into one OpenTelemetry platform spanning IT, OT, and grid
- Agentic diagnostics piloted with MCP-connected LLMs, gated by human-in-the-loop guardrails
- BSI C5, NIS2, and KRITIS-compliant governance with DSFA documentation
How we decided what an agent may do on its own, and what the guardrails around it looked like, is written up in The AI-native Platform Playbook.
2M+ monthly LLM interactions at 92% accuracy. 45% faster model deployment. 6 → 1 monitoring tools consolidated.
