AIOps is forecasting for your infrastructure. Learn to collect the signals, model what normal looks like, turn hundreds of alarms into one incident, and fix the common failures automatically, safely.
610 ms against an expected 250 ms. The static rule stays silent: it's under 700.
Grouped by time, service topology and a deploy at 02:58: cart-svc v2.4.1.
Known failure, reversible fix, high confidence: automated, with an audit trail.
Detected in 2 minutes, resolved in 9. Nobody was paged at 3am.
Illustrative data, modelled on Capstone 1's brief: a retailer whose checkout incidents averaged 51 minutes to resolve, with targets of under 2 minutes to detect and under 10 to fix.
Counted from the lesson files. Every chapter ends with an assessment: 60 of 25 questions and one of 10, all with explained answers.
A regional telecom loses a fibre trunk to a digger. Every modem downstream reports losing signal: 340 alarms, and 340 tickets. The operations centre spends the first 25 minutes sorting tickets instead of finding the cut. Making sure that doesn't happen again is the AIOps engineer's job. This is Capstone 3, from the course.
Metrics, logs and traces from every layer, through OpenTelemetry, Prometheus and CloudWatch, into streams and stores.
Lessons 1–2Model normal, with seasonality, so you catch what's unusual for 3am, not just what's high.
Lesson 3Group alerts by time and topology, add context, and measure alert precision and recall.
Lesson 4Runbooks, state machines, auto-scaling, probes and canaries. Automate only what's safe to.
Lesson 5Blameless post-mortems, SLOs and cost signals feed back into what you detect next.
Ch 4.8 · Lessons 7–8Few job ads say "AIOps engineer". The work is hired under these titles. Each card lists the lessons it draws on.
AIOps sits on top of DevOps. This bootcamp doesn't teach cloud, containers or CI/CD from scratch. Tick what you've actually used.
Used at work or in real projects, not just read about.
Eight core lessons, from raw signals to running an AIOps team. Pick a lesson to see its chapters, then a chapter to see its sections.
10 chapters setting up the toolchain, ending with Killercoda, a free browser sandbox for Kubernetes and container practice.
Three client scenarios, each in three phases with a reference solution. See them ↓
Eight modules that turn your capstone work into case studies, a post-mortem, SLOs and a public repo.
Chapter 3.2, Statistical Methods for Anomaly Detection: four written sections and a 25-question assessment. Try its central idea yourself on the right.
48 readings of database wait time. One is a 3,000 ms instrumentation glitch. Another, at 230 ms, is a real slowdown: about twice the normal 120 ms.
Each capstone is a client with real constraints. You work through it in three phases (a current-state assessment and signal inventory, the design, then a rollout and validation plan), then compare your reasoning with a reference solution: four decisions, each with the option chosen and the one rejected.
Eight modules make your capstone work something a hiring manager can judge, with guidance on anonymising anything confidential first.
Gross annual base salary. Salary sites rarely track "AIOps engineer", so these bands are based on site reliability and DevOps data, the titles this work is usually hired under. Read them as ranges, not promises.
Self-paced doesn't mean on your own. Mentor Bob is an AI study assistant in the corner of every section. It has already read the section you're on, so you can ask about it in your own words.
No. Lesson 3 teaches the ML you need from the operations side: forecasting, statistical and ML anomaly detection, ARIMA, Prophet and DeepAR, correlation and causality, log tokenisation, failure prediction and experiment tracking. Being able to read and change Python scripts helps a lot.
MLOps is about running machine-learning models in production: pipelines, deployment, monitoring the models. AIOps uses data, ML and automation to run IT systems: detecting anomalies, cutting alert noise and fixing incidents automatically. If you want the first, see MLOps Fundamental.
It's AWS-heavy: CloudWatch, Kinesis, Redshift, SageMaker, Systems Manager, Cost Anomaly Detection, Compute Optimizer. It also covers the open and vendor tools AIOps teams use alongside it: OpenTelemetry, Prometheus and PromQL, Elasticsearch and OpenSearch, PagerDuty and Opsgenie, ServiceNow and Jira, ArgoCD and FluxCD, and multi-cloud observability.
It's written material with real configuration, queries and code throughout, and a 25-question assessment at the end of every chapter. Chapters don't have separate practice sections: the hands-on design work is in the three capstones. Lesson 9 walks through installing the toolchain and points to Killercoda, a free browser sandbox.
It isn't exam prep and doesn't follow a certification syllabus. It overlaps with what AWS operations and observability certifications test, so it's useful background, but you'd still want an exam guide.
That depends on your pace and background, so we don't promise a number of weeks. For scale, there are about 820,000 words across the lessons, code included. Access is for life, so there's no deadline.
Yes. Lesson 7 covers cloud financial management and FinOps principles, AWS Cost Anomaly Detection, Compute Optimizer and Trusted Advisor, traffic and growth prediction, Lambda memory-versus-duration tuning, S3 Intelligent-Tiering and VPC endpoints.
Every lesson unlocks straight away in your dashboard, with Mentor Bob in each section. It's a one-time payment with lifetime access, so no subscription and no renewal.
All 8 lessons and 61 chapters, plus the installation guide, three capstones and the portfolio lessons, unlocked as soon as you enroll.
✓ You already own this bootcamp. It's waiting in your Active Bootcamps.