AIOps Fundamental BootcampFundamental · DevOps experience required

See the incident coming.Fix it before the pager does.

AIOps is forecasting for your infrastructure. Learn to collect the signals, model what normal looks like, turn hundreds of alarms into one incident, and fix the common failures automatically, safely.

€379€640−41% One-time payment
Lifetime access
See the 8-day forecast
LIVE checkout · p95 latency Wed 00:00
03:00Out of forecast

610 ms against an expected 250 ms. The static rule stays silent: it's under 700.

03:0123 alerts → 1 incident

Grouped by time, service topology and a deploy at 02:58: cart-svc v2.4.1.

03:02Runbook: roll back

Known failure, reversible fix, high confidence: automated, with an audit trail.

03:09Back inside the band

Detected in 2 minutes, resolved in 9. Nobody was paged at 3am.

Static rule: 0 · 0 incidents missedForecast band: 0 · 0 false pages

Illustrative data, modelled on Capstone 1's brief: a retailer whose checkout incidents averaged 51 minutes to resolve, with targets of under 2 minutes to detect and under 10 to fix.

0Core lessons
0Chapters
0Written sections
0Assessment questions
0Capstones
0Portfolio modules
0Install-guide chapters
0Words, code included

Counted from the lesson files. Every chapter ends with an assessment: 60 of 25 questions and one of 10, all with explained answers.

Bulletin 01 · What the job is

One storm. 340 alarms. One incident.

A regional telecom loses a fibre trunk to a digger. Every modem downstream reports losing signal: 340 alarms, and 340 tickets. The operations centre spends the first 25 minutes sorting tickets instead of finding the cut. Making sure that doesn't happen again is the AIOps engineer's job. This is Capstone 3, from the course.

0Tickets created
0Incidents seen by the NOC
–Grouped correctly?

The loop you'll build, and where each part is taught

Observe

Metrics, logs and traces from every layer, through OpenTelemetry, Prometheus and CloudWatch, into streams and stores.

Lessons 1–2
Detect

Model normal, with seasonality, so you catch what's unusual for 3am, not just what's high.

Lesson 3
Correlate

Group alerts by time and topology, add context, and measure alert precision and recall.

Lesson 4
Act

Runbooks, state machines, auto-scaling, probes and canaries. Automate only what's safe to.

Lesson 5
Learn

Blameless post-mortems, SLOs and cost signals feed back into what you detect next.

Ch 4.8 · Lessons 7–8
Bulletin 02 · Where it takes you

Six stations, one skill set.

Few job ads say "AIOps engineer". The work is hired under these titles. Each card lists the lessons it draws on.

Bulletin 03 · Who it's for

Check your readiness forecast.

AIOps sits on top of DevOps. This bootcamp doesn't teach cloud, containers or CI/CD from scratch. Tick what you've actually used.

Conditions on the ground

Used at work or in real projects, not just read about.

Your outlook

Clear skies if you…

  • work in DevOps, SRE or cloud operations and you're tired of firefighting the same incidents
  • run Prometheus, CloudWatch or Grafana and want smarter alerting than fixed thresholds
  • want to add machine learning to your operations work, applied to real operational data
  • want to build self-healing automation that leadership will actually trust
  • want written material you'll reopen on the job, and capstones to show employers

Stormy if you…

  • have never worked with cloud, containers or CI/CD. Start with DevOps Beginner
  • want a pure data-science course. This is ML applied to IT operations
  • want to operate ML models in production. That's MLOps Fundamental
  • need video, live classes or a cohort. Everything here is written and self-paced
Bulletin 04 · Curriculum

The 8-day forecast.

Eight core lessons, from raw signals to running an AIOps team. Pick a lesson to see its chapters, then a chapter to see its sections.

Lesson 9 · Installation guide

10 chapters setting up the toolchain, ending with Killercoda, a free browser sandbox for Kubernetes and container practice.

Lesson 10 · Capstones

Three client scenarios, each in three phases with a reference solution. See them ↓

Lesson 11 · Portfolio

Eight modules that turn your capstone work into case studies, a post-mortem, SLOs and a public repo.

Bulletin 05 · Inside one chapter

One chapter, taken apart.

Chapter 3.2, Statistical Methods for Anomaly Detection: four written sections and a 25-question assessment. Try its central idea yourself on the right.

Chapter 3.2 · 5 sections16,560 words
Try it · from Section A

Which detector catches the slowdown?

48 readings of database wait time. One is a 3,000 ms instrumentation glitch. Another, at 230 ms, is a real slowdown: about twice the normal 120 ms.

–Centre
–Spread
–Score of the 230 ms slowdown

Bulletin 06 · Capstones

Three advisories. Your call.

Each capstone is a client with real constraints. You work through it in three phases (a current-state assessment and signal inventory, the design, then a rollout and validation plan), then compare your reasoning with a reference solution: four decisions, each with the option chosen and the one rejected.

Lesson 11 · Portfolio

Then turn them into a portfolio.

Eight modules make your capstone work something a hiring manager can judge, with guidance on anonymising anything confidential first.

Bulletin 07 · Salaries

The pay outlook.

Gross annual base salary. Salary sites rarely track "AIOps engineer", so these bands are based on site reliability and DevOps data, the titles this work is usually hired under. Read them as ranges, not promises.

Bulletin 08 · Mentor Bob

Stuck at 2am? Bob's on air.

Self-paced doesn't mean on your own. Mentor Bob is an AI study assistant in the corner of every section. It has already read the section you're on, so you can ask about it in your own words.

  • Answers from the section you're reading, not from the whole internet
  • Explains a concept another way, or with a new example
  • There in every lesson, at any hour. Included, not an upsell
MENTOR BOBChapter 3.2 · Section A
Example conversations, written from the chapters they're set in.
Bulletin 09 · FAQ

Questions from the desk.

Do I need machine learning experience?

No. Lesson 3 teaches the ML you need from the operations side: forecasting, statistical and ML anomaly detection, ARIMA, Prophet and DeepAR, correlation and causality, log tokenisation, failure prediction and experiment tracking. Being able to read and change Python scripts helps a lot.

How is this different from MLOps?

MLOps is about running machine-learning models in production: pipelines, deployment, monitoring the models. AIOps uses data, ML and automation to run IT systems: detecting anomalies, cutting alert noise and fixing incidents automatically. If you want the first, see MLOps Fundamental.

Is it all AWS?

It's AWS-heavy: CloudWatch, Kinesis, Redshift, SageMaker, Systems Manager, Cost Anomaly Detection, Compute Optimizer. It also covers the open and vendor tools AIOps teams use alongside it: OpenTelemetry, Prometheus and PromQL, Elasticsearch and OpenSearch, PagerDuty and Opsgenie, ServiceNow and Jira, ArgoCD and FluxCD, and multi-cloud observability.

Is it hands-on?

It's written material with real configuration, queries and code throughout, and a 25-question assessment at the end of every chapter. Chapters don't have separate practice sections: the hands-on design work is in the three capstones. Lesson 9 walks through installing the toolchain and points to Killercoda, a free browser sandbox.

Will it prepare me for a certification?

It isn't exam prep and doesn't follow a certification syllabus. It overlaps with what AWS operations and observability certifications test, so it's useful background, but you'd still want an exam guide.

How long will it take?

That depends on your pace and background, so we don't promise a number of weeks. For scale, there are about 820,000 words across the lessons, code included. Access is for life, so there's no deadline.

Does it cover cost and FinOps?

Yes. Lesson 7 covers cloud financial management and FinOps principles, AWS Cost Anomaly Detection, Compute Optimizer and Trusted Advisor, traffic and growth prediction, Lambda memory-versus-duration tuning, S3 Intelligent-Tiering and VPC endpoints.

What happens after I pay?

Every lesson unlocks straight away in your dashboard, with Mentor Bob in each section. It's a one-time payment with lifetime access, so no subscription and no renewal.

Tomorrow's forecast

Build systems that fix themselves.

All 8 lessons and 61 chapters, plus the installation guide, three capstones and the portfolio lessons, unlocked as soon as you enroll.

AccessLifetime
PaceYour own
MentorBob, 24/7
Pages at 3amFewer
Today's price
€379€640 · −41% · one-time