The outage classes, AI edition
| Class | Signature | Degraded mode |
|---|---|---|
| Copilot service down | Chat errors, timeouts, dead sidebars (the June 2026 pattern) | Work continues without AI — the card says so plainly, and the org discovers its real dependency level |
| A pipeline stage down | Grounding broken (answers lose tenant context) or specific surfaces failing while others work | 'Web tab works, work tab degraded' — teach the tab distinction (chat-surfaces concept) as a diagnostic |
| Dependency outage | Teams/EXO/SPO incident starving Copilot of sources — the constellation signature again | Route the ticket to the REAL outage; Copilot is the symptom |
| Quality regression | Available but worse: model/feature changes shifting behaviour overnight | The AI-specific class — see below |
| Your config self-inflicted | Correlates to a change: DLP policy overshoot, connector dead, agent disabled | The drift log answers it — serv365's before/after IS the rollback spec |
The AI-specific twist: quality incidents
Model and feature updates ship continuously under the pipeline (foundations concept) — behaviour can change with no advisory and no config drift. Symptoms: prompt patterns that worked now don't; output format shifts breaking downstream habits. Playbook: a small golden prompt set run on cadence (your Copilot canary — the Teams curriculum's probe logic applied to answers), Message Center watched for Copilot-tagged changes (this site's core loop — the change-aware banner above this article is that watch, running), and comms that name it honestly: 'behaviour changed upstream; here's the adjusted pattern'.
The playbook artifacts (pre-written, tested)
Degraded-mode cards per class (one paragraph each: what works, what to tell users), the non-AI fallback statement for critical workflows that quietly grew AI dependencies (drafting, triage — inventory them BEFORE the outage does), status comms via the channel that isn't down, and the post-incident habit: YOUR user-impact timeline vs the advisory, feeding the canary set.
What to watch (proofs)
- Advisory-vs-felt latency: when service health posted vs when tickets started — if users beat the advisory, the golden-prompt canary earns its slot in the scheduler.
- The canary run log: golden prompts, scored pass/fail, dated — quality regressions become detected events, not vibes-drift.
- Dependency inventory: workflows with AI in the critical path, reviewed quarterly — the list that sizes each outage class's real impact.
- Drill evidence: one tabletop per half-year including a QUALITY incident — the class teams haven't rehearsed anywhere yet.