LearnMicrosoft 365 Copilot › 8 · Operations & adoption

Reliability and outage playbooks for AI

June 2026 settled the argument: Copilot outages are now office-availability incidents. The Teams resilience discipline applies — outage classes, degraded-mode cards, comms — with AI's own twist: quality can fail without availability failing.

The outage classes, AI edition

Class Signature Degraded mode
Copilot service down Chat errors, timeouts, dead sidebars (the June 2026 pattern) Work continues without AI — the card says so plainly, and the org discovers its real dependency level
A pipeline stage down Grounding broken (answers lose tenant context) or specific surfaces failing while others work 'Web tab works, work tab degraded' — teach the tab distinction (chat-surfaces concept) as a diagnostic
Dependency outage Teams/EXO/SPO incident starving Copilot of sources — the constellation signature again Route the ticket to the REAL outage; Copilot is the symptom
Quality regression Available but worse: model/feature changes shifting behaviour overnight The AI-specific class — see below
Your config self-inflicted Correlates to a change: DLP policy overshoot, connector dead, agent disabled The drift log answers it — serv365's before/after IS the rollback spec

The AI-specific twist: quality incidents

Model and feature updates ship continuously under the pipeline (foundations concept) — behaviour can change with no advisory and no config drift. Symptoms: prompt patterns that worked now don't; output format shifts breaking downstream habits. Playbook: a small golden prompt set run on cadence (your Copilot canary — the Teams curriculum's probe logic applied to answers), Message Center watched for Copilot-tagged changes (this site's core loop — the change-aware banner above this article is that watch, running), and comms that name it honestly: 'behaviour changed upstream; here's the adjusted pattern'.

The playbook artifacts (pre-written, tested)

Degraded-mode cards per class (one paragraph each: what works, what to tell users), the non-AI fallback statement for critical workflows that quietly grew AI dependencies (drafting, triage — inventory them BEFORE the outage does), status comms via the channel that isn't down, and the post-incident habit: YOUR user-impact timeline vs the advisory, feeding the canary set.

What to watch (proofs)

  • Advisory-vs-felt latency: when service health posted vs when tickets started — if users beat the advisory, the golden-prompt canary earns its slot in the scheduler.
  • The canary run log: golden prompts, scored pass/fail, dated — quality regressions become detected events, not vibes-drift.
  • Dependency inventory: workflows with AI in the critical path, reviewed quarterly — the list that sizes each outage class's real impact.
  • Drill evidence: one tabletop per half-year including a QUALITY incident — the class teams haven't rehearsed anywhere yet.

Discussion

No messages yet — start the thread.

Sign in with your email to join the discussion — we send a one-time link, no password.