The instrument stack
- Copilot usage reports (M365 admin): active users per app surface, trend lines, last-activity — the utilisation layer, with the standard reporting caveats (lag, pseudonymisation switch — the Teams curriculum's usage-reporting rules apply verbatim).
- Copilot analytics/dashboard (Viva-powered surfaces): adoption segmentation, feature-level usage, and — where enabled — impact views correlating usage with collaboration patterns. Powerful; treat correlation claims carefully in exec material.
- Your own task math (the credible layer): pilot-cohort measurements of specific tasks — recap minutes per meeting-heavy role, first-draft cycle time, search-to-found time. Small n, real numbers, YOUR data.
The ROI narrative that survives finance
Bad: 'AI saves 14 hours/month' (whose? doing what? says who?). Good: three task lines with before/after from your pilot ('recap: 20 min -> 4 min × 9 meetings/week × 140 meeting-heavy staff'), utilisation showing the assumption holds at scale, seat-recycling showing discipline, and the honest column: where it DIDN'T help (Excel-heavy roles returning seats is a credibility asset, not a failure).
Two structural rules: measure per role family, not org-average (dilution hides both wins and waste — the pilot-to-scale wave logic), and separate licence value from agent value (credits burn has its own ROI line — cost-and-credits concept — or it silently eats the seat story).
What to watch (proofs)
- Weekly-active per seat by role family: THE utilisation cut; org-wide averages are for press releases.
- The recycling funnel: inactive-30d list -> nudge -> reclaim counts, monthly — discipline the CFO can see.
- Task-math ledger: each measured task with method + n + date — the renewal appendix that ends 'prove it' meetings.
- Dashboard-vs-reality spot checks: analytics claims sampled against the cohort's lived reports — calibrate the instrument before quoting it upward.