Managed IT · Modern Workplace Management

Microsoft 365 Service Health Monitoring

When email stops flowing or people cannot sign in, the help desk’s phones light up and the pressure is immediate.

13 min read
Content owner
Insyto Content Team
Editorial reviewer
Ritesh Mhatre
Next review
To be scheduled
Technical reviewer
Navish Ansari
Last reviewed
Review pending
Technical level
Intermediate · IT directors, Microsoft 365 administrators

Modern Workplace Management · Microsoft 365 Service Health Monitoring

Executive Summary

When email stops flowing or people cannot sign in, the help desk’s phones light up and the pressure is immediate. The single most valuable thing an IT team can do in that moment is answer one question fast: is this our problem, or Microsoft’s? If it is a Microsoft service incident, every minute spent troubleshooting the tenant is wasted, and the right response is to communicate to users and track Microsoft’s resolution. If it is a local problem, the clock is ticking on a fix that is genuinely yours. Service health monitoring is the discipline that answers that question in minutes rather than hours — and it turns an anxious, blame-filled outage into a calm, managed event.

Microsoft publishes the health of every service in real time. The Service Health dashboard in the Microsoft 365 admin center shows which services are healthy, which have active incidents or advisories, whether an issue originated at Microsoft or in your environment, and a live feed of updates as Microsoft works toward resolution. Alongside it, a set of complementary sources — a public status page for when you cannot even sign in, the Message center for upcoming changes, and Windows release health — cover the scenarios the dashboard does not. Used well, and wired to notifications, these mean IT learns about an incident before the first support ticket arrives.

This guide explains how to monitor Microsoft 365 service health: where to look, how to tell an incident from an advisory, what the status states mean, how to set up notifications, and how to respond and communicate when Microsoft has an issue. It is deliberately narrow — the holistic monitoring of your tenant’s own security, identity, and usage health is covered in the companion Microsoft 365 health monitoring guide; this one is about tracking Microsoft’s service status and changes. Because the admin center evolves, verify specifics against the linked documentation.

Who should read this:

  • CIOs, CTOs, and IT directors accountable for service availability and communication
  • Microsoft 365 administrators and help-desk teams who triage outages
  • Service and operations managers who own incident communication
  • SMB decision-makers evaluating managed service-health monitoring

Where do you check service health?

The primary source is the Service Health dashboard in the Microsoft 365 admin center (Health > Service health), which shows the health state of every service and any active incidents or advisories for your tenant. But it is not the only source, and each covers a different scenario — most importantly, what to do when you cannot sign in to the admin center at all.

Where to watch and how alerts reach you

Where to watch and how alerts reach you: sources include the Service Health dashboard (admin center, your tenant’s issues), status.cloud.microsoft (when you can’t sign in), the Message center (upcoming changes and maintenance), Windows release health (Windows update issues), and @MSFT365Status (broad events on X); IT or the MSP monitors the dashboard and notifications with the Service Support or Helpdesk admin role, receiving email alerts (up to two addresses per admin), M365 Admin mobile app push notifications, and per-issue subscriptions for major incidents.

SourceWhat it showsWhen to use it
Service Health dashboardYour tenant’s service state, active incidents & advisoriesThe default — daily monitoring and outage triage
status.cloud.microsoftBroad, public service statusWhen you cannot sign in to the admin center
Message centerUpcoming changes and planned maintenanceTo prepare for changes before they land
Windows release healthKnown issues with Windows updatesWhen a Windows update causes problems
@MSFT365Status (on X)Notices of certain broad eventsAs a secondary, at-a-glance signal

Viewing service health requires an appropriate role — the Service Support admin and Helpdesk admin roles can view it, among others — and the Microsoft 365 Admin mobile app provides the same view with push notifications, which is often the fastest way to learn of an incident.

How do you tell an incident from an advisory?

Not every service issue is an outage. Microsoft classifies problems as either advisories or incidents, and the distinction shapes your response: an advisory usually means a workaround exists and most people can keep working, while an incident means something is genuinely down.

AttributeAdvisoryIncident
SeverityLower — a known problem, service still availableCritical — service or a major function is unavailable
User impactSome users; often intermittent or limited in scopeNoticeable; e.g., cannot send/receive email or sign in
WorkaroundOften availableTypically none until resolved
Your responseNote it, share the workaround, monitorCommunicate proactively, track closely to resolution

Reading this correctly prevents both over-reaction (declaring a crisis over a minor advisory) and under-reaction (dismissing an incident as noise).

What do the status states mean?

Once an issue is posted, Microsoft attaches a live status that changes as the investigation and fix progress. Reading the status tells you exactly where things stand without opening a support case — whether Microsoft is still investigating, actively restoring, or done.

How a service issue progresses

How a service issue progresses: investigating (potential issue, gathering info) leads to service degradation (slow or intermittent) or service interruption (can’t access the service), then restoring service (fix identified, applying it), extended recovery (rolling out to all systems), service restored (confirmed healthy), and finally a post-incident report (root cause and prevention); other statuses include investigation suspended (Microsoft needs more info) and false positive (service confirmed healthy), while an advisory means a problem affecting some users with the service still available — and planned maintenance is not shown here, but tracked in the Message center.

StatusWhat it meansWhat you do
InvestigatingMicrosoft is aware and gathering information on scopeWatch for the next update; hold on troubleshooting
Service degradationConfirmed issue — slow, intermittent, or a feature brokenWarn affected users; expect variable impact
Service interruptionUsers cannot access the service; reproducibleCommunicate proactively; this is a real outage
Restoring serviceCause identified; corrective action under wayReassure users; resolution is in progress
Extended recoveryFix rolling out but taking time to reach all systemsSet expectations that recovery is gradual
Investigation suspendedMicrosoft needs more information from customersProvide any requested data or logs
Service restoredCorrective action resolved the problem; service healthyConfirm with users; await the post-incident report
False positiveInvestigation found the service healthy; no real impactLook elsewhere — the cause is not this service
Post-incident reportA PIR with root cause and prevention has been publishedRead it; note any actions for your side

An important detail: planned maintenance is deliberately not shown on the Service Health dashboard. Upcoming changes and maintenance are tracked in the Message center — filtering to messages categorized as “Plan for change” — which is why monitoring both is necessary to have the full picture.

How do you get notified — and respond?

The goal of service health monitoring is to hear about an incident before your users do, so you can get ahead of it. That requires configuring notifications rather than relying on someone happening to look at the dashboard.

Is it us or Microsoft? When users report a problem, check the Service Health dashboard (or status.cloud.microsoft if you are locke

Is it us or Microsoft? When users report a problem, check the Service Health dashboard (or status.cloud.microsoft if you are locked out): if the issue appears under “Active issues Microsoft is working on,” stop troubleshooting and instead communicate and track resolution; if it appears under “Issues for your organization to act on,” it is yours to remediate following the guidance provided; and if it is not listed anywhere, use “Report an issue” or troubleshoot locally.

Set up email notifications (up to two addresses per admin, filterable by incidents or advisories and by service), use the mobile app for push, and subscribe to updates on individual major incidents. When an alert arrives and confirms a Microsoft issue, the triage question resolves immediately — and the response becomes a clean sequence rather than a scramble.

Responding to a Microsoft incident

Responding to a Microsoft incident: detect (alert or dashboard, confirm it’s Microsoft), assess impact (who and what is affected, read the User Impact field), communicate (tell users proactively and share any workaround), track (follow status updates to resolution), and close and review (confirm restored and read the post-incident report) — proactive communication is the whole value, because users forgive an outage they were warned about.

The response sequence is detect, assess impact, communicate, track, and close. The value is overwhelmingly in the communication step: users forgive an outage they were warned about and kept informed on, and they lose confidence in one that IT seemed not to know about. A one-line proactive message — “Microsoft is investigating an Exchange Online incident affecting email; we’re monitoring and will update you” — transforms the experience and cuts the support-ticket volume dramatically.

How is service health monitoring run as a routine?

Service health monitoring is a light but constant discipline: a daily rhythm of watching, triaging, communicating, and tracking, plus regular attention to the Message center so upcoming changes never arrive as surprises.

Service health monitoring as a routine

Service health monitoring as a routine: watch daily (dashboard and alerts), triage (us or Microsoft?), communicate (to users), track to close, and review changes (Message center) — a daily rhythm plus planned-change reviews keeps the business ahead of both incidents and updates.

SituationActionToolOwner
Daily health checkScan dashboard for active issuesService Health dashboardIT / MSP service desk
New incident alertConfirm origin; assess impactDashboard + email/pushIT / MSP service desk
Confirmed Microsoft incidentCommunicate to users; track to closeDashboard + comms channelIT / MSP service desk
Issue not listedReport an issue; troubleshoot locally“Report an issue” formIT / MSP engineering
Upcoming change / maintenanceReview and planMessage center (“Plan for change”)IT / MSP + vCIO
Post-incidentRead the PIR; note follow-upsService Health historyIT / MSP + vCIO

Managed service-health monitoring model (RACI)

Delivered as a managed service, service health monitoring is an accountable, continuously operated capability. This RACI defines who does what, the tool, the cadence, and the impact.

ActivityResponsible (MSP/IT)Accountable (CIO)ToolCadenceSLA / impact
Configure notifications & rolesMSP service deskCIOAdmin centerOn onboardingAlerts reach the right people
Monitor the dashboardMSP service deskCIOService Health dashboardDaily / on alertIncidents caught early
Triage origin (us vs Microsoft)MSP service deskCIODashboardOn every incidentNo wasted troubleshooting
Communicate to usersMSP service deskCIOEmail / Teams / status pageDuring incidentsCalm, informed users
Track changes in Message centerMSP + vCIOCIOMessage centerWeeklyNo surprise changes
Review post-incident reportsMSP vCIOCIOService Health historyPer PIRLessons captured

Implementation checklist

  • Service health viewing roles (Service Support / Helpdesk admin) are assigned
  • Email notifications are configured for incidents and advisories
  • The Microsoft 365 Admin mobile app is installed for push notifications
  • The team knows to use status.cloud.microsoft when locked out
  • A triage step confirms whether an issue is Microsoft’s or the tenant’s
  • A user-communication channel and template are ready for incidents
  • The “Report an issue” process is understood for unlisted problems
  • The Message center is reviewed weekly for upcoming changes
  • Windows release health is monitored for update-related issues
  • Post-incident reports are read and any follow-up actions captured
  • A daily dashboard check is part of the operational routine

Best practices

  • Configure notifications so IT hears about incidents before users do.
  • Triage every outage against the dashboard first — is it us or Microsoft?
  • Distinguish incidents from advisories and respond proportionately.
  • Communicate proactively to users during any real incident.
  • Use status.cloud.microsoft as the fallback when you cannot sign in.
  • Monitor the Message center separately for planned changes and maintenance.
  • Watch Windows release health for update-caused problems.
  • Read post-incident reports and act on any customer-side follow-ups.
  • Make a daily service-health check a standing routine.

Common mistakes

  • Troubleshooting the tenant for hours before checking service health.
  • Relying on users to report outages instead of configuring notifications.
  • Confusing a minor advisory with a full incident, or vice versa.
  • Going silent during an incident, leaving users to assume IT is unaware.
  • Forgetting the public status page exists for sign-in outages.
  • Ignoring the Message center, then being caught out by a planned change.
  • Never reading post-incident reports, so lessons are lost.
  • Treating service health as something to check only during a crisis.

Frequently asked questions

What is Microsoft 365 service health monitoring?

It is the practice of tracking the status of Microsoft’s own services through the Service Health dashboard and related sources — detecting incidents and advisories, telling whether a problem is Microsoft’s or yours, communicating to users, and tracking resolution.

How is it different from health monitoring?

Service health monitoring is specifically about Microsoft’s service status and changes. The broader Microsoft 365 health monitoring discipline covers your tenant’s own security, identity, usage, and data health. This guide is the narrower, incident-focused one.

How do I know if an outage is Microsoft’s fault or ours?

Check the Service Health dashboard. If the issue appears under “Active issues Microsoft is working on,” it is Microsoft’s. If it appears under “Issues for your organization to act on,” it is yours. If it is not listed, you can report it or investigate locally.

What is the difference between an incident and an advisory?

An advisory is a lower-severity, known problem where the service is still available, often with a workaround. An incident is a critical issue where the service or a major function is unavailable, with noticeable user impact.

What do I do if I can’t sign in to the admin center?

Use the public status page at status.cloud.microsoft, which shows broad service status without requiring sign-in, and follow @MSFT365Status for notices of certain events.

Where do I find out about upcoming changes?

In the Message center, not the Service Health dashboard — planned maintenance and changes are not shown on the service health page. Filter Message center to “Plan for change” to see what is coming and how to prepare.

How do we get notified of incidents?

Configure email notifications in the Service Health dashboard (up to two addresses per admin, filterable by service and by incident or advisory), use the Microsoft 365 Admin mobile app for push notifications, and subscribe to updates for individual major incidents.

Conclusion

When a Microsoft 365 service goes down, the difference between chaos and calm is knowing, quickly, that it is Microsoft’s incident and not yours — and telling your users so before they tell you. Service health monitoring provides that: the Service Health dashboard and its companion sources reveal what is happening, whether the cause is Microsoft or your environment, and how the resolution is progressing, while notifications ensure IT hears first. The response then becomes a clean sequence of detect, assess, communicate, track, and close, with proactive communication doing most of the work of keeping users confident.

The path forward is simple and high-value: assign the viewing roles, turn on notifications, build a triage-and-communicate habit, watch the Message center for changes, and read the post-incident reports. Run this way — as a light daily routine rather than a crisis-only scramble — service health monitoring turns Microsoft’s inevitable incidents into managed, well-communicated events. For monitoring the health of your own tenant across security, identity, and usage, see the companion Microsoft 365 health monitoring guide.

Authoritative references

All sources are official Microsoft documentation. Verify current features before acting; the admin center changes frequently. Source access date: 28 July 2026.

Next step

Discuss your environment with Insyto

Talk through the practical next steps for your Microsoft and IT environment.