Availability SLA
Overview
Managers and customers with a service level agreement ask one question every month: what was the uptime? Alerts say when an instance restarted or could not be reached, but not how much of the month it was available.
The Availability SLA report answers that. It reads the availability history the Database Health Monitor Service keeps for each instance it monitors and works out, for the month you pick:
- the uptime percent, against a target you choose (99% to 99.99%),
- the number of outages, the total downtime and the longest outage,
- the mean time between failures,
- a heat strip with one cell per day, colored by that day’s availability,
- a grid of every outage, restart and monitor gap in the month, with the reason.
The Estate view lists every instance connected in the server tree with this month’s and last month’s uptime, worst first, and exports to Excel with one row per instance for SLA reporting.
Where to find it
| Route | How |
|---|---|
| Server tree | Right-click the server → Instance Level Reports → Availability SLA |
| Instance reports navigator | Recovery group, after Availability Groups |
| Related Links bar | From Error Log |
| Report arrows | Previous is Availability Groups, next is Backup Status |
Where the history comes from
The Database Health Monitor Service checks every instance it monitors every thirty seconds and only stores changes, in the instance’s own DBHealthHistory database (tables AvailabilityMonitor and AvailabilityOutage, added in history version 1590):
| Type | How it is found | From and to |
|---|---|---|
| Unreachable | The service could not connect | The first failed check to the next check that connected |
| Restart | sqlserver_start_time moved on between two checks, including a restart quick enough to fall between them |
The shutdown message in the previous error log (or the last check that connected) to the new start time |
| Service Stopped | A restart whose previous error log says SQL Server was stopped on request | As for Restart |
| Monitor Gap | No check connected and none failed for more than three minutes: the service was stopped or restarting | Not counted as up or down |
A restart and the failed checks around it are one outage, not two, so stopping an instance for ten minutes shows as one outage of about ten minutes. The shutdown reason (a Windows shutdown, a stop request, or an unexpected restart with no clean shutdown message) needs securityadmin to read the error log; without it the restart is still recorded and the reason says it could not be read.
The history is kept for at least 400 days by the history cleanup, or longer if HistoricRetentionDays is set higher (0 keeps everything).
Reading the page
The banner gives the month’s uptime percent against the target, colored green at or over the target, amber under it but at 99% or better, and red below that. Under it: the outages and downtime, how much downtime the target allows over the time monitored and how much of it is left (or how far over it is), the planned outages, and any part of the month that was not monitored.
The tiles repeat the figures: uptime, downtime, outages, longest outage, mean time between failures, last month’s uptime, and how long the month has been monitored.
The heat strip has one cell per day. Hover a cell for that day’s figures; click it to select that day’s outages in the grid. Gray days were not monitored.
The grid lists every outage, restart and monitor gap that touches the month: start, end, duration, type, planned, the failed checks and the SQL Server error number of the last one, the reason, and the machine running the service that recorded it.
Planned outages
Right-click an outage and choose Mark as planned… to record it as planned maintenance (a patch window or a change ticket, for example), with an optional note. The flag is saved in DBHealthHistory, so everyone who opens the report sees it. Mark as unplanned takes it back.
The uptime percent is 1 - (unplanned downtime / time monitored). Turn on Count planned on the toolbar to include planned downtime as well.
How the numbers are worked out
- Unreachable is from the monitor’s point of view. The service could not connect from the machine it runs on, so a network outage between that machine and the instance counts as downtime unless it is marked planned.
- Outages that overlap (a restart and the failed checks around it, or two machines running the service) are joined and counted once.
- Time nobody was watching (a Monitor Gap, or the time since the last check when the service is not running now) is left out of the month, so it does not count as up or down. A restart found by its start time still counts where it falls in such a gap.
- The month runs from the later of its first day and when monitoring began, to the earlier of its last day and now. Times are the server’s own clock.
- The percent is shown to three places and never rounded up, so a month at 99.8996% reads 99.899%, not 99.9%.
Permissions and versions
- The report reads
DBHealthHistory.dbo.AvailabilityOutageandAvailabilityMonitoron the instance. If DBHealthHistory is missing, cannot be opened by your login, or is older than version 1590, the page says which. - Marking an outage planned needs UPDATE on
DBHealthHistory.dbo.AvailabilityOutage. - The service needs to connect to master on the instance, and to write to DBHealthHistory. Reading
sqlserver_start_timeneeds VIEW SERVER STATE; without it the tempdb creation time is used, which is a few seconds after startup.