Exploring Service Observability
Summarize
Summary of Exploring Service Observability
Service Observability in ServiceNow helps operations teams efficiently triage and manage incidents in complex, distributed production environments by consolidating telemetry from external observability monitoring systems with related Configuration Management Database (CMDB) data. This combined information is displayed within the Service Operations Workspace (SOW) to provide a unified view of service health across applications, infrastructure, and network components.
Show less
It supports integration with leading observability vendors such as Amazon CloudWatch, AppDynamics, Datadog, Dynatrace, New Relic, Splunk, and others, as well as MetricBase for data sourcing. Metrics from these vendors are mapped to CMDB services using tags, enabling operators to see detailed health data and related configuration item (CI) information in one workflow for effective incident diagnosis and resolution.
Users and Roles
- System Admins: Configure users and teams, register services for monitoring, connect to observability vendors, map services to data, and view data in SOW.
- Service Observability Admins: Similar to system admins but additionally customize dashboard templates and manage configurations specific to Service Observability.
- Operators/Operations Managers: Use Service Observability within the SOW to triage incidents by viewing service health metrics, related alerts, incidents, and changes. They analyze detailed metrics to identify root causes and ownership for remediation.
Service Observability Workflow
- Admins: Select critical services to monitor, connect observability vendor instances, map services to observability data via tags, and customize metric display templates.
- Operators/Managers: Detect service issues through alerts or dashboards, review overall and detailed service health metrics, investigate related entities causing issues, and initiate remediation based on ownership.
Key Benefits
- Consolidated Data View: Integrates data from multiple monitoring and network health tools, cloud providers, and third-party sources to provide a comprehensive full-stack service health overview within the SOW.
- Improved Incident Resolution: Enables quicker root cause analysis by displaying combined metrics from related entities, helping reduce mean time to resolution (MTTR).
- Contextual Awareness: Shows related incidents, alerts, and system changes alongside health metrics, empowering operators with actionable insights.
- AI-Driven Analysis: Supports generative AI capabilities to analyze metric data and assist in understanding service health.
- Customizable Dashboards: Allows admins to tailor dashboard templates to suit organizational monitoring needs.
- Seamless Incident Management: Integrates Service Observability data directly into Incident Management workflows for streamlined operations.
Practical Takeaway for ServiceNow Customers
By leveraging Service Observability, ServiceNow customers can unify disparate monitoring data and CMDB information to gain a holistic view of service health within the familiar SOW interface. This enables faster identification of issues, efficient collaboration across teams, and improved incident response times. Configurable data mappings and dashboard templates provide flexibility to align observability with business priorities and existing monitoring tools. Additionally, AI-powered insights help operators better understand complex metrics, driving proactive and informed decision-making in service operations.
Service Observability helps operations teams triage and manage incidents in a complex and distributed production system. It combines external observability monitoring systems' telemetry with related data from the Configuration Management Database (CMDB) and displays both in a single workflow in the Service Operations Workspace (SOW).
Service Observability overview
Service Observability displays application, infrastructure, and network health metrics in the SOW related to a given service. Metrics can be ingested from an external observability vendor (application, network, and cloud monitors) and displayed alongside information for related configuration items in the CMDB.
Service Observability supports the following observability vendors:
- Amazon CloudWatch
- AppDynamics
- Cisco ThousandEyes synthetic tests
- Datadog
- DynatraceSaaS and on-premise (both Classic and Grail environments)
- LogicMonitor
- Microsoft Azure Monitor
- New Relic
- Prometheus on-premise
- SolarWinds on-premise
- Splunk Observability and logs from Splunk Enterprise
- Zabbix on-premise
Service Observability also supports MetricBase as a data source. For setup instructions, see Create and manage MetricBase data mappings.
- MySQL
- PostgreSQL (not supported with Splunk)
- RDS (Relational Database Service) for Amazon CloudWatch
After connecting an observability vendor to Service Observability, you map services in the CMDB to observability metrics using existing tags.
For example, say you use Dynatrace to monitor your checkout service, databases, and hosts, and that metrics from all these entities use the tag checkout-service to denote requests coming from that service.
By mapping the checkout service CI to the Dynatrace data tagged with checkout-service, Service Observability retrieves metrics for those databases and hosts and CIs related to the service, then displays them together. Operators can pinpoint issues on entities related to the service and narrow down the
mitigation process without having to leave the SOW.
Service Observability users
| User | Description |
|---|---|
| System admin |
System admins configure users and teams, register services to be monitored, connect Service Observability to observability vendors, and then map those services to that data. They can also view the data in the SOW |
| Service Observability admin |
Service Observability admins can configure users and teams, connect Service Observability to observability vendors, and then map services to that data. They can also view the data in the SOW. Admins can also customize dashboard templates used to display metrics and related information. |
| Operator/operations manager Note: These users must belong to an srm group type to see all data. |
Operators use Service Observability when triaging incidents in the SOW. They can view basic health metrics for a service, along with related incidents, alerts, and changes. They can get more detailed information by navigating to the Observability tab to view additional service metrics, along with metrics from related entities, such hosts, networks, or databases. |
Service Observability workflow
Admins configure Service Observability by creating a connection to an observability vendor and then mapping CI services to that data. Operators use Service Observability to determine if another related entity is causing issues surfaced by the service's performance.
As an admin, you:
- Determine the services to be monitored by Service Observability based on business criticality.
- Connect existing observability vendor instances to Service Observability.
- Map services to observability metric data using vendor-based tags attached to that data.
- Customize the templates used to display metric charts.
As an operator or manager, you:
- Spot an issue with a service while working in the SOW, for example, from an alert, the Service dashboard, or Express List, then navigate to the Service Details page.
- View overall health metrics for the service, along with related incidents, alerts, and changes. If one of the metrics seems unhealthy, navigate to the Observability tab.
- View more detailed service metrics, as well as information from related entities, to start root cause investigation. When finding that the issue is further down the system's stack, identify the ownership for that entity to start remediation.
Service Observability benefits
| Benefit | Feature | Users |
|---|---|---|
Consolidate data from existing monitoring tools, network health tools, cloud providers, ServiceNow agents, and third-party tools for a full-stack view of service health:
|
. | Admins |
| Increase efficiency and reduce mean time to resolution (MTTR). View combined metrics from entities associated with a service to begin to determine blast radius and ownership of an incident. | View service health metrics | Operators |
| See related changes to the system and alerts associated with a service in one place. | View overall service health. | Operators |
| Use generative AI to analyze metric data and find insights to help determine service health. | Operators | |
| See Service Observability data as part of Incident Management workflows | Operators | |
| Customize dashboard templates. | Customize Service Observability dashboard templates | Admins |