APM Metric Detection¶
Document Scope
This document is the second step in the detection rule configuration workflow. After completing the configuration, return to the main document to continue with step three: Event Notification.
Data scope: Trace (T), used to monitor key APM metric data in the workspace. The system counts the number of traces that meet the conditions within the specified time period. When the count exceeds the custom threshold, an anomaly event is triggered.
Detection Configuration¶
Detection Frequency¶
Set the time interval at which detection is executed.
-
Preset options: 1 minute, 5 minutes, 10 minutes, 15 minutes, 30 minutes, 1 hour
-
Crontab mode: Click “Switch to Crontab mode” to configure a custom schedule. Scheduled task execution can be configured with periods based on seconds, minutes, hours, days, months, and weeks.
Detection Range¶
Set the time range of data queried for each detection. (❗️The detection range must be greater than or equal to the detection frequency and must match the actual data reporting interval to avoid missed detections or false alarms.)
| Detection Frequency | Detection Range (Dropdown Options) |
|---|---|
| 30s | 1m/5m/15m/30m/1h/3h |
| 1m | 1m/5m/15m/30m/1h/3h |
| 5m | 5m/15m/30m/1h/3h |
| 15m | 15m/30m/1h/3h/6h |
| 30m | 30m/1h/3h/6h |
| 1h | 1h/3h/6h/12h/24h |
| 6h | 6h/12h/24h |
| 12h | 12h/24h |
| 24h | 24h |
- Custom format: Enter a custom detection range, for example:
20m(last 20 minutes),2h(last 2 hours),1d(last 1 day).
Detection Metrics¶
Set the metrics used for detection. Two detection modes are supported:
-
Service metrics
-
Trace statistics
Note
Avoid selecting high-cardinality fields as detection dimensions. If configured improperly with overly loose trigger conditions, frequent alerts may occur. The current query returns a maximum of 100,000 records.
Service Metrics¶
Monitor the Application Performance Monitoring (APM) services in the current workspace.
| Configuration Item | Description |
|---|---|
| Service | Select the APM services in the current workspace. You can select "All" or a specific service. |
| Metrics | The specific detection metrics, including: request count, error request count, request error rate, average requests per second, average response time, P50 response time, P75 response time, P90 response time, P99 response time. |
| Filter Conditions | Filter the detection data based on metric tags to limit the detection scope. You can add one or more tag filters, and both fuzzy match and fuzzy no-match conditions are supported. |
| Detection Dimensions | Any string-type (keyword) field in the data can be selected as a detection dimension. Up to three fields are supported. A combination of multiple detection dimension fields identifies a specific detection target (for example, {service: svc1, host: host1}). |
| Additional Information | Select fields to be displayed as supplementary information to enrich the event content. |
Trace Statistics¶
Count the number of traces (Spans) that meet the conditions within the specified time period. When the count exceeds the custom threshold, an anomaly event is triggered. This mode can be used to notify about anomalies and errors in service traces.
| Configuration Item | Description |
|---|---|
| Source | Select the source (service) of the trace data to be counted. Keyword filtering is supported. |
| Filter Conditions | Filter traces (span) by tags to limit the data range of the detection. You can add one or more tag filter conditions. |
| Aggregation Algorithm | "*" is selected by default, and the corresponding aggregation function is count (counts the number of spans). If another field is selected, the aggregation function automatically changes to count distinct (counts the number of data points in which the keyword appears, that is, deduplicated counting). |
| Detection Dimensions | Any string-type (keyword) field in the data can be selected as a detection dimension. Up to three fields are supported. A combination of multiple detection dimension fields identifies a specific detection target. |
Trigger Conditions¶
Configure the trigger conditions and severity levels. When the query returns multiple values, an event is generated if any value meets the trigger condition.
Four threshold levels are supported: Fatal, Critical, Important, and Warning, as well as a Normal recovery condition.
| Level | Configuration | Description |
|---|---|---|
| Fatal | When Result >= [value] |
Highest severity alert, requires immediate handling |
| Critical | When Result >= [value] |
High severity alert, requires priority handling |
| Important | When Result >= [value] |
Medium severity alert, requires attention |
| Warning | When Result >= [value] |
Low severity alert, requires monitoring |
| Normal | No events generated in [N] consecutive detections |
If the detected metric triggers a "Fatal", "Critical", "Important", or "Warning" anomaly event, and the subsequent N consecutive detections are all normal, a "Normal" event is generated. This is used to determine whether the anomaly has recovered. It is recommended to configure this. |
For more details, see Event Level Description.
Advanced Options¶
Consecutive Trigger Evaluation¶
When enabled, an event is generated only when the trigger condition is continuously met, preventing false alarms caused by transient fluctuations. (❗️The maximum configurable limit is 10.)
Bulk Alert Protection¶
Enabled by default.
When the number of alerts generated by a single detection exceeds the preset threshold, the system automatically switches to a status-based aggregation strategy: instead of processing each alert target individually, it generates and pushes a small number of summary alerts based on event status.
This ensures timely notifications while significantly reducing alert noise and avoiding the risk of timeouts caused by processing too many alerts.
When this toggle is enabled, the event details generated after the monitor detects an anomaly will not display history records or related events.
Recovery Conditions¶
Configure the recovery conditions and severity levels. When the query returns multiple values, a recovery event is generated if any value meets the trigger condition.
Set independent recovery thresholds for different levels to enable tiered recovery. For example, a Critical alert recovers only when the value drops below 70, while an Important alert can recover when the value drops below 80.
Default Recovery Logic
When the level-based recovery condition configuration is not enabled, the default behavior is to automatically recover when detection results no longer meet the trigger conditions.
Data Gap¶
Strategy for handling empty query results within the detection range:
| Option | Description |
|---|---|
| No event triggered (default) | No alert is generated when there is no data. Suitable for scenarios where missing data is acceptable. |
| Treat query result as 0 | Empty data is treated as a value of 0 for threshold evaluation. |
| Trigger a data gap event | No data is treated as an anomaly and a data gap event is triggered. |
| Trigger a Fatal event | A Fatal-level event is triggered when there is no data. |
| Trigger a Critical event | A Critical-level event is triggered when there is no data. |
| Trigger an Important event | An Important-level event is triggered when there is no data. |
| Trigger a Warning event | A Warning-level event is triggered when there is no data. |
| Trigger a recovery event | A recovery event is triggered when there is no data. |
When trigger conditions, data gap handling, and information generation are configured together, the triggering priority is evaluated as follows: Data gap > Trigger condition > Information event generation.
That is: first determine whether a data gap exists, then determine whether the threshold is triggered, and finally determine whether an information event is generated.
Information Generation¶
When this option is enabled, you must configure the information generation condition. The system writes an "Info" event only when the detection result does not trigger any of the "Fatal", "Critical", "Important", or "Warning" thresholds and the information generation condition is met.
This is suitable for scenarios where normal status changes or low-priority information need to be recorded.
Next Steps¶
After completing the detection configuration above, continue with the following configuration:
- Event Notification: Define the event title, content, notification members, data gap handling, and linked incidents;
- Alert Configuration: Select the alert policy, and configure notification targets and mute periods;
- Link: Link a dashboard for quick navigation to view data;
- Permissions: Set operation permissions to control who can edit or delete this monitor.