Grafana Monitoring¶
Grafana is used for managed alerts that need to evaluate operational signals outside the CloudWatch alarm flow. These alerts should be treated as production health signals: they identify conditions that need investigation, then hand off to internal automation for notification and follow-up.
Long Running Crawler Alert¶
The Long Running Crawler alert detects crawler executions that run longer than expected for their historical behavior. It is managed as a Grafana alert rule and is intended to surface crawlers that may be stuck, degraded, blocked by external dependencies, or running unusually slowly.
The alert rules are configured in Grafana under Alerts & IRM > Alerting > Alert rules. They share the same Grafana folder, evaluation group, datasource, evaluation interval, threshold buckets, and alert-state behavior, but evaluate different schedule repeat units.
| Field | Value |
|---|---|
| Folder / namespace | New Long Crawler |
| Evaluation group | NewLongCrawlerEvaluationGroup |
| Rule type | Grafana-managed alert rule |
| Datasource | grafanacloud-grepsr-prom |
| Evaluation interval | Every 1 minute |
| Rule | Repeat unit | Historical average window | Summary annotation |
|---|---|---|---|
Long Crawler Alert Rule - Tier Hourly - New |
hour |
Successful hourly runs over the previous 2 weeks, offset by 1 day | The crawler has exceeded threshold value |
Long Crawler Alert Rule - Tier Daily - New |
day |
Successful daily runs over the previous 2 weeks, offset by 1 day | The crawlers have exceeded more than threshold value |
Long Crawler Alert Rule - Tier Weekly - New |
week |
Successful weekly runs over the previous 2 weeks, offset by 1 day | The crawlers have exceeded threshold value |
Long Crawler Alert Rule - Tier Monthly - New |
month |
Average of two 4-week successful-run windows: one offset by 1 day and one offset by 4 weeks plus 1 day | The crawlers have exceed their threshold value |
Each rule evaluates the longcrawler Prometheus metric for its schedule repeat unit. It compares
currently processing crawler runtime with the average runtime of successful runs for the same
project_id, report_id, and schedule_id.
| Signal | Meaning |
|---|---|
crawl_time |
Current runtime for longcrawler{status="PROCESSING", repeat_unit="<repeat-unit>"}. |
average_crawl_time |
Historical average of successful crawler runs for the same repeat unit and schedule identity. |
excess_crawl_time |
Difference between the current processing runtime and the historical successful-run average. |
The same bucket thresholds apply to the hourly, daily, weekly, and monthly long-crawler alert rules. The alert fires when any threshold bucket evaluates to true:
| Bucket | Average crawl time | Excess crawl time | Example | What it means |
|---|---|---|---|---|
Bucket_1 |
2 hours or less | More than 4 hours | Average: 1 hour; excess: 5 hours; total runtime: 6 hours | Crawl is usually fast, but the current extra crawl time is over 4 hours. |
Bucket_2 |
More than 2 hours and up to 12 hours | More than 8 hours | Average: 6 hours; excess: 9 hours; total runtime: 15 hours | Crawl usually takes a few hours, and the current extra crawl time is over 8 hours. |
Bucket_3 |
More than 12 hours and up to 48 hours | More than 12 hours | Average: 24 hours; excess: 13 hours; total runtime: 37 hours | Crawl usually takes half a day to 2 days, and the current extra crawl time is over 12 hours. |
Bucket_4 |
More than 48 hours | More than 24 hours | Average: 60 hours; excess: 25 hours; total runtime: 85 hours | Crawl usually takes more than 2 days, and the current extra crawl time is over 24 hours. |
The main condition is:
Bucket_1 || Bucket_2 || Bucket_3 || Bucket_4
If the query returns no data or only null values, Grafana keeps the last alert state. If rule
execution errors or times out, Grafana treats the rule as OK. The rule does not use a pending
period or keep-firing window.
When the rule fires, Grafana sends a webhook to an internal Temporal workflow. The workflow owns the downstream automation path, including any internal notification or follow-up handling configured for the alert.
Use this alert as an investigation signal rather than a final diagnosis. Start by checking the affected crawler run, recent execution history, crawler logs, and any related infrastructure or third-party dependency issues before taking remediation action.