Skip to content

Grafana Monitoring

Grafana is used for managed alerts that need to evaluate operational signals outside the CloudWatch alarm flow. These alerts should be treated as production health signals: they identify conditions that need investigation, then hand off to internal automation for notification and follow-up.

Long Running Crawler Alert

The Long Running Crawler alert detects crawler executions that run longer than expected for their historical behavior. It is managed as a Grafana alert rule and is intended to surface crawlers that may be stuck, degraded, blocked by external dependencies, or running unusually slowly.

The alert rules are configured in Grafana under Alerts & IRM > Alerting > Alert rules. They share the same Grafana folder, evaluation group, datasource, evaluation interval, threshold buckets, and alert-state behavior, but evaluate different schedule repeat units.

Field Value
Folder / namespace New Long Crawler
Evaluation group NewLongCrawlerEvaluationGroup
Rule type Grafana-managed alert rule
Datasource grafanacloud-grepsr-prom
Evaluation interval Every 1 minute
Rule Repeat unit Historical average window Summary annotation
Long Crawler Alert Rule - Tier Hourly - New hour Successful hourly runs over the previous 2 weeks, offset by 1 day The crawler has exceeded threshold value
Long Crawler Alert Rule - Tier Daily - New day Successful daily runs over the previous 2 weeks, offset by 1 day The crawlers have exceeded more than threshold value
Long Crawler Alert Rule - Tier Weekly - New week Successful weekly runs over the previous 2 weeks, offset by 1 day The crawlers have exceeded threshold value
Long Crawler Alert Rule - Tier Monthly - New month Average of two 4-week successful-run windows: one offset by 1 day and one offset by 4 weeks plus 1 day The crawlers have exceed their threshold value

Each rule evaluates the longcrawler Prometheus metric for its schedule repeat unit. It compares currently processing crawler runtime with the average runtime of successful runs for the same project_id, report_id, and schedule_id.

Signal Meaning
crawl_time Current runtime for longcrawler{status="PROCESSING", repeat_unit="<repeat-unit>"}.
average_crawl_time Historical average of successful crawler runs for the same repeat unit and schedule identity.
excess_crawl_time Difference between the current processing runtime and the historical successful-run average.

The same bucket thresholds apply to the hourly, daily, weekly, and monthly long-crawler alert rules. The alert fires when any threshold bucket evaluates to true:

Bucket Average crawl time Excess crawl time Example What it means
Bucket_1 2 hours or less More than 4 hours Average: 1 hour; excess: 5 hours; total runtime: 6 hours Crawl is usually fast, but the current extra crawl time is over 4 hours.
Bucket_2 More than 2 hours and up to 12 hours More than 8 hours Average: 6 hours; excess: 9 hours; total runtime: 15 hours Crawl usually takes a few hours, and the current extra crawl time is over 8 hours.
Bucket_3 More than 12 hours and up to 48 hours More than 12 hours Average: 24 hours; excess: 13 hours; total runtime: 37 hours Crawl usually takes half a day to 2 days, and the current extra crawl time is over 12 hours.
Bucket_4 More than 48 hours More than 24 hours Average: 60 hours; excess: 25 hours; total runtime: 85 hours Crawl usually takes more than 2 days, and the current extra crawl time is over 24 hours.

The main condition is:

Bucket_1 || Bucket_2 || Bucket_3 || Bucket_4

If the query returns no data or only null values, Grafana keeps the last alert state. If rule execution errors or times out, Grafana treats the rule as OK. The rule does not use a pending period or keep-firing window.

When the rule fires, Grafana sends a webhook to an internal Temporal workflow. The workflow owns the downstream automation path, including any internal notification or follow-up handling configured for the alert.

Use this alert as an investigation signal rather than a final diagnosis. Start by checking the affected crawler run, recent execution history, crawler logs, and any related infrastructure or third-party dependency issues before taking remediation action.