Skip to main content

Alerts from your other tools

Send the alerts your existing tools raise to InfraSage. They land on the incident InfraSage already has open for that service, so one page shows both. They also give the day-30 report something to measure lead time against (GET /api/v1/reports/precision).

InfraSage records these alerts but never pages for them: your tool already routed its own page. A resolved message resolves the alert.

Every receiver takes an API key with ingestion or full scope. Send it as X-API-Key, as Authorization: Bearer <key>, or, where the sender can only put credentials in the URL, as the Basic-auth password.

ToolEndpointSetup
Prometheus AlertmanagerPOST /api/v1/receiver/alertmanagera webhook_configs receiver with send_resolved: true and authorization: {type: Bearer, credentials: <key>}
Grafana AlertingPOST /api/v1/receiver/grafanaa Webhook contact point. Set the Authorization header scheme to Bearer and the credentials to the key.
DatadogPOST /api/v1/receiver/datadoga Webhooks integration with the payload below and a custom header {"X-API-Key": "<key>"}. Add @webhook-<name> to the monitors.
Amazon CloudWatchPOST /api/v1/receiver/cloudwatchan HTTPS subscription on the alarm's SNS topic to https://x:<key>@<console host>/api/v1/receiver/cloudwatch. InfraSage confirms the subscription itself.

Datadog payload​

Use this JSON in the webhook's payload. Datadog fills in the $ variables:

{"id": "$ID", "title": "$EVENT_TITLE", "transition": "$ALERT_TRANSITION",
"alert_id": "$ALERT_ID", "cycle_key": "$ALERT_CYCLE_KEY",
"priority": "$ALERT_PRIORITY", "tags": "$TAGS", "date": "$DATE",
"link": "$LINK", "hostname": "$HOSTNAME", "body": "$EVENT_MSG"}

Recovered resolves the alert that the same $ALERT_CYCLE_KEY opened. Priorities P1 and P2 are recorded as critical, P4 and P5 as info, and everything else as a warning.

Which service an alert belongs to​

InfraSage checks the alert's labels or tags in this order: service_id, service, service_name, app_kubernetes_io_name, app, k8s_app, job, deployment, statefulset, daemonset, container. It prefers a value that names a service it already sees in your telemetry.

  • Datadog: tags are read as labels (service:checkout). kube_deployment, kube_service and ecs_service also name the service.
  • CloudWatch: the first of these alarm dimensions names the service: ServiceName, FunctionName, DBInstanceIdentifier, DBClusterIdentifier, QueueName, LoadBalancer, TargetGroup, TableName, CacheClusterId, ClusterName, DomainName, AutoScalingGroupName, InstanceId, StreamName, TopicName or ApiName. A load balancer's name is taken from app/<name>/<id>.
  • CloudWatch states: an INSUFFICIENT_DATA state is not recorded. An alarm whose name or description says "critical" is recorded as critical.

With no usable label, the alert's rule, monitor or alarm name becomes its service.