Alerts from your other tools
Send the alerts your existing tools raise to InfraSage. They land on the incident InfraSage already has open for that service, so one page shows both. They also give the day-30 report something to measure lead time against (GET /api/v1/reports/precision).
InfraSage records these alerts but never pages for them: your tool already routed its own page. A resolved message resolves the alert.
Every receiver takes an API key with ingestion or full scope. Send it as X-API-Key, as Authorization: Bearer <key>, or, where the sender can only put credentials in the URL, as the Basic-auth password.
| Tool | Endpoint | Setup |
|---|---|---|
| Prometheus Alertmanager | POST /api/v1/receiver/alertmanager | a webhook_configs receiver with send_resolved: true and authorization: {type: Bearer, credentials: <key>} |
| Grafana Alerting | POST /api/v1/receiver/grafana | a Webhook contact point. Set the Authorization header scheme to Bearer and the credentials to the key. |
| Datadog | POST /api/v1/receiver/datadog | a Webhooks integration with the payload below and a custom header {"X-API-Key": "<key>"}. Add @webhook-<name> to the monitors. |
| Amazon CloudWatch | POST /api/v1/receiver/cloudwatch | an HTTPS subscription on the alarm's SNS topic to https://x:<key>@<console host>/api/v1/receiver/cloudwatch. InfraSage confirms the subscription itself. |
Datadog payload
Use this JSON in the webhook's payload. Datadog fills in the $ variables:
{"id": "$ID", "title": "$EVENT_TITLE", "transition": "$ALERT_TRANSITION",
"alert_id": "$ALERT_ID", "cycle_key": "$ALERT_CYCLE_KEY",
"priority": "$ALERT_PRIORITY", "tags": "$TAGS", "date": "$DATE",
"link": "$LINK", "hostname": "$HOSTNAME", "body": "$EVENT_MSG"}
Recovered resolves the alert that the same $ALERT_CYCLE_KEY opened. Priorities P1 and P2 are recorded as critical, P4 and P5 as info, and everything else as a warning.
Which service an alert belongs to
InfraSage checks the alert's labels or tags in this order: service_id, service, service_name, app_kubernetes_io_name, app, k8s_app, job, deployment, statefulset, daemonset, container. It prefers a value that names a service it already sees in your telemetry.
- Datadog: tags are read as labels (
service:checkout).kube_deployment,kube_serviceandecs_servicealso name the service. - CloudWatch: the first of these alarm dimensions names the service:
ServiceName,FunctionName,DBInstanceIdentifier,DBClusterIdentifier,QueueName,LoadBalancer,TargetGroup,TableName,CacheClusterId,ClusterName,DomainName,AutoScalingGroupName,InstanceId,StreamName,TopicNameorApiName. A load balancer's name is taken fromapp/<name>/<id>. - CloudWatch states: an
INSUFFICIENT_DATAstate is not recorded. An alarm whose name or description says "critical" is recorded as critical.
With no usable label, the alert's rule, monitor or alarm name becomes its service.