Files
nexus/sreweekly/markdown/82/06-dead-man-s-switch-with-aws-cloudwatch-freshness-alerting-for-backups-a-zh.md

38 lines
2.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 用 AWS CloudWatch 实现死人手柄:备份等的时效性告警
- **期号**: SRE Weekly Issue #82(2017-07-23)
- **作者**: —
- **链接**: https://ruempler.eu/2017/06/26/dead-mans-switch-backup-alerting-aws-cloudwatch/
## 简介
这里有一篇很不错的小教程:
> 我目前参与的一个团队最近遇到的挑战,是找到一种在 AWS CloudWatch 中实现的方法:当指标超过指定阈值时发出告警。当某个指标在指定时间间隔内没有发送到 CloudWatch 时发出告警。
## 正文
我目前参与的一个团队最近遇到的挑战,是找到一种在 AWS CloudWatch 中实现的方法:
1. 当指标超过指定阈值时发出告警。
2. 当某个指标在指定时间间隔内**没有**被发送到 CloudWatch 时发出告警。
虽然第一种基本上是 CloudWatch 的标准功能,但后者要棘手一些。在 Nagios/Icinga 的世界里,这叫做"freshness"(时效性/新鲜度)。你也可以把它称为周期任务/cronjob 的一种特殊的"[死人手柄](https://en.wikipedia.org/wiki/Dead_man%27s_switch)"(dead man's switch)。
举例来说,在我们的场景中,我们想监控并告警:备份作业是否每天运行一次。
下面是我们做的(CloudFormation 代码片段见下):
- 把检查周期(check period)设为该指标应该被发送的时间间隔。例如,如果指标应该每天发送,就设为 `86400`。这会让 CloudWatch 每天检查一次。
- 把评估周期(evaluation periods)设为 `1`:我们希望在没有任何数据写入或阈值被突破时立即收到告警。
- 现在是关键一步:我们必须把缺失数据(missing data)当作 `breaching`(已突破)处理,这样如果在评估周期内没有任何记录,告警就会被触发。
CloudFormation 语法示例:
| 1 2 3 4 5 6 7 | HealthCheckAlarm: Type: AWS::CloudWatch::Alarm Properties: Period: 86400 EvaluationPeriods: 1 TreatMissingData: breaching ... |
###
#### 喜欢这篇文章?
你可以[雇用我](https://cv.ruempler.eu/)或[通过 PayPal 向我捐款](https://www.paypal.me/s0enke)!