Files
nexus/sreweekly/markdown/8/02-high-availability-at-massive-scale-building-google-s-data-infrastructu.md
2026-09-12 17:23:01 +08:00

1.9 KiB
Raw Blame History

High-Availability at Massive Scale: Building Google’s Data Infrastructure for Ads

简介

This is a really awesome paper. Two Googlers describe in detail the pitfalls of failover-based systems and explain how they design multi-homed active/active services. If Google has learned a lesson, we’d all do well to learn from it, too:

Our experience has been that bolting failover onto previously singly-homed systems has not worked well. These systems end up being complex to build, have high maintenance overhead to run, and expose complexity to users. Instead, we started building systems with multi-homing designed in from the start, and found that to be a much better solution. Multi-homed systems run with better availability and lower cost, and result in a much simpler system overall.

正文

Google’s Ads Data Infrastructure systems run the multi-

billion dollar ads business at Google. High availability and strong consistency are critical for these systems. While most distributed systems

handle machine-level failures well, handling datacenter-level failures is

less common. In our experience, handling datacenter-level failures is critical for running true high availability systems. Most of our systems (e.g.

Photon, F1, Mesa) now support multi-homing as a fundamental design property. Multi-homed systems run live in multiple datacenters all the time, adaptively moving load between datacenters, with the ability to handle outages of any scale completely transparently.

This paper focuses primarily on stream processing systems, and describes our general approaches for building high availability multi-homed systems, discusses common challenges and solutions, and shares what we

have learned in building and running these large-scale systems for over ten years.

×