Files
nexus/sreweekly/markdown/235/06-production-testing-with-dark-canaries.md
2026-09-12 17:23:01 +08:00

6.2 KiB
Raw Blame History

Production testing with dark canaries

简介

I like this idea: it’s like a normal canary, except that you only send it a copy of traffic and discard the result, so as to avoid impacting users.

正文

Production testing with dark canaries

Dark cluster architecture diagram  

While there are many ways to dispatch traffic to the Service A dark canary cluster, LinkedIn uses a client-side library that is able to discover the dark cluster(s) that Service A should fork traffic to.

In this case, we store a self-updating mapping from a source cluster (“ServiceACluster”) to a set of corresponding Dark Clusters ([“DarkServiceACluster”]) in Apache ZooKeeper. In addition, we can store other metadata such as how to multiply traffic to the Dark Cluster and other configurations we need for forking traffic. Now, whenever Service A receives a request, an inbound request filter will read this ZooKeeper data, detect that it needs to send dark traffic to DarkServiceACluster, and fork the request appropriately.

Sample D2 Configuration, simplified for brevity. Configs can be structured in XML/JSON/etc.

Using Apache ZooKeeper in this case is not strictly necessary, but it does help because it allows us to know the exact count of both regular clusters and dark clusters. This enables us to scale up and down both regular and dark clusters independently with the dispatching instances adjusted automatically. One strategy for sending dark traffic is to make the queries per second (QPS) between a dark cluster instance and a regular cluster instance identical (or proportional according to some multiplier), so they can be compared by validation tooling. If we add instances to DarkServiceACluster, then we would expect the total outbound dark requests from ServiceACluster to increase, but the QPS on each dark cluster instance to stay the same. Conversely, if we added instances to ServiceACluster, then we’d expect the per-instance QPS to ServiceACluster to drop (as the total QPS remains the same and becomes spread over more instances). As a result, we’d want the per-instance DarkServiceACluster QPS to drop a corresponding amount.

To follow this strategy, let’s illustrate how a Service A instance can calculate the ratio of requests it should send.

First, a definition:

Multiplier: number used to proportionally scale the traffic a dark cluster instance receives compared to a regular cluster instance. Setting the multiplier as 1.0 means the dark cluster per-instance QPS should be the same as the regular per-instance QPS, while 1.2 means the dark cluster per-instance QPS should receive 20% more QPS than the regular per-instance QPS.

All QPS fields below are per-instance unless explicitly specified otherwise.

And

This means that if the inbound QPS is 100, the multiplier is 1, and there are 2 DarkServiceACluster instances and 100 regular ServiceACluster instances, then the outbound DarkRequest QPS per regular ServiceA instance is:

With this formula, the traffic on the dark cluster instances will stay proportional to your regular service instances, and you don’t have to worry about overloading them just by scaling the respective clusters up or down.

Ongoing work

Many teams at LinkedIn are working to realize the full potential of dark canary clusters, including approaches to automate the onboarding and spinning up of dark canary clusters, orchestrate the testing of AI models on dark clusters, and build better validation metrics and test suites that weren’t practical earlier to run in production.

Conclusion

Dark canary clusters are something we believe many companies can benefit from, regardless of whether the cluster management is done using D2. Not only is both onboarding and maintenance of dark clusters easy with this approach, but the strategy to compare regular instances against dark cluster instances can also be used in automated validation, helping to speed up the development process by removing manual validation steps. By removing error-prone manual validation steps with automated validation, you remove the expertise required to correctly validate code (especially against graphs) and build the confidence needed to speed up development velocity. We hope that this can help others establish best practices for validation and deployment within their companies.

Acknowledgements

Many thanks to Sean Sheng and Chris Zhang from the Service Infrastructure team for their development support and brainstorming, to Erik Krogen and Sumit Rangwala for reviewing this blog post, and to the brave first adopters: Srividya Krishnamurthy, Ramon Garcia, Peter Chng, and Yafei Wang. Thanks to the Pro-ML Leadership team for their continued support and investment in this cross discipline area: Eing Ong, Josh Hartman, Joel Young, and Josh Walker.

Topics:

            [Developer Experience/Productivity](https://www.linkedin.com/blog/engineering/developer-experience-productivity)
          
          
            [Artificial intelligence](https://www.linkedin.com/blog/engineering/artificial-intelligence)
          
          
            [A/B Testing/Experimentation](https://www.linkedin.com/blog/engineering/ab-testing-experimentation)
          
          
            [Automation](https://www.linkedin.com/blog/engineering/automation)
          
          
            [Infrastructure](https://www.linkedin.com/blog/engineering/infrastructure)

Related articles