# Mitigate Connection Leaks in Production via Proxies - **期号**: SRE Weekly Issue #247(2020-12-06) - **作者**: Utsav Shah - **链接**: https://reliability.substack.com/p/mitigate-connection-leaks-in-production ## 简介 Hitting file descriptor limits is such an annoying kind of outage. Some good tips here, clearly coming from hard-won experience. ## 正文 Every [socket connection in Unix/Linux systems is represented by a file](https://unix.stackexchange.com/a/116616). Files opened by a process are represented by [file descriptors](https://stackoverflow.com/a/5256705/3399432) - integer numbers that are used in I/O syscalls like read/write. By default, the per process limit for file descriptors for a user on [Ubuntu is 1024](https://serverfault.com/a/122682) - which implies a process cannot have more than 1024 open files (or connections) simultaneously. This is a safe default so that any process cannot exhaust system wide limits, but it’s often set too low for servers that need to make a lot of connections, and the [guidance](https://istio.io/latest/docs/ops/common-problems/network-issues/#envoy-is-crashing-under-load) is to [bump this limit](https://www.mongodb.com/blog/post/tuning-mongodb--linux-to-allow-for-tens-of-thousands-connections) in many cases. However, it’s easy to deploy a bug that leaks connections, causes the system to run out of file descriptors, and prevent new connections. Failures are [often hard to debug](https://github.com/kubernetes/kubernetes/issues/62334), warrant their own [war story blog posts](https://hasura.io/blog/debugging-tcp-socket-leak-in-a-kubernetes-cluster-99171d3e654b/), and happen across [many](https://issues.apache.org/jira/browse/HTTPCORE-567) different [domains](https://jira.atlassian.com/browse/JRASERVER-39677). By the nature of the problem, failures tend to occur only under load (in production) and are hard to reproduce in unit/integration tests. To make things worse, a connection leak in a large client can overwhelm a small service and cause downtime. For example, we might have a large monolith application that runs on a lot of nodes, and a tiny auxiliary service that runs on a few nodes. A leak in the monolith where it creates too many auxiliary clients might not trigger any alarms for the monolith, but will take down the auxiliary service. It’s difficult to solve this problem in a truly general sense due to the various code paths that may have connection leaks. However, if we’re looking to solve only for internal service communication, we control internal RPC frameworks and deployments, and we can make it simple to debug these failures, catch them before full production rollout, and possibly eliminate them. We describe some approaches that can be used in tandem to provide mitigations to this problem. Each approach has its own tradeoffs. ## Singletons/Dependency Injection One technique is purely related to code structure - to design RPC client creation in a manner where constructing multiple clients in the same thread is unidiomatic. One way to do this is to expose RPC client creation via a [singleton pattern](https://en.wikipedia.org/wiki/Singleton_pattern). A simple implementation in Go might use [Once](https://golang.org/pkg/sync/#Once) - a struct that ensures a function runs at most once in a thread safe way - to provide an RPC client that’s shared across the process. It’s important to return a client that can automatically reconnect on connection failures/drops here. ``` // global vars var rpcClientOnce sync.Once var rpcClient rpc.Client func MyRpcClient() { rpcClientOnce.Do(constructRpcClient) // now rpcClient will be non nil return rpcClient } func constructRpcClient() { // construct client here rpcClient = ... } ``` Since singletons are often hard to mock out for tests, we can also use [dependency injection](https://en.wikipedia.org/wiki/Dependency_injection), where RPC client classes (or a corresponding [builder](https://en.wikipedia.org/wiki/Builder_pattern)) is injected and cannot be instantiated manually after startup. This approach doesn’t require any additional operational overhead, but might require extensive refactoring to work for an existing codebase, and doesn’t prevent all regressions. In practice, if the codebase convention is to use one of these approaches, then it’s likely that new callers will follow this pattern and avoid a leak. ## Client Count Metrics We can add a [gauge](https://prometheus.io/docs/concepts/metric_types/#gauge) that is incremented every time a new client is instantiated. These should be added to RPC client creation libraries so that new callers don’t have to explicitly opt into these metrics. This gives us instant visibility when we deploy a leak, since we can see the source of the leak via client metrics. Finally, if we have a canary analysis system like [Kayenta](https://netflixtechblog.com/automated-canary-analysis-at-netflix-with-kayenta-3260bc7acc69), we can detect a large increase in connection count during canary compared to the baseline, and can stop a full rollout of a leak to production. One downside of this approach is that we can’t set up simple alerts based on these metrics - often there’s no appropriate threshold to alert on, or we have to set up some convoluted alerts like client counts by node, and tweak it when we modify process count per node - which might happen when we modify node size. Another downside is that this might be flagged too late to prevent an outage. For example, a slow connection leak that happens on low traffic routes might not cause canary analysis to go off until after production rollout. ## File Descriptor Count Alerts We can use node exporters, like the [Prometheus node exporter](https://github.com/prometheus/node_exporter), to automatically monitor file descriptor count by node. This allows us to write a simple alert that can monitor used file descriptor percent by service/deployment. These can be combined with a rollout pipeline that validates alerts, for example, we can validate that