diff --git a/Pasted image 20260912131326.png b/Pasted image 20260912131326.png deleted file mode 100644 index 7659bee9..00000000 Binary files a/Pasted image 20260912131326.png and /dev/null differ diff --git a/sreweekly/markdown/532/02-unethical-ways-to-manage-technical-debt.md b/sreweekly/markdown/532/02-unethical-ways-to-manage-technical-debt.md deleted file mode 100644 index a55be61a..00000000 --- a/sreweekly/markdown/532/02-unethical-ways-to-manage-technical-debt.md +++ /dev/null @@ -1,13 +0,0 @@ -# Unethical Ways to Manage Technical Debt - -- **期号**: SRE Weekly Issue #532(2026-08-31) -- **作者**: Thomas A. Limoncelli — ACM Queue -- **链接**: https://queue.acm.org/detail.cfm?ref=rss&id=3830399 - -## 简介 - -Ethics are relative, right? This article is full of genuinely useful tips and framings. - -## 正文 - -> ⚠️ 抓取失败:HTTP 403 diff --git a/sreweekly/markdown/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.md b/sreweekly/markdown/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.md index 316220f0..c584fe2a 100644 --- a/sreweekly/markdown/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.md +++ b/sreweekly/markdown/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.md @@ -10,10 +10,6 @@ Whoa. It’s been quite a few years since my last run-in with an overfull connt ## 正文 -[Knowledge Hub](https://www.adyen.com/knowledge-hub) - -Article - # Inside Cilium CNI: solving mysterious Kubernetes pod setup timeouts A deep dive into how Adyen's Data Platform Engineering team investigated and resolved linear-time scaling bottlenecks in Cilium CNI connection tracking garbage collection to fix mysterious Kubernetes pod setup timeouts on high-resource nodes. diff --git a/sreweekly/markdown/532/05-the-rise-of-cognitive-observability.md b/sreweekly/markdown/532/05-the-rise-of-cognitive-observability.md index 1873f275..9cac7225 100644 --- a/sreweekly/markdown/532/05-the-rise-of-cognitive-observability.md +++ b/sreweekly/markdown/532/05-the-rise-of-cognitive-observability.md @@ -10,4 +10,330 @@ ## 正文 -> ⚠️ 抓取失败:HTTP 403 +## How LLMs are exposing the limits of traditional software observability + +![Observability for LLMs](https://miro.medium.com/v2/resize:fit:3396/format:webp/1*TM1szQyDptkYEXCrvdUFHQ.jpeg) + +Observability for LLMs — evolving traditional deterministic observability for non-deterministic behavior. ( This image was created using an AI image creation program.) + +## Why Traditional Observability Fails for LLM Systems + +Imagine a customer support chatbot confidently explaining a company’s refund policy to thousands of users. The dashboards look perfect: latency is stable, traces complete successfully, and every service appears healthy. + +*But there is a problem. The chatbot is wrong.* + +Not catastrophically wrong. Not obviously broken. The responses sound polished, professional, and believable. Most users would trust them immediately. The problem is subtler: a recent prompt update slightly changed how the model interprets edge cases, causing it to invent policy details in borderline scenarios. + +From the perspective of traditional observability, nothing failed. And yet, the response is wrong. + +**Large Language Models (LLMs)** are introducing a different kind of failure altogether. Not infrastructure failure, **reasoning failure.** That distinction matters. + +> The operational assumptions behind modern observability were built for deterministic systems. **LLMs are probabilistic systems** that can behave unpredictably while appearing completely healthy from the outside. + +The industry is now discovering something uncomfortable: we can monitor servers extremely well, but we still struggle to monitor machine reasoning. + +## Modern Observability: The Comfort of Determinism + +For more than a decade, the industry has become exceptionally good at observing systems. As architectures evolved from monoliths to distributed microservices, observability emerged through metrics, logs, and traces to help engineers reconstruct system behavior externally. + +At first glance, LLM systems seem no different. They still run on infrastructure, expose APIs, emit telemetry, and sit behind load balancers. So applying traditional observability feels natural. But traditional systems fail in predictable ways: databases crash, requests timeout, and services degrade under load. LLMs do not. + +> In traditional systems, cause and effect are linked via deterministic paths. +> **LLMs do not operate that way.** They don’t execute logic. They generate it. + +### An Example of When the System Works but the Answer Doesn’t + +Consider a RAG system inside a financial institution. An analyst requests guidance on a compliance procedure. The retrieval layer works correctly, embeddings match relevant documents, and the model responds within latency thresholds. + +But one retrieved document is outdated. + +The model confidently generates a coherent, polished, and factually incorrect answer based on stale information. + +No errors appear. No alerts trigger. Every infrastructure signal remains healthy. + +That is the disconnect. + +Traditional observability asks: + +- Is the system available? +- Are requests failing? +- Is latency acceptable? + +LLM systems introduce harder questions: + +- Is the answer correct? +- Is the reasoning grounded? +- Has behavior drifted over time? + +> Traditional observability monitors execution. LLM observability must monitor behavior. + +That is a fundamentally different engineering challenge. And the industry is only beginning to appreciate how significant that shift really is. + +![Traditional observability vs LLM observability](https://miro.medium.com/v2/resize:fit:2000/format:webp/1*1GFW2XpzkSyNJeovDARr0Q.jpeg) + +Traditional observability vs LLM observability. ( This image was created using an AI image creation program.) + +## Non Deterministic Observability + +In traditional software, behavior is defined by code. If something changes, engineers can usually trace it back to a deployment or configuration update. + +LLMs blur that boundary. In AI systems, the intelligence layer itself becomes part of production infrastructure. + +Take prompts, for example. A single sentence added to a system prompt can shift tone, alter reasoning patterns, increase hallucinations, or change how the model responds under uncertainty. + +At first, prompts feel like configuration. In practice, they behave far more like production code. + +The challenge is that prompt changes rarely produce obvious failures. A tweak intended to make responses “more conversational” may quietly reduce factual reliability. + +This creates a new kind of drift. + +> Not infrastructure drift. **Semantic drift.** + +## Observing LLMs: The Illusion of Visibility + +Many teams respond to this by instrumenting LLM calls — logging prompts, responses, token usage. + +On the surface, this feels like observability. But the real issue isn’t visibility. It’s interpretation. + +Imagine an agentic system responsible for automating infrastructure remediation. During an outage, the agent enters a loop. It repeatedly calls the wrong API endpoint, misinterpreting the error signal and retrying indefinitely. + +A standard trace will show you everything: + +- The sequence of API calls +- The retry logic +- The timing and latency + +But it won’t tell you *why* the agent made that decision. + +- Why did it select that tool? +- Why did it misinterpret the error? +- Why did it fail to break the loop + +> Traditional observability measures operational correctness. +> LLM observability increasingly attempts to measure cognitive correctness. + +Cognitive correctness is far more difficult to quantify. + +![Challenges in implmenting observability of LLMs](https://miro.medium.com/v2/resize:fit:3072/format:webp/1*uDIsatygUmtKNoQ4vymDsQ.jpeg) + +Challenges in implmenting observability of LLMs. ( This image was created using an AI image creation program.) + +### The Fetish of “Correctness Defines Reliability” + +Traditional systems treat correctness as binary: a request succeeds or fails, a container is healthy or crashed. + +LLMs operate in a far stranger space. A response can be partially correct, subtly misleading, or confidently wrong while sounding entirely believable. + +That makes evaluation subjective. A healthcare summary may appear accurate to one reviewer, incomplete to another, and misleading to a third. + +And that is the problem. + +Many LLM tasks lack a universal ground truth, making simple pass/fail observability impossible. + +Instead, teams must build approximation layers: + +- **Semantic similarity scoring:** To compare generated responses against reference answers using embeddings. For example, *“Reset your password from the security settings page”* and *“Password resets are available under account security settings”* differ in wording but remain semantically aligned. This helps evaluate relevance without requiring exact matches. +- **Groundedness checks:** To verify whether a response is actually supported by retrieved context. If a RAG system retrieves documentation stating employees can carry over *10* vacation days, but the model responds with *15*, the answer may sound fluent while remaining unsupported by evidence. +- **LLM-as-a-judge pipelines:** To use one model to evaluate another for factual consistency, reasoning quality, or policy compliance. This approach is increasingly necessary because manually reviewing millions of AI interactions is operationally impossible. +- **Heuristic evaluation frameworks:** To apply deterministic guardrails such as schema validation, citation enforcement, or rule-based rejection of unsupported financial calculations and malformed outputs. + +But each of these introduces its own uncertainty. + +> You are no longer measuring system state. +> You are estimating semantic quality. + +### Troubleshooting Without Reproducibility + +Troubleshooting LLM systems is fundamentally different from debugging traditional software. + +Conventional systems are reproducible: given the same inputs, they usually produce the same outputs. LLMs are not strictly deterministic. Responses can vary based on prompt structure, retrieval ordering, hidden model updates, or even subtle context changes. + +That creates a strange operational problem. An engineer investigating a production incident may never be able to reproduce the exact output that caused it. + +***The failure becomes ephemeral.*** + +It existed. It caused impact. But it cannot be recreated precisely. + +### Observing LLMs is a Data Problem + +There is also a second-order challenge many teams underestimate: scale. + +LLM observability generates enormous amounts of data. Every interaction may include prompts, completions, embeddings, retrieval metadata, evaluation scores, and intermediate reasoning traces. + +At production scale, querying and processing this telemetry quickly becomes a serious infrastructure problem. + +But unlike traditional logs, this data carries semantic meaning. It often contains customer conversations, financial records, legal documents, or proprietary knowledge. + +> LLM observability is not just a telemetry challenge. +> It is also a privacy, governance, and compliance challenge. + +You are not just observing systems. You are observing intelligence interacting with real-world data. + +## Emerging Tools, Incomplete Answers + +The industry is beginning to respond. + +## Get Barnadeep Bhowmik’s stories in your inbox + +Join Medium for free to get updates from this writer. + +New tooling is emerging to make sense of this complexity: tracing frameworks for LLM workflows, evaluation pipelines, integrated observability platforms. + +Some tools focus on tracing execution paths across prompts, tools, and agents — allowing engineers to inspect intermediate reasoning steps. + +Others focus on evaluation — measuring hallucination rates, grounding quality, and output consistency over time. + +There is also a growing effort to integrate LLM observability into existing telemetry ecosystems, rather than treating it as a separate stack. + +Here are some that you can use today: + +### LangSmith + +LangSmith is one of the most widely adopted platforms in the space, particularly among teams building applications with LangChain. + +Its core strength is tracing. + +At first glance, tracing LLM systems may sound similar to tracing microservices. But the value becomes more obvious in agentic workflows where a single request may involve: + +- multiple prompts, +- retrieval steps, +- tool invocations, +- memory updates, +- and reasoning chains. + +Without visibility into intermediate behavior, debugging these systems becomes extremely difficult. + +A simple LangSmith trace looks like this: + +```cs +from langsmith import traceable + +@traceable +def generate_answer(question): + return llm.invoke(question) +``` + +Once instrumented, developers can inspect: + +- prompt flows, +- execution timing, +- intermediate outputs, +- and tool interactions. + +LangSmith becomes particularly useful when engineers need to answer questions like: ***Why did the agent make this decision?*** + +### Langfuse + +Langfuse approaches observability from a more operational perspective. + +One reason it has gained attention is because many enterprises remain uncomfortable sending sensitive AI telemetry into third-party SaaS platforms. + +Langfuse offers open-source and self-hosted capabilities, making it attractive for organizations with strict compliance requirements. + +A simple integration looks like this: + +```cs +from langfuse import Langfuse + +langfuse = Langfuse() +trace = langfuse.trace( + name="customer-support-chatbot" +) +generation = trace.generation( + name="llm-response", + model="gpt-4" +) +``` + +At first, this appears similar to traditional application telemetry. + +But the real value comes from observing behavioral trends over time: + +- Which prompts produce poor outcomes? +- Which workflows generate excessive token costs? +- Which retrieval pipelines correlate with hallucinations? + +Langfuse attempts to operationalize these questions through tracing and evaluation layers. + +### OpenLLMetry + +OpenLLMetry takes a particularly interesting architectural approach. + +Rather than building an entirely separate observability ecosystem for AI, it extends OpenTelemetry concepts into LLM systems. + +That may sound subtle, but it is strategically important. + +Most enterprises already have mature observability pipelines: + +- dashboards, +- collectors, +- tracing backends, +- alerting systems, +- and telemetry standards. + +Completely replacing those systems for AI workloads would be operationally unrealistic. + +OpenLLMetry instead tries to bridge the gap. + +Example instrumentation: + +```cs +from traceloop.sdk import Traceloop + +Traceloop.init() +response = client.chat.completions.create( + model="gpt-4", + messages=[ + {"role": "user", "content": "Hello"} + ] +) +``` + +The framework can automatically capture: + +- prompts, +- responses, +- token usage, +- latency, +- and execution metadata. + +That interoperability matters because many organizations do not want isolated AI observability silos. They want AI telemetry integrated into existing operational ecosystems. + +And over time, that may become one of the most important architectural directions in the space. + +### Why Adoption Is Slower Than Expected + +At first glance, AI observability feels inevitable. If organizations are deploying LLMs into production, they need visibility into hallucinations, reasoning failures, and agent behavior. + +But adoption has been slower than expected. + +Traditional observability matured over decades around stable operational patterns. LLM observability remains immature, with rapidly evolving frameworks, evaluation methods, and unclear definitions of reliable AI behavior. + +The deeper problem is that AI failures are difficult to measure. Infrastructure failures are objective. Reasoning failures are not. + +At enterprise scale, privacy, compliance, and operational complexity make the challenge even harder. Like most reliability disciplines, AI observability is evolving reactively through production failures. + +**Final Thought** + +We spent years learning how to observe machines. Now we are learning how to *observe machine thinking*. + +And the tools, both conceptual and technical, are still catching up. + +The dashboards may still be green. But the real system, the one that matters, is no longer just infrastructure. + +It is behavior. And that is much harder to monitor. + +> ***“The most dangerous AI failures are not the ones that crash. They are the ones that sound convincing.”*** + +## Resources and Citations + +### Platforms + +- [LangSmith](https://www.langchain.com/langsmith/observability) +- [Langfuse](https://langfuse.com/) +- [OpenLLMetry GitHub](https://github.com/traceloop/openllmetry?) + +### Research Papers + +- [HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models](https://arxiv.org/abs/2305.11747) +- [A Survey on Hallucination in Large Language Models](https://arxiv.org/abs/2311.05232) diff --git a/sreweekly/markdown/533/07-a-tale-of-two-flink-autoscalers.md b/sreweekly/markdown/533/07-a-tale-of-two-flink-autoscalers.md index 139aa7f0..1aa1f25b 100644 --- a/sreweekly/markdown/533/07-a-tale-of-two-flink-autoscalers.md +++ b/sreweekly/markdown/533/07-a-tale-of-two-flink-autoscalers.md @@ -10,4 +10,4 @@ Switching from their custom-written autoscaler to the new off-the-shelf option m ## 正文 -> ⚠️ 抓取失败:HTTP 403 +[[A Tale of Two Flink Autoscalers]] diff --git a/sreweekly/markdown/533/A Tale of Two Flink Autoscalers.md b/sreweekly/markdown/533/A Tale of Two Flink Autoscalers.md new file mode 100644 index 00000000..81c43919 --- /dev/null +++ b/sreweekly/markdown/533/A Tale of Two Flink Autoscalers.md @@ -0,0 +1,98 @@ +--- +title: "A Tale of Two Flink Autoscalers" +source: "https://netflixtechblog.com/a-tale-of-two-flink-autoscalers-e9f6a1b1492b" +author: + - "[[Netflix Technology Blog]]" +published: 2026-08-22 +created: 2026-09-14 +description: "How Netflix is moving from our in-house Flink autoscaler to an open-source solution, and what we learned about metrics, cost, stateful scaling, and the hidden price of maintaining infrastructure." +tags: + - "clippings" +--- +[Samuel Yeboah](https://www.linkedin.com/in/samueltyeboah/), [Francesco Di Chiara](https://www.linkedin.com/in/dichiarafrancesco/) and [Mingliang Liu](https://www.linkedin.com/in/liuml07/) + +Today, Netflix runs two Flink autoscalers. That is exactly one more than we want. We built the first one in-house years ago, when there was no mature option suited to our platform. The second came from the Apache Flink community, and it can scale workloads our homegrown system was never designed for. We now run both in production and are steadily converging on the open-source one. Along the way we learned some hard lessons about metrics, cost, and the real price of maintaining infrastructure you could instead adopt, and we hope they are useful whether you run a handful of Flink jobs or tens of thousands. + +## Why autoscaling is not optional at our scale + +Netflix has run stream processing on Apache Flink since 2017. As of 2026 we operate more than 30,000 Flink jobs across multiple AWS regions. Most are not deployed by hand; they are generated by our managed platform [Data Mesh](https://netflixtechblog.com/data-mesh-a-data-movement-and-processing-platform-netflix-1288bcab2873), so the majority of users never touch a Flink job directly. A smaller but growing set are custom jobs, built and operated by teams across the company for use cases like personalization, Ads, and Live events. They range from single-operator jobs that shuttle records between Kafka topics to stateful pipelines with branches, joins, and terabytes of state, and their load swings with daily cycles, launches, and regional failovers. + +Provisioning every one of those jobs for its peak is wasteful; provisioning for the average causes lag during surges. And in our platform a scaling action is not free: by default it means taking a savepoint, stopping the job gracefully, and restarting it at the new size, which for a large stateful job can take minutes. That leaves a genuinely hard question: *how do you give each job the resources it needs, when it needs them, without a human in the loop and without breaking anything?* + +## The first autoscaler: watching from outside + +Our first answer, built around 2019, was an autoscaler shaped like a stream-processing job. It ran on [Mantis](https://netflixtechblog.com/open-sourcing-mantis-a-platform-for-building-cost-effective-realtime-operations-focused-5b8ff387813a), consuming a live feed of cluster-level metrics from [Atlas](https://netflixtechblog.com/introducing-atlas-netflixs-primary-telemetry-platform-bd31f4d8ed9a), our telemetry platform, including CPU, network, Kafka lag, input-rate, and consume-rate signals for every job. The scaler combined lag-derived catch-up time, CPU/network utilization thresholds, observed performance history, and regression over recent input rate to decide when to scale up or whether a smaller cluster could handle the lookahead window. Because the autoscaler operates independently of the Flink platform, it remains unaffected by issues within Flink itself. Building it as a streaming job also made it easy to scale. Each autoscaler node handled the metrics for a subset of Flink jobs, and we never had to write custom sharding or coordination logic to keep up with a growing Flink fleet. It reliably cut resource usage by 25–45% across thousands of managed pipelines. Check our previous talk at [Flink Forward 2020](https://www.youtube.com/watch?v=NV0jvA5ZDNc). + +But watching from outside has a ceiling. The system reasoned about a whole cluster through coarse container metrics, and it scaled a single knob, the total TaskManager count, so every operator in a job moved together. That fit the simple, single-operator pipelines it was built for, but not the multi-operator, stateful DAGs that teams were increasingly bringing to us for Ads, recommendations, and games. Those were exactly the jobs it could not reason about, and supporting each new case meant more custom logic rather than any general capability. + +The autoscaler is only as good as the metrics served by external systems beneath it. Those metrics could miss real trouble: a job could be completely busy without any of it showing up as CPU utilization, leaving the job stuck in a degraded state the scaler had no way to see. Recently a networking migration quietly changed how some traffic was reported, and a subset of the Atlas metrics the scaler relied on stopped capturing everything accurately. The gap stayed invisible until it surfaced in production much later. + +It was time to reconsider *build* versus *buy*. + +## The second autoscaler: reasoning from inside + +When we started, the Flink community had no mature autoscaler to offer. By the time we re-evaluated, it did: the [Apache Flink Autoscaler](https://nightlies.apache.org/flink/flink-kubernetes-operator-docs-main/docs/custom-resource/autoscaler/). Instead of watching containers from outside, it reasons from inside the job. + +![](https://miro.medium.com/v2/resize:fit:1400/format:webp/1*Rn81hdaUHY94Sf8S-GOHZA.png) + +Figure 1: Architecture of the two Flink autoscalers + +Its key idea is to estimate each operator’s [true processing rate](https://www.usenix.org/conference/osdi18/presentation/kalavri) (TPR): the throughput it could sustain if it were fully busy. Flink reports, per subtask, the fraction of each second spent doing actual work, separate from time spent backpressured or idle. Dividing observed throughput by that busy fraction extrapolates capacity to full utilization: an operator handling 700 records/sec while busy 70% of the time has a TPR of 700 / 0.7 = 1,000 records/sec. Starting from the sources, the autoscaler walks the job graph and uses each operator’s TPR, its input/output ratios, and a target utilization to compute the parallelism every vertex needs so that no operator becomes the bottleneck, rather than resizing the whole cluster as a unit. + +![](https://miro.medium.com/v2/resize:fit:1400/format:webp/1*5vX9EEmUr1r5EDgWB8U8ew.png) + +Figure 2: Flink job DAG: current → desired parallelism per vertex, based on busyness + +The two approaches make a different contract, summarized below. + +![](https://miro.medium.com/v2/resize:fit:1400/format:webp/1*nXaMxuXoCByEBitEs7iw-A.png) + +Table 1: Comparison of the two Flink autoscalers + +The decisive difference for us is the last two rows: the OSS autoscaler can scale exactly the stateful, multi-operator jobs our homegrown system could not, and it lets each job carry its own configuration — stabilization periods, thresholds, and other scaling behavior tuned to the workload. That made it the natural fit for the custom jobs teams had been scaling by hand. + +## Making it work at Netflix scale + +Adopting the *algorithm* was straightforward; the community had done the hard part. The work for us was running it reliably across our own jobs, and this is where our system differs most from the stock open-source deployment. + +## Get Netflix Technology Blog’s stories in your inbox + +Join Medium for free to get updates from this writer. + +*Firstly*, the OSS autoscaler was originally architected to reside within the Kubernetes Operator for Flink, but our Flink platform runs on its own control plane, not that operator (see our previous talk at [Current Conference 2024](https://current.confluent.io/2024-sessions/building-a-scalable-flink-platform-a-tale-of-15-000-jobs-at-netflix)). The community later made a fantastic decision to keep the core logic as a standalone library. They refactored four generic interfaces that made it easy to plug directly into our internal ecosystem: a context carrying job metadata and REST API info, a state store, an event handler, and a realizer that applies scaling decisions. + +That service is a Spring Boot application whose orchestration runs on [Temporal](https://netflixtechblog.com/how-temporal-powers-reliable-cloud-operations-at-netflix-73c69ccb5953), the durable workflow engine. An orchestrator workflow polls our Flink control plane about once a minute for the jobs with autoscaling enabled, and starts one long-running workflow per job. Each per-job workflow pulls that job’s per-vertex metrics from its Flink JobManager, runs the OSS evaluation algorithm, and, when a scaling decision results, hands it to a *realizer* that actuates the change through our Flink control plane. + +![](https://miro.medium.com/v2/resize:fit:1400/format:webp/1*ujsQFlUsGVC-qMycK4TjSw.png) + +Figure 3: The OSS-based Flink Autoscaler architecture with Temporal workflows + +The workflow-per-job design was a direct response to pain. We first ran evaluations in a single batch loop over the whole set of jobs, and it was fragile: one slow or misbehaving job could stall metric collection and scaling for every job behind it. Giving each job its own durable workflow isolated that blast radius, so a single problematic job now fails and retries on its own, and the runtime scales out as we onboard more jobs. + +*Secondly*, three engineering gaps stood between “works in community” and “works at Netflix scale”: + +- **Metric collection at high parallelism.** On big jobs, pulling metrics from the JobManager became a bottleneck, and part of the cause was in Flink’s runtime. To address that, we changed the JobManager to cache transient metric names and clean them up once instead of rescanning on every fetch, and we added server-side filtering so the autoscaler asks only for the metrics it needs. This let the autoscaler work on jobs up to 3,000 Flink subtasks, where it had previously struggled above roughly 1,000. Those are in our internal fork of Flink release, while some are contributed upstream such as [FLINK-36172](https://issues.apache.org/jira/browse/FLINK-36172). +- **Preserving forward chaining.** Two separate vertices joined by a *forward* connection must run at the same parallelism, because records are handed over in memory on a fixed local channel. Scale one of them alone and Flink does not fail; it silently converts that edge into a network shuffle. Our fork detects forward-connected subgraphs and scales each as a unit. +- **Respecting sink limits.** Some sinks have finite write capacity, so we added detection for async-sink backpressure (also a fork change) to keep the autoscaler from scaling a job up into a sink that cannot absorb more. + +Before it actuates anything, the realizer runs a set of safety checks. For example, it refuses to scale a job down in a region being evacuated during a company-wide region failover. It also verifies there is enough disk for the new cluster to hold the job’s checkpoint state, and it adds a small standby buffer for larger clusters. + +## The road to one autoscaler + +Last year, the OSS-based autoscaler achieved general availability for custom jobs at Netflix, yielding promising initial outcomes. For instance, our client telemetry and logging team achieved a 58% reduction in its annualized Flink compute expenditures, saving approximately $1.1 million annually. This efficiency is driven by three key factors. First, whereas static provisioning must always account for peak loads, autoscaling dynamically adapts to daily cycles, capturing the drop in traffic during nights and weekends compared to weekday peaks. Second, rather than relying on teams to manually optimize resources following performance improvements or post-holiday slowdowns, the autoscaler continually adjusts capacity. Finally, adopting uniform container dimensions enables superior bin-packing and more granular scaling increments. + +Additionally, scaling down too eagerly is its own trap. Cut too deep and CPU saturates, lag spikes, and the system cannot react instantly because its metric window and stabilization period have to rebuild after each restart. We now run a target utilization of 0.45, below the community default of 0.7, deliberately trading a little efficiency for stability. Fewer and calmer rescales are worth the marginal cost for large stateful jobs. + +While our scaler provides fine-grained signals and vertex-level decision units for stateful DAGs, fast rescaling still heavily depends on Flink Core’s state restoration performance. Today, the biggest remaining cost in scaling a stateful job isn’t the scaler’s logic — it’s the restart and state recovery process itself. Flink 2 addresses this through its [disaggregated state](https://nightlies.apache.org/flink/flink-docs-master/docs/ops/state/disaggregated_state/) architecture, keeping state in external storage rather than on local disk, which can sharply reduce how much a rescale or recovery depends on total state size. Having started supporting Flink 2.2 at Netflix, we plan on experimenting with this new state backend to see if it can help eliminate state recovery bottlenecks when scaling large stateful jobs. + +Looking ahead, we aim to migrate all internal scaler use cases onto the new one based on OSS autoscaler to simplify our operational surface area. + +## Key Takeaways + +Along the way, three lessons that generalize beyond Flink: + +- **Metric choice matters more than algorithm sophistication.** Our most useful debugging was rarely about the scaling math; it was about which signal to trust most. Understand your metrics before you tune your algorithm. +- **Set sensible defaults, but leave room to tune.** Our managed jobs are similar enough that one good default covers most of them untouched, which is the point of a platform. But forcing a single configuration on every job punishes the ones that do not fit, so we pair defaults with per-job overrides and deliberately hide the knobs that need deep expertise. Most teams should never have to think about the autoscaler. +- **Adopt, then extend.** We built in-house because in 2019 nothing mature fit our platform. When a strong community project appeared, the right move was neither to defend our investment forever nor to rip it out overnight, but to adopt it for new workloads, contribute fixes back, and plan a deliberate migration. + +*Thanks to the Flink and Data Mesh teams for the control-plane changes this work depended on, to the Temporal team and our early pilot teams, and to the Apache Flink autoscaler maintainers whose foundation we built on. Special thanks to Andy Zhang, Calvin Cheung, Daniel Trager, Guil Pires, Mark Cho, Matthew Kornitsky, Nikhil Sulegaon, Sujay Jain, and Tom Lee.* \ No newline at end of file diff --git a/temp/sylvainkalache.com/ai-handles-incidents-engineers-lose-touch-with-their-systems/ai-handles-incidents-engineers-lose-touch-with-their-systems.md b/temp/ai-handles-incidents-engineers-lose-touch-with-their-systems-zh/ai-handles-incidents-engineers-lose-touch-with-their-systems.md similarity index 100% rename from temp/sylvainkalache.com/ai-handles-incidents-engineers-lose-touch-with-their-systems/ai-handles-incidents-engineers-lose-touch-with-their-systems.md rename to temp/ai-handles-incidents-engineers-lose-touch-with-their-systems-zh/ai-handles-incidents-engineers-lose-touch-with-their-systems.md diff --git a/temp/wan27-test.png b/temp/wan27-test.png deleted file mode 100644 index 9f805d13..00000000 Binary files a/temp/wan27-test.png and /dev/null differ