SRE weekly 所有文章

This commit is contained in:
2026-09-12 17:23:01 +08:00
parent 409b40ddcb
commit af7633f9dc
8486 changed files with 4489990 additions and 7 deletions

View File

@@ -0,0 +1,129 @@
<!doctype html><html lang=en-us><head><meta charset=utf-8><meta name=viewport content="width=device-width,initial-scale=1"><meta http-equiv=x-ua-compatible content="IE=edge"><meta name=generator content="Source Themes Academic 4.4.0"><meta name=author content="Vallery Lancey"><meta name=description content="What Happened In the most recent Kubernetes SIG-Network meeting (2020-04-02), a meta issue was raised. People had noted a decrease in new network-related issues. The reason, it was discovered, wasn&rsquo;t that new issues weren&rsquo;t being filed. It was that new issues weren&rsquo;t being labelled in the way they were normally labelled.
The Kubernetes Github repo has a fixed set of issue labels, which are managed by a bot (through automatic triggers, and comment commands)."><link rel=alternate hreflang=en-us href=https://timewitch.net/post/2020-04-04-cronjob-retro/><meta name=theme-color content="hsl(339, 90%, 68%)"><link rel=stylesheet href=https://cdnjs.cloudflare.com/ajax/libs/academicons/1.8.6/css/academicons.min.css integrity="sha256-uFVgMKfistnJAfoCUQigIl+JfUaP47GrRKjf6CTPVmw=" crossorigin=anonymous><link rel=stylesheet href=https://use.fontawesome.com/releases/v5.6.0/css/all.css integrity=sha384-aOkxzJ5uQz7WBObEZcHvV5JvRW3TUc2rNPA7pe3AwnsUohiw1Vj2Rgx2KSOkF5+h crossorigin=anonymous><link rel=stylesheet href=https://cdnjs.cloudflare.com/ajax/libs/fancybox/3.2.5/jquery.fancybox.min.css integrity="sha256-ygkqlh3CYSUri3LhQxzdcm0n1EQvH2Y+U5S2idbLtxs=" crossorigin=anonymous><link rel=stylesheet href=https://cdnjs.cloudflare.com/ajax/libs/highlight.js/9.15.6/styles/github.min.css crossorigin=anonymous title=hl-light><link rel=stylesheet href=https://cdnjs.cloudflare.com/ajax/libs/highlight.js/9.15.6/styles/github.min.css crossorigin=anonymous title=hl-dark disabled><link rel=stylesheet href="https://fonts.googleapis.com/css?family=Montserrat:400,700|Roboto:400,400italic,700|Roboto+Mono&display=swap"><link rel=stylesheet href=/css/academic.min.90ec218aa3be0bd4ce472018635fbb3e.css><script>window.ga=window.ga||function(){(ga.q=ga.q||[]).push(arguments)};ga.l=+new Date;ga('create','UA-147393375-1','auto');ga('require','eventTracker');ga('require','outboundLinkTracker');ga('require','urlChangeTracker');ga('send','pageview');</script><script async src=https://www.google-analytics.com/analytics.js></script><script async src=https://cdnjs.cloudflare.com/ajax/libs/autotrack/2.4.1/autotrack.js integrity="sha512-HUmooslVKj4m6OBu0OgzjXXr+QuFYy/k7eLI5jdeEy/F4RSgMn6XRWRGkFi5IFaFgy7uFTkegp3Z0XnJf3Jq+g==" crossorigin=anonymous></script><link rel=manifest href=/index.webmanifest><link rel=icon type=image/png href=/img/icon-32.png><link rel=apple-touch-icon type=image/png href=/img/icon-192.png><link rel=canonical href=https://timewitch.net/post/2020-04-04-cronjob-retro/><meta property=twitter:card content=summary><meta property=twitter:site content=@vllry><meta property=twitter:creator content=@vllry><meta property=og:site_name content="Vallery Lancey"><meta property=og:url content=https://timewitch.net/post/2020-04-04-cronjob-retro/><meta property=og:title content="Kubernetes CronJob Failed For 24 Days: a Retrospective | Vallery Lancey"><meta property=og:description content="What Happened In the most recent Kubernetes SIG-Network meeting (2020-04-02), a meta issue was raised. People had noted a decrease in new network-related issues. The reason, it was discovered, wasn&rsquo;t that new issues weren&rsquo;t being filed. It was that new issues weren&rsquo;t being labelled in the way they were normally labelled.
The Kubernetes Github repo has a fixed set of issue labels, which are managed by a bot (through automatic triggers, and comment commands)."><meta property=og:image content=https://timewitch.net/img/icon-192.png><meta property=twitter:image content=https://timewitch.net/img/icon-192.png><meta property=og:locale content=en-us><meta property=article:published_time content=2020-04-04T19:00:00-07:00><meta property=article:modified_time content=2020-04-04T19:25:08-07:00><title>Kubernetes CronJob Failed For 24 Days: a Retrospective | Vallery Lancey</title></head><body id=top data-spy=scroll data-offset=70 data-target=#TableOfContents><aside class=search-results id=search><div class=container><section class=search-header><div class="row no-gutters justify-content-between mb-3"><div class=col-6><h1>Search</h1></div><div class="col-6 col-search-close"><a class=js-search href=#><i class="fas fa-times-circle text-muted" aria-hidden=true></i></a></div></div><div id=search-box><input name=q id=search-query placeholder=Search... autocapitalize=off autocomplete=off autocorrect=off spellcheck=false type=search></div></section><section class=section-search-results><div id=search-hits></div></section></div></aside><nav class="navbar navbar-light fixed-top navbar-expand-lg py-0 compensate-for-scrollbar" id=navbar-main><div class=container><a class=navbar-brand href=/>Vallery Lancey</a>
<button type=button class=navbar-toggler data-toggle=collapse data-target=#navbar aria-controls=navbar aria-expanded=false aria-label="Toggle navigation">
<span><i class="fas fa-bars"></i></span></button><div class="collapse navbar-collapse" id=navbar><ul class="navbar-nav mr-auto"><li class=nav-item><a class=nav-link href=/#about><span>Home</span></a></li><li class=nav-item><a class=nav-link href=/#posts><span>Posts</span></a></li><li class=nav-item><a class=nav-link href=/#talks><span>Talks</span></a></li><li class=nav-item><a class=nav-link href=/#contact><span>Contact</span></a></li></ul><ul class="navbar-nav ml-auto"><li class=nav-item><a class="nav-link js-search" href=#><i class="fas fa-search" aria-hidden=true></i></a></li></ul></div></div></nav><article class=article itemscope itemtype=http://schema.org/Article><div class="article-container pt-3"><h1 itemprop=name>Kubernetes CronJob Failed For 24 Days: a Retrospective</h1><meta content="2020-04-04 19:00:00 -0700 -0700" itemprop=datePublished><meta content="2020-04-04 19:25:08 -0700 -0700" itemprop=dateModified><div class=article-metadata><span class=article-date>Last updated on
<time>Apr 4, 2020</time></span>
<span class=middot-divider></span><span class=article-reading-time>6 min read</span><div class=share-box aria-hidden=true><ul class=share><li><a href="https://twitter.com/intent/tweet?url=https://timewitch.net/post/2020-04-04-cronjob-retro/&amp;text=Kubernetes%20CronJob%20Failed%20For%2024%20Days:%20a%20Retrospective" target=_blank rel=noopener class=share-btn-twitter><i class="fab fa-twitter"></i></a></li><li><a href="mailto:?subject=Kubernetes%20CronJob%20Failed%20For%2024%20Days:%20a%20Retrospective&amp;body=https://timewitch.net/post/2020-04-04-cronjob-retro/" target=_blank rel=noopener class=share-btn-email><i class="fas fa-envelope"></i></a></li><li><a href="https://www.linkedin.com/shareArticle?url=https://timewitch.net/post/2020-04-04-cronjob-retro/&amp;title=Kubernetes%20CronJob%20Failed%20For%2024%20Days:%20a%20Retrospective" target=_blank rel=noopener class=share-btn-linkedin><i class="fab fa-linkedin-in"></i></a></li><li><a href="https://web.whatsapp.com/send?text=Kubernetes%20CronJob%20Failed%20For%2024%20Days:%20a%20Retrospective%20https://timewitch.net/post/2020-04-04-cronjob-retro/" target=_blank rel=noopener class=share-btn-whatsapp><i class="fab fa-whatsapp"></i></a></li><li><a href="https://service.weibo.com/share/share.php?url=https://timewitch.net/post/2020-04-04-cronjob-retro/&amp;title=Kubernetes%20CronJob%20Failed%20For%2024%20Days:%20a%20Retrospective" target=_blank rel=noopener class=share-btn-weibo><i class="fab fa-weibo"></i></a></li></ul></div></div></div><div class=article-container><div class=article-style itemprop=articleBody><h1 id=what-happened>What Happened</h1><p>In the most recent Kubernetes SIG-Network meeting (2020-04-02),
a meta issue was raised.
People had noted a decrease in new network-related issues.
The reason,
it was discovered,
wasn&rsquo;t that new issues weren&rsquo;t being filed.
It was that new issues weren&rsquo;t being labelled in the way they were normally labelled.</p><p>The Kubernetes Github repo has a fixed set of issue labels,
which are managed by a <a href=https://github.com/k8s-ci-robot target=_blank>bot</a>
(through automatic triggers, and comment <a href=https://prow.k8s.io/command-help target=_blank>commands</a>).
In particular,
all issues are labelled with either <code>needs-sig</code>,
or one or more <code>sig-___</code> labels.
SIG-Network issues are labelled with <code>sig-network</code>.</p><p>Starting in early 2019,
SIG-Network made a point of gradually burning down the issue backlog.
We starting applying an &ldquo;old&rdquo; issue label,
<code>triage/unresolved</code>,
to all open SIG-Network issues that had not yet been inspected and verified/prioritized.
Once an open issue was confirmed to be legitimate,
a member would comment <code>/remove-triage unresolved</code> to remove the label.
This allowed us to easily filter between the inbound and triaged issues.</p><figure><img src=github-issue-filter.png></figure><p>This process was a bit tedious,
and relied on individuals regularly stepping up.
To alleviate this,
I created a simple Github bot,
called <a href=https://github.com/athenabot/k8s-issues target=_blank>athenabot</a>.
Athenabot has several functions,
however the primary function of athenabot is to label new <code>sig-network</code> issues
with <code>triage/unresolved</code>.
This functionality had broken, for a total of 24 days.</p><figure><img src=athenabot-comment.png></figure><p>Athenabot runs as a <a href=https://kubernetes.io/docs/concepts/workloads/controllers/cron-jobs/ target=_blank>Kubernetes CronJob</a>,
on a tiny cluster that I use for medium to long term experimentation.
Kubernetes CronJobs are configured with a traditional cron stanza,
and launch a pod when it is time to run.</p><p>Capacity in Kubernetes is measured by <a href=https://kubernetes.io/docs/concepts/configuration/manage-compute-resources-container/ target=_blank>requests</a> of CPU and memory.
All pods can specify requests (EG 1 CPU, 1 GB memory).
When pod scheduling occurs,
the default scheduler behavior is to only consider nodes where the
sum of requests would not exceed the total machine resources.
Athenabot&rsquo;s pods were set to request 50m CPU (5% of a CPU), and 64Mi memory when launched.</p><p>The cluster was both over capacity,
and using an ephemeral VM.
Periodically,
the VM would shut down,
and there would be no node available to run workloads.
When a new node booted,
all existing pods would need to be re-scheduled on to the node.
As not all pods could fit on the node,
due to CPU requests exceeding capacity,
some would remain unscheduled.
There are multiple possible combinations to fill a node.
As a CronJob, athenabot&rsquo;s pod would not always exist
(it only exists for roughly 30 seconds, every hour),
which meant that it often would not even be considered in initial node scheduling.
The athenabot pod may not be scheduled in a given combination,
or the scheduling may not leave a 50m CPU &ldquo;hole&rdquo; for a future pod to claim.
This would cause periodic scheduling failures.
There is not enough data to indicate exactly how often this occurred.</p><p>CronJobs have a <a href=https://github.com/kubernetes/kubernetes/blob/ccded1494116d6aa1ac3f4612b4a613b56a2044a/pkg/controller/cronjob/utils.go#L123 target=_blank>known design flaw</a>
that will cause repeated scheduling failures to become &ldquo;permanent&rdquo;
in specific conditions.
CronJob schedule parsing is inefficient at large scale,
when checking over long timespans and many iterations.
When deciding if a pod should be launched,
the CronJob controller iterates over all cron schedules times,
from the most recent run time to the present.
There is a cap on checking at most 100 consecutive schedule times.</p><p>The &ldquo;100 missed starts&rdquo; failure was recognizable by mentally pattern-matching against CronJob and child
Job statuses, and seeing 24 days since a run had occurred.
It was verifiable by checking <a href=https://kubernetes.io/docs/tasks/debug-application-cluster/ target=_blank>events</a>
and seeing the error.</p><pre><code>$ kubectl get cronjobs -A
NAMESPACE NAME SCHEDULE SUSPEND ACTIVE LAST SCHEDULE AGE
athenabot athenabot-k8sissues 12 * * * * False 0 24d 166d
</code></pre><p>The CronJob initially failed to schedule because a pending pod already existed
(and the CronJob had a concurrency policy that forbid concurrency).</p><p>In order to hit the terminal failure,
the CronJob must have not run scheduled for at least 100 schedule times,
and one of the following must be true:</p><ul><li><p>There is no starting deadline (<code>startingDeadlineSeconds = nil</code>)</p></li><li><p>The starting deadline exceeds the span of the last 100 missed schedules
(e.g. schedule is minutely, starting deadline is 120 minutes)</p></li></ul><p>Athenabot hit the former case.
At some point,
some combination of &ldquo;unlucky scheduling&rdquo; and ephemeral VM unavailability led to
(at least) 100 consecutive schedules where an athenabot pod could not run
(which may have been as low as &ldquo;just over 99 hours&rdquo;, depending on exact timing).
At that point,
the CronJob could never run again without human intervention.
There are multiple possible resolutions
(setting <code>startingDeadlineSeconds</code>, setting a fake recent <code>lastScheduleTime</code>, or recreating the CronJob).
I opted to set a <code>startingDeadlineSeconds</code>.
Setting this time allowed the CronJob to continue running.
Issue labelling has resumed,
and I am making sure that the backlog is labelled.</p><h1 id=takeaways>Takeaways</h1><p>There were multiple, unideal decisions made:</p><ul><li>Running what is functionally a &ldquo;production&rdquo; workload in a fragile test environment.</li><li>A lack of any monitoring.</li><li>Non-robust configuration of the CronJob.</li></ul><p>Perhaps, using a Kubernetes CronJob altogether was the wrong approach.
Kubernetes CronJobs are complex (adding 2 layers of abstraction, themselves and Kubernetes Jobs),
plus pod scheduling and management concerns.
A traditional cron would still be vulnerable to similar neglect,
but as a less mechanically complex system,
has far fewer failure modes.</p><p>Additionally, the athenabot CronJob has unnecessarily large requests.
In practice, less memory, and <em>substantially</em> less CPU than the requested values is needed.
Setting these lower would not have prevented this particular incident,
but is prudent nonetheless.
<em>Removing</em> the requests <em>would</em> have prevented this incident.
Requests are only for scheduling bookkeeping
- they do not represent a guarantee, or enforcement.
Setting no requests would allow the CronJob to schedule on to a node regardless of capacity,
and the athenabot process is too small to cause undue strain on the node&rsquo;s resources.
However, making a decision like this at larger scale (with larger and/or more workloads)
<em>can</em> become impactful.</p><p>There is an <a href=https://github.com/kubernetes/enhancements/pull/1554 target=_blank>open proposal</a>
gaining traction to build the triage label automation into the official Kubernetes Github automation.
Regardless of having the athenabot code be open-source,
there is risk in productivity tooling being maintained and hosted by a single individual.
Additionally, the triage workflow has been useful within the SIG,
and will likely be helpful for others.</p><h2 id=additional-commentary>Additional Commentary</h2><p>I was quite exasperated to hit this issue in a personal setup,
as this is something we patched at Lyft,
and I am working on porting into upstream Kubernetes (a draft upstream PR is <a href=https://github.com/kubernetes/kubernetes/pull/89397 target=_blank>here</a>).
As that&rsquo;s some work-work content,
I&rsquo;ll save commentary on that for another venue.</p></div><div class="media author-card" itemscope itemtype=http://schema.org/Person><img class="portrait mr-3" src=/authors/admin/avatar_hu51f7efccd86af403139bad264a1ec51e_99131_250x250_fill_lanczos_center_2.png itemprop=image alt=Avatar><div class=media-body><h5 class=card-title itemprop=name><a href=https://timewitch.net>Vallery Lancey</a></h5><h6 class=card-subtitle>Software/Reliability Engineer</h6><p class=card-text itemprop=description>I work on distributed systems and reliability.</p><ul class=network-icon aria-hidden=true><li><a itemprop=sameAs href=/#contact><i class="fas fa-envelope"></i></a></li><li><a itemprop=sameAs href=https://linkedin.com/in/vallery/ target=_blank rel=noopener><i class="fab fa-linkedin"></i></a></li><li><a itemprop=sameAs href=https://twitter.com/vllry target=_blank rel=noopener><i class="fab fa-twitter"></i></a></li><li><a itemprop=sameAs href=https://github.com/vllry target=_blank rel=noopener><i class="fab fa-github"></i></a></li></ul></div></div></div></article><script src=https://cdnjs.cloudflare.com/ajax/libs/jquery/3.4.1/jquery.min.js integrity="sha256-CSXorXvZcTkaix6Yvo6HppcZGetbYMGWSFlBw8HfCJo=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/jquery.imagesloaded/4.1.4/imagesloaded.pkgd.min.js integrity="sha256-lqvxZrPLtfffUl2G/e7szqSvPBILGbwmsGE1MKlOi0Q=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/jquery.isotope/3.0.6/isotope.pkgd.min.js integrity="sha256-CBrpuqrMhXwcLLUd5tvQ4euBHCdh7wGlDfNz8vbu/iI=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/fancybox/3.2.5/jquery.fancybox.min.js integrity="sha256-X5PoE3KU5l+JcX+w09p/wHl9AzK333C4hJ2I9S5mD4M=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/highlight.js/9.15.6/highlight.min.js integrity="sha256-aYTdUrn6Ow1DDgh5JTc3aDGnnju48y/1c8s1dgkYPQ8=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/highlight.js/9.15.6/languages/go.min.js></script><script>hljs.initHighlightingOnLoad();</script><script>const search_index_filename="/index.json";const i18n={'placeholder':"Search...",'results':"results found",'no_results':"No results found"};const content_type={'post':"Posts",'project':"Projects",'publication':"Publications",'talk':"Talks"};</script><script id=search-hit-fuse-template type=text/x-template>
<div class="search-hit" id="summary-{{key}}">
<div class="search-hit-content">
<div class="search-hit-name">
<a href="{{relpermalink}}">{{title}}</a>
<div class="article-metadata search-hit-type">{{type}}</div>
<p class="search-hit-description">{{snippet}}</p>
</div>
</div>
</div>
</script><script src=https://cdnjs.cloudflare.com/ajax/libs/fuse.js/3.2.1/fuse.min.js integrity="sha256-VzgmKYmhsGNNN4Ph1kMW+BjoYJM2jV5i4IlFoeZA9XI=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/mark.js/8.11.1/jquery.mark.min.js integrity="sha256-4HLtjeVgH0eIB3aZ9mLYF6E8oU5chNdjU6p6rrXpl9U=" crossorigin=anonymous></script><script src=/js/academic.min.4e51017b38c7ebaadd7e25fc9503f88c.js></script><div class=container><footer class=site-footer><p class=powered-by>&copy; 2026 Vallery Lancey &middot;
Powered by the
<a href=https://sourcethemes.com/academic/ target=_blank rel=noopener>Academic theme</a> for
<a href=https://gohugo.io target=_blank rel=noopener>Hugo</a>.
<span class=float-right aria-hidden=true><a href=# id=back_to_top><span class=button_icon><i class="fas fa-chevron-up fa-2x"></i></span></a></span></p></footer></div><div id=modal class="modal fade" role=dialog><div class=modal-dialog><div class=modal-content><div class=modal-header><h5 class=modal-title>Cite</h5><button type=button class=close data-dismiss=modal aria-label=Close>
<span aria-hidden=true>&times;</span></button></div><div class=modal-body><pre><code class="tex hljs"></code></pre></div><div class=modal-footer><a class="btn btn-outline-primary my-1 js-copy-cite" href=# target=_blank><i class="fas fa-copy"></i>Copy</a>
<a class="btn btn-outline-primary my-1 js-download-cite" href=# target=_blank><i class="fas fa-download"></i>Download</a><div id=modal-error></div></div></div></div></div></body></html>