129 lines
20 KiB
HTML
129 lines
20 KiB
HTML
<!doctype html><html lang=en-us><head><meta charset=utf-8><meta name=viewport content="width=device-width,initial-scale=1"><meta http-equiv=x-ua-compatible content="IE=edge"><meta name=generator content="Source Themes Academic 4.4.0"><meta name=author content="Vallery Lancey"><meta name=description content="What Happened In the most recent Kubernetes SIG-Network meeting (2020-04-02), a meta issue was raised. People had noted a decrease in new network-related issues. The reason, it was discovered, wasn’t that new issues weren’t being filed. It was that new issues weren’t being labelled in the way they were normally labelled.
|
|
The Kubernetes Github repo has a fixed set of issue labels, which are managed by a bot (through automatic triggers, and comment commands)."><link rel=alternate hreflang=en-us href=https://timewitch.net/post/2020-04-04-cronjob-retro/><meta name=theme-color content="hsl(339, 90%, 68%)"><link rel=stylesheet href=https://cdnjs.cloudflare.com/ajax/libs/academicons/1.8.6/css/academicons.min.css integrity="sha256-uFVgMKfistnJAfoCUQigIl+JfUaP47GrRKjf6CTPVmw=" crossorigin=anonymous><link rel=stylesheet href=https://use.fontawesome.com/releases/v5.6.0/css/all.css integrity=sha384-aOkxzJ5uQz7WBObEZcHvV5JvRW3TUc2rNPA7pe3AwnsUohiw1Vj2Rgx2KSOkF5+h crossorigin=anonymous><link rel=stylesheet href=https://cdnjs.cloudflare.com/ajax/libs/fancybox/3.2.5/jquery.fancybox.min.css integrity="sha256-ygkqlh3CYSUri3LhQxzdcm0n1EQvH2Y+U5S2idbLtxs=" crossorigin=anonymous><link rel=stylesheet href=https://cdnjs.cloudflare.com/ajax/libs/highlight.js/9.15.6/styles/github.min.css crossorigin=anonymous title=hl-light><link rel=stylesheet href=https://cdnjs.cloudflare.com/ajax/libs/highlight.js/9.15.6/styles/github.min.css crossorigin=anonymous title=hl-dark disabled><link rel=stylesheet href="https://fonts.googleapis.com/css?family=Montserrat:400,700|Roboto:400,400italic,700|Roboto+Mono&display=swap"><link rel=stylesheet href=/css/academic.min.90ec218aa3be0bd4ce472018635fbb3e.css><script>window.ga=window.ga||function(){(ga.q=ga.q||[]).push(arguments)};ga.l=+new Date;ga('create','UA-147393375-1','auto');ga('require','eventTracker');ga('require','outboundLinkTracker');ga('require','urlChangeTracker');ga('send','pageview');</script><script async src=https://www.google-analytics.com/analytics.js></script><script async src=https://cdnjs.cloudflare.com/ajax/libs/autotrack/2.4.1/autotrack.js integrity="sha512-HUmooslVKj4m6OBu0OgzjXXr+QuFYy/k7eLI5jdeEy/F4RSgMn6XRWRGkFi5IFaFgy7uFTkegp3Z0XnJf3Jq+g==" crossorigin=anonymous></script><link rel=manifest href=/index.webmanifest><link rel=icon type=image/png href=/img/icon-32.png><link rel=apple-touch-icon type=image/png href=/img/icon-192.png><link rel=canonical href=https://timewitch.net/post/2020-04-04-cronjob-retro/><meta property=twitter:card content=summary><meta property=twitter:site content=@vllry><meta property=twitter:creator content=@vllry><meta property=og:site_name content="Vallery Lancey"><meta property=og:url content=https://timewitch.net/post/2020-04-04-cronjob-retro/><meta property=og:title content="Kubernetes CronJob Failed For 24 Days: a Retrospective | Vallery Lancey"><meta property=og:description content="What Happened In the most recent Kubernetes SIG-Network meeting (2020-04-02), a meta issue was raised. People had noted a decrease in new network-related issues. The reason, it was discovered, wasn’t that new issues weren’t being filed. It was that new issues weren’t being labelled in the way they were normally labelled.
|
|
The Kubernetes Github repo has a fixed set of issue labels, which are managed by a bot (through automatic triggers, and comment commands)."><meta property=og:image content=https://timewitch.net/img/icon-192.png><meta property=twitter:image content=https://timewitch.net/img/icon-192.png><meta property=og:locale content=en-us><meta property=article:published_time content=2020-04-04T19:00:00-07:00><meta property=article:modified_time content=2020-04-04T19:25:08-07:00><title>Kubernetes CronJob Failed For 24 Days: a Retrospective | Vallery Lancey</title></head><body id=top data-spy=scroll data-offset=70 data-target=#TableOfContents><aside class=search-results id=search><div class=container><section class=search-header><div class="row no-gutters justify-content-between mb-3"><div class=col-6><h1>Search</h1></div><div class="col-6 col-search-close"><a class=js-search href=#><i class="fas fa-times-circle text-muted" aria-hidden=true></i></a></div></div><div id=search-box><input name=q id=search-query placeholder=Search... autocapitalize=off autocomplete=off autocorrect=off spellcheck=false type=search></div></section><section class=section-search-results><div id=search-hits></div></section></div></aside><nav class="navbar navbar-light fixed-top navbar-expand-lg py-0 compensate-for-scrollbar" id=navbar-main><div class=container><a class=navbar-brand href=/>Vallery Lancey</a>
|
|
<button type=button class=navbar-toggler data-toggle=collapse data-target=#navbar aria-controls=navbar aria-expanded=false aria-label="Toggle navigation">
|
|
<span><i class="fas fa-bars"></i></span></button><div class="collapse navbar-collapse" id=navbar><ul class="navbar-nav mr-auto"><li class=nav-item><a class=nav-link href=/#about><span>Home</span></a></li><li class=nav-item><a class=nav-link href=/#posts><span>Posts</span></a></li><li class=nav-item><a class=nav-link href=/#talks><span>Talks</span></a></li><li class=nav-item><a class=nav-link href=/#contact><span>Contact</span></a></li></ul><ul class="navbar-nav ml-auto"><li class=nav-item><a class="nav-link js-search" href=#><i class="fas fa-search" aria-hidden=true></i></a></li></ul></div></div></nav><article class=article itemscope itemtype=http://schema.org/Article><div class="article-container pt-3"><h1 itemprop=name>Kubernetes CronJob Failed For 24 Days: a Retrospective</h1><meta content="2020-04-04 19:00:00 -0700 -0700" itemprop=datePublished><meta content="2020-04-04 19:25:08 -0700 -0700" itemprop=dateModified><div class=article-metadata><span class=article-date>Last updated on
|
|
<time>Apr 4, 2020</time></span>
|
|
<span class=middot-divider></span><span class=article-reading-time>6 min read</span><div class=share-box aria-hidden=true><ul class=share><li><a href="https://twitter.com/intent/tweet?url=https://timewitch.net/post/2020-04-04-cronjob-retro/&text=Kubernetes%20CronJob%20Failed%20For%2024%20Days:%20a%20Retrospective" target=_blank rel=noopener class=share-btn-twitter><i class="fab fa-twitter"></i></a></li><li><a href="mailto:?subject=Kubernetes%20CronJob%20Failed%20For%2024%20Days:%20a%20Retrospective&body=https://timewitch.net/post/2020-04-04-cronjob-retro/" target=_blank rel=noopener class=share-btn-email><i class="fas fa-envelope"></i></a></li><li><a href="https://www.linkedin.com/shareArticle?url=https://timewitch.net/post/2020-04-04-cronjob-retro/&title=Kubernetes%20CronJob%20Failed%20For%2024%20Days:%20a%20Retrospective" target=_blank rel=noopener class=share-btn-linkedin><i class="fab fa-linkedin-in"></i></a></li><li><a href="https://web.whatsapp.com/send?text=Kubernetes%20CronJob%20Failed%20For%2024%20Days:%20a%20Retrospective%20https://timewitch.net/post/2020-04-04-cronjob-retro/" target=_blank rel=noopener class=share-btn-whatsapp><i class="fab fa-whatsapp"></i></a></li><li><a href="https://service.weibo.com/share/share.php?url=https://timewitch.net/post/2020-04-04-cronjob-retro/&title=Kubernetes%20CronJob%20Failed%20For%2024%20Days:%20a%20Retrospective" target=_blank rel=noopener class=share-btn-weibo><i class="fab fa-weibo"></i></a></li></ul></div></div></div><div class=article-container><div class=article-style itemprop=articleBody><h1 id=what-happened>What Happened</h1><p>In the most recent Kubernetes SIG-Network meeting (2020-04-02),
|
|
a meta issue was raised.
|
|
People had noted a decrease in new network-related issues.
|
|
The reason,
|
|
it was discovered,
|
|
wasn’t that new issues weren’t being filed.
|
|
It was that new issues weren’t being labelled in the way they were normally labelled.</p><p>The Kubernetes Github repo has a fixed set of issue labels,
|
|
which are managed by a <a href=https://github.com/k8s-ci-robot target=_blank>bot</a>
|
|
(through automatic triggers, and comment <a href=https://prow.k8s.io/command-help target=_blank>commands</a>).
|
|
In particular,
|
|
all issues are labelled with either <code>needs-sig</code>,
|
|
or one or more <code>sig-___</code> labels.
|
|
SIG-Network issues are labelled with <code>sig-network</code>.</p><p>Starting in early 2019,
|
|
SIG-Network made a point of gradually burning down the issue backlog.
|
|
We starting applying an “old” issue label,
|
|
<code>triage/unresolved</code>,
|
|
to all open SIG-Network issues that had not yet been inspected and verified/prioritized.
|
|
Once an open issue was confirmed to be legitimate,
|
|
a member would comment <code>/remove-triage unresolved</code> to remove the label.
|
|
This allowed us to easily filter between the inbound and triaged issues.</p><figure><img src=github-issue-filter.png></figure><p>This process was a bit tedious,
|
|
and relied on individuals regularly stepping up.
|
|
To alleviate this,
|
|
I created a simple Github bot,
|
|
called <a href=https://github.com/athenabot/k8s-issues target=_blank>athenabot</a>.
|
|
Athenabot has several functions,
|
|
however the primary function of athenabot is to label new <code>sig-network</code> issues
|
|
with <code>triage/unresolved</code>.
|
|
This functionality had broken, for a total of 24 days.</p><figure><img src=athenabot-comment.png></figure><p>Athenabot runs as a <a href=https://kubernetes.io/docs/concepts/workloads/controllers/cron-jobs/ target=_blank>Kubernetes CronJob</a>,
|
|
on a tiny cluster that I use for medium to long term experimentation.
|
|
Kubernetes CronJobs are configured with a traditional cron stanza,
|
|
and launch a pod when it is time to run.</p><p>Capacity in Kubernetes is measured by <a href=https://kubernetes.io/docs/concepts/configuration/manage-compute-resources-container/ target=_blank>requests</a> of CPU and memory.
|
|
All pods can specify requests (EG 1 CPU, 1 GB memory).
|
|
When pod scheduling occurs,
|
|
the default scheduler behavior is to only consider nodes where the
|
|
sum of requests would not exceed the total machine resources.
|
|
Athenabot’s pods were set to request 50m CPU (5% of a CPU), and 64Mi memory when launched.</p><p>The cluster was both over capacity,
|
|
and using an ephemeral VM.
|
|
Periodically,
|
|
the VM would shut down,
|
|
and there would be no node available to run workloads.
|
|
When a new node booted,
|
|
all existing pods would need to be re-scheduled on to the node.
|
|
As not all pods could fit on the node,
|
|
due to CPU requests exceeding capacity,
|
|
some would remain unscheduled.
|
|
There are multiple possible combinations to fill a node.
|
|
As a CronJob, athenabot’s pod would not always exist
|
|
(it only exists for roughly 30 seconds, every hour),
|
|
which meant that it often would not even be considered in initial node scheduling.
|
|
The athenabot pod may not be scheduled in a given combination,
|
|
or the scheduling may not leave a 50m CPU “hole” for a future pod to claim.
|
|
This would cause periodic scheduling failures.
|
|
There is not enough data to indicate exactly how often this occurred.</p><p>CronJobs have a <a href=https://github.com/kubernetes/kubernetes/blob/ccded1494116d6aa1ac3f4612b4a613b56a2044a/pkg/controller/cronjob/utils.go#L123 target=_blank>known design flaw</a>
|
|
that will cause repeated scheduling failures to become “permanent”
|
|
in specific conditions.
|
|
CronJob schedule parsing is inefficient at large scale,
|
|
when checking over long timespans and many iterations.
|
|
When deciding if a pod should be launched,
|
|
the CronJob controller iterates over all cron schedules times,
|
|
from the most recent run time to the present.
|
|
There is a cap on checking at most 100 consecutive schedule times.</p><p>The “100 missed starts” failure was recognizable by mentally pattern-matching against CronJob and child
|
|
Job statuses, and seeing 24 days since a run had occurred.
|
|
It was verifiable by checking <a href=https://kubernetes.io/docs/tasks/debug-application-cluster/ target=_blank>events</a>
|
|
and seeing the error.</p><pre><code>$ kubectl get cronjobs -A
|
|
NAMESPACE NAME SCHEDULE SUSPEND ACTIVE LAST SCHEDULE AGE
|
|
athenabot athenabot-k8sissues 12 * * * * False 0 24d 166d
|
|
</code></pre><p>The CronJob initially failed to schedule because a pending pod already existed
|
|
(and the CronJob had a concurrency policy that forbid concurrency).</p><p>In order to hit the terminal failure,
|
|
the CronJob must have not run scheduled for at least 100 schedule times,
|
|
and one of the following must be true:</p><ul><li><p>There is no starting deadline (<code>startingDeadlineSeconds = nil</code>)</p></li><li><p>The starting deadline exceeds the span of the last 100 missed schedules
|
|
(e.g. schedule is minutely, starting deadline is 120 minutes)</p></li></ul><p>Athenabot hit the former case.
|
|
At some point,
|
|
some combination of “unlucky scheduling” and ephemeral VM unavailability led to
|
|
(at least) 100 consecutive schedules where an athenabot pod could not run
|
|
(which may have been as low as “just over 99 hours”, depending on exact timing).
|
|
At that point,
|
|
the CronJob could never run again without human intervention.
|
|
There are multiple possible resolutions
|
|
(setting <code>startingDeadlineSeconds</code>, setting a fake recent <code>lastScheduleTime</code>, or recreating the CronJob).
|
|
I opted to set a <code>startingDeadlineSeconds</code>.
|
|
Setting this time allowed the CronJob to continue running.
|
|
Issue labelling has resumed,
|
|
and I am making sure that the backlog is labelled.</p><h1 id=takeaways>Takeaways</h1><p>There were multiple, unideal decisions made:</p><ul><li>Running what is functionally a “production” workload in a fragile test environment.</li><li>A lack of any monitoring.</li><li>Non-robust configuration of the CronJob.</li></ul><p>Perhaps, using a Kubernetes CronJob altogether was the wrong approach.
|
|
Kubernetes CronJobs are complex (adding 2 layers of abstraction, themselves and Kubernetes Jobs),
|
|
plus pod scheduling and management concerns.
|
|
A traditional cron would still be vulnerable to similar neglect,
|
|
but as a less mechanically complex system,
|
|
has far fewer failure modes.</p><p>Additionally, the athenabot CronJob has unnecessarily large requests.
|
|
In practice, less memory, and <em>substantially</em> less CPU than the requested values is needed.
|
|
Setting these lower would not have prevented this particular incident,
|
|
but is prudent nonetheless.
|
|
<em>Removing</em> the requests <em>would</em> have prevented this incident.
|
|
Requests are only for scheduling bookkeeping
|
|
- they do not represent a guarantee, or enforcement.
|
|
Setting no requests would allow the CronJob to schedule on to a node regardless of capacity,
|
|
and the athenabot process is too small to cause undue strain on the node’s resources.
|
|
However, making a decision like this at larger scale (with larger and/or more workloads)
|
|
<em>can</em> become impactful.</p><p>There is an <a href=https://github.com/kubernetes/enhancements/pull/1554 target=_blank>open proposal</a>
|
|
gaining traction to build the triage label automation into the official Kubernetes Github automation.
|
|
Regardless of having the athenabot code be open-source,
|
|
there is risk in productivity tooling being maintained and hosted by a single individual.
|
|
Additionally, the triage workflow has been useful within the SIG,
|
|
and will likely be helpful for others.</p><h2 id=additional-commentary>Additional Commentary</h2><p>I was quite exasperated to hit this issue in a personal setup,
|
|
as this is something we patched at Lyft,
|
|
and I am working on porting into upstream Kubernetes (a draft upstream PR is <a href=https://github.com/kubernetes/kubernetes/pull/89397 target=_blank>here</a>).
|
|
As that’s some work-work content,
|
|
I’ll save commentary on that for another venue.</p></div><div class="media author-card" itemscope itemtype=http://schema.org/Person><img class="portrait mr-3" src=/authors/admin/avatar_hu51f7efccd86af403139bad264a1ec51e_99131_250x250_fill_lanczos_center_2.png itemprop=image alt=Avatar><div class=media-body><h5 class=card-title itemprop=name><a href=https://timewitch.net>Vallery Lancey</a></h5><h6 class=card-subtitle>Software/Reliability Engineer</h6><p class=card-text itemprop=description>I work on distributed systems and reliability.</p><ul class=network-icon aria-hidden=true><li><a itemprop=sameAs href=/#contact><i class="fas fa-envelope"></i></a></li><li><a itemprop=sameAs href=https://linkedin.com/in/vallery/ target=_blank rel=noopener><i class="fab fa-linkedin"></i></a></li><li><a itemprop=sameAs href=https://twitter.com/vllry target=_blank rel=noopener><i class="fab fa-twitter"></i></a></li><li><a itemprop=sameAs href=https://github.com/vllry target=_blank rel=noopener><i class="fab fa-github"></i></a></li></ul></div></div></div></article><script src=https://cdnjs.cloudflare.com/ajax/libs/jquery/3.4.1/jquery.min.js integrity="sha256-CSXorXvZcTkaix6Yvo6HppcZGetbYMGWSFlBw8HfCJo=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/jquery.imagesloaded/4.1.4/imagesloaded.pkgd.min.js integrity="sha256-lqvxZrPLtfffUl2G/e7szqSvPBILGbwmsGE1MKlOi0Q=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/jquery.isotope/3.0.6/isotope.pkgd.min.js integrity="sha256-CBrpuqrMhXwcLLUd5tvQ4euBHCdh7wGlDfNz8vbu/iI=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/fancybox/3.2.5/jquery.fancybox.min.js integrity="sha256-X5PoE3KU5l+JcX+w09p/wHl9AzK333C4hJ2I9S5mD4M=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/highlight.js/9.15.6/highlight.min.js integrity="sha256-aYTdUrn6Ow1DDgh5JTc3aDGnnju48y/1c8s1dgkYPQ8=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/highlight.js/9.15.6/languages/go.min.js></script><script>hljs.initHighlightingOnLoad();</script><script>const search_index_filename="/index.json";const i18n={'placeholder':"Search...",'results':"results found",'no_results':"No results found"};const content_type={'post':"Posts",'project':"Projects",'publication':"Publications",'talk':"Talks"};</script><script id=search-hit-fuse-template type=text/x-template>
|
|
<div class="search-hit" id="summary-{{key}}">
|
|
<div class="search-hit-content">
|
|
<div class="search-hit-name">
|
|
<a href="{{relpermalink}}">{{title}}</a>
|
|
<div class="article-metadata search-hit-type">{{type}}</div>
|
|
<p class="search-hit-description">{{snippet}}</p>
|
|
</div>
|
|
</div>
|
|
</div>
|
|
</script><script src=https://cdnjs.cloudflare.com/ajax/libs/fuse.js/3.2.1/fuse.min.js integrity="sha256-VzgmKYmhsGNNN4Ph1kMW+BjoYJM2jV5i4IlFoeZA9XI=" crossorigin=anonymous></script><script src=https://cdnjs.cloudflare.com/ajax/libs/mark.js/8.11.1/jquery.mark.min.js integrity="sha256-4HLtjeVgH0eIB3aZ9mLYF6E8oU5chNdjU6p6rrXpl9U=" crossorigin=anonymous></script><script src=/js/academic.min.4e51017b38c7ebaadd7e25fc9503f88c.js></script><div class=container><footer class=site-footer><p class=powered-by>© 2026 Vallery Lancey ·
|
|
Powered by the
|
|
<a href=https://sourcethemes.com/academic/ target=_blank rel=noopener>Academic theme</a> for
|
|
<a href=https://gohugo.io target=_blank rel=noopener>Hugo</a>.
|
|
<span class=float-right aria-hidden=true><a href=# id=back_to_top><span class=button_icon><i class="fas fa-chevron-up fa-2x"></i></span></a></span></p></footer></div><div id=modal class="modal fade" role=dialog><div class=modal-dialog><div class=modal-content><div class=modal-header><h5 class=modal-title>Cite</h5><button type=button class=close data-dismiss=modal aria-label=Close>
|
|
<span aria-hidden=true>×</span></button></div><div class=modal-body><pre><code class="tex hljs"></code></pre></div><div class=modal-footer><a class="btn btn-outline-primary my-1 js-copy-cite" href=# target=_blank><i class="fas fa-copy"></i>Copy</a>
|
|
<a class="btn btn-outline-primary my-1 js-download-cite" href=# target=_blank><i class="fas fa-download"></i>Download</a><div id=modal-error></div></div></div></div></div></body></html> |