4410 lines
1.7 MiB
4410 lines
1.7 MiB
<!doctype html><html lang=en><head><meta charset=UTF-8><meta name=viewport content="width=device-width,initial-scale=1"><title>Choosing Between Count and For-Each | Ned In The Cloud</title>
|
||
<meta name=description content="The world of technology is constantly shifting and evolving. Stay up to date on the latest concepts and conversations with these posts from Ned in the Cloud."><script defer src=/js/alpine.js></script><script src=/js/scrollreveal.js></script><link rel=stylesheet href="/css/style.min.2a57464b9db9ddd89351d40599b390a4de39492ed4687dc948a3a5e2eaa04d81.css" integrity="sha256-KldGS5253diTUdQFmbOQpN45SS7UaH3JSKOl4uqgTYE="><style>html.sr .load-hidden{visibility:hidden}</style><script>ScrollReveal({reset:!1})</script><link rel=stylesheet href=https://cdn.jsdelivr.net/npm/swiper@11/swiper-bundle.min.css><script src=https://cdn.jsdelivr.net/npm/swiper@11/swiper-bundle.min.js></script><script defer src=/fontawesome/js/brands.js></script><script defer src=/fontawesome/js/solid.js></script><script defer src=/fontawesome/js/regular.js></script><script defer src=/fontawesome/js/fontawesome.js></script><link rel=apple-touch-icon sizes=180x180 href=/apple-touch-icon.png><link rel=icon type=image/png sizes=32x32 href=/favicon-32x32.png><link rel=icon type=image/png sizes=16x16 href=/favicon-16x16.png><link rel=manifest href=/site.webmanifest><meta name=msapplication-TileColor content="#da532c"><meta name=theme-color content="#ffffff"><link href=https://nedinthecloud.com/blog/index.xml rel=alternate type=application/rss+xml title="Ned In The Cloud"><link href=https://nedinthecloud.com/blog/index.xml rel=feed type=application/rss+xml title="Ned In The Cloud"><meta property="og:title" content="Choosing Between Count and For-Each"><meta property="og:description" content="The world of technology is constantly shifting and evolving. Stay up to date on the latest concepts and conversations with these posts from Ned in the Cloud."><meta property="og:type" content="article"><meta property="og:url" content="https://nedinthecloud.com/2022/01/27/choosing-between-count-and-for-each/"><meta property="og:image" content="https://nedinthecloud.com/2022/01/27/choosing-between-count-and-for-each/Count-and-For-Each.png"><meta property="article:section" content="blog"><meta property="article:published_time" content="2022-01-27T00:00:00+00:00"><meta property="article:modified_time" content="2022-01-27T00:00:00+00:00"><meta property="og:site_name" content="Ned In The Cloud"><meta name=twitter:card content="summary_large_image"><meta name=twitter:image content="https://nedinthecloud.com/2022/01/27/choosing-between-count-and-for-each/Count-and-For-Each.png"><meta name=twitter:title content="Choosing Between Count and For-Each"><meta name=twitter:description content="The world of technology is constantly shifting and evolving. Stay up to date on the latest concepts and conversations with these posts from Ned in the Cloud."><script async src="https://www.googletagmanager.com/gtag/js?id=G-L0TLSTZELF"></script><script>window.dataLayer=window.dataLayer||[];function gtag(){dataLayer.push(arguments)}gtag("js",new Date),gtag("config","G-L0TLSTZELF")</script></head><body><header class="absolute top-0 z-50 w-full bg-transparent"><nav x-data=accordion(6) class="mx-auto flex w-full max-w-7xl flex-wrap items-center justify-between px-4 py-0 tracking-wide lg:px-6 lg:py-0" x-cloak><div class="w-48 self-center"><a href=/><img src=/images/logo.png class="my-2 mr-3 h-12 w-auto sm:h-20" width=146 height=48 alt="Ned In The Cloud Logo"></a></div><div @click=handleClick() x-data="{open : false}" class="mt-2 block cursor-pointer self-center text-white lg:hidden"><button @click="open = ! open" class="h-6 w-6 text-lg" aria-label="Toggle menu"><svg x-show="! open" viewBox="0 0 48 48" fill="none" xmlns="http://www.w3.org/2000/svg" :clas="{'transition-full each-in-out transform duration-500':! open}"><rect width="48" height="48" fill="#fff" fill-opacity=".01"/><path d="M7.94977 11.9498H39.9498" stroke="currentcolor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"/><path d="M7.94977 23.9498H39.9498" stroke="currentcolor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"/><path d="M7.94977 35.9498H39.9498" stroke="currentcolor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"/></svg><svg x-show="open" xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24" fill="none" stroke="currentcolor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-x"><line x1="18" y1="6" x2="6" y2="18"/><line x1="6" y1="6" x2="18" y2="18"/></svg></button></div><div x-ref=tab :style=handleToggle() class="relative z-40 max-h-0 w-full overflow-hidden rounded bg-[#185b7e] shadow-xl backdrop-blur-md transition-all duration-500 lg:hidden"><div class="my-6 flex flex-col gap-5 text-base text-gray-600"><a href=/ class="block w-full border-b-2 border-transparent pl-6 text-left text-base font-semibold uppercase leading-none tracking-wider text-white/95"><span>Home</span></a>
|
||
<a href=/blog/ class="block w-full border-b-2 border-transparent pl-6 text-left text-base font-semibold uppercase leading-none tracking-wider text-white/95"><span>Blog</span></a><div x-data="{ open: false }" class="relative inline-block w-full"><button @click="open = !open" class="flex items-center">
|
||
<span class="mr-3 w-full pl-6 text-base font-semibold uppercase leading-none tracking-wider text-white/95">Media
|
||
</span><span class="transform rounded bg-white/20 p-1 transition-transform duration-500"><svg class="h-4 w-4 fill-white" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 20 20"><path d="M9.293 12.95l.707.707L15.657 8l-1.414-1.414L10 10.828 5.757 6.586 4.343 8z"/></svg></span></button><div x-show=open x-on:click.away="open = false" x-transition:enter="transition ease-out duration-300" x-transition:enter-start="opacity-0 transform scale-90" x-transition:enter-end="opacity-100 transform scale-100" x-transition:leave="transition ease-in duration-300" x-transition:leave-start="opacity-100 transform scale-100" x-transition:leave-end="opacity-0 transform scale-90" class="absolute left-0 z-50 ml-4 mt-2 w-48 min-w-max overflow-hidden rounded bg-[#3f87b9]/60 text-white shadow-xl backdrop-blur-md"><a href=/podcast/ class="block px-4 py-3 text-base font-medium leading-none text-white/95 hover:bg-black/20 hover:text-white">Podcast</a>
|
||
<a href=/videos/ class="block px-4 py-3 text-base font-medium leading-none text-white/95 hover:bg-black/20 hover:text-white">Videos</a>
|
||
<a href=/books/ class="block px-4 py-3 text-base font-medium leading-none text-white/95 hover:bg-black/20 hover:text-white">Books</a>
|
||
<a href=/courses/ class="block px-4 py-3 text-base font-medium leading-none text-white/95 hover:bg-black/20 hover:text-white">Courses</a></div></div><a href=/sponsorship/ class="block w-full border-b-2 border-transparent pl-6 text-left text-base font-semibold uppercase leading-none tracking-wider text-white/95"><span>Sponsor</span></a>
|
||
<a href=/about/ class="block w-full border-b-2 border-transparent pl-6 text-left text-base font-semibold uppercase leading-none tracking-wider text-white/95"><span>About</span></a></div><div class="mx-auto mt-1 flex w-48 justify-center gap-10 py-6"><form id=search class="relative mx-auto w-full max-w-3xl" action=https://nedinthecloud.com/search/ method=get><label hidden for=search-input>Search site</label>
|
||
<input type=text id=search-input name=query placeholder=Search class="h-auto w-full cursor-pointer rounded-md px-5 py-2 text-sm ring-nblue/60 focus:ring-1" placeholder=Search></form></div></div><div class="hidden w-full lg:flex lg:w-auto lg:items-end"><div class="list-reset mt-0 flex-1 items-center justify-center gap-10 text-base text-gray-500 lg:flex lg:pt-0 xl:gap-6"><a class="focus:shadow-outline text-white rounded px-2 py-2 text-base font-semibold uppercase leading-none tracking-normal transition duration-200 ease-in hover:bg-white hover:text-nblue focus:text-opacity-80 focus:outline-none" href=/>Home </a><a class="focus:shadow-outline text-white rounded px-2 py-2 text-base font-semibold uppercase leading-none tracking-normal transition duration-200 ease-in hover:bg-white hover:text-nblue focus:text-opacity-80 focus:outline-none" href=/blog/>Blog</a><div x-data="{ open: false }" @mouseleave="open = false" class="relative inline-block" :class="{'text-white': open, 'text-white': !open }"><span class="flex items-center" @mouseover="open = true"><span class="text-white mr-2 cursor-pointer rounded px-2 py-2 text-base font-semibold uppercase leading-none tracking-wider transition duration-200 ease-in hover:bg-white hover:text-nblue focus:text-opacity-80 focus:outline-none">Media</span>
|
||
<span @click="open = !open" :class="open ? '-rotate-180' : ''" class="transform transition-transform duration-500"><svg class="h-4 w-4 fill-white" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 20 20"><path d="M9.293 12.95l.707.707L15.657 8l-1.414-1.414L10 10.828 5.757 6.586 4.343 8z"/></svg></span></span><div x-cloak x-show=open x-transition:enter="transition ease-out duration-300" x-transition:enter-start="opacity-0 transform scale-90" x-transition:enter-end="opacity-100 transform scale-100" x-transition:leave="transition ease-in duration-300" x-transition:leave-start="opacity-100 transform scale-100" x-transition:leave-end="opacity-0 transform scale-90" class="absolute left-0 w-32 min-w-max divide-y divide-white/10 rounded-md bg-[#2576b8]/60 text-white shadow-xl backdrop-blur-md"><a href=/podcast/ class="block px-4 py-3 text-base leading-none text-white hover:bg-black/10 hover:text-white">Podcast</a>
|
||
<a href=/videos/ class="block px-4 py-3 text-base leading-none text-white hover:bg-black/10 hover:text-white">Videos</a>
|
||
<a href=/books/ class="block px-4 py-3 text-base leading-none text-white hover:bg-black/10 hover:text-white">Books</a>
|
||
<a href=/courses/ class="block px-4 py-3 text-base leading-none text-white hover:bg-black/10 hover:text-white">Courses</a></div></div><a class="focus:shadow-outline text-white rounded px-2 py-2 text-base font-semibold uppercase leading-none tracking-normal transition duration-200 ease-in hover:bg-white hover:text-nblue focus:text-opacity-80 focus:outline-none" href=/sponsorship/>Sponsor </a><a class="focus:shadow-outline text-white rounded px-2 py-2 text-base font-semibold uppercase leading-none tracking-normal transition duration-200 ease-in hover:bg-white hover:text-nblue focus:text-opacity-80 focus:outline-none" href=/about/>About</a></div></div><div class="hidden w-48 gap-6 xl:flex xl:justify-end xl:self-center"><form id=search class="relative mx-auto w-full max-w-3xl" action=https://nedinthecloud.com/search/ method=get><label hidden for=search-input>Search site</label>
|
||
<input type=text id=search-input name=query placeholder=Search class="h-auto w-full cursor-pointer rounded-md px-5 py-2 text-sm ring-nblue/60 focus:ring-1" placeholder=Search></form></div></nav><script>document.addEventListener("alpine:init",()=>{Alpine.store("accordion",{tab:0,dropdownOpen:!1,overflowState:"hidden"}),Alpine.data("accordion",e=>({init(){this.idx=e},idx:-1,handleClick(){this.$store.accordion.dropdownOpen=!this.$store.accordion.dropdownOpen,this.$store.accordion.dropdownOpen?setTimeout(()=>{this.$store.accordion.overflowState="visible"},700):this.$store.accordion.overflowState="hidden",this.$store.accordion.tab=this.$store.accordion.tab===this.idx?0:this.idx},handleToggle(){const e=this.$store.accordion.tab===this.idx?`${this.$refs.tab.scrollHeight}px`:"";return`overflow: ${this.$store.accordion.overflowState}; max-height: ${e}`}}))})</script></header><div><div class="relative flex h-full min-h-[50vh] items-center justify-center sm:aspect-auto sm:h-[60vh]"><picture class="absolute inset-0 z-10 h-full w-full"><source type=image/webp srcset="/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_320x0_resize_q75_h2_box.webp 320w,/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_640x0_resize_q75_h2_box.webp 640w,/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_960x0_resize_q75_h2_box.webp 960w,/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_1280x0_resize_q75_h2_box.webp 1280w,/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_1600x0_resize_q75_h2_box.webp 1600w,/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_1920x0_resize_q75_h2_box.webp 1920w"><source type=image/jpeg srcset="/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_320x0_resize_q75_box.jpg 320w,/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_640x0_resize_q75_box.jpg 640w,/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_960x0_resize_q75_box.jpg 960w,/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_1280x0_resize_q75_box.jpg 1280w,/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_1600x0_resize_q75_box.jpg 1600w,/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_1920x0_resize_q75_box.jpg 1920w"><img fetchpriority=high class="z-10 h-full min-h-[50vh] w-full object-cover object-center sm:aspect-auto sm:min-h-[60vh]" src=/images/cloud_hubf6df6462c4cf3c04ef00fd0059cc0df_808860_1280x0_resize_q75_box.jpg width=1280 height=711 alt=slide></picture><div class="absolute inset-0 z-20 bg-black/30"><div class="mx-auto flex h-full w-full max-w-7xl flex-col items-center justify-center px-6 pt-8 lg:items-start lg:pt-16"><h1 class="load-hidden fade-in z-30 text-center text-4xl font-bold capitalize leading-tight text-white lg:text-left lg:text-6xl lg:leading-tight">Choosing Between Count and For-Each</h1><h2 class="load-hidden seq fade-in z-30 mt-4 text-left text-2xl font-medium leading-snug text-white lg:text-3xl lg:leading-tight"><span class=font-normal>Ned Bellavance</span>
|
||
<span class="hidden lg:inline" aria-hidden=true>·</span><br class="inline lg:hidden"><time class=font-light datetime="2022-01-27 00:00:00 +0000 UTC">January 27, 2022</time>
|
||
<span aria-hidden=true>·</span>
|
||
<span class=font-light>9 min read</span></h2></div></div></div></div><section class="bg-white py-12 sm:py-16 lg:py-20"><div class="mx-auto max-w-7xl px-4 sm:px-6 lg:px-8"><div class="grid grid-cols-1 gap-y-8 lg:grid-cols-6 lg:gap-x-12 xl:gap-x-20"><div class=lg:col-span-4><div class="flex w-full justify-center"><img src=Count-and-For-Each.png alt=Cover class=rounded-2xl></div><div class="mb-4 mt-12 flex items-center justify-center"><div class="flex flex-wrap gap-1"><a href=/tags/hashicorp-terraform class="z-20 inline-block rounded-full bg-slate-100 px-2 py-1 text-xs font-medium uppercase tracking-wider text-gray-900 no-underline hover:bg-slate-200 lg:px-3 lg:py-2">hashicorp-terraform</a>
|
||
<a href=/tags/hashicorp-terraform-tutorial class="z-20 inline-block rounded-full bg-slate-100 px-2 py-1 text-xs font-medium uppercase tracking-wider text-gray-900 no-underline hover:bg-slate-200 lg:px-3 lg:py-2">hashicorp-terraform-tutorial</a></div></div><div class="prose mx-auto mt-10 sm:prose-lg hover:prose-a:text-nblue lg:mt-14"><p>Terraform has two looping mechanisms for creating multiple resources, <code>count</code> and <code>for_each</code>. The <code>count</code> meta-argument has been around for a long time, but <code>for_each</code> is a relative newcomer (introduced in version 0.12). Each meta-argument allows you to create more than one resource or module with a single configuration block.</p><iframe width=560 height=315 src=https://www.youtube.com/embed/y5u0nixenuk title="YouTube video player" frameborder=0 allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen></iframe><p>A common question is when to use <code>count</code> versus <code>for_each</code>. I would make the argument that <code>for_each</code> is almost always preferred, and in this post I hope to show you why.</p><h2 id=looping-basics>Looping Basics</h2><p>Before I explain why <code>for_each</code> is generally superior, it would be useful to understand what is actually happening when you add a looping meta-argument to a Terraform configuration. Let’s start with a simple example.</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#66d9ef>resource</span> <span style=color:#e6db74>"local_file"</span> <span style=color:#e6db74>"count_int_loop"</span> {
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>count</span> = <span style=color:#ae81ff>3</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>content</span> = <span style=color:#e6db74>"This is file number </span><span style=color:#e6db74>${</span>count.<span style=color:#a6e22e>index</span><span style=color:#e6db74>}</span><span style=color:#e6db74>"</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>filename</span> = <span style=color:#e6db74>"</span><span style=color:#e6db74>${</span><span style=color:#a6e22e>path</span>.module<span style=color:#e6db74>}</span><span style=color:#e6db74>/int-</span><span style=color:#e6db74>${</span>count.<span style=color:#a6e22e>index</span><span style=color:#e6db74>}</span><span style=color:#e6db74>.count"</span>
|
||
</span></span><span style=display:flex><span>}
|
||
</span></span></code></pre></div><p>Running <code>terraform apply</code> will generate three files: int-1.count, int-2.count, int-3.count. If we take a look at the state data:</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-bash data-lang=bash><span style=display:flex><span>$> terraform state list
|
||
</span></span><span style=display:flex><span>
|
||
</span></span><span style=display:flex><span>local_file.count_int_loop<span style=color:#f92672>[</span>0<span style=color:#f92672>]</span>
|
||
</span></span><span style=display:flex><span>local_file.count_int_loop<span style=color:#f92672>[</span>1<span style=color:#f92672>]</span>
|
||
</span></span><span style=display:flex><span>local_file.count_int_loop<span style=color:#f92672>[</span>2<span style=color:#f92672>]</span>
|
||
</span></span></code></pre></div><p>We have three resources created by the <code>count</code> meta-argument with an integer based index. The object <code>local_file.count_int_loop</code> is an ordered list of <code>local_file</code> resource objects.</p><p>If we want four files instead of three, all we have to do is increase the count value by one and Terraform will create a fourth file. This is what <code>count</code> was meant for, undifferentiated resource creation based on an integer.</p><p>But what if instead of an integer, we were dealing with a list of items?</p><h3 id=parsing-a-list>Parsing a List</h3><p>Before the introduction of <code>for_each</code>, the <code>count</code> argument was all we had. <em>(And we liked it.)</em> You could use a list to create mulitple resources by finding the number of elements in the list and using that for the <code>count</code> value. The <code>length()</code> function does an admirable job of accomplishing this.</p><p>Consider the following configuration with a list of toppings defined as a local value.</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#a6e22e>locals</span> {
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>toppings</span> = [<span style=color:#e6db74>"lettuce"</span>,<span style=color:#e6db74>"tomatoes"</span>,<span style=color:#e6db74>"jalapenos"</span>]
|
||
</span></span><span style=display:flex><span>}
|
||
</span></span><span style=display:flex><span><span style=color:#66d9ef>
|
||
</span></span></span><span style=display:flex><span><span style=color:#66d9ef>resource</span> <span style=color:#e6db74>"local_file"</span> <span style=color:#e6db74>"count_loop"</span> {
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>count</span> = length(<span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>toppings</span>)
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>content</span> = <span style=color:#e6db74>"</span><span style=color:#e6db74>${</span><span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>toppings</span>[count.<span style=color:#a6e22e>index</span>]<span style=color:#e6db74>}</span><span style=color:#e6db74>"</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>filename</span> = <span style=color:#e6db74>"</span><span style=color:#e6db74>${</span><span style=color:#a6e22e>path</span>.module<span style=color:#e6db74>}</span><span style=color:#e6db74>/</span><span style=color:#e6db74>${</span><span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>toppings</span>[count.<span style=color:#a6e22e>index</span>]<span style=color:#e6db74>}</span><span style=color:#e6db74>.count"</span>
|
||
</span></span><span style=display:flex><span>}
|
||
</span></span></code></pre></div><p>Running <code>terraform apply</code> will generate three files: lettuce.count, tomatoes.count, and jalapenos.count. If we take a look at the state:</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-bash data-lang=bash><span style=display:flex><span>$> terraform state list
|
||
</span></span><span style=display:flex><span>
|
||
</span></span><span style=display:flex><span>local_file.count_loop<span style=color:#f92672>[</span>0<span style=color:#f92672>]</span>
|
||
</span></span><span style=display:flex><span>local_file.count_loop<span style=color:#f92672>[</span>1<span style=color:#f92672>]</span>
|
||
</span></span><span style=display:flex><span>local_file.count_loop<span style=color:#f92672>[</span>2<span style=color:#f92672>]</span>
|
||
</span></span></code></pre></div><p>Once again, we have three resources created by the <code>count</code> meta-argument with a number based index. Just like before, the object <code>local_file.count_loop</code> is an ordered list of <code>local_file</code> resources.</p><p>So far, so good, right? What’s the point of a <code>for_each</code> arguemnt if we can simply use the <code>count</code> argument with a <code>length</code> function? Let’s try the same thing with <code>for_each</code> instead:</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#66d9ef>resource</span> <span style=color:#e6db74>"local_file"</span> <span style=color:#e6db74>"for_each_loop"</span> {
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>for_each</span> = toset(<span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>toppings</span>)
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>content</span> = <span style=color:#e6db74>"</span><span style=color:#e6db74>${</span>each.<span style=color:#a6e22e>value</span><span style=color:#e6db74>}</span><span style=color:#e6db74>"</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>filename</span> = <span style=color:#e6db74>"</span><span style=color:#e6db74>${</span><span style=color:#a6e22e>path</span>.module<span style=color:#e6db74>}</span><span style=color:#e6db74>/</span><span style=color:#e6db74>${</span>each.<span style=color:#a6e22e>value</span><span style=color:#e6db74>}</span><span style=color:#e6db74>.foreach"</span>
|
||
</span></span><span style=display:flex><span>}
|
||
</span></span></code></pre></div><p>Looking at our state now:</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-bash data-lang=bash><span style=display:flex><span>$> terraform state list
|
||
</span></span><span style=display:flex><span>
|
||
</span></span><span style=display:flex><span>local_file.for_each_loop<span style=color:#f92672>[</span><span style=color:#e6db74>"jalapenos"</span><span style=color:#f92672>]</span>
|
||
</span></span><span style=display:flex><span>local_file.for_each_loop<span style=color:#f92672>[</span><span style=color:#e6db74>"lettuce"</span><span style=color:#f92672>]</span>
|
||
</span></span><span style=display:flex><span>local_file.for_each_loop<span style=color:#f92672>[</span><span style=color:#e6db74>"tomatoes"</span><span style=color:#f92672>]</span>
|
||
</span></span></code></pre></div><p>We have three resources created by the <code>for_each</code> meta-argument with a key based reference. The object <code>local_file.for_each_loop</code> is a <strong>map</strong> (aka hashtable). The keys will be the strings in the set, or if you submit a map it will be the keys of the map. The map values are the <code>local_file</code> resources.</p><p>From a practical standpoint, we have essential generated three files with the same content. Either argument seems to do the trick, so why would you prefer one over the other? Two reasons: consistency and referencing.</p><h2 id=consistency>Consistency</h2><p>Something that’s not immediately obvious is how much the order of the items used by <code>count</code> matters. Let’s say we want to add a topping to our list. How about some onions?</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#a6e22e>locals</span> {
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>toppings</span> = [<span style=color:#e6db74>"lettuce"</span>,<span style=color:#e6db74>"tomatoes"</span>,<span style=color:#e6db74>"onions"</span>,<span style=color:#e6db74>"jalapenos"</span>]
|
||
</span></span><span style=display:flex><span>}
|
||
</span></span></code></pre></div><p>What do you think will happen when we run <code>terraform plan</code> against the <code>count</code> example? Notice the <strong>order</strong> of the toppings. Onions is now in index <strong>2</strong> and jalapenos is in index <strong>3</strong>.</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-bash data-lang=bash><span style=display:flex><span>$> terraform plan
|
||
</span></span><span style=display:flex><span>
|
||
</span></span><span style=display:flex><span>Terraform will perform the following actions:
|
||
</span></span><span style=display:flex><span>
|
||
</span></span><span style=display:flex><span> <span style=color:#75715e># local_file.count_loop[2] must be replaced</span>
|
||
</span></span><span style=display:flex><span>-/+ resource <span style=color:#e6db74>"local_file"</span> <span style=color:#e6db74>"count_loop"</span> <span style=color:#f92672>{</span>
|
||
</span></span><span style=display:flex><span> ~ content <span style=color:#f92672>=</span> <span style=color:#e6db74>"jalapenos"</span> -> <span style=color:#e6db74>"onions"</span> <span style=color:#75715e># forces replacement</span>
|
||
</span></span><span style=display:flex><span> ~ filename <span style=color:#f92672>=</span> <span style=color:#e6db74>"./jalapenos.count"</span> -> <span style=color:#e6db74>"./onions.count"</span> <span style=color:#75715e># forces replacement</span>
|
||
</span></span><span style=display:flex><span> ~ id <span style=color:#f92672>=</span> <span style=color:#e6db74>"626451b23e9097d6a2c081959703df63424602bf"</span> -> <span style=color:#f92672>(</span>known after apply<span style=color:#f92672>)</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#75715e># (2 unchanged attributes hidden)</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#f92672>}</span>
|
||
</span></span><span style=display:flex><span>
|
||
</span></span><span style=display:flex><span> <span style=color:#75715e># local_file.count_loop[3] will be created</span>
|
||
</span></span><span style=display:flex><span> + resource <span style=color:#e6db74>"local_file"</span> <span style=color:#e6db74>"count_loop"</span> <span style=color:#f92672>{</span>
|
||
</span></span><span style=display:flex><span> + content <span style=color:#f92672>=</span> <span style=color:#e6db74>"jalapenos"</span>
|
||
</span></span><span style=display:flex><span> + directory_permission <span style=color:#f92672>=</span> <span style=color:#e6db74>"0777"</span>
|
||
</span></span><span style=display:flex><span> + file_permission <span style=color:#f92672>=</span> <span style=color:#e6db74>"0777"</span>
|
||
</span></span><span style=display:flex><span> + filename <span style=color:#f92672>=</span> <span style=color:#e6db74>"./jalapenos.count"</span>
|
||
</span></span><span style=display:flex><span> + id <span style=color:#f92672>=</span> <span style=color:#f92672>(</span>known after apply<span style=color:#f92672>)</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#f92672>}</span>
|
||
</span></span><span style=display:flex><span>
|
||
</span></span><span style=display:flex><span> <span style=color:#75715e># local_file.for_each_loop["onions"] will be created</span>
|
||
</span></span><span style=display:flex><span> + resource <span style=color:#e6db74>"local_file"</span> <span style=color:#e6db74>"for_each_loop"</span> <span style=color:#f92672>{</span>
|
||
</span></span><span style=display:flex><span> + content <span style=color:#f92672>=</span> <span style=color:#e6db74>"onions"</span>
|
||
</span></span><span style=display:flex><span> + directory_permission <span style=color:#f92672>=</span> <span style=color:#e6db74>"0777"</span>
|
||
</span></span><span style=display:flex><span> + file_permission <span style=color:#f92672>=</span> <span style=color:#e6db74>"0777"</span>
|
||
</span></span><span style=display:flex><span> + filename <span style=color:#f92672>=</span> <span style=color:#e6db74>"./onions.foreach"</span>
|
||
</span></span><span style=display:flex><span> + id <span style=color:#f92672>=</span> <span style=color:#f92672>(</span>known after apply<span style=color:#f92672>)</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#f92672>}</span>
|
||
</span></span><span style=display:flex><span>
|
||
</span></span><span style=display:flex><span>Plan: <span style=color:#ae81ff>3</span> to add, <span style=color:#ae81ff>0</span> to change, <span style=color:#ae81ff>1</span> to destroy.
|
||
</span></span></code></pre></div><p><em>Terraform is going to destroy our jalapenos!</em> And that is because when Terraform runs through the <code>count</code> loop, it sees the value <em>onions</em> in index 2 and that value used to be <em>jalapenos</em>. Terraform has to destroy the original <code>local_file.count_loop[2]</code> resource and replace it with the new value. Then it will create a new resource called <code>local_file.count_loop[3]</code> using the <em>jalapenos</em> value.</p><p>The <code>for_each</code> loop doesn’t have this problem. Since it is using a key based reference, it doesn’t care about order. In fact, the <code>for_each</code> argument requires a set or map as input, neither of which are ordered. Terraform simply sees the key <em>onion</em> without a corresponding entry in <code>local_file.for_each_loop</code> and decides to create one. The jalapenos file is not touched. Thank goodness! I like my jalapenos.🌶️</p><p>Count is going to unnecessarily destroy and recreate the jalapenos file, which might not be a problem for a text file. But imagine that’s a Kubernetes cluster running 100+ applications, and you just destroyed it because you added a new cluster in the wrong order. That’s… bad. Possibly a resume generating event.</p><p>Of course, that would never happen becuase you run <code>terraform plan</code> first, right? Right???</p><h2 id=referencing>Referencing</h2><p>As pointed out by the previous section, using <code>count</code> results in an ordered list and <code>for_each</code> results in a map. When you need to reference the resources somewhere else in your configuration, you might find that being able to refer to a resource by key instead of index is much easier. Let’s look at a slightly more advanced example where we are trying to create users and groups in Terraform Cloud.</p><p>Users are created with the resource type <code>tfe_organization_membership</code>. I could create users with a <code>count</code> like this:</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#66d9ef>resource</span> <span style=color:#e6db74>"tfe_organization_membership"</span> <span style=color:#e6db74>"org_members"</span> {
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>count</span> = length(<span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>users</span>)
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>organization</span> = <span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>organization_name</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>email</span> = <span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>users</span>[count.<span style=color:#a6e22e>index</span>]
|
||
</span></span><span style=display:flex><span>}
|
||
</span></span></code></pre></div><p>Or with a <code>for_each</code> like this:</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#66d9ef>resource</span> <span style=color:#e6db74>"tfe_organization_membership"</span> <span style=color:#e6db74>"org_members"</span> {
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>for_each</span> = toset(<span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>users</span>)
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>organization</span> = <span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>organization_name</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>email</span> = each.<span style=color:#a6e22e>value</span>
|
||
</span></span><span style=display:flex><span>}
|
||
</span></span></code></pre></div><p>To add a user to a team on Terraform Cloud, the resource type <code>tfe_team_organization_member</code> is used. The two arguments <code>team_id</code> and <code>organization_membership_id</code> both require a value that is an attribute of the previously generated team or user. It is a value we must look up using a reference. If we’re using the <code>for_each</code> loop to create users, the reference for <code>organization_membership_id</code> looks like this:</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#a6e22e>organization_membership_id</span> = <span style=color:#a6e22e>tfe_organization_membership</span>.<span style=color:#a6e22e>org_members</span>[each.<span style=color:#a6e22e>value</span>[<span style=color:#e6db74>"member_name"</span>]].<span style=color:#a6e22e>id</span>
|
||
</span></span></code></pre></div><p>We are using the key <code>member_name</code> to find the correct instance of <code>tfe_organization_membership</code> and returning the <code>id</code> attribute of that instance. While the code looks a little confusing at first, trust me it works. The full block is shown below.</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#66d9ef>resource</span> <span style=color:#e6db74>"tfe_team_organization_member"</span> <span style=color:#e6db74>"team_members"</span> {
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>for_each</span> = { <span style=color:#66d9ef>for</span> <span style=color:#a6e22e>member</span> <span style=color:#66d9ef>in</span> <span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>team_members</span> <span style=color:#f92672>:</span> <span style=color:#e6db74>"</span><span style=color:#e6db74>${</span><span style=color:#a6e22e>member</span>.<span style=color:#a6e22e>team_name</span><span style=color:#e6db74>}</span><span style=color:#e6db74>_</span><span style=color:#e6db74>${</span><span style=color:#a6e22e>member</span>.<span style=color:#a6e22e>member_name</span><span style=color:#e6db74>}</span><span style=color:#e6db74>"</span> => <span style=color:#a6e22e>member</span> }
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>team_id</span> = <span style=color:#a6e22e>tfe_team</span>.<span style=color:#a6e22e>teams</span>[each.<span style=color:#a6e22e>value</span>[<span style=color:#e6db74>"team_name"</span>]].<span style=color:#a6e22e>id</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>organization_membership_id</span> = <span style=color:#a6e22e>tfe_organization_membership</span>.<span style=color:#a6e22e>org_members</span>[each.<span style=color:#a6e22e>value</span>[<span style=color:#e6db74>"member_name"</span>]].<span style=color:#a6e22e>id</span>
|
||
</span></span><span style=display:flex><span>}
|
||
</span></span></code></pre></div><p>If we had used a <code>count</code> argument to create the users, we would need to use a <code>for</code> expression with a filter to look up the membership id value. Something along the lines of this:</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#a6e22e>organization_membership_id</span> = ([<span style=color:#66d9ef>for</span> <span style=color:#a6e22e>user</span> <span style=color:#66d9ef>in</span> <span style=color:#a6e22e>tfe_organization_membership</span>.<span style=color:#a6e22e>org_members</span> <span style=color:#f92672>:</span> <span style=color:#a6e22e>user</span>.<span style=color:#a6e22e>id</span> <span style=color:#a6e22e>if</span> <span style=color:#a6e22e>user</span>.<span style=color:#a6e22e>email</span> =<span style=color:#f92672>=</span> each.<span style=color:#a6e22e>value</span>[<span style=color:#e6db74>"member_name"</span>]])[<span style=color:#ae81ff>0</span>]
|
||
</span></span></code></pre></div><p>Would it work? Sure. Is it efficient? Nope. The <code>for</code> expression has to loop through all the <code>tfe_organization_membership</code> resources to find the one that matches. No big deal if we have three users. Pretty big deal if we have three thousand. Reference by key is one of the best things about maps/hashtables/dictionaries, whatever you want to call them.</p><h2 id=when-to-use-count>When to use count</h2><p>You can use <code>count</code> if you don’t care about uniqueness or references in your configuration. If every item created by a loop is ephemeral and functionally identical, then there’s probably no benefit to using <code>for_each</code>. If you don’t need to refer to anything by the key, then using <code>count</code> could be fine. On the other hand, it’s about the same amount of work to use either, and <code>for_each</code> has some serious benefits.</p><h3 id=what-about-conditionals>What about conditionals?</h3><p>One use that still seems relevant is using count with a zero value to make the creation of a resource optional. What do I mean? Consider this:</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#66d9ef>resource</span> <span style=color:#e6db74>"local_file"</span> <span style=color:#e6db74>"count_optional"</span> {
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>count</span> = <span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>create_file</span> <span style=color:#f92672>?</span> <span style=color:#ae81ff>1</span> <span style=color:#f92672>:</span> <span style=color:#ae81ff>0</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>content</span> = <span style=color:#e6db74>"Hello!"</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>filename</span> = <span style=color:#e6db74>"</span><span style=color:#e6db74>${</span><span style=color:#a6e22e>path</span>.module<span style=color:#e6db74>}</span><span style=color:#e6db74>/count-create.txt"</span>
|
||
</span></span><span style=display:flex><span>}
|
||
</span></span></code></pre></div><p>The creation of a new organization will happen if the variable <code>create_new_organization</code> is set to <code>true</code> and not if its set to <code>false</code>. You’re only ever creating one or zero of an item. Is there a way to replace this with <code>for_each</code>? If so, is there any benefit?</p><p>To answer the first question. Yes, you can do it.</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#66d9ef>resource</span> <span style=color:#e6db74>"local_file"</span> <span style=color:#e6db74>"for_each_optional"</span> {
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>for_each</span> = <span style=color:#a6e22e>local</span>.<span style=color:#a6e22e>create_file</span> <span style=color:#f92672>?</span> toset([<span style=color:#e6db74>"any_value"</span>]) <span style=color:#f92672>:</span> toset([])
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>content</span> = <span style=color:#e6db74>"Hello!"</span>
|
||
</span></span><span style=display:flex><span> <span style=color:#a6e22e>filename</span> = <span style=color:#e6db74>"</span><span style=color:#e6db74>${</span><span style=color:#a6e22e>path</span>.module<span style=color:#e6db74>}</span><span style=color:#e6db74>/for-each-create.txt"</span>
|
||
</span></span><span style=display:flex><span>}
|
||
</span></span></code></pre></div><p>Is there any benefit? Not that I can think of. The <code>count</code> version feels more intuitive, but I don’t think it would be any more effective than the <code>for_each</code> loop.</p><h3 id=what-about-using-the-index-value>What about using the index value?</h3><p>I started out the looping overview by using an integer with <code>count</code>. This is the one time when <code>count</code> has a clear advantage. The count argument takes a number and counts up to that number. You can access the current iteration using the <code>count.index</code> expression. <code>Count</code> makes more sense if I am creating resources based off a number, instead of set, list, map, or other object.</p><p>Could you replace it with a <code>for_each</code> argument? Would there be any benefit?</p><p>To answer the first question, yes you can do it.</p><div class=highlight><pre tabindex=0 style=color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4><code class=language-terraform data-lang=terraform><span style=display:flex><span><span style=color:#a6e22e>for_each</span> = toset(<span style=color:#a6e22e>range</span>(<span style=color:#ae81ff>3</span>))
|
||
</span></span></code></pre></div><p>The <code>range</code> function creates a list of integers starting with 0 and going to the max, non-inclusive. The <code>toset()</code> function turns that list of integers into a set. The <code>each.value</code> value will be the integer of the current iteration, so it’s basically the same as <code>index.count</code>.</p><p>Is there any benefit? I’d have to say no. The syntax is clunky and requires the execution of two functions to get it working. Contrasted to the <code>count</code> argument, there is no tangible benefit of using <code>for_each</code> in this situation.</p><h2 id=conclusion>Conclusion</h2><p>To sum up, here’s the general advice. If you are creating multiple resources based off an integer, and the resources are undifferentiated, then <code>count</code> works just fine. Any time you are using an input that is more complex than an integer, the proper answer is going to be <code>for_each</code>. Conditional resource creation is a bit of a toss up, so just do what feels intuitive to you or aligns with your team’s conventions.</p></div></div><div class=lg:col-span-2><h2 class="text-2xl font-semibold lg:text-3xl">Latest<span class=text-nblue>Articles</span></h2><div class="relative mt-1"><hr class="border-t-2 border-gray-100"><hr class="absolute inset-y-0 left-0 w-10 border-t-2 border-nblue"></div><div class="mt-6 space-y-6"><div class="relative flex items-start"><div class=flex-1><p class="text-sm font-bold text-gray-900"><a href=/2026/08/29/prepping-my-ai-pc/ title="Prepping my AI PC">Prepping my AI PC
|
||
<span class="absolute inset-0" aria-hidden=true></span></a></p><p class="mt-3 text-sm font-medium text-gray-500">August 29, 2026</p></div><picture><source srcset=/2026/08/29/prepping-my-ai-pc/PrepAIPC_hud2fd8ac4c9ddca2dcf6b7ac37f3be075_1246435_250x0_resize_q75_h2_box_3.webp type=image/webp><img class="ml-5 aspect-video w-32 shrink-0 rounded object-cover" src=/2026/08/29/prepping-my-ai-pc/PrepAIPC_hud2fd8ac4c9ddca2dcf6b7ac37f3be075_1246435_250x0_resize_box_3.png width=250 height=131 alt="Choosing Between Count and For-Each" loading=lazy></picture></div><div class="relative flex items-start"><div class=flex-1><p class="text-sm font-bold text-gray-900"><a href=/2026/05/23/trail-running-stuff/ title="Trail Running Stuff">Trail Running Stuff
|
||
<span class="absolute inset-0" aria-hidden=true></span></a></p><p class="mt-3 text-sm font-medium text-gray-500">May 23, 2026</p></div><picture><source srcset=/2026/05/23/trail-running-stuff/TrailRunning_hu6cd9071a31bbdb54760d0db5f5c81617_1085853_250x0_resize_q75_h2_box_3.webp type=image/webp><img class="ml-5 aspect-video w-32 shrink-0 rounded object-cover" src=/2026/05/23/trail-running-stuff/TrailRunning_hu6cd9071a31bbdb54760d0db5f5c81617_1085853_250x0_resize_box_3.png width=250 height=131 alt="Choosing Between Count and For-Each" loading=lazy></picture></div><div class="relative flex items-start"><div class=flex-1><p class="text-sm font-bold text-gray-900"><a href=/2026/05/11/developers-are-people-the-psychology-of-software-teams/ title="Developers Are People: The Psychology of Software Teams">Developers Are People: The Psychology of Software Teams
|
||
<span class="absolute inset-0" aria-hidden=true></span></a></p><p class="mt-3 text-sm font-medium text-gray-500">May 11, 2026</p></div><picture><source srcset=/2026/05/11/developers-are-people-the-psychology-of-software-teams/DayTwoDevOps-CatHicks_huf577f96777049894eadc12689172fa4b_740399_250x0_resize_q75_h2_box_3.webp type=image/webp><img class="ml-5 aspect-video w-32 shrink-0 rounded object-cover" src=/2026/05/11/developers-are-people-the-psychology-of-software-teams/DayTwoDevOps-CatHicks_huf577f96777049894eadc12689172fa4b_740399_250x0_resize_box_3.png width=250 height=141 alt="Choosing Between Count and For-Each" loading=lazy></picture></div><div class="relative flex items-start"><div class=flex-1><p class="text-sm font-bold text-gray-900"><a href=/2026/05/04/jet-plane-blues/ title="Jet Plane Blues">Jet Plane Blues
|
||
<span class="absolute inset-0" aria-hidden=true></span></a></p><p class="mt-3 text-sm font-medium text-gray-500">May 4, 2026</p></div><picture><source srcset=/2026/05/04/jet-plane-blues/JetPlaneBlues_hu6a4d5a77639c606fb963e798eccd5793_939384_250x0_resize_q75_h2_box_3.webp type=image/webp><img class="ml-5 aspect-video w-32 shrink-0 rounded object-cover" src=/2026/05/04/jet-plane-blues/JetPlaneBlues_hu6a4d5a77639c606fb963e798eccd5793_939384_250x0_resize_box_3.png width=250 height=131 alt="Choosing Between Count and For-Each" loading=lazy></picture></div></div></div></div></div></section><footer class=bg-zinc-800><div class="mx-auto max-w-7xl px-6 py-6 md:flex md:items-center md:justify-between lg:px-8"><div class="flex justify-center space-x-6 md:order-2"><a href=https://twitter.com/ned1313/ class="text-white hover:opacity-80" rel="noopener noreferrer" target=_blank><span class=sr-only>Twitter</span>
|
||
<i class="fa-brands fa-x-twitter text-2xl"></i>
|
||
</a><a href=https://www.linkedin.com/in/ned-bellavance/ class="text-white hover:opacity-80" rel="noopener noreferrer" target=_blank><span class=sr-only>LinkedIn</span>
|
||
<i class="fa-brands fa-linkedin-in text-2xl"></i>
|
||
</a><a href=https://www.youtube.com/c/NedintheCloud/ class="text-white hover:opacity-80" rel="noopener noreferrer" target=_blank><span class=sr-only>YouTube</span>
|
||
<i class="fa-brands fa-youtube text-2xl"></i>
|
||
</a><a href=/blog/index.xml class="text-white hover:opacity-80" rel="noopener noreferrer" target=_blank><span class=sr-only>RSS</span>
|
||
<i class="fa-solid fa-rss text-2xl"></i></a></div><div class="mt-4 md:order-1 md:mt-0"><p class="text-center text-sm leading-5 text-white md:text-base">© 2026 Ned In The Cloud LLC. All rights reserved.</p></div></div><script>window.store={"https://nedinthecloud.com/blog/":{title:"Blog",tags:[],content:"",summary:"",date:"29 Aug, 2026",url:"https://nedinthecloud.com/blog/",image:"",readingTime:"0"},"https://nedinthecloud.com/2026/08/29/prepping-my-ai-pc/":{title:"Prepping my AI PC",tags:["IntelArc","AMDRadeon","Nvidia","AI"],content:`I’ve been building an AI PC for myself as part of the preparation for a series of courses on Pluralsight focusing on running Small Language Models. Local AI is an area that I’ve wanted to dig into for a while, and building a course for others is a great way to force myself to learn. It’s also a fantastic excuse to spend money on hardware - like I needed one.
|
||
This post is from my notes taken while standing up the assembled AI PC with Ubuntu 26.04 LTS. This is not necessarily the “best” way to do things, but I wanted to document the process in case I needed to rebuild. Maybe you’ll find it interesting too.
|
||
Hardware overview It would be helpful to explain what hardware I’m supporting in the AI PC. The motherboard is an MSI X870E WiFi mega extreme, something, something, something. Honestly, these names are obnoxious and I can’t even pretend to memorize them. It’s an AMD system board with a Ryzen 9 9900X CPU. I dropped in 64GB of RAM, which I may someday bump, but I doubt it. System RAM is not all that interesting for local AI workloads. The main job of the system RAM is to load the model weights and then copy them over to the GPU. As long as I can fit the weights, I’m good.
|
||
For the GPUs, I’ve got three, but I can only have two in at a time:
|
||
Nvidia GeForce RTX 3090 Intel Arc B70 AMD Radeon R9700 The system board has three PCIe5 slots, but the RTX 3090 is freakin’ huge. The AMD and Intel cards take up two slots on the back, while the 3090 takes up at least three. So I can have the 3090 in the top PCIe slot, and the Intel or AMD card in the bottom slot. Basically, the 3090 will stay where it is, and I will swap out the AMD and Intel cards depending on what I’m trying to test.
|
||
OS install Due to the ever changing nature of my hardware, I wanted to minimize the amount of software installed on the box itself. Beyond the basic drivers for the GPUs and some essentials, I would rather run everything else in containers. Honestly, the software for local AI is changing so much every week that using container images is really the only way to go.
|
||
I started with the AMD and Intel GPUs installed, as I didn’t have the Nvidia card just yet. Base operating system is Ubuntu 26.04 LTS
|
||
The basic installation didn’t use the full drive. I ran this to fix it:
|
||
sudo lvextend -l +100%FREE /dev/ubuntu-vg/ubuntu-lv -r After logging in, I updated the system:
|
||
sudo apt update sudo apt upgrade Now I need to verify that the correct device drivers are available for both cards. Intel uses the xe driver and AMD uses amdgpu.
|
||
lsmod | grep xe xe 4419584 2 drm_gpusvm_helper 57344 1 xe intel_vsec 24576 1 xe gpu_sched 69632 2 amdgpu,xe drm_gpuvm 57344 1 xe drm_buddy 28672 2 amdgpu,xe drm_ttm_helper 20480 3 amdgpu,xe ttm 135168 3 amdgpu,drm_ttm_helper,xe drm_exec 12288 3 drm_gpuvm,amdgpu,xe drm_suballoc_helper 24576 2 amdgpu,xe drm_display_helper 303104 2 amdgpu,xe cec 106496 3 drm_display_helper,amdgpu,xe i2c_algo_bit 16384 2 amdgpu,xe video 77824 2 amdgpu,xe lsmod | grep amdgpu amdgpu 21569536 0 amdxcp 12288 1 amdgpu drm_panel_backlight_quirks 12288 1 amdgpu gpu_sched 69632 2 amdgpu,xe drm_buddy 28672 2 amdgpu,xe drm_ttm_helper 20480 3 amdgpu,xe ttm 135168 3 amdgpu,drm_ttm_helper,xe drm_exec 12288 3 drm_gpuvm,amdgpu,xe drm_suballoc_helper 24576 2 amdgpu,xe drm_display_helper 303104 2 amdgpu,xe cec 106496 3 drm_display_helper,amdgpu,xe i2c_algo_bit 16384 2 amdgpu,xe video 77824 2 amdgpu,xe lspci | grep VGA 03:00.0 VGA compatible controller: Intel Corporation Battlemage G31 [Intel Graphics] 08:00.0 VGA compatible controller: Advanced Micro Devices, Inc. [AMD/ATI] Navi 48 [Radeon AI PRO R9700] (rev c0) 78:00.0 VGA compatible controller: Advanced Micro Devices, Inc. [AMD/ATI] Granite Ridge [Radeon Graphics] (rev c2) Looks good. The Intel Arc is the Battlemage G31 card and the Navi 48 is the Radeon Pro 9700. Don’t you love it hardware has like six different names depending on which tool you use? Don’t worry, it gets more fun!
|
||
I can grab more info with lshw:
|
||
sudo lshw -C video *-display description: VGA compatible controller product: Battlemage G31 [Intel Graphics] vendor: Intel Corporation physical id: 0 bus info: pci@0000:03:00.0 version: 00 width: 64 bits clock: 33MHz capabilities: pciexpress msi pm vga_controller bus_master cap_list rom configuration: driver=xe latency=0 resources: iomemory:280-27f iomemory:180-17f irq:169 memory:2800000000-2800ffffff memory:1800000000-1fffffffff memory:dda00000-ddbfffff memory:2801000000-2804ffffff memory:2000000000-27ffffffff *-display description: VGA compatible controller product: Navi 48 [Radeon AI PRO R9700] vendor: Advanced Micro Devices, Inc. [AMD/ATI] physical id: 0 bus info: pci@0000:08:00.0 version: c0 width: 64 bits clock: 33MHz capabilities: pm pciexpress msi vga_controller bus_master cap_list rom configuration: driver=amdgpu latency=0 resources: iomemory:e00-dff iomemory:e80-e7f irq:172 memory:e000000000-e7ffffffff memory:e800000000-e80fffffff ioport:f000(size=256) memory:dde00000-dde7ffff memory:dde80000-dde9ffff *-display description: VGA compatible controller product: Granite Ridge [Radeon Graphics] vendor: Advanced Micro Devices, Inc. [AMD/ATI] physical id: 0 bus info: pci@0000:78:00.0 logical name: /dev/fb0 version: c2 width: 64 bits clock: 33MHz capabilities: pm pciexpress msi msix vga_controller bus_master cap_list fb configuration: depth=32 driver=amdgpu latency=0 mode=1920x1080 resolution=1920,1080 visual=truecolor xres=1920 yres=1080 resources: iomemory:f80-f7f irq:82 memory:f810000000-f81fffffff memory:dd400000-dd5fffff ioport:e000(size=256) memory:dd900000-dd97ffff Awesome. So both cards are there and the correct drivers are loaded. Which version of the drivers am I using? Good question! Let’s collect more about the system and drivers.
|
||
Linux ai-pc 7.0.0-28-generic
|
||
To get kernel drivers I can use lspci -k -s <domain:bus:slot>
|
||
Here’s what I get for the two cards:
|
||
03:00.0 VGA compatible controller: Intel Corporation Battlemage G31 [Intel Graphics] Subsystem: Intel Corporation Device 1701 Kernel driver in use: xe Kernel modules: xe 08:00.0 VGA compatible controller: Advanced Micro Devices, Inc. [AMD/ATI] Navi 48 [Radeon AI PRO R9700] (rev c0) Subsystem: ASRock Incorporation Device 5413 Kernel driver in use: amdgpu Kernel modules: amdgpu Still doesn’t tell me version.
|
||
modinfo xe | grep version will tell me, but it just matches the Linux kernel version, so if I want to upgrade the drivers, I probably need to update my Linux kernel.
|
||
Device paths can be found by running ls against the path /dev/dri/by-path. The symlink will show which card maps to which render device.
|
||
Intel Arc - /dev/dri/renderD128 Radeon 9700 - /dev/dri/renderD129 I think that should stay pretty consistent.
|
||
Software install Next, I need to install the software I plan to use to run the GPUs. Here’s a short list of items:
|
||
Podman to run containers Python3 - included with Ubuntu build uv - although this might end up being optional Git - included GitHub CLI - installed GPU tools for monitoring (nvtop) There are a few other pieces of software I ended up installing, but these were the big ones. Like I said, the goal is to minimize what’s installed locally and leverage containers for things like uv, python, GPU software, llama.cpp, and vLLM.
|
||
I chose to use podman b/c I hate myself and want everything to be more difficult. It’s the same reason I bought an Intel GPU!
|
||
Permissions To run containers and pass the GPU device, my user needs to be a member of render and video:
|
||
sudo usermod -aG render,video $USER
|
||
I needed to log out and back in for this to take effect.
|
||
Intel and AMD testing I am going to try and use podman where possible. Here’s what I’ve discovered. When running podman, I have to include the switch --group-add keep-groups. Essentially that grants the container the same group access as me. I could probably be more explicit and use the render and video groups.
|
||
podman run --rm -it --group-add keep-groups --device /dev/dri/renderD128 docker.io/intel/oneapi-basekit:latest bash Trying with llama.cpp
|
||
podman run --rm -it --group-add keep-groups --device /dev/dri/renderD128 ghcr.io/ggml-org/llama.cpp:full-intel bash I’m going to try running Qwen3.5-9B-Q4_K_M. First I will download it to a new directory called models:
|
||
mkdir models cd models wget https://huggingface.co/unsloth/Qwen3.5-9B-GGUF/resolve/main/Qwen3.5-9B-Q4_K_M.gguf Next I can launch a container to run the model:
|
||
podman run --rm -it \\ --group-add keep-groups \\ --device /dev/dri/renderD128 \\ -v /home/ned/models:/models \\ ghcr.io/ggml-org/llama.cpp:full-intel \\ --run -m /models/Qwen3.5-9B-Q4_K_M.gguf That worked well for interactive chats. What about benchmarking and perplexity? The “full” container image uses llama-cli as it’s entry point, so anything after the docker image name will be passed as arguments to llama-cli.
|
||
You can run a benchmark by using --bench -m <model_file>
|
||
For example, the Qwen3.5-9B-Q4_K_M ran at 2616t/s for pp512 and 71t/s for tg128.
|
||
It can also quantize models for us. I am pulling down the Qwen3.5-9B-BF16.gguf model and will quantize it.
|
||
--quantize "/models/Qwen3.5-9B-BF16.gguf" "/models/Qwen3.5-9B-Q8_0.gguf" Q8_0 The quantize time was ~7 seconds!
|
||
Checking out the various sizes:
|
||
ls -lh /users/ned/models -rw-rw-r-- 1 ned ned 18G Jul 22 15:12 Qwen3.5-9B-BF16.gguf -rw-rw-r-- 1 ned ned 5.3G Jul 22 14:02 Qwen3.5-9B-Q4_K_M.gguf -rw-r--r-- 1 ned ned 9.2G Jul 22 15:21 Qwen3.5-9B-Q8_0.gguf What if we want to use the AMD card? The image is ghcr.io/ggml-org/llama.cpp:full-rocm
|
||
podman run --rm -it \\ --group-add keep-groups \\ --device /dev/dri \\ --device /dev/kfd \\ -v /home/ned/models:/models \\ ghcr.io/ggml-org/llama.cpp:full-rocm \\ --run -m /models/Qwen3.5-9B-Q4_K_M.gguf Note that /dev/dri and /dev/kfd both have to be included, and it doesn’t work if you just pass the specific device /dev/dri/renderD129. Why? Not sure.
|
||
I’d like to try out the Granite model from IBM:
|
||
ibm-granite/granite-4.1-30b-GGUF:Q4_K_M
|
||
podman run --rm -it \\ --group-add keep-groups \\ --device /dev/dri \\ --device /dev/kfd \\ -v /home/ned/models:/models \\ ghcr.io/ggml-org/llama.cpp:full-rocm \\ --run -m /models/granite-4.1-30b-Q4_K_M.gguf After running the benchmark for the Granite model on both cards, here’s the results:
|
||
| model | size | params | backend | ngl | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | --------------: | -------------------: | | granite ?B Q4_K - Medium | 16.29 GiB | 28.87 B | SYCL | -1 | pp512 | 986.08 ± 2.30 | | granite ?B Q4_K - Medium | 16.29 GiB | 28.87 B | SYCL | -1 | tg128 | 26.03 ± 0.02 | | granite ?B Q4_K - Medium | 16.29 GiB | 28.87 B | ROCm | -1 | pp512 | 1053.69 ± 6.92 | | granite ?B Q4_K - Medium | 16.29 GiB | 28.87 B | ROCm | -1 | tg128 | 29.75 ± 0.03 | The AMD card is slightly faster, but it also uses more power. The Intel card uses ~230W and the Radeon pulls 300W.
|
||
What’s interesting is that the benchmark only takes up 16.29GB of VRAM, but when I load the model for chat, it takes up 30GB. Is that a consequence of no context being set for the benchmark testing?
|
||
I should try setting an explicit context and some other arguments to see if it changes how the model runs.
|
||
podman run --rm -it \\ --group-add keep-groups \\ --device /dev/dri \\ --device /dev/kfd \\ -v /home/ned/models:/models \\ ghcr.io/ggml-org/llama.cpp:full-rocm \\ --run -m /models/granite-4.1-30b-Q4_K_M.gguf \\ -c 4096 That did it. The VRAM usage dropped way down when loading the model, so it must be something about the defaults.
|
||
Granite requests a context window of 131,072 tokens. I’m not sure which file that is stored in, but when you load the model with verbosity set to level four, this line shows you the value:
|
||
0.00.416.815 I llama_model_loader: - kv 7: granite.context_length u32 = 131072 llama.cpp tries to set the context window to the what the model requests without going beyond the available VRAM on the card. For example, if I don’t specify a context value for Granite, then the model itself takes up 16.7GB of VRAM. That leaves about 15.3GB of VRAM for the context window plus headroom.
|
||
The calculation for the KV cache has to do with the number of layer in the model, the number of KV heads, and the head dimensionality.
|
||
First you need to know how many values are stored per context token. That is a function of the attention layer.
|
||
Values per token = 2 x layers x heads x head-dimension
|
||
For the Granite 4.1 model, there are 131,072 values per token. If we are using FP16, then that is 2-bytes per value, giving a total of 262,144 bytes or 0.25MB per token.
|
||
So if the context is set to 65k, that will take up ~16GB of space in VRAM. The requested 131k tokens would take up almost 32GB!
|
||
Basically llama.cpp looks at how much VRAM is left over and divides by the storage per token, with a goal of leaving 1GB of head space.
|
||
14,400/0.25 = 57,600 context window.
|
||
If you want to reduce the space used by the context window, you can apply quantization to the storage of K and V. -ctk q8_0 and -ctv q8_0 will use 8-bit quantization to store the values. That should drop the amount of storage by 50%.
|
||
Even at that level of quantization, we still can’t quite reach the 131K context, and instead have to settle for 106K. If I use q4_0 then the whole model and context can fit in about 26GB. Which is great, but 4-bit quantization might mess up the accuracy of the model too badly. More testing would be required.
|
||
Oh, and the benchmark runs? They only use 512 tokens for the context window, so that’s why the memory usage is only 16GB. It’s basically just the model.
|
||
Conclusion Those are the notes from my first attempt with the AI PC. I’ve built out a system with minimal software, which will make it easy to manage and rebuild going forward. All of the quickly changing components, like llama.cpp and PyTorch packages, live in the container images.
|
||
My next adventure is trying out Unsloth Studio and hosting Jupyter notebooks. Catch you then!
|
||
`,summary:"I’ve been building an AI PC for myself as part of the preparation for a series of courses on Pluralsight focusing on running Small Language Models. Local AI is an area that I’ve wanted to dig into for a while, and building a course for others is a great way to force myself to learn. It’s also a fantastic excuse to spend money on hardware - like I needed one.",date:"29 Aug, 2026",url:"https://nedinthecloud.com/2026/08/29/prepping-my-ai-pc/",image:"PrepAIPC.png",readingTime:"11"},"https://nedinthecloud.com/2026/05/23/trail-running-stuff/":{title:"Trail Running Stuff",tags:["trailrunning","techfree"],content:`This post has basically nothing to do with technology. If you’re looking for that stuff, check out the rest of (looks around) this whole site. I thought of this post while out trail running this weekend, and I didn’t have anywhere else to put it. I’m not going to start a totally separate website to post one page of stuff. Who knows? If this becomes a more regular thing, maybe I will. For now, it’s going to live here.
|
||
I started trail running for real about five years ago. Prior to taking on trail running, I had been road running since 2008-ish. In that time, I finished my fair share of half-marathons and full marathons, as well as a half Iron Man in 2015. What’s I’m trying to say is that I already had a solid aerobic base to work off of. But I was becoming increasingly bored with the same old running routes, especially the long weekend runs, and I was ready to try something new. Trail running was that new thing.
|
||
Since I started trail running, I have a renewed enjoyment of running in general. I’ve finished several 50ks and last year I ran 1,773 miles. That is by far a personal best for me. After spending countless hours enjoying the trail, I’ve learned a bit of what to do and what not to do, and I thought I would share it here. If you’re considering picking up trail running, I hope that this list helps you a bit. Trails are rough enough already, let’s see if I can smooth out your transition from road running a bit.
|
||
This advice is good for all trail running, but my primary focus is on more technical trails. What does technical mean? It simply means that the trail is more challenging due to its ruggedness. On a more technical trail, you can expect to find lots of obstacles and uneven terrain. There may be steep climbs, boulder fields, or creek crossings. The more technical a trail is, the less it resembles a nice asphalt ribbon winding through the woods and the more it resembles this.
|
||
I don’t want this advice to scare you off from trail running. Far from it! The trail is for everyone: walkers, runners, hikers, bikers, bipeds, quadrapeds, hell even octopods if they can manage it. In our over-digitalized age, there’s nothing so lovely as spending some time in nature. I’m not a spiritual person, but if I were, the trails would be my church.
|
||
This laundry list of advice is intended to help you have the best time possible on the trail without a bunch of unexpected surprises. The most important piece of advice is to get out there and enjoy yourself. It doesn’t matter if you only hike a single mile or run a 50k, just go do it. You’ll thank yourself later!
|
||
Go Slow, Not That You Have A Choice My first piece of advice is to slow down. Like waaaay down. Honestly, trails aren’t going to give you much of a choice, but I want you to understand how much slower you’ll be when you first start trail running. Back in 2021, my average road pace was around an 8 minute mile. My first trail run that year? My pace was 11 minutes per mile over 7.5 miles.
|
||
Let that sink in. My running pace dropped by 38%. It’s not like I wasn’t working hard. My average heart-rate on that run was 151 bpm. For comparison, the weekend prior I ran 12 miles at an 8:15/mile pace and my heart-rate was 147 bpm.
|
||
The simple truth is that trails make you work harder. The terrain is uneven, soft, and sometimes slippery. Each step will require more energy to push off. There are going to be roots, rocks, creeks, and other obstacles you have to navigate. The path can be full of twists and turns that your GPS doesn’t quite catch, creating a discrepancy between how far you actually ran versus what was recorded.
|
||
When you’ve finished your first trail run and your pace seems abysmally slow, I don’t want you to be disheartened. That’s 100% normal, and it will get better as you learn to navigate the more challenging terrain. Speaking of which…
|
||
Pick Up Your Feet When you’re running on smooth, even asphalt, you can get away with a shuffling run. Your feet barely have to rise above the height of the road. Try that on a trail and you will be eating dirt shortly. Most trails are uneven, with rocks and roots sticking up out of the ground. If you don’t pick up your feet, your toes are going to catch those protrusions and at best you’ll stumble. More likely, you are going to get up close and personal with the ground.
|
||
There’s lot of advice out there around improving your running form. The best and easiest piece I’ve found for trail running is to simply pick you feet up. Make sure you are kicking off and your foot is snapping back, almost like you’re trying to kick your own ass. Maybe not that extreme, but not far from it.
|
||
During especially long runs, this little gem has proven invaluable to me. Four hours into a run, I’m tired and I don’t feel like picking up my feet. In fact, I’d like to lie down and have someone carry me to the finish line. But if I stop picking up my feet, I’ll be eating grass or shattering a toe. So I keep stepping lively to avoid injury and bloodied up palms.
|
||
Case in point. A couple years ago, I was at the tail end of a 16 mile run and I stopped picking my feet up like I should. The result? Going down a slope, my left foot slammed into a rock and brought me to a full stop.
|
||
At the time I didn’t know it, but I had just fractured my pinky toe. Through a fog of adrenaline and lug-headed stubbornness that only comes from being a long-distance runner, I managed to run the remaining two miles back to my car. Only once I cooled down did I realize how bad my toes really were.
|
||
One trip to the urgent care facility and 8 weeks of recovery later, you can bet your ass I started picking up my feet.
|
||
Still though, there’s no getting around it…
|
||
You’re Going To Fall At some point in your trail running adventures, you are absolutely, 100%, no doubt, going to unceremoniously eat shit. It happens to everyone. Newbies and seasoned runners alike. In fact, I don’t think you can call yourself a real trail runner until you have stumbled, awkwardly fallen, and bruised both your hip and your ego.
|
||
Since you are inevitably going to trip over some rogue root and have a date with the dirt, I do have a few pieces of wisdom I’d like to impart to you:
|
||
Try to fall with grace. You might be inclined to try and save yourself from a tumble. Chances are you’ll make it worse. Accept gravity as the harsh mistress it is. Don’t use your hands to stop yourself. You’re cruising at a decent pace. There’s a lot of momentum behind you. Do you really think your wrists are going to be happy about trying to halt that inertia? Scraped up palms and sore forearms have taught me otherwise. You’re probably better off trying to tuck and roll, protecting your head and torso. Stop and assess. My first impulse after getting up from a fall is to keep running. It’s a mix of adrenaline and embarrassment, neither of which help me determine how bad the fall was. Give yourself a moment and make sure nothing is bloody, broken, or bent wrong. Like a toe perhaps. With any luck, you’ll survive your first few falls with minimal damage. In case you don’t…
|
||
Be Prepared I’m not going to tell you that you need to pack a full first-aid kit for every trail run. Frankly, that’s cumbersome and not realistic. On the other hand, it couldn’t hurt to have a couple basics. I recommend bringing a couple band-aids and maybe some gauze packed in something waterproof. It’s not just falls that may bloody you up. Nature is full of unexpected surprises and many of those surprises involve thorns. So many thorns. It’s like the world is a vampire after your blood.
|
||
Quick story. Last year I was running a trail I’d been on dozens of times. Some smaller trees had branches dipping down into the path. As I ducked to avoid a branch, I failed to notice that a pricker bush had also invaded the path at a slightly lower height. One of the thorns caught my left ear and snagged. I wasn’t looking for a new ear piercing that day, but nature had other plans.
|
||
After coming to a screeching (it was me doing the screeching) halt and dislodging the barb from my cartilage, I now had a minor flesh wound that, given it’s proximity to my head, was bleeding profusely. Stupidly, I did not bring any gauze or band-aids with me, so now I had gore dripping down my ear onto my neck and shirt.
|
||
The nearest port-o-john was about a mile away, so I ran there, doubtlessly terrifying any other travelers on the path with the grisly scene festering on the left side of my head. Fun times.
|
||
I did make it to the johnny-on-the-spot, which had toilet paper and hand sanitizer. Once I cleaned myself up a bit, the actual damage to my ear was minimal. I could have avoided the whole ordeal if I had simply packed some basics. Of course it could have been much worse, which is why you should…
|
||
Tell Someone Where You’re Going If you get dehydrated, pull a calf muscle, or collapse on the side of the road, chances are some good samaritan is going to happen along and render assistance. When you’re 10 miles deep on a remote forest trail, the chances of someone stumbling across your exhausted form drop precipitously.
|
||
Additionally, many remote trails are not going to have the best cell service. Don’t depend on your phone to be able to make an emergency call if you’re stranded and unable to make it back to your car.
|
||
For all those reasons and more, when you go on a trail run, tell someone. Tell them where you’re going and when you expect to be back. It doesn’t have to be a big dramatic thing. When I leave for my Sunday morning long runs, I tell my wife something like, “I’m heading out to Lake Nockamixon and I’ll be back around noon.” If 1PM rolls around and I’m still not back, she can call in the calvary, or at least call me.
|
||
I also have my Garmin watch set up to allow my wife to track my runs, so she has a good idea of where I’m at anyway. But like I said, cell service can be spotty, and I wouldn’t rely on it. She’s texted me in the past because my marker didn’t move for 30 minutes due to bad cell service.
|
||
Mud Is Happening Unless you’re planning to run exclusively in the desert, you’re going to get wet and muddy. My second trail run ever was near a lake that had horse trails. The combination of runoff and chewed up turf made for a freakin’ mud pit to rival Woodstock ‘96.
|
||
Trails tend to involve a decent amount of mud, puddles, creek crossings, and all manner of dust. When you get back to your car, you’re probably going to be a hot mess. Personally, I always pack a towel and a shirt to change into. I use the towel to clean off the worst of it, and lay it down on the driver’s seat. In the winter, I usually pack a hoodie instead of a shirt to change into.
|
||
You may also want to consider wearing gaiters over your shoes. It’s one thing for your shoes and socks to get wet, it’s quite another for them to get filled with mud, rocks, and sand. Gaiters will stop the worst of it from getting in, and you get to look super-stylish in the process.
|
||
The World Is Your Bathroom Here’s a bit of good news sprinkled in with all the doom and gloom. Are you tired of waiting for port-a-potties or searching for a restroom in a crowded urban environment? Well good news friend! The wait it over. When you’re trail running, the world is your bathroom. Although, it doesn’t stock its own toilet paper.
|
||
If you’re the kind of person who tends to get runner trots- a delightful euphemism for poop attacks- then you’re going to want to bring some reinforcements with you. I recommend folding up some toilet paper and putting it in a zip-lock bag. That’ll keep it from getting wet from your sweat or weather conditions. That and some hand sanitizer is all you need.
|
||
At least you don’t have to buy a soda or ask for the key!
|
||
Try Going Without Headphones Ok, hear me out. This is not going to be some holier-than-thou lecture on how listening to podcasts or music ruins the running experience and how dare you not be a puritan punishing yourself with a hair shirt and leather flail on each and every stride? Some folks like to get high and mighty about how they never run with headphones and how it blemishes the experience, marring the enjoyment of running.
|
||
Those people can fuck 👏 all👏 the👏 way👏 off.
|
||
I love music, podcasts, and audiobooks. I love running. Why wouldn’t I want to do them together? When it comes to the trail, I think it is worth considering going without.
|
||
First of all, there’s the practical consideration of being safe and aware on the trail. You are going to find yourself on single-track paths with no room to pass. When a runner or biker comes up behind you, they shouldn’t need to holler at the top of their lungs to get your attention. Running trails also requires a high-level of attentiveness to your surroundings and the terrain. Adding music, podcasts, or audiobooks on top of that might be too much, especially when you’re first getting started or trying to navigate especially tricky paths.
|
||
Even now, five years in, I will sometimes pause my music when the trail gets too complicated or technical. I need all my attention focused on not injuring myself, and I find any other stimulus too distracting.
|
||
The other major component is simply enjoying nature itself. I know this might sound a bit trite, but the trails offer more than just beautiful sights. There’s something to engage all of your senses. The smell of flowers and decomposing leaves. The feeling of the path under your feet and wind on your face. The sound of chirping birds, rushing water, and susurration of insects in the meadow.
|
||
I’m not saying you need to leave the headphones at home. I’m just saying you should try starting your run without them and see how it goes. Personally, I tend to do about 1 in 4 of my long runs without headphones in. I still bring them with me, and will pop them in if I need a little motivational boost. Last year, I did a 50k with no headphones and I found it delightful. Your mileage may vary.
|
||
Be Courteous And Kind There’s already volumes on proper trail etiquette and I’m not going to rehash them here. That’s what the Google machine is for. But I will lay down five essentials.
|
||
Everyone belongs on the trail. People of all sizes, backgrounds, ethnicities, genders, whatever. All are welcome. Be considerate of others. Don’t blast your music out of a bluetooth speaker or be a loud asshole. Give a heads up. Let people know you’re coming up behind them when passing. Do this earlier than you think you should. Take nothing. Leave nothing. There are exceptions, of course, but this generally holds. Be a good steward of the trail. Smile and wave. Smiling makes running easier. So does being friendly and getting a wave back. Bring WATER. ALL THE WATER. Hydrate, hydrate, hydrate.
|
||
You may have been able to get away with running water bottle free on the road. That is not going to be the case on the trail. There are no random water fountains or convenience stores. Whatever liquid you need, you’ll be carrying with you. For my longer runs (2+ hours) I always fill up my 1.5L water pack. It might seem like overkill, especially in the winter or fall when you aren’t sweating as much. I’m here to tell you that I don’t care what season it is, I’m still bringing lots of water with me.
|
||
One of my absolute worst experiences happened because I let myself get severely dehydrated. The original plan was to go out for a 14-ish mile run at a lake near me. I wasn’t super familiar with the trails, but I figured it would be fine. I brought maybe a liter of water with me in my hydration pack.
|
||
Well, I ended up getting lost. The day was hotter and more humid than expected, and I was sweating a ton. I finished my water around mile 11 and was getting thirstier by the minute. What was supposed to be a 14 mile run was now up to over 16 miles and I still wasn’t sure how close I was to the trail head. I knew I was really in trouble when I had to stop and walk for the last mile. I simply could not run anymore.
|
||
Easing into my car, I felt shaky and my head was cloudy. It was hard to focus properly and assemble my thoughts in a coherent order. All I wanted to do was get home, roughly a 25 minute drive. 15 minutes into the drive, my feet and calves started cramping. I had to pull over the car and stretch before I could drive again.
|
||
Once I got home, I dragged myself to the couch and collapsed. I was trembling and cramping all over. My wife gave me two bottles of gatorade and went to the store to buy more. She wanted to take me to the hospital, which probably would have been a wise decision. I insisted I would be okay; all I needed was to rehydrate.
|
||
Over the next four hours, I drank 4x 32oz gatorade bottles. My pee was basically brown. I ate bananas and pretzels to get back some potassium and sodium. The cramping stopped after two hours. It was fucking hell. It could have been so much worse.
|
||
So yeah. Bring the water. All the water. Drink more than you think you need to. Be consistent and methodical. Because dehydration sucks ass.
|
||
Bring Fuel Too As a quick corollary, you’re also going to want some fuel with you as well. You’re not just losing water, you’re also losing electrolytes and burning through calories. If you’re already a long distance runner, you probably use gels or something similar for those long runs. Bring that with you on trails runs too, plus a bit more.
|
||
Personally, I stopped using gels a couple years ago and switched to Tailwind. Gels were not sitting well with my tummy, especially after the fourth or fifth packet. Also, if you bring gel packs with you, then you have to bring them back out. Sticky gel packs that get all over your fingers and running vest. With Tailwind, the nutrition is in the water I’m already carrying with me. There’s no sticky residue and no trash to stash somewhere on my person.
|
||
That’s what works for me. Find something that works for you and bring extra.
|
||
Buy New Stuff The last thing I want to mention is the need for new gear when you start trail running. This is me giving you permission to spend money. You’re welcome!
|
||
At a minimum, you need to buy a pair of running shoes meant for the trails. Shoes that have bigger lugs for traction, better protection for your toes, and something for gaiters to grab onto. My shoe of choice is the Altra Lone Peak. Lots of people swear by Kona, but I just don’t like that level of padding in my shoe. Find the one that works for you.
|
||
You’re probably also going to want the following gear:
|
||
Hydration pack that carries a minimum of 1.5L (I like my Osprey vest) Bright orange running hat (for hunting areas) Gaiters for when things get muddy Running watch with great GPS (Garmin Forerunner is my pick) Have Fun! Trail running is an amazing hobby. It’s truly changed my life for the better, and I hope it does for you as well. My first few runs were memorable and not entirely for good reasons. I wish I had read a post like this before I headed out the door. Well friend, now this post exists for you and I hope it helps.
|
||
If you see me on the trail, be sure to smile and wave. You know I’ll wave back.
|
||
`,summary:"This post has basically nothing to do with technology. If you’re looking for that stuff, check out the rest of (looks around) this whole site. I thought of this post while out trail running this weekend, and I didn’t have anywhere else to put it. I’m not going to start a totally separate website to post one page of stuff. Who knows? If this becomes a more regular thing, maybe I will.",date:"23 May, 2026",url:"https://nedinthecloud.com/2026/05/23/trail-running-stuff/",image:"TrailRunning.png",readingTime:"17"},"https://nedinthecloud.com/2026/05/11/developers-are-people-the-psychology-of-software-teams/":{title:"Developers Are People: The Psychology of Software Teams",tags:["day-two-devops","psychology","software-teams","developer-experience","leadership"],content:`There is a stubborn myth in tech that good software work is mostly about tools, velocity, and individual brilliance. If the metrics look good and features ship on time, everything else is secondary. People will adapt. Teams will sort themselves out. Culture is nice to have, but the real work is the code.
|
||
I don’t think that story has ever been true.
|
||
In an upcoming episode of Day Two DevOps, Kyler and I talk with Dr. Cat Hicks, a psychological scientist who studies software teams, about her upcoming book, The Psychology of Software Teams. What I appreciated most about this conversation is that Cat did not treat psychology as a soft add-on to technical work. She treated it as a practical tool for understanding why teams get stuck, why certain myths refuse to die, and how better systems get built when we stop pretending developers are machines.
|
||
That idea showed up early and often in the conversation. Cat said she kept a note in her office while writing the book that read: developers are people. A wild concept, I know.
|
||
Developers Are Not Build Pipelines One of Cat’s core arguments is that technical culture has accumulated a lot of bad stories about how work gets done.
|
||
Some of those stories are familiar:
|
||
the best engineers are the fastest engineers the highest performers can sprint forever real breakthroughs come from lone geniuses emotions, identity, conflict, and belonging are distractions from the work Those ideas can feel intuitive, especially in environments obsessed with throughput. But Cat makes the case that they are often distortions. They push teams toward brittle productivity, shallow measures of performance, and organizational habits that burn people out while teaching leaders the wrong lessons.
|
||
Most of us have seen some version of play out in real time. A team pushes hard, ships something impressive, and then collapses under the weight of stress, rework, or strained relationships. From a distance it can look like success. Up close, it is often a warning sign.
|
||
More Bands, Fewer Rockstars One of my favorite ideas from the episode is Cat’s challenge to the rockstar myth.
|
||
In tech, we love the story of the singular genius. It is neat, flattering, and easy to repeat. It also leaves out how software actually gets built. Real systems come from teams. They come from coordination, trust, shared understanding, and all the invisible glue work that rarely fits inside a heroic narrative.
|
||
Cat talks about this with a phrase that stuck with me: more bands and fewer rockstars.
|
||
That is a much healthier model for software work.
|
||
Bands still need talented people. They still need practice, standards, and leadership. But they also need listening, timing, adaptation, and a sense of how individual contributions fit the larger sound. That feels a lot closer to real engineering than the fantasy that one 10x wizard is going to carry the whole effort.
|
||
Why are we so obsessed with individual contributions and the rockstar mythos? This might be a slight American skew, but I think it comes down to our infatuation with the Hero’s Journey. One specially selected individual who overcomes adversity to achieve greatness through grit, perseverance, and innate talent.
|
||
We celebrate heroes; hold them up on a pedestal. It’s a toxic relationship for both the idolized and those doing the worship. Individual excellence is great, but team productivity is what actually matters. Well, what actually matters is achieving business goals.
|
||
The Best Teams Learn How to Understand Themselves Another thread running through the conversation is that strong teams do not just solve technical problems. They get better at understanding how they solve problems.
|
||
Cat described this as becoming an organization that wants to understand itself, which is a phrase I have been turning over in my head ever since.
|
||
That kind of self-understanding matters because context matters. A behavior that is useful in one environment can be destructive in another. A survival strategy from a large enterprise might create friction in a ten-person startup. A process that protects one team might suffocate another. If you never step back and examine the stories, assumptions, and habits driving the work, you can end up preserving the wrong things simply because they once worked somewhere else.
|
||
Kyler had a great anecdote that illustrates this perfectly. You’re just going to have to listen to the episode to hear it.
|
||
Why This Matters Even More in the AI Era One of the more interesting parts of the discussion is how well this topic maps onto the current AI moment.
|
||
Cat noted that even though her book is not primarily about AI, it ends up being deeply relevant to it. That makes sense. AI is changing how teams write, learn, explore, and evaluate software, but it is not changing the fact that teams are made of humans. If anything, it raises the stakes.
|
||
If AI lets us move faster, then bad assumptions can scale faster too. If tools make it easier to generate output, then judgment, reflection, and learning matter even more. If organizations start believing that humans are interchangeable because machines can fill in some of the gaps, then the need for a more evidence-based and humane approach becomes urgent.
|
||
Final Thoughts I came away from this episode thinking that software teams do not just need better tooling. They need better theories about people. I’ve known Cat for a few years, and I can say without hesitation that she is a voice worth listening to.
|
||
Amongst all the hype and puffed-up LinkedIn posts, it’s refreshing to hear someone talk about a topic they have thought deeply about. And Cat brings the research chops and professional experience to back it up. That’s what Cat is offering with The Psychology of Software Teams: a way to bring science, empathy, and sharper observation to a field that too often runs on folklore.
|
||
If you have ever worked in an environment shaped by burnout, hero culture, shallow productivity metrics, or the quiet assumption that technical people should function like robots, this episode will probably resonate.
|
||
And if you want to go deeper, Cat’s book The Psychology of Software Teams officially launches on July 14, and her newsletter Fight for the Human is already worth following.
|
||
Keep an eye on Day Two DevOps for the full episode on May 27.
|
||
Written with help from AI
|
||
`,summary:`There is a stubborn myth in tech that good software work is mostly about tools, velocity, and individual brilliance. If the metrics look good and features ship on time, everything else is secondary. People will adapt. Teams will sort themselves out. Culture is nice to have, but the real work is the code.
|
||
I don’t think that story has ever been true.
|
||
In an upcoming episode of Day Two DevOps, Kyler and I talk with Dr.`,date:"11 May, 2026",url:"https://nedinthecloud.com/2026/05/11/developers-are-people-the-psychology-of-software-teams/",image:"DayTwoDevOps-CatHicks.png",readingTime:"5"},"https://nedinthecloud.com/2026/05/04/jet-plane-blues/":{title:"Jet Plane Blues",tags:["techburnout","AI"],content:`Neditor’s note: When I wrote this a few months ago, I really wasn’t sure I was going to publish it. It felt too raw, too personal, and too self-pitying. I shared it with a friend, who gave me constructive feedback that boiled down to:
|
||
Tech ate your life a bit. Time for a break, and to go touch some grass for a while!
|
||
She was 100% right, and that’s what I did. I pulled way back on my level of commitments and went for a few long walks (and long runs.)
|
||
After taking a lot of deep breaths, I found some fun new things to do. I’ve been doing a weekly livestream for the Terraform Pro Exam that is just kinda fun to do. Low pressure and no editing. I’ve also been messing around with Swamp in my home lab. And I rolled out an instance of OpenClaw to see what I could do with it. I’m having fun with tech again, and that’s a massive relief.
|
||
With the perspective of a some months and a better mental space, I reread the post and I still think it’s worth publishing. We tend to only show our best and most polished selves online. We hide the truth for others and from ourself. If you’ve ever thought someone was a rockstar and that nothing bothered them, boy do I have news for you.
|
||
Just as likely, you’ve been told you’re a rockstar and you feel like a failure if you can’t take the pressure. That’s some bullshit. We all need a break sometimes. We all crack under pressure. We all need to go touch grass for a while.
|
||
This is me in the midst of a meltdown, and maybe it can help some of you recognize it in yourselves and find comfort knowing you’re not the only one.
|
||
I’ve Got The Jet Plane Blues There’s something that has been bugging me for a while. A general feeling of dissatisfaction with what I’m doing for work.
|
||
Of all the projects I’m currently working on, I’m not especially excited about any of them. Last year I gave up doing some of my favorite things in favor of focusing on things that make money: training, Pluralsight, and podcasting.
|
||
This all came to a head when I attended the recent Cloud Field Day 25 in Santa Clara, CA. Before I get into it, I want to say up front that everyone at Gestalt IT and the whole Tech Field Day crew are nothing but awesome. I don’t want any of this to reflect negatively on what Tech Field Day does or the value it brings to its clients.
|
||
The event was being held over two days, March 11 and 12th, with the welcome dinner on Tuesday March 10th. Before I even left for the event, things were already going poorly. Daylight Savings Time messed up everyone’s rhythm, our power went out for three hours, and two of my kids were getting sick. At that point, I should have just thrown in the towel, yet I persisted.
|
||
On my trip to California, my connecting flight sat at the gate for three hours after boarding due to computer issues. I arrived in California tired, disheveled, and travel worn. This was just the beginning of the event.
|
||
During the event we had two presenting companies showing off their products. What they could do and why it was important, and I couldn’t care less. That’s not to say that either company makes bad products or that their solutions aren’t useful. They are! I just felt deeply uninterested.
|
||
While I was away, things progressively got worse at home with meltdowns, mad dashes, and logistical nightmares. Who knew having three active kids would be so difficult?
|
||
I didn’t want to be in California. I didn’t want to hear about the latest cloud-native, networking solution powered by AI. I didn’t want to write posts on LinkedIn about how excited I was to be there, or how amazing everything was. I just didn’t give a shit. I wanted to be home.
|
||
I’m not gonna lie, this was a bit of a revelation to me. Not the homesickness, but the complete apathy towards technology. That’s never been true. All my life I’ve been excited about cool new tech and what it can do. Sure, sometimes I would get cynical about a particular line of tech -- web3 anyone? -- but about tech as a whole? Never.
|
||
What is causing this apathy? This general disregard for tech stuff? It’s probably a confluence of things conspiring together. First off, we have AI. Fucking AI. I’m so tired of hearing about AI, and I’m deeply conflicted about the tech. It seems like it can do amazing things, and at the same time it seems like it can ruin our minds, our psyche, and our planet.
|
||
Then there’s Terraform. Once again, I want to be clear, I think Terraform is great. It has transformed the way we define and manage infrastructure. However, it has put me in a bit of a rut and I’m starting to feel it.
|
||
In the last 12 months, I’ve published 10 courses on Pluralsight about Terraform. TEN COURSES. Essentially, I spent 365 days writing about Terraform constantly. On top of that, I also deliver training for Terraform and HashiCorp Vault. In the last 12 months, I think I taught about 10 Terraform classes; that’s about one class every month. Oh, and I can’t forget the alpha and beta testing I did for the Terraform Pro exam and the question writing and review I did for the Terraform Associate 004 course.
|
||
I’ve been living, breathing, and thinking about Terraform constantly for over a year, and I’m tired. So, so tired y’all. Despite all this, when Pluralsight came to me and asked if I could create a six course learning path for the Terraform Associate 004 certification, I agreed to do so.
|
||
That means I am locked in to creating more Terraform content until at least mid-June.
|
||
I love Terraform.
|
||
I’m so, so tired of Terraform.
|
||
Then there’s the fact that I’ve been creating content and delivering training for the last 7 years. Before that, I was the Director of Cloud Solutions for a consulting company for a year and a half. I haven’t been hands-on with tech that matters for almost a decade.
|
||
Part of me wonders if I’d be happier working on consulting projects, or going to work for a company as an individual contributor. Slinging code and building cool infrastructure that does stuff. Lately, I just don’t find content creation exciting or fulfilling. It feels like a drag.
|
||
On the other hand, it could be that I am burned out and need to take a break. In the last year, I have been pushing super hard to engage with projects that actually make Ned in the Cloud money. That means three things: Pluralsight courses, live training, and sponsored podcast content. I fear I have over-indexed on money making ventures that pay the bills, and haven’t spent enough time on stuff that is simply fun.
|
||
What’s fun? Chaos Lever was fun. Me and Chris being ridiculous and snarky jerks on a weekly basis. Instead, we started writing books for O’Reilly cause that makes actual money.
|
||
What else is fun? I used to do a once a month Hashicorp Ambassador chat where we just shot the breeze on a livestream. No sponsors, no pressure.
|
||
Marino Wijay and I did a weekly livestream messing around with AI, and that was fun until I tried to make it serious and bring on lots of guests.
|
||
If I am going to make it long-term in this game, I need to get back to having fun. I need to find technologies that interest me without any financial motivation. I need to take time to myself that doesn’t involve tech.
|
||
Or maybe I need to take a break from all of this and go get a normal job for a while. Being a minor celebrity in the world of HashiCorp was kinda fun for a while, but it’s just not as fulfilling as it once was.
|
||
Being a Microsoft MVP, an AWS Community Builder, a HashiCorp Ambassador, it all feels more like an obligation than a benefit. This year I saw the head of the AWS Community get laid off and the HashiCorp Ambassador program be absorbed into the IBM Champions group. I haven’t attended an MVP Summit in years and I don’t participate in the product group interactions or use much Microsoft tech. I’m glad that others find value and community in these programs, I’m not sure I do.
|
||
I dunno. Writing is how I tend to process how I’m feeling, and this post hasn’t resolved itself into a decision. But I think maybe I have a better grasp on what’s wrong. I think that I need to chat with someone who’s been in a similar position and get their perspective. That’s a pretty tiny group of folks. Yet, they exist and it’s time for me to get out of my own head and talk to other people.
|
||
I’m scared to do it. In part because I worry about how it will sound. Am I just being a whiny, entitled brat? I’m also scared to admit that I’m struggling. And I’m worried what others will think about my doubts and issues. I have an image to maintain. My income is reliant on my image. It’s difficult to be vulnerable. And yet, I think it’s the only thing that will help, unless I want to be a miserable old bastard who pushes everyone away.
|
||
I’m going to be honest. I have no idea if I will even publish this. It feels too personal, too raw. On the other hand, maybe that’s something else that’s been missing. I’ve become too polished, too professional, too good at what I do.
|
||
At a time when everything has the tinge of the AI airbrush, there’s never been a better time to be a genuine, messy human being. And goddamn, am I messy. I’m writing this as I fly back home to be with my family and try to bring some order to the chaos. We’ll see how I feel about this tomorrow.
|
||
`,summary:`Neditor’s note: When I wrote this a few months ago, I really wasn’t sure I was going to publish it. It felt too raw, too personal, and too self-pitying. I shared it with a friend, who gave me constructive feedback that boiled down to:
|
||
Tech ate your life a bit. Time for a break, and to go touch some grass for a while!
|
||
She was 100% right, and that’s what I did.`,date:"4 May, 2026",url:"https://nedinthecloud.com/2026/05/04/jet-plane-blues/",image:"JetPlaneBlues.png",readingTime:"9"},"https://nedinthecloud.com/2026/05/01/how-to-set-up-an-exchange-online-mailbox-for-openclaw/":{title:"How to Set Up an Exchange Online Mailbox for OpenClaw",tags:["openclaw","exchange-online","microsoft-graph","shared-mailbox","entra-id","automation"],content:`I wanted to give OpenClaw its own mailbox in Microsoft 365 so it could read mail, send messages, and generally act like a useful automation agent without piggybacking on my personal account. The trick was doing it in a way that was secure, auditable, and constrained.
|
||
After working through the options, I landed on a model that I think is the right fit for this kind of use case:
|
||
Use a shared mailbox instead of a licensed user mailbox Authenticate with a Microsoft Entra app registration using a certificate, not a password Scope mailbox access with Exchange RBAC for Applications Enforce outbound safety with Exchange transport rules Optionally use a mail-enabled group as a preflight recipient allowlist This post walks through that setup and the reasoning behind it.
|
||
Why I Chose a Shared Mailbox When you first start thinking about “mailbox for an agent,” there are three obvious Exchange object types that come to mind:
|
||
A normal user mailbox A shared mailbox A room/resource mailbox A room mailbox is the wrong object type. It exists for scheduling resources, not for autonomous mail handling.
|
||
A regular user mailbox is the easiest thing to understand because it behaves like a normal account. But it usually means consuming a license and it encourages the wrong authentication pattern: acting like a person signed in to Outlook.
|
||
A shared mailbox turned out to be the better fit.
|
||
It gave me:
|
||
a mailbox identity with its own address no need for a human-style username/password workflow a cleaner separation between “the mailbox” and “the app that uses it” a path to certificate-based, app-only authentication That last point is what really matters. I did not want an automation agent depending on stored user credentials.
|
||
The Security Model Here is the model I ended up with:
|
||
Shared mailbox
|
||
This is the mailbox identity, such as agent@yourdomain.com.
|
||
Entra app registration
|
||
This is the application identity OpenClaw uses to authenticate.
|
||
Certificate credential
|
||
This is the app credential. No client secret, no password, no interactive user login. The alternative was setting up OIDC, but that’s a bit more complicated than I thought the situation warranted.
|
||
Exchange RBAC for Applications
|
||
This limits what the app can do in Exchange and which mailbox it can touch.
|
||
Transport rules
|
||
These are the hard backstop for outbound email restrictions.
|
||
That split matters because mailbox access and outbound delivery control are not the same thing.
|
||
The app needs to be allowed to act on the mailbox, but even after that I still want Exchange itself deciding who that mailbox is allowed to send to.
|
||
Why Exchange RBAC for Applications Matters If you have been reading older Microsoft 365 blog posts, you have probably seen advice that looks like this:
|
||
grant Graph app permissions like Mail.ReadWrite and Mail.Send then use Application Access Policies to narrow access to a mailbox or group of mailboxes That was the older pattern.
|
||
The newer model is Exchange RBAC for Applications. Microsoft is replacing Application Access Policies with this RBAC-based approach, so it makes more sense to start there instead of building on the legacy model and migrating later.
|
||
The important idea is that Exchange Online can now express:
|
||
which application can perform which Exchange role and on which scoped set of mailboxes That is exactly what I wanted.
|
||
Setting Things Up With all that context in mind, here are the steps for setting things up.
|
||
Step 1: Create the Shared Mailbox In Exchange Admin Center:
|
||
Go to Recipients → Mailboxes Create a new shared mailbox Give it an address such as agent@yourdomain.com You do not need to enable direct sign-in for the mailbox. In fact, the whole point of this design is to avoid signing in to the mailbox as if it were a person.
|
||
Step 2: Create the App Registration In Microsoft Entra admin center:
|
||
Go to App registrations Create a new app registration Make it single tenant You do not need a redirect URI for this scenario After you create it, note these values:
|
||
Application (client) ID Directory (tenant) ID Later on, you will also need the Enterprise Application object ID, which is not the same thing as the app registration object ID.
|
||
That distinction matters more than it should, and it is easy to get wrong. Thanks once again to Microsoft for making the naming and IDs in Entra endlessly confusing.
|
||
Step 3: Use a Certificate, Not a Secret For an unattended service, I strongly prefer certificate-based auth over a client secret. Although I didn’t store the certificate in Key Vault for this deployment, you could certainly do so if you’re running your OpenClaw VM in Azure.
|
||
On the OpenClaw host, I created a self-signed certificate with OpenSSL:
|
||
mkdir -p state/m365 openssl req -x509 -newkey rsa:4096 -sha256 -days 365 -nodes \\ -subj "/CN=OpenClaw Shared Mailbox App" \\ -keyout state/m365/openclaw-mail.key \\ -out state/m365/openclaw-mail.crt openssl pkcs12 -export \\ -out state/m365/openclaw-mail.pfx \\ -inkey state/m365/openclaw-mail.key \\ -in state/m365/openclaw-mail.crt chmod 600 state/m365/openclaw-mail.* Then I uploaded the public certificate (openclaw-mail.crt) to the app registration under Certificates & secrets.
|
||
The private key stayed only on the host running OpenClaw. Make sure you’ve set permissions appropriately for the private key, I put it in the home directory used by OpenClaw.
|
||
Step 4: Get the Right Service Principal Object ID When you use Exchange RBAC for Applications, Exchange wants a reference to the Enterprise Application service principal, which is not the object ID from the App Registrations blade.
|
||
So in Entra:
|
||
Go to Enterprise applications Locate your app Copy the Object ID You now need three IDs in hand:
|
||
Tenant ID Client ID / App ID Enterprise Application Object ID Step 5: Create the Exchange Service Principal Pointer In Exchange Online PowerShell:
|
||
Connect-ExchangeOnline New-ServicePrincipal \\ -AppId "<CLIENT_ID>" \\ -ObjectId "<ENTERPRISE_APP_OBJECT_ID>" \\ -DisplayName "OpenClaw Shared Mailbox App" This creates the Exchange-side reference to the Entra app.
|
||
Step 6: Scope the App to One Mailbox I wanted the app to access only one mailbox, not every mailbox in the tenant.
|
||
The cleanest way I found was:
|
||
set a custom attribute on the shared mailbox create a management scope that matches that attribute assign Exchange application roles against that scope For example:
|
||
Set-Mailbox agent@yourdomain.com -CustomAttribute15 "OpenClawMailbox" New-ManagementScope \\ -Name "OpenClaw Shared Mailbox Scope" \\ -RecipientRestrictionFilter "CustomAttribute15 -eq 'OpenClawMailbox'" Using a custom attribute makes the scope explicit and easy to reason about later.
|
||
Step 7: Assign the Exchange Application Roles The mailbox roles I needed were:
|
||
Application Mail.ReadWrite Application Mail.Send To assign them:
|
||
New-ManagementRoleAssignment \\ -Name "OpenClaw Mail ReadWrite" \\ -Role "Application Mail.ReadWrite" \\ -App "<CLIENT_ID>" \\ -CustomResourceScope "OpenClaw Shared Mailbox Scope" New-ManagementRoleAssignment \\ -Name "OpenClaw Mail Send" \\ -Role "Application Mail.Send" \\ -App "<CLIENT_ID>" \\ -CustomResourceScope "OpenClaw Shared Mailbox Scope" If you want calendar access as well, you can add an Exchange application calendar role later. I would not pre-grant more than you actually need.
|
||
Step 8: Test Exchange Authorization Before wiring anything into OpenClaw, I wanted to know whether Exchange understood the scope correctly.
|
||
Test-ServicePrincipalAuthorization \\ -Identity "<CLIENT_ID>" \\ -Resource agent@yourdomain.com And just as importantly, test a mailbox it should not have access to:
|
||
Test-ServicePrincipalAuthorization \\ -Identity "<CLIENT_ID>" \\ -Resource someoneelse@yourdomain.com That negative test is worth doing.
|
||
Step 9: Integrate the Mailbox into OpenClaw On the OpenClaw side, I stored mailbox metadata and certificate material locally, something like this:
|
||
{ "tenantId": "YOUR_TENANT_ID", "clientId": "YOUR_CLIENT_ID", "mailbox": "agent@yourdomain.com", "certPath": "state/m365/openclaw-mail.pfx" } Then I asked OpenClaw to build a small helper script that:
|
||
creates a client assertion JWT from the private key requests a token from https://login.microsoftonline.com/<tenant>/oauth2/v2.0/token calls Microsoft Graph with the resulting bearer token The important Graph call for mailbox reads looks like this:
|
||
GET /v1.0/users/agent@yourdomain.com/mailFolders/inbox/messages And for sending:
|
||
POST /v1.0/users/agent@yourdomain.com/sendMail With the Exchange RBAC setup in place, OpenClaw was able to read from and send as the shared mailbox without any delegated user login.
|
||
Step 10: Add Hard Outbound Controls with Mail Flow Rules Mailbox access is only half the problem. The other half is making sure the mailbox cannot blast mail anywhere it wants.
|
||
For that, I used an Exchange transport rule.
|
||
The pattern is simple:
|
||
apply the rule when the sender is the shared mailbox allow only approved recipients or domains reject everything else This gives you a genuinely useful second line of defense.
|
||
Even if the agent tries to send to someone it should not, Exchange will stop the message.
|
||
I also added a second action that sent a notification to me if OpenClaw tried to send to a non-approved address. If OpenClaw suddenly starts spamming hundreds of people with a Nigerian Prince scam, I want to know even if the emails are all bounced by the transport rule.
|
||
For the approved recipients, I created an Exchange group and added contact and users to it that I wanted OpenClaw to be able to contact. That makes it easy for me to update the allowed recipients without altering the transport rule each time. It also lets OpenClaw due a pre-flight check, as we’ll see shortly.
|
||
What I Observed During Testing One behavior surprised me a little, but it makes sense once you see it.
|
||
When I sent mail through Microsoft Graph to a blocked external recipient:
|
||
Graph accepted the sendMail submission the message appeared in Sent Items Exchange transport rules later rejected delivery the mailbox received an NDR/undeliverable message In other words, submission success is not the same as delivery success.
|
||
That is an important operational detail if you plan to automate around this.
|
||
Do not treat “message exists in Sent Items” as proof that the message was delivered.
|
||
Optional: Use a Group as a Preflight Allowlist Transport rules are the hard enforcement layer, but I also wanted OpenClaw to do a quick preflight recipient check before it even tried to send.
|
||
For that, I created a mail-enabled group containing the approved recipients, something like:
|
||
openclawallowed@yourdomain.com
|
||
Then OpenClaw reads the members of that group and refuses to send if any requested recipient is not in the allowlist.
|
||
That sounds straightforward, but it required one more set of permissions.
|
||
Mailbox permissions and directory permissions live in different planes.
|
||
Exchange RBAC for Applications gave the app mailbox access, but it did not give the app permission to read Entra group membership.
|
||
To make the allowlist group readable, I had to add Microsoft Graph application permissions.
|
||
The smallest set that worked in my test case was:
|
||
GroupMember.Read.All Group.Read.All User.ReadBasic.All OrgContact.Read.All Why all four?
|
||
Because the group could contain:
|
||
internal users guest users mail contacts GroupMember.Read.All let me enumerate the members, but I initially only got object IDs and types back, with the useful fields set to null. The additional permissions were needed to resolve those member objects into usable addresses.
|
||
That let me build a safe wrapper around the raw send path:
|
||
read allowlist group membership normalize recipient addresses refuse to send if any recipient is missing from the group submit the message only if everything passes I still kept the Exchange transport rule in place.
|
||
That is the right split of responsibilities:
|
||
Group membership check = convenience and fast feedback
|
||
Transport rule = actual enforcement boundary
|
||
Why I Like This Design This setup gave me a few properties I care about a lot:
|
||
No user password stored on disk No human mailbox identity tied to the automation Mailbox access scoped to one mailbox Outbound mail constrained by Exchange itself Optional recipient allowlist that OpenClaw can read directly Full separation between “can submit mail” and “mail is actually deliverable” It is more work than creating a licensed mailbox and signing in like a user, but I think it is the right kind of work.
|
||
It moves the design away from “pretend this automation is a person” and toward “treat this automation like a service with explicit permissions and guardrails.”
|
||
Final Thoughts I’m trying to adopt OpenClaw slowly and deliberately. As I grant it access and add functionality, I’m trying to keep with the principle of least privilege and have active reporting going back to me when it starts acting hinky. Giving OpenClaw access to send mail as me is a step too far, but giving it a tightly controlled mailbox allows me to set up mail-based workflows without having to worry about it abusing that power.
|
||
`,summary:`I wanted to give OpenClaw its own mailbox in Microsoft 365 so it could read mail, send messages, and generally act like a useful automation agent without piggybacking on my personal account. The trick was doing it in a way that was secure, auditable, and constrained.
|
||
After working through the options, I landed on a model that I think is the right fit for this kind of use case:
|
||
Use a shared mailbox instead of a licensed user mailbox Authenticate with a Microsoft Entra app registration using a certificate, not a password Scope mailbox access with Exchange RBAC for Applications Enforce outbound safety with Exchange transport rules Optionally use a mail-enabled group as a preflight recipient allowlist This post walks through that setup and the reasoning behind it.`,date:"1 May, 2026",url:"https://nedinthecloud.com/2026/05/01/how-to-set-up-an-exchange-online-mailbox-for-openclaw/",image:"OpenClawMailbox.png",readingTime:"10"},"https://nedinthecloud.com/2026/04/29/actually-implementing-ai/":{title:"Actually Implementing AI",tags:["day-two-devops","AI","agentic-workflows","software-engineering","product-management","security"],content:`There is no shortage of AI content right now. Unfortunately, a lot of it lives somewhere between breathless hype and thinly disguised marketing copy. Every vendor claims AI will transform software delivery, replace half your team, and solve all your operational problems if you just buy the platform and turn the crank.
|
||
That is not how real implementation works.
|
||
In this episode of Day Two DevOps, Kyler and I spoke with Enrico Teotti, an independent consultant with deep experience in software development, product management, and practical AI adoption. What I appreciated most about Enrico’s perspective is that it was grounded in actual delivery work. Not theory. Not investor-speak. Not “agents will do everything.” Just a clear-eyed look at where AI helps, where it creates risk, and what teams need to change in order to use it responsibly.
|
||
The Most Useful AI Work Is Usually Pretty Boring One of the big takeaways from the conversation is that AI does not have to do something flashy to be valuable.
|
||
Enrico described a workflow where he gave an AI coding tool access to a monolithic application, a production data replica, and supporting observability tools. That combination let him ask questions that would normally require a bunch of manual stitching across systems:
|
||
Why are these records appearing in the wrong order? Why is this page slow for a specific workflow? What business logic is affecting the data beyond what the schema says? Is anyone actually using this feature we are about to spend time fixing? That is not the kind of demo that gets applause on a keynote stage. But it is exactly the kind of thing that makes teams faster and more effective.
|
||
The interesting part is not just that the AI could generate a SQL query. Plenty of tools can help with that. The real value came from combining multiple sources of context:
|
||
application code database structure and data analytics tooling performance data product understanding When those pieces come together, AI becomes less of a toy and more of a force multiplier.
|
||
Read-Only First, Then Maybe More Another theme that came through loud and clear was the importance of guardrails.
|
||
Enrico was very explicit that he does not trust AI systems enough to let them operate freely in production environments. He talked about using read-only access to production replicas instead of write access, and keeping a human firmly in the loop when investigating bugs or performance issues.
|
||
That matches my own experience. If you are going to let AI interact with real systems, the safest starting point is:
|
||
dedicated agent account read-only where possible narrow permissions explicit boundaries human review before meaningful changes This is one of those areas where teams get into trouble because the easy path is to hand the model your existing credentials and hope for the best. That is not a safety model. That is optimism wearing a trench coat.
|
||
The better approach is deliberate constraint. Give the tool enough access to be useful, but not enough to cause a disaster because it confidently made the wrong call.
|
||
Product Context Matters More Than Ever A lot of AI discussion still treats software development as if the hard part is typing code.
|
||
It is not.
|
||
Enrico’s product background gave him a useful lens on the whole conversation. He kept returning to questions like:
|
||
Why are we building this? What problem are we trying to solve? What does done actually look like? Should we fix this at all? That last question is especially important.
|
||
One of the traps AI creates is that it lowers the cost of making changes, which means it becomes dangerously easy to change things just because you can. If a tool helps you identify an issue in minutes, the natural temptation is to go fix it immediately. But if no one uses the feature, or the issue is not tied to an important workflow, then it may not deserve the effort.
|
||
This is where product thinking becomes even more valuable. Faster code generation does not eliminate the need for judgment. It increases the need for judgment.
|
||
In other words: if AI makes implementation cheaper, prioritization becomes more important, not less.
|
||
Curiosity Is the Real Career Moat Enrico made a blunt observation that I think will make some people uncomfortable: the people most at risk are not necessarily the people whose skills are obsolete, but the people who stop being curious.
|
||
That resonates. There’s a reason I use the title Curious Human. It encapsulates my outlook and reminds me to maintain my curiosity as I get older and more curmudgeonly.
|
||
The most effective use of AI I have seen does not come from blind trust or total resistance. It comes from people who are curious enough to experiment, skeptical enough to verify, and experienced enough to connect what the model says back to the actual business problem.
|
||
That combination is hard to automate, and it is probably the real skill set teams should be cultivating right now.
|
||
Final Thoughts If you are trying to sort signal from noise in the AI conversation, this episode is worth your time.
|
||
Enrico brought a pragmatic perspective that I think more teams need to hear. AI can absolutely help with debugging, analysis, implementation, and discovery. But it works best when it is surrounded by good constraints, strong product thinking, testing discipline, cost awareness, and people who know enough to challenge the output.
|
||
That may not be as exciting as the promise of fully autonomous software development. But it sounds a lot more like the real world.
|
||
And in the real world, that is usually where the value lives.
|
||
If you want to listen to the full episode, check out D2DO301: Actually Implementing AI. You can also find more of Enrico’s writing on his blog, which is well worth reading if you care about practical AI implementation instead of hand-wavy futurism.
|
||
`,summary:`There is no shortage of AI content right now. Unfortunately, a lot of it lives somewhere between breathless hype and thinly disguised marketing copy. Every vendor claims AI will transform software delivery, replace half your team, and solve all your operational problems if you just buy the platform and turn the crank.
|
||
That is not how real implementation works.
|
||
In this episode of Day Two DevOps, Kyler and I spoke with Enrico Teotti, an independent consultant with deep experience in software development, product management, and practical AI adoption.`,date:"29 Apr, 2026",url:"https://nedinthecloud.com/2026/04/29/actually-implementing-ai/",image:"DayTwoDevOps-EnricoTeotti.png",readingTime:"5"},"https://nedinthecloud.com/2026/04/25/commvault-cloud-expands-to-google-cloud-protecting-bigquery-and-beyond/":{title:"Commvault Cloud Expands to Google Cloud: Protecting BigQuery and Beyond",tags:["Commvault","GoogleCloudNext","GoogleCloud","CloudResilience","Sponsored"],content:`Your workloads are probably spread across multiple clouds. That’s nothing new, but what I hadn’t heard before is that 84% of organizations are using two or more cloud providers for critical services. That’s not shocking, but I also didn’t realize the number was that high.
|
||
Now you know me, I come from a cloud and infrastructure background. One of the things I quickly realized was the importance of having a single tool that worked across all the different environments I manage. It’s one of the reasons I like Terraform so much. You learn one tool and it works across any cloud or on-prem environment you find yourself responsible for.
|
||
I’ve certainly consulted at places that had a different tool for every environment and it was always a mess. One place I remember clearly- though I cannot mention the name- was using Commvault for their VMware environment, NetBackup for Windows bare-metal servers, and Veritas for their Linux workloads. That was three different backup products just for their on-premises stuff. They also had a small presence in AWS, and of course that was being protected using the native AWS tools.
|
||
This organization had four different data protection products, supported by four different teams, with zero consistency or transparency. Of course there were some organizational politics going on, and no one wanted to give up their tool of choice, but as an outsider looking in, I could see how inefficient and confusing the whole mess quickly became. Don’t even get me started on their DR runbook!
|
||
If I were telling that story today, the mess would be even worse. Back then “the cloud” meant a handful of EC2 instances. Now it’s BigQuery warehouses, S3 data lakes, Cloud Storage buckets full of training data, DynamoDB tables powering production apps, and every one of those is a potential blast radius if something goes wrong. The number of surfaces you need to protect has exploded, and the teams protecting them haven’t.
|
||
In my ideal world, all of those workloads could be protected by the same solution. And as new platforms arose, that solution could adapt and expand. That’s exactly what something like Commvault Cloud can do today.
|
||
For those who don’t know, Commvault Cloud is not your traditional Commvault deployment with media agents and a CommCell server that you host in your datacenter. It is a SaaS solution born in the cloud to support the dynamic nature of modern infrastructure, without you having to deploy a whole bunch of servers and storage to support it. Commvault cloud supports consumption-based billing, automated discovery, and protection across a wide-range of platforms.
|
||
When it first launched, you could spin up an instance of Commvault Cloud from the Azure or AWS marketplaces. Since it was SaaS, you weren’t responsible for managing the infrastructure underneath the solution, but the actual supporting services were running on either Azure or AWS. This week at Google Cloud Next, Commvault has expanded the solution to offer native support in the Google Cloud Marketplace as well. I’m always a fan of more choices, and if you need to ensure that all your data protection solutions reside in a particular cloud, Commvault now has you covered with the big three.
|
||
Commvault Cloud also supports protecting a ton of different workloads in Google Cloud: Compute Engine, Cloud SQL, BigQuery, and Spanner. With Clumio now joining the mix, Google Cloud Storage gets the same air-gapped, immutable-vault treatment that AWS customers have been using for S3. Personally, I think the support for BigQuery (and now GCS) is huge, especially with how much time and effort organizations are put into building AI platforms. According to Thomas Kurian, Google Cloud’s CEO, 60% of AI startups are using Google Cloud, and I’d bet most of them are housing training data in Cloud Storage and querying it from BigQuery. Losing either one would ruin their week, if not their quarter.
|
||
One feature of the solution I especially want to highlight is automatic discovery of unprotected resources. Back in the dark days when I was a VMware admin, automatic discovery of unprotected VMs was a godsend. The place I was working at allowed a lot of people to create virtual machines, and not all of them knew to put in a ticket to request data protection. After a couple of VMs got deleted by accident with no protection in place, we set up automatic data protection to prevent it from happening again.
|
||
My VMware environment was not that dynamic, so having automatic discovery was nice but not necessary. If you think about how fluid cloud environments are, automatic discovery and protection becomes an absolute essential. Commvault lets you tweak and customize the discovery to suit your needs, exempting some accounts, workloads, and services. I can’t imagine trying to manually protect assets in the cloud across the dizzying number of services available, and now I don’t have to.
|
||
Commvault Cloud has also branched out into supporting SaaS applications, which is not something I immediately think of when considering data protection. Usually I’m focused on protecting internally hosted applications and infrastructure, but given how prevalent the use of SaaS is- I mean Commvault Cloud is SaaS- it’s not surprising that Commvault sees that as a massive area of growth. If you’re using Google Workspace for your office applications, Commvault Cloud supports protecting that data as well.
|
||
When the cloud started becoming a real thing over a decade ago, I desperately wanted a single tool I could use for data protection. At the time, the landscape was an absolute hodge-podge of solutions with overlapping capabilities that never quite covered everything. Even if my unnamed and slightly shamed organization wanted to replace their four solutions, they simply could not have, at the time. Now it seems like Commvault Cloud, together with Clumio, could give them a single solution to cover all their data protection needs now and into the future (across AWS, Azure, and now Google Cloud), and grant a little consistency to their chaotic world.
|
||
If this has piqued your interest, check out this stuff from Commvault to learn more.
|
||
If you’re at Google Cloud Next this week, they’re at Booth #3617 running demos, including the new Clumio-for-GCS stuff in early access.
|
||
Thanks for reading and thanks to Commvault for sponsoring this post!
|
||
`,summary:`Your workloads are probably spread across multiple clouds. That’s nothing new, but what I hadn’t heard before is that 84% of organizations are using two or more cloud providers for critical services. That’s not shocking, but I also didn’t realize the number was that high.
|
||
Now you know me, I come from a cloud and infrastructure background. One of the things I quickly realized was the importance of having a single tool that worked across all the different environments I manage.`,date:"25 Apr, 2026",url:"https://nedinthecloud.com/2026/04/25/commvault-cloud-expands-to-google-cloud-protecting-bigquery-and-beyond/",image:"Commvault-GoogleCloud-Next.png",readingTime:"5"},"https://nedinthecloud.com/2026/04/14/open-source-malware-npm-and-the-risk-of-helpful-ai/":{title:"Open Source Malware, NPM, and the Risk of Helpful AI",tags:["day-two-devops","open-source","malware","npm","ai"],content:`I don’t think most practitioners spend a lot of time worrying about malware hidden inside an open source package. We worry about vulnerable code, sure. We worry about breaking changes, unplanned upgrades, and the occasional dependency rabbit hole. But malware? That still feels like something that happens to someone else, somewhere else, through an obviously sketchy email attachment.
|
||
Unfortunately, that’s not the world we live in anymore.
|
||
In an upcoming episode of Day Two DevOps, Kyler and I talk with Jenn Gile about the sharp rise in open source malware and why the problem is getting worse, not better. Jenn is the co-founder of Open Source Malware, and she has spent the last several years working in application security, platform operations, and open source ecosystems. She brought a mix of hard data, war stories, and practical advice that made one thing painfully clear: modern software supply chain risk is no longer just about bugs. It is also about malicious intent.
|
||
Why NPM Keeps Showing Up One of the first things Jenn explained is that more than 90% of tracked open source malware is showing up in NPM, and that most of that activity has happened recently. That is an alarming number, but it also makes a lot of sense when you think about how the JavaScript ecosystem works.
|
||
JavaScript applications tend to pull in a massive web of dependencies and transitive dependencies. You may choose one package intentionally, but that package can drag in dozens or hundreds more. In some cases, a seemingly simple project ends up relying on thousands of packages. That creates a very large attack surface, and most teams are understandably not doing deep manual vetting on every transitive dependency that shows up in the tree.
|
||
Jenn also pointed out that NPM has historically optimized for low friction. That is great for growth and terrible for security. When the barrier to publishing is low and account protections are weak, it becomes easier for bad actors to slip malicious packages into the ecosystem or compromise existing maintainer accounts and push poisoned updates.
|
||
Then there is the package manager behavior itself. Lifecycle scripts and post-install hooks can turn a compromised package into an automatic delivery mechanism. In other words, consuming the package may be all it takes to trigger the bad behavior.
|
||
And just because we focused on npm, that doesn’t mean other package managers are off the hook. Attacks on PyPi have been on the rise as well.
|
||
AI Is Making the Problem Weirder The part of the conversation that really stuck with me was Jenn’s explanation of how AI is changing the attack chain.
|
||
We tend to think of AI as a productivity layer sitting on top of our tools, but it is quickly becoming part of the operational environment. If an AI coding assistant or local agent is installed on a developer workstation and granted broad permissions, it becomes another thing an attacker can potentially abuse.
|
||
Jenn walked us through the example of the Nx compromise, where attackers were able to publish malicious versions of widely used packages. A post-install script then dropped a file that looked for locally installed AI tools such as Claude, Gemini, and Amazon Q. If it found them, it attempted to coerce those agents into becoming “helpful” to the attacker by using commands and flags that loosened or bypassed normal safety boundaries. Once that happened, the attacker could use the AI tools to help scrape secrets from developer machines.
|
||
That is a nasty evolution in the threat model. The AI is not the original compromise, but it can become a force multiplier once an attacker gets a foothold. The very thing that makes these tools useful—their broad access to local context, credentials, files, and workflows—also makes them dangerous when something goes wrong.
|
||
Jenn also noted that AI is helping attackers in more traditional ways. Phishing is more polished. Fake vendor emails look more convincing. Package campaigns are easier to generate and scale. Even malware analysis now reveals little fingerprints of AI-assisted creation, including things like suspicious emoji usage in malicious code.
|
||
The details are funny right up until they are not.
|
||
What Practitioners Can Actually Do Jenn offered practical advice that I think lands well with infrastructure and platform folks because it is realistic rather than absolute.
|
||
The first recommendation was to pin dependencies once you have decided they are safe. Not forever, and not blindly, but enough to avoid instantly consuming the newest release before anyone has had time to notice whether something is wrong.
|
||
The second recommendation was to introduce a cooldown period for new package versions. Waiting 24 to 72 hours before adopting a fresh release may sound annoying, but Jenn made a compelling point: malware is often a time game. Attackers want you to install quickly, before the ecosystem or community catches on and yanks the package. A one-day delay would have blocked some of the real-world incidents she described.
|
||
Of course, that introduces a tradeoff. Delaying updates can also delay vulnerability patches. But that is the job, isn’t it? Risk is almost never binary. It is a matter of deciding whether you are more worried about a known but not yet exploited bug, a breaking change, or a malicious package that intends to exploit you immediately.
|
||
For organizations, Jenn argued for a broader response than just handing developers another annual security training slide deck. Open source malware touches multiple teams:
|
||
Developers who choose and consume packages Application security teams that can automate checks and policy Incident responders who need to understand this style of compromise Threat hunters looking for signs that malicious code already landed Cyber threat intelligence teams tracking campaigns and infrastructure In other words, this is not only a developer problem. It is a cross-functional security problem.
|
||
Be a Skeptical Engineer If there was one phrase from the episode that I think deserves to stick, it is Jenn’s advice to be a skeptical engineer.
|
||
Do not assume a package is safe because it is popular. Do not assume a skill marketplace has done meaningful vetting. Do not assume an AI assistant is operating in a tidy sandbox. Do not assume only software engineers are at risk. If people in finance, marketing, sales, and operations are using AI tools to generate code or automate workflows, then your actual developer population is much larger than your org chart says it is.
|
||
That does not mean panic and uninstall everything.
|
||
It does mean slowing down a bit, checking what you install, being thoughtful about permissions, using sandboxing where you can, and recognizing that “helpful” AI is still software running with your access.
|
||
That is a lot to carry, but the episode did not end on total doom. Jenn talked about the researchers, maintainers, and providers who are actively working to identify, report, and take down malicious infrastructure. There is real collaboration happening behind the scenes. There are better controls coming. There are people paying attention.
|
||
That should make all of us feel at least a little better.
|
||
If you want to learn more, check out OpenSourceMalware.com and keep an eye out for Jenn’s episode of Day Two DevOps on April 15.
|
||
Written with help from AI
|
||
`,summary:`I don’t think most practitioners spend a lot of time worrying about malware hidden inside an open source package. We worry about vulnerable code, sure. We worry about breaking changes, unplanned upgrades, and the occasional dependency rabbit hole. But malware? That still feels like something that happens to someone else, somewhere else, through an obviously sketchy email attachment.
|
||
Unfortunately, that’s not the world we live in anymore.
|
||
In an upcoming episode of Day Two DevOps, Kyler and I talk with Jenn Gile about the sharp rise in open source malware and why the problem is getting worse, not better.`,date:"14 Apr, 2026",url:"https://nedinthecloud.com/2026/04/14/open-source-malware-npm-and-the-risk-of-helpful-ai/",image:"DayTwoDevOps-JennGile.png",readingTime:"6"},"https://nedinthecloud.com/2026/04/13/the-state-of-platform-engineering-and-devex/":{title:"The State of Platform Engineering and DevEx",tags:["day-two-devops","platform-engineering","devex","terraform","ai"],content:`Platform engineering can be a slippery term because it means different things depending on where you sit. For some teams it means golden paths, paved roads, and internal platforms. For others it means the group that owns all the infrastructure glue no one else wants to think about. And somewhere in the middle is developer experience, which may or may not be a separate function depending on the size and maturity of the organization.
|
||
In this episode of Day Two DevOps, Kyler and I spoke with Annem Shah about how she thinks about platform engineering, DevEx, infrastructure as code, and the growing role of AI in day-two operations. Annem is a cloud platform engineer with experience in government, consulting, and product-led engineering, and that range gives her a useful perspective on what good platform work actually looks like in practice.
|
||
What I appreciated about this conversation is that it never drifted into hand-wavy theory. Annem kept bringing it back to the practical realities of supporting engineers, choosing tools, and making systems easier to operate without locking everyone into bad decisions forever.
|
||
Platform Engineering Is Also Culture Work One of the strongest themes in the episode was that platform engineering is not just about tooling. It is also about creating a safe and useful interface between the platform team and the engineers using that platform.
|
||
Annem talked about setting up an infrastructure guild inside her organization as a way to spread knowledge across engineering teams. Rather than treating infrastructure as something mysterious or locked away behind a ticket wall, the guild creates a regular place for engineers to learn, discuss architecture, and work through system design ideas together.
|
||
I especially liked her use of an “engineering Kafka,” which is essentially a fictional architecture exercise. Instead of arguing over the current state of a production system and all the ego or baggage that can come with that, the team works through a hypothetical scenario together. In the episode, Annem described a sandwich shop with online ordering and lunchtime traffic spikes. From there, the group explores questions around scaling, latency, APIs, payments, and architecture choices.
|
||
That is clever for two reasons. First, it gives people a safe playground where there is no single correct answer. Second, it builds shared language and trust that teams can carry back into their real systems. That is culture work as much as platform work.
|
||
Internal Customers Are Still Customers Another theme that came through clearly is Annem’s view that software engineers, engineering leads, managers, and on-call responders are all internal customers of the platform.
|
||
That sounds obvious when you say it out loud, but not every platform team behaves that way. Plenty of organizations still operate with an old infrastructure-request mindset where one team throws tickets over the wall and another team grudgingly fulfills them. That model rarely creates a good developer experience.
|
||
Annem described a much healthier approach when her team onboarded a new CI/CD workflow tool. Rather than selecting a tool in isolation and forcing it on everyone, she worked directly with the engineers who would be using it most heavily. She gathered feedback on what mattered day to day, what features they relied on, and what would actually help them provision and manage infrastructure more effectively.
|
||
That kind of working group matters. If you have ever had a beloved tool ripped away and replaced by executive fiat, you know how much friction and resentment that creates. When engineers are invited into the decision-making process, they are much more likely to trust both the tool and the team introducing it.
|
||
AI Helps Most With the Ugly Middle When the conversation turned to AI, I think all three of us lit up a little bit because this is where platform engineering gets especially interesting.
|
||
There is a lot of marketing noise around AI writing code from scratch, building entire applications, or somehow eliminating the need for expertise. That is not the part of the workflow I find most compelling. What stood out in Annem’s examples was how AI helps with the ugly middle: imports, migrations, bootstrapping scripts, feedback loops, and deciphering inscrutable errors.
|
||
Annem talked about using AI assistants to accelerate the painful process of migrating infrastructure definitions and writing supporting scripts. She also described experimenting with MCP-style workflows and giving an AI assistant access to enough context to generate GitHub Actions pipelines that mimic the human-friendly experience of HCP Terraform.
|
||
That is a more grounded use case than “AI will run your platform now.” It is about reducing toil.
|
||
AI Might Make Platforms More Portable One idea I mentioned on the episode that probably warrants more thought is that SaaS vendors have a good reason to be nervous about AI.
|
||
Historically, a managed platform could keep customers partly because reproducing its behavior somewhere else required a lot of specialized time and effort. But if AI makes it cheaper and easier to recreate useful workflow patterns, generate migration scripts, and stitch together replacement automation, then some of that lock-in weakens.
|
||
That does not mean every managed service is doomed. Managed platforms still provide a lot of value, especially when it comes to reliability, support, and abstraction. But AI does appear to be shortening the distance between “we depend on this product” and “we can probably build enough of this ourselves.”
|
||
Platforms Are For Building If there was a single takeaway from this episode, it is that platform engineering is at its best when it combines thoughtful tooling, developer empathy, and a willingness to revisit earlier decisions. The job is not just to build the platform. The job is to make it usable, adaptable, and survivable on day two and beyond.
|
||
Platforms are for building after all. The shape of the platform dictates what can be built atop it. Being intentional and mindful of the platform accelerates development and empowers DevOps teams on day two and beyond.
|
||
Written with help from AI
|
||
`,summary:"Platform engineering can be a slippery term because it means different things depending on where you sit. For some teams it means golden paths, paved roads, and internal platforms. For others it means the group that owns all the infrastructure glue no one else wants to think about. And somewhere in the middle is developer experience, which may or may not be a separate function depending on the size and maturity of the organization.",date:"13 Apr, 2026",url:"https://nedinthecloud.com/2026/04/13/the-state-of-platform-engineering-and-devex/",image:"DayTwoDevOps-AnnemShah.png",readingTime:"5"},"https://nedinthecloud.com/2026/02/04/passing-the-ai-900-exam/":{title:"Passing the AI-900 Exam",tags:[],content:`As I mentioned in my planning for 2026 post, I wanted to get the AI-900 and AI-102 certifications from Microsoft as part of a larger goal of learning more about AI Engineering and being able to deliver Microsoft Training in that area. Good news! Last week I sat the exam for the Azure AI Fundamentals (AI-900) certification and passed. I thought I would share my experience and what I did to prepare.
|
||
Preparing for the Exam Let me start with what knowledge I had prior to studying for the exam. Like everyone else, I have been using AI for a while now and have been tangentially aware of how it all works. I knew that there were different types of models and that they used matrix math to find probabilistic solutions to queries. I had read about neural networks and how training could be supervised or unsupervised. And I was aware that generative AI was used to like generate things, and that the big advancement that led to modern generative models was transformers.
|
||
I guess what I’m saying is that I had a lot of general background knowledge regarding AI, just from being a technologist in the 2020s and reading news articles. But that knowledge lacked structure and was fragmented at best. The point of preparing for the exam was to fill in the gaps in my knowledge in a structured format.
|
||
What’s Being Tested You can go read up on the Microsoft Learn site yourself, but here’s a breakdown of the objectives being tested:
|
||
Objective Percentage Describe Artificial Intelligence workloads and considerations 15-20% Describe fundamental principles of machine learning on Azure 15-20% Describe features of computer vision workloads on Azure 15-20% Describe features of Natural Language Processing (NLP) workloads on Azure 15-20% Describe features of generative AI workloads on Azure 20–25% Although this is a Microsoft exam, a lot of the content was foundational to machine learning and AI models and not specific to Microsoft products. That being said, you absolutely should be familiar with the various Azure AI solutions and Microsoft Foundry. The study materials and exercises provided by Microsoft will cover all that ground, but I think you should spend a little extra time getting to know the various services, their capabilities, use cases, and how they tie together.
|
||
How I Studied I only used two resources to prepare for the exam. The first was the official, self-led training on the Microsoft Learn site. Over two weeks, I went through all 14 modules in the training and did all the exercises in each module.
|
||
Personally, some of the theory introduced early on in the course felt a bit esoteric. The Fundamentals of Machine Learning module gets into the various types of machine learning available, including regression, binary classification, multiclass classification, and clustering. The content gets into the weeds exploring the ways to enhance the predictive capabilities of regression models, like MAE, MSE, and RMSE. While this is interesting to know about, it isn’t really relevant to the certification. You are not going to be asked when it’s appropriate to use the Root Mean Squared Error or the Coeffecient of Determination for a regression plot. That is some deep data science stuff that is outside the scope of the certification. Is it interesting to read about? Yes. Do you need to understand it? No.
|
||
I did my best to dedicate about an hour a day to work on modules in the course. My goal was to complete at least one module a day, although that varied depending on the length of the module. I also avoided trying to cram too much into a single day, since I knew I would not retain the information if I did that.
|
||
I also made sure to do all the exercises in each module and not rush through them. To get to know the Azure AI services and platform, you really should take your time and explore during each exercise. Don’t limit yourself to the instructions. Being curious helped me fill in some gaps that weren’t covered by the training.
|
||
When I finished the course it was time to test my knowledge. The Microsoft Learn site has a sample quiz you can take as many times as you like and the questions are selected at random from a question bank. Each quiz run is composed of 50 questions that cover all the objectives of the course. There’s two things I noticed about the sample quiz:
|
||
Questions were often repeated in the same quiz run, sometimes one right after another. That seems like a programming issue on Microsoft’s part. Several questions had typos or grammatical issues. It’s pretty clear Microsoft is not prioritizing question quality in the sample quiz. Given the sorry state of the sample questions, what’s an enterprising student to do? Why, turn to AI of course!
|
||
When I was studying for the AZ-104 exam last year, I used VS Code with GitHub Copilot to create a quiz application and prompt for generating test questions. You can find the project here.
|
||
The quiz application itself is a simple webpage that loads a question bank from a JSON file you select. Once loaded, you can take the quiz in Practice Mode or Exam Mode. Practice mode simply means that each question is graded as you answer it, and exam mode waits till the end.
|
||
To produce the question bank, the repository has a sample prompt that you can tweak as needed. It also includes a script you can run that will generate the necessary prompt for you by asking questions about the exam you’re studying for. Once you have a solid prompt in markdown, you can then use the GenAI model of your choosing to generate the questions. I’ve had good luck with Claude Sonnet 4.5, but any model that supports MCP should work. You really do want MCP available so it can access the official documentation and study materials of the exam you’re preparing for.
|
||
You can ask it to generate questions for the entire exam, or focus on specific objectives. The prompt I used for the AI-102 exam is here. I used it twice to generate two questions banks with 50 questions each. What I really appreciate about the quiz is that it provides references for each question, so if I got something wrong I could go read more about it.
|
||
The combination of the official Microsoft training and the quiz grader app prepared me to take the exam and pass. I really do recommend using the quiz grader as a supplement. There were several questions it generated that included material or details that were not covered by the course, but did appear on the exam. I can’t tell you exact questions, only that the Microsoft course does have some gaps in its coverage.
|
||
Taking the Exam The exam itself is your standard online, proctored format through PearsonVUE. Despite a lengthy waiting period, the check-in process was pretty painless. I’ve done these exams before so I know how to set up my testing environment to meet their requirements. My preferred location is the kitchen table with the shades drawn and everything electronic moved or facing away from the testing environment. I also took the test during the day when no one else is in the house, thereby avoiding interruption.
|
||
I did make the mistake of scheduling the exam around lunchtime, which is probably a pretty busy time for the testing centers. I was 17th in the queue when I joined, and it took about 10-15 minutes for the queue to clear. Bear that in mind if you plan to sit the test and maybe sign-on in a bit early to maximize your time with the exam- you can sign in up to 30 minutes before your exam time.
|
||
It took me about 25 minutes to finish the exam and my final score was in the mid-800s. I can’t comment on the question content itself, other than to say that it aligned well with the objectives and the study materials. This is an entry level exam meant for technical and non-technical folks alike, so don’t overthink it.
|
||
Next Steps for Me Taking this exam served two goals:
|
||
Achieve the AI-900 certification so I can teach the class as an MCT Begin my journey to achieve AI-102 and beyond The natural next step is to start preparing for the Azure AI Engineer Associate track. I plan to follow the same study process as before, using the materials on the Microsoft Learn site and my quiz grader app for test questions. I may also take the AI-102 prep courses on Pluralsight, since I have free access as an author.
|
||
Beyond the AI-102 certification, I am also planning to pursue more AI Engineering knowledge through non-Microsoft material. Marina Wyss pointed me to the AI Engineer for Developers course on datacamp, and that looks pretty interesting. I’m not a developer by trade, but I think I can stumble through the Python enough to check this out.
|
||
`,summary:"As I mentioned in my planning for 2026 post, I wanted to get the AI-900 and AI-102 certifications from Microsoft as part of a larger goal of learning more about AI Engineering and being able to deliver Microsoft Training in that area. Good news! Last week I sat the exam for the Azure AI Fundamentals (AI-900) certification and passed. I thought I would share my experience and what I did to prepare.",date:"4 Feb, 2026",url:"https://nedinthecloud.com/2026/02/04/passing-the-ai-900-exam/",image:"Passing-the-AI-900-Exam.png",readingTime:"8"},"https://nedinthecloud.com/2026/01/07/beginning-ai-engineering/":{title:"Beginning AI Engineering",tags:[],content:`Recently I’ve become interested in AI Engineering as a discipline. As an infrastructure guy, data science and analysis has always fascinated me, but I felt woefully unprepared to work with it. When I was in college, I took a course on relational databases and it left my head spinning. Building out a relational table design, worrying about normal forms, and writing SQL queries was not my forte.
|
||
However, I have to reckon with the fact we live in a world awash with data. Being able to wrangle the data monster is an essential skill for any information worker, and I suppose that includes me. AI Engineering is a new-ish discipline in the realm of data science that focuses on solving practical problems through the use of AI models. Given it’s practical nature, AI engineering holds an attraction for me that more theoretical disciplines do not. What can I say? I like seeing where the rubber meets the road.
|
||
But where to start? Like I said, I’m not a data scientist. I’m not deeply steeped in multi-dimensional data cubes, data lakes, or data warehouses. I don’t live in Jupyter notebooks or spend days tinkering with Amazon Sagemaker. How can I get started in AI engineering and what do I need to know?
|
||
When in doubt, ask an expert, and that’s exactly what I did. In this episode of Day Two DevOps, Kyler and I talk to Marina Wyss about AI Engineering. Marina is a Senior Applied Scientist over at Twitch, and she’s well versed in the world of AI engineering and machine learning. She helped me understand what an AI engineer does and the challenges they face.
|
||
AI Engineering Defined Marina frames AI engineering as a distinct role from traditional machine learning or data science. Drawing on Chip Huyen’s definition, she explains that AI engineers primarily build products using pretrained models, often accessed via APIs, rather than training models from scratch. The core value of the role is not inventing new models, but selecting, adapting, and operationalizing existing ones to solve real product problems.
|
||
Where machine learning engineers and data scientists focus on data collection, labeling, and training custom models, AI engineers focus on:
|
||
Model selection (text, image, audio, multimodal) Prompt engineering Fine-tuning pretrained models Building reliable pipelines and applications AI engineers need to know enough about how models are developed and trained to leverage them successfully, but they aren’t going to be doing the actual training themselves. They need to understand the problem space, the goals of the applications, and what inputs will be available. Based on that information, they can take an off-the-shelf model and tweak it using some combination of RAG, fine-tuning, and prompt engineering to adapt it for a particular application.
|
||
Knowledge and Tooling One of the things I loved about the episode was discovering tools I’ve never heard of before. I live in a world of Terraform, Kubernetes, and YAML. Marina casually brought up Airflow and Dagster and I had to stop her and ask what those things are- turns out they’re tools to manage data pipeline orchestration. It would appear I’ve got some reading and experimenting to do.
|
||
Beyond the pipeline orchestration tools, she also stressed the need to be comfortable in Python and SQL, along with the public cloud offerings around machine learning. Since Twitch is part of Amazon, she is using AWS Sagemaker for a lot of her work, but really any of the major public clouds have ML tools as a service.
|
||
Learning how to use these tools effectively will require a ton of work on my part, and that’s something else Marina stressed. She didn’t become an AI engineer overnight, it was a continual process that involved self-study and off-hours experimentation. We didn’t get a chance to cover it, but Marina has a couple of excellent videos that review AI engineering courses and good starting points for folks of different backgrounds. Fortunately, I’m not start from scratch. I have a CS degree and a good grasp on data fundamentals, what’s missing is training on machine learning and model tweaking.
|
||
My Plan for 2026 I really want to dig deeper into the world of AI engineering and build something in the process. I’ve started my journey by taking the Microsoft Learn courses around their AI-900 certification. Once I’m done that, I plan to start studying for the AI-102 certification for AI engineers. I don’t always enjoy certifications as a way to learn, but in this case I believe it’s the right approach. Marina actually just released a video on AI certifications and she gets the value of certs spot-on. She actually recommends the AI-102 Azure AI Engineer Associate cert or the AWS and Google Cloud equivalents.
|
||
Once I get those two certifications, I plan to take one of the courses Marina recommended, preferably one that involves building out an actual application. I’ve always been a hands-on learner and I’ll be more invested if it’s a project I care about. Ultimately, I’m not looking to get a job as an AI Engineer, just expand my sphere of knowledge and potentially be able to provide training in the discipline sometime in the future. I do have my MCT now, and I bet that the courses teaching Azure AI engineering are in high demand.
|
||
`,summary:`Recently I’ve become interested in AI Engineering as a discipline. As an infrastructure guy, data science and analysis has always fascinated me, but I felt woefully unprepared to work with it. When I was in college, I took a course on relational databases and it left my head spinning. Building out a relational table design, worrying about normal forms, and writing SQL queries was not my forte.
|
||
However, I have to reckon with the fact we live in a world awash with data.`,date:"7 Jan, 2026",url:"https://nedinthecloud.com/2026/01/07/beginning-ai-engineering/",image:"D2DO291-Artwork.png",readingTime:"5"},"https://nedinthecloud.com/2026/01/02/planning-for-2026/":{title:"Planning for 2026",tags:[],content:`Plans. What are these silly things we make? Each year I write a planning post for the coming 12 months for the sole purpose of having something to laugh at when the annum ends. As this has become a tradition, I see no reason to stop with 2026. What bold and completely wrong plans shall I make for the coming year? Stay tuned and find out.
|
||
Things to Accomplish Doing things has never been a problem for me. I cannot stop doing things. I’m always thing-a-ning. In fact, I often agree to do too many thing-a-nings and immediately regret my decisions. Since the boundary of December 31st is artificial at best, I’ve already promised to do many thing-a-nings in 2026. What have I got cooking?
|
||
YouTube Videos I’ve already committed to producing four more livestreams for Terrateam in promotion of their ebook, From Bare Metal to Cloud. I really like the ebook and the conversations I’ve already had with Malcolm Matalka and Amber Britton have been great.
|
||
In addition to those livestreams, I’d also like to bring back my Terraform Tuesday content in a big way starting in June of 2026. I used to publish 2-3 a month, but that dropped way off in 2025 due to my workload with Pluralsight. By June I should be done my current crop of promised courses, and then I’d like to spend some time focused on building the YouTube channel again.
|
||
Terraform is not the only thing I’m interested in. I’d love to dig into other topics. Here’s a few that I’m thinking about:
|
||
WASM - A perennial favorite, but this year I am seriously going to try and dig into WebAssembly and use it with my cluster of Raspberry Pis that are just sitting around. Maybe I can do something with local AI inferencing? I do have an overabundance of web cams and other random hardware. Kubernetes - I haven’t touched K8s much in the last couple years. It’s time to immerse myself back in the technology and build something interesting. Maybe I can combine that with the WASM thing too. OTel - OpenTelemetry was kind of a big deal at KubeCon NA. I think I need to dig more into the topic and understand it holistically. Overall it sounds like what I need to build is some type of AI inferencing project that uses K8s, WASM, and OTel on my home lab gear. That’s the sort of thing that could keep me busy and produce some interesting content!
|
||
Day Two DevOps I’d really like to work on building the audience for Day Two DevOps in 2026. We have amazing guests and informative content, but how do we get that in front of the ears and eyes of the people? I’d like to break this into a few steps:
|
||
Creating interesting content - This is already the case. Kyler and I have awesome conversations with interesting people. Packaging the content - D2DO is primarily distributed through traditional podcasting channels. We are adding full-video episodes in 2026, but I think we also need to focus on creating short form bites. We’re talking audiograms, video shorts, short articles, and blog posts. Each recorded episode need to be repackaged into at least 10 assets. Distributing the content - Again, D2DO is primarily released using traditional podcast channels along with an audiogram on the Packet Pushers YouTube channel. We need to be where the people are, and that’s not just LinkedIn. If I want D2DO to grow and spread, I think I need to build out a presence on TikTok, Insta, and Bluesky. (I think that’s also the correct order of priority). I don’t know a ton about any of these platforms, so I think I need to start with TikTok and work from there. Track stats - How can you tell if your changes are helping if you can’t measure them? While we do track download numbers and video watches, I think I need more robust analytics to track what is working and what isn’t across all the platforms. The goal is two-fold. First, and most importantly, I want to help people advance in their tech career by providing useful information. If we’re not doing that, then everything else is pointless. Secondly, I want to attract more sponsors to help support Day Two DevOps, so I can continue to make this good stuff. The audience will always come first, but sponsorships are ultimately what keeps the whole enterprise afloat.
|
||
Pluralsight Courses For 2026, I’ve already committed to producing nine more courses for Pluralsight. I need to finish the last two courses that are part of the Terraform learning path. Once those are complete, I still owe them a Vault Associate (003) course that is all about preparing for the exam. That’s going to be a really short one, and I should be able to bang it out in a couple weeks.
|
||
After that, I am going to create a series of courses that are solely focused on the Terraform Associate Certification Version 004. That version dropped at the end of 2025, and is now available for exam scheduling. There have been some significant updates from Version 003 of the exam, as detailed here, and so it make sense to produce a series on Pluralsight that covers all the objectives.
|
||
How is this different from the Terraform learning path I’ve been working on for the past year? The Terraform learning path is not focused on the Associate exam, and goes well beyond the content tested for the certification. If you are purely interested in learning Terraform from the ground up and advancing to a professional level, I think the learning path is for you.
|
||
The certification path is going to be completely focused on the exam objectives, and they will be explicitly mapped to courses to help you decide which courses you need to take to fill in any knowledge gaps or refresh your memory. The certification courses don’t have a storyline or anything like that, they are meant to be quick, informative, and focused on the certification. If you are looking to pass the associate exam and get your cert, then I think the certification path makes more sense.
|
||
Ultimately, it’s about giving the learner choices. You aren’t locked into one path of another. You could easily build your own path from the collection of courses out there.
|
||
Once I finish those Terraform certification courses, I’m not sure what’s next in the Pluralsight world. I’ll probably sign on to do some Azure-related courses, as those have been good to me in the past. Or maybe it’s time to create some HashiCorp Boundary content. I have to see what is planned by Pluralsight and what I feel like doing.
|
||
Certification Guides I have two certification guides published: Terraform Associate and Vault Associate. Both of the guides need to be updated for versions 004 and 003 of the certifications respectively. My goal in 2026 is to get both of these guides updated.
|
||
More than that, I also need to promote these guides. I do virtually no promotion of these at all, and they still do decent sales. Imagine if I actually let people know they exist? What a concept!
|
||
Live Instruction The good news is that HashiCorp’s training partner, LearnQuest, is scheduling classes every month. There are currently three classes that I can teach:
|
||
Terraform Foundations Vault Enterprise Operational Infrastructure as Code That last class is a brand new one, and it’s a replacement for the old Terraform Advanced class. The new class includes the use of HCP Terraform, HCP Packer, and other advanced tooling and workflows that go well beyond the standard Terraform Foundations course. I highly recommend checking it out.
|
||
My current plan is to try and teach one class a month. I’m also involved with another training partner for private Terraform instruction. There will probably be 2-3 classes delivered for them through the year as well.
|
||
Beyond that, I’d like to start delivering 1-2 Microsoft classes a quarter. I’m now an official MCT, so I can do that. The only certification I have right now is AZ-104, so that’s the only class I can teach. I think I’m going to try and get a few more certifications this year, so I can teach a broader range of classes. Maybe I’ll focus on the Azure AI stuff? I bet that’s pretty popular. I think I’ll try and get the AI-900 and AI-102 certifications. That will force me to dig back into Python, and build some actual applications.
|
||
Website and Blog In 2025, my blog was incredibly quiet. Just like my YouTube channel, certification guides, and other projects, the overwhelming demands of Pluralsight took all of my writing mojo. This year I will make an effort to publish twice a month. It shouldn’t be too hard. There’s at least two Day Two DevOps episodes each month that would make for good blog posts, plus anything else I happen to be working on.
|
||
I do have a tendency to make blogging more complicated than it needs to be, writing 3500 words when 1500 would suffice. What can I say? Part of writing for me is the processing of working through my thoughts and ideas. What I really need to do is clean up a bit afterwards. It’s like I cooked you a fine dining meal, but also took you grocery shopping, walked you through the prep, and forced you to watch me cook. You just wanted a nice meal, you didn’t need all that extraneous stuff.
|
||
This might be one place that LLMs could help me out. Take my ridiculous ramblings and make them slightly more concise. That’s what a good editor does, and while I think LLMs aren’t very good authors, I think they can make excellent editors.
|
||
Tech Field Day and Gestalt IT I’ve already signed up to go to Cloud Field Day 25 in March of 2026. Beyond that event, I expect to do a few paid gigs for Gestalt IT over the course of the year.
|
||
While I don’t relish flying out to California or the red-eye flight back, I’ve found my time in the Tech Field Day community invaluable. They keep me connected to interesting vendors, spark my interest in new topics, and challenge my current ideas. It wouldn’t be a stretch to say that Ned in the Cloud wouldn’t exist as it does today without Tech Field Day.
|
||
Conferences In 2025 I went to HashiConf, KubeCon, and IBM TechXchange. I’ll probably do the same in 2026. Although, KubeCon is in SLC and I have my own issues with Utah. Instead of KubeCon, I may switch things up and attend re:Invent. For all my misgivings about Las Vegas, re:Invent continues to be the place where the maximum number of tech people I like congregate. That alone makes it worthwhile.
|
||
I’d also like to attend DevOpsDays Philly again. I missed it in 2025 due to scheduling conflicts, but hopefully that won’t happen in 2026.
|
||
Other Opportunities Are there other things popping up on the horizon? Sure thing. There’s a cloud native company that wants me to start building out some content for them in 2026, so I’ll be reaching out in January to start the planning process.
|
||
Beyond that, there is another O’Reilly project that I don’t have many details on. It has something to do with Linux, and that’s all I know at the moment.
|
||
If the past is any indication, there will be several other opportunities coming across my desk in the next 12 months. As always, the challenge is deciding which ones make the most sense to engage with and which ones I should pass on.
|
||
Conclusion 2025 didn’t exactly go as planned, but no year ever does. One of my pillars is to embrace discomfort, and the unpredictability of the future will always result in discomfort. In it’s own strange way, I find that thought comforting in and of itself.
|
||
I have no doubt 2026 will be full of challenges, opportunities, and failures. As long as I approach it with an open mind and a clear sense of purpose, I’m sure 2026 will be a great year!
|
||
`,summary:`Plans. What are these silly things we make? Each year I write a planning post for the coming 12 months for the sole purpose of having something to laugh at when the annum ends. As this has become a tradition, I see no reason to stop with 2026. What bold and completely wrong plans shall I make for the coming year? Stay tuned and find out.
|
||
Things to Accomplish Doing things has never been a problem for me.`,date:"2 Jan, 2026",url:"https://nedinthecloud.com/2026/01/02/planning-for-2026/",image:"Planning-for-2026.png",readingTime:"10"},"https://nedinthecloud.com/2025/12/30/2025-year-in-review/":{title:"2025 Year in Review",tags:[],content:`Welcome to the annual review for Ned in the Cloud! Believe it or not, this is the eighth installment of the “Year in Review” series. Frankly, I’m impressed with myself for having kept this up for almost a decade. Going back to that first post, my goals for 2018 are hauntingly familiar:
|
||
Build a successful cloud practice Learn to program in Go or Python (or both!) Keep creating new Pluralsight courses Blog on a weekly basis Podcast on a weekly basis Speak at more user groups and conferences Granted, that first bullet point Build a successful cloud practice doesn’t really apply anymore. In 2018, I was working at a VAR and trying to build up our cloud services offerings. I quit that job in 2019 to work for myself. Everything else? Still applicable! I supposed the fact that I am a full-time content creator and technical educator makes things like creating content and speaking in public a pretty consistent set of goals.
|
||
If the theme of my 2024 was acceptance, and the theme of 2025 was commitment. In some ways, I had left the revenue generating portions of the Ned in the Cloud business on autopilot while I tried out new things. In 2025, it was time to refocus on building revenue and updating my aging content. As we’ll see shortly, I ended up doing a LOT.
|
||
High-level Review I won’t share actual revenue numbers, but Ned in the Cloud saw a revenue drop of 19% YoY. Since 2022, revenue has dropped every year. If I were a publicly traded company, people would be selling my stock and talking about hostile take-overs. But the thing is, despite these revenue drops, I am still making a very comfortable living from the company. As it’s only employee and sole owner, I really only have to answer to myself. Am I still happy with what I’m doing and making enough money?
|
||
Yes, yes I am.
|
||
What accounted for another drop in revenue? The biggest factor was Pluralsight revenue, with a drop of 8%. Day Two DevOps revenue also dropped in half! I’ll address each of these in their own sections.
|
||
The overall lessons are these:
|
||
Pluralsight revenue is pretty important and I’ve produced a ton of new courses this year to course correct (pun intended) Podcast sponsorships are harder to acquire due to an abundance of supply My YouTube channel needs some TLC Let’s dig into these topics and more in detail.
|
||
Pluralsight Courses Since the biggest chunk of my revenue comes from Pluralsight, I should probably start with that. First, a little context about how Pluralsight works for the uninitiated. When producing courses for Pluralsight, there are two primary ways to get paid. There is an upfront payment for completing a course, and a quarterly revenue share for the course based on viewership.
|
||
The general idea is that Pluralsight brings in X amount of money in revenue for subscriptions. Of that revenue pool, they determine how much of that revenue is from video courses. That revenue is divided by the total number of hours watched on the platform. You get a percentage of that revenue based on how many hours your course was watched and your revenue share percentage, which is negotiated when you sign up to build the course.
|
||
The bottom line? The more popular the course - i.e. the more hours it is watched; the more money you make from a course. The more popular Pluralsight is, the more revenue they generate, and the bigger the revenue pool is to distribute to authors.
|
||
I’m sure I’ve explained this before, but I won’t make you navigate to another link.
|
||
A few years ago, Pluralsight made two big changes. First, due to financial issues, they cut all author’s revenue by 25%. That was a sweeping cut across the board for all authors. Was I pissed? Hell yeah, I was pissed. You can read my feelings in this post.
|
||
The other big change was a lifetime earnings cap for a single course. The general idea was learners want fresh and new courses, and there were authors who had created a popular course many years ago and have been coasting on the revenue ever since. Pluralsight felt this was unfair to new authors, as these old courses were taking a significant share of the revenue pool without contributing back.
|
||
If we’re being totally honest, I think Pluralsight just wanted to force these authors to create a new version of the course with a lower revenue percentage. You can dress it up however you want, but the simple truth was authors created content for Pluralsight. People watch those courses on Pluralsight as part of their subscription. Pluralsight is making money from those courses, and they want to limit or eliminate how much authors were earning based on an arbitrary cap.
|
||
From a business perspective, I get it. Pluralsight has been through a lot in the last 5 years. Going public, going private, TechEd market consolidation, multiple rounds of layoffs, etc. It has been a bumpy ride financially and organizationally. The fact that they are looking to cut costs is not surprising, and the author rev share is a big cost, especially for courses that have been around for years and received no updates.
|
||
From a moral perspective, I don’t think it’s the right thing to do. These courses are still delivering value to Pluralsight, shouldn’t the authors continue to get paid?
|
||
The lifetime earning limit is actually pretty high, and only the most popular courses are likely to hit it. If you don’t have a course in the top 100, then there’s nothing to worry about. But you know who does have a course in the top 100 and wasn’t paying close attention? THIS GUY.
|
||
My course, Terraform - Getting Started, had been in the top 100 courses pretty much since I created the first version in 2017. I have refreshed the course about every 2 years ever since: v2 in 2019, v3 in 2021, and v4 in 2023. The v4 course was the first version affected by the lifetime earning limit, and wouldn’t you know it, in spring of 2025 I hit the cap.
|
||
There was no klaxon, no warning, no email or reminder. I looked at my quarterly payment and it was way off. That’s how I found out. Part of that is on me. I should have been keeping a tally on my own and tracking how close I was to the limit. But I also think it’s pretty fucked up that no one at Pluralsight thought to set up an early warning system for courses that are approaching the limit. For a single course that represents 50% of my income from Pluralsight, having that amount drop to $0 without warning was quite a shock.
|
||
I was already in the middle of finishing the Vault Associate Certification courses and couldn’t really stop to update the Terraform - Getting Started course. I had contractual deadlines I had already agreed to. So, I basically lost out on three months of revenue from my most popular course due to this cap. You want to know why my Pluralsight revenue was down 8%? This is why.
|
||
And it would have been much worse if I hadn’t kicked my course production schedule into overdrive this year. In 2025, I created 18 new courses for Pluralsight. That’s not a typo, EIGHTEEN new courses. That was nine courses for the HashiCorp Vault Associate certification, eight courses for the Terraform learning path, and one course for the foundations of computing.
|
||
Every single one of those courses had an upfront completion payment, and it was those payments that kept me from having a far larger dip in my revenue.
|
||
Entering into Q4 of 2025, things had stabilized. My new Terraform - Getting Started course was in place and the older version was retired. The course itself is shorter and I’ve taken topics from it and my Terraform - Deep Dive courses and spread them across the new Terraform learning path. It’s unlikely I’ll hit the lifetime learning cap for the course as quickly, although I am now keeping my own tally just in case.
|
||
I do not expect to keep up this frantic pace of course creation in 2026, although I have agreed to create courses for the updated Terraform Associate (004) certification. To be completely honest, I kinda let Pluralsight sit on cruise control for all of 2024, and in 2025 I had to make up for that.
|
||
YouTube Videos You may have noticed that my YouTube video output dropped off severely in 2025. We’ll talk about actual stats momentarily, but essentially my focus on Pluralsight course creation kind of sucked all the air out of the room for other projects. I still managed to create a couple of videos around Terraform, but most of what I posted was interviews and livestreams. Things that do not require me to write a script or create a demo.
|
||
Don’t get me wrong, I was still writing a ton of scripts and demos. I was just writing them for Pluralsight and not YouTube. I have all my scripts in markdown files, and the total word count is about 350,000 words. I wrote 350k words about HashiCorp Vault and Terraform in 2025. I’m like the Brandon Sanderson of HashiCorp. Where’s my goddamn Shardblade?
|
||
ANYWAY
|
||
My point is that I had to shift my focus to Pluralsight to bring my revenue number back up. YouTube simply doesn’t pay well enough, even with sponsored content. Now that I’ve righted the Pluralsight ship, I can return my attention to creating YouTube stuff about Terraform and other things.
|
||
In terms of my top performing videos released in 2025, they are as follows:
|
||
Title Release Date Total Views No More Secrets in State! Write-Only Args in Terraform 1.11 7/1/2025 2940 Terraform, OpenTofu, and the Future of IaC 4/22/2025 2836 Terraliths: Breaking Up Is Hard To Do 4/2/2025 2457 Mastering Data Transformation in Terraform 1/14/2025 2223 The Future of Terraform - Announcements from HashiConf 2025 9/27/2025 1560 What I glean from this is that people are most interested in practical Terraform demos and use cases, not theoretical stuff. That’s always been true.
|
||
My top videos overall for 2025 are mostly from previous years:
|
||
Title Release Date Total Views Terraform Basics: Modules 6/1/2021 12,652 Four Years Later: Is Terragrunt Worth It? 11/27/2024 6,453 Managing Multiple Environments with Terraform 10/7/2023 5,055 What is PXE Boot? 11/19/2020 3,541 Exploring the Import Block in Terraform 1.5 7/12/2023 3,235 I still find it hilarious that my What is PXE Boot? video from five years ago continues to rack up the views. It was before I upgraded all my A/V equipment and put together a professional looking YouTube set. It’s just me, sharing what I’d learned about PXE boot because I was trying to get it set up in my home lab.
|
||
Just goes to prove, you never know what will be successful, but 101 content is always a safe bet.
|
||
As my overwhelming workload from Pluralsight dies down, you’ll begin to see more YouTube content being published. I really want to dig into Terraform Actions and Terraform Search. There’s also the Terraform Basics video series that could use some love. Stay tuned.
|
||
Livestream on MAIN At the end of 2024, my buddy Marino Wijay posted about doing a learning livestream about AI. I said I’d be game to do one together, and thus Livestream on M.A.I.N. was born. The M.A.I.N. stands for Marino, AI, and Ned. I thought I was being clever. I’ll leave you to make your own judgement on that account.
|
||
Starting on January 16th, we began our weekly livestream to talk and learn about AI. The idea was pretty simple. Let’s integrate AI into an application and use AI to help us write it. In the process we can learn more about AWS Bedrock, Python, and Agentic AI. For the first seven episodes, that was all we worked on. Building out a podcast processing applications that used Amazon Bedrock, Lambda, and Python to process MP3 files from a podcast episode.
|
||
As Marino and I started getting busy with work things, we invited guests to come on and talk about their experiences building AI stuff. We had my Day Two DevOps cohost Kyler Middleton chat about the AI Slack bot she’s building for work. Later we brought on John Capobianco, Nnenna Ndukwe, Rizel Scarlett, Kat Morgan and others to talk about what they’re doing in AI. I learned a ton in the process and I hope you did to.
|
||
As for the application we were building? That never really got past the basic functionality stage, but it served its purpose. Marino and I learned a ton about how Bedrock functions and how to build out Lambda functions to interact with the service and different models.
|
||
Our last episode was on September 11 with Cat Hicks. We didn’t realize it was going to be the last episode, but I suppose we shouldn’t have been surprised. Conference season was getting into full swing, and Marino and I were super busy with our own stuff.
|
||
Will Livestream on MAIN come back in 2026? Maybe? Honestly, Marino and I haven’t talked about it and I’ve been too in the weeds with Pluralsight stuff to say for sure. It might actually be a good fit for the Packet Pushers network, which doesn’t currently have an AI focused show. Speaking of which…
|
||
Day Two DevOps Day Two DevOps continues unabated for a sixth full year! We have racked up over 2.2M downloads since the podcast started in January 2019 and there’s no sign of stopping. While the focus of the podcast has changed somewhat over the years- as cloud computing has become less of a new idea and more of a given- we continue to discuss practitioner focused challenges.
|
||
Going through the episode downloads for 2025, I wanted to do a fair comparison to determine which episodes were most popular, so I used the number of downloads after 90 days. That necessarily means anything released after September won’t make the list. I’ll include them as contenders for 2026.
|
||
Episode Title Guest Downloads in first 90 days D2DO282: Simplifying Complex Kubernetes Deployments With kro Islam Mahgoub 3349 D2DO271: Public Vs. Private Cloud In 2025 Mark Boost 3035 D2DO268: Solving Big Problems By Solving Small Problems Merritt Baer 2955 D2DO276: MCP: Capable, Insecure, and On Your Network Today Dan Barr 2945 D2DO278: The Future of HashiCorp Inside IBM Armon Dadgar 2930 Looking at this list and the other episodes that attracted the most interest, I have a few conclusions:
|
||
People are still looking for explorations of deep technical topics focused on the practitioner Security is still a major concern in 2025, especially with the rise of AI AI by itself is not all that interesting As Kyler and I plan for episodes in 2026, I think we need to focus less on AI for AI’s sake and more on practical solutions and dissecting practitioner problems. If you have a solution, problem, or technology you’d like us to explore, be sure to reach out.
|
||
Starting in 2026, Day Two DevOps will begin publishing full video for each episode as well. We’ve been capturing the video for a while and using some of it for YouTube Shorts, but didn’t have the necessary resources to edit and publish full episodes. Packet Pushers has recently hired two new full-time employees to assist with editing, and all the shows will begin publishing video.
|
||
That’s actually part of a larger trend with Packet Pushers expanding out to twelve podcasts in the network. It includes amazing shows like N is for Networking, The Cloud Gambit, and Life in Uptime. It’s exciting to be part of a growing network!
|
||
Unfortunately, the general explosion of podcasts across the industry has led to a drop in sponsorships for Day Two DevOps. There’s so many shows out there, all vying for a piece of the advertising pie while marketing budgets seem to be relatively static. The good news is that Packet Pushers is also bringing on a full time outbound sales person, something that is sorely needed with so many shows.
|
||
Outbound sales is not one of my core strengths. Hell, I’m not even very good at inbound sales, i.e. people coming to me asking to sponsor the show. I am intrinsically bad at sales. While at KubeCon this year, I heard Ethan pitch Packet Pushers to potential sponsors and I realized how unprepared I am for that portion of things. This is why you hire people. To do things you’re not good at, and honestly don’t want to be.
|
||
When it comes to Day Two DevOps, I think we’ve got the right hosts, the right topics, and a great reputation. I need someone else to tell that story better than I can.
|
||
Chaos Lever Much like my YouTube presence, Chaos Lever fell victim to the over-abundance of work I had on projects that actually make money. Chaos Lever has always been a labor of love and has never earned any real money for me and Chris. We did it because we both like talking about technology and talking to each other.
|
||
In June of 2025, we decided to put Chaos Lever on hiatus for the summer. That hiatus has extended into the fall and now winter. We’ve been doing the podcast on a weekly basis for over three years and I just didn’t have the time to keep doing it. In 2024, I hired Humble Pod to help me edit and promote Chaos Lever, hoping that we would build up a big enough presence to attract sponsors. That didn’t happen and I had to cut Humble Pod loose due to costs.
|
||
As I said then, Humble Pod was doing an amazing job and it is no fault of theirs that Chaos Lever didn’t become an overnight success. At a certain point I just have to face the fact that two dudes rambling about technology into microphones is just not that unique or fascinating in 2025. There are a ton of other podcasts out there doing something similar, and we simply haven’t won the attention economy lottery. Maybe we should try being fascist douchebags and kissing the Trump Administration’s ass? That seems to be working for a lot of bro-pods.
|
||
Then again, I also have to be able to live with myself, so maybe not.
|
||
While we’re on hiatus, I am trying to pitch Chaos Lever to a few publishing companies to see if they would want to adopt us like the adorable little puppies that we are. Realistically, this is the only way I see the show coming back in the near-term.
|
||
Training Classes After Pluralsight, the biggest part of my revenue comes from delivering live training. That’s right! You could have me as an instructor if you want to learn about HashiCorp Vault or Terraform and sign up for the right class.
|
||
For the last several years, I have been teaching classes through River Point Technology, delivering both custom private trainings and public trainings through HashiCorp Academy. With the acquisition of HashiCorp by IBM, all HashiCorp Academy training is now delivered through IBM’s partners. That means I had to negotiate with a new entity to keep teaching those classes.
|
||
All public classes were put on hold during the transition, so I basically didn’t teach any Academy classes for three months. But I’m glad to say that things are back on track and I’ve taught two classes for this new company. They seem to be offering the public classes once a month, giving me the opportunity to apply to teach up to three classes every month if I wanted to. That’s a bit much, so I’ll be sticking to about one class a month or so.
|
||
I also had another training company reach out to me about teaching private Terraform classes. So far, I have taught two classes for them and there’s a distinct possibility that is going to ramp up a lot in 2026.
|
||
This year I finally decided to become a Microsoft Certified Trainer (MCT) as well. This makes is possible for me to teach Microsoft classes through training partners. I just got the MCT badge in December, and I plan to start looking into training partners this January.
|
||
Out of all the revenue areas for Ned in the Cloud, training is the only one that grew year over year. It makes sense for me to put more time and energy into the thing that is actually growing at the moment.
|
||
Books I’m very excited to say that in 2025, I became an official O’Reilly author. I can’t even remember what the first O’Reilly book was that I read, but I can tell you they have been integral to my career development. Fifteen years ago, I would have never thought that I would be an O’Reilly author too, but last year that changed with a simple LinkedIn message.
|
||
One of their Acquisitions Editors reached out to me about a sponsored guide Microsoft was looking to do around running Linux on Azure. She wanted to know if I’d be interested in writing the guide. I said yes, on the condition that I could have Chris Hayner co-author it with me. Of the two of us, he’s more versed in Linux and I’m deeper into Azure. Between us, I thought we could churn out a pretty good book. O’Reilly gave the OK and we were off to the races.
|
||
I guess we did a decent job, because before the book wrapped, they asked if we wanted to write another guide about picking a Linux flavor. That’s about 75% done at this point, with Chris doing the lion’s share of the work while I tried to get my Pluralsight stuff done on time.
|
||
In 2026, they have another project already lined up for us, but we don’t have a ton of details yet. My hope is to keep growing this relationship and eventually publish a book with a cool animal on it.
|
||
Outside of the O’Reilly stuff, it appears that I need to update the certification guides for both Vault Associate (003) and Terraform Associate (004). I helped write and review the questions for the Terraform Associate (004) exam and I’ll be creating six courses on Pluralsight about it, so updating the guide shouldn’t be very hard.
|
||
Other Appearances Did I do other things in 2025? Yeah probably. Here’s a short list of stuff I can recall:
|
||
HashiTalks 2025 Cloud Field Day 22 HashiCorp Ambassador Thought Leader Series Platform and Cloud Cafe IaC Conf Iac Conf Live The Cube at KubeCon HashiCorp Ambassador Roundtable Edge Field Day Showcase Okay, wow, so it was a lot of other things. Funny. You forget how busy you actually were while in the thick of things.
|
||
Deprecations and Failures This is a fun section with a fun name. What did I fail at in 2025? And I don’t mean that in a super-negative way. One of my three pillars is embracing failure, so it’s time to give my failures a big ole hug!
|
||
YouTube Channel I had planned in the beginning of 2025 to publish Terraform Tuesday content three times a month and also explore some other topics, like Home-labbing and WASM. Needless to say I utterly failed on all counts. If you’ve read this far, you know I sacrificed building the channel in lieu of creating Pluralsight content. No regrets on that front, but still counts as a failure.
|
||
Chaos Lever 2025 was supposed to be a year that Chaos Lever grew up and started making some money. While I feel like the quality of the show never faltered, we still weren’t able to monetize it. Considering how much time it took up in my week to write, edit, and produce; I couldn’t justify continuing with the show due to other priorities that actually make money. Still a failure though.
|
||
WebAssembly Another year has passed and I am no further in my process to use or really understand WebAssembly. Maybe 2026? I wouldn’t bank on it.
|
||
Certification Guides I meant to update the Vault Associate guide this past year, and never quite got around to it. Now that the new version of Terraform Associate is out, it’s time for me to update both. Why didn’t I get the Vault one done? Once again, Pluralsight strikes. I did create a ton of courses for the Vault Associate (003) exam, so I am well suited to update the guide.
|
||
Looking Forward What’s coming for 2026? That’s what the Planning for 2026 post is for silly! Just as a preview though, I’m planning to finish about 10 Pluralsight courses, start delivering some Microsoft Training, and renew my focus on the YouTube channel.
|
||
`,summary:`Welcome to the annual review for Ned in the Cloud! Believe it or not, this is the eighth installment of the “Year in Review” series. Frankly, I’m impressed with myself for having kept this up for almost a decade. Going back to that first post, my goals for 2018 are hauntingly familiar:
|
||
Build a successful cloud practice Learn to program in Go or Python (or both!) Keep creating new Pluralsight courses Blog on a weekly basis Podcast on a weekly basis Speak at more user groups and conferences Granted, that first bullet point Build a successful cloud practice doesn’t really apply anymore.`,date:"30 Dec, 2025",url:"https://nedinthecloud.com/2025/12/30/2025-year-in-review/",image:"2025-Year-in-Review.png",readingTime:"20"},"https://nedinthecloud.com/2025/07/21/vault-provider-and-ephemeral-values/":{title:"Vault Provider and Ephemeral Values",tags:["hashicorp","terraform","azure"],content:`In my last post about ephemeral values and write-only arguments, I showed how you could use the updated AzureRM provider to leverage the new azurerm_key_vault_secret ephemeral resource with write-only arguments that have been added to other resources. But the AzureRM provider is not the only provider to introduce ephemeral resources and write-only arguments! Version 5 of the Vault provider also has these new capabilities, as does the Random provider. In this post I thought we could explore bother providers to see how they have implemented ephemerality.
|
||
Vault Provider Support Version 5.1.0 of the Vault provider has two ephemeral resources in it: vault_kv_secret_v2 and vault_database_secret. The vault_kv_secret_v2 is meant as a replacement for the current data source of the same name, and in fact if you are using the data source in your configuration, you’ll now get deprecation message:
|
||
│ Warning: Deprecated Resource │ │ with data.vault_kv_secret_v2.burrito_recipe, │ on main.tf line 61, in data "vault_kv_secret_v2" "burrito_recipe": │ 61: data "vault_kv_secret_v2" "burrito_recipe" { │ │ Deprecated. Please use new Ephemeral KVV2 Secret resource \`vault_kv_secret_v2\` instead Don’t worry if you’re currently using the vault_kv_secret_v2 data source, it’s still available for now. I’d expect it to be retired in the next major version of the provider.
|
||
You can use the new vault_kv_secret_v2 ephemeral resource to retrieve a secret value from a KV V2 secrets engine on a Vault server. The value will not be written to state or to plan files. Ephemeral value rules apply, so you can only use the value with a write-only argument or another ephemeral value type.
|
||
Let’s check out an example of using the ephemeral resource to copy a secret value from Vault to Azure Key Vault:
|
||
ephemeral "vault_kv_secret_v2" "burrito_recipe" { mount = "burrito_secrets" name = "burrito-recipe" version = 1 } # Write the secret to Azure Key Vault resource "azurerm_key_vault_secret" "write_only" { name = "burrito-recipe" value_wo = jsonencode(ephemeral.vault_kv_secret_v2.burrito_recipe.data) value_wo_version = 1 key_vault_id = azurerm_key_vault.example.id } During plan, Terraform will retrieve the secret value from Vault, but will not include it in the execution plan file. When apply is run, Terraform will retrieve the secret value again, and this time write it to Azure Key Vault. Critically, the value_wo argument is not written to state, so the actual value is not recorded in state. Yay!
|
||
Version Confusion There are two version related arguments that I want to tease apart here, as I think they could be a point of confusion. The version argument in the vault_kv_secret_v2 ephemeral resource refers to the version of the secret you want to retrieve from Vault. I could potentially have several versions of the burrito-recipe secret stored in Vault, so the version argument tells Vault which one to get. If I omit the version argument, the Vault provider will retrieve the latest version.
|
||
The value_wo_version argument in azurerm_key_vault_secret has nothing to do with the version of the secret pulled from Vault. The purpose of the vault_wo_version argument is to let Terraform know when the value of the secret has changed. Since Terraform doesn’t have the current value of the secret stored in state, it has no way of knowing if the secret value retrieved from Vault is different from what is stored in Azure Key Vault. The value_wo_version argument value IS recorded in state, and when it changes, Terraform will overwrite the current Azure Key Vault value with whatever is retrieved from Vault.
|
||
The number set for version and value_wo_version do not have to be the same. And in fact, there is a version property of the Key Vault secret that can also be different. You could have the following:
|
||
Vault Secret Version Azure Key Vault Value WO Version Azure Key Vault Secret Version 7 2 3 That would be totally valid. You are retrieving version 7 of the Vault secret and writing it as version 3 of the Azure Key Vault secret, using the value_wo_version argument set to 2 to make Terraform overwrite the current value.
|
||
Confusing? Yes. What can you do? Here’s my humble recommendation.
|
||
Set the version argument for vault_kv_secret_v2 and the value_wo_version to the same value. If the latest version of your Vault secret is 7, then use that value for both resources. That way you can know exactly what version you’re pulling from Vault, and you can use a single input variable (var.secret_version) to set both arguments, like so:
|
||
ephemeral "vault_kv_secret_v2" "burrito_recipe" { mount = var.secret_mount_path name = var.secret_name version = var.secret_version } # Write the secret to Azure Key Vault resource "azurerm_key_vault_secret" "write_only" { name = "burrito-recipe" value_wo = jsonencode(ephemeral.vault_kv_secret_v2.burrito_recipe.data) value_wo_version = var.secret_version key_vault_id = azurerm_key_vault.example.id } When it’s time to update to version 8 of the secret, you only have to change one input value to retrieve the correct secret version and tell Terraform to update the Key Vault secret.
|
||
What about the version of the Key Vault secret? That doesn’t really factor into this deployment. Someone else is consuming the secret, and it’s up to them how they want to manage the version being retrieved.
|
||
If you want a full example to take for a spin, check out this folder in my Terraform Tuesdays repository. The hashi_vault configuration will set up Vault with a secret, and the hashi_vault_access will retrieve the secret and write it to an Azure Key Vault.
|
||
I fully expect more ephemeral resources to be added to the Vault provider over the next few release cycles. In fact, there are already several logged issues and pull requests to add them. If the ephemeral resource type you want is not yet supported, keep your eye on future releases. Maybe even watch the repository?
|
||
Random Provider Version 3.7.0 of the Random provider added the random_password ephemeral resource. This is kind of a weird one. The syntax is straightforward enough:
|
||
ephemeral "random_password" "password" { length = 16 special = true override_special = "!#$%&*()-_=+[]{}<>:?" } The use case is less clear. Why do I say that? Bear in mind that Terraform doesn’t store the resulting value in state or plan. But unlike the vault_kv_secret_v2 or azurerm_key_vault_secret ephemeral resources, the value of the random_password isn’t being retrieved from a persistent source.
|
||
Every time the random_password ephemeral resource is invoked the value returned will be different. If I use this resource to set the database administrator password for an azurerm_mssql_server, I will have no idea what the value was actually set to. Maybe that’s good if I never plan to log on as the administrator and just need the password set to something sufficiently complex. In fact, no one would know the value, so I guess that’s a bonus.
|
||
I suppose you could use the random_password to set a value of a Key Vault secret, and then use that Key Vault secret to set the password for the Azure MSSQL server. That would solve the secret zero problem of how a password is generated before it’s stored. When you need to roll the password, just change the value_wo_version argument for the Key Vaule secret. The random_password ephemeral resource is creating a different value every time it’s invoked.
|
||
Okay, so maybe I have convinced myself there’s some utility here. I suppose we’ll have to wait and see.
|
||
Conclusion As more providers implement ephemeral resources and write-only arguments, the ability to keep sensitive data our of state will steadily rise. The write-only version argument is going to become increasingly important. You should start thinking now about how you want to manage that value for different use cases.
|
||
`,summary:"In my last post about ephemeral values and write-only arguments, I showed how you could use the updated AzureRM provider to leverage the new azurerm_key_vault_secret ephemeral resource with write-only arguments that have been added to other resources. But the AzureRM provider is not the only provider to introduce ephemeral resources and write-only arguments! Version 5 of the Vault provider also has these new capabilities, as does the Random provider. In this post I thought we could explore bother providers to see how they have implemented ephemerality.",date:"21 Jul, 2025",url:"https://nedinthecloud.com/2025/07/21/vault-provider-and-ephemeral-values/",image:"VaultProviderEphemeral.png",readingTime:"6"},"https://nedinthecloud.com/2025/07/15/write-only-arguments-in-terraform/":{title:"Write-only Arguments in Terraform",tags:["hashicorp","terraform","azure"],content:`Ephemeral Values and Write Only Arguments Ephemeral resources were introduced in Terraform 1.10, but they were missing a key component. Write-only arguments are a new feature in Terraform 1.11, and they help to complete the goals of ephemeral resources.
|
||
Ephemeral Values I’ve already written a whole post about ephemeral values, but allow me to summarize it here as well. Ephemeral values were created to help address the issue of sensitive values being stored in state and plan files. State data holds the attribute values of all data sources and resources, and some of those attributes might be sensitive data you’d rather not appear in state.
|
||
An additional concern is sensitive information found in saved Terraform plan files. Plan files not only contain the entirety of the configuration, but also the planned changes and variable values used to create the plan.
|
||
Values that are marked as ephemeral will never be persisted in state or in a plan file. The value is accessed and stored in memory for the period during which it is needed, and then it is flushed. You can get ephemeral values from input variables, output values, or ephemeral resources. Input variables and output values can be marked as ephemeral by setting the ephemeral argument to true. Ephemeral resources are a distinct block type and must be implemented by a provider.
|
||
Use Cases for Ephemeral Values So how can you use ephemeral values? When the feature was released in 1.10, the primary use cases were for provider settings and provisioner credentials. That’s because the values used for providers and provisioners are not persisted in the plan file or in state. What about resource arguments though? By default, all resource arguments are recorded in the plan and state, so you could not use ephemeral values, that is until the introduction of write-only arguments.
|
||
Write-Only Arguments Write-only arguments are a special type of resources argument that is not persisted in state and not written to the plan file. You can write a new value to the resource, but you cannot retrieve the current value through state data. Now you can’t simply mark an argument in a resource as write-only, instead a new argument needs to be added to the resource type in the provider to support the write-only functionality.
|
||
For example, version 4.34.0 of the azurerm provider includes a write-only argument for the value of a Key Vault Secret, and a write-only password argument for several database related resources like azurerm_mssql_server. As new releases of the provider roll out, additional resources will have write-only arguments added to them.
|
||
You might be wondering, if Terraform doesn’t know the value of the write-only argument, how does it know when that value changes? The answer is that along with the new write-only argument is a version argument to signal to Terraform that the current value should be overwritten.
|
||
For instance, consider the following azurerm_mssql_server resource:
|
||
resource "azurerm_mssql_server" "db" { name = local.name resource_group_name = azurerm_resource_group.db.name location = azurerm_resource_group.db.location version = "12.0" administrator_login = "mssqladmin" administrator_login_password_wo = var.db_password_info.value administrator_login_password_wo_version = var.db_password_info.version } The write-only password argument is called administrator_login_password_wo and the version argument is called administrator_login_password_wo_version. That seems to be the pattern for all write-only arguments I’ve seen. The write-only argument has the suffix wo and the corresponding version argument has the suffix wo_version.
|
||
Key Vault Example Why don’t we walk through a few examples so you can see these write-only arguments in action? If you want to play along, the code for these examples is stored in my Terraform Tuesdays repository.
|
||
Consider a configuration where we want to store an Azure Key Vault secret value. Before write-only arguments, the resource block would look like this:
|
||
resource "azurerm_key_vault_secret" "normal" { name = "db-password-normal" value = var.db_password_regular key_vault_id = azurerm_key_vault.example.id } After deploying the configuration, we can inspect the state data for the Key Vault secret, and sure enough, the value of the secret is stored in plaintext.
|
||
{ "mode": "managed", "type": "azurerm_key_vault_secret", "name": "normal", "provider": "provider[\\"registry.terraform.io/hashicorp/azurerm\\"]", "instances": [ { "schema_version": 0, "attributes": { "value": "T@COSsoG00dY'All", ... That’s no good! If someone gains unauthorized access to my state data, they can find out my innermost secrets.
|
||
Sidebar: I know that some people want to protect state by encrypting it locally before pushing to the remote backend. I guess that’s fine if your state backend is compromised, but anyone with permissions to run a plan would be able to view the encrypted state data. Better to not have your sensitive data in there to begin with.
|
||
Now, let’s check out the same resource type, but with write-only arguments:
|
||
resource "azurerm_key_vault_secret" "write_only" { name = "db-password-wo" value_wo = var.db_password_ephemeral value_wo_version = var.db_password_version key_vault_id = azurerm_key_vault.example.id } Instead of setting the value argument, I’m using the value_wo argument. When you use that argument, you also have to include the vault_wo_version argument. The input variable for the value_wo is marked as ephemeral, so I don’t accidentally use it somewhere else in the configuration.
|
||
variable "db_password_ephemeral" { description = "The ephemeral database password to be stored in the Key Vault." type = string sensitive = true ephemeral = true } variable "db_password_version" { description = "The version of the database password to be stored in the Key Vault." type = number } If I try and use db_password_ephemeral with a regular argument or in a non-ephemeral context, I’ll get back the error:
|
||
Ephemeral values are not valid for "value", because it is not a write-only attribute and must be persisted to state. In my terraform.tfvars I have the following values set:
|
||
prefix = "sopes" db_password_regular = "T@COSsoG00dY'All" db_password_ephemeral = "T@COSsoG00dY'All" db_password_version = 1 Obviously, I wouldn’t want to store the secret values in a terraform.tfvars file, as that would defeat the purpose of hiding them from state. But for the purpose of demonstration, we’ll let that one go. You’d probably want to use an environment variable to pass the value in a production context.
|
||
If I check out the state data for my new Key Vault secret:
|
||
{ "mode": "managed", "type": "azurerm_key_vault_secret", "name": "write_only", "provider": "provider[\\"registry.terraform.io/hashicorp/azurerm\\"]", "instances": [ { "schema_version": 0, "attributes": { "value": "", "value_wo": null, "value_wo_version": 1, The value field is empty, and the value_wo has a value of null. Our secret value is no longer stored in state!
|
||
Updating Write-only Arguments Now what if I wanted to update the secret? I’ll change the value for the write-only secret in my terraform.tfvars file:
|
||
prefix = "sopes" db_password_regular = "T@COSsoG00dY'All" db_password_ephemeral = "BUrr!t0sR@lso" db_password_version = 1 Then I’ll run a terraform plan:
|
||
... No changes. Your infrastructure matches the configuration. Terraform has compared your real infrastructure against your configuration and found no differences, so no changes are needed. Because I didn’t change the value of the value_wo_version argument, Terraform has no way of knowing that the value of the secret should be updated. And sure enough, we get back a plan with no changes.
|
||
Now I will update the db_password_version to 2, and run a new plan:
|
||
Terraform will perform the following actions: # azurerm_key_vault_secret.write_only will be updated in-place ~ resource "azurerm_key_vault_secret" "write_only" { id = "https://sopes-ephemeral-83610.vault.azure.net/secrets/db-password-wo/582c7570578c42dd80d0c6badf2a938b" name = "db-password-wo" tags = {} ~ value_wo_version = 1 -> 2 # (8 unchanged attributes hidden) } Plan: 0 to add, 1 to change, 0 to destroy. Since the version has changed, now Terraform knows it needs to update the value as well, and the plan we get back has one change listed.
|
||
A few quick notes about the value_wo_version argument.
|
||
First, the value for the argument must be a whole number greater than zero. So it might make sense to update the input variable with some validation:
|
||
variable "db_password_version" { description = "The version of the database password to be stored in the Key Vault." type = number validation { condition = var.db_password_version > 0 && var.db_password_version == floor(var.db_password_version) error_message = "The db_password_version must be a non-negative integer." } } Second, I thought I would be clever and combine the password value and version into a single input variable:
|
||
variable "db_password_ephemeral" { description = "The ephemeral database password to be stored in the Key Vault." type = object({ value = string version = number }) sensitive = true ephemeral = true } But this marks the entire input variable as ephemeral. When you try to use the version value with the value_wo_version argument, Terraform throws an error because the argument is not write-only and the value is ephemeral. So until you can mark individual attributes of an input variable as ephemeral, you’ll need to keep the value and version separate.
|
||
Using Ephemeral Resources with Write-only Arguments The first example showed how we might get a secret value into Key Vault without storing it in state data. What about using the secret? That’s where ephemeral resources come into play.
|
||
In a new configuration, I have an ephemeral block for a Key Vault secret:
|
||
ephemeral "azurerm_key_vault_secret" "db_password_ephemeral" { name = var.db_password_info.key_vault_secret_name key_vault_id = var.db_password_info.key_vault_id } The syntax of the block is exactly the same as the azurerm_key_vault_secret data source, but the block keyword is ephemeral instead of data.
|
||
To use this ephemeral resource, I have an azurerm_mssql_server using the argument administrator_login_password_wo and administrator_login_password_wo_version:
|
||
resource "azurerm_mssql_server" "db" { name = local.name resource_group_name = azurerm_resource_group.db.name location = azurerm_resource_group.db.location version = "12.0" administrator_login = "mssqladmin" administrator_login_password_wo = ephemeral.azurerm_key_vault_secret.db_password_ephemeral.value administrator_login_password_wo_version = var.db_password_info.version } Just like when we created the Key Vault secret, we need to include a version here so Terraform knows if it needs to update the login password, since that value isn’t stored in state.
|
||
Let’s see what happens when I run a plan:
|
||
$ terraform plan ephemeral.azurerm_key_vault_secret.db_password_ephemeral: Opening... ephemeral.azurerm_key_vault_secret.db_password_ephemeral: Opening complete after 3s ephemeral.azurerm_key_vault_secret.db_password_ephemeral: Closing... ephemeral.azurerm_key_vault_secret.db_password_ephemeral: Closing complete after 0s During the plan process, Terraform opens the contents of the ephemeral resource in memory. Then it uses that information with the MSSQL server, and once the plan is complete, the ephemeral resource is closed. Terraform will do the same thing during apply, accessing the ephemeral resource, using the value for creation, and closing the resource.
|
||
The planned creation of the azurerm_mssql_server.db also shows the write-only nature of the attribute:
|
||
# azurerm_mssql_server.db will be created + resource "azurerm_mssql_server" "db" { + administrator_login = "mssqladmin" + administrator_login_password_wo = (write-only attribute) + administrator_login_password_wo_version = 1 Feel free to deploy this configuration yourself if you want to see the resulting state data.
|
||
Unsupported Resources The addition of write-only arguments is absolutely fantastic and I’m excited to see more resource types implement them. But what if the resource type you want to manage isn’t supported? There is a solution! The AzAPI provider.
|
||
My fellow HashiCorp Ambassador, Stu Mace, has written a whole blog post on using the AzAPI provider and its sensitive_body argument. Give it a read and a share!
|
||
Conclusion The introduction of write-only arguments closes the loop on the ideas first introduced with ephemeral values in Terraform 1.10. Between write-only arguments in azurerm resources and the sensitive_body in azapi resources, you can now keep your secrets out of Terraform state entirely. And I think that’s a good thing!
|
||
If I had one gripe, it’s the management of the new version argument to signal when the secret value changes. As a community, we’ll need to develop some guidelines and best practices on managing the version value effectively.
|
||
`,summary:`Ephemeral Values and Write Only Arguments Ephemeral resources were introduced in Terraform 1.10, but they were missing a key component. Write-only arguments are a new feature in Terraform 1.11, and they help to complete the goals of ephemeral resources.
|
||
Ephemeral Values I’ve already written a whole post about ephemeral values, but allow me to summarize it here as well. Ephemeral values were created to help address the issue of sensitive values being stored in state and plan files.`,date:"15 Jul, 2025",url:"https://nedinthecloud.com/2025/07/15/write-only-arguments-in-terraform/",image:"WriteOnlyArguments.png",readingTime:"9"},"https://nedinthecloud.com/2025/07/01/ephemeral-values-in-terraform/":{title:"Ephemeral Values in Terraform",tags:["hashicorp","terraform","azure"],content:`Terraform 1.10 has introduced the concept of ephemeral values for outputs and input variables, and as a new object type. Before I dig into the details of how they work, first I think it’s important to understand the problem ephemeral values are trying to solve.
|
||
We all know that Terraform state can contain sensitive information. After all, it contains all the attributes of each resource and data source and any output values defined in the configuration. Terraform uses these values to calculate an execution plan when updates are made to the configuration.
|
||
An additional concern is sensitive information found in saved Terraform plan files. Plan files not only contain the entirety of the configuration, but also the planned changes and variable values used to create the plan.
|
||
How can you remove sensitive information from state data and saved plans? That is what ephemeral values are meant to do. You can mark input variables and output values as ephemeral, and there is a brand new object type: ephemeral resources.
|
||
Unlike data sources or managed resources, an ephemeral resource and its attributes are never written to state or an execution plan. Additionally, ephemeral resources and their properties cannot be used in the context of a non-ephemeral object. If they could, then the values of the ephemeral resource would end up in state or a plan, and would defeat the purpose behind them.
|
||
Input variables and outputs marked as ephemeral have similar restrictions. You can only use an ephemeral value for an argument that will not be recorded in state.
|
||
Given that you can’t use an ephemeral value inside of a typical managed resource or data source, what good are they? Well there are a few use cases:
|
||
Inside a provider block Inside a provisioner block With write-only arguments (new in Terraform 1.11!) We will explore write-only arguments in their own blog post, so for now let’s focus on the first two use cases.
|
||
Ephemeral Values Syntax Input variables and output values can be marked as ephemeral simply by adding the ephemeral argument and setting it to true:
|
||
variable "identity_token" { type = string description = "Token used for Azure authentication." sensitive = true ephemeral = true } The syntax for an ephemeral resource is basically the same as a managed resource:
|
||
ephemeral "<ephemeral_resource_type>" "<name_label>" { <identifier> = <expression> } Ephemeral resources are a new object type in Terraform 1.10, and they appear as a separate category in the provider documentation. Often they mirror existing data source types, but the implementation is different. Currently, the AzureRM, AWS,GCP, and Vault providers have all added support for ephemeral resources, with more coming in the future.
|
||
The current version (4.35.0) of the azurerm provider supports the following ephemeral resource types:
|
||
azurerm_key_vault_secret azurerm_key_vault_certificate If you have a secret stored in Key Vault that you want to use, the syntax for the ephemeral resource would be:
|
||
ephemeral "azurerm_key_vault_secret" "example" { name = var.key_vault_secret_name key_vault_id = var.key_vault_id } Now what can you do with this secret? Can you use it in a regular resource?:
|
||
resource "azurerm_container_group" "example" { name = "ephemeral-continst" location = azurerm_resource_group.example.location resource_group_name = azurerm_resource_group.example.name #... container { name = "hello-world" image = "mcr.microsoft.com/azuredocs/aci-helloworld:latest" cpu = "0.5" memory = "1.5" secure_environment_variables = { KEY_VAULT_SECRET = ephemeral.azurerm_key_vault_secret.example.value } } } Nope! When you run terraform validate you’ll get the error:
|
||
Error: Invalid use of ephemeral value │ │ with azurerm_container_group.example, │ on main.tf line 36, in resource "azurerm_container_group" "example": │ 36: KEY_VAULT_SECRET = ephemeral.azurerm_key_vault_secret.example.value │ │ Ephemeral values are not valid in resource arguments, because resource instances must persist between Terraform phases. To generate a proper execution plan, Terraform would have to persist the secret value beyond the planning phase. Since we’re trying to avoid that, Terraform throws an error.
|
||
So if we can’t use ephemeral values with regular resource or data source arguments, where can we use them? You can use them with the follow object types:
|
||
Local values: The local value will inherit the ephemeral marker. Provider blocks: You can use ephemeral values to configure the properties of a provider. Ephemeral resources: You can reference the attribute of one ephemeral resource in another since neither persists. Ephemeral variables: You can pass an ephemeral value to a child module if the input variable is marked as ephemeral. Ephemeral outputs: You can pass an ephemeral value to a parent module if the output is marked as ephemeral. Provisioner and connection blocks: You can use an ephemeral value to configure a provisioner. Write-only arguments: You can use ephemeral values with resource and data source arguments that have been marked write-only. Note that ephemeral outputs cannot be used in the root module, because then they would be written to state.
|
||
Also, since ephemeral values are not written to the plan file, they are calculated on each Terraform run. So you might get a different value for plan than you do for apply. That’s a pretty important thing to note. If your ephemeral value is non-deterministic, the value may be different between the plan and apply.
|
||
Using with Providers Honestly, the biggest use case with the release of Terraform 1.10 was provider configuration. Say for instance, you are using the Kubernetes provider to deploy a manifest and you have the Kubernetes credentials stored in Azure Key Vault secret. You could grab the secret with an ephemeral resource and pass it to the Kubernetes provider block:
|
||
ephemeral "azurerm_key_vault_secret" "k8s" { for_each = toset(["client-certificate", "client-key", "cluster-ca-certificate"]) name = each.key key_vault_id = var.key_vault_id } provider "kubernetes" { host = "https://\${var.aks_cluster_host}" client_certificate = base64decode(ephemeral.azurerm_key_vault_secret.k8s["client-certificate"].value) client_key = base64decode(ephemeral.azurerm_key_vault_secret.k8s["client-key"].value) cluster_ca_certificate = base64decode(ephemeral.azurerm_key_vault_secret.k8s["cluster-ca-certificate"].value) } Because the provider configuration is never stored in state or the plan, this is a valid use of an ephemeral resource. You can also use meta-arguments with an ephemeral resource, as shown above with the for_each argument. Ephemeral resources also support:
|
||
depends_on count for_each provider lifecycle The other potential use case is for provisioners, but here I would remind you that provisioners are considered an anti-pattern and not recommended. If you must though, you could run a remote-exec every time a member is added to a cluster like this:
|
||
ephemeral "azurerm_key_vault_secret" "example" { name = var.key_vault_secret_name key_vault_id = var.key_vault_id } resource "terraform_data" "cluster" { triggers_replace = azurerm_linux_virtual_machine.cluster[*].id connection { type = "ssh" user = "clusteradmin" private_key = ephemeral.azurerm_key_vault_secret.example.value host = azurerm_linux_virtual_machine.cluster[0].private_ip_address } provisioner "remote-exec" { inline = [#...] } } Provisioners also do not write their connection information to state or plan.
|
||
The introduction of ephemeral values in Terraform 1.10 is a huge step towards getting secrets out of state and plan, but there is one key thing missing that closes the loop with resources and data sources. Write-only arguments. I’ll address those in my next post.
|
||
`,summary:`Terraform 1.10 has introduced the concept of ephemeral values for outputs and input variables, and as a new object type. Before I dig into the details of how they work, first I think it’s important to understand the problem ephemeral values are trying to solve.
|
||
We all know that Terraform state can contain sensitive information. After all, it contains all the attributes of each resource and data source and any output values defined in the configuration.`,date:"1 Jul, 2025",url:"https://nedinthecloud.com/2025/07/01/ephemeral-values-in-terraform/",image:"EphemeralValues.png",readingTime:"6"},"https://nedinthecloud.com/2025/03/03/on-hashicorp-ibm-and-acceptance/":{title:"On HashiCorp, IBM, and Acceptance",tags:[],content:`I originally started writing this post as a detailed history of my relationship with HashiCorp, how it grew as a company, and its IPO in December 2021- which happened at the exact wrong time. After laying down about 750 words, I deleted the whole thing. Why? Because that’s not the post I wanted to write. This one is.
|
||
A Year Of Acceptance For me, 2024 was a year of acceptance. I had to accept a lot of things I didn’t necessarily like. Ned in the Cloud had seen two consecutive years of dropping revenue. Pluralsight had gone IPO, been acquired, and now was changing hands again, with people being laid off in the process and author payments being slashed by 25%. HashiCorp had changed to Business Source Licensing and now was being acquired by IBM. And Donald Trump was amazingly not in jail, but rather our next president.
|
||
It was A FUCKING LOT. I spent 2024 grieving, working through my feelings, and slowly approaching acceptance. I’m not a religious person, but I appreciate the serenity prayer:
|
||
Grant me the power to change the things I can, to accept what I cannot, and the wisdom to know the difference.
|
||
I can’t change most of the things that impacted me in 2024. Donald Trump is our first felon president. Pluralsight will never be the company it once was. HashiCorp is owned by IBM. Admittedly, some of these things are worse than others- read: Trump.
|
||
I can’t change these things, but I CAN control how I react and think about them. 2024 was a year of coming to terms with what life was throwing at me, and working the mental gymnastics necessary to find a way through.
|
||
2025 is about moving forward.
|
||
Misaligned Expectations There’s this idea in manufacturing of a digital twin. It’s a virtual model of a real world system that you can use to test out changes and monitor reactions. The digital twin is limited by its fidelity, i.e. how closely it mirrors its real world counterpart.
|
||
Our mind also carries a digital twin of what we perceive, and it has the same struggles with fidelity. I assumed Americans would reject a convicted felon for president. I was wrong. My mental model of the electorate was flawed.
|
||
When a mental model- our digital twin- has an impedance mismatch with reality, we cannot entirely blame reality for being what it is. Our mental model was incorrect, probably due to a lack of information or incorrect assumptions. What we have to do is recalibrate that model by questioning our assumptions and gathering more accurate information.
|
||
That’s not to say that a lack of fidelity is our fault. People can actively lie to you or intentionally withhold useful information. I’m once again tempted to reference Donald Trump, but I think I’ve adequately flogged that dead horse. If you don’t get that he is a lying, liar, who lies about everything all the time, I doubt this is the post that will help you.
|
||
Instead, let’s talk about HashiCorp.
|
||
HashiCorp Is What It Is I’ve read a lot of posts that impugn the folks over at HashiCorp. Words like betrayal get thrown around. And I get that. When the BSL change was announced, I did feel a sense of betrayal. Despite the fact that the licensing change impacts almost no one from a financial perspective, it does feel like a betrayal of deeply held beliefs.
|
||
If you were an open-source purist who had a mental model of HashiCorp being an open-source first company, the BSL announcement created cognitive dissonance. Here is a company you understood to be a good steward of open-source software, and now they are betraying that belief by altering the licensing their core products.
|
||
Your feelings of betrayal are legitimate. Your mental model of what HashiCorp is and the decisions they would make was wrong. That sucks, and I’d be lying if I didn’t say I felt the same way. What I had, and you probably didn’t, was access to people inside the company who were able to place this decision inside of a larger context.
|
||
Neither you or I were at the table or inside the heads of HashiCorp’s leadership when they made this difficult decision. Circumstances had changed for the company, and as a publicly traded company they had a fiduciary duty to address risks to their shareholders. There was a growing contingent of companies that were directly competing with HashiCorp’s cloud services using HashiCorp’s own software. How do you protect against that?
|
||
Conflict Avoidance I want to be very clear that these are my opinions and feelings, and not those of HashiCorp employees or the community at large. I’ve chosen not to engage in the various flame wars and trolling that have occurred on Reddit and LinkedIn, largely because I am a conflict avoidant person. That’s why I like creating blogs and videos, but hate dealing with feedback and comments. So I’m going to lay out my thought process, and leave it at that. Do not expect a reply, since I seldom have enough spoons to read the comments, let alone respond to them.
|
||
The cash cow for HashiCorp is Vault. It’s not their most popular product, but it’s the one they successfully monetized. Sure Terraform Enterprise has been out there for a long time, but people don’t need it to run Terraform. Honestly, a solid CI/CD pipeline in GitHub and state storage in one of the cloud providers solves 80% of what Terraform Enterprise gets you.
|
||
Terraform, despite being the lingua franca of infrastructure, simply was not profitable for HashiCorp. However, since it was the most popular product from a consumption standpoint, the company had to dedicate serious resources to maintain and improve the product. In essence, Vault paid for Terraform to be a success.
|
||
Now, along come these IaC automation companies- direct competitors to Terraform Enterprise- and they are using Terraform open-source. HashiCorp is literally developing a product for their competitors, and that kinda sucks.
|
||
At the same time, there were a non-zero number of companies that were using Vault open-source to create competing platforms to HashiCorp’s HCP Vault service. Again, HashiCorp is literally developing a product for their competitors.
|
||
So we’ve got a publicly traded company, beholden to shareholders, who is struggling to turn an profit and has direct competitors using their R&D against them. What would you do?
|
||
No, seriously, put yourself in the shoes of HashiCorp leadership and think about how you would fix this.
|
||
This is the context that was probably missing from your mental model when the BSL change was announced. If you knew all this, you might have felt less betrayed, or at least less surprised.
|
||
Early Decisions There’s a great talk that Adam Jacob gave and I’m just going to drop the link right here. Go watch it. I’ll wait…
|
||
He explicitly calls out the business model of companies like HashiCorp and Redis, who were born of the same early 2010s era, where open core seemed to be the road to success. The thing is that open core kinda sucks for everyone. The developer of an open core solution has to make a critical decision:
|
||
Neuter the functionality of the open-source version so people will pay you money, but won’t adopt the product. Make the open-source version robust so people will adopt the product, but have no reason to pay you money. (Anchor your ASP to $0, as Adam would say.) Let’s say you choose Path 1. You’re probably dead in the water from the get go. People will try your neutered open-source version, realize it doesn’t really solve their problems and feel tricked into needing to buy the paid version. Adoption is slow, if at all, and you run out of funding.
|
||
Now let’s go with Path 2. Assuming you’ve found product market fit, your open-source version is massively successful, and people don’t really have an incentive to pay for the additional features. Adoption is massive, sales are tiny, and you run out of funding.
|
||
The majority of companies go with Path 2, and rely on two things to save their bacon:
|
||
You provide a hosted version of the software that is better than people hosting it themselves. You provide professional services and support for the software, and this is something people need. The hosting solution assumes that your software is something that needs to be hosted, and that someone else won’t host it better. Go ahead and ask Elastic how well that worked out. Turns out cloud providers are really good at hosting open-source projects and operationalizing them.
|
||
Professional services and support assumes that your software is sufficiently complex and important that folks will need help setting it up and want some level of support to keep things running. As a former consultant, I can tell you that professional services is a VERY competitive market and established consulting firms will beat you to the punch any day of the week. So selling support is probably your best bet.
|
||
HashiCorp chose Path 2 for their products and had two runaway hits: Vault and Terraform. Vault had the additional benefit of being complex and difficult to implement and manage, so HashiCorp was able to make a bunch of money off of professional services and support.
|
||
Terraform though, not so much.
|
||
It’s The Capitalism Dummy If HashiCorp had tried to charge for Terraform from the get-go, I doubt it would have been such a success. We’d all be using some other IaC DSL. As Adam pointed out, the market had already decided that config management tools started at the selling price of FREE. Terraform won because it was free, easy to get started with, and met a real need. It had excellent product market fit.
|
||
Path 2 was probably the only viable option, but now we come back to, how does HashiCorp make money off the success of Terraform? The thing about Terraform is that it doesn’t really need to be hosted. It’s not an application like Vault or Kubernetes or WebSphere. It’s a programming language and workflow tool that can easily be automated by any pipeline tool. Demand for hosting is going to be low.
|
||
Nevertheless, low is not zero, so HashiCorp developed Terraform Enterprise and eventually HCP Terraform. Still, with such a small serviceable addressable market (SAM) of customers, any competition would further erode the narrow margins of hosting software. Even worse, HashiCorp was competing against IaC automation companies that were still in startup mode and didn’t have to worry about stock prices or investor sentiment.
|
||
Terraform also doesn’t really need professional services and support. If you’re a huge enterprise that is using Terraform across all your business units, then paying HashiCorp a licensing fee for support probably makes good sense. But for the SMB? Nah. You’ll probably hire a consultant to build some modules for you, set up your pipelines, and then bugger off.
|
||
It looks like HCP Terraform and Terraform Enterprise are the way to make money on Terraform, but what about all those competitors? HashiCorp is doing all the R&D for the core product that people love, and these IaC automation companies are reaping the benefit and eating into HashiCorp’s admittedly limited SAM.
|
||
HashiCorp chose an open core model from the beginning, and that choice has come back to bite them in the ass. And this is why HashiCorp found itself at a crossroads— struggling with its monetization strategy, battling competitors, and ultimately making a move that have been inevitable: selling to IBM. Or at least some big company with deep pockets that could take them private.
|
||
What About IBM? I’ve been working in technology for almost 25 years, and during that entire stretch years IBM has seemed like the past. A dinosaur of a company that lumbers through the IT ecosystem, crushing innovation in its wake, and inexorably marching into extinction. Is that a fair assessment? Maybe not. But that was my impression from anyone who worked with IBM, and every data center I went into had either replaced their IBM gear or was desperately looking to get off that last AS/400.
|
||
My point is- fair or not- I did not have a positive impression of IBM.
|
||
And then the whole Red Hat thing happened.
|
||
IBM bought Red Hat for $34B in 2019, and rather than stifling Red Hat’s culture and profitability, IBM has successfully ridden that wave back to relevance. Seriously, IBM had become more of a finance company than anything to do with technology prior to acquiring Red Hat. I’m not saying that IBM didn’t have any R&D around emerging technologies, they just weren’t the go-to for anything remotely modern.
|
||
Acquiring Red Hat and focusing on hybrid cloud was the right move for IBM, and I think that HashiCorp is a continuation of that story.
|
||
Still, I was pretty uneasy about the acquisition, and that is because I worked with some Red Hat and IBM products recently and it was not a good experience. Actually, I should be more specific. Working with Red Hat was great! Good docs, easy licensing, and simple deployment of OpenShift on a VMware cluster.
|
||
The IBM product I was meant to deploy and test was an absolute nightmare. Unreadable docs, arcane licensing, and arduous deployment. It was very clear to me that just because IBM and Red Hat are the same company, there is still a massive gulf between the two cultures.
|
||
I was worried that if HashiCorp got mired down in the traditional IBM culture, the simplicity and usefulness of HashiCorp products would be jettisoned in favor of impossible licensing portals and inane requirements.
|
||
Based on what IBM and HashiCorp have said, I don’t see that happening. HashiCorp is going to remain fairly independent from the larger IBM organization. What HashiCorp gets is access to funding for R&D and a massive salesforce in Red Hat and IBM professional services. IBM gets another cloud native company to further build out their hybrid cloud story.
|
||
I’m not sure if you’ve noticed, but since the BSL change, the development velocity on Vault and Terraform have both increased. New features are being added rapidly, and persistent problems are finally getting real solutions. Now imagine what happens when IBM’s R&D money lets them double the engineering team. I’m not kidding. There are currently 16 open reqs for engineering positions on the Vault and Terraform teams.
|
||
What About OpenTofu? I’ve been asked many times what I think about OpenTofu, and I hope I’ve been clear. I like the OpenTofu project and I’m glad that it exists. It serves a real need in the community.
|
||
For open-source aficionados, it provides an alternative to Terraform. For the IaC automation companies impacted by the licensing change, it provides a solution they can use without paying HashiCorp licensing fees. For those of us feeling a little worried about the IBM acquisition, it’s a back up plan if things go south with Terraform. And for the R&D team at HashiCorp, it serves as both incentive to innovate and another testing ground for interesting ideas and solutions.
|
||
At first I was worried about the long-term viability of the project and the splintering of the IaC community. But we’re more than a year in, and I’m seeing plenty of active development on the project, including former HashiCorp engineer Martin Atkins joining the engineering team.
|
||
While there has been a bit of acrimony in the IaC community, the impression I get is that most people are content to keep their heads down and get actual work done with whichever tool works best. The fact that it’s trivial (for now) to migrate from one to the other is great, and I hope both communities continue to prioritize interoperability.
|
||
There is a small segment of each community that likes to snipe at the other and generally act like children. It’s best to ignore them and focus on the actual solutions and your own needs. Don’t feed the trolls, as it were.
|
||
How I’m Feeling Now 2024 was a year of acceptance. I’ve accepted that HashiCorp is now part of IBM. I can’t change that outcome, but I can control how I respond to it and the mental model I use to think about it. My digital twin tells me to be cautiously optimistic, while acknowledging that I am working with incomplete information.
|
||
Regardless of what happens with HashiCorp and Terraform, there is still a massive DevOps community that I am lucky enough to be part of. That’s not something which can be bought, and it isn’t for sale.
|
||
HashiCorp helped to build that community, but they don’t own it. No one does. I hope they remain active and robust contributors, and I value the people I’ve met at the company, but I refuse to be overly precious about the company itself.
|
||
Companies come and go. Community remains. That’s what really matters. People, process, and technology? There’s a reason people are first in that list.
|
||
`,summary:`I originally started writing this post as a detailed history of my relationship with HashiCorp, how it grew as a company, and its IPO in December 2021- which happened at the exact wrong time. After laying down about 750 words, I deleted the whole thing. Why? Because that’s not the post I wanted to write. This one is.
|
||
A Year Of Acceptance For me, 2024 was a year of acceptance. I had to accept a lot of things I didn’t necessarily like.`,date:"3 Mar, 2025",url:"https://nedinthecloud.com/2025/03/03/on-hashicorp-ibm-and-acceptance/",image:"HashiCorpIBM.png",readingTime:"14"},"https://nedinthecloud.com/2025/01/09/the-science-and-magic-of-network-mapping-and-measurement/":{title:"The Science and Magic of Network Mapping and Measurement",tags:[],content:`Exploring Internet Equity and Network Magic with Dr. Nick Feamster
|
||
In the latest episode of Day Two DevOps we covered some seriously cool stuff about networking, data science, and internet equity with Dr. Nick Feamster. Nick is a computer science professor at the University of Chicago, director of their Internet Equity Initiative, and someone who’s been knee-deep in network measurements and analysis for years. He’s fascinated by how networks work (and don’t work), and his curiosity has led to some groundbreaking projects in internet performance and accessibility. Let’s break down what we talked about and why it matters.
|
||
Listen Here:
|
||
Networking: It’s Basically Magic Nick kicked things off by telling us how his fascination with networking started back when he built his first socket-based network. Just being able to type something on one machine and see it appear on another blew his mind. He called it “magic” — and honestly, it kind of is. Fast forward to today, and Nick’s work focuses on unraveling the layers of that magic by measuring network performance and applying machine learning to it.
|
||
One of the coolest things Nick shared was how early on, even at places like Google, applying machine learning to predict network response times felt like sci-fi. Now, it’s almost standard practice. But back then, people thought it was hocus-pocus. His takeaway? Stay curious. The things that seem like magic today might be common sense tomorrow — but only if we take the time to dig in and understand how they work.
|
||
The Internet Equity Initiative: More Than Maps One of Nick’s most impressive projects is the Internet Equity Initiative, which works to identify and close the gaps in broadband access across the country. The idea came to him after a pretty frustrating experience trying to get internet installed at his home in Chicago. Despite living in a relatively affluent neighborhood near the university, the ISP had no equipment nearby, and the process to get connected was ridiculously slow and outdated.
|
||
When Nick looked at the FCC’s broadband maps, he noticed the data didn’t match reality — and that wasn’t just a “his neighborhood” problem. This is happening across the country, and those inaccurate maps impact where funding for better broadband goes. So Nick and his team started working with city governments and advocacy groups to create more accurate measurements and help secure funding to improve connectivity.
|
||
Why Internet Speed Tests Lie We all love to run speed tests when our internet seems slow, but Nick pointed out just how misleading those numbers can be. The device you’re using, whether you’re on Wi-Fi, what else is happening on your network — it all affects the results. And yet, these tests often get cited as hard data for major decisions. Nick shared how, in one case, speed test data was being used to prove people were getting ripped off by their ISPs. But when he broke down the data by device type, it turned out that a bunch of people were running tests on outdated phones with older network capabilities, which capped their speeds at 100 Mbps.
|
||
Getting Involved This episode was a reminder that even seemingly simple questions, like “How fast is my internet?” can lead to massive, world-changing projects. Nick and his team are solving real problems that impact how people access education, jobs, and even basic services.
|
||
If you want to help, Nick is always open to working with folks who care about making the internet better for everyone. You can find him on LinkedIn or email him at feamster[at]uchicago[dot]edu for research-related projects or nick[at]netmicroscope[dot]com for commercial collaborations. And hey, if you’re an ISP or a city looking to make sense of your data, seriously—reach out to him!
|
||
Thanks again to Nick for being a guest on the show!
|
||
`,summary:`Exploring Internet Equity and Network Magic with Dr. Nick Feamster
|
||
In the latest episode of Day Two DevOps we covered some seriously cool stuff about networking, data science, and internet equity with Dr. Nick Feamster. Nick is a computer science professor at the University of Chicago, director of their Internet Equity Initiative, and someone who’s been knee-deep in network measurements and analysis for years. He’s fascinated by how networks work (and don’t work), and his curiosity has led to some groundbreaking projects in internet performance and accessibility.`,date:"9 Jan, 2025",url:"https://nedinthecloud.com/2025/01/09/the-science-and-magic-of-network-mapping-and-measurement/",image:"D2DO262-artwork.png",readingTime:"3"},"https://nedinthecloud.com/2025/01/02/planning-for-2025/":{title:"Planning for 2025",tags:[],content:`You know it’s funny. For 2024, I theoretically started with a plan. But that plan was in the form of a YouTube video instead of an actual blog post. The problem with creating a video instead of a post is that it’s hard to scan through it and check on how I did. It’s also hard to copy and paste the format for the next year’s blog post. So this year I’m doing my future self a favor and sticking to the blog post format.
|
||
With that preamble out of the way, I thought I would take some time and create a planning list for 2025. This list isn’t meant to be comprehensive and I certainly cannot predict everything that will come down the pike in 2025, but it’s a solid starting place for activities I am relatively certain I will continue to do in the new year.
|
||
Let’s get started!
|
||
Things to Accomplish This is the list of things that I want to get done in 2025. Where possible, I am trying to put some numbers against these goals to track my progress over the year and see how I did.
|
||
YouTube Videos Last year I published 116 videos on my YouTube channel! But that is counting all the Chaos Lever videos, of which there’s two a week. I’d like to get into a solid cadence with my Terraform Tuesdays stuff and maybe get back to doing some Cloud Perspectives and other topics. In terms of actual numbers, here’s what I’m hoping to do:
|
||
Terraform Tuesdays - 3x videos a month for a total of 36 videos Cloud Perspectives - 1x video a month for a total of 12 videos Home lab stuff - 4x videos over the course of the year, topic TBD WebAssembly Ops - 4x videos over the course of the year I have a home lab with bunch of Raspberry Pis, a Turing 2 compute board, and four Supermicro servers. I really should do something with all that gear beyond letting it take up space in the basement. What possible topics do I have in mind? Here’s a few: Proxmox, NodeWeaver, Azure Arc, and AI. I’ll have to see which one tickles my fancy this year and focus up on that. If you’re a vendor looking to sponsor something with my home lab, please reach out!
|
||
WebAssembly is one of those technologies that I’ve been meaning to dig into, but I keep putting it off. This year I’d like to really understand how wasm works and how one goes about managing it from an operational perspective. This may end up being a livestream learning thing.
|
||
Livestream Thing Oh look! I’ve added a new category of shit to do. Well done me. I’ve always enjoyed learning in public, but that has mostly been through blogs and carefully planned videos. This year I’d like to do some livestreams of me working on new technology, and I’ve found a partner to do that. Marino Wijay and I are working on a series of livestreams wherein we will try to build an Python-based solution that leverages AI to automate… something. That something right now is my podcast publishing workflow, but we’re still actively planning and brainstorming. This is subject to change!
|
||
I’d also like to rope Marino or someone else in the second half of the year to do some WebAssembly livestreaming. I may reach out to people active in that area, where they can be the expert and I’ll be the audience proxy. Are you that somebody? Reach out and ping me!
|
||
I’ve not done a ton of livestreaming before, so this is uncharted territory for me. I also don’t watch a lot of other livestreams, so I may need to do a little research and learn some of the etiquette and audience expectations. Or I’ll just wing it and make my own norms! Yeah, it’s probably going to be that second one.
|
||
Day Two DevOps Day Two DevOps went through a lot of changes in 2024, and I don’t really want to introduce anything big in 2025. In 2024, we changed the podcast name, changed the cohost, and switched to a fortnightly cadence for non-sponsored content. Oh, and we now have a merch store and cool sticker! What I’m saying is that 2024 was A LOT. I’d like 2025 to be LESS.
|
||
Kyler and I will keep finding interesting guests and chat with them about cool stuff. In fact, we already have amazing episodes in the can for January and February covering security anthropology, networking science, and the state of serverless. As always, we are open to sponsors and advertisers, so if you are a vendor who wants to get their message out and have a great time doing it, hit us up!
|
||
Chaos Lever Chris and I have settled into a nice routine when it comes to Chaos Lever. The podcast also saw significant change in 2024, with the introduction of HumblePod for editing, migrating to a proper website, migrating to a proper feed host, and publishing content on YouTube. We also experimented with guests to great success! Our shows with Sarah Autumn and Dylan Beattie were some of my favorites for the year. I’d like to have more guests in the coming year, so they can bring their unique expertise about tech history to our adoring fans.
|
||
Chris has also been making some noise about restarting the newsletter. We’ll see if he follows through on that. I’d like to also work with AI to convert out scripts to blog posts. Believe it or not, Chris and I write out most of the episode content and it should be pretty easy to convert that to a blog post. You’d think AI would be able to do this easily… You would be incorrect (so far).
|
||
Pluralsight Courses Last year I did a grand total of ONE course for Pluralsight, which is my lowest output since 2017. This was in large part due to the tumultuous year Pluralsight had and the fact that I was still feeling a bit raw about the whole slashing-author-payments-by-25%-with-little-to-no-notice thing. Also that Azure Virtual Desktop set of courses was extremely difficult to produce and I was absolutely burned out on course creation for a while.
|
||
2024 was a year of acceptance and coming to terms with things, and I have come to terms with the fact that Pluralsight has changed and I need to change with it. They still represent a large portion of my revenue, and if I want to sustain that revenue, I need to put work back into the machine. Keep the flywheel going as it were.
|
||
My plan for 2025 is to refresh several of my courses and create a couple new ones:
|
||
Vault Associate certification - A new version of the exam has been published, so the courses all need to be updated to reflect the change. Terraform Getting Started and Deep Dive - I update these on a regular cadence to keep up with developments and new releases of Terraform. Terraform Pro certification - I need to create a set of courses around the content of the Terraform Pro cert along with labs. Those items should keep me plenty busy in 2025!
|
||
Certification Guides The Terraform Associate guide has been updated for the 003 version of the exam. I don’t expect a new version of the exam to drop for at least another year, and more likely two years.
|
||
My Vault Associate guide needs to be updated for version 002 of the exam. I haven’t looked too closely at the updated objectives, but I don’t think they’ve changed too dramatically. I’ll try and get this done in Q2 of the year.
|
||
As for the Terraform Pro certification, I started writing a guide in 2024. But then Mattias Fjellstrom published his and I am just pointing everyone at that guide. Maybe I’ll change my mind, but I doubt it.
|
||
Live Instruction In 2024, I averaged one HashiCorp Academy class a month. Given the dip in Pluralsight revenue, I am probably going to increase that a bit in 2025. I’m aiming to do about four classes per quarter, or 16 for the year. I’ve already signed up to do eight classes through June, which puts me on track for my goal. It’s a nice mix of Terraform Foundations, HCP Terraform, and Vault Enterprise, so I won’t be teaching the same class eight times (whew).
|
||
Website and Blog The migration of my site off Wordpress was completed in 2024, and I am very happy about that. I didn’t blog as much as I wanted to last year, and I’d like to change that. I know, I KNOW. Every blogger says they mean to blog more and few rarely do it. However writing by itself is not the problem. I write thousands of words every week for various things, it just doesn’t always make its way onto the blog. For all the podcasts and videos that I do, each one could probably become a blog post. Doing so requires transforming the media to a written form that is easy to consume, and has all the attendant graphics, sections, and tags.
|
||
My hope is that I can leverage AI to help transform my other media content into blog posts. If ChatGPT or similar could get it 80% of the way there, I can tailor it to meet my standards and move it across the finish line. I’m hoping to have at least one blog post a week. Whether its something original or adapted from Terraform Tuesday, Chaos Lever, or Day Two DevOps, content is not the problem, time and effort is.
|
||
Tech Field Day and Gestalt IT I’ve signed up to attend Cloud Field Day 22 in February, and I’ll probably do another one this year in the Fall. These events really do help me stay abreast of what’s going on with technology and network for new clients and sponsorships.
|
||
Conferences For 2024, I kept things pretty chill when it comes to conferences. I went to the DevOpsDays Philly in the spring and then both the NYC and DC AWS Summits in the summer. For the autumn, I only went to HashiConf, despite almost going to KubeCon and re:Invent. Ultimately, there was just too much going on at home and I didn’t see a huge benefit to going to either event. It also doesn’t help that I would have attended the events on my own dime, and conference passes, travel, and hotels ain’t cheap. Considering it was a down year for Ned in the Cloud as a whole, adding a bunch of additional expenses didn’t seem financially sound.
|
||
In 2025, I want to continue focusing on smaller conferences that are closer to home. Although there’s nothing on the books yet, I hope DevOpsDays Philly comes back for the spring. I’ll definitely come back to the AWS NYC Summit in the summer and HashiConf in the autumn. I think I also need to go to KubeCon or re:Invent, but which one? Honestly, it’s going to come down to whether I can get someone else to foot the bill for travel. Maybe I’ll submit some talks to both and see if I get accepted at either.
|
||
If you have suggestions for other conferences, or maybe even want me to keynote yours, let me know! I delivered the keynote for IntegraONE’s ONE CON conference this year and it was a great success!
|
||
Other Opportunities There were a few surprise opportunities that popped up in 2024 that I will be pursuing in 2025.
|
||
O’Reilly Book I have agreed to write a shorter book for O’Reilly around running Linux on Microsoft Azure. The first chapter is already done and submitted, leaving another six to go. The book is slated for publishing in June. I’m coauthoring it with Chris Hayner, which makes the workload manageable. This book has a big upfront payment with no royalties. Given my previous experience with book royalties, upfront payments are the way to go.
|
||
Terraform Consulting A former Cloud Field Day presenter reached out to me to develop some Terraform modules for their edge solution. Building them was a lot of fun and they have already indicated they’ll have more work for me in 2025. It was nice to do some real-world consulting and not just create demos and training.
|
||
Multiverse Several of the people I worked with at Pluralsight have gone to work for a training company called Multiverse. I’ve already built out some instructional materials for them in 2024, and I expect to do more in 2025. Much like the O’Reilly book, this is upfront payments for work delivered and no royalties. The level of effort is nicely aligned with the compensation- after a little haggling- so I’m happy to keep doing this work for people I trust.
|
||
Conclusion If I’m being totally honest, I’m entering 2025 with a mix of trepidation and excitement. The trepidation comes from both the political reality of being a US citizen, and also the economic uncertainty of having an ill-tempered, petulant toddler for a US President being led around by another ill-tempered, petulant toddler, neither of which seem to have ever learned impulse control. That could reek havoc on the economy and the tech sector, or it could not.
|
||
My excitement stems from all the new projects I’m going to embark on this year. Learning Python and AI, digging into wasm, doing actual Terraform consulting, and who knows what else by the end of the 2025. As long as I stick to my core principles of embracing discomfort, planning to fail, and leading with kindness, this coming year is going to be a good one.
|
||
Happy New Year to all of you and may you have an interesting and enjoyable 2025!
|
||
`,summary:"You know it’s funny. For 2024, I theoretically started with a plan. But that plan was in the form of a YouTube video instead of an actual blog post. The problem with creating a video instead of a post is that it’s hard to scan through it and check on how I did. It’s also hard to copy and paste the format for the next year’s blog post. So this year I’m doing my future self a favor and sticking to the blog post format.",date:"2 Jan, 2025",url:"https://nedinthecloud.com/2025/01/02/planning-for-2025/",image:"Planning-for-2025.png",readingTime:"11"},"https://nedinthecloud.com/2024/12/30/2024-year-in-review/":{title:"2024 Year in Review",tags:[],content:`The overall theme of 2024 was acceptance. If you go through my 2023 year in review, I was generally pretty negative about the year. Not to say that nothing good happened in 2023, but it was a major down year in terms of revenue and I failed at a number of things. That’s okay though! One of my central pillars is to embrace failure, and I did that. Even though 2024 saw another dip in the Ned in the Cloud revenue, I have made my peace with that and moved into the phase of acceptance and planning. There’s going to be a whole separate post about planning for 2025, so let’s focus on what happened in 2024.
|
||
High-level Review As a company, Ned in the Cloud saw a revenue drop of 4%. That is not nearly as drastic as the 30% drop seen in 2023! The major cause of the drop was Pluralsight, a topic which I have more to say about in the Pluralsight section. I also wrote a whole post about what is going on with Pluralsight as a whole, which happens to be one of my most popular posts this year.
|
||
Despite my revenue from Pluralsight dropping by 17%, I had a host of new clients come in to lessen the impact. Terramate, Firefly, IntegraONE, and Cast.AI all helped me make up for lost revenue. Additionally, I did more live training for HashiCorp Academy. My projection for 2025 actually has another revenue drop of about 8%, but we can talk about that more in the planning section.
|
||
Technical Education The central principle behind Ned in the Cloud is technical education. I know I say this every year, but the whole point and mission behind Ned in the Cloud is to provide educational material to folks in the information technology field. Almost every project I undertake has some educational component to it, whether that’s writing documentation, creating videos, or writing courses. What can I say? I love learning and sharing with others!
|
||
YouTube Videos In 2024, I published 116 new videos to the channel. Which is a lot more than 50 last year! Here’s the thing though, YouTube added the ability to publish your podcast directly to your YouTube channel through a feed. I added Chaos Lever, which mean that all Chaos Lever episodes (two a week!) were now being added to my YouTube channel automatically.
|
||
In fact, the import process adds all the older episodes too, which means I added 235 videos to the channel. But I am only going to count new episodes, which is how we get to the 116 number.
|
||
I’ll deal with the Chaos Lever of it all in a later section, so let’s talk about the other YouTube videos. First, some stats on my best performing videos released this year:
|
||
Title Date Views Using Azure Storage for Terraform State - Best Practices 2024-01-04 3,973 Getting Started with Terraform Stacks 2024-10-29 3,518 Migrating From Terraform To OpenTofu 2024-06-18 3,199 Using the Terraform Test Framework 2024-05-29 2,980 Introducing Azure Verified Modules for Terraform 2024-02-14 2,393 What can I glean from these top performers? Well, Terraform continues to perform well for me. Anything Azure related also seems to get a nice bump, but in fairness, I don’t often do videos that are AWS or GCP specific. I don’t really have anything to compare my Azure videos to, except for a lack of Azure.
|
||
The other thing I should note about these videos is the fact that videos released earlier in the year have had more time to accumulate views. A better metric for comparison would probably be views in the first 30 days, if I’m trying to compare apples to apples. As far as I can tell, there’s no easy way to do that with the YouTube analytics engine. If you know a way to do it, please let me know!
|
||
My top videos watched videos in 2024 are all from previous years, which is a little depressing I guess:
|
||
Title Date Views Terraform Basics: Modules 2021-06-01 13,581 Exploring the Import Block in Terraform 1.5 2023-07-12 7,059 Managing Multiple Environments with Terraform 2023-10-17 5,562 What is PXE Boot? 2020-11-19 5,481 Azure DevOps Pipeline with Terraform 2021-05-18 5,125 I had no idea my Terraform Basics: Modules video was going to blow up like that. I mean, it’s a good video and I’m proud of it, but it’s not like it’s earth shattering. Sometimes 101 content is just what people the need. It gets them going!
|
||
Case in point, my What is PXE Boot? video. I made that video four years ago and it still performs like a champ. At the time, I was just trying to get my Raspberry Pi to PXE boot properly and I ended up doing a deep dive on the tech. The resulting video has almost 22k views, making it my 5th most popular video of all time. What’s funny is that I’m not a PXE boot expert by any stretch of the imagination. I just wanted to figure out how to configure my DHCP server to support PXE boot and configure a TFTP server to support booting from LAN. Turns out I wasn’t the only one!
|
||
The other trend I am noticing is that my Terraform Basics series in general does really well. Which is why I have more in the series planned for 2025. In fact, here are a few key takeaways I’m going to use to inform my videos in 2024:
|
||
People are always looking for 101 content People want to know about new products and features Follow your passion, especially on other projects That translates into more coverage for the Terraform Basics series, covering new products in the IaC space, and maybe getting back into the homelabbing stuff of yesteryear. I still have a Turing Pi 2 board collecting dust and Supermicro servers that needs a purpose. I know I said I was going to dig into Proxmox in 2024 and I didn’t. Maybe I should in 2025!
|
||
Day Two DevOps (nee Day Two Cloud) There were some big changes over on the Day Two Cloud podcast. A new name and a new cohost! What didn’t change was the amazing content and guests we had on the show. Let’s first talk about the host change.
|
||
Day Two Cloud started out with just me at the beginning of 2019, and I hosted the first 21 episodes solo. After that, Ethan Banks joined me for episode 22 and we changed the schedule of the show from bi-weekly to weekly.
|
||
Having a cohost is awesome! In part because you’re no longer responsible for sourcing every guest. While Ethan and I didn’t keep close track of who found which guest, I would say we split things about 50/50. Ethan also assisted with formulating questions, keeping editor notes, and asking questions that I wouldn’t think of. Given his deep background in networking and mine in cloud, we could play the audience surrogate for different roles.
|
||
Ethan is also the co-founder of Packet Pushers along with Greg Ferro. That makes him responsible for keeping the company running. He also cohosts least two other podcasts on the network.
|
||
When Greg opted to retire in 2024, Ethan suddenly had a lot more on his plate. I could tell it was weighing on him. He dropped a couple subtle hints about it being okay if I wanted to have a different cohost on some of the episodes, and I took the cue. It was time to lighten Ethan’s load.
|
||
Thus the search for a new cohost commenced. But where to start? How do you find a cohost that you think you’ll gel with and that the audience will enjoy listening to? My mind immediately went to previous guests on the podcast that I really enjoyed talking to. I also wanted to bring some diversity to the podcast, rather than it being yet-another-two-middle-age-white-guys-talking-tech dot com. Based on that criteria, I drew up a short list. At the top of that list was Kyler Middleton.
|
||
I found Kyler through the HashiCorp Ambassador program. We were both Ambassadors in 2021, and at the time I was looking for potential new guests for Day Two Cloud and for my YouTube channel. What better place to look than the list of HashiCorp Ambassadors?! It’s a self-selected group of community-minded individuals who excel at communicating and sharing the interesting tech they’re working on. I did the same thing with Microsoft MVPs when I was getting Day Two Cloud off the ground in early 2019.
|
||
Kyler was a guest on episode 128 of Day Two Cloud, and if you go back and listen, I think you can tell that Kyler, Ethan, and I are having a blast. Not only was it a great conversation, but we also gelled as people. I immediately put Kyler in my list of cool people I want to collaborate more with. Yes this is a mental list I keep, no you cannot see it.
|
||
Since she was at the top of my potential cohost list, I reached out to her first and immediately she said yes. Just like that, my “long” search for a new cohost was over. Kyler made her debut in episode 236 on March 6, 2024 and the rest is history.
|
||
Now about that name change thing…
|
||
When I started Day Two Cloud in 2019, the idea was simple. Find practitioners who had done interesting stuff in the cloud and find out what challenges they faced beyond Day One. Over time I started to think that Cloud was too limiting of a term. We weren’t just talking about cloud technologies, but well beyond into the practices, processes, and people that make modern infrastructure work. People? Processes? Practices? Hmmmm… that sounds suspiciously like DevOps!
|
||
I figured since we were shaking things up with cohosts, it might be the ideal opportunity to change the podcast’s name to reflect the topics we actually cover. We worked through a few different ideas and ultimately landed on Day Two DevOps. It took a few months to prepare the branding updates and figure out the naming change with the website, but finally in July of 2024 the name change became official. Day Two Cloud was now Day Two DevOps, where the DevOops is in the details.
|
||
In terms of episodes we published this year, here is a selection of my favorites:
|
||
D2DO247: Chocolate or Carrots? How Humor Can Foster Good DevOps Relationships - Kyler and I chatted with Ashish and Shilpi from the Cloud Security Podcast and had a grand old time. D2C229: Standing Out From The Crowd With Tim Banks - Anytime you have a chance to chat with Tim Banks, I suggest you take it. D2C240: The Duality of Enterprise AI - Jonah and Greg from Heavy Strategy join us to talk about the state of Enterprise AI. D2DO248: Using Creativity and Empathy to Ease the Pain of Compliance Audits - You might think that auditors have it out for you and compliance is a total drag. We’re here to change your mind. D2DO253: Breaking Into Tech: Job Hunting Realities for Recent Graduates - Katrina Janeczko shares her experience breaking into tech as a recent college grad. D2DO258: How System Initiative Rethinks Infrastructure as Code - Adam Jacob paints a compelling alternate vision for managing infrastructure that is not Terraform. D2DO257: Love is in the Firmware - Dr. Cat Hicks and Dr. Ashley Juavinett share their scientific studies when it comes to developers and operations. Keeping myself to seven episodes was a real challenge!! What about the audience? What episodes were most popular with them? Here’s the top five:
|
||
Title Date Downloads D2C242: Data Engineering and its Streams, Rivers, and Lakes 2024-05-08 8,043 D2C240: The Duality of Enterprise AI 2024-04-17 7,442 D2C245: Don’t Fear Database DevOps 2024-06-19 7,090 D2C235: Building Modern Apps In GovCloud 2024-02-28 6,142 D2C234: What to Do About VMware 2024-02-21 5,130 You might notice that these are all from the first half of the year, and that’s probably because they’ve had the most time to rack up downloads. Most episodes have a big initial spike of downloads when published and then a long tail of downloads over time. The more recent episodes simply haven’t had time to rack up that long tail.
|
||
Still, based on the topics that resonate most with folks, it certainly seems that AI, data wrangling, and non-public cloud are of interest. If I expand the view out beyond the top five, that pattern doesn’t really hold. So I guess formulating a content strategy may be an exercise in futility. Instead, Kyler and I will keep finding interesting people and inviting them to be on the show. We’ve already got some great stuff recorded for 2025 with more to come!
|
||
Chaos Lever The year for Chaos Lever is a bit hard to describe. I guess you could say it’s the year we grew up a little? As I mentioned in my 2023 post, I hired HumblePod to help with production on Chaos Lever. They took over the editing, publishing, and promotion of the podcast. In the process, they got me to move hosting from my home-grown Azure Storage solution to Transistor and the website off an Azure Static Web App and onto Podpage.
|
||
The result is that Chaos Lever looks way more professional and I can actually collect reliable statistics! HumblePod also put together YouTube shorts using our recordings to help promote the episodes and showed me how to integrate the podcast feed with my YouTube channel. I really can’t say enough nice things about HumblePod and all they did for the show.
|
||
Unfortunately, that amazing service also comes with a hefty price tag, and as of right now Chaos Lever effectively makes no money. My hope was that the more refined Chaos Lever would attract sponsors and we would start getting some sweet, sweet ad revenue floating in to pay for all the excellent work HumblePod was doing. After eight months, I made the difficult decision to drop HumblePod and go back to doing things myself. It’s not too bad, and my hope is that Chris and I eventually manage to build up enough of a following that we’ll be able to afford HumblePod sometime in the future. Until then, I’ll just have to deal with editing the episodes myself.
|
||
Speaking of episodes, I think we really embraced the historical slant of Chaos Lever this year and in the process created some amazing episodes. The most popular episodes according to the audience were:
|
||
Title Date Downloads HashiCorp Under IBM’s Wing 2024-05-02 770 The Reality of ‘Secure by Design’ and the Future of Cybersecurity 2024-03-07 456 DNS: The Backbone Of Browsing (Part 1) 2024-05-16 358 Break the Glass and Walk Away: A (VERY) Brief Overview of BGP 2024-06-20 301 Going Deeper into BGP with Doug Madory 2024-07-18 278 What’s interesting is that most of the downloads and view are coming from YouTube and not the podcast feed itself. Funny that. I guess adding Chaos Lever to YouTube was a good move. The download numbers are still nothing to write home about, and I’m not sure what else I could be doing to promote the podcast.
|
||
I don’t know if Chaos Lever will ever be the smash hit I think it should be, or if Chris and I will continue to toil in obscurity for the next X number of years. I suppose as long as we both enjoy the work and the process of discovery that comes with it, then it continues to be worth doing on its own merits.
|
||
HashiCorp Academy After Pluralsight, my biggest revenue driver this year was teaching HashiCorp Academy courses through River Point Technology. If my notes are to be believed, I taught eleven classes in total: six Terraform Foundations classes, three Terraform Advanced classes, one Terraform Fundamentals private training, and one Vault Enterprise class. My goal for 2024 was to stick to about one class a month and that’s what I did. There was actually one more class I was signed up to teach that got cancelled, so I would have been at twelve exactly.
|
||
Pluralsight Courses Pluralsight continues to be the bulk of my revenue here at Ned in the Cloud. However, the company as a whole has been experiencing some significant financial woes, as I documented here along with my feelings on the matter. I’m not going to rehash all that in this post- really you should just go read what I wrote- the tl;dr is that Pluralsight is still an excellent platform for learners that is slowly being ruined by private equity. My hope is that the current executive team can realign on doing things right and not just for short-term gain.
|
||
In terms of courses for Pluralsight, I published two course updates and one actual course. The Azure Virtual Desktop certification course needed to be updated to reflect the new objectives in the certification. I also updated my Implementing Terraform on Microsoft Azure course with a new Advanced Terraform with Azure course. But there’s some weirdness going on with the two courses that I need to explain.
|
||
I published the Advanced Terraform with Azure course in June on A Cloud Guru. The older Implementing course continues to live on the Pluralsight platform and is still available. Since publishing six months ago, the older Implementing course has 2,286 viewed hours and my newer Advanced course has 155 hours. What this tells me is that of the two platforms, Pluralsight is wildly more popular for me.
|
||
Why was my updated course published on A Cloud Guru (ACG) and not on Pluralsight? Here’s a little background. Pluralsight acquired A Cloud Guru back in 2021. At the time, they were two separate platforms with different catalogs and experiences. Since that time, Pluralsight has been trying to meld the two together into a single cohesive platform. But that is still not really the case!
|
||
The two platforms are still fairly different and you still have to purchase plans separately. Which means if someone signed up on Pluralsight to take my Terraform Getting Started and Deep Dive courses and then wanted to take an Azure-focused Terraform course, they would only have access to my older Implementing course! That’s why the older course is so popular, because its the only option for Pluralsight subscribers.
|
||
However, Pluralsight made the decision that all cloud-related content would now be published on the ACG platform. Since my Azure-focused Terraform course was indeed focused on a cloud platform, Azure- it needed to be published on the ACG side, where no one who has a Pluralsight-only course can access it. The result is that learners are forced to take an older course because they don’t have access to the newer one. Why they couldn’t just publish the new course across both platforms is beyond me.
|
||
Anyway, I’ll have more to say in my planning for 2025 post, so I think I’ll just stop here.
|
||
Books I did not write any new books or update any existing books in 2024. I did start writing a Terraform Professional certification guide, by my buddy Mattias beat me to it and I’m just recommending that people go buy that instead.
|
||
Other Appearances Naturally I appeared on a few other things this year:
|
||
HashiTalks 2024 HashiConf 2024 IntegraONE ONECON Keynote DevOpsDays Philly Workshop Cloud To Cloud Ops’N’Hops Veeam Partner Perspective PA NUG Panel (no link yet) HashiTalks Build There might be some more, but looking through my notes and calendar, I don’t see anything.
|
||
Deprecations and Failures This is a fun section with a fun name. What did I fail at in 2024? And I don’t mean that in a super-negative way. One of my three pillar is embracing failure, so it’s time to give my failures a big ole hug!
|
||
Failure 1 - Proxmox Boy howdy did I think I was going to do a whole series of videos about Proxmox. This was spurred by the VMware acquisition by Broadcom and a general sense that people might be looking for another virtualization platform. I actually did end up building a Proxmox server in the homelab and tinkering around with it. But then I got busy with other things and Tim Warner published a Proxmox course on Pluralsight, so I just kinda let it go to the side. Will I pick it back up in 2025? Unlikely.
|
||
Failure 2 - WebAssembly I also thought I might spend some more time with WebAssembly and get to know it from an operations standpoint. I also didn’t do that. I didn’t even spin up a basic proof of concept. I just shunted wasm off to the side and continued on with my life.
|
||
Failure 3 - Crossplane Crossplane is an interesting alternative to CloudFormation or Terraform. I tried getting started with it on Azure using kind for Kubernetes clusters, and I pretty much stopped there. The learning curve and admin curve of Crossplane is pretty steep and I just don’t see the benefit. Maybe I’m ready to tackle this one in 2025.
|
||
Failure 4 - Pluralsight Courses I had planned to do some updates to my HashiCorp Vault courses, the Implementing Terraform on AWS course, and maybe do an OPA and Terraform course. I did none of those. Mostly because of the strained relationship with Pluralsight and the fact that most of the people I worked with there left. Still Pluralsight is kind to me when it comes to revenue, so maybe it’s time to put some effort back into the platform. The new version of the Vault certification dropped and that alone should require a review of the existing material.
|
||
Looking Forward What’s coming for 2025? A lot more videos, blog posts, and podcasts. Of that you can be sure. For specifics, I guess you’ll just have to eagerly await my 2025 planning post/video!
|
||
`,summary:"The overall theme of 2024 was acceptance. If you go through my 2023 year in review, I was generally pretty negative about the year. Not to say that nothing good happened in 2023, but it was a major down year in terms of revenue and I failed at a number of things. That’s okay though! One of my central pillars is to embrace failure, and I did that. Even though 2024 saw another dip in the Ned in the Cloud revenue, I have made my peace with that and moved into the phase of acceptance and planning.",date:"30 Dec, 2024",url:"https://nedinthecloud.com/2024/12/30/2024-year-in-review/",image:"2024-Year-in-Review.png",readingTime:"18"},"https://nedinthecloud.com/2024/11/15/resourcely-guardrails-and-blueprints/":{title:"Resourcely Guardrails and Blueprints",tags:["resourcely","terraform"],content:`I’ve spent the last eight years working with Terraform on an almost constant basis. Sometimes I think in HCL, and dream of waves of resource creation messages flying past my eyes. But for most people, that is not the case. While I’ve cultivated an expertise in Terraform, I cannot expect others to have done the same. That includes Ops folks, but also developers, networking teams, SREs, and security peeps. They all have responsibilities and skills that aren’t thinking about Terraform 24/7.
|
||
Scaling Terraform Beyond Day One When organizations get started on their IaC journey, it’s usually an outside consultant or an internal platform team that introduces and implements Terraform. Initial projects are successful because of the deep expertise and knowledge those folks bring to the table. But how do you generalize that knowledge? How do you spread the wealth of expertise around to other teams that are less familiar with Terraform and infrastructure in general?
|
||
You could try and make everyone learn Terraform. I’m sure there’s nothing a bunch of harried developers would like more than to learn yet another DSL (YADSL?). Learning Terraform isn’t easy. I’ve been working with it since 2016, and I’m still learning about new features and functionality. And that’s just the IaC, developers also need to learn the intricacies of the cloud platforms you’ve been using, including applying proper security and sane defaults.
|
||
What about the network, security, and identity teams? They all have their own fires to put out. Asking everyone in the organization to adopt and integrate Terraform into their daily practices is just not realistic. So what’s the solution? Usually it’s to build a platform team that crafts well-defined standards, golden paths, and templatized environments for other teams to use. They can lean on specialized teams for their expertise and translate that expertise into practical implementations, like codifying infrastructure in Terraform modules and applying policy as code to ensure adherence to best practices.
|
||
Those Terraform modules are still using Terraform though. One problem platform teams are trying to solve is how to help others interact with Terraform modules without having an in-depth knowledge of the language. If you’ve ever tried to use a sufficiently complex one (AKS I’m looking at you), you realize how quickly the abstraction starts to break down. Additionally, combining modules in a coherent composition still requires an understanding of the Terraform language itself and how modules interact with each other. Platform teams can try and put a pretty UI in front of that, but it doesn’t solve the underlying challenge of root module authorship.
|
||
This is a problem that managed solutions are trying to solve. One such solution, and sponsor of this post is Resourcely. They are a configuration platform for defining and deploying cloud infrastructure at scale in an efficient and secure manner.
|
||
Resourcely offers Blueprints for deployment and Guardrails for enforcing policy and compliance, both of which can be authored by your platform team. Developers can consume these artifacts from behind a UI that abstracts the underlying Terraform code, while still leaving it accessible for break-glass scenarios. Your developers can build the infrastructure they need quickly and without that deep working knowledge of Terraform.
|
||
In this post, I’ll run through how Resourcely is set up and how you can author and utilize Blueprints and Guardrails to deploy infrastructure.
|
||
Introducing Resourcely As I just mentioned, Resourcely is a platform to aid in defining and deploying infrastructure. It integrates with source control to manage your infrastructure as code, and it leverages Terraform under the covers to perform the actual deployment and management of your environments.
|
||
The two main concepts to understand in Resourcely are Blueprints and Guardrails. Blueprints are an abstracted form of a Terraform module, providing an interpretation layer between a developer-facing form and a rendered Terraform configuration. Guardrails are a way to define policy as code using the Really domain specific language. Guardrails are applied to Blueprints and verified during the planning stage to ensure that proposed deployments are in compliance. That is all a bit esoteric, so let’s walk through a demonstration to see it all in action.
|
||
Google Cloud Storage Demo For this demo, let’s suppose I’ve been charged with creating a self-service option for developers to provision a Google Cloud Storage bucket. The module should create a basic storage bucket and ensure that public access is not enabled. Additionally, if the bucket is used for production it should be limited to regions in the US.
|
||
We can use a combination of Blueprints and Guardrails to expose the right fields to the developers while ensuring that we’re in compliance with policies. First, we’ll create a Blueprint using a starter template. Next we’ll create a Guardrail to prevent public access for the bucket. Then, we’ll deploy a storage bucket using GitHub integration. Finally, we’ll add a second Guardrail to restrict the region selection and verify it works.
|
||
Creating a Blueprint Once you’re signed into Resourcely, you’ll see the Foundry, Blueprints, and Guardrails options in the side menu.
|
||
Foundry is where you can author Blueprints and Guardrails from scratch using Resourcely’s built-in IDE. You can also start the creation process from the Blueprints or Guardrails view.
|
||
Blueprints are defined using the file type .tft, which I can only assume stands for TerraForm Template. When someone wants to consume a Blueprint, they are presented with a form to fill out based on the contents of the Blueprint file. The values from the form are combined with the Terraform template in the Blueprint to create valid Terraform code that can be checked into source control.
|
||
When you author a Blueprint, you get to decide what fields are available to the consumer and what values are allowed for that field. Later, we’ll see how we can use Guardrails to further enforce settings based on context and policy.
|
||
In the video below, I am creating a Blueprint from a starter supplied by Resourcely. You can also opt to import a Blueprint from a source control repository or from the modules in the Terraform public registry.
|
||
There should have been a video here but your browser does not seem to support it. Blueprint Content The content for the Blueprint looks very similar to Terraform code, except there are some variables and syntax that need to be interpreted. The Blueprint file can start with YAML-based front-matter that includes the definition of constants and variables to use inside the template.
|
||
--- constants: __name: "{{ name }}_{{ __guid }}" This code defines a constant called __name that is a combination of the variable name and a GUID created by the special expression __guid. We can now use the __name constant throughout the rest of the template as a unique value.
|
||
You may have already noticed that variables and other special expressions start and end with double curly braces {{ }}. If you’re familiar with Go’s templating language, you’ll feel right at home here.
|
||
Below the constant, we define a variable called location and its properties.
|
||
variables: location: type: string desc: "GCP Location to use for the bucket." required: true suggest: "US-CENTRAL1" --- Resourcely calls the properties tag parameters. In addition to setting the type and description, we are also providing a suggestion for its initial value.
|
||
You might notice that the name variable used in our constant hasn’t been defined yet. That’s because you can define variables in-line as well.
|
||
Below the front matter is the Terraform code to create our bucket. It uses the {{ }} to signify that some interpolation needs to occur when this code is rendered.
|
||
resource "google_storage_bucket" "{{ __name }}" { name = "{{ name | desc: "Name of bucket" }}" location = "{{ location }}" force_destroy = true public_access_prevention = "enforced" uniform_bucket_level_access = true lifecycle_rule { condition { age = 7 } action { type = "AbortIncompleteMultipartUpload" } } } The resource block name label is "{{ __name }}", which will use the value stored in the constant __name for the resource name in the configuration. This helps with uniqueness of names once the template is rendered.
|
||
The location and name arguments both use template variables. Variables can be defined in-line or as part of the front-matter. The name variable is defined in-line using the syntax {{ name | desc: "Name of bucket" }}. The | character after the variable’s name is used to set tag parameters, with a pipe between each tag parameter.
|
||
The template variables are turned into a form the consumer can fill out, which is shown on the Developer Experience tab.
|
||
In our form, both the Name and Location appear along with the description we included for each variable. The location value has also been pre-populated with US-CENTRAL1.
|
||
Resourcely is smart enough to understand the context of the location value, and so it creates a drop-down list of GCP regions to select instead of a free-form text field.
|
||
Nothing like preventing invalid input!
|
||
The official documentation for Blueprints includes several other supported tag parameters, and other useful interpolation syntax.
|
||
Beyond the content of the template, there are three other sections that go into a complete Blueprint: Guardrail customization, metadata definition, and context attachment. We don’t have any Guardrails yet, so we can skip that section.
|
||
Define Metadata The Define Metadata section allows us to specify information about the Blueprint itself.
|
||
The fields include the Name of the Blueprint, a Description, which Provider is being used, and metadata tags to associate with the Blueprint to aid in discovery.
|
||
Attach Context The settings shown in the Attach Context section come from the Global Context defined at the top-level of the organization.
|
||
Each context value will be added to the form presented to the consumer, and the value can be referenced by any Guardrails being applied. We’ll touch on this later when we use the context to control which regions are allowed for a bucket.
|
||
Previewing the Terraform Code In addition to configuring the Blueprint settings and viewing the developer experience, the Terraform tab shows the resulting Terraform configuration based on the form input.
|
||
You can toggle between the Settings, Developer Experience, and Terraform tabs to refine and test the Blueprint until it meets your requirements.
|
||
Publishing the Blueprint Once all the sections are complete, we can create our and publish our Blueprint. Now it will appear in the list of available Blueprints for use.
|
||
Blueprints can be set to an Unpublished status, hiding them from consumers.
|
||
Creating a Guardrail Based on our requirements, we need to prevent the Google Cloud Storage bucket from being publicly accessible. We will achieve this by creating a Guardrail.
|
||
The argument that prevents public access in our code is public_access_prevention. We set that argument to enforced, which prevents public access. But what if someone alters the configuration after it’s rendered? We can detect and potentially block deployment of the altered configuration using Guardrails.
|
||
Guardrails are verified both when the form is being filled out and when the Blueprint is queued for deployment. Resourcely can check the contents of the Terraform plan and verify it is in compliance with the attached Guardrails.
|
||
We’ll start over in the Foundry to create a Guardrail.
|
||
There should have been a video here but your browser does not seem to support it. In the video, I select a starter Guardrail that already has the settings I’m interested in. The resulting code is expressed using Really, which reads a lot like a SQL query.
|
||
GUARDRAIL "[Storage] Bucket Public Access Prevention Enabled" WHEN google_storage_bucket REQUIRE public_access_prevention = "enforced" OVERRIDE WITH APPROVAL @ned1313 Essentially, the Guardrail looks for any instances of the google_storage_bucket resource type and checks to make sure the public_access_prevention argument is set to "enforced". If it’s not, then the Guardrail fails, but I can override it because I’m super special.
|
||
Similar to the Blueprint creation, the Guardrail authoring environment has a Define Metadata section where we can specify things like the Guardrail’s name and description.
|
||
All of the fields are pre-filled based on the starter Guardrail.
|
||
The Set Activation Policy section is where the rubber meets the road. It determines where the Guardrail is applied and how it should be evaluated.
|
||
A Status of Inactive means the Guardrail isn’t applied or evaluated. The Evaluate Only status will include an informational message about the status of the Guardrail evaluation, but it won’t block the process from moving forward. When a Guardrail is Active, it will block deployment on failure, unless someone with permission overrides it.
|
||
For the Repository setting, you can specify a single repository or set of repositories using wildcards. We left the field blank, which means the Guardrail will apply to all repositories in the organization.
|
||
Now that we have our Blueprint and Guardrail, we’re almost ready to deploy some infrastructure. First we’ve got to get our GitHub organization connected, and set up GitHub Actions to integrate with Resourcely.
|
||
Connecting to GitHub There are two steps to connecting Resourcely to GitHub and integrating it with your GitHub Actions. First you need to grant Resourcely access to your GitHub organization and configure a webhook that fires on pull requests and pull request reviews. The video below demonstrates the process:
|
||
There should have been a video here but your browser does not seem to support it. You can also follow the instructions documented here.
|
||
Once Resourcely is connected to GitHub, you can update your GitHub Actions to incorporate Resourcely into the CI process. There is an example GitHub Action workflow provided by Resourcely you can leverage. Let’s look at some of the key portions of my configuration.
|
||
The Resourcely CI process needs a copy of the plan output in JSON format. In the terraform-plan job section of my workflow, I save the plan to a file, convert it to JSON, and save it as an artifact:
|
||
- name: Terraform Plan run: terraform plan -var "project_id=\${{ secrets.PROJECT_ID }}" -out=plan.tfplan - name: Convert plan to JSON id: convert_plan run: terraform show -json plan.tfplan > plan.json - name: Upload Terraform Plan JSON Output uses: actions/upload-artifact@v4 with: name: plan-file-json path: plan.json The resourcely-ci job only fires when the terraform-plan job is complete and if the event is a pull request:
|
||
resourcely-ci: needs: terraform-plan if: github.event_name == 'pull_request' The job checks out the source code, retrieves the plan file, and kicks off the Resourcely action:
|
||
steps: - name: Checkout uses: actions/checkout@v4 - name: Download Terraform Plan Output uses: actions/download-artifact@v4 with: name: plan-file-json path: tf-plan-files/ - name: Resourcely CI uses: Resourcely-Inc/resourcely-action@v1 with: resourcely_api_token: \${{ secrets.RESOURCELY_API_TOKEN }} resourcely_api_host: "https://api.resourcely.io" tf_plan_directory: "tf-plan-files" You might notice the secrets.RESOURCELY_API_TOKEN in the above code. You’ll need to generate a Resourcely API token and save it in the GitHub Actions Secrets for the repository.
|
||
You can check out the full workflow file, as well as the rest of the configuration by looking at this repository. Now it’s time to create a resource!
|
||
Creating a Resource with a Blueprint To create a resource with a Blueprint, we are essentially going to add Terraform code to an existing repository through a pull request. Resourcely will take the inputs for the Blueprint, render a Terraform configuration, and then create a pull request targeting the repository we specify.
|
||
The pull request will kick off the GitHub Action workflow we just reviewed. Once the change is approved and merged, the actual resource will be deployed through the same workflow.
|
||
Let’s start by going to the Resources section of the Resourcely page and use the Google Cloud Storage Blueprint.
|
||
There should have been a video here but your browser does not seem to support it. I hope you noticed that I went in and manually changed the public_access_prevention argument to inherited. That is still a valid value for the argument, but it doesn’t match our Guardrail. Let’s see what happens during the pull request review.
|
||
Once the pull request has been created, the Resourcely CI process will kick off as part of the workflow. We can check out the results by clicking on the link to the pull request:
|
||
There should have been a video here but your browser does not seem to support it. The Resourcely Guardrails check doesn’t pass because our public access Guardrail failed! Digging into the details, we can see which Guardrail failed and why. Since our Guardrail is set to evaluate only, we can still move forward with the deployment, but I’d rather be in compliance.
|
||
Back in the Resourcely portal, I’ll update my pull request and switch the public_access_prevention argument back to enforced.
|
||
There should have been a video here but your browser does not seem to support it. Now all the checks in the PR pass, and we can move forward with the apply.
|
||
There should have been a video here but your browser does not seem to support it. When the deployment is complete, we now have our compliant resource deployed in Google Cloud. But that was just for enforcing the public access piece, what about limiting the region based on environment?
|
||
Updating the Guardrail and Blueprint Before we add a new Guardrail and apply it to the Blueprint, I need to introduce a new concept, the Global Context.
|
||
In the Global Context area, you can create questions that will be included in the form filled out by the consumer. The questions can be multiple choice, single choice, or free-form text.
|
||
We want to limit production workloads to the US, so we need to know what type of environment is being targeted. I’ve created a single choice Global Context called deploy_environment and added some values to it.
|
||
These options will appear as a drop-down menu in the Blueprints we apply it to.
|
||
With our new Global Context applied to our existing Blueprint, we can create a Guardrail that leverages it.
|
||
Creating a Context Specific Guardrail I want a Guardrail that requires my bucket location to be in the US if the deploy_environment is Production.
|
||
The code for the Guardrail looks like this:
|
||
GUARDRAIL "GCP Storage Allowed Regions" WHEN google_storage_bucket AND CONTEXT deploy_environment MATCHES ANY ["production"] REQUIRE location MATCHES "US-*" OVERRIDE WITH APPROVAL @default The WHEN statement is looking for resources of type google_storage_bucket and checking if the deploy_environment context is in the list ["production"]. If I want to add the same restriction for another environment later, I can simply add it to the list. The REQUIRE statement limits the location to regions that match the string "US-*".
|
||
Like our previous Guardrail, this Guardrail is being applied to all repositories, but this time we have it set to an Active status. Let’s take it for a test drive on our Blueprint.
|
||
Testing the Guardrail Over in the Blueprint, we get to see a different way that Guardrails are applied. I’ll jump over to the Developer Experience tab, and test things out.
|
||
There should have been a video here but your browser does not seem to support it. There is now a Global Context section to the Blueprint. If I set the environment to development and check the location drop-down, it shows me all the regions. Once I change it to production, now the location selection is limited to only US regions. I can override the Guardrail by unlocking it, however during the pull request review it will require an approver to allow the exception.
|
||
Guardrails are not just checked during plan, they also dynamically impact what the consumer sees in the Blueprint form. That shortens the feedback loop for creating compliant infrastructure.
|
||
Summary One of the big challenges of driving IaC adoption is making it easy to consume for others. Sure your platform team knows Terraform inside and out, but how do you get your application teams on board? Give them a simple interface for vending infrastructure and tie it into the source control systems and developer platforms they’re already using. Make it easy to compose infrastructure from Blueprints and enforce best practices through the use of Guardrails. At the same time, include an escape hatch for exceptions, so you don’t block developer productivity unnecessarily.
|
||
I’ve only spent a few days working through what Resourcely has to offer and I know that the product is still evolving and growing. It’s great to see that they already have integrations with developer portals, CI/CD platforms, and ITSM products. Resourcely is not a product that lives on an island, but rather a way to integrate your IaC management into other tools and workflows.
|
||
If you’d like to take Resourcely for a spin yourself, you can use my referral link. I don’t get any kickbacks beyond this sponsored post, but it does let them know you read it and I appreciate that.
|
||
Sponsored Note: This blog post was sponsored by Resourcely. While I did receive compensation, the opinions and content of the post are entirely created and written by me.
|
||
`,summary:"I’ve spent the last eight years working with Terraform on an almost constant basis. Sometimes I think in HCL, and dream of waves of resource creation messages flying past my eyes. But for most people, that is not the case. While I’ve cultivated an expertise in Terraform, I cannot expect others to have done the same. That includes Ops folks, but also developers, networking teams, SREs, and security peeps. They all have responsibilities and skills that aren’t thinking about Terraform 24/7.",date:"15 Nov, 2024",url:"https://nedinthecloud.com/2024/11/15/resourcely-guardrails-and-blueprints/",image:"res-header.png",readingTime:"17"},"https://nedinthecloud.com/2024/11/12/deploying-azure-landing-zones-with-terraform/":{title:"Deploying Azure Landing Zones with Terraform",tags:["hashicorp","terraform","azure"],content:`A couple years ago I discovered the Azure Landing Zone module on the Terraform registry, and I was aghast. What was this nightmare tangle of nested modules with hundreds of resources? What was its purpose? What is an Azure Landing Zone anyway? Finally, I have answers for all these questions and more thanks to Kevin Evans over at Code to Cloud.
|
||
Azure Landing Zones It’s really easy to create an Azure subscription with a credit card and start deploying resources willy-nilly. That’s fine if you’re just experimenting with Azure or spinning up resources for a small company. It doesn’t work so well for an enterprise scale organization that will need multiple subscriptions, separation of duties, and a well-defined hierarchy. That’s the problem that Azure Landing Zones are meant to solve.
|
||
Essentially, Landing Zones create the underlying substrate for you to deploy applications. They involve setting up Management Groups to delegate permissions, applying Azure Policies to enforce best practices, and provisioning subscriptions that are dedicated to identity, management, and connectivity. These form the platform landing zones layer, and provide shared services to application landing zones that will contain your actual applications.
|
||
The chart Microsoft provides is admittedly a bit cluttered. So I’ve simplified it.
|
||
Trust me when I say that there’s a LOT more to Landing Zones than what’s being shown here, but the essentials are pretty much covered.
|
||
ALZ Module The Terraform module to deploy an Landing Zones is absolutely gigantic, with a README that spans into multiple pages and scenarios. There’s a basic 100 level deployment that basically just sets up the Management Group structure and places your subscriptions in the correct place in the hierarchy. That jumps all the way up to a 300 level deployment with a hub and spoke Azure Virtual WAN and zero trust networking.
|
||
To learn more about Landing Zones and deploying them using the Terraform module, I joined Kevin Evans on his YouTube channel Code to Cloud.
|
||
Kevin gave me a primer on what Landing Zones are, how the module is structured, and some prerequisites to get started. Then we went ahead and deployed the whole damn thing. We even ran into a few hiccups that Kevin left in, so you can see that even an Azure expert and a Terraform guru get stymied from time to time.
|
||
Terraform Stacks One of the big issues we ran into was the sequencing of resources. Kevin deliberately chose to roll out the changes incrementally, otherwise they would fail due to resource dependencies. This situation and the complexity of the module made me think of Terraform Stacks, and how they would be a great fit for this exact scenario. Is that something you’d be interested in seeing? Hit me up on LinkedIn or Bluesky and let me know. I think I could adapt the current ALZ Terraform Module to use Stacks without too much difficulty.
|
||
Final Thoughts Thanks to Kevin for inviting me to be on his show and for dispelling the mystery around Azure Landing Zones. That’s been on my to-do list for a while, but honestly I don’t know when I would have gotten to it. Kevin’s invitation served as a forcing function and now I feel comfortable talking about Azure Landing Zones and approaching the module with slightly less trepidation - slightly less.
|
||
Please give the video a watch and give Kevin a subscribe and like!
|
||
`,summary:`A couple years ago I discovered the Azure Landing Zone module on the Terraform registry, and I was aghast. What was this nightmare tangle of nested modules with hundreds of resources? What was its purpose? What is an Azure Landing Zone anyway? Finally, I have answers for all these questions and more thanks to Kevin Evans over at Code to Cloud.
|
||
Azure Landing Zones It’s really easy to create an Azure subscription with a credit card and start deploying resources willy-nilly.`,date:"12 Nov, 2024",url:"https://nedinthecloud.com/2024/11/12/deploying-azure-landing-zones-with-terraform/",image:"azure-landing-zones.png",readingTime:"3"},"https://nedinthecloud.com/2024/10/18/hashiconf-2024-wrap-up/":{title:"HashiConf 2024 Wrap-Up",tags:["hashicorp","terraform","vault","hashiconf"],content:`As I write this, I’m on my way home from HashiConf 2024, hosted in the great city of Boston.
|
||
This year’s HashiConf was the 10th of its kind and the first to be hosted on the East Coast- also know as the best coast or beast coast. I thought I would take a moment to collect my thoughts on the event as a whole, the announcements made during the keynote, and particularly awesome sessions I attended. Since I’m guessing you’re already chomping at the bit- or chopping, I’m not judging what you do with your bits- I’ll go over the big announcements.
|
||
Major Announcements HashiConf took place over two days, with Tuesday being focused on Infrastructure Lifecycle Management (ILM) and Wednesday shifting left into Security Lifecycle Management (SLM). I truly appreciate how HashiCorp has focused and simplified its messaging in the last few years. It’s helped to sharpen the product portfolio and direct feature development.
|
||
Infrastructure Lifecycle Management The biggest announcement of day 1 was the public beta of Terraform Stacks on HCP Terraform. Stacks was announced as an upcoming feature at HashiConf 2023, and over the last 12 months the Stacks team has been working tirelessly to get the feature ready for public testing.
|
||
I was part of the private testing and due to the restrictions involved, I was absolutely not allowed to talk about it. Which is sad, because I think Stacks is a game changer for anyone using Terraform at scale and I really wanted to talk about it. But now, I can! What are these mysterious Stacks? That might take a moment to explain.
|
||
HCP Terraform has the concept of workspaces to separately manage environments. Within a workspace, a Terraform configuration describes the intended state of the world and an instance of state data records the results of the most recent apply. A workspace is essentially a logical unit of work and an administrative boundary on HCP Terraform.
|
||
There is a tension when it comes to workspace design. On the one hand, you’re trying and keep related resources together. The contents of a workspace should share a common lifecycle and be tightly coupled. Within a Terraform configuration it’s trivial to express dependencies through implicit references or explicit dependencies.
|
||
One the other end of the tension is a desire to: minimize the size of state, include only resources that share a common administrative boundary, and limit blast radius due to changes. For instance, shared networking resources probably shouldn’t be in the same workspace as the applications that consume it. There’s probably a different team managing the shared network; the lifecycle of the network is not tightly coupled to the applications; and a change in an application workspace should not blow everything up for other apps or the network as a whole.
|
||
Thus it makes sense to separate resources into different workspaces based on blast radius, admin boundary, and separation of duties. However, there are still dependencies between resources in separate workspaces, and it is far more difficult to express those dependencies and references across the workspace boundary. For instance, if the subnet ID of a network resource changes in workspace A, how does workspace B know it needs to update the network cards that use the subnet? And how does it get the value of the new subnet ID?
|
||
You can use remote state or query via data sources to get actual values. And you can use run triggers to watch for changes to the state of one workspace and have that cascade to other workspaces. That feels a little like a reactive kludge though, and orchestrating changes across multiple workspaces is a fraught endeavor.
|
||
Compounding these challenges is the need to have multiple instances of the same configuration for different environments, e.g. Dev, QA, Prod etc. Each follows the same patterns, but needs to be configured separately.
|
||
On top of all that, there are times when a roll out has to be sequenced through glue scripts, because values are not yet known. One obvious example is provisioning a Kubernetes cluster and then trying to bootstrap the cluster. Since the cluster doesn’t exist during the initial deployment, the planning phase will throw an error or timeout trying to reach a Kubernetes API endpoint that doesn’t exist.
|
||
Stacks solves all of these problems by introducing a new organizing principle- the titular Stack- two new core constructs: Components and Deployments. Each component is a module reference with inputs, providers, and dependencies. Components are combined in a declarative way that allows you to express the relationship between the layers of the Stack. Deployments are an instance of the stack you want to deploy with the opportunity to supply unique inputs for that environment.
|
||
The Stacks engine on HCP Terraform understands that sometimes a full plan for the entire stack is not possible, since what is in one layer may rely on resources that do not yet exist in another layer. The resulting plan will contain “deferred changes” which will be planned and applied on following runs once the missing resources are in place. You can also create orchestration rules that automate the application of deferred changes.
|
||
That’s a very quick breakdown of Stacks, and if you want a more thorough explanation, check out this video from Sarah Hernandez.
|
||
As I said, Stacks is now in public beta, so go kick the tires yourself! It should be available to any organization on one of the current crop of plans (including Free.)
|
||
Security Lifecycle Management Terraform Stacks may have stolen the show on the ILM side, but the SLM side had two products that I think deserve equal airtime: Vault Radar and Vault Secrets.
|
||
As a quick side note, HashiCorp Vault has morphed from a single stand-alone product to a suite of products under the Vault Umbrella. A similar thing happened to Microsoft’s Azure Stack product. I didn’t like it then, and I don’t think I like it now, but naming is hard and I doubt I could do any better. The original HCP Vault product is now referred to as HCP Vault Dedicated. I guess OG Vault is Vault Server now? It’s unclear.
|
||
Vault Radar Before I talk about Vault Radar, allow me to tell a little story about when I did something stupid. As you may already know, I’ve created several courses and demos around using Terraform with AWS. In some of those demos, I show how you can hard-code your AWS credentials into a Terraform configuration, and then I tell you to NEVER DO THAT.
|
||
To my great embarrassment, after creating those demos I have then accidentally committed my code to GitHub, including those hard-coded credentials 🤦🤦🤦. Even worse? I’ve done it at least three times.
|
||
Fortunately, GitHub constantly scans repository commits for AWS keys and alerts AWS when it happens. You’ll quickly get a nastygram from AWS that your credentials have been compromised and the user involved will have their permissions and roles reduced until you revoke the credentials, chastise the user, and make a proper blood sacrifice under a harvest moon.
|
||
Checking access credentials into source control is hardly an uncommon event, and it’s only one of many ways that secrets are placed in vulnerable locations. If you’re a security professional trying to get your arms around secrets protection and lifecycle management, you need something to help you find and identify secrets across your organization wherever they might live.
|
||
Basically, that’s what Vault Radar does. It constantly scans the sources you provide it access to, and it finds things that look like secrets and lets you know. Vault Radar is officially in public beta now, so you can go try it out for yourself.
|
||
As part of the public beta, HashiCorp announced that Vault Radar will support agents to scan for secrets behind your firewalls on internal systems. Vault Radar also includes a pre-commit hook that developers can implement to catch potentials secrets before they are even committed to git. I sure could have used that a couple years ago!
|
||
Vault Secrets Once you’ve identified where secrets are being stored, you may wish to remove them or manage them with a more robust solution than asking Developer Rick to change them once a quarter. You know Rick is pretty busy and is probably going to forget.
|
||
One option is to roll out Vault Server or HCP Vault Dedicated and force your developers to store all secrets securely in Vault. I’m sure they have nothing more pressing and can dedicate the next couple months integrating Vault into all their workflows and application logic. Seems like a piece of cake. I know Rick said he’s busy, but we both know he’s just addicted to Candy Crush.
|
||
Okay, so maybe forcing all your developers to drop what they’re doing and implement the Vault API in their code isn’t very feasible. What if you could meet them in the middle? Or even better, what if you could manage and rotate their secrets for them and they didn’t have to change anything? That’s what Vault Secrets is all about.
|
||
Vault Secrets is a service on HCP that allows you to manage secrets centrally and synchronize them to targets like AWS Secrets Manager, Azure Key Vault, and Kubernetes. Hopefully your developers are using one of these services to store their secrets and not simply writing them to a random CSV on a file share. Vault Radar can help you figure that out.
|
||
Assuming that they are using a service like GitHub Secrets, they will not need to change their workflow or application. Vault Secrets can rotate secrets on demand and synchronize the new value to the service being used by an application or pipeline. Developers can also pull the secret value directly if they aren’t using a secrets service today, which they should be, but Rick is like so close to hitting level 600.
|
||
HCP Vault Secrets has been GA since last year. At HashiConf 2024, they added two new features I think are super cool. The first is the general availability of auto-rotation for certain secret types. While Vault Secrets could already rotate secrets, it was an on-demand affair. Auto-rotate will automatically create a new version of the secret and keep both in force for a period of time you choose, giving applications a chance to refresh to the new secret value. When the period elapses, the older version of the secret is revoked.
|
||
The second big feature is dynamic secrets for certain secret types. Just like dynamic secrets in Vault Server, dynamic secrets on Vault Secrets are created on demand for an application and have a limited lifetime. For instance, you may have a pipeline that needs temporary credentials for AWS. Vault Secrets can provision credentials on demand that are good for the next 30 minutes, after which time they expire. If another pipeline process requests credentials, it will get a separate, unique set of credentials that are also only good for 30 minutes. This helps limit the blast radius for leaked credentials and simplifies tracking credential leaks. Dynamic secrets is now in public beta and supports AWS and GCP with Azure support coming later this year.
|
||
But Wait, There’s More Terraform Stacks, Vault Radar, and Vault Secrets are not the only things announced during the two keynotes. If you want a full rundown, you can check out the official HashiCorp blog posts for ILM and SLM. As far as I’m concerned, though, these three items are the most significant and impactful.
|
||
Great Sessions HashiConf is always a very busy time for me. I struggle to walk from one end of the expo to the other without being stopped a half-dozen times. Nevertheless, I did manage to attend more sessions that just the two keynotes. All the sessions I mention below will be posted to the HashiConf site and YouTube in due time, and I will add links to the post once they’re available.
|
||
Microsoft Azure Stuff I went into the Microsoft Azure and Terraform session thinking I knew what was going on with Azure and Terraform. And I was pleasantly surprised to discover I was wrong. Turns out Microsoft has been hella busy while I wasn’t looking.
|
||
Mark Gray and Steven Ma demoed several features that are going into private preview shortly. There are three I’d like to focus in on: Terraform Export, Portal Copilot, and VSCode Extensions.
|
||
Terraform Export is a feature being added to the portal to help export an existing resource or set of resources to Terraform. When you’re looking at a resource today, there’s an Automation section of the left-hand menu that includes the Export option. That option will generate an ARM template that matches the deployed resource. But ARM templates are awful and I hate them. The preview feature Mark demoed included tabs for Bicep and Terraform. You could already do something similar with the Azure Export for Terraform command line tool, but now it will be directly in the portal and allegedly give more robust results.
|
||
Speaking of better results, Copilot in Azure has now been trained on the AzureRM and AzAPI providers and Terraform documentation. Generic LLMs can give pretty inconsistent results- read: terrible- when it comes to generating Terraform code, and by inconsistent I mean completely unusable and hopelessly out of date. That’s because ChatGPT hasn’t specifically been trained to produce valid Terraform and doesn’t have the resource definitions handy to guide it. Copilot in the Azure portal now has that additional context, and should be able to generate better Terraform code that actually includes proper input variables and real resource naming.
|
||
In a similar vein, the updated VSCode Extension will offer a conversational interface to get the same results without having to leave your IDE. I try to avoid the portal when I can, so this is a welcome update.
|
||
If you’re interested in trying out these preview features and more, join the Azure Terraform Community at the link!
|
||
Boundary Self-service with GitHub Actions Mattias Fjellstrom gave an excellent presentation on how he leverages GitHub Actions and GitHub Issues to enable self-service for Boundary access. In case you didn’t already know, Boundary provides reverse proxy access to remote systems with dynamically injected credentials and scoped permissions.
|
||
When a user requests access to a remote system, Boundary verifies they have permission to the host and then works with Vault to dynamically generate and inject credentials into the session. But what if a user doesn’t have access to a system and they need it?
|
||
They could raise a ticket and wait for someone to grant access, and then remind that person to revoke their access when they’re done. Or, you could set up a self-service system that provides just-in-time access and revokes that access after a defined period. That would be pretty cool right?
|
||
That is exactly what Mattias built. When a user needs access to a system, they create a GitHub issue. That kicks off a GitHub Action granting them access via a Terraform run. When the time period expires, another GitHub Action fires to remove access. Alternatively, the user can close the issue, which will trigger access termination as well. Mattias even set up an approval workflow for systems marked as sensitive.
|
||
If you want to check it out for yourself, he’s written a thorough blog post.
|
||
Kubernetes Vault Operator at Scale When I was walking out of the Speaker Room on Wednesday morning, I held the door for Anthony Ralston, and I could just tell by the look in his eye that he was going ot deliver a killer talk. Call it confirmation bias, but I was absolutely right.
|
||
Anthony works for Canva- which I use on a daily basis- and he wanted to talk about how they deployed the Vault Operator on their Kubernetes clusters to leverage dynamic secrets for Datadog.
|
||
But he didn’t start with that, instead he started with a problem statement. Every time a new Kubernetes cluster was built, a Datadog API key had to be generated and added to the cluster so they could collect and send metrics through the Datadog agent. To rotate the API key, they had to generate and replace it manually for each cluster. Both processes were time consuming and inefficient. Canva was already using Vault heavily for storing and generating secret values, so they thought perhaps Vault could provision and distribute Datadog API keys dynamically.
|
||
Anthony walked through their decision process and included some of the dead ends they hit along the way. It was a nuanced talk that not only highlighted the power of Vault, but also the importance of understanding your requirements and limitations ahead of time and working towards a solution that addresses the actual business problem. Deploying tech for tech’s sake is not a valid justification for a project.
|
||
I also learned about the differences between the Vault CSI, Vault Secrets Injector, and Vault Operator. Each solution has its use cases, and it was informative to see them compared in a real world context. Who won out for Anthony’s team? You’ll have to watch the presentation to find out.
|
||
General Thoughts I look forward to HashiConf every year. Honestly, its the highlight of conference season for me, and the only conference I’ll be attending this fall. Why? Community and scale.
|
||
I’ve been to KubeCon and re:Invent and Microsoft Ignite. And you know what? Conferences with 20k+ attendees kinda suck. Everything is overwhelming: the crowds, the sessions, the venue, the crowds, the lines, the crowds. Are you getting the sense I’m not big on crowds? You would be correct.
|
||
There is something about the massive scale of those events that feels necessarily impersonal. You are a badge to be scanned. An attendee to be counted. Sheep to be herded around from session to session. Smaller conferences like HashiConf are nothing like that. People know who you are. You can easily get to every session you want to. You don’t have to show up 15 minutes early to get a seat. You can make real connections with other people.
|
||
I’m not trying to yuck anyone’s yum here. If you love going to re:Invent, then more for you. I’ve always been one to prefer a more intimate setting. I went to a small high school, small college, and worked at small businesses. I prefer a show at a 200 person venue to a stadium concert. I’d rather do a trail race with 100 people than the Broad Street Run in Philly with 30k strangers. Every time I try joining some large scale organization, it just doesn’t fit me. Hell, I’ve worked for myself for over five years. You cannot work for a small organization than that.
|
||
My main point is that I enjoy conferences with a personal feel and a vibrant caring community. I value substance over spectacle and quality over quantity. My two favorite events this year were HashiConf and DevOpsDays Philly. Both are built on a foundation of a supportive community that is accepting of all kinds.
|
||
As HashiCorp says in their code of conduct: For Everyone, Everywhere.
|
||
Sponsored Note: This blog post was sponsored by HashiCorp. The opinions and information in the post are mine alone and where not reviewed or edited by HashiCorp.
|
||
`,summary:`As I write this, I’m on my way home from HashiConf 2024, hosted in the great city of Boston.
|
||
This year’s HashiConf was the 10th of its kind and the first to be hosted on the East Coast- also know as the best coast or beast coast. I thought I would take a moment to collect my thoughts on the event as a whole, the announcements made during the keynote, and particularly awesome sessions I attended.`,date:"18 Oct, 2024",url:"https://nedinthecloud.com/2024/10/18/hashiconf-2024-wrap-up/",image:"hashiconf-2024.png",readingTime:"16"},"https://nedinthecloud.com/2024/08/27/whats-new-in-the-azurerm-provider-version-4/":{title:"What's New in the AzureRM Provider Version 4?",tags:["hashicorp","terraform","opentofu","azure"],content:`Major version 4.x of the AzureRM provider was just released! What’s new, what’s different, and what do you need to know? That’s what I’m going to cover in this post.
|
||
Introduction The AzureRM provider has been on major version 3 for two and a half years! Since that release, a ton of new resources and data sources have been added to the provider and they’ve kept up the pace of releasing a new minor version of the provider roughly every two weeks. As such, the most recent minor version for major version 3 is 116. As in one-hundred and sixteen minor release 🤯. A minor release does not have breaking changes, meaning all significant changes to the provider have been waiting until now to be implemented!
|
||
So what are the big, breaking changes? They fall into three categories:
|
||
Provider property changes Removed resources and data sources Removed or altered properties of resources and data sources There’s also some new functionality added to the provider which doesn’t break existing code, but enhances the use of the provider. In particular, there are two changes:
|
||
Provider-defined functions Resource provider registration Why don’t we start with the new functionality and then get into breaking changes?
|
||
New Functionality Provider-defined Functions Provider-defined functions were introduced to Terraform in version 1.8. Of the providers adopting this feature, azurerm was notably absent at launch. That’s not because HashiCorp didn’t want to add provider-defined functions to AzureRM, but rather because the versions of the provider SDK being used made it difficult. Major version 4 adopts the Terraform plugin framework for some portions of the provider, which brings with it support for provider-defined functions.
|
||
There are two new provider-defined functions: normalise_resource_id and parse_resource_id. As you probably already know, Azure resource IDs are really long strings that uniquely identify a resource. They include the subscription, resource group, resource provider type, parent resource, etc. While you could try and manipulate with these strings using the built-in functions of Terraform, the two new provider-defined functions make it a bit easier.
|
||
The normalise_resource_id function takes a resource ID and attempts to normalize the capitalization of its elements to match what the Azure API is expecting. The Azure API is case sensitive (sometimes) and if you’re constructing a resource ID by hand or receiving it from another source- like an input variable- you may want to normalize it before using it.
|
||
variable "vnet_resource_id" { type = string description = "Resource ID of VNet" } locals { vnet_resource_id = provider::azurerm:normalise_resource_id(var.vnet_resource_id) } The above code would create a normalized version of a VNet resource ID for use in the rest of your code. This is also super helpful for import blocks, where you have to supply a resource ID to successfully import your resource:
|
||
import { to = azurerm_container_app.example id = provider::azurerm::normalise_resource_id(var.container_app_id) } The parse_resource_id function breaks a resource ID into its component parts. The output is a map with the following keys:
|
||
full_resource_type parent_resources resource_group_name resource_name resource_provider resource_scope resource_type subscription_id Depending on the resource type, some of the values may be set to an empty value. For instance, an azurerm_virtual_network would have the following fields set:
|
||
parsed_id_vnet = { "full_resource_type" = "Microsoft.Network/virtualNetworks" "parent_resources" = tomap({}) "resource_group_name" = "function-example-rg" "resource_name" = "function-example-vnet" "resource_provider" = "Microsoft.Network" "resource_scope" = "" "resource_type" = "virtualNetworks" "subscription_id" = "XXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXX" } Since a VNet doesn’t have parent resources, the parent_resources field is empty. This function is very similar in nature to the arn_parse function in the aws provider.
|
||
Resource Provider Registrations The Azure Resource Manager control plane is built on the idea of Resource Providers (RP) which are responsible for the lifecycle of resources under a certain category. For instance, the VNet shown above is of resource type Microsoft.Network/virtualNetworks, which means it uses the Microsoft.Network Resource Provider.
|
||
When you create a subscription in Azure, a subset of Resource Providers are registered for use by default. You can find the full list here, just look for any namespace that has “registered by default” appended to it. Most of the default services are core to the platform, like microsoft.support or Microsoft.Billing. Bonus points for inconsistent capitalization! Good thing there’s a function to help with that 😂.
|
||
All other Resource Providers need to be registered before you can create or manage resources that leverage the RP. When you deploy a resource using the Portal or an ARM template, Azure checks to see if the necessary RP is registered. If it isn’t, Azure Resource Manager registers the RP for you (assuming you have permissions to do so.)
|
||
To avoid potential deployment errors, the default behavior of the azurerm provider is to register every single RP that isn’t already registered. That’s… unnecessary and actually goes against Microsoft’s guidance:
|
||
“Register a resource provider only when you’re ready to use it. This registration step helps maintain least privileges within your subscription. A malicious user can’t use unregistered resource providers.”
|
||
Not only that, but if the account you’re using doesn’t have permissions to register RPs, you’ll get an error back from Terraform. The workaround for previous versions of the azurerm provider was the argument skip_provider_registration. When set to true, Terraform wouldn’t try to register any RPs, which could result in a deployment error.
|
||
Clearly this was a very blunt solution, and version 4 of the azurerm provider takes a slightly more nuanced approach. There are two new arguments for the provider: resource_provider_registrations and resource_providers_to_register.
|
||
The resource_provider_registrations argument can be set to a pre-defined group of RPs to automatically register. Currently the allowed values are core, extended, all, none, and legacy. I would imagine more may be added over time. The current list of RPs included in each group can be found on GitHub.
|
||
The resource_providers_to_register argument takes a list of Resource Providers to register. This can be used in tandem with resource_provider_registrations to register additional RPs that are not part of the grouping. If you want to be REALLY specific, you could set resource_provider_registrations to none and only list the RPs your want registered with resource_providers_to_register. Is it worth doing that? That’s for you to decide.
|
||
With the introduction of the two new arguments, the skip_provider_registration argument has been removed. If you want the same functionality, simply set resource_provider_registrations to none.
|
||
This is a breaking change if you were using the skip_provider_registration before. Speaking of breaking changes…
|
||
Breaking Changes Subscription ID Required Possibly the biggest breaking change is making Subscription ID a required argument in the provider. For anyone who uses Azure CLI based authentication for their configurations, you’ll now get an error when you attempt to run terraform plan or terraform apply:
|
||
Planning failed. Terraform encountered an error while generating this plan. ╷ │ Error: \`subscription_id\` is a required provider property when performing a plan/apply operation When using the Azure CLI for authentication, the default behavior was to use the credentials and currently selected subscription from the CLI for Terraform deployments. If you wanted to use a different subscription, you could run az account set -s SUBSCRIPTION_NAME and Terraform would use that one.
|
||
Did this lead to accidental deployments in the wrong subscription? Almost certainly. The implicit relationship between the Azure CLI and the provider made it all too easy to run terraform apply on the wrong subscription and inadvertently take down Prod when you though the Dev subscription was selected.
|
||
Version 4 makes the relationship explicit by requiring that you set the target subscription either through the subscription_id argument in the provider block or through the ARM_SUBSCRIPTION_ID environment variable. This only impacts Azure CLI based authentication, as every other authentication type requires you to set the subscription ID explicitly.
|
||
I can hear the clatter of a thousand Terraform tutorials all being updated at once, and I probably have to go back and make sure none of my example code is broken by this change. This is why its so important to pin the version of your provider to a at least a major version if not more specific.
|
||
Resources and Data Sources I’m not going to list out every single resources or data source that has been removed or altered in a breaking way. The 4.0 Upgrade Guide does an excellent job of that. But I do want to point out some significant changes that I think are worth mentioning.
|
||
Azure Kubernetes Cluster There’s two big things to note with AKS. First, the azurerm_kubernetes_cluster resource will only use the stable version of the AKS API. There are two versions of the API: stable and preview. Previously, the AKS cluster resource used a mix of the two APIs to allow cutting-edge and experimental features to be available to the resource. The AKS team politely asked Microsoft and HashiCorp to knock it off and only use the stable API. You know, because it’s the stable one. This means that certain arguments will no longer be available in version 4.0 of the provider.
|
||
If you wish to keep using the preview AKS API, you can deploy your AKS cluster using the azapi provider. That gives you full access to any Azure APIs exposed to the public, including experimental and preview versions.
|
||
The other big change is a bunch of removed and changed arguments in the azurerm_kubernetes_cluster resources. There are 23 changes in total including removals, renaming, and accepted values. If you’re managing an AKS cluster with Terraform, you’re going to want to pay special attention to the provider changes!
|
||
Database Services There are a ton of database related resources that are being replaced. Basically, all resources that start with azurerm_sql are being replaced with resources that start with azurerm_mssql, and all the resources starting with azurerm_mariadb or azurerm_mysql are being replaced with azurerm_mysql_flexible resources. If you happen to be using these older resources, some have been superseded and others are being retired entirely.
|
||
List to Set The azurerm_subnet and azurerm_virtual_network resources have arguments that are changing the accepted value type from a list to a set. As a quick reminder, a set in Terraform is an unordered collection of unique elements. Since it is unordered, you cannot refer to an individual element using an index, e.g. local.my_list[0]. Instead, you’ll be using a for expression to loop through the contents of the set. Also, when passing a list to these arguments, you’ll need to use the toset() function to convert it to a set.
|
||
These changes may require rewriting some of your existing code to handle the difference between a set and list.
|
||
Conclusion It has been two and a half years since the last major version of azurerm! 30 MONTHS! I’m pretty sure in that time the aws provider had two major version releases. Clearly there was a lot of work and thought put into this newer version of the provider, with an eye towards managing change going forward.
|
||
I’m really excited to see the AzureRM provider use the newer plugin framework, and I think it will simplify the maintenance of the provider in the future. With any luck that means quicker additions of new resources and properties when they’re released by Microsoft.
|
||
`,summary:`Major version 4.x of the AzureRM provider was just released! What’s new, what’s different, and what do you need to know? That’s what I’m going to cover in this post.
|
||
Introduction The AzureRM provider has been on major version 3 for two and a half years! Since that release, a ton of new resources and data sources have been added to the provider and they’ve kept up the pace of releasing a new minor version of the provider roughly every two weeks.`,date:"27 Aug, 2024",url:"https://nedinthecloud.com/2024/08/27/whats-new-in-the-azurerm-provider-version-4/",image:"azurerm-version-4.png",readingTime:"9"},"https://nedinthecloud.com/2024/08/20/debugging-the-azurerm-provider-with-vscode/":{title:"Debugging the AzureRM Provider with VSCode",tags:["hashicorp","terraform","opentofu","azure"],content:`Recently I’ve been learning Go and applying it to technologies that I work on everyday, like Terraform. For instance, check out my Terrahash CLI tool for validating modules. It’s neat!
|
||
I’ve also taken an interest in dipping my toe into the world of Terraform providers, and what better provider to start with than the azurerm provider I use for sooooooo many demos?
|
||
Of course, working on provider issues means debugging the provider itself and I had to learn how to set up debugging for VSCode. I thought I would document the process so you can see what I set up and possibly debug some providers yourself.
|
||
Shoutout to Drew Mullen and his excellent post on debugging the aws provider with VSCode. The azurerm provider is a little different, but the process is very similar.
|
||
Prerequisites Before you get started with provider debugging on VSCode, you’ll need a few things in place:
|
||
VSCode (I feel like that’s obvious, but 🤷) Go (Also kind of a given) VSCode Go extension (this includes delve which is necessary for debugging) Forked instance of the azurerm provider cloned to the correct directory The correct directory being $GOPATH/src/github.com/hashicorp/terraform-provider-azurerm Even though I run Windows, I prefer to do my Go development in WSLv2. In my experience, most Go tooling is intended for Mac and Linux with Windows as an afterthought. My examples will be using WSLv2, so fair warning on that.
|
||
How It Works Since you’re debugging the provider and not the Terraform binary itself, you need to let Terraform know it should send provider requests to a listener on the debugger and not to the usual provider plugin. Terraform uses gRPC to communicate with providers, so it’s a simple matter of launching the provider in debugging mode and delve will handle setting up a gRPC listener on a UNIX socket.
|
||
Once the debugger finishes launching, it will produce a value for you called TF_REATTACH_PROVIDERS, which you can set as an environment variable for simplicity’s sake. Terraform looks for that environment variable and uses it to redirect provider plugin requests to the debugger.
|
||
After that, you can simply run Terraform commands like normal and if that command calls the provider and hits breakpoints in your code, it will pause and wait for you to resume.
|
||
The official Terraform docs for debugging providers uses a lot of fancy terminology that was mostly foreign to me. So I tried to simplify it down for someone who is not super familiar with Go or debugging, e.g. me.
|
||
Setting Up Debugging VSCode gets its debugging settings from a file stored inside the .vscode directory in your repository. So first we need to create that directory and then a couple files.
|
||
From the root of the azurerm repository, run the following commands:
|
||
mkdir .vscode touch .vscode/launch.json touch .vscode/private.env The launch.json file should contain the following:
|
||
{ "version": "0.2.0", "configurations": [ { "name": "Debug AzureRM Provider", "type": "go", "request": "launch", "mode": "debug", // this assumes your workspace is the root of the repo "program": "\${workspaceFolder}", "env": {}, "args": [ "-debuggable", ], "showLog": true, "envFile": "\${workspaceFolder}/.vscode/private.env" } ] } And the private.env should contain this:
|
||
TF_ACC=1 TF_LOG=INFO GOFLAGS='-mod=readonly' You can change the TF_LOG level if you want, but you’ll be getting all your info from the debugging process.
|
||
Using the Debugger Once you have the files created, VSCode will show the name of the debug configuration in the debug tab:
|
||
Click on the green arrow to start the debugging process, and then bring up the DEBUG CONSOLE from the tabs on the integrated terminal.
|
||
You should see Delve starting up and listening on a local port:
|
||
Starting: /home/ned1313/go/bin/dlv dap --log=true --log-output=debugger --listen=127.0.0.1:44335 --log-dest=3 from /home/ned1313/go/src/github.com/hashicorp/terraform-provider-azurerm DAP server listening at: 127.0.0.1:44335 It will take a while for the debugger to actually start, but once it does, you should see output like this:
|
||
TF_REATTACH_PROVIDERS='{"registry.terraform.io/hashicorp/azurerm":{"Protocol":"grpc","ProtocolVersion":5,"Pid":293831,"Test":true,"Addr":{"Network":"unix","String":"/tmp/plugin1338197622"}}}' This is the environment variable you need to set to let Terraform know that it should redirect requests for the azurerm provider to the plugin listening at /tmp/plugin#######. The actual values will vary depending on the version of the provider and your system.
|
||
From the configuration you’d like to troubleshoot, first export the TF_REATTACH_PROVIDERS variable:
|
||
export TF_REATTACH_PROVIDERS='{"registry.terraform.io/hashicorp/azurerm":{"Protocol":"grpc","ProtocolVersion":5,"Pid":293831,"Test":true,"Addr":{"Network":"unix","String":"/tmp/plugin1338197622"}}}' Then run your Terraform commands as usual. For instance, I have the following configuration:
|
||
terraform { required_providers { azurerm = { source = "hashicorp/azurerm" version = "~>3.0" } } } provider "azurerm" { features { } } resource "azurerm_resource_group" "example" { name = "debug-testing" location = "eastus" } I’ll run the following commands:
|
||
terraform init
|
||
$ terraform init Initializing the backend... Initializing provider plugins... Terraform has been successfully initialized! You may now begin working with Terraform. Try running "terraform plan" to see any changes that are required for your infrastructure. All Terraform commands should now work. If you ever set or change modules or backend configuration for Terraform, rerun this command to reinitialize your working directory. If you forget, other commands will detect it and remind you to do so if necessary. terraform plan
|
||
$ terraform plan azurerm_resource_group.example: Refreshing state... [id=/subscriptions/4d8e572a-3214-40e9-a26f-8f71ecd24e0d/resourceGroups/debug-test] I added a breakpoint in the provider.go file to interrupt the process.
|
||
Once I continue through the breakpoint, the process continues and finishes.
|
||
Terraform will perform the following actions: # azurerm_resource_group.example will be created + resource "azurerm_resource_group" "example" { + id = (known after apply) + location = "eastus" + name = "debug-testing" } Plan: 1 to add, 0 to change, 0 to destroy. And that’s it! You’re now debugging the azurerm provider. From here you can add breakpoints, inspect values, and look at the stack calls. I’ve found this invaluable for tracing down where an issue is occurring between Terraform core, the provider, and the Azure API.
|
||
Good luck and happy bug squashing!
|
||
`,summary:`Recently I’ve been learning Go and applying it to technologies that I work on everyday, like Terraform. For instance, check out my Terrahash CLI tool for validating modules. It’s neat!
|
||
I’ve also taken an interest in dipping my toe into the world of Terraform providers, and what better provider to start with than the azurerm provider I use for sooooooo many demos?
|
||
Of course, working on provider issues means debugging the provider itself and I had to learn how to set up debugging for VSCode.`,date:"20 Aug, 2024",url:"https://nedinthecloud.com/2024/08/20/debugging-the-azurerm-provider-with-vscode/",image:"debugging-azurerm-provider.png",readingTime:"5"},"https://nedinthecloud.com/2024/08/01/state-encryption-with-opentofu/":{title:"State Encryption with OpenTofu",tags:["hashicorp","terraform","opentofu"],content:`Wondering how to encrypt your state data using OpenTofu? Then this is the post for you!
|
||
If you’d prefer your information in video format, check out the Terraform Tuesday video instead.
|
||
Introduction For this post, we are going to return once again to OpenTofu. If you’re not familiar with OpenTofu and how it differs from Terraform, then you may want to check out the blog post I’ve been maintaining that compare the differences between the two.
|
||
In this post, I wanted to dig into the data encryption feature that was introduced with OpenTofu 1.7. We’ll look at the feature’s history, what problem it is attempting to solve, and how to add or remove encryption on your state data and plan files.
|
||
If all you care about is the actual nuts and bolts, jump to this section, but I do think you need to consider whether this feature is right for you first. That requires some context and history.
|
||
Why Encrypt Your State Data? Before OpenTofu launched, one of the most popular requests on the Terraform list of open issues was to support encryption of state data and plan files at rest. And the issue languished for years, despite some possible solutions being put forth. HashiCorp, as maintainers of the Terraform repository, argued that using Terraform to encrypt state data directly was not a sound idea in practice and plan files should be ephemeral. Rather than implement a solution they chose to leave the issue open.
|
||
One of the big promises of OpenTofu was that this feature would finally be added and the issue resolved. And in OpenTofu 1.7 it was! That’s quite an accomplishment and I have to give the OpenTofu team their flowers for delivering. I also feel I should mention that this feature is not compatible with Terraform and you would have to remove encryption if you wanted to migrate to Terraform or have a third-party tool interact with your state data or plan files directly.
|
||
Why were people asking for this feature? Mostly because state data and plan files can contain sensitive information that you don’t want others to have access to. That could be obvious stuff like API keys, passwords, and tokens, but it could also be the less obvious fact that anyone who gets a copy of your state data or a plan file has a map of your deployed infrastructure. They might be able to use that map to develop a plan of attack or find weaknesses to be exploited.
|
||
Having OpenTofu encrypt state data and plan files at rest means that even if someone gains access to wherever your state data is stored, they still won’t be able to read it without the necessary keys.
|
||
What other options are there? In defense of HashiCorp, there are lots of ways to encrypt state data at rest that do not rely on managing your own encryption keys. The most popular state backends all support encryption of state at rest and in transit, and allow you to apply robust access permissions.
|
||
Azure storage, for instance, encrypts data by default, and you can bring your own key for both data encryption and infrastructure encryption. You can apply access control through a number of mechanisms, whether its SAS tokens, Azure AD RBAC, or attribute-based access control. If you want to know just how much you can lock down Azure storage, check out my video on the topic.
|
||
So aside from adding another layer of complexity and more keys to manage, what does this additional encryption facility buy you? Well, if you don’t trust the backend you’re using to properly store and protect your state data, then encrypting state with OpenTofu removes the onus from the platform and places it squarely on you. Especially if you’re using some generic HTTP backend for state, you may was to manage the encryption yourself.
|
||
It also provides a way for you to encrypt your state data if its being stored locally. I know we should all be storing our state on a remote backend, but that’s not always the reality. You could use another encryption solution to manage encryption at rest on your filesystem, but OpenTofu makes it so you can rely on the encryption provided directly by OpenTofu itself.
|
||
The option to encrypt plan files is the real benefit in my estimation. When you save a plan file for review and execution, it contains a ton of information, including planned changes, new values, and updated outputs. If that plan file falls into the wrong hands, even if the plan itself is out of date, an attacker will have more than enough information to launch a attack on your infrastructure. And that’s assuming the API keys weren’t passed directly as input variables.
|
||
Again, there are existing ways to protect against this. Depending on the state backend you’re using, you could stash you plan files next to your state data. The same encyprtion and security you’ve applied to your state would now apply to those plan files. Or you could stash it on any respectable object-based storage and apply the necessary encryption and permissions. However, that’s not baked into Terraform in any way. That’s something you would need to automate yourself or get from a TACO platform.
|
||
All that being said, if I had to pick a winning use case for OpenTofu’s encryption feature, it’s encrypting plan files. Let’s see how it all works!
|
||
Encrypting State Data With OpenTofu First let’s lay down some terminology and syntax. The encryption settings for state and plan data can be configured in two ways: using an encryption block inside of the terraform block, through the environment variable TF_ENCRYPTION, or a combination of the two. You can use either JSON or HCL to describe the encryption settings, though I think you already know my general feelings on JSON 🤢.
|
||
Encryption requires two components, a key provider and a method. The key provider defines where OpenTofu is getting the key material to perform the encryption, and the method describes how the key is used to encrypt the data. Wow, it’s like right there in the name isn’t it?
|
||
The method is used to apply encryption to state data and plan files. State and plan encryption is managed separately, so you can choose to use one form of encryption for state and another for plan. Or you can choose to apply encryption to one and not the other. Hooray for options!
|
||
So what kind of key providers are available right now? As of writing there are four:
|
||
pbkdf2 - is used to generate a key locally using a passphrase. The passphrase is stored with the configuration or injected with the TF_ENCRYPTION environment variable. aws_kms - uses the AWS KMS service for the encryption key. gcp_kms - uses the Google Cloud KMS service for the encryption key. openbao - uses the transit engine in OpenBao for the encryption key. And if you’re wondering what the hell OpenBao is, it’s the open source fork of HashiCorp Vault. The OpenBao key provider will also work with Vault versions under the MPL license, which is 1.14 or older.
|
||
Azure’s Key Vault is not yet on the list, but I’m sure that’s being worked on right now. Plus, this is open source, so if your preferred encryption service isn’t on the list, you can implement it yourself.
|
||
There are only two methods available for encryption, and really there’s only one that actually encrypts things. aes_gcm uses, well, AES-GCM, which is an industry standard encryption process that uses symmetric keys to encrypt and decrypt data.
|
||
The other method is called unencrypted and is used for migration scenarios.
|
||
Let’s put this all into practice with a simple example. If you want to follow along, check out this directory on my Terraform Tuesdays repository.
|
||
Encrypting Locally Example We’ll start with a basic example that encrypts data locally using the pbkdf2 key provider. Let’s take a look at the terraform.tf file. In here, I’ve got the encryption block nested inside the terraform block.
|
||
terraform { encryption { key_provider "pbkdf2" "passphrase" { passphrase = "tacos-are-delicious-and-nutritious" key_length = 32 iterations = 600000 salt_length = 32 hash_function = "sha512" } To define a key provider, I have a key_provider block type, which takes two labels. The first is the provider type, which I have set to pbkdf2 and the second is a name label called passphrase.
|
||
Inside the block I have some arguments that set the passphrase, key length, and some other options. Many of these arguments have default values which you can go with instead of adding them explicitly.
|
||
Once I have my key provider, I need a method block to define which encryption method I’m using and which key to use for it.
|
||
method "aes_gcm" "passphrase_gcm" { keys = key_provider.pbkdf2.passphrase } The method block has two labels. The method type, which in my case is aes_gcm, and a name label, for which I’m using passphrase_gcm.
|
||
Inside the block, the only argument is the keys argument, which I set to the key provider block identifier.
|
||
To use this method in my configuration, I can specify a state and/or plan block.
|
||
state { method = method.aes_gcm.passphrase_gcm } plan { method = method.aes_gcm.passphrase_gcm } Neither of these block types take labels. Inside the block, there is the argument method, which is set to the method identifier.
|
||
I have both state and plan set, so any state data or plan files I create should be encrypted using this method. Here’s my main.tf file.
|
||
resource "local_file" "main" { content = "Encrypt state and plan!" filename = "\${path.module}/testplan.txt" } output "test" { value = local_file.main.filename } I’m simply creating a local file and an output set to that file name. I haven’t defined a state backend in this configuration, so we’ll be using the local file backend.
|
||
We’ll start by running tofu init as you normally would.
|
||
$ tofu init Initializing the backend... Initializing provider plugins... - Reusing previous version of hashicorp/local from the dependency lock file - Installing hashicorp/local v2.5.1... - Installed hashicorp/local v2.5.1 (signed, key ID 0C0AF313E5FD9F80) Providers are signed by their developers. If you'd like to know more about provider signing, you can read about it here: https://opentofu.org/docs/cli/plugins/signing/ OpenTofu has been successfully initialized! Now I’m going to run a plan and save the output to a file.
|
||
$ tofu plan -out plan.tfplan ... Plan: 1 to add, 0 to change, 0 to destroy. Changes to Outputs: + test = "./testplan.txt" ──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── Saved the plan to: plan.tfplan Usually a plan file is in an unreadable binary form that you need to use tofu show to express as plaintext or JSON. What’s in the plan.tfplan file?
|
||
{ "meta":{ "key_provider.pbkdf2.passphrase":"eyJzYWx0IjoiR2drM3BtZlJpTWEycS9ET3BJNkJrZlgvSk5xZUV0eW9FMX..."}, "encrypted_data":"s4s...cYDi9Q==", "encryption_version":"v0" } I truncated the immensely long strings, but basically we have a JSON file that defines the key provider, encrypted data, and encryption version. I’m assuming newer versions of the encryption process will be released, so encryption_version helps OpenTofu know what version encrypted this payload. The encrypted payload also appears to be base64 encoded, making it safe for HTTP transmission.
|
||
Can I still view the execution plan? Sure! If I run tofu show, I will get back the unencrypted contents of the plan.
|
||
$ tofu show plan.tfplan OpenTofu used the selected providers to generate the following execution plan. Resource actions are indicated with the following symbols: + create OpenTofu will perform the following actions: # local_file.main will be created + resource "local_file" "main" { + content = "Encrypt state and plan!" + content_base64sha256 = (known after apply) + content_base64sha512 = (known after apply) + content_md5 = (known after apply) + content_sha1 = (known after apply) + content_sha256 = (known after apply) + content_sha512 = (known after apply) + directory_permission = "0777" + file_permission = "0777" + filename = "./testplan.txt" + id = (known after apply) } Plan: 1 to add, 0 to change, 0 to destroy. Changes to Outputs: + test = "./testplan.txt" The catch is that if I want to evaluate the plan information, I need to do so through OpenTofu. Third party tools that read an execution plan directly will need OpenTofu to do the decryption for them.
|
||
Let’s apply this plan by running tofu apply:
|
||
$ tofu apply plan.tfplan local_file.main: Creating... local_file.main: Creation complete after 0s [id=857d8bfe0af9c5a506e28f3a48e2dad06427e4d6] Apply complete! Resources: 1 added, 0 changed, 0 destroyed. Outputs: test = "./testplan.txt" Once everything is created, I now have a terraform.tfstate file in my current working directory. Looking at it’s contents:
|
||
{ "serial": 1, "lineage": "400979ee-83b6-8cb0-3a04-70f6f885ce14", "meta": { "key_provider.pbkdf2.passphrase": "eyJzYW...GVuZ3RoIjozMn0=" }, "encrypted_data": "zoRtfr0Sa...dB6Rw==", "encryption_version": "v0" } It’s very similar to the plan file. The serial number and lineage are still in plaintext, so OpenTofu can still do a comparison with a stored plan without decrypting the file. Beyond that is the encryption information and then the state data payload. I can still interact with the state by using tofu commands, but we cannot view the state data directly.
|
||
That’s a basic example, but it relies on setting a passphrase, and I don’t feel great about that in an automation setting. Let’s take a look at an example that uses AWS KMS and has existing state data I want to encrypt.
|
||
AWS KMS Example Alright, let’s say I already have a deployed configuration and I want to add encryption on top of it. How do I migrate to encrypted state?
|
||
To start with, we need an instance of AWS KMS and an S3 bucket to hold our state data. Here’s a configuration that will create those resources:
|
||
provider "aws" { region = "us-west-2" } resource "random_integer" "bucket_suffix" { min = 10000 max = 99999 } data "aws_caller_identity" "current" {} // Create a KMS key resource "aws_kms_key" "tofu_key" { description = "Tofu encryption key" enable_key_rotation = true deletion_window_in_days = 10 key_usage = "ENCRYPT_DECRYPT" customer_master_key_spec = "SYMMETRIC_DEFAULT" policy = jsonencode({ Version = "2012-10-17" Id = "key-default-1" Statement = [ { Sid = "Enable IAM User Permissions" Effect = "Allow" Principal = { AWS = "\${data.aws_caller_identity.current.arn}" }, Action = "kms:*" Resource = "*" } ] }) } module "terraform_state_backend" { source = "cloudposse/tfstate-backend/aws" version = "1.4.1" force_destroy = true bucket_enabled = true dynamodb_enabled = true name = "encrypted\${random_integer.bucket_suffix.result}" environment = "test" namespace = "tofu" } output "bucket_name" { value = module.terraform_state_backend.s3_bucket_id } output "dynamodb_table_name" { value = module.terraform_state_backend.dynamodb_table_name } output "kms_id" { value = aws_kms_key.tofu_key.id } You might notice that the key usage is set to ENCRYPT_DECRYPT and the key spec is SYMMETRIC_DEFAULT. Those are your only options at the moment. OpenTofu is using symmetric encryption to protect the state and plan files. That means the same key is used for both encryption and decryption versus having a public/private keypair.
|
||
As of right now, you cannot select an RSA or ECC key spec because those are asymmetric, meaning that the private key never leaves AWS KMS. Tofu would have to send the encrypted data up to AWS KMS to perform a decrypt operation and then get the unencrypted data back. That’s not how state data encryption works at the moment.
|
||
For this example, let’s assume I’ve already deployed a configuration using the S3 bucket as a backend. My terraform.tf file looks like this:
|
||
terraform { backend "s3" { region = "us-west-2" bucket = "tofu-test-encrypted70785" key = "terraform.tfstate" encrypt = "true" dynamodb_table = "tofu-test-encrypted70785-lock" } } And the main.tf is the same as the basic example:
|
||
resource "local_file" "main" { content = "Encrypt state and plan again!" filename = "\${path.module}/testplan2.txt" } output "test" { value = local_file.main.filename } Looking at the state data stored up in S3, it is unencrypted:
|
||
{ "version": 4, "terraform_version": "1.7.2", "serial": 1, "lineage": "9b6fee0e-b62d-1599-4a36-058391b98fe1", "outputs": { ... } Now I want to apply encryption to state, but maybe I don’t want to hardcode the KMS key ID into the terraform block. I can use the TF_ENCRYPTION environment variable instead. That’s useful for an automation scenario where you want the same pipeline code to support multiple environments.
|
||
Here’s a PowerShell script that uses heredoc syntax to populate the TF_ENCRYPTION environment variable.
|
||
$env:TF_ENCRYPTION = @" key_provider "aws_kms" "tofu" { kms_key_id = "62b4299c-0e4e-4323-98dc-e185d8dfe7b9" region = "us-west-2" key_spec = "AES_256" } method "aes_gcm" "tofu" { keys = key_provider.aws_kms.tofu } method "unencrypted" "tofu" {} state { method = method.aes_gcm.tofu fallback { method = method.unencrypted.tofu } } plan { method = method.aes_gcm.tofu fallback { method = method.unencrypted.tofu } } "@ The content of the heredoc string is basically HCL with my encryption settings. That includes the key provider of type aws_kms with the key id, region, and key_spec.
|
||
key_provider "aws_kms" "tofu" { kms_key_id = "62b4299c-0e4e-4323-98dc-e185d8dfe7b9" region = "us-west-2" key_spec = "AES_256" } The method is aes_gcm and the state and plan blocks are both using that method for encryption.
|
||
method "aes_gcm" "tofu" { keys = key_provider.aws_kms.tofu } But what’s this other method? Unencrypted?
|
||
method "unencrypted" "tofu" {} This method is specifically for when you want to set up or remove encryption. Unencrypted simply means the method doesn’t encrypt data.
|
||
Looking in the state block:
|
||
state { method = method.aes_gcm.tofu fallback { method = method.unencrypted.tofu } } There is a fallback block that uses the unencrypted method. When we run a tofu plan, it wants to use the aes_gcm method for handling state, but the existing state data doesn’t match that method. Normally OpenTofu would stop here and error out. It would look at the state data, not find a matching encryption method, and halt.
|
||
The fallback block gives it a second method to try if the first one doesn’t pan out. Currently our state data is not encrypted, so the appropriate fallback method type would be unencrypted. When it writes the updated state data back out during an apply, it will use the aes_gcm method, and our data will be encrypted.
|
||
Let’s try it! After running the PowerShell script to set up my TF_ENCRYPTION environment variable, I’ll run an apply.
|
||
$ tofu apply -auto-approve local_file.main: Refreshing state... [id=5d049985f49cb38a2ebf7c38f2b1b185cc65f7c7] ... Apply complete! Resources: 1 added, 0 changed, 1 destroyed. Outputs: test = "./testplan2.txt" Once the apply is complete, let’s take a look at the contents of our state data:
|
||
{ "serial":2,"lineage":"9b6fee0e-b62d-1599-4a36-058391b98fe1", "meta":{ "key_provider.aws_kms.tofu":"eyJjaX...MEE9PSJ9" }, "encrypted_data":"0pOk5...0Zojj0=", "encryption_version":"v0" } It appears that the state data is now encrypted, and based on the metadata, it’s using AWS KMS for the key provider. Neat!
|
||
If I wanted to go back to the unencrypted format, I would swap the method and fallback method. OpenTofu will read in the state data using the fallback method aes_gcm and write it back using the unencrypted primary method. You can do the same thing to change key providers or encryption methods as more become available.
|
||
Precaution and Warnings State data is like, really, really important. If you lose access to your state data, you’ve got a long day ahead of you. Sure import blocks and discovery tools make it a little easier, but it’s going to make for a bad time, especially if you’re using the same encryption key for multiple state data instances and you lose it. AWS KMS keys are cheap, so you should probably use a different key for each environment. There’s a tradeoff though. It creates another thing for your team to juggle and if someone gains access to those AWS KMS keys, they can really wreck your day!
|
||
Second, although you’re protecting your state data and plan files from an attacker who somehow gained access to where they’re stored, you aren’t protecting your state data and plan files from someone who needs to work with the configuration on a regular basis. That person needs to have access to state data, and they can easily make an unencrypted copy of state data or a plan file using tofu state pull or tofu show. It’s not storing your data in a secure enclave that no one can ever access or see inside.
|
||
I think you need to ask yourself, is adding this level of encryption is actually worthwhile? If you’ve secured your state backend properly, unauthorized people shouldn’t be able to get to your state data anyway. Whether you’re using an automation platform like GitLab or env0 or self-managing your state with S3 or Azure storage, you have the ability to lock down access and encrypt data at rest and in transit. Adding another layer of encryption might buy you a sliver of additional security, but is that worth the additional cost of managing keys and potentially losing access to state data? I can’t answer that for you, that’s something you’ll need to decide with your Security and Risk teams.
|
||
My personal feeling is that native encryption of state and plan files is not a huge boon to security, but it’s nice to have the option. And it’s a feather in the cap of OpenTofu that they can offer this feature for those who have been waiting forever to see it implemented. I’m curious to hear what you think. Will you be using this feature to encrypt your state or plan? Why or why not? Hit me up on LinkedIn or use the contact form.
|
||
`,summary:`Wondering how to encrypt your state data using OpenTofu? Then this is the post for you!
|
||
If you’d prefer your information in video format, check out the Terraform Tuesday video instead.
|
||
Introduction For this post, we are going to return once again to OpenTofu. If you’re not familiar with OpenTofu and how it differs from Terraform, then you may want to check out the blog post I’ve been maintaining that compare the differences between the two.`,date:"1 Aug, 2024",url:"https://nedinthecloud.com/2024/08/01/state-encryption-with-opentofu/",image:"opentofu-state-encryption.png",readingTime:"17"},"https://nedinthecloud.com/2024/07/24/migrating-to-opentofu/":{title:"Migrating to OpenTofu",tags:["hashicorp","terraform","opentofu"],content:`Are you thinking about migrating from Terraform to OpenTofu? What are the possible benefits? What are the drawbacks? And can you migrate back? I’ll try to answer those questions in this blog post.
|
||
If you’d prefer your information in video format, check out the Terraform Tuesday video instead.
|
||
Why OpenTofu? In August of 2023, HashiCorp decided to change the licensing for their core products from Mozilla Public License (MPL) 2.0 to Business Source License (BSL) 1.1. The change was introduced via a commit on each GitHub repository and took effect with the next subsequent release of each product. Terraform was included in the list of impacted products and some folks were nonplussed about the change.
|
||
I’m not here to re-litigate the whole thing. This is not a blog for a bunch of drama. If you want to hear people spilling the tea on HashiCorp and the state of open source in general, there’s plenty of folks happy to ramble at you. Just scoot on over to reddit or the bird site.
|
||
HashiCorp’s interpretation of BSL was simply that you could continue to use their community products for free if you weren’t creating a product that competed with their paid solutions. And the use had to be like for like. So you couldn’t use Vault community edition to create a competitor to HCP Vault or Vault Enterprise. And you couldn’t use Terraform community edition to create a competitor to Terraform Enterprise or HCP Terraform (previously known as Terraform Cloud). You could, however, use Terraform community edition to create a competitor to HCP Consul.
|
||
If you really want to dig into the details, you can check out their blog post and BSL FAQs. And if that isn’t adequate, there’s a licensing questions email address you can ping for your unique situation. For 99% of you, the license change doesn’t require any change from you. If licensing is your only concern, you’re probably in the clear (#notalawyer).
|
||
Terraform is by far HashiCorp’s most popular software, and many companies and projects have worked hard to fill in the gaps not covered by the community version of Terraform. Whether that was automation, state management, policy enforcement, or testing, there is a rich ecosystem of Terraform adjacent projects and products that suddenly felt under threat by the change to BSL.
|
||
At the same time, there were some open-source purists who rightly identified that BSL is not true open-source software (OSS), and now Terraform was under a source-available license. Organizations like the CNCF, that require their projects only use OSS tools and components found themselves in a jam.
|
||
A group of folks in the IaC ecosystem got together and wrote a manifesto demanding that HashiCorp revert back to an OSS license or they would fork Terraform. HashiCorp was disinclined to acquiesce to their request, and thus OpenTofu was born.
|
||
Like I said before, there is a non-zero amount of drama surrounding the birth of OpenTofu, HashiCorp’s response, and general internet sturm and drang. I am not going to get into it. I’ve said my piece elsewhere and I stand by it.
|
||
OpenTofu Now OpenTofu was forked from the latest version of Terraform that still carried the previous MPL license, equating to roughly 1.5.something. The current version of OpenTofu as of this post is 1.7.3, so everything I’m going to mention is about that version. Likewise, the current version of Terraform right now is 1.9.2, so when I make comparisons between the two that is the version I’m referring to.
|
||
The folks at OpenTofu have been trying to keep feature parity with newer releases of Terraform, but that is starting to diverge a bit. Any configuration that was valid before Terraform 1.6, should work just fine with OpenTofu. If you’re on Terraform 1.6 or newer, then you may have to do some small amount of rework depending on what features you’ve leveraged in the newer versions.
|
||
Comparing OpenTofu and Terraform I have a blog post I’ve tried to keep up to date as new minor versions of each project are released. If that’s tl;dr, here are the big highlights.
|
||
Both Terraform and OpenTofu have introduced a testing framework that makes use of run blocks, .tftest.hcl files, and a test command. Terraform includes mock data, which OpenTofu doesn’t yet. Other than that, the two are basically the same.
|
||
Both Terraform and OpenTofu introduced removed blocks, looping for import blocks, and provider defined functions. The removed block syntax is a little different, but otherwise I believe all of these features are identical.
|
||
OpenTofu released state and plan encryption with 1.7. This feature does not exist in Terraform, and I doubt it ever will. There’s an interesting WIP about ephemeral resources and values for Terraform that will allow you to omit certain values from state and plans, but I wouldn’t expect any implementation of that until 1.10 at the earliest.
|
||
Version 1.9 of Terraform introduce the ability to include other object values in an input variable validation block. As far as I can tell, this feature is not planned for OpenTofu at the moment.
|
||
Version 1.8 of OpenTofu will introduce support for .tofu files to enable a configuration to support both platforms while allowing you to take advantage of new features in either. The core idea is to have two files with the same name, but one ends in .tf and the other ends in .tofu. Terraform will ignore the .tofu file, and OpenTofu will ignore the .tf file. You can leverage new Terraform functionality in the .tf file and new OpenTofu stuff in the .tofu file. It’s a neat idea!
|
||
While there has been some divergence, OpenTofu and Terraform are largely the same from an operational standpoint. The code sitting behind it will be slightly different, but unless you’re a developer trying to contribute, that shouldn’t concern you too much.
|
||
That’s enough chit-chat. How about we check out a migration?
|
||
Migrating from Terraform to OpenTofu The steps to perform the migration are as follows:
|
||
Verify no pending changes Backup state data Remove incompatible features Change provider and module source addresses Run tofu init to switch providers and modules Run tofu plan to verify no changes Run tofu apply to update state data The Example Configuration I’m going to start with a simple project that deploys an Azure resource group and virtual network using Azure storage as the backend. Looking inside the configuration, we’ve got a terraform.tf file that defines the required providers and backend:
|
||
terraform { required_providers { azurerm = { source = "hashicorp/azurerm" version = "~>3.0" } } backend "azurerm" { resource_group_name = "tacoTruck" storage_account_name = "opentofu1313" container_name = "tfstate" key = "terraform.tfstate" } } And there’s a main.tf that defines the azurerm provider, resource group, and uses a module from the public registry to deploy a Virtual Network.
|
||
provider "azurerm" { features {} } resource "azurerm_resource_group" "test" { name = "opentofu-test" location = "East US" tags = { managed_by = "Terraform" } } module "vnet" { source = "Azure/vnet/azurerm" version = "4.1.0" resource_group_name = azurerm_resource_group.test.name use_for_each = true vnet_location = azurerm_resource_group.test.location address_space = ["10.0.0.0/16"] subnet_names = ["subnet1", "subnet2"] subnet_prefixes = ["10.0.0.0/24", "10.0.1.0/24"] } Let’s assume I’ve already deployed this configuration to Azure using Terraform. I can pull up the terminal and run terraform state list to see the resources under management.
|
||
$ terraform state list azurerm_resource_group.test module.vnet.azurerm_subnet.subnet_for_each["subnet1"] module.vnet.azurerm_subnet.subnet_for_each["subnet2"] module.vnet.azurerm_virtual_network.vnet The good folks at OpenTofu have published a migration guide that walks you through the process depending on which version of Terraform you’re currently using. I’m on version 1.9.2, so I’m following those directions.
|
||
Verify No Pending Changes The first step is to make sure there are no pending changes from the Terraform side. I’m going to run terraform plan to make sure.
|
||
$ terraform plan azurerm_resource_group.test: Refreshing state... ... No changes. Your infrastructure matches the configuration. Terraform has compared your real infrastructure against your configuration and found no differences, so no changes are needed. It comes back as no changes, so we’re good there. Had there been pending changes, we would want to apply those before attempting a migration. The goal is to switch to OpenTofu without impacting the target environment(s) in any way.
|
||
Backup Your State Data Next up, I am going to back up my state file just in case something goes sideways. If you’re using Azure storage, make sure you have versioning enabled and maybe even take a snapshot. You can also run terraform state pull and direct it to a file for a quick and dirty backup, but I wouldn’t recommend it for production workloads where there might be sensitive information in the state data.
|
||
We can use this copy of the state data for comparison as well once we migrate to OpenTofu.
|
||
Remove incompatible features You’re going to want to check your code for any features that OpenTofu doesn’t support. Here are the ones I know of at the moment:
|
||
Mock data in the testing framework Provider defined functions for the built-in Terraform provider Input variable validation referencing other objects As always, you should check the release notes from the most recent version of OpenTofu. Version 1.8 is going to add support for mock data.
|
||
Change Provider and Module Source Addresses OpenTofu had to set up their own registry for providers and modules because HashiCorp changed the licensing terms of their public registry to prevent non-Terraform binaries from accessing it. Again, I’m not here to point fingers or get catty. The practical upshot for you is that moving to OpenTofu means using their registry instead of the Terraform public registry.
|
||
The Terraform public registry is essentially an abstraction over the provider and module repositories on GitHub. It adds support for version constraints, supplies nicely formatted documentation, and allows for easy discovery. But the actual provider binaries and module files live on GitHub. The OpenTofu registry points at the exact same GitHub repositories without violating the Terraform registry license.
|
||
When you define a source for your provider or module, you can choose to use a shorthand like hashicorp/aws or the longer form registry.terraform.io/hashicorp/aws. When using the shorthand form, Terraform prepends registry.terraform.io automatically. OpenTofu does the same, but prepends registry.opentofu.org instead.
|
||
If any of the providers or modules in your configuration are using the longer form, you’ll need to switch to the short form. That includes modules and providers inside of child modules as well. You only need to do this for modules and providers from the Terraform public registry.
|
||
OpenTofu will still work if you pull providers from the Terraform public registry, but you might technically be violating the Terms and Conditions of the registry. And you’re not a rule breaker, are you?
|
||
Run tofu init to switch providers and modules With that out of the way, it’s time to initialize the configuration with OpenTofu. From the terminal, I’ll run tofu init:
|
||
$ tofu init Initializing the backend... Initializing modules... Downloading registry.opentofu.org/Azure/vnet/azurerm 4.1.0 for vnet... - vnet in .terraform\\modules\\vnet Initializing provider plugins... - Finding hashicorp/azurerm versions matching "~> 3.0, >= 3.11.0, < 4.0.0"... - Installing hashicorp/azurerm v3.113.0... - Installed hashicorp/azurerm v3.113.0 (signed, key ID 0C0AF313E5FD9F80) Providers are signed by their developers. If you'd like to know more about provider signing, you can read about it here: https://opentofu.org/docs/cli/plugins/signing/ OpenTofu has made some changes to the provider dependency selections recorded in the .terraform.lock.hcl file. Review those changes and commit them to your version control system if they represent changes you intended to make. OpenTofu has been successfully initialized! The output shows that the vnet module is being pulled from registry.opentofu.org/Azure/vnet/azurerm. Looking at the updated .terraform.lock.hcl file:
|
||
provider "registry.opentofu.org/hashicorp/azurerm" { version = "3.113.0" constraints = "~> 3.0, >= 3.11.0, < 4.0.0" hashes = [ "h1:beBFmLUcm/WPYeAX+E40sP3g6kn4Flpv2kO8DIwxj6c=", ...] } The azurerm provider is also pointing at the OpenTofu registry. Additionally, in the .terraform directory, there are now two folders for the azurerm provider:
|
||
$ tree .terraform/providers . ├───registry.opentofu.org │ └───hashicorp │ └───azurerm │ └───3.113.0 │ └───windows_amd64 └───registry.terraform.io └───hashicorp └───azurerm └───3.113.0 └───windows_amd64 You may want to go back after the migration and clean out the provider plugins from the Terraform public registry if you’re cramped for space.
|
||
Let’s check out the state data too, I’ll run tofu state pull and that will print the state to the terminal.
|
||
$ tofu state pull { "version": 4, "terraform_version": "1.7.3", "serial": 11, ... Looking at the terraform_version, it says 1.7.3. However the source for the azurerm providers still says registry.terraform.io:
|
||
"resources": [ { "mode": "managed", "type": "azurerm_resource_group", "name": "test", "provider": "provider[\\"registry.terraform.io/hashicorp/azurerm\\"]", ... Once we run a successful tofu apply, that will update as well.
|
||
Run tofu plan To Verify No Changes With the initialization is done, I’m going to run a tofu validate to make sure OpenTofu is cool with my config and then run tofu plan. If any issues occur it might come down to the slight differences in implementation of features.
|
||
$ tofu validate Success! The configuration is valid. My validate comes back clean, so I’m ready to run plan to see if any changes are listed. There shouldn’t be- we did check with Terraform before switching- but I want to make sure.
|
||
$ tofu plan azurerm_resource_group.test: Refreshing state... ... No changes. Your infrastructure matches the configuration. OpenTofu has compared your real infrastructure against your configuration and found no differences, so no changes are needed. My plan comes back with no changes, but if you do see changes, you may want to roll back to Terraform or try and troubleshoot what is different about the configuration. I’ll address rolling back later in the post.
|
||
Run tofu apply to update state data Although OpenTofu did change the Terraform version logged in state data, it left it otherwise unchanged. There are no pending changes for our actual infrastructure, so all tofu apply will do is alter the registry references in state data.
|
||
$ tofu apply -auto-approve tofu apply -auto-approve azurerm_resource_group.test: Refreshing state... ... No changes. Your infrastructure matches the configuration. OpenTofu has compared your real infrastructure against your configuration and found no differences, so no changes are needed. Apply complete! Resources: 0 added, 0 changed, 0 destroyed. As predicted, there are no changes to our resources, but there should be some changes in state. I’ll run tofu state pull to check:
|
||
{ "version": 4, "terraform_version": "1.7.3", "serial": 12, "lineage": "a084a4f3-7aee-5943-dc21-613e4d1a88c1", "outputs": {}, "resources": [ { "mode": "managed", "type": "azurerm_resource_group", "name": "test", "provider": "provider[\\"registry.opentofu.org/hashicorp/azurerm\\"]", The serial has incremented from 11 to 12 and the provider reference now points to registry.opentofu.org. Otherwise, the actual resources are exactly the same.
|
||
If I want to test out a change, I can update the managed_by tag on the resource group from Terraform to OpenTofu and run an apply.
|
||
resource "azurerm_resource_group" "test" { name = "opentofu-test" location = "East US" tags = { managed_by = "OpenTofu" } } This will prove that OpenTofu is now managing our infrastructure.
|
||
$ tofu apply -auto-approve azurerm_resource_group.test: Refreshing state... ... ... Apply complete! Resources: 0 added, 1 changed, 0 destroyed. After a few moments the apply is complete and now our infra is managed with OpenTofu.
|
||
Not too tough eh?
|
||
The YOLO Option The above migration process is cautious and methodical. For anything production-ish, I’d recommend that approach. But if you’re moving development stuff that you don’t really care about, the whole thing can be boiled down to two steps:
|
||
Run tofu init Run tofu apply -auto-approve And that’s it. The rest of the steps are about protecting the state and preventing issues. While I can’t recommend the YOLO approach, it really can be that simple to migrate.
|
||
Migrating Back to Terraform In case you were wondering, migrating back to Terraform is essentially the same process. Run terraform init, plan, and then apply and you’ll be back where you started. If you’ve started using new OpenTofu features, that might complicate things, especially the state data encryption. You’ll need to remove that before migrating back.
|
||
Your actual infrastructure in either case isn’t altered by the migration, and I think that’s the most important part. If you want to experiment with OpenTofu and then move back to Terraform, you can! No harm, no foul.
|
||
Should You Migrate? This is a tougher question, and it’s one that I struggle to answer because honestly, and I know this is trite, IT DEPENDS™. So rather than giving you a straight answer, since I don’t think there’s a one size fits all response, instead here are some questions to consider:
|
||
Are you running on HCP Terraform today? OpenTofu doesn’t work with HCP Terraform, so this is a non-starter. Are you running on a Terraform automation platform like Spacelift or env0? BSL versions of Terraform are not-available on these platforms, so if you want some of the new features introduced after the BSL change, you might want to move to OpenTofu. Are you working on a project that requires OSS tools? I guess you’re moving to OpenTofu! Are you building an HCP Terraform competitor for paid public consumption? Yeah, you’re going to need to move to OpenTofu. Are you using Terraform to build out your infrastructure and don’t really care about the whole OSS battle? I guess my follow-up question would be whether there is a killer feature of OpenTofu that you really want. Right now that’s just state data encryption. If that doesn’t move the needle for you, then I don’t see a reason to migrate. Has your legal team decided that the BSL is no good for your organization? Well I guess that’s out of your hands and above my pay-grade. To OpenTofu it is! You can also opt to stay on Terraform pre-BSL changes and just excuse yourself from the whole debate. You won’t be getting any security patches- that ended on December 31st 2023- but my point is that you can sit on 1.5 and wait to see what happens with the whole IBM acquiring HashiCorp thing, which is another can of worms I don’t want to get into right now.
|
||
Final Thoughts I hope in this post I’ve shown you how easy it is to move from Terraform to OpenTofu. While its not quite a drop-in replacement- you will need to do a little homework and analysis- it’s also a relatively low-risk and reversible operation. You could easily pilot OpenTofu on some development environments and see if it meets your needs.
|
||
That being said, my bottom line is that I would stick with the status quo unless there is some compelling reason to migrate. Most engineers and admins I know already have enough on their plate, and migrating to a new tool (even one as simple as Terraform to OpenTofu) is just one more thing to worry about. Unless there’s a clear business benefit or legal compulsion, I wouldn’t make the move.
|
||
That’s not to denigrate or cheapen what the OpenTofu folks have done. They’ve successfully forked an incredibly popular project and made the migration process as easy as possible. Kudos to everyone working on the project! As both projects evolve and inevitably diverge, I’d like to revisit the topic and see if my thinking has changed.
|
||
`,summary:`Are you thinking about migrating from Terraform to OpenTofu? What are the possible benefits? What are the drawbacks? And can you migrate back? I’ll try to answer those questions in this blog post.
|
||
If you’d prefer your information in video format, check out the Terraform Tuesday video instead.
|
||
Why OpenTofu? In August of 2023, HashiCorp decided to change the licensing for their core products from Mozilla Public License (MPL) 2.0 to Business Source License (BSL) 1.`,date:"24 Jul, 2024",url:"https://nedinthecloud.com/2024/07/24/migrating-to-opentofu/",image:"migrating-to-opentofu.png",readingTime:"15"},"https://nedinthecloud.com/2024/07/18/using-provider-defined-functions-in-terraform/":{title:"Using Provider Defined Functions in Terraform",tags:["hashicorp","terraform"],content:`Have you ever wished there were custom functions in Terraform? Well now there are! Sort of. Let me explain.
|
||
If you’d prefer your information in video format, check out the Terraform Tuesday video instead.
|
||
Introduction Whenever I am delivering a course on Terraform and we get to functions, someone inevitably asks if they can write their own custom functions. And that makes perfect sense. In a general purpose programming language, you can write your own methods and functions to help with reusability and consistency. It’s the DRY mentality.
|
||
My answer, up until now, was no, you cannot write your own functions. If you want a new function in Terraform, you could have to try and get it added to the Terraform binary or you could write a module that does something similar to what the function should do. But now there’s another option! With the release of Terraform 1.8, HashiCorp added provider defined functions.
|
||
Terraform Functions Up until now, all the functions in Terraform were built into the core binary. They are part of the compiled code and they execute super fast. If you wanted to write your own custom repeatable logic, you could do that with modules, but those are slower and don’t have access to the full power of the Go programming language.
|
||
But there’s another set of binaries that do use Go and are compiled, and those are provider plugins. So why not allow those plugins to define functions too? That’s exactly what provider defined functions are.
|
||
Provider Function Syntax Provider defined functions are written and compiled as part of a provider plugin. The syntax might look a little funky, but it makes sense. Here’s the generalized format:
|
||
provider::provider_name::function_name(arguments) Terraform needs to know that you’re invoking a provider-defined function, which plugin the function is coming from, and the function name. That is what the syntax is trying to communicate. Just like a built-in function, the arguments for the function go inside the parentheses at the end. If this syntax looks familiar, it’s what PowerShell uses to reference classes from the .Net framework. I’m sure the same syntax is used in other programming languages, but that’s the one that immediately came to my mind.
|
||
To demonstrate a more concrete example, the AWS provider has a function called arn_parse that breaks up an ARN into its constituent pieces. It’s something you could have done with a combination of other functions or a dedicated module, but this is way easier and possibly better than whatever you may have written in the past.
|
||
The function takes a single arn as an argument and returns a map with keys for each component like partition, service, and account_id. In fact, why don’t we see this and a few other provider functions in action.
|
||
Using Provider Functions Here’s a configuration I’ve put together that creates an AWS VPC.
|
||
provider "aws" { region = "us-west-2" } resource "aws_vpc" "my_vpc" { cidr_block = "10.0.0.0/16" } Once the configuration is applies, the VPC will have an ARN that I can retrieve and parse. Just like regular functions, I can try out the provider defined functions from the console.
|
||
$ terraform console > provider::aws::arn_parse(aws_vpc.my_vpc.arn) { "account_id" = "XXXXXXXXXXX" "partition" = "aws" "region" = "us-west-2" "resource" = "vpc/vpc-077f5e200f0ca2c62" "service" = "ec2" } And I get back the ARN broken up like I would want. I can take that same expression and use it to build a policy that references a partial arn.
|
||
HashiCorp has also added some functions into the built-in terraform provider, specifically: encode_tfvars, decode_tfvars, and encode_expr. You can read more about them in the documentation, but they’re for fairly uncommon situations and some of them can be replaced with the built-in templatestring function.
|
||
As a quick aside, I happened to notice that core functions that are part of the binary can now be referenced using their extended name, core::function_name. Not sure how long that’s been the case, but if I run core::max(1,3,5) at the console, it renders properly.
|
||
> core::max(1,3,5) 5 Not super important, but neat!
|
||
I want to mention here that if you plan to use provider defined functions, you’re going to want to set a lower bound for the provider plugins and Terraform version you’re using. Older versions of the provider won’t have the functions, and so Terraform will return an error. And earlier versions of Terraform won’t have any idea what the provider-defined function syntax is.
|
||
In particular, if you’re writing modules for others to consume, you want to make sure to specify the minimum provider versions and set the minimum Terraform version to 1.8 or newer.
|
||
Finding Provider Functions I started looking through the most popular providers on the registry and here’s some of the providers that have some functions today: aws, google, kubernetes, local, and time. Right now, there’s no easy way to discover functions outside of looking in a particular provider.
|
||
Which I guess is kind of the point. The functions are supposed to be something specific to the provider. The function direxists in the local provider checks to see if a directory exists. The function rfc3339_parse in the time provider breaks apart an rfc 3339 timestamp. That alone is hugely useful and might actually belong in Terraform proper.
|
||
While I appreciate the addition of provider defined functions, I’m a little concerned over discoverability for generic utility functions. I would expect to find the arn_parse function in the aws provider. I might not think to check the local provider for a function that checks for a directory’s existence. I’d probably look at the built-in functions, find that it isn’t there, and kludge something else together.
|
||
Writing Your Own On the topic of discoverability, I wouldn’t be surprised for providers to spring up that are purely for functions. In fact, check out the recently published assert provider. It’s got a ton of functions that are meant to make testing and validation easier in Terraform. If you want to know how, check out this awesome post from Bruno Schaatsbergen, the author of the provider.
|
||
I bet he’s looking for folks to chip in, so if you have suggestions, leave an issue or try your own hand at adding a function.
|
||
Conclusion Provider defined functions are a way for Terraform to support functions outside of what is baked into the core binary. This is an important step forward for Terraform and makes it easier for you to develop you own logic or leverage providers to bring additional functionality to Terraform.
|
||
`,summary:`Have you ever wished there were custom functions in Terraform? Well now there are! Sort of. Let me explain.
|
||
If you’d prefer your information in video format, check out the Terraform Tuesday video instead.
|
||
Introduction Whenever I am delivering a course on Terraform and we get to functions, someone inevitably asks if they can write their own custom functions. And that makes perfect sense. In a general purpose programming language, you can write your own methods and functions to help with reusability and consistency.`,date:"18 Jul, 2024",url:"https://nedinthecloud.com/2024/07/18/using-provider-defined-functions-in-terraform/",image:"provider-defined-functions.png",readingTime:"6"},"https://nedinthecloud.com/2024/07/08/variable-validation-improvements-in-terraform-1.9/":{title:"Variable Validation Improvements in Terraform 1.9",tags:["hashicorp","terraform"],content:`The release of Terraform 1.9 brings with it a welcome improvement regarding input variable validation. In this post I’ll review the change in functionality and provide a few examples for reference.
|
||
If you’d prefer your information in video format, check out the Terraform Tuesday video instead.
|
||
Variable Validation A core tenet of programming is to sanitize your inputs, and a big portion of sanitization is making sure the input structure and values match what you expect. In the world of Terraform, there are two controls you can leverage in the variable block to perform validation of input values.
|
||
The first is the type argument, allowing you to define the data structure you expect the input value to match. You can actually get pretty sophisticated with type by using the tuple and object structural data types to include required and optional keys and elements. But what about the contents of the value? That’s what the validation block is for.
|
||
Inside the variable block you can include one or more validation blocks. Those validation blocks include two arguments:
|
||
condition - with a value that is true or false error_message - the message to print if validation fails The condition argument can verify that the value(s) submitted match what you’re expecting to receive. You can compare the value to a list of allowed values, a regular expression, a range, or anything else that will result in a bool value.
|
||
Here are a few quick examples:
|
||
variable "region" { type = string description = "The region where the resources will be deployed" validation { condition = can(regex("^us-.*", var.region)) error_message = "Region must start with 'us-'" } } variable "instance_type" { type = string description = "The type of instance to be launched" validation { condition = contains(["t2.micro", "t3.micro"], var.instance_type) error_message = "Invalid instance type" } } variable "instance_count" { type = number description = "The quantity of instances to deploy" validation { condition = var.instance_count >= 1 && var.instance_count <= 20 error_message = "Instance count must be between 1 and 20" } } Validation blocks are pretty cool and I recommend using them.
|
||
Limitations The problem with variable validation blocks is that prior to Terraform 1.9, they could only reference the variable value being tested. You couldn’t reference other variables, local values, data sources, resources, etc. This made validation somewhat less dynamic than you might want. Consider the following:
|
||
Here’s that input variable example that checks if an instance_type is in an allowed list.
|
||
variable "instance_type" { type = string description = "The type of instance to be launched" validation { condition = contains(["t2.micro", "t3.micro"], var.instance_type) error_message = "Invalid instance type" } } The allowed list is stored with the variable block. So if I want to change that list, that’s a code change to the config. Wouldn’t it be nice if I could pull that list from somewhere?
|
||
Here’s an input variable that uses a subnet ID:
|
||
variable "subnet_id" { type = string description = "ID of subnet to use for application" } Wouldn’t it be nice to check and see if that subnet actually exists?
|
||
Or what about this configuration, where I want to use two different regions and make sure they aren’t the same value?
|
||
variable "primary_location" { type = string description = "Primary location for the resource group" } variable "secondary_location" { type = string description = "Partner location to use for deployment" } The good news is that with the release of Terraform 1.9, the validation block can refer to any other object in the same module. The only restriction is that Terraform needs to know the value during the plan if you want the validation block to fire. More on that later.
|
||
Let’s give it a try!
|
||
Location Example If you’d like to try this out yourself, all my examples can be found on the terraform tuesdays repository.
|
||
Let’s take the previous location example for a spin. The first variable is primary_location and this would be my first or primary location for a deployment. The second is called secondary_location and I probably don’t want it to be the same as my primary location. That would be silly!
|
||
To check that, I’ve simply added a validation block with the condition that the secondary location is not equal to the primary location. If it is, I’ll get back an error message.
|
||
variable "primary_location" { type = string description = "Primary location for the resource group" } variable "secondary_location" { type = string description = "Partner location to use for deployment" validation { condition = var.secondary_location != var.primary_location error_message = "Secondary location must be different from the primary location" } } To test it, I have a terraform.tfvars file that has eastus for the location and westus for the partner location.
|
||
primary_location = "eastus" secondary_location = "westus" I’ll run terraform plan, and once it finishes its process, everything comes back green.
|
||
Terraform used the selected providers to generate the following execution plan. Resource actions are indicated with the following symbols: + create Terraform will perform the following actions: # azurerm_resource_group.main will be created + resource "azurerm_resource_group" "main" { + id = (known after apply) + location = "eastus" + name = "main-resources" } # azurerm_resource_group.partner will be created + resource "azurerm_resource_group" "partner" { + id = (known after apply) + location = "westus" + name = "partner-resources" } Plan: 2 to add, 0 to change, 0 to destroy. Now I’ll change the secondary location to be eastus and run terraform plan a second time. This time the validation fails and I get a helpful error back.
|
||
│ Error: Invalid value for variable │ │ on terraform.tfvars line 2: │ 2: secondary_location = "eastus" │ ├──────────────── │ │ var.primary_location is "eastus" │ │ var.secondary_location is "eastus" │ │ Secondary location must be different from the primary location │ │ This was checked by the validation rule at main.tf:24,3-13. Neat!
|
||
What about using a data source?
|
||
Network Example For this example, let’s assume I’ve got an existing Virtual Network in Azure with three subnets in it: web, app, and db. In my deployment I want to use one of the subnets, and I want to make sure that the subnet actually exists.
|
||
I can use a data source to get the list of subnets for an existing VNet like so:
|
||
data "azurerm_virtual_network" "main" { name = var.vnet_name resource_group_name = var.resource_group_name } For the subnet_name input variable, I’ve added a validation block that uses the contains function to check and see if the subnet_name value is in the list of subnets from the data source.
|
||
variable "subnet_name" { type = string description = "Name of the subnet" validation { condition = contains(data.azurerm_virtual_network.main.subnets, var.subnet_name) error_message = "Subnet name must be in the list of subnets from the virtual network." } } In the terraform.tfvars file I have the correct Virtual Network and resource group in it.
|
||
vnet_name = "nettest-vnet" resource_group_name = "nettest-resource-group" At the command line, I’ll run terraform plan -var subnet_name=“web” and after a few moments the plan comes back successful.
|
||
data.azurerm_virtual_network.main: Reading... data.azurerm_virtual_network.main: Read complete after 0s [id=/subscriptions/4d8e572a-3214-40e9-a26f-8f71ecd24e0d/resourceGroups/nettest-resource-group/providers/Microsoft.Network/virtualNetworks/nettest-vnet] No changes. Your infrastructure matches the configuration. Terraform has compared your real infrastructure against your configuration and found no differences, so no changes are needed. The data source is queried before the validation for the input variables is run, so Terraform can reference the contents of the subnet attribute for the VNet data source. In the subnet attribute it found the web subnet, thus all is well.
|
||
Now let’s try a subnet name that’s not in the virtual network, like tacos for instance.
|
||
$ terraform plan -var subnet_name="tacos" ... Planning failed. Terraform encountered an error while generating this plan. ╷ │ Error: Invalid value for variable │ │ on main.tf line 31: │ 31: variable "subnet_name" { │ ├──────────────── │ │ data.azurerm_virtual_network.main.subnets is list of string with 3 elements │ │ var.subnet_name is "tacos" │ │ Subnet name must be in the list of subnets from the virtual network. │ │ This was checked by the validation rule at main.tf:35,3-13. This time it comes back with a failure and the error message. Sadly, there is no tacos subnet. 😢
|
||
Pretty useful stuff!
|
||
Unknown Values The validation block improvement in Terraform 1.9 can reference any object in the same module. But what happens when the value for an object is not known during plan? For example, let’s say I reference the attribute of a resource that isn’t known until the resource is created:
|
||
variable "partner_resource_group_id" { type = string description = "Partner resource group to use for deployment" validation { condition = azurerm_resource_group.main.id != var.partner_resource_group_id error_message = "Secondary resource group must be different from the primary resource group" } } The attribute azurerm_resource_group.main.id won’t be known until the resource group is created. So if I run terraform plan before any resources are created, what will Terraform do?
|
||
I had assumed that Terraform would do one of two things:
|
||
Issue a warning that the validation block couldn’t be processed and produce an execution plan. Throw an error that all referenced values must be known during plan. It turns out that Terraform does neither of these things. Instead, it simply ignores the validation block and produces an execution plan. No error, no warning, just produces the plan.
|
||
If you apply the plan and the resource is created, Terraform will evaluate the validation block on the next plan run as expected.
|
||
I think Terraform should at least issue a warning so you know the validation block was skipped, or the docs should explicitly acknowledge this behavior. Guess I’ve got an issue to log. Shout out to Mattias Fjellström for pointing out this odd behavior.
|
||
Final Thoughts While it was possible to do this type of checking by hacking something together with pre or post condition blocks for resources and data sources, I think this catches things earlier on in the evaluation cycle. You can also define acceptable values using locals or reference another input variable in the configuration.Overall this is an excellent addition to the existing variable validation block and I’m excited to see it!
|
||
If you’d like to try this feature out for yourself, the example code is in my Terraform Tuesday repository. Thanks for reading!
|
||
`,summary:`The release of Terraform 1.9 brings with it a welcome improvement regarding input variable validation. In this post I’ll review the change in functionality and provide a few examples for reference.
|
||
If you’d prefer your information in video format, check out the Terraform Tuesday video instead.
|
||
Variable Validation A core tenet of programming is to sanitize your inputs, and a big portion of sanitization is making sure the input structure and values match what you expect.`,date:"8 Jul, 2024",url:"https://nedinthecloud.com/2024/07/08/variable-validation-improvements-in-terraform-1.9/",image:"variable-validation.png",readingTime:"8"},"https://nedinthecloud.com/2024/07/06/pluralsight-problems/":{title:"Pluralsight Problems",tags:["pluralsight"],content:`I came across this Yahoo! Finance article via a LinkedIn post from Don Jones, and much like him I had a visceral reaction while reading it. The pit of my stomach dropped, my muscles tensed, and I felt vaguely nauseous. The tl;dr? Pluralsight was saddled with unsustainable debt when they were acquired by Vista Equity Partners and now the piper has come a calling. All usual, private equity firms continue to destroy everything they touch and should be heavily regulated if not made outright illegal.
|
||
(There’s some adult language in this post. If that’s not your thing, then I guess don’t read it?)
|
||
A Little Pluralsight History You’re probably aware of Pluralsight even if you haven’t used the platform. They are an EdTech company founded back in 2004 to focus on in-person training. Over time they pivoted to online video instruction and quickly grew due to the quality of their materials, excellent instructors, and first-mover advantage.
|
||
Over the next 15 years, Pluralsight continued to grow, building a solid reputation for excellence and attracting talented authors. Investors took notice and started pouring money into the company with funding rounds in 2012, 2014, and 2016 totaling about $200 million.
|
||
The natural next step in the process was to make an initial public offering to pay back the VC investors and drum up additional capital to keep the company growing. So in 2018, Pluralsight had their IPO on the NASDAQ, starting at $15 a share and closing the first day at $20, a solid 33% pop.
|
||
Unfortunately, the performance of the company didn’t follow the market’s expectations and the stock price suffered. Once you’ve take your company public, you are now at the whims of institutional and activist investors. They are generally not thinking about long-term viability. They need to see growth at all costs, and Pluralsight was no longer showing the same dramatic growth seen in pre-IPO.
|
||
In July of 2019, the stock price dropped from a respectable $31 a share, to a low of $15. That perked up the ears of the ravenous private equity firms and soon the wheels were in motion. The stock continued to crater to about $8 in mid-March of 2020, exacerbated by the pandemic and attendant economic turmoil. The share price did recover, as we saw the COVID related bounce of all tech stocks, but that was a temporary effect and in December of 2020, Pluralsight announced that it would be going private again through Vista Equity Partners to the tune of $3.5 Billion.
|
||
Private Equity Context The thing about private equity is that they really don’t like using their own money to buy companies. What they do is borrow a bunch of money using the equity of the company they are purchasing as collateral. Then they take that loan obligation and associate it with the company they just bought. It’s called a leveraged buyout, and it fucking sucks ass.
|
||
The acquiring company- Vista Equity Partners- is making a bet that the company they acquire- Pluralsight- can be restructured in such a way that it can pay off the debt it’s been saddled with, but it’s a pretty low-risk bet for Vista. Best case, Pluralsight takes off, services its debt, and makes the Vista partners a fuck-ton of money. Worst case, Pluralsight fails to service the debt, Vista restructures the company with a massive write down, the debtors take a haircut and Vista partners make a fuck-ton of money.
|
||
See, the thing is that the private equity firm pays itself all kinds of bonuses and compensation for the acquisition itself and each subsequent financial machination. Need to restructure the debt? That’s gonna require investment bankers who charge a shitload of money to do the work. Made your quarterly revenue numbers? That’s a bonus for the private equity partners! Decided to carve up the company and sell it off to a bunch of other firms? Lawyers, bankers, and investors all get their due. And the acquired company? They get fucked.
|
||
Recent Developments With that little context corner out of the way, what’s happening with Pluralsight? Well, things haven’t been going great. There’s been multiple rounds of layoffs since they went private again. Competition in the EdTech space has continued to mount, and with Pluralsight now living with the shackles of a massive debt load, they haven’t been able to focus on executing a coherent long-term growth strategy. As a direct result, they haven’t kept pace with other EdTech companies in terms of growth, and the revenue coming into Pluralsight is insufficient to service the debt.
|
||
In Q1 2023, Vista Equity first started to ask lenders to loosen up their loan covenant, otherwise Pluralsight would be able to honor their commitments. Since then, Vista Equity has been writing down the value of Pluralsight on their books from an initial $3.5B in 2021 to now effectively $0 in May of 2024.
|
||
In April of 2024, CEO Aaron Skonnard stepped down and Chris Walters took over. Aaron was one of the founders of Pluralsight, and his departure signals the last of the old guard vestiges leaving the building.
|
||
Vista Equity also took the odd step of moving Pluralsight’s intellectual property (IP) to a new subsidiary and then borrowing against it to pay loan obligations. This is a naked admission that Pluralsight retains massive value in terms of intellectual property, and it is the financial manipulations of Vista that have presaged the current collapse of Pluralsight as a company.
|
||
Which brings us to the Yahoo! Finance article detailing how Vista is in talks to cede control of Pluralsight to lenders. At best, the lenders will write down some of the debt and give the new CEO time to try and get Pluralsight’s house back in order. At worst, they will layoff 95% of the company’s employees, and shop around the IP bundle to other private equity firms to try and recoup some portion of their investment.
|
||
To quote a great philosopher, “That’s fucking bullshit, man.”
|
||
My History With Pluralsight If I seem a little emotional, it’s because I am. That’s because I’ve been working with Pluralsight to produce courses since 2016. At this point, I’ve worked for Pluralsight longer than any other employer since I got my first job at 14. Technically I don’t work for them, I work with them, since I am a contract author and not an FTE. Technicalities aside, I’ve been working with the people at Pluralsight for eight years and developed deep and abiding relationships with many of them.
|
||
At the same time, Pluralsight has also been my main source of income since the inception of Ned in the Cloud LLC in 2019. I started with a single course in 2017 and over time have built up a stable of high-performing courses that earn me residuals every quarter. The number varies by year, but about 70% of my revenue comes from Pluralsight, and that’s making me very nervous at the moment.
|
||
I’ve always recognized that having such a large portion of my revenue come from a single source was a major risk. I also thought that the chance of Pluralsight going out of business was fairly low, and that I would have adequate time to diversify before the axe fell.
|
||
I’m not sure if I’m ready to say that the axe is falling, but the headsman is definitely sharpening the blade and doing some stretches. I would be a fool to bank on Pluralsight’s continued existence for the bulk of my income.
|
||
How I’m Feeling Now (Apologies to Charlie XCX) Frankly? I’m fucking pissed off. And I want to be clear. I am not angry at the rank and file folks at Pluralsight. Every single person I know who works there or has worked there cares deeply about the learner experience. They want to assist in the creation of high-quality learning materials that help folks progress in their career. The people at Pluralsight are top-notch. If you have a chance to hire one after the next round of layoffs, I highly recommend you do so.
|
||
No, I’m pissed off at the Pluralsight leadership for steering the company into such a quagmire. I’m angry at the private equity firms who couldn’t give one flying fuck about the companies they strip mine to line their pockets. I’m livid that the financial regulators are so completely lax that Pluralsight’s fate is not even an uncommon occurrence. The employees, contractors, and customers of Pluralsight were all fucked over by an industry obsessed with growth at all costs and short-term gains. And just thinking about it makes me fucking furious.
|
||
It didn’t have to be this way. Pluralsight was started by a group of idealistic people trying to educate others about cool tech. It was self-funded for the first eight years of its existence, and I suspect they could have continued in that vein, taking out traditional loans for growth instead of turning to the VC market. Once you start down that path, I’m not saying your fate is sealed, but your control over the company and its future rapidly diminishes with each funding round you accept.
|
||
Once Pluralsight IPO’d, they essentially let go of the reins. The die was cast and the rest was simply going through the motions. I’m not sure if I even believe that entirely, but it really feels that way.
|
||
Pluralsight Is Still Good The irony? If you can divide the weird corporate finance bullshit from what Pluralsight actually does as a company, you’d find a legitimate business that provides a valuable service to a large body of customers.
|
||
I can say, without a doubt, that the content you find on Pluralsight will generally be of higher quality than any other platform. The standards, consistency, and care put into each course shows, and I’ve heard that echoed by literally hundreds of people.
|
||
Like I said, I’ve been making courses on Pluralsight for over eight years now. My courses alone have been viewed by 391k people for a combined 580k hours! If there’s one consistent piece of feedback I receive it’s the level of quality and rigor present in the course materials. And I know I’m not the only one!
|
||
All of that high quality content? It’s still there! Pluralsight’s strength has long been the panoply of dedicated contract authors from a range of backgrounds and technologies. The authors are held to a high standard and well compensated for their efforts. As long as Pluralsight maintains a healthy relationship with their authors, the platform will continue to host some of the best training material in the world.
|
||
As an author, that relationship isn’t feeling to great right now. My monthly viewership has been dipping for months now. Most of the people I felt close to have left the company, and no one has stepped in to fill that void. About 18 months ago, they cut author compensation by 25% across the board, and I’d be lying if I said I wasn’t concerned that more cuts are coming.
|
||
Without the authors, Pluralsight is doomed.
|
||
My Advice Jumping Jesus Christ on a pogo stick, I have no goddamn idea what Pluralsight should do. I’m not a financial guru or a corporate leader. I’ve never planned the long-term strategy for a multi-million dollar company. If it could be fixed with PowerShell and Terraform, I’d be happy to help. But this? Fixing a company that has been financially ruined by ambivalent owners who turn a profit no matter what happens? I got nothing.
|
||
That being said, here’s what I like to see happen.
|
||
I would love for an investor to come along and see that Pluralsight’s best days are in front of it, and scoop it up for pennies on the dollar. Then spend time formulating a plan to fix the relationship with existing authors, while actively recruiting new ones to the platform. I’d like to see Pluralsight embrace new learning modalities and even have mixed media courses to meet learners where they are and in the format that matches their current need.
|
||
Doing so will require significant upfront capital and a willingness to invest in people, rather than laying them off to try and save up enough money to make the next loan payment. I think on a ten year time horizon, buying Pluralsight could be an excellent investment, but the next few years are going to be rough.
|
||
Unfortunately, all the money sloshing around tech at the moment is focused on GenAI. I have it on good authority that Pluralsight’s best performing courses are all GenAI related for the last quarter. No surprise there. I don’t know if there’s another way to capitalize on the GenAI madness, but then again I’m not sure the leadership at Pluralsight has the leeway to do anything besides tread water and ward off the circling sharks with a broken oar.
|
||
Final Thoughts It’s been an amazing eight years with Pluralsight. I’m not going to stop making courses for them entirely, and I will keep my current crop of courses up-to-date. But I also have to recognize the distinct possibility that my efforts will see diminishing- if not disappearing- returns. For the next few years, I need to focus on differentiating my revenue sources and minimizing the impact of the possible total collapse of Pluralsight.
|
||
I sincerely hope that Pluralsight can find a way through this financial morass foisted upon it by callous assholes. I hope that those folks impacted by recent layoffs and probable future ones will end up at awesome companies in well paying positions. Y’all deserve it.
|
||
As for Vista Equity Partners? Go fuck yourselves.
|
||
`,summary:"I came across this Yahoo! Finance article via a LinkedIn post from Don Jones, and much like him I had a visceral reaction while reading it. The pit of my stomach dropped, my muscles tensed, and I felt vaguely nauseous. The tl;dr? Pluralsight was saddled with unsustainable debt when they were acquired by Vista Equity Partners and now the piper has come a calling. All usual, private equity firms continue to destroy everything they touch and should be heavily regulated if not made outright illegal.",date:"6 Jul, 2024",url:"https://nedinthecloud.com/2024/07/06/pluralsight-problems/",image:"pluralsight-problems.png",readingTime:"11"},"https://nedinthecloud.com/2024/04/15/using-azure-container-instances-for-an-azure-devops-self-hosted-agent/":{title:"Using Azure Container Instances for an Azure DevOps Self-hosted Agent",tags:["hashicorp","terraform","azure"],content:`When you create a new organization in Azure DevOps, you’ll quickly discover you’re not able to run any pipeline jobs using Microsoft-hosted agents. In fact, if you kick off a pipeline, you’ll get the error: No hosted parallelism has been purchased of granted. This post will detail how you can host your own pipeline agents for almost free using Azure Container Instances.
|
||
Background In 2021, the good folks over at Microsoft were fed up with people abusing the free Microsoft-hosted agents that come with an Azure DevOps organization. Because you can make the agent machine run pretty much any arbitrary script, shitty people were creating organizations en masse and using the free compute to do all sorts of nefarious and unsavory things (read: cryptocurrency and spambots).
|
||
Microsoft had tried to combat this by restricting how many pipeline jobs were available for use. First they removed the 10 parallel jobs for public projects. Then they restricted private projects to a single parallel job. Nevertheless, the shitty people persisted.
|
||
So Microsoft said ENOUGH! And rescinded the previous grant of one parallel job per organization on private projects for new Free-tier organization. Now when you sign up you get ZERO parallel jobs.
|
||
That’s intensely annoying, and it stymies your ability to try out various functions of Azure DevOps, as well as do some of the exercises on Microsoft’s Learn platform! While I don’t necessarily agree with Microsoft’s draconian approach, I get it.
|
||
So does that mean the shitty people have won and you are SOL? Not necessarily.
|
||
You have one of three options to remedy this situation:
|
||
You can fill out this form and wait 1-3 days for them to grant you a single parallel job for free. 🥱 You can pay for an Azure DevOps subscription that includes parallel jobs. 🤑 You can use a self-hosted agent. 😊 This post is meant to assist you with the third option. The linked repository will spin up an Azure Container Instance (ACI) that has the Azure DevOps agent installed. Through the use of environment variables, you can configure the agent to be part of one of the pools provisioned in your Azure DevOps organization.
|
||
To make things even easier, I’ve added an option to provision the agent pool in AZDO for you as well.
|
||
Prerequisites If you want to follow along, you’ll need the following:
|
||
An Azure subscription An Azure DevOps organization (free tier) The Azure CLI installed on your machine A Personal Access Token (PAT) for your Azure DevOps organization The token should have Full Access The repository folder copied locally By default, the configuration uses my ned1313/azp-agent:1.2.0 image on Docker hub. The image is based on the ubuntu:22.04 image and has a startup script that installs the latest version of the Azure DevOps agent your organization supports.
|
||
It also includes some common tools like curl, jq, and unzip. Here’s a link to the image build repository in case you’d like to see what’s included, or fork it for yourself.
|
||
On startup, the image will grab the latest version of the Azure DevOps agent for your organization and configure it to run as a service. The agent will be added to the pool you specify in the environment variables.
|
||
Let’s walk through a sample deployment shall we?
|
||
Deployment Process The deployment process is simple, but you need to set up a few things first.
|
||
Start by logging into Azure using the Azure CLI and selecting the subscription you want to use:
|
||
az login az account set -s SUBSCRIPTION_NAME Next, create a personal access token (PAT) in Azure DevOps with Full Access and store that token value in an environment variable:
|
||
export TF_VAR_azp_token=TOKEN_VALUE Now fill out the terraform.tfvars file with your desired location and AZDO organization name.
|
||
location = "eastus" azp_org_name = "tacocat" If you want to use an existing agent pool you’ve already provisioned, add that as well:
|
||
azp_pool = "tacopool" If you don’t give a value for the azp_pool input variable, the configuration will create a pool named aci-agents instead.
|
||
Now just do the standard Terraform deployment dance (it’s like the Neutron Dance, but way cooler):
|
||
terraform init terraform apply Once the deployment is complete, you can look at your agent pool and see that it is deployed.
|
||
The configuration is set up to make the pool visible to all projects by default, but you’ll need to go into the desired project and grant pipelines permission to use the pool.
|
||
In the YAML for the pipeline, you can configure a job to use the pool with the following syntax:
|
||
stages: - stage: Apply displayName: Apply jobs: - job: apply pool: aci-agents # or whatever your pool name is... Like I said, the image has all the tooling I would want for a basic Terraform deployment pipeline, but it doesn’t come with build tools for every language. If you need to add more tools, take my Dockerfile and build your own. You might find the public repo for the Microsoft-hosted agents helpful for installation scripts.
|
||
Cost ACI is not free, but it is relatively inexpensive. If you want totally free, spin the image up on your local workstation instead. There’s a start-agent.sh script included in the image repository.
|
||
The cost of an ACI group varies by region, so I recommend consulting the Azure pricing page for the most up-to-date information. My configuration uses 1 vCPU and 2 GB of memory, which if you run it for an hour in US East would come out to about $0.05 or $1.20 per day. You might be able to save a little money by using less vCPU or memory, but I haven’t tested that.
|
||
You can also destroy the ACI instance when you’re not using it and leave the pool in place. The pool doesn’t cost anything to maintain.
|
||
Final Thoughts It sucks that shitty people ruin everything™️. Personally, I think Microsoft has over-corrected on this one. When I went to go create a new Azure DevOps organization, I had to jump through two separate captcha hurdles and a bunch of other nonsense. It seems like they’ve put sufficient controls in place to start granting that single parallel job to Free tier accounts again.
|
||
The form and 1-3 day wait isn’t forever, so I’d recommend filling out the form and using my solution as a stop-gap measure. It’s also a chance to learn about self-hosted agents and options for deploying them. That’s nice and sometimes we can have nice things.
|
||
Share and enjoy!
|
||
`,summary:`When you create a new organization in Azure DevOps, you’ll quickly discover you’re not able to run any pipeline jobs using Microsoft-hosted agents. In fact, if you kick off a pipeline, you’ll get the error: No hosted parallelism has been purchased of granted. This post will detail how you can host your own pipeline agents for almost free using Azure Container Instances.
|
||
Background In 2021, the good folks over at Microsoft were fed up with people abusing the free Microsoft-hosted agents that come with an Azure DevOps organization.`,date:"15 Apr, 2024",url:"https://nedinthecloud.com/2024/04/15/using-azure-container-instances-for-an-azure-devops-self-hosted-agent/",image:"ado-self-hosted-agent.png",readingTime:"6"},"https://nedinthecloud.com/2024/03/04/creating-your-first-terraform-repository/":{title:"Creating Your First Terraform Repository",tags:["hashicorp","terraform"],content:`Terraform is an infrastructure as code tool. The key term here being code. And where do you store your code? In a version control system like GitHub! I recently published a video with the awesome April Edwards on creating your first Terraform repository.
|
||
This blog post is a companion to that video. Let’s dive in!
|
||
Why Use Version Control? If you’re seasoned developer using version control, then I don’t need to convince you of its benefits. But if you’re like me and coming from a sysadmin background, then you might not be as familiar with version control software or its benefits. Here are some reasons why you should use version control:
|
||
Backup - Your code is stored in a remote location Versioning - You can see the history of your code and roll back to previous versions if needed Collaboration - Multiple people can work on the same codebase Code Review - Others can review and approve changes before they’re merged and applied Automation - You can automate testing and deployment of your code on check in With all that in mind, how do you get started with version control for your Terraform code? First we need to go over some terminology and tools.
|
||
Terminology The most popular source control software in the world is git. It’s leveraged by GitHub, GitLab, Bitbucket, and many other platforms. There are other source control systems like Subversion, but I don’t know anything about those, so we’ll focus on using git.
|
||
Git tracks your code through a repository. Your code is checked into a local repository managed by git. The repository tracks the various versions of code through commit IDs and branches. To add a file to git, you first stage the file, then commit it to the repository. The commit process creates a new commit ID and stores the file in the repository.
|
||
There’s a LOT more to it than just that. But for our purposes, the important thing is to understand the process of staging files and committing them to the repository. Just saving a file to the directory being used by git doesn’t mean it’s being tracked by git.
|
||
In addition to your local repository, you can also configure one or more remote repositories. This allows you to push your local code to a remote location, collaborate on the code with others, and pull down updates they have made.
|
||
While you’ll certainly want to commit and track your Terraform files, like main.tf or variables.tf, there are some files that you don’t want to commit. For instance your local state file (if you’re using local state) or the .terraform directory that holds the provider plugins and modules used by your code. You can tell git to ignore these files by creating a .gitignore file in your repository.
|
||
The .gitignore file is a special file that tells git what files to exclude from commits. You can specify files by name, or use wildcards to exclude entire directories or file extensions.
|
||
Now that we have some background in version control and git, it’s time to create a repository!
|
||
Creating a Local Repository If you’ve been using Terraform for a while now, there’s a good chance you have a bunch of directories with your code in them. Maybe you’ve been sharing the code by zipping it up and emailing it, or copying it to a shared drive. Now it’s time to try a better way by storing that code in a git repository. Let’s see how you can start by creating a local repository for your existing code.
|
||
In my example, we have a directory called taco-app that has the following file tree:
|
||
taco-app$ tree -a -L 1 . ├── .terraform ├── .terraform.lock.hcl ├── main.tf ├── outputs.tf ├── terraform.tf ├── terraform.tfstate ├── terraform.tfvars └── variables.tf This is the code that we want to store in a repository. First we’ll run a git command to create the repository in the taco-app directory. In the same way that you initialize a Terraform configuration with terraform init, you start your git journey by running git init.
|
||
taco-app$ git init -b main Initialized empty Git repository in .../taco-app/.git/ A new repository needs a default branch for git to track. We can use the -b flag to set the default branch name for our repository. In this case, we used main as the branch name.
|
||
The command created a hidden folder called .git, which will store information about the files being tracked, commit ids, and configuration of our git client for this repository.
|
||
taco-app$ tree -L 1 .git .git ├── HEAD ├── branches ├── config ├── description ├── hooks ├── info ├── objects └── refs Now we can add our files to the repository. Before we do that, I’m going to set up a .gitignore file to exclude the .terraform directory and the terraform.tfstate file. We don’t want to commit these files to the repository. You can find a good .gitignore file for Terraform on GitHub.
|
||
The command to stage our files is git add followed by the files we want to add. Rather than laboriously typing out all the files and directories, instead we can use the . to add all files and subfolders of the current working directory.
|
||
taco-app$ git add . You won’t see any output from this command if everything has gone well. If we want to see what files have been staged, we can run git status:
|
||
taco-app$ git status On branch main No commits yet Changes to be committed: (use "git rm --cached <file>..." to unstage) new file: .gitignore new file: .terraform.lock.hcl new file: main.tf new file: outputs.tf new file: terraform.tf new file: variables.tf With our files staged, now we can commit them to source control using the git commit command. This command requires that you provide a message about the commit. The message can be anything you like, but it should be descriptive of the changes being introduced by your commit.
|
||
taco-app$ git commit -m "First commit" [main (root-commit) 827b4dd] First commit 6 files changed, 58 insertions(+) create mode 100644 .gitignore create mode 100644 .terraform.lock.hcl create mode 100644 main.tf create mode 100644 outputs.tf create mode 100644 terraform.tf create mode 100644 variables.tf Our files are now part of the repository! But right now they’re only stored on our local machine. If we want to collaborate with others, or have a backup of our code, we need to push it to a remote repository.
|
||
Creating a Remote Repository on GitHub To create a remote repository on GitHub, you’ll need to have an account. If you don’t have one, go to GitHub and sign up.
|
||
You can create a new repository using the web interface or the GitHub command line tool. Since we’re doing everything else with the command line, we might as well use it to create our remote repository too. If you don’t have the GitHub CLI tool installed, you can download it from the GitHub CLI website.
|
||
If you’ve never used the tool before, you’ll need to login into GitHub using the command gh auth login. The wizard will walk you through authenticating with GitHub and procuring a token for the CLI to use.
|
||
We can create the repository using the gh repo create command. Since we already have the local repository, we can let GitHub know that it should use our local repository by specifying the --source flag. We also want to make the repository public, so we’ll use the --public flag. Finally, we want to add a remote entry to our local repository, so we’ll use the --remote flag to specify the name of the remote repository entry.
|
||
taco-app$ gh repo create --public --source=. --remote=upstream ✓ Created repository ned1313/taco-app on GitHub ✓ Added remote https://github.com/ned1313/taco-app.git Awesome! We’ve successfully created a local repository for our Terraform code and a remote repository on GitHub. Now we need to push our local code to the remote repository.
|
||
Pushing Changes to the Remote Repository Our local repository now knows about the remote repository on GitHub. We can view the details of the remote repository by running the git remote show command with the name of the remote entry:
|
||
taco-app$ git remote show upstream * remote upstream Fetch URL: https://github.com/ned1313/taco-app.git Push URL: https://github.com/ned1313/taco-app.git HEAD branch: (unknown) See that HEAD branch: (unknown) line? That’s because we haven’t pushed our local code to the remote repository yet. We can do that using the git push command. The -u flag tells git to set the main branch on the remote repository called upstream to track our local main branch.
|
||
taco-app$ git push -u upstream main Enumerating objects: 9, done. Counting objects: 100% (9/9), done. Delta compression using up to 16 threads Compressing objects: 100% (7/7), done. Writing objects: 100% (9/9), 1.79 KiB | 37.00 KiB/s, done. Total 9 (delta 1), reused 0 (delta 0), pack-reused 0 remote: Resolving deltas: 100% (1/1), done. To https://github.com/ned1313/taco-app.git * [new branch] main -> main Branch 'main' set up to track remote branch 'main' from 'upstream'. This also means that in the future, when we run git push, git will know to push to the remote branch we’ve set in the upstream entry.
|
||
Now if you check the info for the remote repository, you’ll see that the HEAD branch is set to main:
|
||
taco-app$ git remote show upstream * remote upstream Fetch URL: https://github.com/ned1313/taco-app.git Push URL: https://github.com/ned1313/taco-app.git HEAD branch: main Remote branch: main tracked Local branch configured for 'git pull': main merges with remote main Local ref configured for 'git push': main pushes to main (up to date) Awesome! We’ve taken our Terraform code, put it in a local repository, and pushed the code up to GitHub. Now we can collaborate with others, have a backup of our code, and automate testing and deployment of our code on check in.
|
||
Next Steps Using version control for your Terraform code is a great first step. But there’s a lot more you can do with version control. You can create branches to work on new features or bug fixes. You can create pull requests to review and approve changes before they’re merged and applied. You can even automate testing and deployment of your code on check in.
|
||
In a future blog post and video, I’ll dig into how you can use source control to publish your own Terraform modules!
|
||
`,summary:`Terraform is an infrastructure as code tool. The key term here being code. And where do you store your code? In a version control system like GitHub! I recently published a video with the awesome April Edwards on creating your first Terraform repository.
|
||
This blog post is a companion to that video. Let’s dive in!
|
||
Why Use Version Control? If you’re seasoned developer using version control, then I don’t need to convince you of its benefits.`,date:"4 Mar, 2024",url:"https://nedinthecloud.com/2024/03/04/creating-your-first-terraform-repository/",image:"Create-Your-First-Terraform-Repo.png",readingTime:"9"},"https://nedinthecloud.com/2024/01/22/comparing-opentofu-and-terraform/":{title:"Comparing OpenTofu and Terraform",tags:["hashicorp","terraform","opentofu"],content:`[Last updated 2025-03-03 for Terraform 1.11]
|
||
Terraform and OpenTofu are both IaC tools that share a common ancestry. OpenTofu was created when HashiCorp shifted the licensing of Terraform from Mozilla Public License 2.0 to Business Source License 1.1 aka BSL or BUSL.
|
||
OpenTofu is a fork on Terraform before the BUSL licensing change, so anything that has been added in Terraform 1.6 or newer will not be in OpenTofu, or at least the code backing any given feature will not be copied over. The OpenTofu maintainers can implement the same functionality using their own original code, but they cannot simply copy the Terraform code even though it is publicly available.
|
||
Likewise, any new features added to OpenTofu will not be back-ported into Terraform. While HashiCorp can add feature-parity, the actual code will be different.
|
||
Which might leave you wondering, which features are supported in each tool? Are they interchangeable? Can I move from one to the other without rewriting my code? The short answer? If you’re running Terraform 1.5.x or older, you should be able to drop OpenTofu in and run your code without any changes. If you’re running Terraform 1.6 or newer, you will need to check the table below to see if the feature you’re using is supported in OpenTofu. The features included in the table only reflect those that are in official releases and not in alpha or beta builds. I will endeavor to keep this table up to date as new GA features are added to either tool.
|
||
Feature Terraform OpenTofu Feature parity Testing framework 1.6 1.6 Yes Mock data in testing 1.7 1.8 Yes removed block 1.7 1.7 Yes Updated S3 backend 1.6 1.6 As far as I can tell, these appear to be identical Provider defined functions 1.8 1.7 Yes, but the terraform built-in provider functions may not be in OpenTofu State encryption Not present 1.7 No Provider transfer of object 1.8 Not present No Input variable validation with other objects 1.9 Not present No Early variable/locals evaluation Not present 1.8 No Support for .tofu files Not present (of course) 1.8 No Ephemeral values and resources 1.10 Not present No Write-only attributes 1.11 Not present No Native S3 locking 1.10 Not supported No Provider block for_each Not present 1.9 No Plan and apply -exclude flag Not present 1.9 No Those are all the major updates in versions 1.6 and newer. I should note that there maybe minor enhancements and bug fixes that are not included here.
|
||
You can always find the latest releases here:
|
||
Terraform Releases OpenTofu Releases If you think that I have missed something, please let me know via the contact form.
|
||
`,summary:`[Last updated 2025-03-03 for Terraform 1.11]
|
||
Terraform and OpenTofu are both IaC tools that share a common ancestry. OpenTofu was created when HashiCorp shifted the licensing of Terraform from Mozilla Public License 2.0 to Business Source License 1.1 aka BSL or BUSL.
|
||
OpenTofu is a fork on Terraform before the BUSL licensing change, so anything that has been added in Terraform 1.6 or newer will not be in OpenTofu, or at least the code backing any given feature will not be copied over.`,date:"22 Jan, 2024",url:"https://nedinthecloud.com/2024/01/22/comparing-opentofu-and-terraform/",image:"opentofu-vs-terraform.png",readingTime:"3"},"https://nedinthecloud.com/2024/01/17/2023-year-in-review/":{title:"2023 Year in Review",tags:[],content:`If I’m being honest, 2023 was a mixed bag. There were several factors that lead to a serious drop in revenue (-30%!) for Ned in the Cloud versus 2022. However, I connected with several new vendors and further established my YouTube channel. In this post I would like to look back at 2023 and see what went well, how I did against my goals, and where I fell short.
|
||
High-Level Review If I could sum 2023 up in a word, it was less. I did less work, made less money, but ironically had no less stress. For the year, I really tried to embrace the fact that I make up my own schedule and decide on what projects to do. I don’t have to do something just because someone asked me to. I can say “No” to things. Especially things that are going to make me miserable.
|
||
So my main mission in 2023 was to close out existing projects that I was not enjoying and avoid taking on new ones that would make me sad. Why work for yourself if you’re going to make yourself do things you don’t enjoy? That’s the big takeaway, I did less and made less, and I’m okay with that- or at least I’ve come to terms with the latter. Now let’s get down into some details.
|
||
Technical Education The central goal of Ned in the Cloud continues to be providing entertaining and educational content for a technical audience. I do that through a variety of mediums, including YouTube videos, podcasts, Pluralsight courses, and live instruction. By that metric, 2023 was excellent!
|
||
YouTube Videos In 2023, I published 50 videos on my YouTube channel. That’s right, 50!!! That’s an average of one video per week, although I did double up on a few weeks and miss a couple others. Looking at the statistics, I clocked in 178k views and I gained 3k subscribers. Compared to 2022, that’s a increase of 6% over the previous year. My most popular videos from 2023 were Exploring the Import Block in Terraform 1.5 and Terraform Certified Associate - What’s New in Version 003?. That’s not terribly surprising. The updated exam and the new import block were both big news in the Terraform community.
|
||
My most popular videos were still Azure DevOps Pipeline with Terraform and Getting Started with GCP and Terraform with 30.8k and 28.1K lifetime views respectively. Both of those video were published in 2021, and they’re starting to look a little dated. Azure DevOps in particular has changed a lot since I published that video. I think in 2024, I need to publish a follow-up to that one.
|
||
My goal for 2023 was to publish two videos a month and grow my subscriber base to 25k. I may only have 12k subscribers right now, but I smashed my publishing goal out of the park. An open question for 2024 is whether I lean into Terraform more heavily, or try and diversify. You’ll just have to wait for my 2024 planning post to find out!
|
||
Day Two Cloud In 2023, Day Two Cloud kept up the weekly publishing schedule and I thought we had some fantastic episodes. Here are a few that I really enjoyed:
|
||
Making The Most Of Red Teaming With Gemma Moore Coaching For Accidental (And On-Purpose) Managers with Steve Dwire How Did We Get To WebAssembly And What Is It For? with Matt Butcher Despite that, the subscriber base shrunk a little, or at least that’s what the metrics are saying. Our total downloads looks to be around 400k, and I don’t have a subscriber count. Revenue was way down from 2022, dropping by 75%. What happened? I think there are a few factors here.
|
||
First off, the economy wasn’t doing super well in the first half of 2023, and talking to other content creators in the tech sphere, marketing dollars basically dried up in the first half of the year. In addition, in-person conferences have made a big comeback in 2023, and marketing dollars that had been diverted to online content during the pandemic returned to fund in-person events. A smaller pie is being sliced up among more events and content creators, which means less money for allocated for podcasts.
|
||
Which brings me to the next thing, the proliferation of podcasts. There are so, so many podcasts out there now, and I think it’s harder to stand out. At the same time, podcast growth has been slowing down and there are several other media formats competing for people’s eyes and ears. The very public implosion of the podcast rush at Spotify serves to highlight that the podcast market went through a major boom cycle, possibly even a bubble, and now we’re seeing the fallout.
|
||
I’m not trying to blame external factors for the drop in listeners and sponsors for Day Two Cloud. I think we can do better in terms of content, format, and marketing. Packet Pushers, of which Day Two Cloud is a part, just hired a CEO and new sales person in 2023 to help market and promote all the podcasts on the platform. We will be brainstorming new ideas and trying new things in 2024 to help grow the audience and revenue for Day Two Cloud. My primary mission remains the same, to help IT professionals learn and grow in their careers. So while the format and marketing may change, the core of the podcast will remain the same.
|
||
Of course, we’re always looking for feedback from the community, so if you have ideas or comments you can submit them here.
|
||
Chaos Lever In 2023, Chris and I worked to refine the format of Chaos Lever. It started out as a Buffer Overflow clone- a podcast we had been doing together for our former employer. But one thing we noticed is how we enjoy looking at the history behind emerging tech to provide context for what’s happening now. With that in mind, we changed the main episode and tag line to promote the historical lens of the podcast. We also split out the lightning round at the end of each show into it’s own show called Tech News of the Week. As it stands now, Tech News of the Week drops every Tuesday and the main Chaos Lever episode releases on Thursdays.
|
||
Remember how I said that I didn’t have a reliable way to track subscribers and downloads for Chaos Lever? Well, in December I set up PodTrac to track downloads and subscribers. With these more reliable numbers, it looks like Chaos Lever is getting about 450 downloads a month with a 216 subscriber count. That’s… a lot less that what I was tracking before. It also doesn’t mesh with the stats I’m culling from Azure directly. According to my custom Azure report, we’re seeing closer to 750 downloads a week. I guess what I’m trying to say is that I think Chaos Lever has been growing, but I’m not sure. That’s okay though, it’s given us a chance to experiment with the podcast and find our voice.
|
||
Speaking of which, at the end of the year I brought on HumblePod to handle the editing, publishing, and promotion of Chaos Lever. That has freed me up to focus on the content of Chaos Lever, and also work on other projects. I’m really excited to see what we can do with Chaos Lever in 2024.
|
||
HashiCorp Academy In 2022, I started delivering live training for River Point Technology and HashiCorp Academy. Fifteen classes in fact! That was… too much. In 2023, I planned to scale it back to one training a month. I ended up teaching only eight classes in 2023, and that was fine! I was signed up to teach a few more, but they ended up getting cancelled due to low enrollment. I did not meet my goal of one class a month. However, that was an upper limit to keep me from over committing. I’m happy with the amount of live instruction I did in 2023.
|
||
Pluralsight Courses At the beginning of 2023, I was very much still in weeds with the Azure Virtual Desktop courses. What was supposed to be done by the end of 2022, was dragging on into 2023 and I was feeling kinda miserable about it. With three courses still to complete, I was also dealing with patching the completed courses due to changes in the AVD exam objectives. I finished patching those courses in January, and we decided to publish them before the rest of the learning path.
|
||
By the end of March, I finished the Azure Virtual Desktop learning path courses and breathed a deep sigh of relief. Was it worth it? From a financial standpoint, those six courses have garnered me about $20k. That’s an okay return and what’s nice is that this is passive income from here on out. Sure, I’ll need to patch the courses occasionally, but for the most part I can just let them sit there and earn money. From a personal standpoint, I’ve received several comments on LinkedIn about how the courses helped people pass the AVD exam. That’s really rewarding to hear, and I’m glad I was able to help them.
|
||
For the rest of 2023, I went on to refresh the Terraform Getting Started, Deep Dive and Terraform Cloud courses. So it looks like I published 9 courses in 2023, although a few of them were basically done in 2022. I did not create a course for Open Policy Agent or HashiCorp Boundary like I had originally planned. Those are both on the table for 2024, although first I am focusing on redoing my Terraform on Azure and AWS courses on the CloudSkills platform (formerly A Cloud Guru).
|
||
There’s a bit more to the story though. I’m not sure if I’ve talked about this publicly before, but at the end of 2022, Pluralsight told their authors that effective immediately, all author payouts would be cut by 25%. To say I was stunned would be an understatement, and the author community as a whole was pretty pissed. The reason given for the cut was that Pluralsight had been incorrectly calculating video revenue and they had “fixed the glitch.” Since we as authors have no visibility into their internal accounting and they are no longer a publicly traded company, we have no way of verifying that claim. Had they been incorrectly calculating revenue share for years? Or was this due to their acquisition by a private equity company that is trying to squeeze more profit out of the company? I’m not in a position to say for sure.
|
||
What I can say is that Pluralsight had two rounds of layoffs over 2022 and 2023, and it’s clear they’ve been trying to cut costs across the board. We also went into 2023 with a fairly bleak outlook on the economy and a possible recession on the horizon. I understand why Pluralsight had to reign in costs, but I also know that the author payouts were based on revenue generated by the videos. So if revenue is down overall, we as authors would be making less money. That’s not what happened. Pluralsight simply cut the revenue share by 25% and blamed it on bad accounting. I’m not sure I buy that.
|
||
What does this mean for me and my future as an author at Pluralsight? To be frank, Pluralsight represents about 70% of my revenue. That’s not something I can just walk away from. Plus, I still think Pluralsight as a platform is a really great place packed with amazing authors! Is this cut going to be a one time correction? Or is it a harbinger of things to come? I’m adopting a wait and see approach. I’ll continue to create new courses and refresh existing ones, while at the same time expanding other areas of my business.
|
||
Books Just like 2022, I did not write any books in 2023. I did have a few offers, but I turned them down. The juice is just not worth the squeeze when it comes to writing technology books. I updated the Terraform Associate certification guide to coincide with the release of version 003 of the exam. Seems like certification guides are the exception when it comes to books. They’re short, they’re useful, and people actually buy them. I’ll keep my certification guides up to date, and maybe even write a new one in 2024 for the HashiCorp Vault Operator exam. I’m not totally sold on that, since it isn’t nearly as popular of an exam as the Terraform Associate.
|
||
I did contribute to a book in 2023, The Best Kept Secrets Of HashiCorp Vault, compiled by fellow HashiCorp Ambassador Bryan Krausen. He polled the community to see who would be interested in contributing a chapter to the book, and many of the Ambassadors raised their hand, including me. The profits from the book are going to charity, which I think is awesome. If you’re a Vault admin, you should definitely check out the book!
|
||
Other Appearances Beyond the podcasts, YouTube videos, and courses, I did appear a few other places in 2023. Here is a short list of items if you’re interested in checking them out:
|
||
Cloud Field Day 17 - This one was in Boston! It was so nice to stay on the east coast for a change and not have to deal with time zone changes or airports. I took the train to Boston and loved it! Edge Field Day 2 - Edge is big area of interest for me, so I was excited to attend this event even if I couldn’t be there in person. HashiTalks 2023 - This was my fifth time speaking at HashiTalks. My topic was Autoscaling Workers for Boundary. Good stuff. HashiConf 2023 - For HashiConf 2023 I was a speaker on both the main stage and the hallway track. It was an incredibly busy week for me, and I had a fantastic time! This is exactly the size conference I like to attend. Firefly Webinar - Firefly became a sponsor of my YouTube channel in 2023 and also had me on for a webinar. I’m really excited about the work they’re doing in the IaC space. Imposter Syndrome Network Podcast - This is a great podcast hosted by Chris Grundemann and Zoe Rose. I got to let my hair down a little on this show and talk about some more personal stuff and not just Terraform. Screaming in the Cloud podcast - Corey Quinn invited me to be on his podcast while we were chatting at HashiConf, and so I did! We talked about Terraform and teaching tech. There may have been some other appearances, but I didn’t write them all down in one place. So if I missed your thing, I’m sorry! Maybe in 2024 I’ll keep a better list, but probably not.
|
||
Deprecations and Failures I set some modest goals for myself in my 2023 Goals post. Did I accomplish all of them? No, and that’s okay. Still it’s good to reflect.
|
||
Certification Guides - I had planned to update both the Terraform and Vault certification guides in 2023. The Vault Associate certification guide hasn’t been updated, but then again neither has the exam. Live Instruction - I had planned to do one live class a month. I ended up doing eight classes in 2023, which is still a lot. I’m okay with that. Website and Blog - I did not publish a blog post after March, and I did not update the website. At least I didn’t update it publicly. Behind the scenes, I worked with someone from Fiverr to port the site over to Hugo and streamline it. My first post of 2024 was about exactly that! Universal Object Reference Project - I made some contributions to the project in Q1 of 2023, but then I got pulled away on other projects and didn’t make any further contributions. Conferences - I attended two conferences in 2023, HashiConf and KubeCon. I also attended Cloud Field Day in Boston, but that’s not really a conference. I failed in my goal of one conference a quarter. The primary goal of 2023 was to say “No” to more things. I wanted to limit my level of commitment after severely over-committing in 2022. Did I say “No” to projects and opportunities? Yes, I did! There were several offers to create a course, write a book, or do some other project that I turned down. And like I planned at the beginning of 2023, I did my best to redirect opportunities to existing projects. For that reason, I ended up with several sponsored YouTube videos.
|
||
The two failures I want to reflect on are the lack of blog posts and my failure to keep up with the UOR project (now called Emporous.)
|
||
Blog Posts I wrote a grand total of three posts in 2023 and then promptly abandoned my blog. Why? What happened? It’s not like I stopped writing! On any given week, I probably end up writing 10k words or more. I write scripts for my YouTube channel, Pluralsight courses, and podcasts. The amount of words I type into a browser or VS Code is frankly embarrassing. Where’s the disconnect? Why am I not blogging?
|
||
Wordpress. The problem is Wordpress.
|
||
This is hard to describe, but I’ll do my best. I prefer low friction interactions that feel snappy. If a process feels like a slog, I’m going to avoid it. Even if the actual time it takes to accomplish something is minimal, in my mind it feels like an overwhelming burden, so I put it to the side and focus on other tasks that don’t require as much mental overhead.
|
||
The Wordpress version of my website felt that way to me. The interface was slow. Logging in alone felt like a chore. Interacting with the editor was like walking through molasses. I didn’t know what half the plugins were doing, whether they were necessary, and why my disk space kept filling up even though I wasn’t posting anything. The CSS would randomly break when I wasn’t looking, and sometimes images just wouldn’t load. Why? No idea.
|
||
What’s nice about Wordpress is that it has plugins to do everything and it abstracts a bunch of stuff away from you. However, it’s really hard to figure out why things are going wrong. The whole setup feels deliberately opaque. I feel like I’m fighting with Microsoft Word in 2003. And I’m losing.
|
||
Additionally, I wanted to update a bunch of pages and reorganize the site, but I struggled to do it in Wordpress. The theme and layout seemed inscrutable to me. Even a minor change took me hours to figure out. Wow, again, getting some really strong MS Word 2003 vibes here. Why did that line suddenly format itself as bold and 8pt Arial? Because you forgot to burn the correct incense at the last midnight mass to Steve Ballmer. Oh and all the formatting marks are hidden and showing them is arcane magic reserved for only the high priest class.
|
||
Not to go on too deep of a tangent, but this is my blog afterall. One of the reasons I love writing in Markdown so much is that none of this shit is hidden. All the formatting is right there in the text. Sure, you have to learn some basic Markdown syntax, but after that, writing in Markdown is a breeze. I’m writing this in Markdown right now. And my two certification guides were entirely written in Markdown and VS Code. Suck it Word.
|
||
The net result was that I didn’t like interacting with my website, so I simply avoided it. I didn’t write blog posts, I didn’t update the site, and I cringed every time I had to visit it. My own website made me cringe. That’s bad. In summer of 2023, I decided that had to change, and moreover that I wouldn’t be able to do it. Time to bring in some help!
|
||
I contracted with a person on Fiverr to migrate my site to Hugo and help me with the layout. We made some decent progress, but ultimately he wasn’t able to finish the design or do the migration. I suppose that’s the way of Fiverr. So I hired someone else to finish the job. It took a few months, but he got everything migrated over, streamlined the design as I requested, and helped me make a few tweaks that were probably out of scope. I’m really happy with the result!
|
||
The website is now hosted on Netlify and I can write and publish posts using Markdown. VS Code, and GitHub. I understand how the site is compiled and published, and I can make changes without spending hours fighting with weird plugins. The process is now low friction, git-based, and fast. I truly believe I’ll end up posting a lot more in 2024 as a result.
|
||
Emporous I wanted to contribute to an open-source project in 2023, and I was inspired by an episode of the now retired Full Stack Journey podcast to reach out to one of the maintainers of the Universal Object Reference project. I spoke to her and she brought me up to speed on the current status of the project and how I might get involved. The project was changing its name to Emporous and they were in the process of renaming the GitHub repo and rewriting the documentation. Writing docs is something I’m good at, so I offered to help with that.
|
||
The group maintaining Emporous was meeting every other week, and I attended some of those meetings and helped write some of the docs on the GitHub repo. But soon I got busy working on other projects and delivering live training, so I had to skip a few meetings. Then I missed a few more. Then I found out the project was being moved to be part of Ortelius, an open-source, supply chain catalog solution. There were some politics involved, and I didn’t really understand what the future looked like for Emporous. So I stopped attending the meetings and contributing to the project.
|
||
Not every project is going to be a home run. And while the people involved were great, I just didn’t have the time or motivation to stay involved. I wish I had been more upfront about that, instead of just fading away. My failure was not stopping contributions, but rather not communicating that I was leaving the project.
|
||
Looking Forward I’m going to be writing a full post about plans for 2024, so I’m not going to go into too much detail here. This post is already pushing 3900 words! (Told you I write a lot) If you’ve stuck around till the end, thanks! The theme for 2024 will be focus. I need to focus on less things and do better to promote them. My primary drivers will be increasing revenue and engaging with stuff I actually like doing. Wild, I know.
|
||
Anyhow, thanks for reading the post and I hope you have a great 2024!
|
||
`,summary:`If I’m being honest, 2023 was a mixed bag. There were several factors that lead to a serious drop in revenue (-30%!) for Ned in the Cloud versus 2022. However, I connected with several new vendors and further established my YouTube channel. In this post I would like to look back at 2023 and see what went well, how I did against my goals, and where I fell short.
|
||
High-Level Review If I could sum 2023 up in a word, it was less.`,date:"17 Jan, 2024",url:"https://nedinthecloud.com/2024/01/17/2023-year-in-review/",image:"2023-Year-in-Review.png",readingTime:"19"},"https://nedinthecloud.com/2024/01/16/terraform-taint-is-bad-and-heres-why/":{title:"Terraform Taint Is Bad, And Here's Why",tags:["hashicorp","terraform"],content:`The terraform taint command marks an existing resource in state data for replacement. On it’s surface, this seems like a useful feature. However, it’s actually a ticking time bomb that can sabotage your environment. In this post, we’ll explore why taint is bad, and what you should do instead.
|
||
If you prefer a video version of this post, click here.
|
||
What Taint Does Sometimes you have a resource in your Terraform configuration that just didn’t provision quite right or you need to force the replacement for reasons outside of Terraform. Maybe it’s a VM who’s setup script bombed and you want to replace it. Maybe it’s a storage bucket you need to empty out and recreate. Whatever the reason, you want to replace an existing resource without changing the configuration. That’s why terraform taint was created.
|
||
The taint command marks a resource in the Terraform state data as tainted. This means that the next time you run terraform apply, that resource will be destroyed and recreated. The configuration for the resource will not change, but the resource will be replaced. Let’s take a look at an example.
|
||
Taint Example If you’re following along at home, the code for this example is in my Terraform Tuesdays repository.
|
||
In the below configuration, I have a single resource, an Azure resource group.
|
||
terraform { required_providers { azurerm = { source = "hashicorp/azurerm" version = "~> 3.0" } } } provider "azurerm" { features {} } resource "azurerm_resource_group" "main" { name = "tainted-love" location = "eastus" } Let’s assume I’ve already run terraform apply and the resource group has been created. Now I want to replace it. The command is terraform taint followed by the identifier for the resource. In this case it’s azurerm_resource_group.main.
|
||
$ terraform taint azurerm_resource_group.main Resource instance azurerm_resource_group.main has been marked as tainted. Before we run a plan, let’s take a look at what Terraform actually did. Running a terraform state show against the resource shows that it is tainted.
|
||
$ terraform state show azurerm_resource_group.main # azurerm_resource_group.main: (tainted) resource "azurerm_resource_group" "main" { id = "/subscriptions/4d8e572a-3214-40e9-a26f-8f71ecd24e0d/resourceGroups/tainted-love" location = "eastus" name = "tainted-love" If we cat out the state file, the resource group has a property called status and it’s set to tainted. This is how Terraform knows to replace the resource.
|
||
$ cat .\\terraform.tfstate { "version": 4, "terraform_version": "1.6.6", "serial": 10, "lineage": "45c5895a-1117-8a59-8ae9-0a61d1b388db", "outputs": {}, "resources": [ { "mode": "managed", "type": "azurerm_resource_group", "name": "main", "provider": "provider[\\"registry.terraform.io/hashicorp/azurerm\\"]", "instances": [ { "status": "tainted", # ... } Now let’s run a terraform plan:
|
||
$ terraform plan Terraform will perform the following actions: # azurerm_resource_group.main is tainted, so must be replaced -/+ resource "azurerm_resource_group" "main" { ~ id = "/subscriptions/4d8e572a-3214-40e9-a26f-8f71ecd24e0d/resourceGroups/tainted-love" -> (known after apply) name = "tainted-love" - tags = {} -> null # (1 unchanged attribute hidden) } Plan: 1 to add, 0 to change, 1 to destroy. Terraform comes back and let’s us know that the resource group will be deleted and recreated, and it also tells us the reason is because the resource is tainted. Good information to have.
|
||
If you want to undo the taint on a resource, the corresponding command is terraform untaint. And the syntax is the same as taint:
|
||
$ terraform untaint azurerm_resource_group.main Resource instance azurerm_resource_group.main has been successfully untainted. If you look at the state file again, the status property has been removed entirely.
|
||
Which makes me wonder what else the status property is used for? I think it’s used with create before destroy actions to flag an older instance of the resource for deletion. More research is needed for that one.
|
||
Now if you run a terraform plan, Terraform says that no changes are necessary because the taint has been removed and the target environment matches the config.
|
||
terraform plan azurerm_resource_group.main: Refreshing state... [id=/subscriptions/4d8e572a-3214-40e9-a26f-8f71ecd24e0d/resourceGroups/tainted-love] No changes. Your infrastructure matches the configuration. Terraform has compared your real infrastructure against your configuration and found no differences, so no changes are needed. All this seems pretty copacetic, so why is taint bad?
|
||
Why Taint Is Bad If you’ve been following the changes to Terraform over the last couple years, you know that HashiCorp is trying to move away from imperative commands and towards a declarative model for all operations that affect state. They are also trying to ensure that only the terraform apply command is used to make changes to state.
|
||
That’s why there is now a moved block and an import block, and soon there will be a removed block too. These replace the imperative commands terraform state mv, terraform import, and terraform state rm with a declarative counterpart. The idea is that you should be able to make all changes to state declaratively through the configuration, preview those changes, and make them with the apply command.
|
||
Why the change? Well, it’s all about state data and being able to preview changes without impacting other team members. If you run a terraform state mv or terraform taint command, you are altering the state data without making a change to the configuration. In a collaborative environment, this can cause problems.
|
||
For a simple example, let’s say that I need to replace a VM in my environment, but I can’t do it until after hours. So I run a terraform taint command to mark the VM for replacement. But I forget to tell my team members that I did this. One of them is making other changes to the configuration, and they run a terraform plan.
|
||
In the execution plan review, they completely miss my VM change, since there’s no difference in the code, and they aren’t looking for it. They approve the plan and run terraform apply.
|
||
Now my VM has been recreated during regular business hours, and I’m getting a call from my boss. Not good.
|
||
The tainted status is like a ticking time bomb in the state data, waiting to go off at an unexpected time whenever someone decides to run an apply. You are also altering state before you can preview the changes it might make to your environment, including other impacted resources. The normal cycle of plan, apply, alter state is flipped to alter state, plan, apply. That’s why taint is bad.
|
||
What To Do Instead The alternative is simple! Starting in Terraform 0.15, the -replace flag was added to the terraform plan and apply commands. This flag allows you to replace a resource without changing the configuration. It’s the same as running a terraform taint followed by a terraform apply, but it’s all done in one command. You can also repeat the flag to replace multiple resources.
|
||
Since the replace flag can be use with the plan command, you can preview the changes before you make them, and more importantly, you can preview the changes without altering state data. This is a much safer way to replace resources.
|
||
If you save the execution plan, and someone else makes a change to the configuration and applies it, your execution plan will show as invalid when you try to apply it and you’ll know that you need to re-run it. The circle of infrastructure life is maintained.
|
||
Let’s try the replace flag with our previous example.
|
||
Replace Example The flag exists for both the plan and apply commands, so I’ll run terraform plan -replace="azurerm_resource_group.main".
|
||
$ terraform plan -replace="azurerm_resource_group.main" azurerm_resource_group.main: Refreshing state... [id=/subscriptions/4d8e572a-3214-40e9-a26f-8f71ecd24e0d/resourceGroups/tainted-love] Terraform used the selected providers to generate the following execution plan. Resource actions are indicated with the following symbols: -/+ destroy and then create replacement Terraform will perform the following actions: # azurerm_resource_group.main will be replaced, as requested -/+ resource "azurerm_resource_group" "main" { ~ id = "/subscriptions/4d8e572a-3214-40e9-a26f-8f71ecd24e0d/resourceGroups/tainted-love" -> (known after apply) name = "tainted-love" - tags = {} -> null # (1 unchanged attribute hidden) } Plan: 1 to add, 0 to change, 1 to destroy. Terraform comes back and tells us that the resource group will be replaced, and it also tells us that the replace flag was used. If we run a terraform apply with the same flag, we get the same result, with a prompt asking us to confirm the changes. Only then will Terraform make the changes to the target environment and our state data.
|
||
It’s just that easy!
|
||
What About Automation? Now you might be thinking, “Ned, you said that HashiCorp was trying to move away from imperative commands and doing everything through the configuration. But the replace flag is still part of an imperative command. What gives?”
|
||
First-off, I’m impressed that you somehow added code formatting to speech. Second, you’re right.
|
||
The replace flag is still part of the imperative plan and apply commands, and if you’re running everything through an automation pipeline, there’s no easy way to use it. You’d have to kludge something together to make it work. Maybe using commit messages or PR comments? More research is required here as well.
|
||
Do I love this? No. It would be great to have a solid alternative for the replace flag in an automated setting. Hopefully, you don’t need to use the replace flag except in rare, break-glass circumstances, and you can use the declarative model for most of your changes.
|
||
Personally, I tend to use replace when I’m working on a new configuration and debugging some issue with a particular resource. Most of the time, it’s a one-off thing and I don’t need to worry about the automation workflow since I’m still running everything at the terminal. By the time it gets rolled out to a collaborative environment with automation, the configuration is working as expected and I don’t need to use the replace flag.
|
||
Conclusion And that my friends is why terraform taint is bad actually. It makes changes to state data outside of the plan and apply loop, and it can cause problems in a collaborative environment. The replace flag is the preferred method for replacing resources without changing the configuration, but it’s not perfect either since it can clash with your existing automation workflows.
|
||
`,summary:`The terraform taint command marks an existing resource in state data for replacement. On it’s surface, this seems like a useful feature. However, it’s actually a ticking time bomb that can sabotage your environment. In this post, we’ll explore why taint is bad, and what you should do instead.
|
||
If you prefer a video version of this post, click here.
|
||
What Taint Does Sometimes you have a resource in your Terraform configuration that just didn’t provision quite right or you need to force the replacement for reasons outside of Terraform.`,date:"16 Jan, 2024",url:"https://nedinthecloud.com/2024/01/16/terraform-taint-is-bad-and-heres-why/",image:"TT-2024-01-09-TerraformTaint.png",readingTime:"8"},"https://nedinthecloud.com/2024/01/03/new-site-who-dis/":{title:"New Site, Who Dis?",tags:[],content:`If you’re reading this, then I must have done something correct. That’s right dear reader, I’ve migrated my website off of Wordpress and onto Netlify using Hugo. I’m not going to lie, this process took longer than I expected. I hired two different people off of Fiverr to perform the migration for me. The first one was a bust, but the second one was able to get things 99% of the way there. The rest was small tweaks and adjustments that I was able to make myself.
|
||
Hopefully, you didn’t get flooded with a torrent of new posts if you’re using the RSS feed. I tried to keep all the links to existing posts the same. I also tried to keep the same structure for the site, but there are some differences. I’ll get into those in a bit.
|
||
Why the Change? The main reason I decided to migrate off of Wordpress was speed. I received several complaints from readers that the site was super slow to load. I tried a few different caching plugins to try and help, but ultimately the site was just too weighed down by all the plugins, themes, and other junk that added to the load time. I’m sure if I had a team of developers working to optimize my Wordpress site, I could have gotten load times down to a reasonable level, but I don’t have a team of developers and I can’t afford to hire one.
|
||
The next big reason was that of cost. It’s not that I was paying a ton of money for hosting, however it wasn’t $0. Contrast that with the cost of running Chaos Lever, which is effectively nothing. I pay for the domain name, and that’s it. There’s nothing so special about Ned in the Cloud that it needs all the bells and whistles of Wordpress. A static site will suffice. Unless I have a sudden over-abundance of traffic, I can remain on the free tier of Netlify for the foreseeable future.
|
||
The third reason was workflow. I spend a LOT of time in VS Code. It’s where I do all my coding and honestly, most of my writing. A few years ago, I had a part-time gig writing technical docs for a startup. They were using a static site to host the docs, and I had to learn their process. I became familiar with Markdown, Hugo, and the git-based workflow for publishing content. While it felt awkward at first, I quickly grew to love the simplicity of Markdown and the ease of generating and viewing the updated site locally. Since that time, I’ve continued to use Markdown to write scripts, blog posts, and even my certification guides for Terraform and Vault. Writing in Markdown and VS Code has become second nature to me, and I wanted to bring that workflow to Ned in the Cloud.
|
||
What’s Different Now that the migration is complete, what’s different? Well, the biggest change should be the speed of the site. I can already tell that it loads significantly faster. The difference is impressive and I’m sure it will get better as Netlify builds up the caching.
|
||
In terms of the site’s structure, I took this opportunity to lean out and simplify the navigation. There are far fewer sections, and each page is more focused. I condensed the podcasts, videos, books, and courses into the single category media. I also combined the about and contact pages into a single about page. The home page has a lot less cruft on it, and I’ve slimmed the blog page down to less categories.
|
||
Under media, the podcast page now has the three podcasts I want to mention: Day Two Cloud, Chaos Lever, and The Daily Check-in. Gone are the guest appearances and the other podcasts I’ve hosted. The videos page is now a single page with YouTube playlists instead of individual videos posts. I also whittled down the books and courses pages to reflect my most popular offerings.
|
||
The goal of the migration was to speed things up and lower costs, but it was also about a better experience for you, the reader. I hope that I’ve achieved that goal. If you have any feedback, hit up the about page and send me a message. I’d love to hear from you.
|
||
What’s Next I don’t have any major plans to change the structure or content of the site. I am hoping that this streamlined workflow will lead to more frequent posting. I haven’t posted anything since March of 2023, and that’s pretty sad. And it shouldn’t be hard for me to post more often! I write something like 15k words a week across all my projects, so I should be able to carve out a few hundred for a blog post. Or just repurpose some of the content I’ve already written for other projects!
|
||
Thanks to everyone who reads the blog, listens to the podcasts, and watches the videos. I appreciate your support and I hope you enjoy the new site!
|
||
`,summary:"If you’re reading this, then I must have done something correct. That’s right dear reader, I’ve migrated my website off of Wordpress and onto Netlify using Hugo. I’m not going to lie, this process took longer than I expected. I hired two different people off of Fiverr to perform the migration for me. The first one was a bust, but the second one was able to get things 99% of the way there.",date:"3 Jan, 2024",url:"https://nedinthecloud.com/2024/01/03/new-site-who-dis/",image:"New-Website-2024.png",readingTime:"4"},"https://nedinthecloud.com/2023/03/10/hashicorp-cloud-engineer-terraform-associate-exam-update-for-version-003/":{title:"HashiCorp Cloud Engineer: Terraform Associate Exam Update for Version 003",tags:[],content:`Are you preparing for the Terraform Certified Associate exam? Did you know there’s a new version? That’s what we’re going to cover in this post.
|
||
Terraform Certified Associate Exam Summary Let’s start with the basics, if you’re not familiar with the Terraform Certified Associate exam, I covered it in some detail in a couple YouTube videos. The tl;dw is that the exam covers nine primary objectives, is intended for people who have been using Terraform in development/prod for about six months, and is a multiple choice test lasting one hour.
|
||
Since Terraform is a constantly changing technology, the exam has to be updated periodically to stay in line with best practices and new features that have been introduced. The latest update is version 003, launching in March of 2023- so, like right now. Later in the post, we’ll go over the differences between version 002 and 003, but first I want to answer some common questions.
|
||
Common Questions For starters, you will be able to take version 002 of the exam until mid-May of 2023. So if you’ve already started studying and don’t want to adjust or find new study materials for version 003, fret not, you have a month and a half to take the previous version.
|
||
Second, if you take version 002 of the exam, your certification is good for two years. That’s the case for either version, but I wanted to make it clear that taking the older version doesn’t change the expiration date.
|
||
Lastly, both exam versions cost the same. There’s no discount on the older version of the exam, or temporary sale to entice people to take the new version.
|
||
With all that out of the way, let’s dig into the differences between version 002 and 003.
|
||
Updates for Version 003 Comparing version 002 and 003, they look largely the same. Some of the changes are cosmetic in nature, like the names of objectives, and some go deeper into deprecated commands or new features.
|
||
Let’s start with the cosmetic changes. Four of the primary objectives have new names, but honestly the only one that really matters is objective nine, which has changed from “Understand Terraform Cloud and Enterprise capabilities” to “Understand Terraform Cloud capabilities”. I know that seems pretty minor, but the focus is clear. You should know what Terraform Cloud can do, but don’t worry too much about Terraform Enterprise- it is just a self-hosted version of Terraform Cloud after all. They’ve also revised the sub-objectives for this primary objective, and I’ll get to those in a moment.
|
||
Removals Next up let’s deal with what has been removed. All sub-objectives dealing with provisioners are gone. 🥳 No longer do you need to know how to use local-exec and remote-exec. HashiCorp has been advising against using provisioners for a while now, and they’ve been steadily adding functionality to core Terraform that replaces the need for provisioners and the null_resource. Check out the new builtin terraform_data resource released with version 1.4.
|
||
The terraform taint command is deprecated in favor of the -replace flag with terraform plan and apply. Instead of knowing about taint, focus on the existence of the -replace flag and how it works.
|
||
The terraform refresh command has also been deprecated in favor of the -refresh-only flag for terraform plan and apply. Again, focus on the existence of the -refresh-only flag and how it works.
|
||
The terraform workspace command has been removed from the exam, but it is not deprecated. HashiCorp has started advising against the use of Terraform core workspaces, which is a contentious opinion I don’t necessarily endorse, but that’s a topic for another time. The point is that you don’t need to know how to use workspaces, but you do need to know that they exist.
|
||
The objective “Configure resource using a dynamic block” is no longer its own objective, but be aware it is still part of the exam. I’m not sure why it was its own objective in the first place, and now it isn’t. You should still know how a dynamic block works.
|
||
That’s the removals, what about additions?
|
||
Additions A new addition to Terraform since version 002 is the cloud block that connects a Terraform configuration to Terraform Cloud, using it as a remote state backend among other things. The previous method was to use the remote backend type, and now the cloud block is preferred. You should know how to use the cloud block to connect to Terraform Cloud with the CLI workflow.
|
||
To further emphasize the change and also to differentiate the available backends for state data, the objective “Describe remote state storage mechanisms and supported standard backends” has been renamed to “Differentiate remote state back end options”. Instead of knowing how the different backends work, you simply need to know what features they might support, that’s basically remote state lock and workspaces. Also, the name is a lot shorter, which I appreciate on an aesthetic level.
|
||
In lieu of the deprecation of terraform taint and refresh, the new objective “Manage resource drift and Terraform state” is basically telling you to study up on the -replace and -refresh-only flags and how they can be used to deal with changes made outside of Terraform.
|
||
Another new feature of Terraform is the .terraform.lock.hcl file generated when you run terraform init against a new configuration. The lock file records the constraints for providers and modules, and what version of each provider is being used. This lets you lock in a specific version of a provider and check the lock file into source control for consistency. Make sure you know what the lock file is, what it contains, and how to update it.
|
||
Introduced back in version 0.15, sensitive values are now part of the exam, included in the “Demonstrate use of variables and outputs” objective. You should know how to mark variables and outputs as sensitive, what it means functionally for them to be marked as sensitive, and how to remove the sensitive attribute from a variable or output.
|
||
Turning back to the last objective, regarding Terraform Cloud, the original three sub-objectives have been slimmed down to just two:
|
||
Explain how Terraform Cloud helps to manage infrastructure Describe how Terraform Cloud enables collaboration and governance The first objective basically covers workspaces in Terraform Cloud. You should know what a workspace is, how it works to provision infrastructure, and the three available workflows for workspaces.
|
||
The second objective is more about the governance available from Sentinel in Terraform Cloud, cost control, and collabortion via the private registry and teams. My recommendation is to sign up for an account and start the free trial that includes all features. The trial is good for 30 days and doesn’t require a credit card or any other form of payment.
|
||
TL;DR My big takeaway, and the thing I’d recommend focusing on is the new functionality. It’s not going to hurt you to know the deprecated stuff, but you’ll struggle a bit if you don’t know about the new features. If you’re going to study anything, focus on the following:
|
||
The cloud block and how it works The .terraform.lock.hcl file How to use -refresh-only and -replace How to mark a variable or output as sensitive Certification Guide If you want to do a bit of studying before the exam then may I humbly suggest my certification guide on Leanpub? I just finished updating it to be in line with version 003 of the exam, and since I publish it through Leanpub, you’ll get updates to the guide for free. That means when version 004 of the exam comes out and you need to recertify, you’ll have the latest version of the guide available to help you through.
|
||
And if you use this link to buy the guide you’ll get the guide for a mere $10 instead of the suggested $15 until April 7th.
|
||
Summary That’s about all you need to know about version 003 of the exam. If you decide to sit the new version the exam, let me know about your experience pass or fail, I’d love to hear about what did or didn’t help in this post or the guide.
|
||
Thanks for reading and happy Terraforming!
|
||
`,summary:`Are you preparing for the Terraform Certified Associate exam? Did you know there’s a new version? That’s what we’re going to cover in this post.
|
||
Terraform Certified Associate Exam Summary Let’s start with the basics, if you’re not familiar with the Terraform Certified Associate exam, I covered it in some detail in a couple YouTube videos. The tl;dw is that the exam covers nine primary objectives, is intended for people who have been using Terraform in development/prod for about six months, and is a multiple choice test lasting one hour.`,date:"10 Mar, 2023",url:"https://nedinthecloud.com/2023/03/10/hashicorp-cloud-engineer-terraform-associate-exam-update-for-version-003/",image:"TT-2023-03-08-TerraformAssociateUpdate.png",readingTime:"7"},"https://nedinthecloud.com/2023/03/06/kubernetes-monitoring-tutorial-prometheus-grafana-and-robusta/":{title:"Kubernetes Monitoring Tutorial – Prometheus, Grafana, and Robusta",tags:["devops","kubernetes","robusta"],content:`Are you working in the Kubernetes space and looking for a way to tie together your monitoring tools like Prometheus and Grafana? I took Robusta for a spin and here’s what I discovered.
|
||
Background I’ve been working on some kind of technology for over 20 years now (don’t ask how long), and one of the things that has always been a challenge is monitoring of resources. While the tools have evolved from Nagios, Cacti, and even- god forbid- SCOM, the actual challenges have been roughly the same. Here are the big ones for me:
|
||
What do I monitor? What does healthy look like? How do I investigate when something goes wrong? How do I separate the signal from the noise? Most solutions try to do everything for you, and usually that means they don’t do any of it particularly well. For instance, when I was working as a consultant setting up System Center Operations Manager, you would first deploy the service (no small feat in itself), enroll a bunch of servers, and install even more management packs.
|
||
Immediately, you would be overwhelmed with a sea of red alerts telling you about all the things that were wrong with your environment. 90% of the alerts were false positives, rendering the whole installation completely useless until you spent months tuning it. It’s not that the alerts were necessarily wrong, it’s just that the defaults of the management packs were a bit too aggressive or not context aware.
|
||
And context really is key here. The errors I used to get out of SCOM or vRealize were often not helpful, and they didn’t provide or understand the larger context of the environment they were functioning in. That made tracing down root cause a real challenge as I waded through a fog of alerts, playing whack a mole with false positives.
|
||
Once I did find the issue, I needed to know what to do. In the early days of monitoring, you had your Google-Fu and internal knowledge base to go off, but now I would expect any modern monitoring tool to provide me with a knowledge base article or even a best practice guide.
|
||
A complete monitoring system needs to be simple to setup, understand context, and provide helpful insight. Does it sound like I’m asking a lot of a single solution? That’s because I am. And maybe the answer is not a single solution, but a collection of tools that work together to provide a complete solution.
|
||
The Linux Way One of the key tenets of Linux is creating small tools that do one thing well, and then stringing those tools together to create a complete solution. We can extend the Linux philosophy to Kubernetes, since it was certainly founded in that environment.
|
||
I think a lot of people view Kubernetes as a developer-centric solution, and in many ways it is. But it’s important to remember that Kubernetes is a platform for running applications, and that is an operational task. K8s is Ops by design, with a developer friendly API.
|
||
Bearing that in mind, what do us Ops folks need to monitor?
|
||
The Nodes comprising the cluster The Pods running on the cluster The Services running on the cluster The API Server of the cluster itself And any other custom resources you might have deployed. It’s also unlikely there’s only one cluster in most organizations. We need to be able to monitor multiple clusters and possibly compare them to each other, especially when you have dedicated clusters for dev, staging, and production.
|
||
What tools do we have to help us with this? Most commonly, we have Prometheus to collect metrics, Grafana to visualize them, and AlertManager to notify us when things go wrong. But we still need something to investigate the problem, provide context, and possibly automate a solution. One potential tool is Robusta.
|
||
Robusta Robusta is both an open-source project and the company who created it. I want to focus on the open-source project first, and then we can cover the commercial offering.
|
||
The central goal of Robusta is to take the alerts fired off by AlertManager, enhance those alerts with data from Kubernetes, and then produce events that can be sent to a notification system. Out of the box it includes a set of rules for common issues to watch, and you can also create your own rules. Why don’t we walk through an example of deploying a crashing pod and see how Robusta reacts?
|
||
Crash Pod Alert Example I have two AKS clusters running in my Azure subscription, with Robusta deployed on both using Helm. You can try this out for yourself by running through the demo files in my GitHub repository.
|
||
I’ve deployed a crashing pod to the first cluster using the commands stored in the crashing_pod directory. After running the apply command, we now have a pod that will crash every time it runs, creating a crash loop backoff.
|
||
I’ve got Robusta wired into my Slack team, so I get a notification from Robusta when the pod crashes, on a channel I configured earlier. In the notification, I get the name of the cluster, the pod, the namespace, etc. It also includes the logs from the pod, so I don’t have to grab them myself.
|
||
Clicking on Investigate takes me to the Robusta dashboard where I can see a timeline of issues based on the pod.
|
||
This gives me a bunch more context about the application, namespace, and cluster. I know the pod is crashing by design, but this could also show me in the timeline that the pod crashes every hour for some unknown reason. Or that the crash correlates with some other event in the environment.
|
||
Since we’re here, let’s dig into a few other features.
|
||
Robusta Dashboard Before we get too far into the dashboard, I want to mention that we are now wading into the commercial and closed-source side of Robusta. The open-source version includes alert and log correlation, additional context, notification routing like we saw in Slack, the automation engine, and multi-cluster support.
|
||
Their SaaS dashboard is free for up to 20 nodes, which is plenty for a small cluster, but if you want to go beyond that scale, then you’ll need to upgrade to the paid version. There’s also an enterprise option to host your own instance of the Robusta UI, if SaaS doesn’t work for your organization.
|
||
Here in the UI, I can check out the timeline for the entire cluster and create a customized view based on an application, cluster, and namespace and save it as a preset.
|
||
There’s another feature in the UI that I find extremely useful, the Comparison section.
|
||
Let’s say I’ve got two clusters, one for development and one for production, both with the same application deployed. Ideally, I want the two environments to closely mirror each other, and if there are differences I want to be able to spot that quickly. That’s exactly what the comparison feature does. I can compare namespaces from the same or different clusters, view the resources in each, and toggle a button to see only differences.
|
||
If you’re trying to determine drift between clusters, this is an essential feature.
|
||
I mentioned earlier that the open-source version of Robusta includes built-in rules. But it’s more than just sending an alert with some info, there’s a sneaky automation engine under the hood.
|
||
Robusta Rules Okay, it’s not that sneaky, but hear me out. The rules are housed in Robusta playbooks and they are composed of a trigger and actions. The output of an action can be send to a destination called a sink. The trigger determines when the rule is fired, the actions are the things that should happen, and the sink is where the output of the actions should go.
|
||
You can trigger on an event from the Kubernetes API server, Prometheus and AlertManager, on a schedule, from a webhooks and more. The webhook option means you can trigger on just about anything, including your own custom events.
|
||
Moving to actions, Robusta comes with a set of built-in actions it can take, but you can also write your own custom actions in Python and have them run as well.
|
||
The sinks are targets for your output, and there are a lot of possible sinks. You’ve already seen Slack, and you can also add Teams, PagerDuty, Jira, and more; really if the target has a webhook, you can send data there.
|
||
This means the automation engine is ridiculously extensible and customizable. Robusta begins with useful defaults, but you can easily add your own triggers, actions, and sinks to make it do whatever you need it to do. And there’s a community of people who are doing just that and sharing their work. One fun example to play with is Robusta’s implementation of a ChatGPT enhanced bot. You can check out the code on GitHub.
|
||
Try it out! As I mentioned earlier in the post, I’ve put together a repository that uses Terraform (what else would I use?) to deploy two AKS clusters and install Robusta using Helm. It also includes some scripts for deploying differnt applications, including the crashing pod you saw earlier; a full blown twelve-factor, polyglot app app with a traffic generator; and a basic PHP app with a load generator. That should give you plenty to play with and break to see how Robusta handles it.
|
||
Conclusion I’ve barely scratched the surface on Robusta, but my main takeaway is that it’s a solid tool that serves a scoped purpose really well. At its core, it’s a rule processing engine that loves Kubernetes. It takes what Prometheus collects, info from Grafana, alerts from AlertsManager and makes that information more useful to the poor Ops folks who have to deal with it.
|
||
Having been one of those poor Ops people on more than one occasion, I appreciate the context added to alerts. I also appreciate that the automation component is open-source and extensible without using some arcane language or needing to compile binaries in Visual Studio (SCOM!). If you know a little Python, you’re good to go.
|
||
Of course, I don’t know Python all that well, so I’d love to see a low-code/no-code alternative for the Playbook creation. I’d also love to see some integration with the application level of running pods that have been instrumented with OpenTelemtry. Maybe that already exists and I haven’t seen it yet, or maybe that’s beyond what Robusta wants to focus on. I do appreciate focusing on a core problem and delivering a solution that solves it well.
|
||
Thanks to Robusta for sponsoring this post and thanks to you for reading! If you have any questions or comments, please reach out on Twitter or LinkedIn.
|
||
`,summary:`Are you working in the Kubernetes space and looking for a way to tie together your monitoring tools like Prometheus and Grafana? I took Robusta for a spin and here’s what I discovered.
|
||
Background I’ve been working on some kind of technology for over 20 years now (don’t ask how long), and one of the things that has always been a challenge is monitoring of resources. While the tools have evolved from Nagios, Cacti, and even- god forbid- SCOM, the actual challenges have been roughly the same.`,date:"6 Mar, 2023",url:"https://nedinthecloud.com/2023/03/06/kubernetes-monitoring-tutorial-prometheus-grafana-and-robusta/",image:"RobustaVideo1.png",readingTime:"9"},"https://nedinthecloud.com/2023/02/09/terraform-cloud-managing-your-workspaces-with-projects/":{title:"Terraform Cloud - Managing Your Workspaces with Projects",tags:["hashicorp","terraform-cloud"],content:`Guess what?! That thing I’ve been complaining about for the last 2 years is finally here! Terraform Cloud Projects! Before we get into exactly what they are, here’s a little background for those who might not be familiar with my previous rants.
|
||
Background Terraform Cloud separates things into Organizations and Workspaces inside organizations, and that is where the hierarchy ends. As someone who comes from a Microsoft background, I’m used to multiple hierarchical layers for administrative and security reasons. Active Directory OUs, Azure Management Groups, File Servers with NTFS permissions, etc.
|
||
Having a flexible hierarchy construct is useful for organizing things, but also useful for applying permissions that are inherited by the objects lower in the hierarchy. I can make you an administrator of an Azure subscription, and all the resources in that subscription will inherit that permission. That’s super helpful!
|
||
I can also group all of my related resources inside the same subscription or resource group, making it easier for others to suss out which things are related. Bonus, when you’re viewing things through a dashboard or UI, adding a layer of hierarchy makes it easier to find things.
|
||
Previous to the introduction of projects, all workspaces in Terraform Cloud existed in a flat hierarchy. This meant that if you wanted to apply permissions to a group of workspaces, you had to do it one by one. While it is possible to use Terraform to configure Terraform Cloud, that’s not the best solution in my view. Tags for workspaces were introduced to help with filtering and provide additional metadata about a given workspace, but they didn’t do anything to alleviate the pain of managing permissions.
|
||
With all that in mind, I think we’re ready to talk about projects.
|
||
Introducing Terraform Cloud Projects A project in Terraform Cloud is a container for workspaces. Every workspace needs to be a member of a project. Existing workspaces have been automatically placed in the Default Project, which is in fact called “Default Project”, space and all.
|
||
You can create new projects and move workspaces between projects as needed. When you create a new workspace, you will be prompted to select the project it should be a member of. If you don’t specify a project, it will be placed in the Default Project.
|
||
Speaking of the Default Project, you can rename that project, but you cannot delete it. After all, Terraform Cloud needs to put all of your existing and new workspaces somewhere, and if you don’t specify a project when you create a new workspace (perhaps with the CLI) the Default Project is where it lands.
|
||
Creating a Project You can create a project through the UI, API, or using Terraform. The tfe provider has been updated to include the tfe_project resource. You can find the documentation for that resource here.
|
||
Using the CLI Workflow with Projects If you are using the CLI workflow with Terraform Cloud, then you know that you have to use the cloud or remote block to use Terraform Cloud as a backend. As of right now, neither block has any idea about projects.
|
||
In fact, the workspace configuration at the CLI level is completely unaware of the existence of projects. It doesn’t show up in the local terraform.tfstate file, and there’s no option to specify a project in either block type.
|
||
Fortunately, it doesn’t matter if you move the workspace to a different project after it’s created. You local CLI workflow will continue to function without issue, assuming your account still has access to the workspace in the destination project. It bothers me that the cloud and remote block don’t have a project argument, but this approach also kept the addition of projects from breaking existing deployments.
|
||
An interesting side effect is that each workspace still needs to be uniquely named across the entire organization, even if two workspaces with the same name are in a different projects.
|
||
If you create a new workspace through the CLI workflow, it will be placed in the Default Project. You can then move the workspace to a different project through the UI or API. Alternatively, you can pre-create the workspace in the correct project, and then use the CLI workflow for deployments.
|
||
If you’re using the VCS workflow, then all of this is moot. You don’t use the cloud or remote blocks with the VCS workflow. You would assign the correct project when you create the workspace for the VCS connected repo.
|
||
Permissions Now that we have projects, there are a few things to consider when it comes to permissions. Who can create a project? Who can move workspaces between projects? Who can delete a project? Who can view a project?
|
||
Essentially, there are three levels of permissions: Organization, Project, and Workspace.
|
||
Let’s walk each of these in turn.
|
||
Organization Level Permissions Project permissions are managed at the organization level. There are two new permission sets that govern projects and workspaces. The first is View all Projects and Workspaces. It pretty much does what it implies. The second is Manage all Projects and Workspaces. This permission allows you to create, edit, and delete projects and workspaces, as well as move workspaces between projects.
|
||
There remains two vestigial permission set at the Organization level that only govern workspaces. These are the permission sets that existed before the introduction of projects, and HashiCorp chose to introduce new permission sets rather than expand the existing one.
|
||
There first is View workspaces only. Can a person with that permission set see all workspaces regardless of what project they are in? Yes! Yes they can. I assume this is because HashiCorp didn’t want to break anyone’s existing permissions. But I don’t love it. You might move a workspace into a the Super Secret Project thinking you’ve hidden it from view only to discover that’s not the case.
|
||
There’s also the Manage workspaces only permission set that allows you to create, edit, and delete workspaces, but not projects. Someone with that permission set can edit and delete any workspace, regardless of project, but can only create new workspaces in the Default Project. Again, the approach here seems to be to not break existing permissions.
|
||
Again, I don’t love this. I’d rather the existing permissions be restricted to the Default Project, and then have new permissions for all other projects and the workspaces inside. Regardless, when you view a workspace you don’t see the permissions it inherits from the organization or project level. When it comes to security, that type of visibility is important.
|
||
Project Level Permissions Each project can have permissions associated to various teams. As usual the owners team has full permissions. Beyond the owner association, there are two access levels: Read and Admin. Read can read the project name and the workspaces inside the project. Admin can do just about everything, including the following:
|
||
All read permissions Admin for all Workspaces inside the project Move workspaces (more on that in a moment) Edit the project Manage team access to the project Delete the project Create workspaces in the project The Move workspaces permission is interesting. Where can they move a workspace to? The answer is any other workspace that they have Admin permission on. So if Bob has Admin permission on Project A and Project B, he can move a workspace from Project A to Project B. Good job Bob!
|
||
I should note that you cannot delete a project that has workspaces in it. You must first delete or move all the workspaces out of the project before you can delete it.
|
||
Workspace Permissions You remember like 800 words ago when I said that I wanted to assign permissions based on hierarchy? Projects lets you do that, kind of. You cannot assign the workspace level permissions at the project level. For instance, if I wanted to give Bob Plan permissions on all the workspaces in Project A, I can’t do that at the project level. I can grant Bob Read access or Admin access at the project level, but not the workspace level specific permissions. Do I love this? Again, no.
|
||
The project level permissions also do not allow you to set granular permissions for the project or the workspaces inside. This is unlike workspace permissions, that allow you to select pre-canned access levels (like Plan) or create your own custom permissions set. Maybe I want to grant Alice Read access to the project and Admin access to the workspaces inside. The Admin project level permission set would grant Alice too much access, and that’s no bueno for to principle of least privilege.
|
||
Improvement and Closing Thoughts I’m very glad to see projects rolled out to Terraform Cloud and I’m certain that it will evolve and improve over time. In fact, I’m in contact with the product team over at HashiCorp, so if you have any feedback, please let me know!
|
||
For my part, I’d like to see the following improvements:
|
||
Allow workspace level permissions to be assigned at the project level Enable custom permissions sets at the project level Add a project argument to the cloud block for the CLI workflow Restrict older organization level workspace permissions to the Default Project Duplicate workspace names in separate projects (this might seem strange, but I think it would be useful) Inherited permissions shown in the UI for projects and workspaces What are your thoughts on projects? Do you have any feedback for HashiCorp? Let me know via Twitter or LinkedIn!
|
||
`,summary:`Guess what?! That thing I’ve been complaining about for the last 2 years is finally here! Terraform Cloud Projects! Before we get into exactly what they are, here’s a little background for those who might not be familiar with my previous rants.
|
||
Background Terraform Cloud separates things into Organizations and Workspaces inside organizations, and that is where the hierarchy ends. As someone who comes from a Microsoft background, I’m used to multiple hierarchical layers for administrative and security reasons.`,date:"9 Feb, 2023",url:"https://nedinthecloud.com/2023/02/09/terraform-cloud-managing-your-workspaces-with-projects/",image:"Managing-Workspaces-with-Projects.png",readingTime:"8"},"https://nedinthecloud.com/2023/01/02/planning-for-2023/":{title:"Planning for 2023",tags:[],content:`For 2022, I didn’t go into the new year with a solid plan for what I wanted to accomplish, and more importantly what I wasn’t going to do. As a result, I ended up accepting projects that I did not have the bandwidth for and ultimately wasn’t excited to do. In 2023, I’d like to have a solid list of things that I want to do and stick to that list unless there is some overriding condition, which I will also try and quantify. My hope is that having a list to refer back to will stop me from overwhelming myself or taking on projects that I shouldn’t.
|
||
Let’s get started!
|
||
Things to Accomplish This is the list of things that I want to get done in 2023. Where possible, I am trying to put some numbers against these goals to track my progress over the year and see how I did.
|
||
YouTube Videos Last year I produced 13 videos and ended the year with 8k subscribers. In 2023, I would like to produce two videos per month and increase my subscriber base to 25k. Lofty goals? Yes. But doable I think.
|
||
In terms of content, at least half the videos will be Terraform related. That’s my bread and butter and it’s what people tend to watch the most. Additionally, I plan to create monthly videos for ActualTech Media events and cross-post them to my YouTube channel. The remainder of videos will likely be focused on another project that I will get to in a moment.
|
||
Day Two Cloud Day Two Cloud finished the year with 420k downloads and 21k subscribers (roughly). I don’t want to change anything specific about the Day Two Cloud content - I think we’re doing very well in that department. Instead, I want to focus on promoting Day Two Cloud episodes and growing the subscriber base. I also want to increase revenue by 25%. Let’s target a goal of 550k downloads and 30k subscribers.
|
||
Chaos Lever Chaos Lever launched in early 2022, and I didn’t really start tracking metrics until November. We are averaging about 350 downloads per week with a subscriber base that I don’t even know how to quantify. Since I don’t yet have a way to track total subscribers (maybe I can do something with unique RSS feed pulls?), I think I will target a average of 700 downloads per week. I am also starting a newsletter based on the content of each week’s episode. I’d like to get about 150 subscribers by the end of the year. Modest goals, but achievable!
|
||
Pluralsight Courses As I mentioned in my year end post, I really need to finish the learning path for Azure Virtual Desktop. My goal is to finish it by the end of February and breath a DEEP sigh of relief. Aside from that, my goal is to publish updates for four of my Terraform courses, and create a new course focused on either HashiCorp Boundary or Open Policy Agent. I think the audience is bigger for OPA, so that’s what I’m leaning towards. Combined with the three AVD courses, that’s a total of four new courses and four redos. Seems possible.
|
||
Pluralsight is also pushing their labs pretty hard. Maybe I do a couple of these? That’s a stretch goal and I will not prioritize it over other work. (See that Ned? You bolded it so you would remember.)
|
||
I will NOT be updating my existing pure Azure courses. Pluralsight has made it clear that they will be moving this content over to A Cloud Guru and I suspect I won’t have time to recreate the courses for that platform.
|
||
ActualTech Media As I mentioned in the YouTube section, I am going to try and produce one video a month for ActualTech Media’s events. I’ll probably only make 10 in total, and that seems like plenty.
|
||
Certification Guides Both the Terraform Associate and Vault Associate certifications will be updated in 2023. I need to update my certification guides to keep pace. The Terraform one is already 50% done, which is great. My goal is to get that update done by the end of January. Vault will have to wait until March.
|
||
Live Instruction Last year I taught 15 live courses for Terraform and Vault. That was great from an income perspective, but probably too much for me. In 2023, I want to target doing one a month. That’s plenty!
|
||
Website and Blog The Ned in the Cloud website needs a little spit and polish. Buffer Overflow still crops up in pages, even though it hasn’t published an update since April 2021. I would like to reorganize things to reflect the projects I am focused on.
|
||
I should also try and publish a blog post about twice a month. One can be based on a Chaos Lever episode and the other can be based on a Terraform YouTube video. This is an easy win and allows me to recycle my content in different mediums.
|
||
Tech Field Day and Gestalt IT I did some blogging and paid events for Gestalt IT in 2022. I’m okay with doing a few for 2023. The limit will be one project per quarter and no more.
|
||
Universal Object Reference Project This coming year I’d like to get involved in an interesting open-source project from the very beginning. I was listening to an episode of Full Stack Journey with Scott Lowe and I heard about the Universal Object Reference (UOR) project from Kat Morgan. The project piqued my interest based on my works with early versions of Wasm modules that leverage OCI-compliant interfaces. If the Open Container Initiative spec could be expanded to support Wasm and Helm, what else could it do? That’s the question UOR aims to answer.
|
||
The project seems interesting from a technical perspective, and they are actively looking for folks to help contribute to docs as the project develops and matures. My time writing technical docs for Solo.io was a lot of fun and it forced me to learn about new technologies like OCI. I hope that I can help move the UOR project forward, while also learning more about the adjacent tech it will support.
|
||
I’m not sure how to gauge this goal aside from making active contributions to the project. As long as I do that, I’ll consider it a success!
|
||
Conferences In 2022, I got to attend and speak at the HashiConf Global conference. I also attended re:Invent, however briefly. That was it for conferences I attended in person. I also went to Cloud Field Day 14 in Autumn, which is not exactly a conference, but does require some amount of travel.
|
||
For 2023, there are so many conferences I’d like to attend, but I have to balance my desire to go with obligations at home and other projects I want to accomplish. I think limiting myself to one a quarter is fair. In Q1 I am already traveling to Arizona to visit friends and run a trail race. That’s all I want to do for Q1, since I have a lot of other projects on my plate for the first quarter of 2023.
|
||
In Q2, there’s not a ton of conferences going on in the US. KubeCon EU is in April and HashiConf EU is in June. Put me down as a maybe for one of those. I’d also love to go to a local DevOps day, if one ever happens around me again.
|
||
Q3 is when I tend to take personal vacations, so fitting in a conference might be tricky. If there’s a Tech Field Day event around this time, that may be my best bet.
|
||
Q4 is conference palooza! I’m planning to go to HashiConf Global again, and put me down as a maybe for a couple days at re:Invent or KubeCon US. I really enjoy walking the expo floor and chatting with vendors. I think I could spend two days doing that, and then fly home.
|
||
Other potential conferences are from Microsoft. There’s the Global MVP Summit in March which might be in person, Microsoft Build in Spring, and Microsoft Ignite in Autumn. At this point, I have no idea if any of them will be in-person. Best to adopt a wait and see attitude.
|
||
Things to Avoid A general rule for this coming year is “No”. If the request is not part of the above goals, then the answer is a polite, “not at this time”. And I know that something will come up. A training company will ask me to create a course, book, project, etc. for them around Terraform/Vault/K8s/Azure. A vendor will ask me to blog for their site. A startup will ask me to do consulting for them. One of the above organizations will ask me to do something that is outside of my goals with them.
|
||
The answer has to be “No.”
|
||
Unless…
|
||
Exceptions I’m not immune to money. Ned in the Cloud is a business after all. In most cases, I should try and redirect the request to my existing projects. Want me to do a video for your startup? I have a YouTube channel and 24 videos to produce this year. Why don’t you sponsor one? You’d like me to do a webinar about your cool new feature release? Day Two Cloud does sponsored content, including ads, tech bytes, and full episodes.
|
||
The exception will be for projects that I feel will materially impact my top-line revenue for the year. And if I make that determination, then I need to sacrifice one of my other goals to keep my time allocated appropriately. What does a material impact look like? 10% increase in revenue sounds like a good figure. If I expect to pull in $500k in top-line revenue for 2023, and the project will earn me $50k, then I should do it. Otherwise, it’s a pass. Stick with my current goals and refer them to someone else who’s looking for work.
|
||
Conclusion Setting a proper work/life balance and keeping myself from being over-committed is a constant struggle. I’m hoping that setting measurable goals at the beginning of the year can keep me on task and prevent me from agreeing to new projects out of a fear of missing out.
|
||
And I will miss out on stuff. The tech industry is simply too vast and ever-changing for me not to. My mission is to focus on my core content and leave myself a little time to experiment. At the end of the year, we can review my goals together and see how I did.
|
||
`,summary:"For 2022, I didn’t go into the new year with a solid plan for what I wanted to accomplish, and more importantly what I wasn’t going to do. As a result, I ended up accepting projects that I did not have the bandwidth for and ultimately wasn’t excited to do. In 2023, I’d like to have a solid list of things that I want to do and stick to that list unless there is some overriding condition, which I will also try and quantify.",date:"2 Jan, 2023",url:"https://nedinthecloud.com/2023/01/02/planning-for-2023/",image:"Planning-for-2023.png",readingTime:"9"},"https://nedinthecloud.com/2022/12/30/2022-year-in-review/":{title:"2022 Year in Review",tags:[],content:`As 2022 winds to a close, it’s time to reflect back on the year. There were some triumphs, some failures, and a fair bit of frustration as I continue to try and figure out what Ned in the Cloud LLC should focus on and what I should avoid. If there’s one thing to take away from 2022, it’s that in 2023 I should endeavor to do less. This is not exactly a new problem for me and I think I’m handling my overall workload better now than in previous years. I’m no longer the panicking new business owner in my first year of operation just praying that I can keep the lights on. Financially, Ned in the Cloud had a banner year - more on that later - and that tells me I can slow down a little and enjoy the journey.
|
||
I like the format I used for the post last year, so let’s do it again. First I’ll focus on my various technical education endeavors. Then we can look at some of my failures. Finally, I’ll talk about some of the things I’m looking forward to in 2023.
|
||
Technical Education The primary focus of Ned in the Cloud is to provide entertaining and engaging technical education. Measuring success is a function of how well I’ve met that goal. By that measure, I think I can declare 2022 a resounding success!
|
||
YouTube One of my goals for 2022 was to get back to regular video production on my YouTube channel. How did I do? I managed to publish 13 new videos, which comes out to about one a month. My hope was to come closer to two a month, and I did not reach that goal. Nevertheless, the videos I created amassed 21k views! My most successful video of the year was “Choosing Between Count and For-Each”, which tells me that Terraform fundamentals will always do better than niche content like “Using OIDC with GitHub Actions and Terraform”.
|
||
Taking the wider view of all my video content, I had over 167k views for over 15k hours and picked up 4k subscribers, more than doubling my subscriber count in 2021. This is all despite producing less overall videos. I’m starting to suspect that it’s quality over quantity! What was my #1 video for 2022? It was a close tie between “Getting Started with GCP and Terraform” and “Azure DevOps Pipeline with Terraform”.
|
||
I have two key takeaways from my YouTube channel this year. First, I should focus on fundamentals, like “Getting Started” or investigating a specific feature of Terraform. Second, more is not always better and consistency of release cadence is a fool’s errand. I was sporadic at best and still doubled my subscribers.
|
||
Still though, for 2023, I would like to meet the two videos a month goal. I plan to stick with Terraform for one of them, and then maybe start back up with HashiCorp Vault for the other. Or maybe I’ll get into a new solution? (That’s foreshadowing for a future post in case you weren’t sure)
|
||
Day Two Cloud Day Two Cloud had a great year in terms of shows and sponsorships. Revenue increased by almost 70% over the previous year. Seriously, D2C has become a major source of revenue for Ned in the Cloud, and that makes me feel like we’re doing something right. Vendors want to advertise with us and listeners enjoy what we’re creating! We had a total of 420k downloads, with each episode being downloaded an average 10k times- excluding the most recent episodes that haven’t had their “long tail” grow yet.
|
||
There were so many great episodes this year. Here are five that I particularly enjoyed!
|
||
Chris Wahl came on to share his thoughts on leadership versus management. Richard Seroter gave us a much needed update on where Google Cloud is today and where it’s headed. I’ve been hearing so much about Open Policy Agent, it was great to have on Anders Eknert to give us the straight dope. Any time Tanya Janca wants to be on your show, the answer is always “YES”. Multi-cloud is a fraught state of affairs for the budding cloud technologist, Forrest Brazeal helped us construct a path forward. There’s also one more episode I’d really like to recommend. Ethan and I did a mental health episode with just the two of us sharing our experience of IT burnout, panic attacks, and dealing with anxiety. It was impromptu and brutally honest.
|
||
We already have two months worth of episodes in the can for 2023, and already there are some amazing episodes in store. I’m hoping to keep growing the Day Two Cloud audience and finding amazing guests to put in front of your ears.
|
||
The Daily Check-In Is it better to burn out or fade away? For The Daily Check-In podcast, I chose an Irish Goodbye. No big brouhaha, since I’m not going anywhere. But keeping up a daily cadence on a podcast just wasn’t in the cards. The final episode dropped on February 22nd, and ironically it was about bringing back the Buffer Overflow podcast that I used to do with coworkers at my previous employer. Speaking of which…
|
||
Chaos Lever Is it called Buffer Overflow? No. Is it basically the same show? Kinda. Buffer Overflow was originally me and Chris Hayner exploring tech topics and explaining why they were all terrible. At some point we added two more cohosts, Kim and Brenda, to provide flexibility for scheduling and content. Then I left the company and got busy with other things. We shuttered Buffer Overflow in April 2021, and that was that.
|
||
By September of 2021, Chris and I had started talking about maybe resurrecting the podcast with just the two of us for simplicity. We had a clear vision of what we wanted the show to be and where we wanted to take it. But I was busy with other projects and Chris was still working full time for our shared previous employer.
|
||
In 2022, circumstances changed. I stopped doing The Daily Check-In. Chris quit his job and decided to become an independent consultant. And now suddenly resurrecting Buffer Overflow was a real possibility. We decided to call it Chaos Lever instead, to avoid any potential trademark issues, and did a soft launch of the podcast on March 22nd.
|
||
Since that time, we’ve been releasing weekly episodes on a consistent basis and slowly growing the subscriber base. There are no current plans to monetize the show, but I wouldn’t rule it out. Regardless, it keeps the two of us out of trouble- or at least keeps us in a manageable kind of trouble, and that’s enough justification on its own.
|
||
In 2023, I’d like to focus on growing the audience. We’re going to try out a newsletter and maybe we’ll play with video? I dunno, we both have faces for podcasting. If you head to the Chaos Lever website, you should find a form to sign up for the newsletter.
|
||
HashiCorp Academy Sometime in 2008, I attended a training class for Exchange Server 2007. Compared to previous versions of Exchange, the 2007 release was a major departure. Microsoft had completely changed the internals of Exchange, introducing multiple roles and layers, and the interaction model, shifting from a heavyweight console to a web UI and PowerShell. It’s not a stretch to say that the modern era of PowerShell was fueled by Exchange 2007 and also VMware’s PowerCLI.
|
||
What does any of this have to do with HashiCorp? Nothing directly. This was the first live training I had attended as an IT Professional and the instructor was amazing. He had deep knowledge of the product and all of the ancillary components. He kept the delivery engaging, even when discussing fairly dry and mundane concepts. And he was more than happy to field questions and give detailed answers. Somewhere in the back of my head the seed was planted, I wanted to be like this guy.
|
||
At the time I didn’t appreciate how that seed would grow and flower into a full-blown career. I was basing my career path off that of my boss, who had spent about ten years in tech before becoming the head of technology for the SMB I was now employed by. I thought that was what the future held for me. I would spend a few more years honing my technical skills and them move to a leadership role at another company. I even enrolled in an MBA program because I thought it would help me advance my career as I transitioned to management.
|
||
Long story short, since this section has gone on longer than I intended, I never went into management. I ended up in consulting for six years, and during that time I started producing courses for Pluralsight in 2017. It wasn’t classroom training being delivered live, but I was now a trainer of sorts. Clearly, my drive to create educational content for technical folks became a central pillar of my career, and led me to create Ned in the Cloud LLC in 2019. But I still wasn’t doing the live training that inspired me over a decade prior. This year, that changed.
|
||
I mentioned in last year’s post that I had produced some private training videos for River Point Technologies. They are a HashiCorp training partner and deliver live training for HashiCorp products like Terraform, Vault, and Consul. I asked them if they needed help with their live training, and the answer was a resounding and enthusiastic, “Yes!”
|
||
So I started delivering live trainings for Terraform and Vault to private companies and for HashiCorp Academy. Over the course of 2022, I have conducted 15 live training sessions via Zoom for groups of 10-20 people.
|
||
This was my first foray into live instruction, and I was more than a little nervous for the first few trainings in February and March, but I quickly found my footing. I love the interactive nature of live training. The fact that I get feedback and questions from the class is amazing! It forces me to know even more about Terraform and Vault, and also provides fodder for future videos and blog posts. I’m excited to go back and update my Pluralsight courses based on what I’ve learned as an instructor.
|
||
I plan to keep teaching live classes in 2023, and I wouldn’t be opposed to doing one or two in person. Zoom’s nice and all, but I bet it would be even better being in a room with actual humans.
|
||
Manning liveProject In 2021, I signed on to deliver a series of liveProjects for Manning publishing. This is the beginning of a trend wherein Ned commits to a large project and ends up regretting it. Manning liveProjects are a really great idea. Instead of passively reading a book and maybe walking through some pre-built labs and exercises, liveProjects have you building actual solutions from scratch based off requirements and guidance from the author. You’re presented with a goal, criteria for success, and access to some of Manning’s books for reference. If you get stumped, there are partial and full solutions available, but the goal is to think things through and try to design a solution yourself. To be clear, I love this idea.
|
||
I love this idea for the learner. For the author, it was a difficult process, and I signed on to do not just one liveProject, but five of them as part of a series. The topic, that of building and managing an AKS cluster with Terraform, wasn’t even mine. Someone else had written the treatment for the projects, and couldn’t commit to the execution. Manning found me through a cursory Google search and approached me to execute. I could have said no. I probably should have said no. But instead, I said yes.
|
||
The result was that I had to learn the production process of a liveProject, work with the tooling provided by Manning, design and write the code for five different projects, and manage my existing workload. What should have take me a few months, turned into nine months of slogging through work that I wasn’t excited to do.
|
||
This is not the fault of Manning. While I don’t particularly care for their tooling to create the project, they are nothing if not thorough in their process. Once the projects were in a rough draft, they had to be reviewed, alpha tested, updated, and then beta tested. I had to hold live feedback sessions with folks from the alpha and beta groups, and then incorporate the feedback into the final version of the project.
|
||
Having this giant project hanging over my head constantly weighed on me, yet it was difficult to motivate myself to finish it. As someone who occasionally struggles with procrastination, I had to find something I wanted to do less (which I did, more on that in the next section), and motivate myself to finish the liveProjects by avoiding that other thing.
|
||
The good news is that I did finish the five project series and despite my hemming and hawing, the final product is really good. You get to design, build, and manage an AKS cluster using Terraform and Azure DevOps. If you’re working as a platform engineer or on a DevOps team managing infrastructure, then I think you’d get a lot out of the project series.
|
||
Am I proud of what I built? Yes. Would I do it again? No. The level of effort and stress from building the project was not worth the financial benefit I’m earning from creating it.
|
||
What was that other project I was avoiding? Ah, yes, it was a series of Pluralsight courses.
|
||
Pluralsight Early on in 2022, I was asked by the team at Pluralsight to redo the existing learning path for the Azure Virtual Desktop specialty certification (AZ-140). The path is a set of six courses, recorded back when the product was still called Windows Virtual Desktop. A new version of the exam objectives had dropped, the product name had changed, and the original author was no longer creating courses. So they turned to me, and for the second time in 2022, I committed to a large project and immediately began to regret it.
|
||
Azure Virtual Desktop Hell I have spent the entire year working on this learning path on and off. My original goal was to finish by September, which was a cute and totally unrealistic goal. September went sailing by and now I find myself at the end of December and not a single course has been published.
|
||
The original idea was that we would publish all the updated courses together as a cohesive whole. I thought this would be the best experience for the learner, and that it would motivate me to finish quicker. It did not. Although, I was making progress! I was almost done the fourth course in late November, when I discovered that Microsoft had updated the exam objectives once again. And this wasn’t a set of minor tweaks and additions. They had completely overhauled the objective names and goals.
|
||
After a few moments of absolute panic and a bout of desolation, I did a comparison of the old and new objectives, mapped them to the courses I had already created, and tried to figure out how much I would need to change. The result was more than nothing, but also not a complete overhaul. Yes, I was going to have to rewrite and record some portions of the completed courses, but it was not a full redo.
|
||
So in December, I began the arduous process of patching the existing courses with the updated objectives. At this point, the first two courses are done and rather than waiting for the full path to be complete, we are going to publish these in early January.
|
||
Turns Out It’s Fine One major problem with choosing to wait on publishing is that I wasn’t earning anything for the courses I completed. They were just sitting there idle and not generating the passive income they were supposed to. On one hand, it’s nice that I didn’t publish them and then have to go back and patch them almost immediately. On the other, it would have been nice to have that revenue coming in!
|
||
In previous years, I have published six or seven new courses every year. In 2022, I published one. ONE. I wrote and produced five, but only published one, Getting Started with Terraform Cloud. I don’t feel great about that. You better believe I’m going to bang out more courses in 2023.
|
||
Despite not publishing a lot of courses in 2022, my numbers for the year were extremely good. The main takeaway is that people still care a lot about Terraform! Although the year is not quite over, I have racked up 117k hours of view time for my courses, with the standout being Terraform - Getting Started at 58k hours alone. It’s unlikely I’ll post the same numbers for 2023, as 12 of my cloud related courses are being retired. To keep up my active course count, I’m going to need to be a bit more prolific in 2023.
|
||
Alternatively, I’m thinking perhaps I should focus on updating my existing courses that are top performers and look for something new that’s in demand. Both WebAssembly and Open Policy Agent appear to be ascendant, so I may do some courses for both topics from an operational standpoint. I also have to finish the Azure Virtual Desktop courses and get that off my plate. Look for the path to be complete by the end of February (I hope!).
|
||
Books I did not write or contribute to any books in 2022. This was a good decision! I do need to update my Terraform and Vault certification guides for 2023. There’s a new version of both exams coming, in particular Terraform Associate, so this is the perfect time to crank through the update process.
|
||
Other Appearances As usual, in 2022 I did participate in a number of events. Here’s what I could find:
|
||
Cloud Field Day 14 - Chris and I attended this together. That was… an idea. Intel Tech Field Day Showcase - This was mostly about CXL for me, and a reminder that hardware still matters in a cloud-centric world HYCU Tech Field Day Special Event - I’m not usually excited about data protection products, but HYCU really does seem to be on to something. HashiConf Global 2022 - I got to speak at HashiConf Global twice! Only one was recorded, so here it is. HashiTalks 2022 - I also participated in HashiTalks 2022. And before you ask, I’ll be doing HashiTalks 2023 too! ActualTech Media Keynotes - Just like the previous year, I created keynote videos for about ten ActualTech Media events. These video essays are super fun to do, and I should really post them to my YouTube channel. Add that to the 2023 pile. There may have been some other things. Honestly, I didn’t keep good track of my appearances this year and that’s on me.
|
||
Deprecations and Failures Another year full of learnings and missteps. Here’s what didn’t work out so well in 2022:
|
||
Learning Rust - Oh look, I started trying to learn a programming language and then gave up after a couple days. Wonder if there’s a pattern here? Finish the AVD Learning Path - This was supposed to be done in September. It is still not done. Two videos a month - I completely failed to be consistent at publishing videos. Not sure if it mattered. The Daily Check-In - Three podcasts is too many, and I’m more interested in Chaos Lever. Looking forward Taking in the whole of 2022, there were some personal high points and a lot of professional frustration. Since I’m my own boss, I have no one to blame but myself. My desire to avoid conflict and appease others got in the way of choosing what was realistic and healthy. As much as I know that I have to say “No” more often, it’s difficult each time I’m asked to be part of something. I like joining. I like being part of something. I fear missing out. I suspect this is a struggle I’ll continue to deal with for the rest of my life. This year I lost the battle, hopefully next year will be better.
|
||
…
|
||
Hrm, that was a negative note to ride out on. Realistically, 2022 was a mixed bag, with some powerful highs and low lows. At the end of it all I still have a healthy business, an amazing family, and a pretty sweet life.
|
||
Plans for 2023 will be a separate post, as this has gone on for far too long. Thanks for reading! I hope 2022 was a good year for you and that 2023 will be even better!
|
||
`,summary:"As 2022 winds to a close, it’s time to reflect back on the year. There were some triumphs, some failures, and a fair bit of frustration as I continue to try and figure out what Ned in the Cloud LLC should focus on and what I should avoid. If there’s one thing to take away from 2022, it’s that in 2023 I should endeavor to do less. This is not exactly a new problem for me and I think I’m handling my overall workload better now than in previous years.",date:"30 Dec, 2022",url:"https://nedinthecloud.com/2022/12/30/2022-year-in-review/",image:"2022-Year-in-Review.png",readingTime:"17"},"https://nedinthecloud.com/2022/12/16/aws-reinvent-2022-is-aws-ok/":{title:"AWS re:Invent 2022 - is AWS OK?",tags:["aws","reinvent"],content:`As I mentioned in this episode of Chaos Lever, re:Invent 2022 happened. And it was fine? Fine, yes. This is the 11th re:Invent and while it’s not the largest ever, attendance has certainly recovered from the pandemic. Depending on who you ask there was somewhere between 60k and an infinite number of people at the conference. The high end estimate came from an AI trained in crypto finance, so I wouldn’t trust anything it has to say relating to reasonable numbers.
|
||
There were also plenty of people attending virtually or more accurately, sporadically viewing remotely, in some cases in Las Vegas itself. As much as AWS tries to make getting around easy between the four major locations, it’s still a slog to get from the MGM to the Venetian, and sometimes it’s just easier (and more comfortable) to watch from a screen a few miles away from the actual event.
|
||
I was there for a couple days only, less than that in fact. And I focused my efforts on walking the expo hall and chatting with vendors, especially in start-up alley. Eschewing the larger vendor booths in favor of constructions that could actually be called booths and not encampments for a prolonged siege.
|
||
Honestly, I’m not sure what one has to gain from wandering into the dense forest of screens and marketing effluvia that populates a Diamond Sponsor- oh, excuse me, Emerald Sponsor (Emerald, like all the extra green you have to pay over Diamond) booth. For example, VMware. No offense meant, but why would I go there? I know who the company is, I know what they do, and they have battalions of marketing people who already keep me abreast of the company’s latest escapades via multiple mediums and formats. Unless I recognize one of the poor souls assigned booth duty, I am unlikely to go near any booth that takes up more than a few square feet.
|
||
No, I submit to you the actual vendors of interest can be found in the periphery. And they are desperate for human contact after being ignored by swaths of people. I had excellent conversations with XoSphere, Aiven, HYCU, Structura, and Retool.
|
||
But what about AWS itself? What was the big news? Any crazy new services or amazing improvements? Despite well-dressed people talking at Keynotes for many, many hours, there doesn’t seem to be one blockbuster thing people are buzzing about. If I peruse the tech news over the last two weeks, it is as if re:Invent barely happened. The tech zeitgeist is squarely focused on the nightmare of Twitter, the continuing disaster of crypto, and the rise of OpenAI’s weird stuff that will eventually destroy multiple industries and people. Oh, and a new version of Kubernetes (1.26) is out. And maybe that’s the biggest point of all. Open source platforms vs. closed-source primitives.
|
||
Primitives versus Platforms AWS and the public cloud in general have moved into the boring infrastructure category and out of the innovation seat, and no wonder. Cloud computing, depending on how you measure it, is now about 15 years old. While it has enabled incredible transformation of industries and disrupted traditional companies in countless ways, cloud computing and the big public cloud providers have succeeded in becoming staid and boring. That’s not necessarily bad. One thing that the cloud providers did create was self-service access to basic primitives, like IaaS, for companies to build new solutions on. In fact, that has long been AWS preferred mode of operations, creating primitives that developers can combine in new and interesting ways.
|
||
The problem with primitives, is that there can be too many of them and too many ways to combine them. And sometimes developers and infrastructure people appreciate a helping hand. AWS has something ludicrous like 300+ services. If we assume that each of those services represents a handful of primitives, that is an infinite number of possible combinations to get right, or horribly wrong. One of the things I’ve always felt about AWS is that it was like a set of tinker toys without an instruction manual. You can build anything you want, and they’re not going to tell you how. That can be overwhelming to the cloud architect and developer alike.
|
||
Consider the developer experience of 10 years ago to what they need to contend with now. It’s not just choosing which programming language to use, but also which framework to use, and if you’re using Javascript today like most folks, there are about a bazillion last I checked. Next you need to deploy that framework in a target environment. Do you use EC2 instances, containers on ECS or EKS, try to package things as Lambda or run them in Fargate? How do you handle your database backend, are you going to use Postgres or something Postgres compatible? And if so, do you go with a proprietary cloud DB, or an open source DB hosted as a platform, or host it yourself in that EKS cluster? Oh, and you’re going to need ingress controllers, or load balancers, or API gateways. And security! Don’t forget about security! It’s enough to make a dev want to throw their hands up and go back to deploying ASP.Net apps on IIS with a basic MSSQL backend.
|
||
AWS has always excelled at providing primitives that enable programmers to do whatever they want. They are looking for builders! Unfortunately, developers and infrastructure folks alike are overwhelmed with choice. The obvious answer is to conjure up a platform based on those primitives, providing an opinionated environment for application development and deployment. AWS has tried that in the past with things like Elastic Beanstalk, LightSail, and App Runner, but even then the seams stitching together the underlying primitives were showing. EBS in particular was an imperfect, leaky abstraction, and AWS chose to do that by design. It’s like they knew they had to build a platform, but they didn’t want to seem heavy handed or reduce customer choice, and in the end it’s a half-baked mess.
|
||
Platform Innovation Where some see a failure, others see opportunity! And there are several vendors today that specialize in providing a platform for app deployment that lacks choice deliberately. The idea is to get the code deployed as quickly as possible, without the developer needing to know too much about the underlying primitives; ideally nothing. Fly, Fastly, Vercel, and Fermyon are all laser focused on getting you from idea to code to running app in the least amount of time possible. Are there tradeoffs and lock-in? Absolutely. But that’s comparatively cheap compared to the hourly cost of engineers who are spending entirely too much time fiddling with primitives and doing undifferentiated heavy lifting on AWS.
|
||
Many of the platform companies are actually building their solutions on AWS, Fermyon for instance has started out with a fleet of EC2 instances using custom AMIs. AWS isn’t necessarily losing out on revenue as a result of the new platforms, but the innovation is from the new platforms companies using AWS, not from AWS itself. Again, this is fine and to a certain degree expected.
|
||
VMware Voyages I’d compare the journey of public cloud to that of another game changer in the industry, VMware. In its halcyon days, VMware was innovating on what a virtual machine could be and how it might be managed. It’s hard to overstate how groundbreaking vMotion and Storage vMotion were. How DRS could rebalance your hypervisor while you slept soundly at night. How developers could get a development environment spun up in days instead of weeks. VMware sure had some sweet primitives to build on.
|
||
VMware also tried to crack the platform nut when it comes to application development. Their acquisitions of Spring with their framework and later the Pivotal platform, were supposed to give them platforms that developers would embrace, trading complexity and choice for simplicity and speed. While Spring has its adherents and there are certainly fans of Pivotal, as a whole the introduction of containerization and Kubernetes threw a bit of a wrench in VMware’s plans for complete global dominance. In short, they backed the wrong horse, and were slow to get on the right one with Tanzu. Now that we have Broadcom looming in the distance, VMware is unlikely to become a dominant application platform, forever existing at the infrastructure layer.
|
||
That’s a profitable place to be, but it doesn’t drive growth in the way investors have come to expect.
|
||
Changing of the Guard The second major thing to bear in mind about AWS is the change of leadership. Andy Jassy has made the transition to CEO of Amazon proper and handed the mantle of AWS over to Adam Selipsky. And while they might seem like two completely interchangeable, middle-aged, white guys, their presentation style and management style are remarkably different. That difference will be reflected in the way AWS is run going forward. While there is little chance major changes are coming to the organization structure or principles espoused by leadership, there is a certain calcification settling in at AWS. Does that mean their position of cloud dominance could be in danger?
|
||
It’s tempting to equate AWS’ position to that of Microsoft in the early 2000s. Dominant at the desktop and server OS levels, Microsoft was blindsided by the one two punch of mobile and Linux. Web 2.0 is built on the back of open source technologies and not Windows Server with its bloated stack and aggressively expensive licensing. Web 2.0 is also built on browsers, something that Microsoft one had dominance in, but was edged out by another company that had embraced open source, good ole Google with their Chrome browser. At the same time, Microsoft failed to get mobile right not just once, or twice, or even three times. I think I can count at least 5 separate failures of Microsoft to break into the mobile ecosystem. Look up the Kin if you’d like a good laugh. They’ve finally thrown in the towel and gone with Android, another Open source based product, and even then the Surface Duo is… not great.
|
||
Many analysts assumed at the time Microsoft was on the road to become a has-been by the mid-2010s, based on their trajectory and lack of success in so many areas. It took some hard choices and big shake-ups at the company to turn things around and recapture market share in a completely new way. Services instead of licenses, and platforms instead of primitives.
|
||
Whence Microsoft Goes Do I see AWS falling into a similar spiral in the next few years? Not really, like I said, it’s tempting to compare the two, but I don’t think it makes sense. Microsoft’s problems were a misunderstanding of the future and some serious negativity from developers and IT folks. I believe AWS understands the market it operates in and will continue to grow in both revenue and profit for many years to come. Providing great primitives is useful to companies that want to build a platform, and AWS tends to be sticky for platform companies that don’t care to build for multi-cloud or host their own datacenter. We can have a conversation another time about cloud repatriation and the economic break even point of self hosting.
|
||
Directionally, AWS is fine. And they are focusing on what I think the next big frontier is, whether we like it or not, that of machine learning and AI. The data and analytics area of the Expo was booming and it also received its own keynote. I expect AWS will continue to build primitives and host platforms for other vendors, but I wouldn’t expect any major new services to be launched in the next few years. AWS is no longer re:Inventing cloud; they are providing the space for others to do so. The biggest threat to AWS’ dominance is not other cloud providers, it’s a better set of multicloud primitives offered at lower cost.
|
||
`,summary:"As I mentioned in this episode of Chaos Lever, re:Invent 2022 happened. And it was fine? Fine, yes. This is the 11th re:Invent and while it’s not the largest ever, attendance has certainly recovered from the pandemic. Depending on who you ask there was somewhere between 60k and an infinite number of people at the conference. The high end estimate came from an AI trained in crypto finance, so I wouldn’t trust anything it has to say relating to reasonable numbers.",date:"16 Dec, 2022",url:"https://nedinthecloud.com/2022/12/16/aws-reinvent-2022-is-aws-ok/",image:"AWS-reInvent-2022.png",readingTime:"10"},"https://nedinthecloud.com/2022/10/24/supercloud-is-neither-super-nor-a-cloud-discuss./":{title:"Supercloud is neither super nor a cloud, discuss.",tags:["supercloud"],content:`Supercloud - It’s a thing now, so I guess we have to address it.
|
||
Crap. I mean yay? In the not too long ago, Chris and I discussed the Supercloud briefly on Chaos Lever as part of a larger discussion of hybrid, on-premises, and other things. My original take was, this is a stupid term someone invented so they could own and promote it. The term adds nothing substantive to the discussion on the future of technology, conflates multiple terms we already lack a solid definition for, and adds a fluffy confusing buzzword to the marketing hell-scape vendors will glom onto and abuse until the term is more meaningless than it started.
|
||
And yet it persists.
|
||
Since we can’t simply nuke it from orbit, I suppose we must engage head-on with a robust examination on the origin of the term; the evolving use of the term; a working definition that is entirely too long; and the popularity of the term now and in the future. But first a word from our sponsor PedanTree:
|
||
“Tired of being an easy going person people enjoy talking to at social gatherings? Exhausted by the constant bother of interfacing with fellow human beings? Sick of being pleasant and reasonable at company retreats? Try pedantry, a product of Misanthropy Inc. With pedantry, you can annoy even the most well intentioned person by correcting their grammar. You can lose and alienate loved ones by picking apart their sentence structure. You could even be ostracized by critiquing your friend’s pronunciation. Yes pedantry is here to make your life as lonely and miserable as possible. Sign up now and we’ll send you a hefty supply of grammar nazi absolutely free!”
|
||
I’ve reached a point in my life and career where I’ve recognized that being pedantic and exacting about language is not always the best solution. Are there situations that call for using your lexicon with exacting precision? Certainly. Is that situation a casual conversation with a colleague or friend? No. Words mean things, but what they mean is constantly evolving and changing based on how humans use language. There’s no right way, except what we’ve collectively decided is right. Remember, language isn’t physics or mathematics. There’s no natural law or inherent rules, just humans being silly and trying (badly) to communicate with each other.
|
||
For the most part I’ve stopped correcting people on the use of on-premises vs. on premise. I don’t really care whether you’ve capitalized VMware properly (and neither does OneNote BTW). And I’ve just accepted that AI and ML, like literally and figuratively, are going to be used interchangeably, and in most situations that’s fine. It’s FINE. The point of words is to communicate meaning effectively. If your words are doing that, then we can all step back from the bully pulpit.
|
||
We invent terms when we encounter a person, place, or thing that we don’t already have a term for. That could be an existing idea or product that is expressed in a unique way, where the previous terminology no longer applies. Or it could be a wholly new thing that requires the invention of a term. Machine Learning, for instance, was something that didn’t exist until the mid 20th century, and so a new term had to be invented to describe what was happening. New terms are not bad, per se, they are a much needed component of any language. The best new terms are intuitive, provide clarity, and have an established definition we can all agree on.
|
||
And yet Supercloud really chaps my ass. Part of this post is attempting to define Supercloud and how it’s used, but the other part is a personal exploration of why this term bothers me SO MUCH.
|
||
How it got started I’d like to first acknowledge that the term supercloud has been around since at least 2016, and if you Google/DDG it you will get a website from Cornell that is an actual software project to allow for application migration between clouds. The second hit is to an MIT webpage about a project they called supercloud that was meant to enhance collaboration between MIT Lincoln labs, students, and faculty. My point is that the term is not new, but we can trace the origin of its use in the current context back to a post on the SiliconAngle website called The rise of the supercloud, which is not written in titlecase just to poke the pedant in me a little more.
|
||
If you’re not familiar with SiliconAngle (and hey if they can’t capitalize things right, then I can’t be bothered with their name’s capitalization either), they are an industry analysis and news firm probably best known for theCUBE (fine I’ll do it, but I won’t enjoy it) which is usually recorded live at various tech conferences, with hosts like Dave Vallente and John Furrier chatting with vendors, other analysts, or themselves about what’s going on in the industry.
|
||
The folks at siliconANGLE (see, I said I’d do it) have a ton of industry experience, and this post is not meant to impugn on their credentials or disparage their personal brands. They are kind people who I’ve only had positive experiences with. Nevertheless, Supercloud.
|
||
As someone who has done some work as an analyst, your primary job is to make shit up with the intent of promoting your brand or firm. For example, consider the Magic Quadrant from Gartner. You know it, you “love” it? You might even think there’s some hard science behind it. I assure you, there is not. For all the talk of qualitative vs. quantitative measurement of vendor products, most of it is a gut feeling and a guess. Analysts have to take multiple briefings a week from vendors, and they absolutely do not have time to test drive any of these solutions. In most cases, they try to apply a skeptical eye on what each vendor tells them, but not too skeptical because they need to stay in the good graces of the vendor if they want future briefings. So they take the briefing, maybe read some of the white papers, peruse the website and then do a write up.
|
||
Analyst firms make money by selling their analysis to companies who don’t want to or can’t do the research on their own. They also make money by doing research for the same vendors they are supposed to be impartially judging. Look, it’s a weird industry, and I am not accusing anyone of unethical behavior. You can maintain your morals and still get paid, but it still introduces the potential for pay-to-play and other below board behavior.
|
||
Jesus, that was quite a tangent. Why did I even get into all of that? Oh, right, the Supercloud. Analyst firms need you, the customer, to trust what they are saying. They need to be thought leaders. They need to be on the cutting edge of what’s happening in the industry. And there’s no better way to do that than making up a new term and convincing everyone else it’s useful and relevant. Gartner has its Magic Quadrant, Forrester has The Wave, and I guess siliconANGLE has rolled out the Supercloud.
|
||
Terms and Conditions Apply Why Supercloud? Did we not already have enough cloudy terms? Let me give a few examples.
|
||
Hybrid cloud - Running applications in your datacenter and in a public cloud Multicloud - Running applications in multiple clouds Private cloud - Running applications in non-public cloud Public cloud - Running applications on a public-cloud (multitenant) environment Community Cloud - Running applications in a cloud that has really strict rules about who can use them Fuuuuuuuu Cloud - How I feel every morning.
|
||
We’ve got a lot of terms that define ways to deploy and manage applications. All of them seem to have this cloud term attached. What’s that all about? The term cloud has been used into oblivion, but personally I tend to think of it as an operational model more than anything else. It’s about automation, elasticity, self-service, and measuring consumption. You know, that old NIST definition ain’t so bad.
|
||
Supercloud as a term came to exist because siliconANGLE analyst John Furrier was trying to invent a term to describe a trend he and others had observed in the tech landscape. Simply put, there were a bunch of technology vendors building new platforms based on the primitives of the public cloud vendors (aka AWS, Azure, GCP) and offering their higher level abstraction to customers as an alternative to building it yourself. If that sounds suspiciously like SaaS (Software as a Service), you’re not wrong.
|
||
I, too, had been noticing the trend of what I would call cloud-native companies and products that aren’t SaaS in the traditional sense, but are using public cloud primitives to build out their platform. When I say platform, I am talking about a technology or service another company would use to build a product or application leveraged by end users. In the same way that Oracle has a database platform used by thousands of companies to build their applications, Supercloud offerings would do the same but with the added benefit of offering it as a service.
|
||
The cognitive dissonance here lies in the difference between a platform and an application. And I’ll be the first to admit that it is a tenuous distinction that harkens back to the old defense of “I know it when I see it.” Maybe some examples are in order. Snowflake has built a data warehouse and analytics platform that sits astride the public cloud providers. Your average end user would not interface directly with Snowflake. Instead they might use an application that leverages the Snowflake platform, but that is abstracted or hidden from the consumer. Another example is Aviatrix, who is building a multicloud networking abstraction platform. You might use it as the transport platform for your applications, but it would never be customer facing.
|
||
Traditional SaaS, on the other hand, is customer facing. Canva, Office 365, Netsuite, etc. All of these offerings are consumed by non-technology professionals to get their work done. In some cases, solutions like Salesforce.com have also drifted into the platform territory, as I said the line in the distinction is tenuous. You could make the same argument about PowerApps on Microsoft 365 and I would concede the point. As Einstein would say, “It’s all relative, now leave me alone I have a grand unified theory to work on. Also, I’m dead.”
|
||
Work, work, work, work There is a working group definition of Supercloud you can check out which attempts to put some formal constraints around what the term Supercloud is intended to mean. They lay out three key characteristics or essential properties that a Supercloud offering would have to include: Run a set of services across more than one cloud (maybe), purpose-built super-paas layer (a new term! fun.), metadata intelligence.
|
||
The first one is fairly intuitive. The second one includes a new made up term called SuperPaaS, which no. Just no. Stop it. The term is trying to describe a platform that spans multiple clouds and hides the underlying cloud primitives. That’s what a platform is. Platforms abstract the underlying components, whether it’s a set of cloud providers or virtualization hosts. We already have a word. We don’t need a new one. If you want to be more precise, add a descriptor like cross-cloud or multicloud.
|
||
The third essential property simply means that the Supercloud offering is aware of its underlying components and can make informed decisions based on regionality, cost, etc. That’s um. That’s still just a platform, friend. Any platform that was unaware of its underlying components would be a pretty terrible platform.
|
||
There are also three deployment models: single cloud instantiation (which seems to contradict the first essential property), multi-cloud/multi-region, and global instantiation. No mention of transubstantiation. Weird.
|
||
Trying to pick those deployment models apart, the key here is control plane versus data services. The deployment model is mostly referring to how the control plane is deployed. As the consumer, I shouldn’t have to care. Give me the correct number of 9’s and everything else is an implementation detail. Does Office 365 tell me where the control plane is running? No. Do I care? No.
|
||
That’s it. That’s the big definition. Supercloud is a distributed platform used to build applications from abstracted cloud primitives. It’s a platform for building applications. It’s a platform… ever say a word so many times it starts to lose meaning? Platform. Orange. Toy boat. Toy boat. Toy boat.
|
||
Super is as super does Does this new term need to exist? No.
|
||
Sweet, post over.
|
||
JK!
|
||
In the working definition document, they address my question in the FAQ, and I thought I would quote from it:
|
||
“There is broad agreement that clouds in the 2020s are different from clouds of the 2010s. That something new is happening within the AWS (and other) cloud ecosystems, beyond traditional IaaS and PaaS and isn’t just SaaS running in the cloud. Supercloud is an attempt to describe a new architecture that integrates infrastructure, unique platform attributes and software to solve specific problems that public cloud vendors aren’t directly addressing. Supercloud is an evocative term that catalyzes debate, conversation and thought.”
|
||
It’s a term for stirring the pot. It’s a term that elevates the profile of those who invented it. It’s a term that can be trademarked and used for marketing purposes if desired. It’s a term that is both useful for debate and useless for anything practical. It’s the “Change my mind” Stephen Crowder meme, worse yet, it might even be the Ben Shapiro of cloud terms. Too far? Yeah, that might be too far.
|
||
And shit, it’s working! I’m talking about it. Supercloud was discussed in an IT Roundtable at Cloud Field Day 15. I’ve heard it mentioned on podcasts, in articles, and blog posts. Most of them are dismissive of the term at best. But they are engaging. Does that make me culpable? Am I part of the problem? (Sigh). I guess we’re all pundits now.
|
||
If you’re an analyst or a marketing firm, you can use Supercloud to imply or imbue some mystical properties on a product or service. Leverage the Supercloud with PendaTree from Misanthropy Inc!
|
||
But is it useful to practitioners? Probably not. If I am an application architect and I need a data platform, I can evaluate Snowflake, AWS Aurora/RedShift, and Oracle and decide which one meets my requirements without throwing some weird terminology into the mix. If I’m working on the platform team in an enterprise, I might be interested in a supercloud-type product to integrate into the platform I offer to the product teams in my org. But again, the term doesn’t help me in any meaningful way.
|
||
Does it group similar offerings into a category where they can be compared fairly on their merits? Not really either. What does NetApp vs. Snowflake vs. VMware have to do with each other? Supercloud is grouping together disparate product offerings and rolling them into a useless category for comparison and evaluation. I think it fails as an analytical term too.
|
||
At the end of the day, Supercloud is a marketing term with a touch of self-promotion from siliconANGLE. It doesn’t describe a new thing, it doesn’t provide clarity in our communications, and it has no utility for practitioners. And I’m happy to never mention it again.
|
||
`,summary:`Supercloud - It’s a thing now, so I guess we have to address it.
|
||
Crap. I mean yay? In the not too long ago, Chris and I discussed the Supercloud briefly on Chaos Lever as part of a larger discussion of hybrid, on-premises, and other things. My original take was, this is a stupid term someone invented so they could own and promote it. The term adds nothing substantive to the discussion on the future of technology, conflates multiple terms we already lack a solid definition for, and adds a fluffy confusing buzzword to the marketing hell-scape vendors will glom onto and abuse until the term is more meaningless than it started.`,date:"24 Oct, 2022",url:"https://nedinthecloud.com/2022/10/24/supercloud-is-neither-super-nor-a-cloud-discuss./",image:"Supercloud.png",readingTime:"13"},"https://nedinthecloud.com/2022/10/19/my-hashiconf-global-2022-experience/":{title:"My HashiConf Global 2022 Experience",tags:["hashiconf-global","hashicorp"],content:`This post was sponsored by HashiCorp, see the bottom for more details
|
||
Before we get into the announcements and the event, I want to pause for a moment and just acknowledge how amazing everyone was at the conference. Before the keynote each day, the HashiCorp MCs explained the code of conduct for the event, ending with above all be kind. Without exception, the people I met at the conference were overwhelmingly kind, generous with their time, and excited to engage with fellow attendees. I’ve been to plenty of conferences in the past where I felt awkward; unsure of where to be and who to talk to. At no point did I feel this way at HashiConf. I couldn’t walk more than ten feet of the expo floor without landing in a conversation with friends both old and new. The HashiConf attendees and the larger HashiCorp community are incredibly warm, inviting, and above all kind.
|
||
Big Trends Alright, alright, alright. That’s enough of the mushy stuff, what about the tech? I won’t exhaust you by listing out each and every update, instead I’d like to focus on a couple key trends of the conference. Namely, the increasing focus on HashiCorp Cloud Platform (HCP) and the evolving role of the platform team.
|
||
Over the last few years, HashiCorp has rolled out SaaS offerings around their product portfolio starting with Terraform Cloud and expanding to HCP Consul, Vault, and Packer. Each offering is still based on the open-source version of the offering, but helps to offload some of the administrative burden involved in running the solution in production. In particular, both Consul and Vault require a significant amount of effort to deploy, upgrade, and maintain the solution, which represents undifferentiated heavy lifting for your operations team. HCP aims to manage those tasks for you, freeing you and your team up to leverage the features of each solution without having to worry about its care and feeding. You know, Software as a Service!
|
||
The keynote announcements continued this trend with the release of HCP Vault on Azure in Beta, HCP Waypoint in Beta, and HCP Boundary in GA. I’m really excited to see HCP Vault coming to Azure and I’m curious to see what new DR and Performance topologies are supported as the service becomes generally available. I was surprised and more than a little chuffed to see HCP Boundary go GA so swiftly. There’s really no reason for you to manage the control plane of Boundary yourself, so why not let HashiCorp do that heavy lifting for you? For a product that is only a couple years old, Boundary has come sprinting out the gate and I’m excited to see it approach a 1.0 release.
|
||
The other big idea is around the platform team as an operating model. Based on the HashiCorp State of Cloud Strategy report, a majority of organizations are using multicloud and building a platform team to support their DevOps teams by providing a stable platform to deploy on, along with guidance and templates that enable developer self-service. During the Keynote, Armon made sure to point out this emerging trend, but I don’t think he made it completely clear how platform teams and multicloud are going to impact tool selection.
|
||
The reality of hybrid and multicloud means most platform teams are going to be supporting more than one cloud provider. As part of their tool selection and workflow design, those platform teams are going to be on the lookout for solutions that support a consistent look and feel across multiple cloud providers. We also have the tightening of economic purse strings across the world, leading to a slowdown in hiring and a push for automation - aka doing more with less. I’ve noticed these trends on my own, as have several other analysts and industry pundits. It was gratifying to see during the keynote that HashiCorp also understands the problem space and they are focusing on developing features in their products to assist the platform team with multicloud challenges.
|
||
Hallway Track In my post leading up to HashiConf Global, I said that I was going to be spending my time watching the hallway sessions and hanging with people in the hallway track. And I did exactly that. Unfortunately the hallway sessions were not recorded, so I can’t share them with you. But I did have some favorites, and I’d like to point you to the folks behind those sessions.
|
||
Bryan Krausen gave a great talk about what he’s learned about Vault from over five years of consulting. Turns out a lot of folks are using the root token for everyday operations. Stop that! Definitely check out Bryan’s content on Udemy if you want to learn a ton about Vault. Maybe even buy the book he co-authored!
|
||
Steven Ma from Microsoft gave us an update about what’s new with Azure Terrafy. While the name hasn’t yet changed - although they are planning to do so - they released 0.8.0 of the tool just in time for the conference. A couple big highlights for me are the addition of a filter to pull in resources across an entire subscription and improved cross resource dependency resolution. They’ve also got an aggressive roadmap full of awesome features. If you have an opinion on what you’d like to see next, join the Azure Terraform community and let them know!
|
||
While I don’t know a ton about Nomad, I’ve always been curious about how it might stack up against Kubernetes in the long term. Matt Butcher from Fermyon talked about how they gave up trying to shoehorn WebAssembly programs into Kubernetes and instead chose to use Nomad with a custom driver. They have a blog post that lays things out quite nicely if you want to know more.
|
||
You might be wondering if Mitchell Hashimoto actually hung out in the hallway track as promised? I assure you that he did. Unfortunately I was on my way to deliver my breakout session when I saw him, but I know that he was hanging out by the hallway track area for a couple hours, graciously speaking to anyone who wanted to meet him and taking in the excellent energy of the conference.
|
||
And that brings me nicely back to my original point about HashiConf. From the Founders to the CEO to employees at all levels of HashiCorp, everyone I met was friendly and engaging. No one was putting on airs or acting aloof. That positive energy was palpable and extended out to the speakers and attendees. For me, that is the true value of attending a conference, not the technical sessions or the big keynote announcements. It is the opportunity to connect with the community; to meet new friends; to have interesting conversations. By that measure, HashiConf Global was a resounding success and I can’t wait for next year!
|
||
Sponsorship – This post was sponsored by HashiCorp. Although I received compensation, they did not tell me what I could/couldn’t write about. All opinions and thoughts remain mine. Big thanks to the fine folks at HashiCorp for helping me keep the lights on!
|
||
`,summary:`This post was sponsored by HashiCorp, see the bottom for more details
|
||
Before we get into the announcements and the event, I want to pause for a moment and just acknowledge how amazing everyone was at the conference. Before the keynote each day, the HashiCorp MCs explained the code of conduct for the event, ending with above all be kind. Without exception, the people I met at the conference were overwhelmingly kind, generous with their time, and excited to engage with fellow attendees.`,date:"19 Oct, 2022",url:"https://nedinthecloud.com/2022/10/19/my-hashiconf-global-2022-experience/",image:"My-HashiConf-Global-2022-Experience1.png",readingTime:"6"},"https://nedinthecloud.com/2022/09/26/using-optional-arguments-in-terraform-input-variables/":{title:"Using Optional Arguments in Terraform Input Variables",tags:["hashicorp-terraform-tutorial","terraform"],content:`Well hot damn! Terraform 1.3 has introduced an incredibly popular feature for the Terraform community: optional arguments. This is a feature that has been requested for a long time, and it’s finally here. Let’s take a look at how it works.
|
||
Type Constraints Before we dig into optional arguments, it’s probably important to address type constraints, you know the place where you would use them. For the casual Terraform user, you probably haven’t written a lot of type constraints for input variables, so the need for optional arguments may not be immediately obvious. Let’s take a look at an example input variable.
|
||
variable "my_string_var" {} This is all you need to define an input variable. You declare a variable block and give it a name. That’s it. But what if you want to make sure that the value passed in is a string? You can use a type constraint to do that.
|
||
variable "my_string_var" { type = string } Easy peasy. Now if you pass in a number, you’ll get an error. But what if you want to make the variable optional? Every input variable needs a value at run time, so you do that by providing a default value.
|
||
variable "my_string_var" { type = string default = "Hello, Taco!" } That works really well for a basic string, but what if you are dealing with a more complicated data structure? You can use a type constraint to build out a complicated object with multiple arguments. Here’s an example someone sent me recently via Twitter.
|
||
variable "global_secondary_index_map" { type = list(object({ hash_key = string global_name = string non_key_attributes = list(string) projection_type = string range_key = string read_capacity = number write_capacity = number })) } If someone wants to use that variable, they can’t leave out any of the arguments. They have to provide a value for each one.
|
||
global_secondary_index_map = [ { hash_key = "taco-id" global_name = "chicken_taco" non_key_attributes = ["peppers"] projection_type = "INCLUDE" range_key = "meat" read_capacity = 5 write_capacity = 5 } ] That’s great if you want to make sure that the user provides values for all of the arguments, but what if you want to make some of them optional? That’s where optional arguments come in.
|
||
Optional Argument Usage Optional arguments leverage the optional keyword to make arguments optional. Let’s take a look at an example.
|
||
variable "taco_object" { type = object({ meat = string cheese = optional(string, "cheddar") salsa = optional(string) }) } In this example, we are defining a variable that is an object with three arguments. The first one meat is required, but the other two are optional. The meat argument is a string, and the cheese argument is a string with a default value. The third argument (salsa) is a string with no default value.
|
||
The general syntax for an optional argument is optional(type, default). The type is the type of the argument, and the default is the default value for the argument. If you don’t provide a default value, the argument will be set to null.
|
||
I can submit the following values for the variable.
|
||
taco_object = { meat = "chicken" cheese = "jack" } And that will work just fine. Now the salsa argument has a null value, and I need to deal with that when I parse the variable. I can do that by using a conditional expression.
|
||
locals { salsa = var.taco_object.salsa != null ? var.taco_object.salsa : "mild" } Which would set the value of salsa to mild if no value is present in the taco_object.salsa argument. If I didn’t want a default value, I could leave the value set to null. When an argument is set to null for a resource or data source, it will not be included in the request. The provider will use the default value for that argument.
|
||
Conclusion The optional keyword is a great addition to Terraform. However, it’s not the sort of thing you would care about if you are just consuming Terraform. The keyword shines as you develop modules for others to consume, providing flexibility for both the user and the developer. If you found yourself using the any type constraint to make arguments optional, you can now use the optional keyword instead.
|
||
This post was partially written by GitHub Copilot. Initially I turned off Copilot for markdown files, but it actually does help flesh out a sentence quickly and comes up with code examples that are close enough to what I would write. Not bad Copilot, not bad indeed.
|
||
`,summary:`Well hot damn! Terraform 1.3 has introduced an incredibly popular feature for the Terraform community: optional arguments. This is a feature that has been requested for a long time, and it’s finally here. Let’s take a look at how it works.
|
||
Type Constraints Before we dig into optional arguments, it’s probably important to address type constraints, you know the place where you would use them. For the casual Terraform user, you probably haven’t written a lot of type constraints for input variables, so the need for optional arguments may not be immediately obvious.`,date:"26 Sep, 2022",url:"https://nedinthecloud.com/2022/09/26/using-optional-arguments-in-terraform-input-variables/",image:"Using-Optional-Arguments-in-Terraform-Input-Variables.png",readingTime:"4"},"https://nedinthecloud.com/2022/09/20/what-im-excited-about-at-hashiconf-global-2022/":{title:"What I'm Excited About At HashiConf Global 2022",tags:["hashiconf-global","hashicorp"],content:`This post was sponsored by HashiCorp, see the bottom for more details
|
||
As I prepare for HashiConf Global in a couple weeks, I thought I would share the things that I am most excited about. That includes the sessions, the hallway track, and the other hallway track. What do I mean by that? Read ahead to find out.
|
||
The Sessions Looking over the session catalog, there are a few talks that immediately jump out at me. Now, granted I am a little biased towards Terraform and Vault, but that being said I also am curious about Nomad and Boundary. I should also note that some of these sessions will not be streamed to virtual attendees; however, all talks will be accessible about a week after the conference is over.
|
||
What’s New in Terraform for Azure?: In terms of provider support, Azure has always lagged a little behind AWS in terms of features and functionality. The 3.0 release of the azurerm provider represented a huge shift in terms API parity and speed. Microsoft has also been busy developing the AzAPI provider and Azure Terrafy tool. I’m curious to see what they have in store to further improve the Terraform experience on Azure. Terraform AWS Cloud Control Provider - Under the Hood: In a similar vein to the AzAPI provider, the AWS Cloud Control API added a set of common APIs to improve the developer experience of working with the AWS APIs, which if we’re being honest were incredibly inconsistent across their portfolio of services. I’m super curious to see how the use of the AWS CCP will alter usage of the standard AWS provider. Managing Terraform Enterprise At Scale: Want to see how things break in new and unexpected ways? Do it at scale. From this presentation, I’m hoping to gain some insight into how to administer thousands of Terraform Enterprise workspaces effectively, and some best practices and patterns for managing modules, credentials, and variables. Using OIDC with HashiCorp Vault and GitHub Actions: I mean, I’ve got to plug my own talk right? I feel like the title says it all. We’re going to explore the magic of OIDC together, and sprinkle some JWT fairy dust on GitHub Actions and Vault. Nomad: Past, Present, and Future: Nomad is one of the HashiCorp products that is constantly on my list of things to take a closer look at. I’m hoping to get a sense of where things stand today for Nomad, and maybe how I can start integrating it into my existing projects. Validate Infrastructure, Detect Drift, and Enforce OPA Policies with Terraform Cloud: One tool that I am incredibly interested in knowing more about is Open Policy Agent (OPA). So much so, that I am doing a Hallway Track talk about it to force me to learn more. This session is a learning lab to give you hands-on experience around using Terraform Cloud to run OPA. Distributed Flexibility: Nomad and Vault in a post-Kubernetes World: Post-Kubernetes world? I had no idea! Once again, I am curious about Nomad as a product and this presentation promises to show a different architecture than the K8s cluster with service mesh that we’ve just started getting used to. I also see Web Assembly is involved somehow, so color me intrigued! Those are all the main track sessions, but there’s a whole other hallway track. Which brings me to the next two sections, and why you may want to attend in-person if you can.
|
||
The Hallway Track At HashiConf Global, there will be a series of smaller, hallway-style presentations. These presentations will be only 15 minutes in length, will not be recorded or streamed through the virtual conference option, and will literally be in the hallway between the main room and the session halls. The full list of talks is available on the HashiConf Global site, so I won’t go through them exhaustively, but here are a few that jump out at me:
|
||
Configuration Driven Terraform using Generic Reusable Resource Modules Secure Every Data Source Serverless Vault for Dev If you took a gander at the full listing, you may have noticed there are impromptu talks over lunch each day. Got an idea? Want to give a quick talk? Swing by around lunch and see if the slot is available!
|
||
I love these quick, informal talks. They force the presenter to get right down into the meat of the presentation and skip all the effluvia of a longer, more formal presentation. You also tend to get more esoteric or experimental talks that might not work for a longer format. And, I’m not the only one who feels this way.
|
||
I'll be at HashiConf and my plan this year is to mostly hang around the hallway track! Just happy to chat with our community, hear what you all are doing. https://t.co/QBmy02LUg0
|
||
— Mitchell Hashimoto (@mitchellh) September 16, 2022 That’s right, the co-founder of HashiCorp, Mitchell Hashimoto is planning to spend most of his time checking out the hallway presentations. If you want a chance to meet Mitchell, and some other excellent folks, the hallway track is where it’s at!
|
||
The Other Hallway Track Yes, there is an official hallway track with rapid presentations, but there’s also the less formal hallway track just like any other conference. It’s a series of organic conversations that you get to have with like-minded people. This is always my favorite part of any conference. There’s a spontaneity and serendipity to the in-person experience that we have yet to replicate in a virtual context.
|
||
Walking through the expo area, you might catch a fragment of conversation about Vault and auth methods that piques your interest. Getting lunch you may sit down at a table with folks talking about Packer for GCP or get sucked into a chat about the best way to bootstrap Nomad. While waiting to get into a talk about Boundary, you might find yourself talking to a HashiCorp engineer about how to use Terraform with Boundary’s dynamic inventory feature.
|
||
I’m sure the virtual experience will be top-notch, if HashiConf EU is any indication, but there’s still something about being there in-person that we have yet to capture with various chat and streaming tools. If you have the means and are able to, I highly recommend making the trip out to LA. You won’t be disappointed.
|
||
In fact, HashiCorp has given me a handy little discount code you can use at checkout. Just use the coupon code HASHICONFVIP2022 when you start your registration and you’ll pay $599 which is $300 off the list price!
|
||
Hopefully I’ll see you in LA! And if you find me, you just might get a fun little knick-knack to take home.
|
||
Sponsorship - This post was sponsored by HashiCorp. Although I received compensation, they did not tell me what I could/couldn’t write about. All opinions and thoughts remain mine. Big thanks to the fine folks at HashiCorp for helping me keep the lights on!
|
||
`,summary:`This post was sponsored by HashiCorp, see the bottom for more details
|
||
As I prepare for HashiConf Global in a couple weeks, I thought I would share the things that I am most excited about. That includes the sessions, the hallway track, and the other hallway track. What do I mean by that? Read ahead to find out.
|
||
The Sessions Looking over the session catalog, there are a few talks that immediately jump out at me.`,date:"20 Sep, 2022",url:"https://nedinthecloud.com/2022/09/20/what-im-excited-about-at-hashiconf-global-2022/",image:"What-Im-Excited-About-For-HashiConf-Global-2022.png",readingTime:"6"},"https://nedinthecloud.com/2022/06/29/nonsensitive-function-fails-in-terraform/":{title:"Nonsensitive function fails in Terraform",tags:["terraform"],content:`When I was trying to work with a module in Terraform, I came across an interesting issue. The module in question created an Azure AD service principal and optionally a secret for the service principal.
|
||
I wanted to print the service principal application id and secret to the screen so I could use it for testing. Naturally, the output for the service principal secret (client_secret) had been marked as sensitive, so in order to print it to the terminal window I would need to use the nonsensitive function.
|
||
output "client_secret" { value = nonsensitive(module.sp.client_secret) } But it didn’t work! I got the following error message:
|
||
│ Error: Output refers to sensitive values │ │ on main.tf line 33: │ 33: output "client_secret" { │ │ To reduce the risk of accidentally exporting sensitive data that was intended to be only internal, Terraform requires that any root module output containing │ sensitive data be explicitly marked as sensitive, to confirm your intent. │ │ If you do intend to export this data, annotate the output value as sensitive by adding the following argument: │ sensitive = true Super weird right? A quick peek into the module showed me what was happening. The creation of the service principal secret is optional. You might want to create a certificate instead. The code creating the secret looks like this, note the conditional expression for the count meta-argument:
|
||
resource "azuread_service_principal_password" "main" { count = var.enable_service_principal_certificate == false ? 1 : 0 service_principal_id = azuread_service_principal.main.object_id rotate_when_changed = { rotation = time_rotating.main.id } } The count meta-argument means the resulting data structure will be a list of azuread_service_principal_password instances, and the address of the secret value would be azuread_service_principal_password.main[0].value.
|
||
Looking at the actual output for client_secret that’s not what we see:
|
||
output "client_secret" { description = "Password for service principal." value = azuread_service_principal_password.main.*.value sensitive = true } The expression azuread_service_principal_password.main.*.value is going to return a list of the values in the value attribute for all instances of the azuread_service_principal_password.main resource.
|
||
The splat expression above is actually shorthand for this: [ for k in azuread_service_principal_password.main : k.value ]
|
||
In our case, that will be a list with a single element, the client secret. For example, if the client secret is 1234567890, the expression will evaluate to ["1234567890"]. It’s a subtle but critical difference.
|
||
By setting the output’s sensitive argument to true, the list is set as sensitive. The client secret value inside the list is already set as sensitive. When I use the nonsensitive function to remove the sensitivity marker from the list, the sensitivity marker on the value inside the list remains. And thus I got the confusing error.
|
||
To prove that was the issue, I updated the value for my output to be the following:
|
||
output "client_secret" { value = nonsensitive(module.sp.client_secret[0]) } And sure enough the error went away! The nonsensitive function was now being applied to the value inside the list and not the list as a whole.
|
||
A similar situation can arise when you’re dealing with maps. Check out this excellent post by Chad Quinlan that explains the necessary workaround.
|
||
If you’re in a similar situation, I hope this post helped you out!
|
||
`,summary:`When I was trying to work with a module in Terraform, I came across an interesting issue. The module in question created an Azure AD service principal and optionally a secret for the service principal.
|
||
I wanted to print the service principal application id and secret to the screen so I could use it for testing. Naturally, the output for the service principal secret (client_secret) had been marked as sensitive, so in order to print it to the terminal window I would need to use the nonsensitive function.`,date:"29 Jun, 2022",url:"https://nedinthecloud.com/2022/06/29/nonsensitive-function-fails-in-terraform/",image:"Nonsensitive-function-fails-in-Terraform.png",readingTime:"3"},"https://nedinthecloud.com/2022/06/08/using-oidc-authentication-with-the-azurerm-backend/":{title:"Using OIDC Authentication with the AzureRM Backend",tags:["azure","microsoft-azure-terraform"],content:`Microsoft recently announced the general availability of OIDC authentication for GitHub Actions using Azure AD. Naturally, I immediately thought of how I could use this to remove static credentials from my GitHub Actions workflows that deploy Terraform configurations. I could use a service principal and OIDC for deployment of the Terraform configuration as well as the storage of state data in an Azure Storage account.
|
||
The actual process of using OIDC with Azure AD and Terraform will be the subject of an entirely different post. In this post, I wanted to explore how to use the OIDC authentication with the azurerm backend, and some considerations depending on your state data storage structure.
|
||
TL;DR If you’re just looking for the conclusion, here it is. To scope permissions at the container level when using OIDC authentication for the azurerm backend, enable both OIDC and Azure AD authentication using the use_azuread_auth and use_oidc arguments or the ARM_USE_OIDC and ARM_USE_AZUREAD environment variables.
|
||
If you don’t, you will get an access error that the ListKeys action failed. You’re welcome.
|
||
The AzureRM Backend Authentication The azurerm backend stores state data in an Azure storage blob in a container in a storage account. There are many ways you can authenticate to the Azure storage API to execute the necessary actions:
|
||
Azure CLI Service Principal Azure AD OIDC SAS Token MSI (Managed security identity) The OIDC option was introduce in a recent version of Terraform, since the backend code is part of the core Terraform binary and not part of a provider. To use OIDC authentication, you will need to configure the azurerm backend, either by including the information in the backend block or by setting environment variables. Here is an example backend block:
|
||
terraform { backend "azurerm" { resource_group_name = "sa-rg" storage_account_name = "storageaccountname" container_name = "tfstate" key = "terraform.tfstate" use_oidc = true subscription_id = "00000000-0000-0000-0000-000000000000" tenant_id = "00000000-0000-0000-0000-000000000000" client_id = "00000000-0000-0000-0000-000000000000" } } You can replace some of the argument with environment variables:
|
||
Argument Environment Variable subscription_id ARM_SUBSCRIPTION_ID tenant_id ARM_TENANT_ID client_id ARM_CLIENT_ID use_oidc ARM_USE_OIDC The rest of the arguments can be specified at run time when you initialize Terraform using the -backend-config option for each argument.
|
||
Configuring Storage Account Permissions One question you might ask is, how do I properly configure permissions on the storage account to adhere to the principle of least privilege? And the answer - as always - is it depends (TM). To better understand why it depends, you need to know what Terraform is doing when it leverages the azurerm backend.
|
||
Assuming you’re using a configuration block similar to what you see above, Terraform will take the following actions:
|
||
Authenticate to Azure AD using OIDC and get a token Use the token to get a token from the Azure Storage API Use the Azure Storage API token to try and retrieve the access keys for the storage account Use one of the access keys to perform all subsequent operations on the storage account What I want to highlight here is that Terraform is going to retrieve the access keys for the storage account, and then use those keys to perform operations. This means two important things:
|
||
Terraform has access to do ANYTHING in that storage account b/c it has the access keys Terraform must have the ListKeys permission on the storage account to get those keys For those that are not familiar, access keys are the equivalent of having root on the storage account. Those keys can do anything on the account.
|
||
I was… not excited about Terraform needing effectively root access to the storage account to write data to a storage blob. That seems like overkill. In case you don’t believe me, you can check out the source code right here.
|
||
Let’s assume for a second you are using a separate storage account to house state data for each Terraform configuration. Well then, I suppose we don’t really have a problem. I’m still not wild about the godlike power Terraform has over the storage account, but it isn’t going to affect anything but its own configuration.
|
||
What if you wanted to use a single storage account to house state data for multiple instances of a Terraform configuration? Well now we have a possible problem. Each instance now has the permissions to overwrite or delete anything in the storage account, even if you are using different service principals for each instance. Your development environment service principal has the ability to delete production state data. That’s… bad. And you can’t restrict it with Azure IAM, because it will be using the access key, which doesn’t check Azure permissions.
|
||
What’s a sad Azure boy to do?
|
||
Azure AD Authentication If you’re using a service principal or the Azure CLI to authenticate to your azurerm backend, then you will see the behavior I am describing. But there is another way! Actually there’s two:
|
||
Shared Access Signature (SAS) Tokens - These are limited use tokens you can generate for a storage account. They can be scoped based on time, objects, source IP address, and access rights. Azure AD Authentication - Selecting Azure AD authentication uses the Azure IAM permissions associated with the service principal to determine access rights. Since we are using OIDC, there is no option to generate a SAS token, but we can use Azure AD authentication! (I think the Azure Storage API may be using SAS tokens under the covers, but that’s just speculation). In that case, instead of trying to get the access keys for the storage account, the backend simply tries to access the container and blobs in the container. Instead of assigning our service principal rights to ListKeys on the storage account, we can narrow down the scope of permissions to the container level (the lowest level you can scope a permission on Azure Storage).
|
||
You can either add the following argument to your backend block:
|
||
terraform { backend "azurerm" { use_azuread_auth = true } } Or set the environment variable ARM_USE_AZUREAD to TRUE.
|
||
It is now possible to use the same storage account for as many instances of state data you want, with each instance residing in its own storage container. Assuming you’re using a different service principal for each environment (dev, stage, prod, etc.), you can assign narrowly scoped rights to each service principal that only allows access to the state data corresponding to that environment.
|
||
According to the HashiCorp docs, you need to grant Storage Blob Data Owner permissions to the service principal when you’re using Azure AD auth. In practice, I have found that Storage Blob Data Contributor is sufficient. We don’t need our service principal to alter permissions.
|
||
To sum up, enable OIDC and Azure AD authentication on the backend, and assign the service principal the Storage Blob Data Contributor role scoped to the container that will house state data.
|
||
Additional Thoughts Two additional things of note: workspaces and SAS tokens.
|
||
Workspaces The container level is the narrowest scope you can define for Azure IAM on a storage account. You cannot set individual permissions on storage blobs. Savvy readers might take note that if you’re using workspaces, this presents a problem. When using the azurerm backend, each workspace resides in the same container, with the workspace name added to the blob name. Whatever service principal you use for each workspace will have permissions to alter the state data for all workspaces in the container. That is less that ideal.
|
||
Of course, these days I would recommend against using workspaces for most situations. There are better patterns for managing multiple environments with the same configuration, most notably using branches in your source control for each environment. Each environment can use a different storage container and the problem is solved. Be on the lookout for a blog post and video about exactly that.
|
||
SAS Tokens In my humblest of opinions, I would prefer to use a SAS token over a service principal. The SAS token gives you an even higher level of control regarding access. Unlike Azure IAM permissions, you can scope a token to a specific storage blob, meaning you can use it with workspaces and prevent one workspace from accessing the state data of another.
|
||
The SAS token supports setting permissions, start and stop time, and source IP address as additional controls on the token. Given the sensitive nature of state data, I like the idea of being able to grant limited access for a short duration of time.
|
||
The downside is that you’ll need a portion of your automation workflow that can generate a SAS token before a Terraform run starts. There’s plenty of example scripts that do exactly that, but it is another piece of your pipeline that you need to maintain.
|
||
`,summary:`Microsoft recently announced the general availability of OIDC authentication for GitHub Actions using Azure AD. Naturally, I immediately thought of how I could use this to remove static credentials from my GitHub Actions workflows that deploy Terraform configurations. I could use a service principal and OIDC for deployment of the Terraform configuration as well as the storage of state data in an Azure Storage account.
|
||
The actual process of using OIDC with Azure AD and Terraform will be the subject of an entirely different post.`,date:"8 Jun, 2022",url:"https://nedinthecloud.com/2022/06/08/using-oidc-authentication-with-the-azurerm-backend/",image:"Using-OIDC-Authentication-with-the-AzureRM-Backend.png",readingTime:"7"},"https://nedinthecloud.com/2022/05/01/diversity-starts-with-me/":{title:"Diversity Starts with Me",tags:["onalytica"],content:`A couple weeks ago I received a notification from Onalytica that I would be included in their latest Who’s Who in Cloud? report for 2022 under the category of Content Creators. I’ve been aware of Onalytica for a while now, but haven’t paid them much mind. In fact, I wasn’t entirely sure of what they do, but I knew they had put me on some list in 2020 as a cloud influencer. This seemed to be more of the same, and while I appreciate the recognition, I also have no idea if it matters in the grand scheme of things.
|
||
The report was released in the following days, and they tagged me on Twitter with a link to the report. I figured I would give it a gander, and what I saw was… how to put this diplomatically, incredibly homogenous.
|
||
The page I was featured on for content creators was composed entirely, ENTIRELY of men, and only one of color. This was their list of content creators selected based on “consistency of content creation, online activity, and relevance.” Sorry ladies, you’re not consistent or relevant. Needless to say - I am writing a blog post after all - I had some feelings about this list.
|
||
Excuses, Excuses “This is a sample list” is what it says at the top of the page, which, fair enough. Perhaps the rest of their database is brimming with BIPOC folks from all edges of the globe. Perhaps they have an amazing assortment of diverse content creators that simply didn’t make the cut. That’s not a good excuse and in some ways is actually worse. Allow me to enumerate the reasons:
|
||
If you actually have a diverse group, then someone made a decision (conscious or not) to pick this list of white dudes. If you don’t have a diverse group, then you are clearly not doing your homework. I can think of ten women off the top of my head who should be on this list before me. This list is being sent out ostensibly as a marketing tool to entice businesses to engage with Onalytica. Failing to have a single woman on the sample list sends a pretty clear message to the marketing folks at prospective customers. Okay, but it’s like one page In fairness, and balancedness, and balanced fairness, I am looking at one page in a report. What about the rest of the report? Let’s not cherry pick our data!
|
||
The report has 172 entries across eleven categories. Out of the 172 people Onalytica chose (remember this is a sample, so this is what they chose to include) 140 are men. Only 19% are women, and 10% are people of color. That’s um checks notes less than representative.
|
||
Since the report itself states that Onalytica has “an influencer database of 1M influencers across 500+ topical communities”, either the report is not representative of their database, or their database is deeply flawed. Not sure which is worse, but neither is good.
|
||
But the algorithm! At the beginning of the report Onalytica says:
|
||
Onalytica uses a unique combination of our proprietary software, human qualitative analysis and secondary research to analyze online and offline influence, to create the best social influencer lists in the world. Our influence scores are driven by 38 algorithms and this methodology has been continually refined over the past 10 years to evolve with social media developments and how influence is best calculated.
|
||
Oh goodness! That’s a lot of words. They’ve got 38 algorithms! Can’t fault the results, there are so many algorithms, how could it be wrong?
|
||
I would also like to point out, they are using human qualitative analysis. That’s fancy speak for having an actual person making decisions based on factors that would be difficult to quantify with an algorithm. So it’s not just a fancy machine spitting out ranked names. This list is also hand-curated.
|
||
Do better There’s a lot of problems with this report, but the biggest is one of representation. Put in the most blunt way possible, diversity matters and this report fails to meet the minimum standard.
|
||
I don’t want to just leave it at that. There’s some nuance here I want to elucidate.
|
||
Representation matters in two important ways. It matters to the companies receiving this report and to the people reading this report. Sounds like I said the same thing twice, but bear with me.
|
||
Industry see, industry do Representation matters to the companies receiving this report. When I first tweeted about the report, I was credulous of its importance. But within a few minutes, a fella from Dell Technologies confirmed that they use this report to find influencers. Companies are using this report and Onalytica’s service to find influencers and reach out to them.
|
||
On the anecdotal side, I’ve been picking up a lot of new Twitter followers in the last two weeks. I can’t attribute this directly to the report, but I also cannot rule it out.
|
||
We already have representation issues in tech; I don’t think anyone could reasonably argue we don’t. And those who do can go take a jump in Lake Disingenuous. I hear it’s lovely this time of year. Reports like the one from Onalytica help to reinforce the representation issue by amplifying the status quo.
|
||
It’s a classic positive feedback cycle. Onalytica discovers relevant influencers and adds them to their database. Companies search the database for the top influencers and throw money at them to increase the company’s visibility. Which in turn increases the influencer’s clout, making certain they’ll be included and promoted in the Onalytica database. If the influencers Onalytica chooses (chooses) to promote are the status quo (read: cis-het white dudes), then the status quo will not change.
|
||
There’s plenty of blame/responsibility to go around when it comes to promoting change and diversity. The companies searching for influencers can make diversity a priority, and they should. The influencers themselves can use their clout to put a spotlight on a fellow influencer who deserves more attention. And finally, Onalytica should be mindful of who they are promoting and what signal that sends to the industry.
|
||
Increasing diversity is something everyone can help with, and firms like Onalytica have a platform to help more than most.
|
||
See you, see me Representation matters to the people reading this report. The individual people looking at this report are being given a clear signal. Here are the people Onalytica and the larger industry thinks are important and noteworthy. If you aspire to be influential in the tech industry, you should look like this.
|
||
That’s great for someone who looks like me! I thumb through the pages and think, “Hey I bet I could do what all these fellas are doing. In fact, I already look the part!” I also benefit tremendously from an overabundance of confidence and some significant privilege. And would you look at that, I’m already on the list. How about them apples?
|
||
For someone who doesn’t look like me, the message is quite different. The report aggressively says, “You don’t belong in these hallowed halls. You are not and will not be influential. In fact, you shouldn’t even be in tech. Wouldn’t you be happier in marketing?” Okay, I’m being a little hyperbolic. However, when I think about all the other barriers already in place for underrepresented folks, maybe I’m not being hyperbolic. Maybe I’m just plain bolic.
|
||
The central point is that underrepresented folks will be more likely to join the tech industry and become influential if they see other people like them being supported and promoted. Having mentors and role-models that share your experience and background is tremendously helpful. I’ve never had a shortage of folks who look like me, and I’ve definitely received help and advice from them. At the very least, I’ve always felt like I belonged in tech.
|
||
I think it’s hard for anyone in my position to understand how difficult it must be for others. Luckily, some of those underrepresented folks have been candid with me about the tech industry and what a dumpster fire it can be. And I have had the good sense to listen.
|
||
Be the change you want to see Onalytica. Listen to the people who are telling you change is needed. And then do something about it to show you care.I’m going to give you the benefit of the doubt, and assume you do.
|
||
Blameless? Listen, I’m not perfect about this kind of thing. I’m not pretending to be some bastion of diversity who’s never made a misstep. Learning about the need for diversity and how to take action to promote it is an ongoing journey and I don’t expect to ever complete it. The important thing is that I’m trying and I’d like to see others try too.
|
||
I reached out to Onalytica about this issue and got back some vague platitudes about how this was a sample list and they use algorithms, yadda, yadda, yadda. Likewise, a friend of mine also reached out via email with similar concerns and also got the brush off.
|
||
While I can’t force Onalytica to change, there is something I can do. I’ve asked them to remove me from the report and not include me in their future reports until they have addressed their representation issues. If you are included in the report and feel motivated to do the same, I encourage and applaud you. They say we’re influential, so let’s use that influence!
|
||
`,summary:"A couple weeks ago I received a notification from Onalytica that I would be included in their latest Who’s Who in Cloud? report for 2022 under the category of Content Creators. I’ve been aware of Onalytica for a while now, but haven’t paid them much mind. In fact, I wasn’t entirely sure of what they do, but I knew they had put me on some list in 2020 as a cloud influencer.",date:"1 May, 2022",url:"https://nedinthecloud.com/2022/05/01/diversity-starts-with-me/",image:"Diversity-Starts-with-Me.png",readingTime:"8"},"https://nedinthecloud.com/2022/03/03/migrating-state-data-off-terraform-cloud/":{title:"Migrating State Data Off Terraform Cloud",tags:["hashicorp-terraform","terraform-cloud","terraform-tutorials"],content:`Terraform Cloud (TFC) is a pretty cool service with a bunch of excellent features. But what if you decide it’s just not working out? Sorry TFC. It’s not you. It’s me. How do you migrate off of Terraform Cloud and onto another platform? In particular, how do you migrate your state data off TFC to another backend? That is what I’ll address in this post.
|
||
State Data Storage I want to start with a few Terraform basics, since this is going to factor heavily into how state data is managed and migrated. For starters, state data is the mapping between what exists in your Terraform config and what is deployed in the target environment. I’ve written about it extensively in my post of what happens when changes occur outside of Terraform, so I won’t rehash it here.
|
||
State data needs somewhere to live. The default location is the local backend, creating state files in the same directory as your configuration. There are two different scenarios to consider when using the local backend: working with the default workspace only and working with multiple workspaces. Terraform OSS workspaces allow you to use the same configuration to support multiple target environments. Each workspace has its own state data, which means Terraform needs to create multiple state data files when using the local backend.
|
||
When you first spin up a Terraform configuration using the local backend, you have a single default workspace. You cannot remove the default workspace, but you don’t have to deploy anything to it. Simply create a new workspace and then run your terraform apply. A few key questions come to mind:
|
||
Where does Terraform store the workspace listing? Where does Terraform put the data for each workspace? How does Terraform know which workspace is currently active? Starting with the first question, assuming you’re using the local backend, when you create the first non-default workspace, Terraform creates the directory terraform.tfstate.d. Each non-default workspace gets a subdirectory inside the terraform.tfstate.d directory. For instance, if I have the following workspaces: default, development, and production, then I will have the following file tree.
|
||
> tree terraform.tfstate.d terraform.tfstate.d/ ├── development └── production The state data for each workspace will be stored in a terraform.tfstate file inside each workspace directory. You might notice the default workspace doesn’t have a directory. It will store state data in the file terraform.tfstate in the configuration directory.
|
||
When you run terraform workspace list, Terraform looks at the subdirectories inside terraform.tfstate.d and compiles a list from there, adding the default workspace since it won’t have a directory.
|
||
Terraform knows which workspace is currently active by writing it to the file .terraform/environment. Terraform will create this file when you create your first non-default workspace. The file will have a single entry, the name of the currently active workspace. Running terraform workspace select simply changes the entry in this file.
|
||
The way that Terraform organizes and writes state data for workspaces differs based on the backend being used. The information above is specific to the local backend. The azurerm backend we’ll use later takes a different approach, adding env:workspace_name to the end of each state data blob and storing all the blobs in the same storage account container.
|
||
One thing all the backends have in common is the local .terraform/environment file that tells the Terraform CLI which workspace is currently selected.
|
||
Alright! So let’s deploy to Terraform Cloud Yes, I know this is all about migrating off Terraform Cloud, but first we have to get our data on Terraform Cloud to begin with. If you want to follow along, you can check this directory on my terraform-tuesdays GitHub repo. The configuration I’m going to deploy has the following terraform configuration block:
|
||
terraform { required_providers { azurerm = { source = "hashicorp/azurerm" version = "~> 2.0" } } #backend "azurerm" { # key = "webapp" #} cloud { organization = "ned-in-the-cloud" workspaces { name = "tfc-migration-test" } } } You might notice the azurerm backend config commented out. We’ll eventually try and get our config migrated to an Azure Storage Account, but first we’ll look into how to migrate to a local backend.
|
||
After we’ve run a terraform init and terraform apply our local directory does not have a terraform.tfstate file or terraform.tfstate.d directory. That’s because we’re using the cloud backend. However, in the .terraform directory we have an environment file and a terraform.tfstate file. What’s in those files?
|
||
The environment file serves the same function as before, it has a single entry identifying the currently selected workspace. The terraform.tfstate file holds information about the cloud backend:
|
||
{ "version": 3, "serial": 1, "lineage": "8e0c46e8-ef07-a6f3-1558-08d5bba7d574", "backend": { "type": "cloud", "config": { "hostname": null, "organization": "ned-in-the-cloud", "token": null, "workspaces": { "name": "tfc-migration-test", "tags": null } }, "hash": 4214871454 }, "modules": [ { "path": ["root"], "outputs": {}, "resources": {}, "depends_on": [] } ] } The actual state data is securely stored in the Terraform Cloud workspace. We can grab that state data by going to the UI or by running terraform state pull.
|
||
Migrating to the Local Backend If you were going to migrate your state between any other two backends, the process would generally be:
|
||
Update the Terraform configuration with the new backend Run terraform init to update the backend and migrate state data But, as of right now, Terraform Cloud doesn’t work that way. If we comment out the cloud block from our configuration and run terraform init we will get the following error:
|
||
> terraform init Initializing the backend... Migrating from Terraform Cloud to local state. ╷ │ Error: Migrating state from Terraform Cloud to another backend is not yet implemented. │ │ Please use the API to do this: https://www.terraform.io/docs/cloud/api/state-versions.html │ │ ╵ Ouch. And BTW, that link doesn’t really help you migrate via the API.
|
||
That’s okay! We can figure this out with all the knowledge we’ve already gained. So what do we need in place locally to support our tfc-migration-test workspace with the local backend?
|
||
Create a terraform.tfstate.d directory with a tfc-migration-test subdirectory Put our state data in a terraform.tfstate file in the tfc-migration-test directory Remove the terraform.tfstate file in the .terraform directory that points at the cloud config Update the configuration to remove the cloud block Run terraform init to prepare local files That should do it!
|
||
> mkdir -p terraform.tfstate.d/tfc-migration-test > terraform state pull > terraform.tfstate.d/tfc-migration-test/terraform.tfstate > mv .terraform/terraform.tfstate .terraform/terraform.tfstate.old # Remove the cloud block in the config > terraform init Sure enough, if we make a change to the configuration, a plan will run successfully. This works well enough for a single workspace. If you have multiple workspaces, you’ll need to create a directory for each workspace and pull the state data for that workspace.
|
||
Migrating to an AzureRM Storage Account Will the migration process be any easier if we’re moving to another remote backend instead of the local backend?
|
||
> terraform init -backend-config="backend.txt" Initializing the backend... Migrating from Terraform Cloud to backend "azurerm". ╷ │ Error: Migrating state from Terraform Cloud to another backend is not yet implemented. │ │ Please use the API to do this: https://www.terraform.io/docs/cloud/api/state-versions.html │ │ ╵ No. No, it will not.
|
||
We should be able to follow the same basic process, only this time we need to create the necessary files in the target storage account. Here’s a reminder of what is in the azurerm config block:
|
||
backend "azurerm" { key = "webapp" } That’s just the partial configuration, following HashiCorp’s guidance. The rest of the configuration comes from environment variables and a backend.txt file with this content:
|
||
storage_account_name="tfc40300" resource_group_name="tfc-40300" container_name="terraform-state" The container being used in Azure is terraform-state. The key value is webapp. The azurerm backend will add env: and the workspace name to the end of the key value. So our state file will be webappenv:tfc-migration-test. The migration process will be similar to the local migration:
|
||
Copy our state data to webappenv:tfc-migration-test in the storage account container Remove the terraform.tfstate file in the .terraform directory that points at the cloud config Update the configuration to remove the cloud block and add the azurerm block Run terraform init to download files and validate config We can grab the state data with the same terraform state pull command we used before. And then use the Azure CLI to copy it to our storage account.
|
||
> terraform state pull > statedata > az storage copy -s statedata --destination-account-name tfc40300 --destination-container terraform-state --destination-blob "webappenv:tfc-migration-test" Now we’ll update the backend configuration, rename the terraform.tfstate file, and run a terraform init.
|
||
> mv .terraform/terraform.tfstate .terraform/terraform.tfstate.old # Remove the cloud block in the config > terraform init -backend-config="backend.txt" For authentication and access to the storage account, I have Azure Service Principal information stored in environment variables.
|
||
After running terraform init, we can update the config and run a plan or apply to show that our migration was successful.
|
||
Conclusion In this post we’ve seen how to migrate from Terraform Cloud to either the local or azurerm backend. The process for any other backend would be similar, except you’ll need to know how it handles workspaces and the actual state data.
|
||
As an alternative, you could perform a two stage migration from Terraform Cloud to local and then from local to your remote backend of choice. Either way, you’ll still need to pull the state data down to an intermediary location before uploading to the new remote backend.
|
||
In all likelihood, HashiCorp will update the cloud backend to support direct migration in the not too distant future, making this entire post moot. But until then, I hope this has helped you migrate off Terraform Cloud or at least learn a bit more about how state data and workspaces are managed by Terraform.
|
||
`,summary:`Terraform Cloud (TFC) is a pretty cool service with a bunch of excellent features. But what if you decide it’s just not working out? Sorry TFC. It’s not you. It’s me. How do you migrate off of Terraform Cloud and onto another platform? In particular, how do you migrate your state data off TFC to another backend? That is what I’ll address in this post.
|
||
State Data Storage I want to start with a few Terraform basics, since this is going to factor heavily into how state data is managed and migrated.`,date:"3 Mar, 2022",url:"https://nedinthecloud.com/2022/03/03/migrating-state-data-off-terraform-cloud/",image:"Migrating-Off-TFC.png",readingTime:"8"},"https://nedinthecloud.com/2022/02/23/validating-iac-with-terraform-and-github-actions/":{title:"Validating IaC with Terraform and GitHub Actions",tags:["azure-devops-labs","github-actions","hashicorp-terraform","hashicorp-terraform-tutorial"],content:`This post is an accompaniment to my second appearance on the Azure DevOps Labs YouTube channel. My presentation focused on validating your Terraform code as part of a GitOps workflow. During the demonstration I ran into an unexpected policy violation from Checkov. I thought I would review what I presented in the demo, and the resolution of my policy violation. It was not what I expected!
|
||
Why Validate? One thing that host, April Edwards, brought up during my presentation that I completely failed to communicate was the why of validation. I presumed the why was obvious, but she reminded me that not everyone has been in the IaC space long enough to grok why they should be validating their code or what validation even means.
|
||
When you are writing IaC using Terraform, or any other language, the point is to reliably deploy and manage infrastructure using software development practices. The first time you code something up in Terraform, it probably won’t deploy correctly. That’s okay! You’re not bad at this coding thing. Mistakes are common and inevitable, so software developers have created tools to check for common mistakes and catch them before they make their way into production. Starting with the humble linter, moving to more complex syntax and logic validation, and stretching all the way to static and dynamic code testing.
|
||
The goal of all this testing is to catch and resolve issues as early as possible in the development process. It’s much easier to fix something in your local code editor than trying to hot-patch in production while everything is on fire!
|
||
There are many kinds of tests and validation, but for my demo I covered three areas: formatting, syntax, and security.
|
||
Formatting Properly formatted code doesn’t just look pretty, although I do find it deeply satisfying. It also makes it easier to read and catch common errors. When you have multiple developers all working on the same codebase with their own personal style of formatting, reading through the code can be a nightmare. Consistent formatting lowers the noise level and let’s you read through the code regardless of who wrote it. In the world of Terraform, this is accomplished with the terraform fmt command.
|
||
Syntax There are two kinds of syntax that you can validate with Terraform code. One is basic language syntax, like remembering to close a curly brace in a configuration block. The linter in your code editor of choice should pick up on this type of error easily.
|
||
The second syntax error is more akin to logical issues, where you reference an attribute that doesn’t exist or have more than one resource with the same name label. You haven’t messed up the language syntax, but your code is still invalid. Terraform will pick up on these issues when you run a terraform plan or apply, but you can check beforehand using terraform validate.
|
||
Running validate will check both the language syntax and the logic of the code. You must run terraform init first, because validate will also inspect the syntax of providers and modules you use in the code. I make it a habit of running validate early and often when I’m developing my Terraform configurations. It’s much faster than running a plan and requires less inputs.
|
||
Security When you are deploying infrastructure in an organization, they will likely have best practices and policies for how things should be configured. Those policies can be expressed in code and used to evaluate a Terraform configuration. Tools that look at the configuration code or a plan generated by the code are thought of as static code analysis tools. Dynamic code analysis tools will actually provision the infrastructure and test directly against what has been deployed. For the demonstration, I showed how you could use Bridgecrew’s Checkov static code analysis tool to check your Terraform code against their list of best practices for Terraform and Azure.
|
||
Checkov will flag common security issues, like having the remote desktop port 3389 open to the world or not enabling HTTPS on an Azure Web Application. The list of checks is [fairly exhaustive(https://www.checkov.io/5.Policy%20Index/terraform.html)] and growing everyday. In fact, it changes often enough that it caught me out during my demonstration!
|
||
The Demonstration The goal of the demo was to show how formatting, validation, and security checks could be integrated into a GitOps style workflow with GitHub Actions. If you want to check out the code for yourself, go ahead and fork the repo and try it out!
|
||
The starting configuration has GitHub Actions triggers for commits to the non-default branch, pull requests on the default branch, and commits to the default branch.
|
||
When an engineer wants to update the Terraform code, they will create a new branch from the default branch and update the code. Then they will commit the change and push it up to GitHub. The commit to a non-main branch will trigger a GitHub action workflow that runs both terraform fmt and terraform validate against the code. If either of these tasks fails, then the code needs to be revised before a pull request can be created.
|
||
Once the changes are complete and the code has been properly formatted and validated, the next step is to create a pull request to merge the feature branch to the default branch. This triggers a task in GitHub actions which runs terraform plan and Checkov. The results of both are added as a comment to the pull request. The person reviewing the pull request can see what Terraform will change in the target environment by looking at the plan output, and verify that the code is following best practices by looking at the Checkov results. If the results are not satisfactory, they can ping the original engineer to make updates to their code to bring it into spec.
|
||
If the results look good, the pull request reviewer can merge the pull request, which effectively creates a commit on the default branch. This will trigger a task in GitHub Actions to run a terraform apply with the merged code. That should make the desired changes seen in the plan output in the target environment.
|
||
Whoops! During the pull request review portion of my demo, I expected to have a single Checkov policy failure, specifically CKV_AZURE_14 which checks to make sure that all HTTP requests to the Web App are redirected to HTTPS. Then I was going to fix the issue in code and voila all the checks would pass. But of course that is not what happened! Imagine my horror when I looked at the results and there were two failures (CKV_AZURE_14 and CKV_GIT_4). Of course, when I had checked the demo the week before CKV_GIT_4 didn’t exist!
|
||
I think I kept my cool pretty well in the video - things going wrong during a demo is neither unexpected or new to me. But afterwards I wanted to know what this failure was and whether I needed to fix something in the code or add it to the list of ignored policies. The policy in question, CKV_GIT_4 refers to a check against secrets created in a GitHub repository by Terraform. Here’s the relevant code that was found to be in violation:
|
||
resource "github_actions_secret" "actions_secret" { for_each = { STORAGE_ACCOUNT = azurerm_storage_account.sa.name RESOURCE_GROUP = azurerm_storage_account.sa.resource_group_name CONTAINER_NAME = azurerm_storage_container.ct.name ARM_CLIENT_ID = azuread_service_principal.gh_actions.application_id ARM_CLIENT_SECRET = azuread_service_principal_password.gh_actions.value ARM_SUBSCRIPTION_ID = data.azurerm_subscription.current.subscription_id ARM_TENANT_ID = data.azuread_client_config.current.tenant_id TERRAFORM_VERSION = var.terraform_version } repository = var.github_repository secret_name = each.key plaintext_value = each.value } In particular, the plaintext_value argument was considered insecure for use. Naturally, I went to the github_actions_secret resource to learn more about the argument and what alternatives there were. According to the docs, the plaintext_value argument is “(Optional) Plaintext value of the secret to be encrypted.” Hrm, that doesn’t sound terribly insecure. The value is in plaintext, but GitHub is going to encrypt it and send it securely with HTTPS right?
|
||
The alternative was the encrypted_value argument which is “(Optional) Encrypted value of the secret using the Github public key in Base64 format.” So in this case, the secret value is encrypted using the GitHub public key and encoded in Base64 before being sent to GitHub. Ostensibly, GitHub wouldn’t need to encrypt the value once it is received and could simply store it as is, using the private key later to decrypt the value once it is needed.
|
||
But how am I supposed to encrypt a value like the azuread_service_principal_password.gh_actions.value that is generated within the same Terraform configuration? Terraform does not have an encryption function, so I would need to have a null resource that invokes an local-exec provisioner to encrypt the value with a script and then Base64 encode it. Before I went any further, I went to look at the issues in the GitHub repo for the GitHub provider to see what other folks were doing. Sure enough there was an open issue about how to obtain the value for the encrypted_value argument. I added a comment asking if there was any benefit from a security perspective in using the encrypted_value over the plaintext_value.
|
||
The answer I received was that the only reason to use the encrypted_value argument would be for situations where you didn’t want the value to be stored in Terraform state unencrypted. In my case, the values are coming from the other resources in the same configuration, so they’ll be in the state data regardless. If I were going to pass a value as a variable into Terraform, I could encrypt it ahead of time so that it is never stored in plaintext in the state data.
|
||
I mentioned in the comment thread that the plaintext_value was being flagged as insecure by CKV_GIT_4, and one of the maintainers of the GitHub provider filed an issue in the Checkov repo to change the behavior of CKV_GIT_4. The check was adjusted to only trigger when a value is submitted as a variable and not from another resource/data source in the configuration and the issue was closed!
|
||
What’s the point of this story? I want to highlight a few things. I could have simply added CKV_GIT_4 to my list of ignored checks and moved on with my life, but digging deeper to understand the check and the Terraform resource enhanced my knowledge of GitHub secrets and security considerations around submitting data through Terraform. Because both the GitHub provider and Checkov are open source, I was able to bring up the issue and get it resolved quickly. The community around both Terraform and Checkov are incredibly helpful and diligent. The entire process took about a day! Now the check is more targeted and useful and my code no longer fails against it.
|
||
If you are using Terraform or Checkov, rest assured that you have the supported of an amazing community. They can help you conquer whatever challenges you encounter when trying to adopt IaC and automation for your organization.
|
||
Conclusion Hopefully you enjoyed my presentation and I hope that you’ll take the demo for a spin on your own. I also hope that my little victory with CVK_GIT_4 inspires you to dig a little deeper when you encounter an issue in an open-source product. By filing an issue or commenting on an existing one, you can help make the tools we all use a little bit better!
|
||
`,summary:`This post is an accompaniment to my second appearance on the Azure DevOps Labs YouTube channel. My presentation focused on validating your Terraform code as part of a GitOps workflow. During the demonstration I ran into an unexpected policy violation from Checkov. I thought I would review what I presented in the demo, and the resolution of my policy violation. It was not what I expected!
|
||
Why Validate? One thing that host, April Edwards, brought up during my presentation that I completely failed to communicate was the why of validation.`,date:"23 Feb, 2022",url:"https://nedinthecloud.com/2022/02/23/validating-iac-with-terraform-and-github-actions/",image:"Validating-IaC-with-Terraform-and-GitHub-Actions.png",readingTime:"9"},"https://nedinthecloud.com/2022/02/03/managing-terraform-cloud-with-the-tfe-provider/":{title:"Managing Terraform Cloud with the TFE Provider",tags:["hashicorp-terraform-tutorial","terraform-cloud","terraform-enterprise"],content:`While working on my newest Pluralsight course, Getting Started with Terraform Cloud, I learned a lot about how Terraform Cloud functions and the services it includes. As I was bulding out the demonstrations, I kept thinking about real world environments and how you might go about organizing and managing Terraform Cloud in an SMB or a 10k seat enterprise. That led me down a rabbit hole of using the tfe provider and Terraform Cloud to manage Terraform Cloud. Sounds confusing? It’s not! And I even went so far as to create a module to help you with the process.
|
||
What is Terraform Cloud? Terraform Cloud (TFC) is a hosted service from HashiCorp that expands on the capabilities of Terraform OSS, including a graphical UI, managed state data, team based access control, and integrations with other products. There are a few core constructs to understand when it comes to how TFC is managed and organized:
|
||
Organizations - The core management unit for your Terraform configurations. Workspaces - Exist within an organization and associated with a single Terraform deployment. Teams - Used to grant permissions at the organization and workspace levels. TFC Accounts - Accounts in Terraform Cloud that can be members of one or more organizations. Users - Link between a TFC account and an organization. Must be a member of at least one team in an organization. The central management unit in TFC is the organization. It contains workspaces, teams, users, Sentinel policies, variable sets, and more. When you sign up for a TFC account, you have the option to create an organization through the UI. From there you can create all the other resources I just mentioned, like teams and workspaces. Although, I should mention that to use teams you’ll either need to move up to a paid plan or start a 30 day trial.
|
||
Using Terraform on TFC When you’re just getting started, building out TFC by hand in the UI is fine. No big deal. But just like anything else in the world of Terraform and Infrastructure as Code, it’s better if you can shift to defining things declaratively. Naturally, there is a a tfe provider available to configure both Terraform Cloud and Terraform Enterprise using Terraform.
|
||
There are several benefits to using Terraform to configure TFC:
|
||
Consistency - It’s easier to make sure teams are assigned access consistently when you are pushing the changes from an external source. Efficiency - Making the same change across 10 workspaces is much easier when you’re doing it programmatically. Transparency - All changes made to the organization will be documented in the code changes. You can compare audited changes to your Git log. Basically, if you plan on running TFC at scale with tens or hundreds of workspaces, trying to manage through the UI will be a nightmare. So we are going to use Terraform and IaC to manage TFC.
|
||
But that begs a few questions.
|
||
Where does the Terraform code run to configure your organization? Where do you store the configuration about your TFC organization? How many organizations should you be running? Where to run the code? The answer to the first question could simply be running Terraform OSS on your local desktop, but that’s no fun. You have TFC sitting there, just begging to be used. I suggest having a dedicated organization for configuring other organizations running in TFC. We’ll call it a Configuration Organization (CO). Access to the CO should be tightly controlled, and each managed org will have a dedicated workspace. You could compare this to an empty root domain in Active Directory or the management account for an AWS organization. Except there is no management hierarchy across organizations in TFC.
|
||
Each workspace in the CO will use the VCS workflow (more on that in a moment) to configure a Managed Organization (MO). Changes to a MO will be submitted through pull requests on a repository, and then applied when the PR is approved and merged. Since the workflow does not require anyone to log in to TFC (assuming you enable auto-apply on each MO workspace), the number of people that actually need to log into the CO will be minimal.
|
||
The best part is that due to the limited nature of the CO, you can stay on the Free tier of billing. It includes everything you’ll need. Unless of course you want to create teams and grant special access within the CO, but like I said, you should keep access to the CO down to just a few people. Like… five maybe? (That’s the max number of users in the free tier 😉.)
|
||
Where to store the configuration? Probably the best place to store the configuration data and code is in a private repository on version control system (VCS). All changes and updates can be pushed through the standard GitOps process I described earlier. TFC includes a VCS workflow that can be tied back to your VCS repositories. When a commit is made to a tracked branch, that can kick off a run on TFC to apply the change to the target organization. By default, the run will stop before the apply and wait for someone to approve it, but you can enable auto-apply to skip the approval. After all, someone has already reviewed and approved the proposed changes on the VCS side, right?
|
||
The VCS workflow supports many options for triggering a run. It can track the default branch and root directory, a specific branch, and a particular directory. I see three possible setups based on these options.
|
||
Each managed organization could be a represented as a branch in your source control. Each branch would share the same basic Terraform code, but have a different configuration file to build out the MO associated with that branch. While it’s feasible, I would worry about the possibility of overwriting a branch through a merge and the difficulty of getting a clear picture of all your managed organizations.
|
||
An alternative is to have a default branch with a directory for each organization. The root folder would be empty, and each sub-directory would have the code and configuration files for an organization. When you want to make a change to an organization, you would simply update the configuration files in the corresponding directory. TFC would pick up on the change and start a run on the organization’s workspace in your CO. You could use the same code housed in a single directory for all the MOs, but you might want to experiment with changing the code on one organization before rolling it out to others. Chances are you won’t be running that many organizations simultaneously, so I don’t think it will be much of a burden keeping the code base the same across MO directories.
|
||
The final option is to have a separate repository for each organization. If you’re worried about locking down who can make changes to each organization, then this is probably your best option. You can grant different permissions for each repository and restrict who can access the repository, make commits, and approve pull requests. It would be more work to manage and maintain multiple repositories, but you’d get the benefit of granular permissions.
|
||
Of course, this begs the question: how many organizations are you going to have anyway?
|
||
How many organizations? Chances are that even the biggest enterprises won’t need more than one or two organizations. There is effectively no limit on how many workspaces and teams you can have in an organization. With the teams-based permissions model, you can easily have multiple applications coexist in the same organization. Additionally, workspaces that are in the same organization can share their state data with each other, something that is not possible across organizations in TFC. (It is possible in Terraform Enterprise, but that’s not my focus here.)
|
||
There are three probable reasons for multiple organizations in a company:
|
||
You are a Managed Service Provider running TFC as a service, and each client gets their own dedicated organization. Billing is done per organization, and you have internal business units that want dedicated licenses for users. Business units in your company want to administer their own organization for “reasons”. For the first two situations, using a single repository to manage all your organizations works. The MSP is managing all the organizations, so it doesn’t need to worry about separating out permissions. Likewise, if it’s purely a billing issue, the organizations are probably still managed by a single dedicated team.
|
||
If your business units are dead set on administering their own organization, then separating each organization into its own repository will probably make the most sense.
|
||
Ultimately, it’s unlikely you’re going to have tens or hundreds of organizations at your company. Pick the option that minimizes your administrative overhead, while still meeting the business requirements.
|
||
The more important resources to worry about are workspaces and teams in each organization. That is what we want to manage programmatically.
|
||
Using Terraform and the TFE Provider The TFE (Terraform Enterprise) provider in the public Terraform Registry can configure most aspects of a TFC organization, including workspaces and teams. However, you still need to put the pieces together yourself. And that’s why I decided to write a module for managing an organization. The primary idea is that the module should be able to create the following:
|
||
Workspaces with options set for tags, Terraform version, and team access Teams with membership and organization level permissions Users associated with teams My first idea was to craft complicated variable objects to store all this information and apply it. That quickly became a nightmare, and I realized the best thing to do was store the configuration data in JSON and parse it with Terraform. Here’s the abbreviated format:
|
||
{ "workspaces": [ { "name": "workspace_name", "description": "workspace description", "teams": [ { "name": "team_name", "access_level": "access_level" } ], "terraform_version": "1.1.0", "tag_names": ["tag1"] } ], "teams": [ { "name": "team_name", "visibility": "visibility_level", "organization_access": { "manage_policies": true, "manage_policy_overrides": true, "manage_workspaces": false, "manage_vcs_settings": false }, "members": ["user_email_address"] } ] } The list of users can be extrapolated from the list of team members, so we really only need the teams and workspaces.
|
||
Since there’s no sensitive information in the configuration data, it can be safely stored with the rest of the Terraform code. If you’re worried about the user email addresses being exposed, then you’ll need to add them some other way. Or you could skip creating users and simply create the teams you’d want users placed in.
|
||
With the information stored in JSON, I need to use it in Terraform. Importing JSON data is super easy with the jsondecode and file functions. Essentially the JSON is imported as a complex object, and I can use standard Terraform object reference syntax to extract information.
|
||
For instance, to get all the workspaces I simply do this:
|
||
locals { org_data = jsondecode(file("\${var.config_file_path}")) workspaces = local.org_data.workspaces } From there, it’s a simpler matter of parsing the data with for_each loops for every resource that needs to be generated.
|
||
The JSON format is also easy to extend. If I want to add support for VCS connections or configuring policy sets for Sentinel, I can simply add new resources and update the JSON format. But what happens to older configuration files that don’t have the new fields? No problem!
|
||
Using the try function, I can normalize the JSON input in case the new JSON field hasn’t been added. For instance, if I add support for Variable Sets to the module, I can test for that field in the JSON input.
|
||
locals { json_data = jsondecode(file("path_to_data.json)) variable_sets = try(local.json_data.variable_sets, {}) } If the field variable_sets doesn’t exist in the JSON file, the try function will set the local value to an empty map. The for_each argument in the variable sets resource will see an empty map and not create any instances. We can use the same logic to create normalized structures from our JSON data. I’ve implemented that logic in the latest version of my module, so you only need to add the JSON fields that are necessary for your configuration.
|
||
Using try is pretty cool, and I have to admit I had no idea it even existed until I tried to figure out how to deal with missing values. Turns out I am not the first to encounter this problem. You’ll definitely see more about using try in a different post.
|
||
Conclusion If you plan on adopting Terraform Cloud in your organization and you’d like to manage it programmatically, I hope I’ve laid out a compelling argument for using Terraform Cloud to manage Terraform Cloud. Through the mechanism of a Configuration Organization and workspaces using the VCS workflow, you can automate the management of your TFC organizations with a GitOps style process. To save yourself some time, consider checking out the module I wrote on the public registry and let me know what you think!
|
||
`,summary:"While working on my newest Pluralsight course, Getting Started with Terraform Cloud, I learned a lot about how Terraform Cloud functions and the services it includes. As I was bulding out the demonstrations, I kept thinking about real world environments and how you might go about organizing and managing Terraform Cloud in an SMB or a 10k seat enterprise. That led me down a rabbit hole of using the tfe provider and Terraform Cloud to manage Terraform Cloud.",date:"3 Feb, 2022",url:"https://nedinthecloud.com/2022/02/03/managing-terraform-cloud-with-the-tfe-provider/",image:"Managing-TFC.png",readingTime:"11"},"https://nedinthecloud.com/2022/01/27/choosing-between-count-and-for-each/":{title:"Choosing Between Count and For-Each",tags:["hashicorp-terraform","hashicorp-terraform-tutorial"],content:`Terraform has two looping mechanisms for creating multiple resources, count and for_each. The count meta-argument has been around for a long time, but for_each is a relative newcomer (introduced in version 0.12). Each meta-argument allows you to create more than one resource or module with a single configuration block.
|
||
A common question is when to use count versus for_each. I would make the argument that for_each is almost always preferred, and in this post I hope to show you why.
|
||
Looping Basics Before I explain why for_each is generally superior, it would be useful to understand what is actually happening when you add a looping meta-argument to a Terraform configuration. Let’s start with a simple example.
|
||
resource "local_file" "count_int_loop" { count = 3 content = "This is file number \${count.index}" filename = "\${path.module}/int-\${count.index}.count" } Running terraform apply will generate three files: int-1.count, int-2.count, int-3.count. If we take a look at the state data:
|
||
$> terraform state list local_file.count_int_loop[0] local_file.count_int_loop[1] local_file.count_int_loop[2] We have three resources created by the count meta-argument with an integer based index. The object local_file.count_int_loop is an ordered list of local_file resource objects.
|
||
If we want four files instead of three, all we have to do is increase the count value by one and Terraform will create a fourth file. This is what count was meant for, undifferentiated resource creation based on an integer.
|
||
But what if instead of an integer, we were dealing with a list of items?
|
||
Parsing a List Before the introduction of for_each, the count argument was all we had. (And we liked it.) You could use a list to create mulitple resources by finding the number of elements in the list and using that for the count value. The length() function does an admirable job of accomplishing this.
|
||
Consider the following configuration with a list of toppings defined as a local value.
|
||
locals { toppings = ["lettuce","tomatoes","jalapenos"] } resource "local_file" "count_loop" { count = length(local.toppings) content = "\${local.toppings[count.index]}" filename = "\${path.module}/\${local.toppings[count.index]}.count" } Running terraform apply will generate three files: lettuce.count, tomatoes.count, and jalapenos.count. If we take a look at the state:
|
||
$> terraform state list local_file.count_loop[0] local_file.count_loop[1] local_file.count_loop[2] Once again, we have three resources created by the count meta-argument with a number based index. Just like before, the object local_file.count_loop is an ordered list of local_file resources.
|
||
So far, so good, right? What’s the point of a for_each arguemnt if we can simply use the count argument with a length function? Let’s try the same thing with for_each instead:
|
||
resource "local_file" "for_each_loop" { for_each = toset(local.toppings) content = "\${each.value}" filename = "\${path.module}/\${each.value}.foreach" } Looking at our state now:
|
||
$> terraform state list local_file.for_each_loop["jalapenos"] local_file.for_each_loop["lettuce"] local_file.for_each_loop["tomatoes"] We have three resources created by the for_each meta-argument with a key based reference. The object local_file.for_each_loop is a map (aka hashtable). The keys will be the strings in the set, or if you submit a map it will be the keys of the map. The map values are the local_file resources.
|
||
From a practical standpoint, we have essential generated three files with the same content. Either argument seems to do the trick, so why would you prefer one over the other? Two reasons: consistency and referencing.
|
||
Consistency Something that’s not immediately obvious is how much the order of the items used by count matters. Let’s say we want to add a topping to our list. How about some onions?
|
||
locals { toppings = ["lettuce","tomatoes","onions","jalapenos"] } What do you think will happen when we run terraform plan against the count example? Notice the order of the toppings. Onions is now in index 2 and jalapenos is in index 3.
|
||
$> terraform plan Terraform will perform the following actions: # local_file.count_loop[2] must be replaced -/+ resource "local_file" "count_loop" { ~ content = "jalapenos" -> "onions" # forces replacement ~ filename = "./jalapenos.count" -> "./onions.count" # forces replacement ~ id = "626451b23e9097d6a2c081959703df63424602bf" -> (known after apply) # (2 unchanged attributes hidden) } # local_file.count_loop[3] will be created + resource "local_file" "count_loop" { + content = "jalapenos" + directory_permission = "0777" + file_permission = "0777" + filename = "./jalapenos.count" + id = (known after apply) } # local_file.for_each_loop["onions"] will be created + resource "local_file" "for_each_loop" { + content = "onions" + directory_permission = "0777" + file_permission = "0777" + filename = "./onions.foreach" + id = (known after apply) } Plan: 3 to add, 0 to change, 1 to destroy. Terraform is going to destroy our jalapenos! And that is because when Terraform runs through the count loop, it sees the value onions in index 2 and that value used to be jalapenos. Terraform has to destroy the original local_file.count_loop[2] resource and replace it with the new value. Then it will create a new resource called local_file.count_loop[3] using the jalapenos value.
|
||
The for_each loop doesn’t have this problem. Since it is using a key based reference, it doesn’t care about order. In fact, the for_each argument requires a set or map as input, neither of which are ordered. Terraform simply sees the key onion without a corresponding entry in local_file.for_each_loop and decides to create one. The jalapenos file is not touched. Thank goodness! I like my jalapenos.🌶️
|
||
Count is going to unnecessarily destroy and recreate the jalapenos file, which might not be a problem for a text file. But imagine that’s a Kubernetes cluster running 100+ applications, and you just destroyed it because you added a new cluster in the wrong order. That’s… bad. Possibly a resume generating event.
|
||
Of course, that would never happen becuase you run terraform plan first, right? Right???
|
||
Referencing As pointed out by the previous section, using count results in an ordered list and for_each results in a map. When you need to reference the resources somewhere else in your configuration, you might find that being able to refer to a resource by key instead of index is much easier. Let’s look at a slightly more advanced example where we are trying to create users and groups in Terraform Cloud.
|
||
Users are created with the resource type tfe_organization_membership. I could create users with a count like this:
|
||
resource "tfe_organization_membership" "org_members" { count = length(local.users) organization = local.organization_name email = local.users[count.index] } Or with a for_each like this:
|
||
resource "tfe_organization_membership" "org_members" { for_each = toset(local.users) organization = local.organization_name email = each.value } To add a user to a team on Terraform Cloud, the resource type tfe_team_organization_member is used. The two arguments team_id and organization_membership_id both require a value that is an attribute of the previously generated team or user. It is a value we must look up using a reference. If we’re using the for_each loop to create users, the reference for organization_membership_id looks like this:
|
||
organization_membership_id = tfe_organization_membership.org_members[each.value["member_name"]].id We are using the key member_name to find the correct instance of tfe_organization_membership and returning the id attribute of that instance. While the code looks a little confusing at first, trust me it works. The full block is shown below.
|
||
resource "tfe_team_organization_member" "team_members" { for_each = { for member in local.team_members : "\${member.team_name}_\${member.member_name}" => member } team_id = tfe_team.teams[each.value["team_name"]].id organization_membership_id = tfe_organization_membership.org_members[each.value["member_name"]].id } If we had used a count argument to create the users, we would need to use a for expression with a filter to look up the membership id value. Something along the lines of this:
|
||
organization_membership_id = ([for user in tfe_organization_membership.org_members : user.id if user.email == each.value["member_name"]])[0] Would it work? Sure. Is it efficient? Nope. The for expression has to loop through all the tfe_organization_membership resources to find the one that matches. No big deal if we have three users. Pretty big deal if we have three thousand. Reference by key is one of the best things about maps/hashtables/dictionaries, whatever you want to call them.
|
||
When to use count You can use count if you don’t care about uniqueness or references in your configuration. If every item created by a loop is ephemeral and functionally identical, then there’s probably no benefit to using for_each. If you don’t need to refer to anything by the key, then using count could be fine. On the other hand, it’s about the same amount of work to use either, and for_each has some serious benefits.
|
||
What about conditionals? One use that still seems relevant is using count with a zero value to make the creation of a resource optional. What do I mean? Consider this:
|
||
resource "local_file" "count_optional" { count = local.create_file ? 1 : 0 content = "Hello!" filename = "\${path.module}/count-create.txt" } The creation of a new organization will happen if the variable create_new_organization is set to true and not if its set to false. You’re only ever creating one or zero of an item. Is there a way to replace this with for_each? If so, is there any benefit?
|
||
To answer the first question. Yes, you can do it.
|
||
resource "local_file" "for_each_optional" { for_each = local.create_file ? toset(["any_value"]) : toset([]) content = "Hello!" filename = "\${path.module}/for-each-create.txt" } Is there any benefit? Not that I can think of. The count version feels more intuitive, but I don’t think it would be any more effective than the for_each loop.
|
||
What about using the index value? I started out the looping overview by using an integer with count. This is the one time when count has a clear advantage. The count argument takes a number and counts up to that number. You can access the current iteration using the count.index expression. Count makes more sense if I am creating resources based off a number, instead of set, list, map, or other object.
|
||
Could you replace it with a for_each argument? Would there be any benefit?
|
||
To answer the first question, yes you can do it.
|
||
for_each = toset(range(3)) The range function creates a list of integers starting with 0 and going to the max, non-inclusive. The toset() function turns that list of integers into a set. The each.value value will be the integer of the current iteration, so it’s basically the same as index.count.
|
||
Is there any benefit? I’d have to say no. The syntax is clunky and requires the execution of two functions to get it working. Contrasted to the count argument, there is no tangible benefit of using for_each in this situation.
|
||
Conclusion To sum up, here’s the general advice. If you are creating multiple resources based off an integer, and the resources are undifferentiated, then count works just fine. Any time you are using an input that is more complex than an integer, the proper answer is going to be for_each. Conditional resource creation is a bit of a toss up, so just do what feels intuitive to you or aligns with your team’s conventions.
|
||
`,summary:`Terraform has two looping mechanisms for creating multiple resources, count and for_each. The count meta-argument has been around for a long time, but for_each is a relative newcomer (introduced in version 0.12). Each meta-argument allows you to create more than one resource or module with a single configuration block.
|
||
A common question is when to use count versus for_each. I would make the argument that for_each is almost always preferred, and in this post I hope to show you why.`,date:"27 Jan, 2022",url:"https://nedinthecloud.com/2022/01/27/choosing-between-count-and-for-each/",image:"Count-and-For-Each.png",readingTime:"9"},"https://nedinthecloud.com/2022/01/18/replacing-the-template_cloudinit_config-data-source/":{title:"Replacing the template_cloudinit_config data source",tags:["hashicorp-terraform-tutorial","terraform"],content:`I was working on updating some Terraform code as part of a consulting engagement and I came across an EC2 configuration that was using the template_cloudinit_config data source to create user_data to send to the instance. Since I know that the template provider has been archived by HashiCorp and the recommendation is to use the templatefile function, I endeavoured to replace the template_cloudinit_config data source with templatefile and that is where I fell down a rabbit hole of the MIME format, cloud-init picadillos, and nested templates.
|
||
I thought I would write a post about my little adventure and the eventual workaround. If you don’t care and just want the answer, feel free to skip to the end or simply use the module I wrote to solve this problem.
|
||
The Template Provider The template provider has been archived by HashiCorp in favor of the templatefile function. To understand why, you can check out a whole video I did on it, but I can quickly summarize here. The template provider has a data source called template_file which will render text based on a template and variable inputs. Since you’re using a provider, Terraform has to hand off the work to a provider plugin. The templatefile function does exactly the same thing, but because it’s a function, it is included in the Terraform binary.
|
||
The evaluation and execution time for a function is much faster than a provider plugin. With the introduction of the templatefile, the template_file data source was no longer required. However, there is another data source in the template provider that doesn’t have a comparable function, template_cloudinit_config.
|
||
The other data source in the template provider is template_cloudinit_config. To support folks who want to use that data source, HashiCorp created the cloudinit provider with a single data source called cloudinit_config. Essentialy, it functions exactly like the template_cloudinit_config data source, but it’s in a new provider that is being actively maintained by HashiCorp.
|
||
DIY Cloud-Init But wait. If the templatefile function is faster than the template_file data source, wouldn’t the same be true for the template_cloudinit_config data source? Unfortunately, there is no templatecloudinit function, so how can I create the same thing using functions? First we need to understand what is being created by the template_cloudinit_config data source and recreate it.
|
||
MIME The template_cloudinit_config data source creates a multipart MIME configuration for cloud-init. This is the moment I realized I was in for a yak-shaving expidition. What the hell is MIME? And what it multipart about it? And what does it have to do with cloud-ini? MIME at least sounds familiar.
|
||
MIME is the multipurpose internet mail extensions standard created for handling mail messages that use non-ASCII characters and to support attachments. As a former Exchange Admin I remember seeing MIME from time to time in various menus and dropdowns, but I never had to do anything with it.
|
||
Even though MIME was originally intended for email messages, it has been adapted for use in HTTP communcation and the cloud-init standard. In addition to supporting different media types, MIME allows you to construct a single configuration that contains multiple parts, each with their own content type. So there we have it, multipart MIME. Now, what does that have to do with cloud-init? Further down the rabbit hole we go!
|
||
Cloud-init is an industry standard used to initialize compute instances in a cloud by reading in information like cloud metadata, user data, and vendor data. User data is provided by the client to initialize the system after the cloud metadata portion is complete.
|
||
User data must be in a multipart MIME format and optionally gzipped to keep the user-data content under the 16KB limit. Cloud-init supports multiple content types in the MIME configuration including cloud-config, jinja2, and x-shellscript. If you don’t know what any of those are, don’t worry, the official cloud-init docs have you covered.
|
||
To sum up, multipart MIME is a format originally intended for email messages, but adatped for use by cloud-init to assist with configuring compute instances on first boot. The template_cloudinit_config data source creates a multipart MIME configuration. We need to understand the format to recreate it without the data source.
|
||
Multipart MIME Format If I want to natively produce multipart MIME content using Terraform functions, I will need to know what the resulting content looks like. The general format is something like this:
|
||
MIME-Version: 1.0 Content-Type: multipart/mixed; boundary=MIMEBOUNDARY This is the beginning of the cloud-init, followed by a boundary delimiter for the next part. --MIMEBOUNDARY Content-Transfer-Encoding: 7bit Content-Type: text/cloud-config Mime-Version: 1.0 YAML for the cloud-config --MIMEBOUNDARY Content-Transfer-Encoding: 7bit Content-Type: text/x-shellscript Mime-Version: 1.0 Bash script to run --MIMEBOUNDARY-- That should be pretty easy to replicate. I can use the templatefile function for each part of the MIME content and build it inline using standard Terraform constructs like for expressions.
|
||
The last thing to cover is the encoding. The template_cloudinit_config data source gives the option to compress with gzip and encode with base64. Fortuantely, Terrform has a base64gzip function which will take care of that for me.
|
||
Building the MIME Content Here’s the orginal code that uses the template_cloudinit_config data source:
|
||
data "template_file" "cloud_init" { template = file("cloud-init.yaml") vars = { package_update = "true" package_upgrade = "false" } } data "template_file" "x_shellscript" { template = file("startup-script.sh") vars = { name = "Arthur" } } data "template_cloudinit_config" "config" { gzip = true base64_encode = true part { content_type = "text/cloud-config" content = data.template_file.cloud_init.rendered } part { content_type = "text/x-shellscript" content = data.template_file.x_shellscript.rendered } } We’re using the template_file data source twice and the template_cloudinit_config data source once. The goal is to replace all of those with native functions. First we need to build the parts for cloud-config and x-shellscript. Ideally, this should be extensible, so if someone wants to add more parts, it’s pretty easy to do so.
|
||
The parts information includes the content type, content from a file, and variables for that file. We can store that with a list of objects stored in a local value:
|
||
locals { cloud_init_parts = [ { filepath = "cloud-init.yaml" content-type = "text/cloud-config" vars = { package_update = "true" package_upgrade = "false" } }, { filepath = "startup-script.sh" content-type = "text/x-shellscript" vars = { name = "Arthur" } } ] } We can add more parts by adding another object to the cloud_init_parts list. Next we need to render each part into the content used by the multipart MIME format. Using a local value and a for expression, each part can be stored in a list as a string.
|
||
locals { cloud_init_parts_rendered = [ for part in local.cloud_init_parts : <<EOF --MIMEBOUNDARY Content-Transfer-Encoding: 7bit Content-Type: \${part.content-type} Mime-Version: 1.0 \${templatefile(part.filepath, part.vars)} EOF ] } Finally, we need to put it all together with the header and footer of the format. I created a cloud-init.tpl file with a for expression we can pass our rendered parts to:
|
||
Content-Type: multipart/mixed; boundary="MIMEBOUNDARY" MIME-Version: 1.0 %{~ for part in cloud_init_parts ~} \${part} %{~ endfor ~} --MIMEBOUNDARY-- Using a combination of the templatefile and base64gzip functions, we have the final product:
|
||
locals { cloud_init_gzip = base64gzip(templatefile("cloud-init.tpl", {cloud_init_parts = local.cloud_init_parts_rendered})) } Voila! The local value cloud_init_gzip can be used in place of the rendered content from the template_cloudinit_config data source. And I didn’t have to use any provider plugins to do it.
|
||
Results You might be wondering if all this mucking about was worth it. I mean there is a perfectly good cloudinit provider. Why not just use that and call it a day? That’s fair! In part, I just wanted the challenge of doing it with native functions. But there are two other considerations here. First, we’ve removed a dependency on a plugin. That’s one less codebase we have to trust and pull on each terraform init.
|
||
Second, in theory using the native functions should be faster than the provider plugin. So, is it? Yes!
|
||
Here’s a run of the original code:
|
||
$ time terraform apply -auto-approve No changes. Your infrastructure matches the configuration. Terraform has compared your real infrastructure against your configuration and found no differences, so no changes are needed. Apply complete! Resources: 0 added, 0 changed, 0 destroyed. real 0m4.598s user 0m1.267s sys 0m1.039s And here’s a run of the updated code:
|
||
$ time terraform apply -auto-approve No changes. Your infrastructure matches the configuration. Terraform has compared your real infrastructure against your configuration and found no differences, so no changes are needed. Apply complete! Resources: 0 added, 0 changed, 0 destroyed. real 0m0.197s user 0m0.013s sys 0m0.099s So um, yeah. It’s a bit faster. Does saving ~4s matter in the grand scheme of things? Not at my scale. But imagine if you’ve got a large configuration that needs to render multiple template provider data sources on every run. And that configuration is baked into a CI/CD pipeline that runs every time someone opens a PR or makes a commit. The time savings could start to stack up!
|
||
Conclusion Replacing the template_file data source with the templatefile function is a slam dunk in terms of simplicity and support. But getting rid of the template_cloudinit_config data source is less straightforward. While you could use the cloudinit provider, there’s an opportunity to save time and remove a dependency if you’re willing to do a little extra work. And I kinda did the extra work for you!
|
||
In fact, if you’d like to consume this as a module, you can do exactly that: https://registry.terraform.io/modules/ned1313/native/cloudinit/latest.
|
||
Of course that introduces a new dependency on an external module, so that’s entirely up to you. But at the very least, you’ll still get the performance improvements.
|
||
`,summary:"I was working on updating some Terraform code as part of a consulting engagement and I came across an EC2 configuration that was using the template_cloudinit_config data source to create user_data to send to the instance. Since I know that the template provider has been archived by HashiCorp and the recommendation is to use the templatefile function, I endeavoured to replace the template_cloudinit_config data source with templatefile and that is where I fell down a rabbit hole of the MIME format, cloud-init picadillos, and nested templates.",date:"18 Jan, 2022",url:"https://nedinthecloud.com/2022/01/18/replacing-the-template_cloudinit_config-data-source/",image:"Replacing-template_cloudinit_config.png",readingTime:"8"},"https://nedinthecloud.com/2022/01/12/using-the-terraform-cloud-configuration-block/":{title:"Using the Terraform Cloud Configuration Block",tags:["hashicorp-terraform","hashicorp-terraform-tutorial","terraform-cloud"],content:`Terraform 1.1 brings with it some new cool Terraform Cloud management options. Cloud blocks, Tags, and Workspace commands Oh MY! But wait. What was broken about the old system? And why is this better? Let’s dig in.
|
||
Terraform Cloud Workspaces Everything in here is about the CLI workflow for a Terraform Cloud workspace. If you’re using the VCS or API workflow, you can safely ignore most of this post. The only major improvement for you is the proper evaluation of terraform.workspace.
|
||
Terraform Remote Backend Before Terraform 1.1, the way you connected a Terraform configuration to Terraform Cloud in a CLI workflow was through the use of the backend block in a terraform configuration block. The backend type was remote and it came with settings for the hostname, organization, and workspaces.
|
||
The workspace block had two possible arguments:
|
||
name: associated the configuration with a single workspace in TFC with a matching name. prefix: matched your current local workspace to a workspace in TFC by adding a prefix. The two arguments are mutually exclusive. You might be wondering about the prefix, so allow me to illustrate with an example:
|
||
terraform { backend "remote" { hostname = "app.terraform.io" organization = "taconet" workspaces { prefix = "networking-" } } } When you initialize the configuration, it will look for any workspaces in the target organization that have the prefix “networking-”. The next action will depend on what it finds:
|
||
No matching workspace: Terraform will prompt you to create one using the terraform workspace command. One matching workspace: Terraform will automatically select the workspace for you. Multiple matching workspaces: Terraform will prompt you to select a workspace from the list. Since we are starting with an empty organization, there will be no matching workspaces. The following command will create a workspace:
|
||
terraform workspace new dev Listing out the workspaces at the CLI will show the following:
|
||
$ terraform workspace list * dev Looking at the workspaces on Terraform Cloud, you’ll see a workspace called networking-dev. Terraform is adding the prefix for the workspace it generated in Terraform Cloud.
|
||
The main problem with the prefix argument is the cognitive dissonance between what you’re seeing at the command line - a workspace called dev, and in Terraform Cloud - a workspace called networking-dev. This is further compounded by a problem with the terraform.workspace value.
|
||
Before Terraform 1.1, the workspace used by the remote runner was always the default workspace. If you used the terraform.workspace value in your code, it would evaluate to default no matter what the name of the workspace was locally or in Terraform Cloud.
|
||
Terraform 1.1 set out to fix this and add room for future capabilities.
|
||
Terraform Cloud Block Terraform 1.1 introduced the cloud block as an alternative to backend "remote". The arguments were mostly the same including hostname and organization. The main change was with the workspaces block, which now had the name and tags arguments.
|
||
name: associated the configuration with a single workspace in TFC with a matching name. tags: match your current local workspaces to workspaces with matching tags. One of the goals behind the cloud block was to remove the cognitive dissonance between local workspaces and Terraform Cloud workspaces.
|
||
How did it do that? By giving you full control over naming each workspace, but at the same time applying consistent metadata tags to each workspace associated with a configuration. An example would be helpful.
|
||
terraform { cloud { organization = "taconet" workspaces { tags = ["cloud:aws", "security"] } } } When you initialize the configuration, Terraform will look for any workspaces in the target organization that have the tags “cloud:aws” and “security”. The next action will depend on what it finds:
|
||
No matching workspace: Terraform will prompt you to create one directly. One matching workspace: Terraform will automatically select the workspace for you. Multiple matching workspaces: Terraform will prompt you to select a workspace from the list. You might notice that instead of asking you to creating a workspace using the terraform workspace new command, the dialog prompts you to do so as part of the workflow. That’s a small, but appreciated improvement to the experience.
|
||
Let’s say I created a workspace called shared-services-dev during initialization. Running the terraform workspace list command would show me the following:
|
||
$ terraform workspace list * shared-services-dev Looking at the workspaces on Terraform Cloud, I will see a workspace named shared-services-dev with the tags “cloud:aws” and “security”. The dissonance between my local workspaces and what I see in Terraform Cloud is gone.
|
||
Even better, regardless of which workflow you use, Terraform 1.1 will use the actual workspace name on the remote runner. That means the terraform.workspace value will evaluate properly again.
|
||
Migrating from Backend Remote to Cloud What if you’ve gone all in on using the backend "remote" method to manage your workspaces and now you want to move to the cloud block? Whether you are using the name or prefix argument in your backend block, the migration process is essentially the same.
|
||
If you’ve been using the prefix argument, then you will need to decide on tags to apply to the migrating workspace. For the name argument, you can simply use the same value for the name argument in the cloud block.
|
||
Let’s look at an example of the prefix scenario. We’ve got three workspaces in Terraform Cloud: application-dev, application-staging, and application-prod. The current backend block looks like this:
|
||
terraform { backend "remote" { hostname = "app.terraform.io" organization = "taconet" workspaces { prefix = "application-" } } } And a workspace listing on your local workstation would show the following:
|
||
$ terraform workspace list dev * prod staging The first thing to remember is that all the state data and workspace information is stored up in Terraform Cloud. The workspaces you have on your local workstation do not matter. What you’re trying to do is map to the Terraform Cloud workspaces using the new cloud block. Since we have multiple workspaces using the same configuration, we are going to use the tags argument.
|
||
Let’s say we want to use the tag “app:taco” to identify our migrated workspaces. We can update our configuration replacing the backend block with the cloud block:
|
||
terraform { cloud { organization = "taconet" workspaces { tags = ["app:taco"] } } } Because we are changing our backend, we need to run terraform init. You might think you need to go into Terraform Cloud and add the “app:taco” tag to the three workspaces, but you don’t! When you run terraform init, Terraform will recognize you are migrating from the remote backend to the cloud backend. Stored in the local state file is the following information:
|
||
"backend": { "type": "remote", "config": { "hostname": "app.terraform.io", "organization": "taconet", "token": null, "workspaces": { "name": null, "prefix": "application-" } }, "hash": 1338747517 } During the migration process, Terraform will use the prefix information stored in local state and your existing list of local workspaces to find the matching workspaces in Terraform Cloud. Then it will apply the tags list in the cloud block and migrate the state. It will also update your local workspace names to match the names in Terraform Cloud.
|
||
One important caveat! If you have a bunch of existing workspaces in Terraform Cloud, chances are they are set to use an older version of Terraform. The cloud block and migration functionality require that your Terraform Cloud workspace is at Terraform v1.1 or higher. Before you run the migration, go into each impacted workspace and update the Terraform version in the General settings. If you don’t, you’ll get this fun message:
|
||
│ Error: Error migrating the workspace "dev" from the previous "remote" backend │ to the newly configured "cloud" backend: │ Error loading state: │ Remote workspace Terraform version "1.0.1" does not match local Terraform version "1.1.2" Don’t worry! Nothing is broken. Terraform fails gracefully on the migration. Simply go and update the workspaces to the proper Terraform version and run terraform init again.
|
||
Once the migration completes, you’ll see that your local workspace names now match what is in Terraform Cloud, and the Terraform Cloud workspaces have the proper tags.
|
||
Migration complete! Your workspaces are as follows: * application-dev application-prod application-staging Cloud Block Questions HashiCorp could have introduced these improvements without creating a new configuration block type, so why did they do it? In part, I think it comes down to semantics. Terraform Cloud isn’t just a backend, it’s got a lot more services and features, including remote operations. Creating the cloud configuration block makes the difference clear and creates a migration path.
|
||
The other part is future updates and features. Instead of adding more arguments to the backend block that are Terraform Cloud specific, they can leave the backend block alone and introduce new options in the cloud block. What are those new options? No idea. But you can bet they’re coming soon.
|
||
Conclusion The new cloud block in Terraform 1.1 provides an improved experience for those using the CLI workflow. Workspace names match between local and Terraform Cloud, and you can use tags to manage multiple workspaces. This change paves the way for future improvements in Terraform Cloud and the CLI experience. Migration from the remote backend is a simple affair as long as you remember to update the version of Terraform used by your workspaces.
|
||
`,summary:`Terraform 1.1 brings with it some new cool Terraform Cloud management options. Cloud blocks, Tags, and Workspace commands Oh MY! But wait. What was broken about the old system? And why is this better? Let’s dig in.
|
||
Terraform Cloud Workspaces Everything in here is about the CLI workflow for a Terraform Cloud workspace. If you’re using the VCS or API workflow, you can safely ignore most of this post. The only major improvement for you is the proper evaluation of terraform.`,date:"12 Jan, 2022",url:"https://nedinthecloud.com/2022/01/12/using-the-terraform-cloud-configuration-block/",image:"Using-Terraform-Cloud-Blocks.png",readingTime:"8"},"https://nedinthecloud.com/2022/01/03/2021-year-in-review/":{title:"2021 Year in Review",tags:[],content:`Entering 2021, I think we all were happy to say adios to 2020 and harbor a little hope for the coming year. After all, we had several viable vaccines about to roll out, meaning this worldwide pandemic might just end up in our rearview mirror. Oh what sweet summer children we were. Maybe 2021 didn’t work out the way we all hoped, but that doesn’t mean nothing positive happened. Amongst all the turmoil and bad takes were some gems of genuine kindness.
|
||
As is tradition, I thought I would take a moment to reflect on what Ned in the Cloud accomplished in 2021, and what I plan to do in 2022. If you’d like to get a sense for how I entered 2021, just check out last year’s post on the topic.
|
||
Technical Education The primary goal of Ned in the Cloud is to create educational material for a technical audience. I want to help folks learn and grow, gaining the skills they need to excel in their chosen field. This goal is pursued through multiple formats, including videos, courses, books, and podcasts.
|
||
Videos My YouTube channel has gained a pretty decent following since last year. At the beginning of 2021 I was at 1K subscribers and the channel currently has 3.6K subscribers. I wouldn’t call it bombastic growth by any means, but it’s a respectable and steady rate of growth. Things may have slowed down in the last three months, since I was not posting videos as regularly.
|
||
In 2022, I plan to return to a regular release schedule for Terraform Tuesday and add back in content beyond just Terraform. Exactly what that content will be is still up for debate. I’m tinkering with a few ideas including: a weekly tech analysis video, monthly videos on HashiCorp Boundary, and bringing back the Best Career Advice Ever series. I’ve also been tinkering with Raspberry Pi’s again, so who knows what the future holds?
|
||
Courses Pluralsight Even thought I didn’t match my course output from 2020 of seven courses, I still managed to produce six courses! One of the courses was the second update of my Terraform - Getting Started course, this time updated for version 1.0 of Terraform and completely revamped to focus on teaching through challenges.
|
||
When I reflect on how little I knew about Terraform when I developed the first version of the course in 2017, I find it striking how well the course did. I think that is in part due to pent up demand and a lack of available resources. It probably also has to do with my presentation style. Even if I didn’t know everything about Terraform, I presented the material in an engaging and entertaining style. That goes a long way.
|
||
The first revision of the course was in 2019, where I rectified some of the technical shortcomings of the course based on my more extensive understanding of Terraform. The improved technical quality of the information was reflected in improved performance of the course. It had been lingering in the top 200 courses before the first update. After the update was published, it spiked to the top 25 and has rarely left that space.
|
||
For the newest revision, I focused on instructional design and completely overhauled the course. All the technical information is still there, but I changed the order it was presented in and leveraged the story structure of the course to have the learner build alongside me. I believe it’s a more engaging and immersive way to learn the content, and that is born out by the numbers.
|
||
Speaking of numbers, I now have 26 active courses on Pluralsight. In the last year 91K viewers have watched 116K hours of my content on Pluralsight. That is a staggeringly large number to me. I am humbled that folks have found my content useful and worth watching on the platform.
|
||
Looking towards 2022, I am already working on a Getting Started course for Terraform Cloud and I plan to revise my other Terraform courses, including Deep Dive, AWS, and Azure. I’d also like to create a Terraform course centered around GCP. Beyond the world of Terraform, HashiCorp’s new Boundary product will probably warrant a course in the second half of 2022. Some of my Microsoft Azure content is also due for a refresh as well. With 26 courses in the catalog, revisions and refreshes are going to eat up more of my time.
|
||
Labs and Project Pluralsight is also providing labs to their enterprise customers, both on the Pluralsight platform and through A Cloud Guru. I think labs are an ideal way to learn a technology, and I’m hoping to get involved with the creation of labs in 2022.
|
||
In addition to Pluralsight, I am also in the process of developing a series of liveProjects for Manning based on using Terraform to manage Azure Kubernetes Service. You can expect those to drop in Q1 of 2020.
|
||
Books At the end of 2020 I had started writing a certification guide for the Vault Associate exam. I finished the guide in 2021 and updated my Terraform Associate certification guide as well. I plan to keep these up to date in 2022 and possibly start a study guide for the upcoming Vault Operator certification.
|
||
Podcasts At the start of 2021, I was running two podcasts: Day Two Cloud and Buffer Overflow. I had made the difficult decision to retire The Daily Check-In in December of 2020, since it was a lot of work to create, edit, and publish a video every day. That decision was not built to last.
|
||
Day Two Cloud I promised that 2021 would have some banger guests and I wasn’t lying. We recorded some of my favorite episodes ever this year, and I am happy to report that our subscriber base has gown accordingly. Here are my top five episodes from 2021:
|
||
Ep. 88 - The Tech Recruiter with Tyler Desseyn Ep. 100 - Get to Know Crossplane with Daniel Mangum Ep. 114 - Transitioning from a Tech Role to Mgmt Ep. 111 - Infra as Software with Kris Nova Ep. 128 - DevOps’ing All the Things with Kyler Middleton We’ve also seen an uptick in sponsored episodes lately. If you’ve listened to any of our existing sponsored episodes, you know that we don’t just parrot the talking points from a vendor’s marketing department. That would be a disservice to you - the listener - and honestly to the sponsors themselves. It’s only effective marketing if people listen to the episode, and to that end we do our best to make the sponsored content interesting and informative.
|
||
I’ve also been doing my best to transcribe our episodes and post that with the show notes on daytwocloud.io. It seems like I’m always a few weeks behind, and that is because I don’t just feed it through a transcription engine. I actually listen to each episode and correct the auto-generated transcript to be as accurate as possible. Even listening at 1.4x speed, it still takes at least an hour to do each episode. In 2022, I have started blocking out 90 minutes on every Wednesday to get the transcription done. Now I just need to get caught up!
|
||
The Daily Check-In In December of 2020, I posted the last Daily Check-In video on my YouTube channel. Honestly, the daily grind of a video was too much to sustain and I wasn’t sure how much value folks were actually getting from it. After posting the video, I got feedback from a few people that they really enjoyed the Daily Check-In content. All of them were listening to the audio-only version I had been posting through Anchor. At the end of January 2021, I decided that I would revive the Daily Check-In as a podcast only and try to post daily episodes during the week. I got rid of the themed days element, and just focused on talking about things that were important to me. And if there was a day where I had nothing to talk about or life got too crazy, then I just skipped it. That has proved a healthy approach and it has allowed me to continue the Daily Check-In throughout 2021.
|
||
In terms of audience, it’s still fairly small, and that’s okay. People I know and respect are listening, and I think I am doing the episodes as much for me as I am for them. It also provides great fodder for new blog posts and videos. Turns out writing still has a place in this world! And of course, doing a daily podcast keeps my speaking skills sharp, which might come in handy at some point?
|
||
Buffer Overflow Good lord, three podcasts is probably too many. In fact, it is 100% too many, and that is why I made the difficult decision to drop off as a host of the Buffer Overflow podcast. The reasons are many and sundry, but I think it boils down to three key points:
|
||
The podcast was associated with Anexinet, my former employer. That was… weird. There was a significant amount of work to create, produce, and edit an episode. There was no way to monetize the podcast, largely because of point one. My decision to leave the show was essentially the end of the podcast. The rest of the cohosts decided to put the podcast on indefinite hiatus. That was back in April 2021, and not only have there been no new episodes, the latest redesign of the Anexinet website removed the links and feed. (I still have all the original files though.)
|
||
Failures and Deprecations It wouldn’t be an honest accounting of 2021 if I didn’t acknowledge some of my failures and deprecated features. Here’s a fun, bullet pointed list of things I failed at!
|
||
Learning Go: Once again I planned to learn Go this year, and once again I did not. The two primary things standing in my way are time and motivation. Oh well, I guess I’ll try again next year! Passing the CKA Exam: I sat the CKA exam in January of 2021 and failed it. I have declined to take it again. My primary reason for taking the exam was to learn more about the guts of Kubernetes. Despite failing the exam, I achieved my goal, and that’s enough for me. For now. Learn about Consul or Packer: You might have noticed that I am extremely familiar with some of the HashiCorp stack, but I don’t really know much about their Consul, Nomad, or Packer products. I know enough to be dangerous, but I’d like to really dig into them. I went so far as to buy a course on Udemy on Consul by Bryan Krausen. It has languished in my email hoping and praying for a day I make time to take it. Hold strong buddy, maybe 2022 will be your year? Patreon: I started a Patreon b/c it’s what all the cool kids were doing. I am not a cool kid, and now I don’t have a Patreon. Lessons were learned. Being an analyst: This wasn’t something I failed at, quite the contrary it turns out I am quite good at being an analyst and benchmark tester. But I don’t enjoy doing it. So I stopped. How about that? Appearances Over the course of 2021 I appeared on a decent amount of things. This probably won’t be an exhaustive list, but I figured it was worth writing it down.
|
||
Tech Field Day Stephen Foskett and the folks over at Gestalt IT continued to put together awesome events with amazing people. In 2021 I participated in the following events:
|
||
Cloud Field Day 10 Tech Field Day Extras - Scality Tech Field Day Extras - NGINX Spring 2.0 Cloud Field Day 12 While the first three were virtual events, Cloud Field Day 12 was a hybrid event with some presenters and delegates attending in-person and others joining remotely. I was extremely excited to attend in-person, and it was one of the highlights of my year, although it did leave me feeling drained and ready to be a hermit for a bit.
|
||
There really is no replacement for in-person events. My favorite events are smaller and more intimate, and Tech Field Day gives me that in spades. I can’t wait to attend another event, assuming Omicron doesn’t cancel in-person stuff for 2022.
|
||
Other Events Critical Conversations with Bluecat: A great discussion with people I am completely unqualified to share a stage with. HashiTalks 2021: Not only did I present a sessions, I also got to MC a portion of HashiTalks! I think I enjoyed that more that presenting. ActualTech Media events: I’ve been doing short 5-10 minute videos for ActualTech’s various events. They are fun to do and sometimes I repost them on YouTube as well. Microsoft’s DevOps Labs YouTube channel: I did one video on Getting Started with Terraform and GitHub Actions at the end of 2021. It is the first in a series of presentations I have planned for 2022. Apps On Cloud Summit: This was an excellent event put on by the folks at Turbonomic (now an IBM company). I presented one session, and also did a panel with Kenneth Hui and Eric Wright. CTO Advisor: I was a guest on the CTO Advisor podcast, and also did a video with Keith when he came to visit me on his Road Trip! C.R.E.A.M.: Cloud Rules Everything Around Me! This was a Pluralsight and A Cloud Guru event to celebrate acknowledge and celebrate the acquisition of ACG by PS. Random and Sundry There’s a few other notable things that don’t fit nicely into a category, so in no particular order:
|
||
New website! - With the help of BluePixel, I migrated to a completely new website. It’s still Wordpress-based, but now it looks like someone actually designed this thing instead of selecting a default theme and customizing it badly. Private training videos - During 2021 I created a series of videos intended for an internal training site at a major insurance company. It was fun to do and made me think deeply about some of the fundamentals of cloud, DevOps, and Infrastructure as Code. Core Contributor to Boundary - When HashiCorp released Boundary back in 2020, I got excited about engaging with a new product. They put up a reference architecture for AWS, but not one for Azure. So I made one and submitted it. It was merged into their main branch and I became a contributor! AWS Solution Architect Professional - My AWS SA-Pro cert was about to expire, so I decided to renew it. After two weeks of studying, I sat the exam and passed! Turns out I retained more from the first time around than I thought. HashiCorp Ambassador and Microsoft MVP - I was re-awarded both the Microsoft MVP award for Azure and HashiCorp Ambassador for 2021. The MVP award was automatically renewed due to the pandemic, and I hope I’ve done enough this year to warrant being renewed next year! Personal Stuff On a more personal note, there were a few exciting things throughout the year.
|
||
Trail running - I took up trail running this year and I freakin’ love it. Trail running forces you to be present in a way that regular road running doesn’t. Trails are uneven, obstacle-laden, and occasionally slippery. If you stop paying attention you’re likely to get hurt. I ran in my first trail run race in September and finished first in my age group and sixth overall. It was an 18 mile run in the Poconos and I finished it in just under three hours. Broke my toe - Speaking of getting injured, turns out that I let my attention slip while trail running in early October and managed to break my pinkie toe and badly bruise the toe next to it. Of course I was still two miles from my car, so I ran the last couple of miles on a broken toe. It took eight weeks to heal and completely ruined my autumn running plans. Que sera, sera. Went to Nashville - My wife and I went away to Nashville for four days without the kids. This is the longest trip we have taken since our first child was born (over a decade ago) and it was amazing. Now that the kids are all 5 and up, I expect we’ll make this a more regular thing. Conclusion 2021 was a helluva year. Professionally, I was able to narrow my focus on what I truly care about. That increased focus led to an increase in earnings and an increase in satisfaction. Moving into 2022, I plan to continue creating educational technical content. I also plan to keep experimenting with new mediums and pursue new opportunities. The status quo is something I will never accept. There’s a drive deep within me to continue exploring and changing, which serves me well in an industry that seems to have the same restless and relentless drive.
|
||
`,summary:"Entering 2021, I think we all were happy to say adios to 2020 and harbor a little hope for the coming year. After all, we had several viable vaccines about to roll out, meaning this worldwide pandemic might just end up in our rearview mirror. Oh what sweet summer children we were. Maybe 2021 didn’t work out the way we all hoped, but that doesn’t mean nothing positive happened. Amongst all the turmoil and bad takes were some gems of genuine kindness.",date:"3 Jan, 2022",url:"https://nedinthecloud.com/2022/01/03/2021-year-in-review/",image:"2021-Year-in-Review.png",readingTime:"14"},"https://nedinthecloud.com/2021/12/23/terraform-apply-when-external-change-happens./":{title:"Terraform Apply? When External Change Happens.",tags:["hashicorp-terraform","hashicorp-terraform-tutorial"],content:`In an ideal world, all changes to your infrastructure would be managed through Terraform. In reality, that doesn’t always happen. Which can leave you in a bit of a conundrum regarding your next terraform apply. What will happen to the changes made outside of Terraform? Does it matter whether you saved the last terraform plan output? Will Terraform even see the changes made to the infrastructure? Will it overwrite those changes? And what can you do to prevent future occurrences? I set out to discover these answers and more.
|
||
Terraform Plan and Apply Basics Before we go through the various scenarios out there, I first want to go back to basics and consider what Terraform is doing when it runs plan or apply. First we’ll start with plan.
|
||
Terraform Plan Assuming you run terraform plan without any additional arguments, the Terraform executable is going to do the following:
|
||
Look for state data Refresh the state data Compare the state data to the code Find necessary changes to make state data match the code Print the proposed changes out to the terminal window Note the first two steps are to check for state data and refresh it. Terraform wants to work with the newest information available from the target environment. Each resource and data source has two unique attributes: the address used in the code and the unique ID in the target environment. For instance, an Amazon VPC might have the address aws_vpc.my_vpc in the Terraform code, and the unique ID of vpc-12345 in AWS. The state data is simply a mapping of the address to the unique ID, and a list of attributes associated with the resource at that unique ID. In the case of the VPC, the attributes would include things like the cidr_block, owner_id, and main_route_table_id. The full list of attributes for a given resource or data source can be found in the Terraform documentation.
|
||
Within your Terraform code, you can set some of the attribute values through arguments. The arguments can be required or optional, for instance the cidr_block argument is optional for a VPC. Any attributes that are managed by AWS are not available as arguments, the owner_id for example is not a value you can change.
|
||
Terraform plan first checks to see if the target resource exists, e.g. does the VPC vpc-12345 already exist? If the resource doesn’t exist, Terraform will plan to create it. If the resource exists, plan will check and see if the attribute values of the target resource match the argument values in the Terraform code. If they match, awesome, no change needed. If they don’t, Terraform will log a change required. Some attributes can only be changed by recreating the resource, in which case Terraform plans to destroy the resource recreate it.
|
||
When you run terraform plan, you can decide to save the planned changes to a file, usually with the file extension tfplan. That tfplan file can be used with terraform apply to execute the changes.
|
||
I realize this might seem rudimentary for some of you, but it is very important when trying to understand what Terraform will do on apply.
|
||
Terraform Apply Running plan doesn’t make any changes to the target infrastructure, that is the province of terraform apply. If you run apply without any other arguments, Terraform will do the following:
|
||
Look for state data Refresh the state data Compare the state data to the code Find necessary changes to make state data match the code Print the proposed changes out to the terminal window Ask for you to confirm the changes Make the changes if you approve them Update the stata data Wow, that looks awfully similar to terraform plan doesn’t it? You bet your sweet bippy it does. Unless you tell it otherwise, Terraform runs through the full plan process again, including refreshing the state data, and asks at the end if you want to make the changes. If you approve the changes, Terraform will execute the changes on the target environment.
|
||
There are two ways to skip the prompt in terraform apply. The first is to pass the argument -auto-approve which simply answers “yes” for you at the prompt. All the steps happen in the same order, you just don’t need to be involved. The second way is to pass terraform apply a tfplan file saved from a previous plan. The steps will now change to this:
|
||
Load the tfplan file Compare the state data markers Make the changes to the target environment Update the state data Before Terraform executes the change, it does a sanity check to make sure the state data hasn’t changed since the tfplan was generated. If the state data has changed, then the plan is no longer valid and you need to run a new one. Otherwise, the plan is valid and Terraform goes on it’s merry way executing the changes detailed in the plan.
|
||
Changes Outside of Terraform Any resources defined in Terraform code can be referred to as managed by Terraform. Changes to those resources should always happen through updating the code and applying it using Terraform. But that is not always the world we live in. Sometimes you will discover changes have been made outside of the Terraform workflow. Let’s think through those changes and figure out what Terraform is going to do. I have included some examples to enable you to try this out on your own.
|
||
Changes Made to Non-managed Resources We’ll start by first considering what happens if someone alters the target environment by creating a resource Terraform isn’t managing. For instance, let’s say Terraform has deployed a VPC and subnet like so:
|
||
resource "aws_vpc" "vpc" { cidr_block = var.vpc_cidr_block enable_dns_hostnames = var.enable_dns_hostnames tags = { Name = "Taconet" } } resource "aws_subnet" "subnet1" { cidr_block = var.vpc_subnet1_cidr_block vpc_id = aws_vpc.vpc.id tags = { Name = "Taconet-sub1" } } Another admin comes along and uses the AWS CLI to add second subnet to the VPC:
|
||
vpc_id=$(aws ec2 describe-vpcs --filters Name="tag:Name",Values="Taconet" \\ --query 'Vpcs[0].VpcId' --output text) aws ec2 ec2 create-subnet --cidr-block "10.0.10.0/24" --vpc-id $vpc_id Bad admin!
|
||
What will Terraform do? The answer is nothing. That’s right. Nothing.
|
||
The new subnet in the VPC is not a resource being managed by Terraform. Although the subnet is making use of the VPC created by Terraform, the actual subnet is not managed, so Terraform doesn’t care.
|
||
You can import the subnet by adding it as a resource to the code and running the terraform import command, but that’s a whole other story.
|
||
As far as Terraform is concerned, all of the infrastructure it is managing has stayed the same. No changes are necessary.
|
||
Changes made to Non-managed Attributes What if someone makes a change to a managed resource, but the change is on an attribute that we aren’t managing in code? Remember all those optional arguments for a resource? Well, what if someone sets the value of an attribute and we have not supplied a value through an argument in the resource block?
|
||
That might sound a bit confusing, so why don’t we take a look at an example? Again, we will start with the same VPC and subnet configuration.
|
||
resource "aws_vpc" "vpc" { cidr_block = var.vpc_cidr_block enable_dns_hostnames = var.enable_dns_hostnames tags = { Name = "Taconet" } } resource "aws_subnet" "subnet1" { cidr_block = var.vpc_subnet1_cidr_block vpc_id = aws_vpc.vpc.id tags = { Name = "Taconet-sub1" } } Among the list of arguments we can set for the subnet is map_public_ip_on_launch. The argument is optional and has a default value of false. Looking at the attributes of our subnet, we can see that attribute is set to false.
|
||
$> terraform state show aws_subnet.subnet1 # aws_subnet.subnet1: resource "aws_subnet" "subnet1" { id = "subnet-0a0801172b8a3a44f" map_public_ip_on_launch = false vpc_id = "vpc-04469f3e99db93865" } Let’s change that value using the AWS CLI:
|
||
aws ec2 modify-subnet-attribute --subnet-id subnet-0a0801172b8a3a44f --map-public-ip-on-launch Next we’ll run terraform plan to see what Terraform thinks of this situation. I’ll truncate the output to include only the juicy bits.
|
||
$> terraform plan aws_subnet.subnet1: Refreshing state... [id=subnet-0a0801172b8a3a44f] Note: Objects have changed outside of Terraform Terraform detected the following changes made outside of Terraform since the last "terraform apply": # aws_subnet.subnet1 has changed ~ resource "aws_subnet" "subnet1" { id = "subnet-0a0801172b8a3a44f" ~ map_public_ip_on_launch = false -> true + tags = {} # (9 unchanged attributes hidden) } Terraform will perform the following actions: # aws_subnet.subnet1 will be updated in-place ~ resource "aws_subnet" "subnet1" { id = "subnet-0a0801172b8a3a44f" ~ map_public_ip_on_launch = true -> false tags = {} # (9 unchanged attributes hidden) } Plan: 0 to add, 1 to change, 0 to destroy. Even though we haven’t configured the map_public_ip_on_launch attribute with an argument, Terraform wants to use the default value of false. Since someone changed it to true and we didn’t update our code, Terraform will overwrite the value to false.
|
||
What about an optional argument that doesn’t have a default value? For that, we will turn our gaze to an EC2 instance.
|
||
Here is the code for our EC2 instance:
|
||
resource "aws_instance" "nginx1" { ami = nonsensitive(data.aws_ssm_parameter.ami.value) instance_type = var.instance_type subnet_id = aws_subnet.subnet1.id vpc_security_group_ids = [aws_security_group.nginx-sg.id] tags = { Name = "instance-1" } } There is an optional argument for the aws_instance resource called secondary_private_ips with no default value. If you don’t want secondary private IP addresses for your instance, you simply don’t configure the argument. We can confirm this by looking at the state data for the instance:
|
||
$> terraform state show aws_instance.nginx1 # aws_instance.nginx1: resource "aws_instance" "nginx1" { id = "i-08f30e9ef3efe1028" instance_type = "t2.micro" primary_network_interface_id = "eni-099d0bb465273d233" private_dns = "ip-10-0-0-244.ec2.internal" private_ip = "10.0.0.244" secondary_private_ips = [] } There are no secondary private IP addresses set. Now let’s set one using the AWS CLI:
|
||
aws ec2 assign-private-ip-addresses --network-interface-id eni-099d0bb465273d233 --private-ip-addresses=10.0.0.111 Once that completes, we can run terraform plan to see what will happen. The output below has been truncated to include the juicy bits only.
|
||
$> terraform plan Note: Objects have changed outside of Terraform Terraform detected the following changes made outside of Terraform since the last "terraform apply": # aws_instance.nginx1 has changed ~ resource "aws_instance" "nginx1" { id = "i-08f30e9ef3efe1028" ~ secondary_private_ips = [ + "10.0.0.111", ] tags = { "Name" = "instance-1" } Unless you have made equivalent changes to your configuration, or ignored the relevant attributes using ignore_changes, the following plan may include actions to undo or respond to these changes. ─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── No changes. Your infrastructure matches the configuration. Well, how about that? Terraform sees the change, but doesn’t want to do anything about it. We aren’t managing the secondary_private_ips attribute, and there’s no default setting, so Terraform doesn’t want to do anything.
|
||
However, we will keep getting the change message until we run a terraform apply. If we are fine with the change, and just want to get our state data to match it, we can run terraform apply -refresh-only:
|
||
$> terraform apply -refresh-only Note: Objects have changed outside of Terraform Terraform detected the following changes made outside of Terraform since the last "terraform apply": # aws_instance.nginx1 has changed ~ resource "aws_instance" "nginx1" { id = "i-08f30e9ef3efe1028" ~ secondary_private_ips = [ + "10.0.0.111", ] tags = { "Name" = "instance-1" } # (26 unchanged attributes hidden) # (5 unchanged blocks hidden) } This is a refresh-only plan, so Terraform will not take any actions to undo these. If you were expecting these changes then you can apply this plan to record the updated values in the Terraform state without changing any remote objects. Would you like to update the Terraform state to reflect these detected changes? Terraform will write these changes to the state without modifying any real infrastructure. There is no undo. Only 'yes' will be accepted to confirm. Enter a value: yes Apply complete! Resources: 0 added, 0 changed, 0 destroyed. Our state data is up to date with the private IP address we set through the command line. If we want to override what was set, we can simply add a secondary_private_ips block to our code with the proper values and Terraform will overwrite it.
|
||
Changes made after a plan So far we have seen what happens when a change is made outside of Terraform and then we run a plan. But what happens if you run terraform plan and while you are reviewing the output, somebody makes a change to the infrastructure? Let’s think through that setup.
|
||
As we saw earlier, when you run a terraform apply without a plan file Terraform will run a refresh against the target environment and then run a fresh plan before making any changes. If you ran terraform plan and then someone made a change before you ran terraform apply, then Terraform would simply run a fresh plan and present you with the changes it wants to make. If the external change was made on a managed attribute of a managed resource, Terraform would overwrite the change with whatever is defined in the code.
|
||
What about a tfplan file? If you ran terraform plan and saved the output to a tfplan file, and then someone made a change before you ran terraform apply with the tfplan file, what would Terraform do? Running apply with a tfplan file means that Terraform will only make the changes defined in the plan. If the external change was made on a managed attribute of a managed resource, and Terraform is not about to update that resource, Terraform will not overwrite the change.
|
||
That seems important to point out:
|
||
If Terraform wasn’t planning to update the altered attribute, it will not overwrite the change during apply.
|
||
Of course, if you run a new plan later, Terraform will pick up on the altered attribute and make a plan to overwrite it.
|
||
What if you the external change was on a managed attribute of a managed resource, and you were changing that value in the plan? Woah, that’s a lot. Let’s go with an example.
|
||
The instance type for our aws_instance.nginx1 is set with the variable var.instance_type. The default value is t2.micro, which is what our current state shows:
|
||
$> terraform state show aws_instance.nginx1 # aws_instance.nginx1: resource "aws_instance" "nginx1" { ... id = "i-08f30e9ef3efe1028" instance_type = "t2.micro" } Let’s run a plan to change the instance size to t2.medium and save it to a file:
|
||
terraform plan -var instance_type="t2.medium" -out medium.tfplan Next we will change the current instance type to t2.small using the AWS CLI:
|
||
aws ec2 stop-instances --instance-ids i-08f30e9ef3efe1028 aws ec2 modify-instance-attribute --instance-type t2.small --instance-id i-08f30e9ef3efe1028 aws ec2 start-instances --instance-ids i-08f30e9ef3efe1028 Now we can run our apply with the tfplan file.
|
||
terraform apply medium.tfplan The result is that Terraform updates the instance type to be t2.medium. It simply overwrites the attribute. Whew, that’s a relief!
|
||
What have we learned? There are many different ways to approach the question of changes made outside of Terraform. Let’s address the basics. If a resource managed by Terraform is altered outside of Terraform, then Terraform will attempt to overwrite those changes back to the intended state. How does it detect the changes? By refreshing the state. When does it refresh the state? When you run a plan or an apply without a plan file.
|
||
If someone adds something to the infrastructure environment and it is not managed by Terraform, then Terraform will happily ignore it. The import workflow is there to move resources under Terraform’s management.
|
||
If a change is made to a managed resource and Terraform performs a refresh, it will detect the change. If the change is to a managed attribute, Terraform will attempt to overwrite it. If the change is detected during plan, Terraform will plan to overwrite the change. If the change is detected during apply, Terraform will make the actual change to overwrite it.
|
||
If a change is made to a managed resource and Terraform does not perform a refresh, it will not detect the change. When would this happen? Terraform plan can save it’s proposed changes to a file, and that file can be passed to Terraform apply for execution. When a plan file is passed to apply, Terraform does not perform a refresh. This means that changes made to managed resources between the plan and apply actions will not be detected! Unless the change stored in Terraform plan overwrites the change made outside of Terraform, the external change will remain.
|
||
I admit that all of this is pretty confusing, so I made this nice chart for you!
|
||
Conclusion The question of what happens when a change is made outside of Terraform is a difficult question, and the answer is going to depend on several factors. I think the key takeaway is to avoid making any changes outside of Terraform.
|
||
`,summary:"In an ideal world, all changes to your infrastructure would be managed through Terraform. In reality, that doesn’t always happen. Which can leave you in a bit of a conundrum regarding your next terraform apply. What will happen to the changes made outside of Terraform? Does it matter whether you saved the last terraform plan output? Will Terraform even see the changes made to the infrastructure? Will it overwrite those changes?",date:"23 Dec, 2021",url:"https://nedinthecloud.com/2021/12/23/terraform-apply-when-external-change-happens./",image:"2021-12-23.png",readingTime:"13"},"https://nedinthecloud.com/2021/12/14/using-the-moved-block-in-terraform-1.1/":{title:"Using the Moved Block in Terraform 1.1",tags:["hashicorp","hashicorp-terraform","terraform","terraform-tutorials"],content:`The release of Terraform 1.1 has brought with it a new configuration block type called moved. This is a super-cool new block that helps with when you want to refactor your Terraform code without breaking production. There are two primary use cases for the moved block. The first is to refactor versioned modules you have published in a directory. The second, and the one we’ll focus on in this post, is refactoring your code by renaming resources, adding loops, and moving resources into modules.
|
||
https://youtu.be/fDPB7xbckVM
|
||
Essentially, the idea is that you have an existing deployment using your Terraform code. Your code has changed and grown over time, as the needs of your infrastructure have changed and grown. Now you want to update the code with better organization, more efficiency, and reusable components. You might want to change the name field of a resources to be more descriptive, or condense mutliple instances of a resources with a for_each loop. Maybe you want to take a logical grouping of components and turn it into a module.That seems like a reasonable thing to do right?
|
||
You will quickly discover that Terraform doesn’t understand that you have updated the resource address for existing resources, and so it thinks you want to destroy your existing resources and create new ones. That’s not a problem when you’re using an ephemeral environment that is torn down and rebuilt regularly, but it’s markedly less awesome when you accidentally delete all subnets in production and respawn them. People get mad about that kind of thing.
|
||
The way to deal with this prior to the moved block was to use the command terraform state mv to move resources to a new location in the state file. As we all know, mucking around with the state file is fraught with peril. The introduction of the moved block lets you be more deliberate with resource address changes, and also enables you to document changes in code for those who might be using your Terraform code as a module.
|
||
That’s enough theory, why don’t we ground this with some examples?
|
||
Moving resources into a module Let’s say I have Terraform code that defines an AWS VPC, including a subnet, route, and internet gateway:
|
||
resource "aws_vpc" "vpc" {} resource "aws_subnet" "subnet" {} resource "aws_route" "default_route" {} resource "aws_internet_gateway" "igw" {} When I use the code to deploy the VPC to my AWS account, Terraform creates the infrastructure and saves the environment information in state data. The address for my VPC will be aws_vpc.vpc and it will map to the id of the VPC in my AWS account vpc-12345. Just to be clear, the address information is totally internal to Terraform and the state data. It does not impact the resources in AWS.
|
||
Now let’s say I want to create a VPC module to handle the networking for this and other configurations. I can update my code to this:
|
||
module "vpc" { source = "./vpc_module"} And move my networking resources to the module. But what happens when I want to apply this new code to my existing deployment? Running a Terraform plan tells me the following:
|
||
Plan: 4 to add, 0 to change, 4 to destroy. 4 to destroy?! That’s not what I want! If I look at one of the resources being destroyed:
|
||
aws_vpc.vpc will be destroyed As far as Terraform is concerned, the VPC resources in state data with the address aws_subnet.vpc was removed from the code, and so the corresponding VPC in AWS with the ID vpc-12345 should also be removed, i.e. destroyed.
|
||
In the resources being created, I can see:
|
||
module.vpc.aws_vpc.vpc will be created The address for the VPC is now module.vpc.aws_vpc.vpc. Terraform has no way of knowing that the VPC at address module.vpc.aws_vpc.vpc meant to refer to vpc-12345. As far as Terraform is concerned, this is a brand new VPC we’re creating.
|
||
Using terraform state mv Prior to Terraform 1.1, the way to deal with this problem was to use the terraform state mv subcommand to update the state data with the correct mapping. For instance, we could run the following command:
|
||
terraform state mv aws_vpc.vpc module.vpc.aws_vpc.vpc Essentially, this updates the underlying state data with a new mapping of module.vpc.aws_vpc.vpc to vpc-12345. When Terraform runs a plan, it won’t think any changes are necessary for the VPC resource. We would have to repeat the process for the remaining resources being destroyed to ensure nothing in our target environment is actually changed.
|
||
terraform state mv aws_subnet.subnet module.vpc.aws_subnet.subnets[0] terraform state mv aws_internet_gateway.igw module.vpc.aws_internet_gateway.igw terraform state mv aws_route.default_route module.vpc.aws_route.default_route Now when I run a Terraform plan, I get the following output.
|
||
No changes. Your infrastructure matches the configuration. Excellent! The process worked successfully. Unfortunately, the changes happened in an imperative way and are not documented by the code. That means if we were running this through an automation pipeline, we would have to make the state updates manually and then kick off the pipeline. That’s less than ideal. Worse, if we were using workspaces, we’d have to repeat the process for every workspace. Plus, if anyone else happens to be using this code and doesn’t know about the change, it will break their infrastructure. Terraform 1.1 introduces a better way.
|
||
Using a moved block Instead of using the terraform state mv commands, we can instead use the declarative moved block to express where our resources have moved to. In the code we can add the following blocks:
|
||
moved { from = aws_vpc.vpc to = module.vpc.aws_vpc.vpc } moved { from = aws_internet_gateway.igw to = module.vpc.aws_internet_gateway.igw } moved { from = aws_route.default_route to = module.vpc.aws_route.default_route } moved { from = aws_subnet.subnet to = module.vpc.aws_subnet.subnets[0] } If we add the above blocks and run terraform plan, we will get the following output:
|
||
Plan: 0 to add, 0 to change, 0 to destroy. And for each resource we will see something like this:
|
||
aws_vpc.vpc has moved to module.vpc.aws_vpc.vpc Terraform has not yet made any changes; that’s kind of the point in running plan first. Terraform is simply letting us know what changes it will make on apply. Peaking in the state data, the resources have not been updated yet.
|
||
$> terraform state list data.aws_availability_zones.available aws_internet_gateway.igw aws_route.default_route aws_subnet.subnet aws_vpc.vpc Running apply will result in the following:
|
||
Apply complete! Resources: 0 added, 0 changed, 0 destroyed. Sweet! And if we check our state data again, we’ll see that it has been updated.
|
||
$> terraform state list data.aws_availability_zones.available module.vpc.aws_internet_gateway.igw module.vpc.aws_route.default_route module.vpc.aws_subnet.subnets[0] module.vpc.aws_vpc.vpc We were able to refactor our code, keep everything declarative, and verify there was no impact to the target environment before applying the change. Because we kept it all in code, the process could be handled through our regular automation process instead of manually messing with state data. There is also a clear trail of what changes were made and when.
|
||
After the change has been applied, the moved blocks can be removed from code, but HashiCorp recommends that you leave them in until you are certain no one is using the older version of the code. There is no harm to leave the blocks in, so I tend to agree with their assessment.
|
||
Renaming a resource The same process shown above can be used to simply rename a resource in a configuration. Let’s say we have a resource like this:
|
||
resource "aws_subnet" "subnet" {} And we want to change the resource to create multiple subnets using the count meta-argument.
|
||
locals { subnets = ["192.168.0.0/24", "192.168.1.0/24", "192.168.2.0/24"] } resource "aws_subnet" "subnets" { count = length(local.subnets) } Just like the previous example, we could use the terraform state mv command to update our state data to change the address of the subnet from aws_subnet.subnet to aws_subnet_subnets[0].
|
||
Or we could add a moved block like this:
|
||
moved { from = aws_subnet.subnet to = aws_subnet_subnets[0] } And accomplish the same goal without resorting to manual, imperative processes. Plus, we have a record of the change in case it impacts anything tied to our code.
|
||
Important Caveat! The moved block doesn’t solve all your problems. There are a few caveats to keep in mind.
|
||
External module packages UPDATE - Terraform 1.3 now allows moving resources to an external module! Check out my YouTube video for more information.
|
||
Let’s say that instead of moving our resources into a VPC module we wrote, instead we wanted to use the VPC module from the Terraform registry. Unfortunately, that’s a no-go for the moved block. Trying to do so will result in the following error:
|
||
│ Error: Cross-package move statement │ │ on main.tf line 65: │ 65: moved { │ │ This statement declares a move to an object declared in external module package "registry.terraform.io/terraform-aws-modules/vpc/aws". Move statements can be only within a single module package That’s correct friends, moved is limited to refactoring for a single module package. You can move resources into child modules that reside in the same directory structure as the root module, but you can’t move resources to a module in an external location. I’m guessing that might change with future releases, but it’s an important point to bear in mind.
|
||
Why is this the case? I haven’t heard this directly from HashiCorp, but I think it has to do with module refactoring. The other main use case for the moved block is refactoring versioned modules available on a registry. I think there is concern that refactoring a module on a registry with moved blocks that point to different external module package could create too many weird dependencies. You don’t have direct control over the external module and addressing in that module, which means they could use moved blocks in their code which moves resources referenced in your moved blocks and chaos ensues.
|
||
If you plan to migrate to an external module package for resources in your code, you’ll have to stick with terraform state mv for now.
|
||
A note on using sets in for_each Another thing I came across while working on the demo for this is updating a resource block to use a for_each meta-argument with a set of strings. Let’s consider the following:
|
||
locals { subnet_cidr = "192.168.0.0/24" } resource "aws_subnet" "subnet" { cidr_block = local.subnet_cidr } What if you want to update the code to define the same subnet along with several others using a for_each loop and a map or subnets?
|
||
locals { subnets = { subnet1 = "192.168.0.0/24" subnet2 = "192.168.1.0/24" subnet3 = "192.168.2.0/24" } } resource "aws_subnet" "subnet" { for_each = local.subnets cidr_block = each.value } Not a problem. You can easily add the moved block for the original subnet like this:
|
||
moved { from = aws_subnet.subnet to = aws_subnet.subnet["subnet1"] } What if you wanted to use a list instead of a map for your values?
|
||
locals { subnets = ["192.168.0.0/24", "192.168.1.0/24", "192.168.2.0/24"]} resource "aws_subnet" "subnet" { for_each = toset(local.subnets) cidr_block = each.value} Well, now we have a problem. The data type submitted to the for_each argument is a set, not a list, so we cannot refer to the resources created by an element index, i.e. aws_subnet.subnet[0] is not valid. What are we going to use in our moved block to refer to the original subnet?
|
||
The answer lies in what data type is created by a resource with a for_each meta-argument. The data type is a object, and the keys of the object are based on whether the data type submitted was a set or a map. If it’s a map, the keys of the object are set to the keys of the map. If it’s a set, the keys of the object are set to the values of the set.
|
||
Our moved block would look like this:
|
||
moved { from = aws_subnet.subnet to = aws_subnet.subnet["192.168.0.0/24"] } Problem solved.
|
||
If you’ve ever wondered why the for_each meta-argument only accepts strings as the value in the set, this is why. Terraform uses those strings as keys in the object generated. If you tried to give it a set of maps or a set of lists, what would it use for the keys? It is also why the for_each uses a set and not a list for values. The set data type is by definition a collection of unique values. The toset() function will remove any duplicates in the list and return only the unique elements as a set.
|
||
Wrap-up The moved block is going to be hugely useful for situations where you want to refactor your code in a way that supports versioning and automation. It isn’t going to solve every situation, like if you’re moving to an external module package. For those cases, you can always use terraform state mv to manipulate the state data.
|
||
moved blocks also help tremendously when you are using workspaces for Terraform. Imagine the problem of using terraform state mv for each workspace under management by Terraform, versus the simplicity of using a moved block in your code. I definitely prefer the latter.
|
||
Once you have added moved blocks to your code, you should leave them in place as long as anyone is running the older version of the code somewhere. I could see creating a moved.tf file in your code specifically to track the moved blocks. In addition, I’d recommend adding some comments, a date, and maybe even a code commit hash to know when each moved block was added and why.
|
||
`,summary:"The release of Terraform 1.1 has brought with it a new configuration block type called moved. This is a super-cool new block that helps with when you want to refactor your Terraform code without breaking production. There are two primary use cases for the moved block. The first is to refactor versioned modules you have published in a directory. The second, and the one we’ll focus on in this post, is refactoring your code by renaming resources, adding loops, and moving resources into modules.",date:"14 Dec, 2021",url:"https://nedinthecloud.com/2021/12/14/using-the-moved-block-in-terraform-1.1/",image:"Untitled-design3.png",readingTime:"11"},"https://nedinthecloud.com/2021/12/08/github-actions-with-terraform/":{title:"GitHub Actions with Terraform",tags:["github-actions","microsoft-azure","terraform"],content:"Recently, I was a guest on the Azure DevOps Lab YouTube channel, talking about using GitHub Actions with Terraform to deploy infrastructure on Azure. April Edwards was a gracious host and let me ramble on for 10+ minutes about the very basics of GitHub Actions. Due to the short format of DevOps Lab videos, I wasn’t able to really dig into specifics around the GitHub Actions file. I thought I would write a blog post to fill in the gaps, and so here we are.\nhttps://youtu.be/QcBtWX72dRw\nDevOps Lab on Getting Started with GitHub Actions and Terraform\nGetting Started As I mentioned in the video, there’s a lot of scary sounding words for your average infrastructure admin or cloud newbie. Five years ago, you could get by as a Sysadmin without knowing what GitHub, Terraform, and GitOps were. Now? Not so much. Before I dig too deep into the tech, I’d first like to put your potentially troubled mind at ease. If you’re a seasoned DevOps veteran, you can probably skip to the next section.\nLet’s start with GitHub and GitHub Actions. The venerable GitHub is a place for you to store version controlled repositories of code. If you’re a more traditional Sysadmin, you might not have much use for code repositories, since you don’t have code per se. But as you move into the world of Infrastructure as Code, now you do have code you’d like to store somewhere. GitHub is one of several excellent options.\nGitHub Actions is simply a workflow engine. It’s a set of instructions for GitHub to execute when an event occurs, like when you push new code to the repository. It can be used for Continuous Integration and Delivery (there I go use DevOps’y terms again), but you can really have it listen for any type of event and take just about any kind of action that is supported by an API. We joked in the video that you could have GitHub Actions order you a pizza from Dominos using Terraform. But that is a real thing you could do! Not especially useful for standing up infrastructure perhaps, but it’s a fun example.\nOkay, so GitHub is a place to store your code. And GitHub Actions is a workflow engine that can do something with your code. What type of code might you be working with? Terraform code of course! Terraform is an application from HashiCorp that automates the deployment and management of infrastructure. Terraform code is expressed in either JSON or HashiCorp Configuration Language, and it is evaluated and executed by the Terraform binary. Terraform has a general workflow of initialize, plan, and apply. Initialize prepares the code to be executed by downloading providers and modules and setting up the state data backend. Plan evaluates the code against the target environment and shows you a list of changes it would make to have the target environment match the desired state expressed by the code. Apply makes the changes outlined by plan.\nWait. Did someone say workflow? Sweet. We can use GitHub Actions to execute the standard Terraform workflow. But when should the actions kick off? That brings us to the final term, GitOps.\nGitOps is a extension of DevOps, and it is premised on the idea that your workflow should be driven by events in your Git-based repository. Git is the magic behind GitHub and most other source control repository systems.\nWithout getting too deep into the world of Git, we can think of the code we are working with on our workstation as local, and the code stored up on GitHub as remote. Locally, I could be working on some updates for my Terraform code. When I think it’s in a good state, I will commit the change to source control locally, and then push that change to remote. That push is one possible event to listen for on the GitHub actions side.\nWhen developers - like you and me? - are working on a new feature or big change, the common practice in Git is to create a new branch based on the main branch. Once we’ve pushed our new branch to remote and tested things out, we can create a pull request to merge our branch into main. That’s another event we want to listen for!\nFinally, once we’ve decided to accept the pull request, we merge the code into the main branch. A merge is really just a push from one branch to another, which means we have to listen for push events on the main branch.\nI think that covers the background, and hopefully I’ve demystified some of these fancy terms. If not, there are literally whole books dedicated to each topic. Go forth and consume!\nGitHub Actions File GitHub Actions looks for YAML files in the directory .github/workflows. The contents of those files determine what events to listen for and what actions to take. The actions are broken up into jobs that are assigned to a runner, typically hosted by GitHub. Each job is made up of steps for the runner to execute.\nTo summarize from the previous section, we are looking for three different events: A push on non-main branches, a pull request against the main branch, and a push on the main branch. What might we want to do for each event? Pushes on non-main branches could be checked to ensure initialization works and that the code is valid and properly formatted. A pull request against main means that code could be ready for production roll-out, so we probably want to run a plan to see what changes are being made. And then the merge to main would mean the code looks good and we want to use it, meaning we should run an apply. With that in mind, let’s check out the beginning of our GitHub Actions file.\nYou can find the file, as well as the rest of the demo on the ado-labs-github-actions GitHub repository.\nname: 'Terraform' on: [push, pull_request] env: TF_LOG: INFO We have named the action ‘Terraform’ and let GitHub know that it should run this action when there is a push or pull_request event on the repository. We can also set global environment variables for all jobs, which in our case we are setting the level of Terraform logging to INFO.\nNext up we can create our jobs.\njobs: terraform: name: 'Terraform' runs-on: ubuntu-latest # Set the working directory to main for the config files defaults: run: shell: bash working-directory: ./main We are creating a single job here and running it on a GitHub hosted runner using the latest Ubuntu flavor runner machine. We are also going to set some defaults for all the steps in the job, setting the shell to bash and the working directory to ./main. That is where out Terraform code lives.\nWith the defaults established, we are now going to create the steps. First up, we have to do a little prep work:\nsteps: # Checkout the repository to the GitHub Actions runner - name: Checkout uses: actions/checkout@v2 # Install the preferred version of Terraform CLI - name: Setup Terraform uses: hashicorp/setup-terraform@v1 with: terraform_version: 1.0.10 The Checkout step performs a checkout of the code in our repository so the runner can do it’s thing. The Setup Terraform step installs the Terraform binary on the runner, allowing us to run Terraform commands against our code.\nWith our code and Terraform available on the runner machine, we can start the standard Terraform workflow. The first step will be Terraform Init, which needs to run regardless of whether we are in the plan phase or the apply phase.\n- name: Terraform Init id: init env: ARM_CLIENT_ID: ${{ secrets.ARM_CLIENT_ID }} ARM_CLIENT_SECRET: ${{ secrets.ARM_CLIENT_SECRET }} ARM_TENANT_ID: ${{ secrets.ARM_TENANT_ID }} ARM_SUBSCRIPTION_ID: ${{ secrets.ARM_SUBSCRIPTION_ID }} RESOURCE_GROUP: ${{ secrets.RESOURCE_GROUP }} STORAGE_ACCOUNT: ${{ secrets.STORAGE_ACCOUNT }} CONTAINER_NAME: ${{ secrets.CONTAINER_NAME }} run: terraform init -backend-config="storage_account_name=$STORAGE_ACCOUNT" -backend-config="container_name=$CONTAINER_NAME" -backend-config="resource_group_name=$RESOURCE_GROUP" Terraform needs to download the providers and modules used by our code and prepare the state data backend. We are running the terraform init command along with a bunch of arguments that establish where the state data will live.\nThe state data backend is using azurerm and we need to supply credentials to access the Azure storage account. The values for the backend and Azure credentials can be passed using environment variables, which is why we have an env block that creates those environment variables. The environment variables starting with ARM are the Azure credentials, which Terraform knows to check for when the azurerm backend or provider is used.\nWhat about the values? Where are those coming from? Each value is stored as a Secret in the GitHub repository:\nAnd referenced by using the syntax ${{ secrets.NAME }}. The runner will load those values dynamically and store them in environment variables, meaning the values are never written to disk on the runner or stored in the logs. Given the sensitive nature of our Azure credentials, that’s probably a good thing!\nMoving to the Terraform Plan step, we are going to run this step only if the event is a pull request.\n- name: Terraform Plan id: plan env: ARM_CLIENT_ID: ${{ secrets.ARM_CLIENT_ID }} ARM_CLIENT_SECRET: ${{ secrets.ARM_CLIENT_SECRET }} ARM_TENANT_ID: ${{ secrets.ARM_TENANT_ID }} ARM_SUBSCRIPTION_ID: ${{ secrets.ARM_SUBSCRIPTION_ID }} if: github.event_name == 'pull_request' run: terraform plan -no-color Once again we’ll need those Azure credentials, this time for the azurerm provider used in our code, so we are loading the required environment variables. Next, we have an if statement that checks to see if this is a pull request event. If that statement evaluates to true, the step will execute and run terraform plan.\nThe point of running terraform plan is to see what changes Terraform would make to our target environment. The output of the plan command will be in the runner logs, but wouldn’t it be nice if we could capture the output and add it to the pull request as a comment? Yup, it sure would be.\n- name: add-plan-comment id: comment uses: actions/github-script@v3 if: github.event_name == 'pull_request' env: PLAN: "terraform\\n${{ steps.plan.outputs.stdout }}" with: github-token: ${{ secrets.GITHUB_TOKEN }} script: | The comment step uses the GitHub script action to run Bash script on the runner as defined by the script block. This step will only run if the event is a pull request, just like our Terraform Plan step. In the environment variable PLAN we are storing the stdout of the Terraform Plan step. That’s right, you have access to the standard output of any previous steps in the workflow.\nIn the Bash script, we are going to post a comment to the pull request, which requires authentication. In the repository there is a pre-existing secret called GITHUB_TOKEN that the runner can grab and use to execute commands against the repository.\nLet’s take a look at the script being run by this step:\nscript: | const output = `#### Terraform Format and Style 🖌\\`${{ steps.fmt.outcome }}\\` #### Terraform Initialization ⚙️\\`${{ steps.init.outcome }}\\` #### Terraform Validation 🤖${{ steps.validate.outputs.stdout }} #### Terraform Plan 📖\\`${{ steps.plan.outcome }}\\` <details><summary>Show Plan</summary> \\`\\`\\`${process.env.PLAN}\\`\\`\\` </details> *Pusher: @${{ github.actor }}, Action: \\`${{ github.event_name }}\\`, Working Directory: \\`${{ env.tf_actions_working_dir }}\\`, Workflow: \\`${{ github.workflow }}\\`*`; github.issues.createComment({ issue_number: context.issue.number, owner: context.repo.owner, repo: context.repo.repo, body: output }) First we have to construct our output for the comment. Basically it will summarize the results of previous steps taken by the script. We didn’t run an fmt or validate step, but if we had, the outcome or output would be included here.\nThe initialization and plan steps are referenced, and the full output of the plan is hidden using a details HTML tag.\nOnce we have constructed our output, the function createComment is invoked to create the comment in the referenced pull request. Fun fact, a pull request is really just a special type of issue, which is why you see the issue_number argument and context.issue.number reference.\nThe code above is rendered like this in the pull request comments:\nThe format and validation steps are blank because we didn’t run those steps. The initialization and plan were both successful, and the full plan is available by expanding Show Plan.\nAnyone reviewing the pull request for approval doesn’t have to dig into logs to find the plan output, they can just expand the comment in the pull request added by this step.\nFinally we can move on to the last step, Terraform Apply, which applies the code if the pull request is approved.\n- name: Terraform Apply if: github.ref == 'refs/heads/main' && github.event_name == 'push' env: ARM_CLIENT_ID: ${{ secrets.ARM_CLIENT_ID }} ARM_CLIENT_SECRET: ${{ secrets.ARM_CLIENT_SECRET }} ARM_TENANT_ID: ${{ secrets.ARM_TENANT_ID }} ARM_SUBSCRIPTION_ID: ${{ secrets.ARM_SUBSCRIPTION_ID }} run: terraform apply -auto-approve Merging a pull request is a push event on the target branch - main in our case, so we have an if statement checking for that event. Below the if statement, we once again establish the credentials needed to access the target environment in Azure, and run a terraform apply with the -auto-approve flag, which skips the usual prompt on an apply.\nOur GitHub Actions file follows a GitOps workflow from the initial push of a feature branch, to a pull request to merge the feature branch into main, to the merge being approved and pushed to the main branch. Along the way the code is initialized, a Terraform plan is run and verified, and the code is applied to the target environment.\nWrap Up The GitHub Actions file used for the demonstration doesn’t have a lot of bells and whistles, but you can certainly add some! At a minimum, you could add in steps to run fmt and validate on a push or a pull request. You could also integrate a tool like Checkov to do static code analysis of your Terraform code. You could even deploy the code to a test environment and run something like Terratest to validate it functions properly. Starting with the basic framework from this demo, the sky is the limit.\n",summary:"Recently, I was a guest on the Azure DevOps Lab YouTube channel, talking about using GitHub Actions with Terraform to deploy infrastructure on Azure. April Edwards was a gracious host and let me ramble on for 10+ minutes about the very basics of GitHub Actions. Due to the short format of DevOps Lab videos, I wasn’t able to really dig into specifics around the GitHub Actions file. I thought I would write a blog post to fill in the gaps, and so here we are.",date:"8 Dec, 2021",url:"https://nedinthecloud.com/2021/12/08/github-actions-with-terraform/",image:"Untitled-design.png",readingTime:"11"},"https://nedinthecloud.com/2021/12/07/data-transformation-in-terraform/":{title:"Data Transformation in Terraform",tags:[],content:`I was attempting to do something I thought was relatively simple, and it broke my brain. The simple thing? Assigning permissions to namespaces in an Azure Kubernetes Service cluster through Azure RBAC using Terraform. Okay, well that part sounds complicated, but it’s not really important what exactly I was trying to do. The important part is that I was trying to do some data transformation in Terraform and the struggle is real. Which got me thinking about Terraform’s place in a workflow and the need for a general purpose programming language.
|
||
The Problem Let’s start with the core problem. Imagine that you have an AKS cluster and you are using Azure RBAC to control permissions on namespaces. That is an actual thing you can do. Azure AD takes care of authentication and authorization, and constructs the resulting permissions for the Kubernetes cluster. Now let’s assume you want to manage the AKS cluster using IaC and Terraform. You can create role assignments using the azurerm_role_assignment resource and pass it the role, scope, and principal to apply it to:
|
||
resource "azurerm_role_assignment" "namespace_admin" { role_definition_name = "Azure Kubernetes Service RBAC Admin" scope = "\${azurerm_kubernetes_cluster.cluster.id}/namespaces/\${var.namespace}" principal_id = var.principal_id } The above resource block grants the principal stored in var.principal_id the role “Azure Kubernetes Service RBAC Admin” on the namespace stored in var.namespace.
|
||
So far, so good. But you probably want to grant permissions to more than one group and more than one namespace. You also want to make it dynamic, meaning you can submit a list of namespaces and admins, and then create those assignments with a loop. What would that structure look like? Well, why don’t we start with what seems like the simplest structure; a map with keys equal to the namespace and values being a list of admins for the namespace.
|
||
namespace_admins = { namespace1 = ["admin1", "admin2", "admin3"] namespace2 = ["admin1", "admin3", "admin4"] namespace3 = ["admin3", "admin5", "admin2"] } You might think you could use this variable value with a for_each meta-argument for the azurerm_role_assignment, but there’s a problem. The for_each resource will run once for each key in the map, but we have three entries for each key. We need to create nine total resources, not three!
|
||
The Solution What we need is a way to convert this data type to a format that will iterate nine times and not three. We can do this with nested for expressions to first expand the map keys, and then the values in each key. The expression looks like this:
|
||
[ for ns, adms in var.namespace_admins : [ for adm in adms : { namespace = ns admin = adm } ] ] The resulting data structure will be a nested tuple with maps as values. We can further apply the flatten() function to get rid of the nested tuple structure, and the end result is this tuple of maps, each representing the inputs for our azurerm_role_assignment resource!
|
||
[ { "admin" = "admin1" "namespace" = "namespace1" }, { "admin" = "admin2" "namespace" = "namespace1" }, { "admin" = "admin3" "namespace" = "namespace1" }, { "admin" = "admin1" "namespace" = "namespace2" }, { "admin" = "admin3" "namespace" = "namespace2" }, { "admin" = "admin4" "namespace" = "namespace2" }, { "admin" = "admin3" "namespace" = "namespace3" }, { "admin" = "admin5" "namespace" = "namespace3" }, { "admin" = "admin2" "namespace" = "namespace3" }, ] The updated structure is no longer going to work with a for_each argument because that only accepts a map or a set of strings and we have a tuple of maps. We could use the toset() function to convert the tuple to a set, but it’s still going to contain maps instead of strings. The solution is to switch to a count argument and use the length() function on the data structure to determine how many map items we need to process. The updated code will look like this:
|
||
locals { namespace_admins = flatten([ for ns, adms in var.namespace_admins : [ for adm in adms : { namespace = ns admin = adm } ] ]) } resource "azurerm_role_assignment" "namespace_admins" { count = length(local.namespace_admins) role_definition_name = "Azure Kubernetes Service RBAC Admin" scope = "\${azurerm_kubernetes_cluster.cluster.id}/namespaces/\${local.namespace_admins[count.index].namespace}" principal_id = local.namespace_admins[count.index].admin } If you’re looking at that code and you don’t find it intuitive, you are not alone. I can’t even take credit for this solution. The truth is that I was banging my head against a wall for the better part of a morning trying to figure out the proper structure to use for my code, and finally I posted the question on the HashiCorp Ambassador Slack. My fellow Ambassadors swooped in to help and within half an hour I had a solution which would work.
|
||
My primary challenge in manipulating data structures in Terraform is that I am trying to apply imperative logic to a declarative format. I’ve spent years writing for loops in imperative languages like PowerShell. For expressions in Terraform, just like count and dynamic blocks, simply don’t work in a way I find intuitive. If I wanted to parse the initial data structure in PowerShell, I would simply use a nested for loop that looks something like this:
|
||
foreach($entry in $namespace_admins){ foreach($admin in $entry.values){ New-AzRoleAssignment -Scope "$($aksClusterId)/namespaces/$($entry.key)" -PrincipalId $admin -Role "Azure Kubernetes Service RBAC Admin" } } Or something similar to that. (I didn’t test the above, so don’t expect it to actually work). For whatever reason, my brain has a much easier time understanding imperative scripting than declarative code, and I don’t think I’m the only one to find this challenging.
|
||
The Need for a General Purpose Programming Language Terraform is awesome for automating infrastructure deployment and management. That’s what it excels at. It was never meant to deal with complex data structures, and I don’t think it should have to. Many of the more recent updates to Terraform applied principles like data types, expanded functions, and complex objects. While that is appreciated, I think it’s also a double-edge sword. The addition of those features makes Terraform more of a full fledged programming language, instead of a representation of infrastructure in code. We’re adding complexity and potentially making our code harder to parse properly. What’s the alternative?
|
||
Consider CloudFormation for a moment, and I am not talking about the engine that parses CloudFormation. I am simply talking about the JSON or YAML templates. AWS has made a deliberate decision to avoid overloading their template language with functions and data structures. If it doesn’t fit into JSON, it doesn’t fly. There are a handful of functions for convenience, but you’d be hard pressed to call it a Domain Specific Language (DSL). It’s closer to a template format with a sprinkling of logic.
|
||
The result is that CloudFormation is extremely verbose, almost to a fault. If you want to do something programmatic in a template, you need to farm that work out to a custom resource that invokes Lambda. I’ve groused about this in the past. Like, why isn’t there a simple function to set a string to all lowercase so my S3 bucket creation doesn’t fail? Why do I need to list out each EC2 instance individually instead of using a count or loop on a single resource? And the answer is that CloudFormation is simple on purpose. Could AWS bake all that stuff in? Sure. Would that add a lot of complexity to the template language and parsing engine? Yup.
|
||
If you look at something like the CDK, which emits CloudFormation templates, the solution is to have a general purpose programming language (GPPL) create the templates as artifacts. All the logic and functions you might want to use in a template are instead evaluated as part of the GPPL code that emits the template. The fact that the template is verbose or lacking functions and complex data structures is beside the point. The template as artifact is not going to be read or edited by a human; rather, it is going to be fed into the CloudFormation engine for processing.
|
||
Terraform is taking a different tack, where the Terraform code is not being generated by another program, but the HCL is instead being written by a human being. Since that is the case, it requires all the bells and whistle of a GPPL, without actually being a GPPL. In a sense, since Terraform is written in Go, the functions and data structures as expressed by HCL are lifted directly from Go. Why not simply use Go as your IaC platform and have the result be valid Terraform code that is ingested by the Terraform executable for provisioning?
|
||
That’s exactly what the Terraform CDK aims to do. You can still leverage all the providers and processing from Terraform, but rather than writing directly in HCL, you can use a programming language you’re already comfortable with. Of course, that assumes you are comfortable with a programming language beyond basic PowerShell and Bash scripting.
|
||
Looking into my crystal ball, I see Terraform as having a long and fruitful life in the IaC space. I also see CDKs becoming more prevalent, with Terraform becoming the backend engine to get the actual provisioning done. Some projects, like Pulumi, are taking it one step further by leveraging some of the Terraform providers while using their own platform to generate and process artifacts.
|
||
Essentially what we have are two diverging paths for Terraform. One is the expansion of HCL to support more complex programming logic and data structures, slowly making it something closer to a GPPL instead of a DSL. The second is the use of CDK to programmatically generate HCL and feed it to Terraform for processing. Which one will win out? Probably neither. There’s always going to be some contingent of Ops folks who don’t want to learn a GPPL, and they will stick with just Terraform. On the other hand, there will be Devs who would rather use their knowledge of common programming languages, along with the attendant bells and whistles of a GPPL.
|
||
For my part, I plan to spend more time learning about the CDK for Terraform and Pulumi in the next year. In the long run, I believe that will be more beneficial to my career than simply doubling down on pure Terraform.
|
||
`,summary:"I was attempting to do something I thought was relatively simple, and it broke my brain. The simple thing? Assigning permissions to namespaces in an Azure Kubernetes Service cluster through Azure RBAC using Terraform. Okay, well that part sounds complicated, but it’s not really important what exactly I was trying to do. The important part is that I was trying to do some data transformation in Terraform and the struggle is real.",date:"7 Dec, 2021",url:"https://nedinthecloud.com/2021/12/07/data-transformation-in-terraform/",image:"terraform-transform.jpg",readingTime:"8"},"https://nedinthecloud.com/2021/11/23/what-i-learned-building-an-azure-devops-pipeline-with-terratest/":{title:"What I Learned Building an Azure DevOps Pipeline with Terratest",tags:["azure","azure-devops-pipelines","hashicorp-terraform","terraform-tutorials"],content:`I am in the process of creating a series of liveProjects for Manning Publishing, one of which involves deploying and managing an Azure Kubernetes Cluster using Terraform and Azure DevOps Pipelines. As I tried to build a CI/CD workflow that uses several pipelines, I ran into a bunch of roadblocks, some of which were my own doing and others which stemmed from a lack of clear documentation. While I won’t give away the contents of the liveProject - that would sort of defeat the point - I did want to call attention to things I stumbled over in hopes of helping others on a similar journey.
|
||
The Basic Premise The basic premise could be a little tricky to grok, so I’ll start there. We are defining an AKS cluster using Terraform code. We can verify the functionality of the cluster by using Terratest to check both the infrastructure and applications running in the cluster. Terratest is pretty cool like that. My goal was to follow a GitOps workflow, where each stage of the overall pipeline matched an event in a GitHub repository. There are three events that should trigger a pipeline - spoiler alert, I didn’t necessarily grasp how GitHub and Azure Pipelines interact.
|
||
Push to a feature branch Pull request to merge feature branch to main Merge feature branch to main My thought process is that any pushes to a feature branch should have the Terraform code checked for formatting and validity. Ideally, developers - yes that includes you infrastructure people too - would properly format and validate their Terraform code before committing and pushing to origin. In reality, probably not so much. Any code pushed to the GitHub repo that isn’t formatted and valid should fail basic Continuous Integration (CI) tests.
|
||
No one should be committing directly to main, so the next step would be to verify the code in a testing environment when a pull request (PR) is created to merge to main. We know the code is formatted and validated, but that doesn’t mean it will produce functional infrastructure. The PR pipeline will spin up a testing instance, validate functionality, and tear it down when complete. We could call this Continuous Delivery (CD), as the code should now be ready for deployment to production once it is merged to main.
|
||
Assuming the code passes muster in the PR, the last step in the workflow is to merge the branch to main and deploy the updated code to a production environment. This could actually be a multistage pipeline where it flows to staging or QA first, but I’m trying to keep things “simple”.
|
||
That’s the basic premise and I created three pipelines to execute it. Let’s examine the three pipelines a bit and walk through issues I ran into.
|
||
Continuous Integration The first pipeline is CI, and it is checking Terraform formatting and validation using terraform fmt and terraform validate. The trigger should be any push on a non-main branch, which is done by simply excluding the main branch in the trigger block of the YAML file:
|
||
trigger: branches: exclude: - main The formatting check is easy and doesn’t even require initializing Terraform. But validate does, which is where we run into the first roadblock.
|
||
You kind of need to use a remote backend for state data in a CI/CD pipeline. The hosted runner executing your pipeline is ephemeral and if you store your state data there, poof! It’s gone. That means you need to setup a remote backend and pass credentials to access it. Now you wouldn’t put the whole backend config in your Terraform code, right? You’re not an animal. So we go with a partial config using the azurerm backend, and pass the additional information including credentials with environment variables and the -backend-config switch. But where to store those values? Azure Key Vault of course! And wouldn’t you know it? There’s an option in the Pipelines GUI to create a variable group and link it directly to a Key Vault. How convenient! Unfortunately, that option doesn’t exist in the Azure CLI or Terraform provider for Azure DevOps. Yup, you read that correctly, there is an option in the GUI that is not available in the CLI, API, or Terraform provider. Grrrrr…
|
||
There is an undocumented API call you can make to create it, but “Ewwww!” Undocumented APIs are hot garbage. It roughly translates into “an internal facing API that could be changed at a moments notice or removed completely, and we probably won’t notify you because you shouldn’t be using it anyway”. If you want to complain loudly about this injustice, or at least up vote this issue on GitHub, I highly encourage you to do so.
|
||
The solution in this case was to use the AzureKeyVault task in the pipeline to load secrets as environment variables. Not the biggest stumbling block, but 100% annoying. All interfaces in Azure DevOps should be API first.
|
||
Continuous Delivery Now we move into the next pipeline that should be triggered by a pull request. And here is where I ran into some difficulty. Part of this is poor documentation on Microsoft’s part and part of this is my lack of understanding when it comes to Git and GitHub.
|
||
There’s a special pr block you can define which will trigger a pipeline when a PR event happens in GitHub. Here’s what my block looks like:
|
||
pr: branches: include: - main There’s a few things you should know. The include statement refers to the target branch of the merge, not the source branch. This confusion led to a non-insignificant amount of head to brick wall interactions for me. If you want it to fire when you create a pull request to merge a feature branch to main, the include statement should reference main.
|
||
The second thing to know is that the pr block stands alone in the pipeline, it does not get nested in a trigger block. When I was trying to figure out why my pipeline wasn’t firing, I thought maybe it needed to be nested, and that was definitely wrong.
|
||
The third thing to know is what happens when you create a PR on GitHub. Every time I created a new PR, both my CI and PR pipelines would fire. And I didn’t understand why. I wasn’t pushing code or creating a commit right? I’m just creating a PR! Why are both pipelines triggering?
|
||
Well, here’s the thing. When you create a pull request on GitHub, it actually creates a new ref and a commit for that ref. My PR pipeline is looking for pull requests and firing like it should. The CI pipeline is in a different file and it is looking for any commits that aren’t on main. The PR event in GitHub is a commit on a non-main branch, and so my CI pipeline is firing every time. In fact, the Microsoft Docs say so in a roundabout sort of way:
|
||
If no pr triggers appear in your YAML file, pull request validations are automatically enabled for all branches, as if you wrote the following pr trigger. This configuration triggers a build when any pull request is created, and when commits come into the source branch of any active pull request.
|
||
https://docs.microsoft.com/en-us/azure/devops/pipelines/repos/github?view=azure-devops&tabs=yaml#branches
|
||
If you had to read that paragraph like three or four times, you are not alone. Because I have a separate pipeline with no pr trigger, it is firing on every PR event. The way to avoid this is add a conditional statement to the stages in your pipeline like this:
|
||
condition: eq(variables['Build.Reason'], 'IndividualCI') The Build.Reason for the PR will be PullRequest and so the stage will not run when you create a new PR. The pipeline itself will still run, but it will skip all the stages and come back green. The stages will still run if you make new commits to an open PR, and that’s probably what you want!
|
||
You might be wondering if you can combine the pr and trigger blocks in one pipeline definition, and the answer appears to be yes. You can combine both the CI and PR pipelines in to a single pipeline and use a condition for each stage, job, or task to filter on when it runs.
|
||
Terratest and Authentication Another issue I ran into while developing my PR pipeline was using Terratest to validate the AKS cluster and cluster components. The problem is when you attempt to use a Terratest module that leverages the Azure Client to connect. For instance consider the following code snippet to get information about an AKS cluster:
|
||
cluster, err := azure.GetManagedClusterE(t, resourceGroupName, clusterName, subscriptionId) The code needs to create an Azure Client to connect to the Azure API and get the AKS cluster information. If you were running this locally on your desktop, you probably would already be logged into the Azure CLI. Terratest will use that cached login to connect. But if you’re running this in a pipeline, the hosted runner doesn’t have the cached login.
|
||
The same goes for the Terraform commands that leverage the AzureRM provider in a Terratest module. By default, those will also use the cached credentials from your Azure CLI login. The Terratest documentation calls out that you can use the standard environment variables: ARM_CLIENT_ID, ARM_CLIENT_SECRET, and ARM_TENANT_ID for authentication. Great! But hold on a second, that doesn’t seem to work for the Azure Client.
|
||
Turns out that the while the Terraform AzureRM provider uses one set of environment variables, the Azure Client from the Azure SDK uses another. Here’s the equivalent variables:
|
||
AzureRM Provider Environment VariablesAzure SDK Environment VariablesARM_CLIENT_IDAZURE_CLIENT_IDARM_CLIENT_SECRETAZURE_CLIENT_SECRETARM_TENANT_IDAZURE_TENANT_ID Table of environment variables in Terraform and Azure SDK
|
||
Confused yet? Why the two decided to use completely different variables is a mystery to me, but ultimately it’s not a big deal. Since I’m using environment variables for the AzureRM provider, I can simply create environment variables in my Go code by adding the following block:
|
||
// Create env variables for Azure client communication os.Setenv("AZURE_CLIENT_ID", os.Getenv("ARM_CLIENT_ID")) os.Setenv("AZURE_CLIENT_SECRET", os.Getenv("ARM_CLIENT_SECRET")) os.Setenv("AZURE_TENANT_ID", os.Getenv("ARM_TENANT_ID")) You could also use self-hosted runners with a Managed Identity associated with them (assuming you’re running them on Azure). Then both the AzureRM provider and Azure Client would be able to use MSI authentication. Or at least, I think that would work. I haven’t actually tried it.
|
||
Total Destruction Azure’s API can be… problematic when it comes to fully destroying all the infrastructure Terraform has stood up. The PR pipeline builds a test instance of the environment, runs the tests in Terratest, and then tears everything down. Due to the eventual consistency nature of Azure’s APIs, sometimes the destruction fails due to a sequencing error. In my particular case, the AKS cluster was utilizing the Application Gateway Ingress Controller, which spawns a new Application Gateway in a dedicated subnet when the cluster is created.
|
||
That subnet requires a set Network Security Group rules that must be in place before the Application Gateway will be generated, and the rules cannot be deleted until the Application Gateway is removed. Therein lies the problem. Terraform will attempt to destroy the cluster and the rules at the same time, and sometimes the rules destruction will fail because the cluster destruction process has not yet torn down the Application Gateway.
|
||
If you run terraform destroy a second time, the process will complete successfully, since now the Application Gateway is gone. The whole apply and destroy process is happening inside my Go code, so I needed to write a function that would run destroy a second time. The original code looked like this:
|
||
defer terraform.Destroy(t, terraformOptions) Where the defer keyword means that the command runs when all other tests complete, regardless of outcome. This will run destroy a single time, so I wrote a new function called DoubleDestroy:
|
||
func DestroyDouble(t terratest.TestingT, options *terraform.Options) string { out, err := terraform.DestroyE(t, options) if err != nil { out = terraform.Destroy(t, options) } return out } And then in the main testing function I changed the defer line to this:
|
||
defer DestroyDouble(t, terraformOptions) Now the destroy process will run twice, and if that still fails it will error out. I suppose I could have done a while loop up to x number of tries, but really if the destroy fails after the second go around, I’d rather fail the whole thing and see what happened.
|
||
Conclusion There were plenty of other headaches I encountered while trying to build the pipeline, and I don’t want to bore you with them now. I think this post is already sufficiently long. If you are planning to build an Azure Pipeline to automate Terraform deployments, I hope this information helped you. I’d also recommend checking out my videos on YouTube regarding building with Azure DevOps.
|
||
`,summary:"I am in the process of creating a series of liveProjects for Manning Publishing, one of which involves deploying and managing an Azure Kubernetes Cluster using Terraform and Azure DevOps Pipelines. As I tried to build a CI/CD workflow that uses several pipelines, I ran into a bunch of roadblocks, some of which were my own doing and others which stemmed from a lack of clear documentation. While I won’t give away the contents of the liveProject - that would sort of defeat the point - I did want to call attention to things I stumbled over in hopes of helping others on a similar journey.",date:"23 Nov, 2021",url:"https://nedinthecloud.com/2021/11/23/what-i-learned-building-an-azure-devops-pipeline-with-terratest/",image:"ado-problems.png",readingTime:"10"},"https://nedinthecloud.com/2021/10/11/weekly-newsletter-10/11/2021/":{title:"Weekly Newsletter - 10/11/2021",tags:[],content:`Hello everyone!
|
||
Last week was a rough one for me. As you might have read in last week’s newsletter, I broken my pinky toe while out on a trail run. While I knew it was broken last Sunday thanks to an urgent care nearby and the miracle of X-rays, I didn’t know what the recovery process looked like. There is a trail race coming this Friday, and I really hoped that somehow I would be able to run it. That is not going to happen.
|
||
I went to see a podiatrist on Wednesday, and he told me that the recovery time is 4 to 6 weeks. I might be able to comfortably walk on it in a couple weeks, but doing anything strenuous would likely result in additional injury and a longer recovery time. That’s the bad news. The good news is that I get to wear this super fashionable footwear for the next 4 weeks.
|
||
The dreaded surgical shoe.
|
||
Needless to say, I was feeling a bit down last week, and my productivity suffered as a result. But I have pulled myself out of the mire, and I am ready to get some good stuff cooking. I’ve just started multiple projects including a full redo of my Terraform - Getting Started course on Pluralsight, some consulting work with Terraform, and a Manning liveProject using AKS, Terraform, and Azure DevOps. My dance card is now quite full, and I think it will serve as a healthy distraction from my toe.
|
||
Here’s what going on with Ned in the Cloud this week.
|
||
YouTube Channel - https://www.youtube.com/c/NedintheCloud/
|
||
This was an off week for Terraform Tuesdays, so I worked on my Python application for a GCP video I will publish this week. In a previous video, I walked through setting up self-hosted runners for GitHub Actions on GCP. The next step in the process is to have something to deploy on GCP with those runners. To that effect, I have written a simple Flask app that allows you to create polls and vote on them. It uses Flask on the front end and Google Cloud SQL on the backend. The video this week will show how to deploy the application using an instance group, secrets, and Cloud SQL.
|
||
There will be two more videos in the series after this. Once the application is being deployed successfully using Terraform, I will move it into GitHub Actions for deployment. The last video will involve setting up branches and releases to deploy updates to the infrastructure. Should be a good time!
|
||
Day Two Cloud - https://daytwocloud.io
|
||
Last week we spoke to Emily Omier about growing your open-source community. While the podcast mainly focused on open-source, I think the lessons apply more broadly to any project or product you are working to develop. Emily had some key insights around developing your messaging, working with contributors, and listening to user feedback. As someone who has read their fair share of press releases and product websites, I am constantly surprised at how badly most of them communicate what their project or product actually does. Pro tip: If I’m three paragraphs into your product copy and I still have no idea what it does, you have failed at a fundamental level. Clearly communicating what your product does and the value it provides should be top of mind for any content that is aimed at potential users.
|
||
This week we pivot to cloud security with sponsor Valtix. You may not have heard of Valtix, but you probably know some of the products the founders have worked on, including Andiamo, Big Switch, and Cisco ACI. Their basic pitch is to create a centralized security control plane for all your cloud deployments, which can provide visibility, reporting, and policy. The policy gets pushed down to their managed gateways running in your cloud environment, so your data never leaves your boundary. The coolest part is that the reporting and centralized control plane are both free! You only have to pay for any managed gateways you stand up. Basically, you can get analysis of your cloud environments at no cost. With multi-cloud becoming so prevalent, I see cloud agnostic tools as the future, and Valtix definitely sits in that category.
|
||
(See what I did there? I described what a product does and its benefit to the user in a single paragraph. Quick, someone pay me thousands of dollars!)
|
||
Daily Check-In
|
||
I missed two days last week on account of self-pity and general wallowing.
|
||
Facebook Down Schadenfreude - It’s tempting to point and laugh, but that could be you. Terraforming My Mental Health - My attempt to climb out of the hole. Video Killed the Audio Star? - Adding live video to my courses? Still not sure. Highlighted episode - Facebook being down for six hours was an outage that could not be ignored, and brought with it a lot of crappy takes. I tried to make sure mine wasn’t one of them.
|
||
Other Stuff
|
||
Top Five Cloud Posts:
|
||
MSP consolidation through private equity - I’m sure service will improve and costs will go down right? Drinking from the firehose is not healthy - Thanks to Ben for connecting some dots. Intriguing changes in datacenter infrastructure spending - Cloud. Rules. Everything. Around. Me. How IBM Lost Public Cloud - Spoiler, it wasn’t just one thing. It was all the things. (Hat tip: Scott Lowe) Cloudflare analysis on Facebook outage - There’s a LOT of data here. Pretty pictures abound. That’s all for this week. Thanks for reading! If there’s something you’d like to see in my newsletter, please let me know.
|
||
`,summary:`Hello everyone!
|
||
Last week was a rough one for me. As you might have read in last week’s newsletter, I broken my pinky toe while out on a trail run. While I knew it was broken last Sunday thanks to an urgent care nearby and the miracle of X-rays, I didn’t know what the recovery process looked like. There is a trail race coming this Friday, and I really hoped that somehow I would be able to run it.`,date:"11 Oct, 2021",url:"https://nedinthecloud.com/2021/10/11/weekly-newsletter-10/11/2021/",image:"Weekly-Newsletter.png",readingTime:"5"},"https://nedinthecloud.com/2021/10/04/weekly-newsletter-10/4/2021/":{title:"Weekly Newsletter - 10/4/2021",tags:[],content:`Hello everyone!
|
||
You just never know the direction a day will head when you wake up. Sure, you can make plans and hope for the best. But I think we all know what reality thinks of our finely laid plans (cue condescending laughter.) For instance, this Sunday I had planned to go out for a 15 mile trail run, followed by standard Sunday activities – read cleaning, and making fresh pasta for the family.
|
||
The reality is that I bashed my foot on a rock while running and ended up going to urgent care only to discover that I had fractured my pinky toe. Needless to say, this changed our Sunday plans a bit. While I was disappointed and more than a little frustrated, I had to accept the change. To go with the flow as it were.
|
||
If nothing else, the last 18 months have shown how unpredictable and chaotic the future can be. The answer is not to give up or stop making plans. The answer is to be fluid and adaptable with those plans, seeking out opportunities and coming up with clever solutions. I’ve learned to take my lumps, lick my wounds, and come back for more. I hope you have too.
|
||
Here’s what going on with Ned in the Cloud this week.
|
||
YouTube Channel - https://www.youtube.com/c/NedintheCloud/
|
||
There’s a fair amount of confusion when new users of Terraform encounter input variables and local values. When do you use each, and what is the difference? That is the topic of last week’s video, another installment of my Terraform Basics series where we take a close look at one specific aspect of Terraform.
|
||
This week is an off week for Terraform Tuesday. I am working on a demonstration deploying an application to GCP using GitHub Actions and Terraform. I wrote the application myself using Python and Flask, and I’m feeling a bit proud. Not that it’s a particularly interesting application (it’s not), but I’ve never written a web application in Flask from scratch and I was excited to get it working.
|
||
Day Two Cloud - https://daytwocloud.io
|
||
Last week was a sponsored show from Akamai focusing on how they helped IBM Cloud transform their cloud portal application. The portal application originally started out as a monolithic app running on CloudFoundry with a single instance per region. Working with Akamai, IBM Cloud was able to seamlessly break the monolith up into microservices and create a multi-region version of the portal that was more resilient and performant. We had Tony Erwin from IBM Cloud and Pavel Despot from Akamai on the show to tell the story, and they don’t shy away from getting into the details. I found it a compelling case study into how web applications can evolve from a traditional approach to cloud native.
|
||
This week we’re changing it up by talking to a fellow podcaster about building an Open Source Software community. Emily Omier runs a consulting business helping tech startups do exactly that, and during the episode she shares some of her best advice for growing and maintaining a community. While we mostly focus on OSS, I think the information is more broadly applicable to any tech community out there. If you are active in any tech communities, there are some valuable insights from Emily you won’t want to miss.
|
||
Daily Check-In
|
||
We’ve got four solid episodes from last week (Tuesday was a bit of a wild card.)
|
||
Community Recognition Programs - When are they worth trying to join? HashiCorp’s State of the Cloud Report - Skills and security are the key takeaways. Part 2 of HashiCorp’s State of the Cloud Report - Methodology matters and here’s why. Making a Plan for Learning - How do you stay up to date, and what should you stay up to date on? Highlighted episode - Check out both part 1 and part 2 of the HashiCorp State of the Cloud Report.
|
||
Other Stuff
|
||
This week I am starting to revise my Terraform Getting Started course on Pluralsight. The current course is two years old and was created based on version 0.12 of Terraform. While it isn’t as big of a jump as the move from the first version of the course using 0.9, there’s a decent amount that has changed in the last two years. In addition to updating all the exercises to version 1.x, I am also looking at how I can change the demos to make them as relevant as possible without learners getting frustrated. I have a few ideas already, and I’m sure there’s more to come!
|
||
Top Five Cloud Posts:
|
||
AWS Cloud Control API - One API to rule them all and in the Terraform bind them. Let’s Encrypt CA Cert Expires - A planned retirement still causes pain for those who didn’t heed the advice. OpenVSCode from Gitpod - Now you can run VS Code in the browser without MSFT’s special sauce. CloudFlare vs. AWS round 2 - R2 Storage takes aim at S3 and AWS pricing. AWS IAM docs bad? - Ben Kehoe has FEELS, and I agree. That’s all for this week. Thanks for reading! If there’s something you’d like to see in my newsletter, please let me know.
|
||
`,summary:`Hello everyone!
|
||
You just never know the direction a day will head when you wake up. Sure, you can make plans and hope for the best. But I think we all know what reality thinks of our finely laid plans (cue condescending laughter.) For instance, this Sunday I had planned to go out for a 15 mile trail run, followed by standard Sunday activities – read cleaning, and making fresh pasta for the family.`,date:"4 Oct, 2021",url:"https://nedinthecloud.com/2021/10/04/weekly-newsletter-10/4/2021/",image:"Weekly-Newsletter.png",readingTime:"5"},"https://nedinthecloud.com/2021/09/27/weekly-newsletter-9/27/2021/":{title:"Weekly Newsletter - 9/27/2021",tags:[],content:`Hello everyone!
|
||
Last week I spent some time in Draper, UT - about 30 minutes outside of Salt Lake City - at the new Pluralsight HQ building for the annual Author Summit. The summit is only two days long, but the Pluralsight team packed in a ton of great sessions and social time for authors and staff to mingle and learn from each other. I came back from the trip feeling energized and excited to create new educational content for y’all.
|
||
There are three main takeaways I had from the trip:
|
||
Video courses are not enough - People learn by doing, and Pluralsight is focused on making interactive labs that seamlessly weave into their courses. Live video is more engaging - I kind of knew this based on my YouTube experience and a talk from Simon Allardice further cemented that feeling while providing practical guidance for incorporating it into my courses. In-person events aren’t going anywhere - This was the first in-person event I attended since the pandemic started and the difference between Author Summit and virtual conferences was stark and sobering. We might have convinced ourselves that virtual was good enough during the height of the pandemic, but last week showed me what a pale version of in-person these virtual events were. I covered these takeaways in more detail during my Daily Check-In episodes, two of which I recorded while in Utah. Fun fact: my carry-on bag got flagged for additional inspection both ways due to the Rode Podcaster mic looking like something suspicious. On the way home, the TSA agent even did the chemical wipe to test for possible explosives. No offense dude, but it’s pretty clearly a microphone even on the X-ray scan. Ah, the perils of being a traveling podcaster.
|
||
Here’s what going on with Ned in the Cloud this week.
|
||
YouTube Channel - https://www.youtube.com/c/NedintheCloud/
|
||
Last week was an off week for Terraform Tuesdays, so I was working on demos for future videos. I’ve got a few interesting things in the hopper:
|
||
GitHub Actions with GCP part 2 - Wherein we use our self-hosted runners to provision infrastructure and a Python application in GCP using Terraform. Libvirt - Still trying to get all the things working with the Libvirt provider from the Terraform registry. Local Values - As part of the Terraform Basics series, I am going to do a video on local values. This week I will be publishing the Local Values video. It will likely be a relatively short video, because local values are just not that complicated. They are local values you define in your configuration to make life easier and that’s about it. But it’s important to understand the scope of local values, how they are different from input variables, and different strategies for placing them in your files.
|
||
I am writing my own Python app for the GitHub Actions video. This is completely unnecessary. I could just grab any example Flask app and run with it. However, it (a) gives me an excuse to learn more about Python and (b) gives me a chance to write an example app I can use for other demos that is mine and mine alone.
|
||
Day Two Cloud - https://daytwocloud.io
|
||
Making the move from individual contributor to a management position is not an easy path to walk. In previous episodes, Ethan and I shared our stories, and we chatted with two people who have made the transition successfully. This time we’re talking to Shelley Benhoff, who not only spent some time in management, but has also created courses for Pluralsight about being a better manager. Shelley is a burgeoning podcaster in her own right, with her show Tiaras and Tech launching at the end of October. Check out the episode and keep an eye on her forthcoming podcast launch!
|
||
This week we have a sponsored episode with Akamai to talk about how they helped IBM Cloud restructure their public facing portal and API. IBM Cloud went from having region specific portals backed by a monolithic app without any CDN to a single unified portal using microservices leveraging Akamai services for CDN and more. I love it when we get to dig into a real world example of how technology evolves and migrates. Both Tony Irwin from IBM and Pavel Despot from Akamai were candid and transparent about the challenges they faced and how they solved them. These are the type of conversations I love to have on Day Two Cloud!
|
||
Daily Check-In
|
||
I skipped Wednesday last week, since I was on a plane, in the air. Something told me my fellow passengers wouldn’t appreciate me trying to record a podcast mid-flight.
|
||
Brevity and Efficiency - How can I slim down my instruction while still maximizing content? In-person Isn’t Going Anywhere - Virtual doesn’t hold a candle to in-person. Better Learning Through Video - Apparently people engage better with a face than a slide. Thoughts on Tech Startups - I’d recommend embracing SaaS, cloud-native, and cloud-agnostic. You know, ALL the buzz words. Highlighted episode - I thought my episode on Tech Startups was pretty relevant to anyone looking to get into the field or evaluate a product.
|
||
Other Stuff
|
||
Top Five Cloud Posts:
|
||
Canonical Gives Admins a Little More Time - You’ve got Ubuntu 14.04 for a couple more years. Good News! No more passwords at all - Now generally available for everyone. Bad News! Your passwords are compromised - Dammit Autodiscover! You know better. Give the Developers What They Want? - If what they want is a platform… Web 3 and NFTs - I’m sure this thread will age well. That’s all for this week. Thanks for reading! If there’s something you’d like to see in my newsletter, please let me know.
|
||
`,summary:`Hello everyone!
|
||
Last week I spent some time in Draper, UT - about 30 minutes outside of Salt Lake City - at the new Pluralsight HQ building for the annual Author Summit. The summit is only two days long, but the Pluralsight team packed in a ton of great sessions and social time for authors and staff to mingle and learn from each other. I came back from the trip feeling energized and excited to create new educational content for y’all.`,date:"27 Sep, 2021",url:"https://nedinthecloud.com/2021/09/27/weekly-newsletter-9/27/2021/",image:"Weekly-Newsletter.png",readingTime:"5"},"https://nedinthecloud.com/2021/09/20/weekly-newsletter-for-9/20/2021/":{title:"Weekly Newsletter for 9/20/2021",tags:[],content:`Hello everyone!
|
||
Last week I spent five days in Nashville, taking in the sights, listening to live music, and probably drinking a bit too much. This was the first time my wife and I had been away from the kids for more than a night in ten years. Ten years! Although both of us have been on longer trips individually, we hadn’t gone together since 2011. It was supremely strange to be hanging out for multiple days and not be interrupted by a child who wants/needs something. Both of us commented on how strange and liberating it was.
|
||
But alas, as I write this, we are waiting in the airport to fly home. It was an absolutely fantastic trip and I am ready to be back home with the kiddos. It will be a brief respite, as I am heading out again on Monday to go to the Pluralsight Author Summit in Salt Lake City. Eighteen months of basically no travel, and now I have trips back to back! That’s just how things shook out with our schedules and when Pluralsight decided to have their summit. The trip is a short one and I will be back on Wednesday night with no future travel until November.
|
||
Here’s what going on with Ned in the Cloud this week.
|
||
YouTube Channel - https://www.youtube.com/c/NedintheCloud/
|
||
The original plan for this week was to have a demonstration around using the libvirt provider for Terraform to create and configure virtual machines on a remote host. As I tried to put together the demo, I kind of hit a wall with the provider. The problem is entirely mine I’m sure. I don’t know enough about the libvirt API and the underlying components, so when something didn’t configure correctly I had a lot of trouble figuring out the cause. I am not going to abandon the demo, but I will need a bit more time to figure out.
|
||
Instead of the libvirt demo, I threw together a demo of setting up self-hosted runners in GCP for GitHub Actions. And when I say I threw it together, I mean that I had no demo at the beginning of the day on Tuesday and by lunchtime I had self-hosted runners in GCP deployed with Terraform. One of the great things about being familiar with Terraform and GCP is that I didn’t have to fight through a bunch of mistakes, because I’ve already made those mistakes before and learned from them! This is the first of at least two videos. In the next one, we are going to use GitHub Actions and our self-hosted runner to deploy an application to GCP. Once that is all working, we’ll see where else we can go. I also need to put together a demo using GitHub Actions with Microsoft Azure, so this is serving a bit of a dual-purpose.
|
||
Day Two Cloud - https://daytwocloud.io
|
||
Last week we had a sponsored episode with the good folks over at Console Connect PCCW. You might see PCCW and think that it is some consulting firm, but you would be incorrect. PCCW is one of the largest internet service providers in the APAC region. They handle a not insignificant amount of total internet traffic, and they brought Console Connect on to help them manage their internal products as well as provide their service to end customers. The central idea is simplifying the process of provisioning and connecting networks, whether public or private. Our episode with PacketFabric was along the same lines, and there is more than enough room in the market for both of these companies to grow and thrive. What I really liked about the episode was our peek behind the curtain into how the system is structured, including a discussion about APIs and abstractions.
|
||
This week is a new entry in our series about moving from an individual contributor to a management role. So far you’ve heard about Ethan and my personal journeys, from two people we know that have made the move, and now we get to talk to Shelley Benhoff. She not only made the transition herself, but went a step further to create courses on Pluralsight about becoming a more effective manager. I really enjoyed chatting with Shelley about her personal history and her struggles making the move to management. You won’t be surprised to discover that everyone struggles the first time they try to make the move. Focusing on empathy and compassion will take you a long way!
|
||
Daily Check-In
|
||
Since I was traveling last week, there’s only three Daily Check-Ins.
|
||
How Much To Charge? - A difficult question. There are a few shortcuts to help you out. Don’t Want To Be A Pendantic Idiot - Gatekeeping through terminology is a real problem in tech. Don’t be that guy. An End to Patreonizing - I’ve decided to close my Patreon page. Here’s why. Highlighted episode - As much as I’m sure you care about the end of my Patreon, I think the most useful episode is How Much To Charge?.
|
||
Other Stuff
|
||
Top Five Cloud Posts:
|
||
Azure needs to update Linux Agent - Like real bad. In the meantime, you’ll have to do it. AWS DevOps Pro Study Guide Part 1 - AWS Hero Chris Williams is laying down some knowledge. Malware Targeting WSL - Where there’s a WSL there’s a way. SoftBank Investment in Latin America - Latin America is seeing a LOT of love right now. Mirantis Launches Flow DCaaS - It’s basically managed OpenStack and Kubernetes, but OK. That’s all for this week. Thanks for reading! If there’s something you’d like to see in my newsletter, please let me know.
|
||
- Ned
|
||
`,summary:`Hello everyone!
|
||
Last week I spent five days in Nashville, taking in the sights, listening to live music, and probably drinking a bit too much. This was the first time my wife and I had been away from the kids for more than a night in ten years. Ten years! Although both of us have been on longer trips individually, we hadn’t gone together since 2011. It was supremely strange to be hanging out for multiple days and not be interrupted by a child who wants/needs something.`,date:"20 Sep, 2021",url:"https://nedinthecloud.com/2021/09/20/weekly-newsletter-for-9/20/2021/",image:"Weekly-Newsletter.png",readingTime:"5"},"https://nedinthecloud.com/2021/01/25/failing-the-certified-kubernetes-administrator-exam/":{title:"Failing the Certified Kubernetes Administrator Exam",tags:["certified-kubernetes-administrator","cka","kubernetes"],content:`As you may have gathered from the title, I did NOT pass the CKA exam. The title is not some clickbait where I turn things around and explain how I passed the exam by doing X, Y, and Z. Nope, I failed. And failure is important. You learn more from your failures than your successes, as anyone who has tried to build a CI/CD pipeline or write a piece of code can attest. There’s always going to be a lot of red before you see a hint of green. In this post I want to talk about what I did to prepare, what I should have done, and what I plan to do before retaking the exam. Yes, dear reader, I am going to retake the exam and this time I will PASS!
|
||
Before I dive into the CKA exam proper, I think it’s important to give some background. I’ve been taking certification exams for a long time, starting with the Windows 2000 Professional exam back in 2002. Without exception, all of the exams I had taken before the CKA were basically paper exams. There were a few sad attempts by Microsoft and Citrix to provide a “real” environment to do tasks in. But they were severely crippled. You knew you weren’t in a real environment, and functionality was highly limited.
|
||
The CKA exam is not that. You are not answering multiple choice questions, fill in the blank, drag-and-drop, or any other such nonsense. You are given a full server running Ubuntu that is connected to multiple Kubernetes clusters and they give you tasks to accomplish. You have a single terminal session into the environment and no GUI. Everything is going to be done at the command line level. This is a daunting prospect to a former Windows administrator who is used to answering multiple choice questions about esoteric components in Microsoft applications.
|
||
So I tried and failed. But like I said, there’s a lot to learn from failure. Let’s break it down.
|
||
What worked It’s not like I didn’t study before the exam. I started studying back in December. Here’s what I did for studying:
|
||
Went through the whole CKA path from Anthony Nocentino on Pluralsight Did the exercises on this GitHub repo and this one Built two K8s clusters on Ubuntu, one in Azure and one in my home lab In the process I learned a TON about Kubernetes. While I had been working on K8s in one respect or another during the previous year, my knowledge was specific to accomplishing a single task. The courses and practice gave me the background knowledge I needed to fill in the gaps.
|
||
What didn’t There’s a few things I think are worth pointing out. First, while I have been using command line tools for many years, they have mostly been in the Windows context. Linux has always been secondary in my professional life and as a result I’ve never been as familiar with it. I am still not great with basic tools like sed and awk. The practical upshot is that taking an exam entirely at a Linux shell is not a comfortable thought - in fact it strikes terror in my heart.
|
||
There are things that I just don’t know or am not comfortable with, and if you are a Windows admin, you may be in the same boat. Allow me to give you a brief list of things you should probably get comfortable with.
|
||
Autocomplete - You need to know how to set this up, it will save you tons of typing in real life and the exam. The directions are in the K8s official docs. You don’t have to memorize them, just remember where they are. Terminal multiplexing - I used tmux, but this applies to any similar program. You only get one terminal in the exam. Having a tool like tmux is going to save you time and let you reference multiple screens at once. Text editor - The default is vim, which is what I first learned to use when I was introduced to Linux. It’s also the default editor for kubectl when you edit an existing configuration. The default config for vim doesn’t work well with YAML, and that’s a problem. You either need to edit the default config for vim or use a different editor. General Linux knowledge - You should really know how to use the most common Linux tools to parse and manipulate output. Things like more, head, grep, and output redirection are all important. If you’ve worked in Linux for years, these are all probably second nature to you. They are not for me. Ubuntu - The nodes are all going to be running Ubuntu. You need to know a bit about using Ubuntu if you want to get by. Not deep knowledge, but you should know what systemctl is, how systemd organizes stuff, using the apt package manager, and the overall layout of the filesystem. Docker - Despite the fact that Docker is “deprecated”, the exam nodes will be using Docker for the container runtime. You should know a little about using Docker to check on containers. Other software - You should know how to use journalctl and etcdctl to work with the logs on the nodes and the etcd backend Like I said, the main problem here for me is that Linux is not my daily driver. I am not deeply steeped in using bash, tmux, vim. Sure, I’m familiar enough to get by, but there’s still a cognitive penalty to using these tools that I don’t have with most Windows tools and PowerShell commands.
|
||
What I learned There are three key things I am taking from the exam: muscle memory, Linux fluency, and exam strategy.
|
||
Muscle memory The more you use a technology, the more you build up a muscle memory when it comes to commands. I’ve started to develop that muscle memory with kubectl, but it isn’t enough. Basically, I need to spend more time going through exercises with kubectl and the most common commands memorized. I definitely spent too much time looking up command structure during the exam, and that was wasted time. For instance, I should know what objects you cannot create imperatively with kubectl. Instead, I wasted precious time trying to use autocomplete; looking up the command in help; finding it wasn’t there; and then looking up the YAML for a declarative manifest.
|
||
Key takeaway: Spend more time doing exercises until common tasks become ingrained in my memory.
|
||
Linux fluency There’s a certain point you reach with tools where you don’t feel constrained by them. It’s like riding a bike. At first you are just trying to balance on two wheels without falling over. Then you get a bike with gears, and now you’re thinking about shifting. Once that becomes second nature, you start considering the road and how to get where you are going. With many Linux tools, I feel like I am still trying not to fall over, while with others I am trying to figure out all the gears. I need to get to the point where the tool disappears, and now I am using it to arrive at the desired destination. The only way I can think of improving is by practicing and perhaps taking a Linux course with exercises that force me to get better with Linux in general, and Ubuntu specifically.
|
||
There’s also a case to be made for selecting the right tool for the job. I definitely lost time using vim during the exam, because the default in vim is to treat a tab like five spaces and not two. Which means, when you past YAML into vim, the whitespace in the file is completely wrong. Unfortunately, that’s kind of important in YAML. You can change the vim config with a .vimrc file, but the entries are not exactly easy to remember and you can’t look them up during the exam. If you’ve got them memorized, go for it. I plan to use nano next time, since it doesn’t mess with the spacing. It doesn’t help either, but that’s better than actively making my life harder.
|
||
Key takeaway: Become fluent with Linux and your chosen shell to the point they start to disappear.
|
||
Exam strategy I sort of alluded to some of my strategy, which is to reduce wasted time during the exam. I need to be more comfortable with the tooling and common resources, so I don’t spend too much time looking up structures or commands. The tasks in the exam are also weighted differently, and that should inform which tasks I should start with. It probably would have made more sense to flip through all the tasks, knock out the easiest ones first, and then comes back to the ones with the highest weight that I feel most confident about. You don’t have to clear every task, in fact you only need a 66% to pass. I bet that if I had focused on a troubleshooting task that was weighted highly over some of the other tasks, I might have passed.
|
||
The exam environment has some weird quirks that also mess with me. First, there is a visual indicator showing how much time you have left, but you don’t see the actual value. You know the bar is shrinking, but it doesn’t say you have 10 minutes left. Instead, the proctor tells you every once in a while that you have XX amount of time left. I found that disruptive and annoying. You can still see your own system clock, and you can do the math. I’d rather the proctor didn’t break my train of thought to make me panic more than I already am.
|
||
You also cannot use Ctrl+C and Ctrl+V to copy and paste. Why? I dunno. Other browser terminals have figured this out. It adds cognitive dissonance and additional stress to go against your own muscle memory. Plus, the replacement command sequence for Windows is Ctrl+Insert to copy and Shift+Insert to paste. I was using a laptop without a dedicated Insert key (You have to hold down Fn and hit the Del key), adding even more dissonance to my mental process. Mac users, which I am assuming is the bulk of the exam writers, can blissfully use Apple+C and Apple+V without problems. When I retake the exam, I am going to plug a full-size keyboard into my laptop to avoid this issue.
|
||
Some tasks require that you do things declaratively, and there are great examples on the official Kubernetes docs site. However, you need to be able to find those examples without too much effort. Next time I am going to make a mental map of how to find things on the docs site easily and avoid wasting time clicking around. You also have access to the GitHub page for Kubernetes. I didn’t look at it closely, and I plan to investigate it for useful resources before I try the exam again.
|
||
Key takeaway: Review all tasks before starting, use a better setup, and review the online resources in-depth.
|
||
Wrap-up Failing sucks. There’s really no other way to put it. It’s hard to pick yourself up and get back on the bike. I feel I wasted a bunch of effort just to come up short. But that’s the wrong approach, and I know that in my mind, if not in my heart. After a weekend of reflection, I am ready to learn from my mistakes, work hard to improve, and pass this exam.
|
||
I hope this has been helpful to you as well. If you have questions, hit me up on Twitter or leave a comment. Thanks for reading!
|
||
`,summary:"As you may have gathered from the title, I did NOT pass the CKA exam. The title is not some clickbait where I turn things around and explain how I passed the exam by doing X, Y, and Z. Nope, I failed. And failure is important. You learn more from your failures than your successes, as anyone who has tried to build a CI/CD pipeline or write a piece of code can attest.",date:"25 Jan, 2021",url:"https://nedinthecloud.com/2021/01/25/failing-the-certified-kubernetes-administrator-exam/",image:"CKA-Post.png",readingTime:"10"},"https://nedinthecloud.com/2021/01/03/2020-year-in-review/":{title:"2020 Year in Review",tags:[],content:`2020 was… a year. Or more like a decade packed into a single year. Looking back at everything that happened, it seems both too much and not enough. Conferences were cancelled, schools were closed, Zoom proficiency spiked, and we all added several new words to our vocabulary. This post is not meant to rehash our shared trauma of 2020, but I need to at least acknowledge it as a driving factor for the year. At the beginning of the year, I set out some ideas and goals for the year. Let’s see how I fared given all that happened.
|
||
Because I am completely ridiculous, I created a mission statement for Ned in the Cloud at the beginning of the year. That statement?
|
||
Ned in the Cloud’s vision and purpose is to create compelling technical content for IT professionals across multiple mediums based on impactful and emerging technologies.
|
||
That’s a pretty good vision statement! I think it accurately describes what I did this year. Speaking of which, what did I create this year?
|
||
Day Two Cloud The Day Two Cloud podcast continued to roll on into a second year of existence. At the end of 2019, Ethan Banks came on as my cohost and the show shifted to a weekly cadence. In 2020, we published fifty episodes covering cloudy topics like Kubernetes, cloud economics, and hybrid cloud. If I had to pick our five best episodes for the year, I would have to go with the following:
|
||
Episode 79 - K8s is inevitable but not always necessary Episode 70 - The state of multi-cloud networking Episode 67 - Choosing the right applications for the cloud Episode 55 - Securing cloud infrastructure and applications Episode 52 - Moving back home from the cloud That was a tough selection process! We had so many amazing guests this year, including people like Mike Pfeiffer, Tanya Janca, Bobby Allen, April Edwards, and Corey Quinn. Our goal is to create engaging content from interesting people with real-world experience. And I think we did it! We’ve already got a few episodes in the can for 2021, and I’ve got to tell you there are some bangerz.
|
||
The Daily Check-In As the pandemic ramped up, I saw event after event being cancelled, and I was worried about my mental health and others. I work from home 100% of the time, with the sole exception being conferences and events. These are my main social avenues and when I get to interact with others. Removing that outlet was going to hit me hard, and I suspected others might have a similar reaction.
|
||
In fact, I thought it would be worse for those who normally work in an office. They were used to the constant social contact and casual interactions that I had learned to live without. Robbed of those situations, things were going to be difficult psychologically. With all that in mind, I created The Daily Check-In, a short, daily live-stream on YouTube where I talked about a prepared topic and interacted with anyone who joined the chat.
|
||
If you’ve watched any of the episodes, you’ll know at the beginning I check in with the viewers and let them know how I am doing. Sometimes people would join the live-stream and say hi, but mostly it was just me talking to myself. After a lot of streaming issues and a lack of interaction, it became easier to record my daily videos and publish them, instead of trying to live-stream the whole thing. I maintained the check-in at the beginning, and also tended to do every video in one-take.
|
||
The result was 178 videos on topics ranging from tech industry analysis, to professional development, to Terraform Tuesdays. Each day of the week ended up having a theme to make it easier for me to figure out a topic. From April to December, I managed to amass about one thousand subscribers, which is not a ton I know, but it was good enough for me.
|
||
At the end of the December, I decided that it was time to shift things on my YouTube channel. Looking at the analytics for my videos, I could see that my best performing content was around HashiCorp technologies and how-to videos with home lab stuff. I had also started to do Best Career Advice Ever with fun guests, and that seemed to be picking up steam. Rather than try to pump out half-decent content everyday, I decided it would make more sense to slow down and increase the quality of my output. Since it would no longer be daily, I would drop the Daily Check-in from the moniker and simply be Ned in the Cloud.
|
||
To increase the quality of videos will require investing some money into the channel for better equipment, editing help, and probably some other things. With that in mind, I have started a Patreon to help fund my efforts. If you’ve got a couple dollars to spare, why not join up? You’ll get regular updates on what is going on with me and a chance to ask me questions and influence future episodes!
|
||
Certification Guides Last year I started writing a Terraform Associate certification Guide with Adin Ermie. The guide was meant to prepare readers already familiar with Terraform for the certification exam. We took the approach of breaking out each major objective into its own chapter, and including some key takeaways at the end of each chapter to really hammer the point home. At this point we have sold just under one thousand copies! I’ve heard from multiple people on LinkedIn that they used a combination of our certification guide, Bryan Krausen’s prep questions, and my courses on Pluralsight to prepare for the exam and pass.
|
||
With the success of the Terraform cert guide, I decided to also write a guide for the HashiCorp Vault Associate certification. While the writing process is still ongoing, the guide is about 95% complete. I am currently proofreading the guide, and I have sent it to a few folks for technical review and accuracy. I imagine I will write one for Consul since the certification for that product just went GA.
|
||
As a companion to the Vault cert guide, I also created a series of videos on YouTube to help you prepare for the exam. They still have the standard daily check-in introduction, but I am working on adding timestamps to each one so you can skip straight to the content.
|
||
Pluralsight I published seven new courses in 2020, bringing my total course count up to 20. That’s not as many courses as 2019, but I also got to stretch out a bit. The Advanced AWS Networking course allowed me to expand my AWS expertise and get a little more familiar with the more esoteric concepts in AWS VPCs. I also updated my Terraform Deep Dive course to use Consul for remote state instead of AWS, and I updated the content to support Terraform 0.12. Recently I had to update all the exercise files to support Terraform 0.14 and the way it handles provider versioning. That’s the price you pay for creating content on a software platform that keeps evolving.
|
||
Speaking of evolving, my poor Azure Security Center content is now out of date again because Microsoft decided to change the product names to Microsoft Defender and completely reorganize the ASC menus. It was a long overdue update, but it also means I’ve got my work cut out for me in 2021 to update any courses that reference ASC. The good news is that under the covers, everything still works basically the same, so it’s more that I need to record the videos over again than actually create anything original.
|
||
GigaOm I continued working with GigaOm as an analyst, but my focus has shifted from writing Key Criteria reports to performing benchmarking tests. Although I was able to write the Key Criteria and Radar reports, they weren’t especially enjoyable for me. I like to be hands-on with technology, and a lot of the analysis stuff was taking vendor briefings and never actually using the product. By performing benchmarking tests, I can get my hands dirty and learn more about the product at a basic level. While that information might not be as valuable to a CEO or CIO - who is looking at tech trends, it is going to be useful for the IT Director or SysAdmin who is trying to make purchasing decisions about a tool or product and wants to know more about the technical implementation, verified capabilities, or actual cost of operation.
|
||
Docs Writing I spent all of 2020 writing docs for Solo.io part-time, and it was an excellent opportunity. Working with a startup that is on the front-line of cloud-native technologies, I got to see how products are developed, redesigned, and expanded. It also gave me a chance to work with the tech first-hand and write documentation that attempted to clearly explain difficult concepts to a new user. Being a new user myself, it was easier to capture my initial confusion and stumbling blocks, and then write docs that would hopefully make things easier for the next person. It was also fascinating and occasionally frustrating to write docs on products that were under active development. As I tried to write guides, I also became a bit of a QA tester and bug finder.
|
||
Starting in 2021, I will be leaving Solo.io to pursue a new opportunity that came up at the end of 2020. It has been a pleasure working with their team, and I highly recommend checking out their products if you have the need.
|
||
Various and Sundry There were lots of other random things that came up in the past year. Webinars, video content, speaking at virtual conferences, and more. I appreciate all the opportunities that cropped up as the year progressed. When I started Ned in the Cloud as a full-time venture, I was naturally concerned about keeping enough work in the pipeline to feed the family. Even with the pandemic, I did not have that worry this year. Instead, I had to be selective about what opportunities to accept and balance the cost of pursuing an opportunity versus the benefit of doing so.
|
||
And it’s not just about money. There is a satisfaction component and a desire to give back to the community. Basically everything I did for my YouTube channel this year was at the cost of doing something else that would be making me money. The satisfaction of creating content for you and trying to give back to the community with engaging professional and technical content was worth more than the money I could have made doing something else. I want to make that content even better, which means sacrificing even more time and potential earnings. But I think there is a net benefit in the long run for myself, the community, and vendors that I work with.
|
||
Conclusion There was no way for me to know what 2020 had in store, and all things considered, Ned in the Cloud did very well. I believe I stayed true to my vision statement for the company, and I’ve set a course to continue building on that vision in 2021. However, this post is long enough, so I’ll save those thoughts for another time. Thanks for reading and thanks for being you. I hope you emerged from 2020 relatively unscathed, and that 2021 will be a better year for all of us.
|
||
`,summary:"2020 was… a year. Or more like a decade packed into a single year. Looking back at everything that happened, it seems both too much and not enough. Conferences were cancelled, schools were closed, Zoom proficiency spiked, and we all added several new words to our vocabulary. This post is not meant to rehash our shared trauma of 2020, but I need to at least acknowledge it as a driving factor for the year.",date:"3 Jan, 2021",url:"https://nedinthecloud.com/2021/01/03/2020-year-in-review/",image:"2020-Year-in-Review.png",readingTime:"9"},"https://nedinthecloud.com/2020/10/19/hashicorp-boundary-dev-error-on-windows/":{title:"HashiCorp Boundary Dev Error on Windows",tags:[],content:`This is going to be a quick one. You’ve probably already heard about HashiCorp’s new Boundary project announced at HashiConf. If not, you can check out my YouTube video all about it.
|
||
When I tried to fire up the dev instance to take Boundary for a test drive, I immediately got an unpleasant error:
|
||
Error creating dev database container: unable to start dev database with dialect postgres: could not start resource: : Post "http://localhost:2375/images/create?fromImage=postgres&tag=12": dial tcp [::1]:2375: connectex: No connection could be made because the target machine actively refused it. I learned two things quickly:
|
||
The dev server for Boundary uses Docker Boundary was having trouble talking to my install of Docker Desktop If you’re in a similar boat, the fix is super easy! Open up the settings for Docker Desktop and tick this box:
|
||
Then click Apply & Restart. That’s it.
|
||
Told you it was easy! Have fun playing with Boundary! I know I will.
|
||
`,summary:`This is going to be a quick one. You’ve probably already heard about HashiCorp’s new Boundary project announced at HashiConf. If not, you can check out my YouTube video all about it.
|
||
When I tried to fire up the dev instance to take Boundary for a test drive, I immediately got an unpleasant error:
|
||
Error creating dev database container: unable to start dev database with dialect postgres: could not start resource: : Post "http://localhost:2375/images/create?`,date:"19 Oct, 2020",url:"https://nedinthecloud.com/2020/10/19/hashicorp-boundary-dev-error-on-windows/",image:"1602281050-boundaryhashiconf.png",readingTime:"1"},"https://nedinthecloud.com/2020/08/29/use-hashicorp-vault-aws-engine-with-multiple-accounts/":{title:"Use HashiCorp Vault AWS engine with multiple accounts",tags:["hashicorp-vault"],content:`I received a question recently on how to properly configure the AWS secrets engine on HashiCorp Vault to work with multiple AWS accounts. It took me a bit, but I did figure out how to do it and what the limitations are. In this post, I will break down how the secrets engine works and how to use it to dynamically create credentials across multiple AWS accounts using the assume_role feature.
|
||
I’m going to assume for the purposes of this article you are already familiar with HashiCorp Vault at a basic level. Like, you know what secrets engines and policies are. If not, check out my course on Pluralsight! And I’ll assume you know a little bit about AWS Identity and Access Management (IAM). Not expert level of course, IAM still makes my head hurt on the best days, but you know what an IAM role, policy, and user are at the very least. With that all out of the way, let’s talk a little bit about the AWS secrets engine in Vault.
|
||
AWS Secrets Engine The AWS secrets engine in Vault allows you to dynamically generate credentials in AWS through Vault. There are three credential types: iam_user, assumed_role, and federation_token. We’ll get back to those in a moment. When you enable an instance of the AWS secrets engine, you need to configure Vault access to AWS so it can generate these dynamic credentials. For this post, we will call this user vault-account. You’ll do this by creating an IAM user and generating an access key and secret key. The IAM user assigned to Vault needs sufficient permissions to perform actions that relate to the type of credential being generated. For iam_user credentials, it will need permission to perform actions like CreateUser and CreateAccessKey. For the assumed_role type, the vault-account really only needs the sts:AssumeRole action for any AWS roles it will be creating credentials against. I’m deliberately going to ignore the federation_token for the purpose of this post. You can find an example set of permissions for each credential type in the official AWS secrets engine docs.
|
||
Let’s talk about the first two credential types: iam_user and assumed_role. The iam_user type is probably the more typical and better documented. In Vault, you create a role on the AWS secrets engine for each iam_user type you want. When a user requests credentials, an IAM user is dynamically created on the same AWS account as the vault-account and an access key is returned for the user. Revoking the credentials in Vault destroys the IAM user on AWS. This works really well for a single AWS account and secrets engine, but what if you were working with multiple AWS accounts? What approaches are available to you?
|
||
Create an AWS secrets engine for each AWS account - This is feasible, but could quickly become a management nightmare, especially when it comes to policies. Grant IAM users cross-account roles - You can handle this entirely on the AWS side by creating roles in each AWS account and granting a group permission to assume that role. When the IAM user is created dynamically, it will be assigned as a member of the group. Use the assumed_role credential type, and grant vault-account permission to assume roles in different accounts. All three of these are viable options, but I would argue the cleanest and easiest is probably using the assumed_role credential type. The first option suffers from engine sprawl as the number of accounts increases. The second spreads the management of permissions and policies across both AWS and Vault. The third keeps the permissions management options entirely on the Vault side and has less overall configuration. The generated credentials also have a lower TTL, which helps with maintaining security.
|
||
The assumed_role credential type essentially has the vault-account requesting credentials for an AWS role defined in the AWS secrets engine role. Yes, both AWS and Vault use the word roles. And yes, it can be confusing. The AWS role can be in the same account as the vault-account or in a different account. You can create multiple Vault roles and specify multiple AWS roles in a single Vault role definition. The requestor - assuming they have proper permissions on Vault - performs a write operation on the Vault role and receives AWS credentials that are good for a limited amount of time, typically 60 minutes or less. You can restrict who has access to request credentials using Vault policies. For instance you could have a policy allowing all developers to request credentials to development AWS accounts, but not have any access to production AWS accounts.
|
||
With all that in mind, let’s actually get to the meat and potatoes of what we need to configure to get cross-account credentials from the assumed_role credential type.
|
||
Setting up cross-account credentials We are going to be setting up our AWS environment and a dev instance of Vault server to get the cross-account credentials working. If you want to follow along, you will need the following:
|
||
Two AWS accounts - primary and secondary Admin permissions in each AWS account The Vault executable The AWS CLI We are going to be performing the following steps to get things working:
|
||
Create the vault-account IAM user in the primary AWS account Create the IAM role in the secondary AWS account Grant the vault-account the AssumeRole permissions to the IAM role Start up a dev instance of the Vault server Enable the AWS secrets engine and configure it Create the Vault role and test it Let’s start by getting our AWS environment set up.
|
||
Configure the AWS environment First, we are going to set up our two AWS profiles, primary and secondary. Each will refer to a separate AWS account that you have admin access to.
|
||
aws configure --profile primary aws configure --profile secondary After each command, enter the Access Key, Secret Access Key, and default region for the profile.
|
||
Now we are going to create the vault-account IAM user and store the ARN in a variable for later use:
|
||
# Create the vault-account IAM user on the primary account vaultacct=$(aws iam create-user --user-name=vault-account --profile=primary) vaultarn=$(echo $vaultacct | jq .User.Arn -r) In the secondary account, we will create a role called ec2-admin that has full admin permissions on EC2, and attach an assume role policy that grant the vault-account permission to assume this role:
|
||
# Create the role with an assume policy in the secondary account cat << EOF > assume_policy.json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "AWS": "$vaultarn" }, "Action": "sts:AssumeRole", "Condition": {} } ] } EOF ec2admin=$(aws iam create-role --role-name=ec2-admin --assume-role-policy-document=file://assume_policy.json) # Grant the role the AmazonEC2FullAccess permission aws iam attach-role-policy --role-name=ec2-admin --policy-arn=arn:aws:iam::aws:policy/AmazonEC2FullAccess --profile=secondary We are capturing the ec2-admin role ARN in a variable as well, since we will need it for a policy that will be attached to the vault-account.
|
||
# Create the allow policy in the primary account ec2adminarn=$(echo $ec2admin | jq .Role.Arn -r) cat << EOF > allow_role.json { "Version": "2012-10-17", "Statement": { "Effect": "Allow", "Action": "sts:AssumeRole", "Resource": "$ec2adminarn" } } EOF allow_policy=$(aws iam create-policy --policy-name=allow-vault-ec2-admin --policy-document=file://allow_role.json --profile=primary) allow_policy_arn=$(echo $allow_policy | jq .Policy.Arn -r) aws iam attach-user-policy --user-name=vault-account --policy-arn=$allow_policy_arn --profile=primary Now we have an IAM user in the primary account with permissions to assume a role in the secondary account. Lastly, we will need to generate an access key for the vault-account. That will be used when we configure the AWS secrets engine on Vault.
|
||
# Get an access token for the vault-account to use with vault access_key=$(aws iam create-access-key --user-name=vault-account --profile=primary) key=$(echo $access_key | jq .AccessKey.AccessKeyId -r) secret=$(echo $access_key | jq .AccessKey.SecretAccessKey -r) Configure Vault We now have everything setup correctly on the AWS side. It’s time to configure Vault. In this example, we are going to fire up a dev instance of Vault server. Open a separate terminal window and run the following command:
|
||
vault server -dev Make note of the root token as we will need that to log into the Vault server. Back in your original terminal, run the following to set the address of the Vault server and login using the root token:
|
||
export VAULT_ADDR='http://127.0.0.1:8200' vault login It’s probably a good idea to point out that we are using the root token for these operations since this is a dev server instance. In a real world scenario, you would never use the root token for these operations. You probably already knew that, but I feel that it bears repeating.
|
||
Now we are going to enable the AWS secrets engine on the default path and configure it to use the vault-account:
|
||
vault secrets enable aws vault write aws/config/root \\ access_key=$key \\ secret_key=$secret \\ region=us-east-1 With the secrets engine configured, we can create a Vault role on the engine of type assumed_role and specify the ARN of the ec2-admin role we created in the secondary account:
|
||
vault write aws/roles/ec2-admin\\ role_arns=$ec2adminarn \\ credential_type=assumed_role Finally, we can request credentials from this role on the path aws/sts/ by issuing a write command. You could create a Vault policy that restricts who has permissions to execute write against this path or this specific role. We are running as the root login, so we can do anything we like. Let’s request a credential!
|
||
vault write aws/sts/ec2-admin ttl=60m You should receive a response similar to this:
|
||
Key Value --- ----- lease_id aws/sts/ec2-admin/KN8XkkYfHT0bLHbGBqLx9aQG lease_duration 1h lease_renewable false access_key ASIAQWO6FK2M3DWVP4KL secret_key 3070VHhjHNKvVVidiQ/5FQno2KzzJLjmdWMdwUV0 security_token FwoGZXIvYXdzEEgaDGfxR/2ttX6FIw7dCSLIAXAM7eBF32NXZUD3E5rkPKa/XQVJgc4rZfMWaiaSNIMkehOzGsjdy008befX20mHYlANiYaYDLd2Jp66ceSa/FPR4ev5GAgt8+mNjNrPYmSCx3VbZ5Gygi72XmvS0/T4GhDPuaflHz9nHsUdGeun1hjoWAK4VtISN1oB/xuTCz9cZ+nvgKDX73q9ueomtFExvgDYQmg/bJfkNnloHHo+pDK6x24x4OT5NAlAyZtNr3x2ExIq9N4IFbomz6KL/mhgXH6EN6f69duaKMDWqfoFMi0sqzp0soP+aYF537YleaaAayIF5X3DtIKEuDgavZIe8KQtLxg80cSbnMYm9jM= The lease duration specifies one hour, so these credentials will no longer be valid in 60 minutes, and the credentials are not renewable.
|
||
We should be able to use these credentials to do something like list all the VPCs in the region for the secondary account. Let’s try and do that.
|
||
export AWS_ACCESS_KEY_ID=ACCESS_KEY export AWS_SECRET_ACCESS_KEY=SECRET_ACCESS_KEY export AWS_SESSION_TOKEN=SESSION_TOKEN aws ec2 describe-vpcs --region=us-east-1 The easiest way to use the credentials is to set environment variables for the access key, secret key, and session token. You should receive a valid response including your default VPC in the region at the very least.
|
||
Conclusion In this post we examined the different credential types available to the AWS secrets engine and explored how to configure the assumed_role type for cross-account access. While the example was simple, involving only two AWS accounts, you could easily expand this pattern out to tens or hundreds of accounts with little additional effort. Vault policies would be used to govern account access, and you could use a tool like Terraform to configure the necessary accounts and permissions in each AWS account.
|
||
You would also want to setup logging on both Vault and AWS to correlate who is requesting access and what they are doing with it.
|
||
That does it for this post. If you have additional questions or thoughts, please leave them in the comments. And check out my weekly Vault certification videos on YouTube.
|
||
`,summary:"I received a question recently on how to properly configure the AWS secrets engine on HashiCorp Vault to work with multiple AWS accounts. It took me a bit, but I did figure out how to do it and what the limitations are. In this post, I will break down how the secrets engine works and how to use it to dynamically create credentials across multiple AWS accounts using the assume_role feature.",date:"29 Aug, 2020",url:"https://nedinthecloud.com/2020/08/29/use-hashicorp-vault-aws-engine-with-multiple-accounts/",image:"Vault-and-AWS.png",readingTime:"9"},"https://nedinthecloud.com/2020/05/07/ned-in-the-cloud-first-year-review/":{title:"Ned in the Cloud - First Year Review",tags:[],content:`One year ago I decided to quit my consulting job at a respectable company and go rogue as the Founder and sole employee of Ned in the Cloud LLC. I thought that now might be a good time to reflect on how I feel about that decision a year in, how things are going personally and financially, and where I see the company going in the future. Come with me friend on this introspective journey to examine the real-world of an independent content producer.
|
||
My Decision to Leave I guess I’ll start with how I feel about my decision to leave a stable position at a good company for the great unknown. It was with a fair amount of trepidation that I left the cozy confines of an employer. It meant I no longer had the comfort of a bi-weekly paycheck of a set amount. If I wanted to get paid, I was going to have to do things that people valued enough to pay me. That being said, I had developed a fair number of clients and revenue streams. So while money was a concern, it turned out not to be the primary issue.
|
||
People are social creatures, and as much as we IT folk like to joke about being introverts sequestered away in our cubicles doing our best to ignore end users and co-workers, the truth is all of us need social contact. The current pandemic puts this into sharp focus for all of us, but even before the pandemic, my decision to work for myself from home created a self-quarantine of sorts. I was not entirely psychologically prepared for the lack of interaction throughout the day. Small trips to the coffee machine. Dropping in on someone in their office. Chatting outside of a meeting room while you wait for the current occupants to leave. All of those small interactions create a background hum that is eerily satisfying, and one you only notice by its absence. Sort of like the thrum of an A/C unit that you only realize was there when it stops.
|
||
I miss those daily social interactions, and I struggled for a while to find a suitable alternative. A year in - pandemic and quarantine notwithstanding - I am still learning the balance.
|
||
But do I miss the job I left? Hmmmmm… No. I was working in a position I didn’t really enjoy. I felt unmotivated and unfulfilled. And now I don’t! That’s a win.
|
||
I give my decision to leave 4.5 stars out of 5.
|
||
How Things Are Going If I was concerned about finding enough work, I needn’t have been. When speaking to a colleague yesterday, I said something along the lines of, “You and I have accrued very specific skill-sets over the last decade, and those skills are in high demand.” Being able to think about technology critically, communicate clearly, and get work done on time are a rare combination of traits. Rarer than you might think. Through a certain amount of determination, foresight, and sheer luck, I have found myself in a position where people want to hear what I have to say. While I occasionally find this baffling, I also appreciate it for the gift that it is.
|
||
All of that is a long-winded way of saying that business is good. Very good. Rather than scrambling to find opportunities, instead I find myself turning things down. I simply don’t have the bandwidth. I’ve even conscripted a buddy of mine to help out with a couple projects. He’s not an employee in an official capacity, but I am going to be paying him money to do things.
|
||
People get weird about numbers, so I’ll just put it this way. In my first year of freelancing, I have already exceeded my salary at my previous employer and next year I expect about a 30% increase year-over-year.
|
||
You might be wondering how I found all this work? I think it all boils down to two primary things. The first is networking and the second is doing the thing. If you don’t mind, I’ll expand a bit.
|
||
Networking Remember that whole social creature thing I mentioned earlier in the post? Yeah, turns out that since we are social creatures, we tend to work with people that we know socially. That’s all networking is. When I hear a career development person - or, god help you, a “life coach” - talk about networking, it sounds like some esoteric process where you merge with the Borg collective by accepting their core tenets. And then you become one of us, the network anointed. The reality is so much simpler.
|
||
Networking is talking to people you know and the people they know. Not selling, pandering, pan-handling, evangelizing, proselytizing, pretending, faking, insinuating, genuflecting, or self-aggrandizing. It’s just talking to people and not being an asshole. You might recall that one of my three pillars is be nice, and I think you’ll find that being a pleasant person and talking to other people is a great way to expand your network - aka meet more people - and find new opportunities without selling yourself.
|
||
That’s not to say that you shouldn’t put yourself out there. The corollary to networking is to promote yourself. I think there’s a tasteful way to do this without sounding like a blowhard egomaniac. Partly this will be your blog, podcast, social media, etc. All of those platforms are important. People need to be able to see your work and remember that you exist. Even more important is the word of mouth effect. Once you’ve done a good job for one person, they are likely to recommend you to colleagues. That is a virtuous cycle.
|
||
You can think of your public presence as your resume, and your network as a recruiter. But more importantly, don’t treat them that way. Make friends. Care about people. Help out when you can. People will notice that you’re a genuine human being, and respond in kind.
|
||
Doing the Thing My buddy Stephen Foskett has this hierarchy of things you need to do to be successful. The first is to show up. Yes, it’s that simple. If you want to be successful, you actually have to show up. Seems like a fairly simple thing, but I can tell you from my experience managing people, they often fail at showing up.
|
||
The second is do the thing. Whatever your job is, do it. You don’t even have to do it especially well - more on that in a moment - you just have to do the thing. I know it sounds reductive, but again it is shocking how many people fail to do the thing they were hired to do.
|
||
The third is do the thing well and the fourth is do more than the thing. I’ll reserve those for another, more in depth post.
|
||
When someone hires me to make write content, create a course, make a video, I do the thing. Being reliable is incredibly powerful. The next time someone needs a video made for their webinar, an article written about their technology, a guest for a podcast, they are going to think of you. Because the last time they asked, you showed up and did the thing. I’ve been showing up and doing the thing for the last year, and now I am constantly being asked to do more of the things.
|
||
The Future The future is now! In the short term I need to deal with the pandemic and my new working situation. Working from home alone is a lot different than working from home with three kids and a spouse. And I’m having trouble focusing on things for long periods of time due to general stress and anxiety. That’s fine. I mean it’s not fine, but I can deal with it. The whole household is adjusting and adapting, and I’m still getting completed work out the door.
|
||
In the long term, I need to start thinking about ways that I can increase my capacity by farming out work to others. For example, I hired an editor for my last two courses on Pluralsight. It probably takes me 4-6 hours to edit one hour of recorded content. By having someone else do the editing, it frees me up to do more valuable work. Editing is something I can do, but it’s not the core value of my company. In the same vein, I can do accounting, but I hire an accountant because they are better at it than me and it saves me time.
|
||
If I want to increase my capacity for work, I need to take a serious look at where I can provide the most value and farm out the rest. I’m about to start a couple writing projects, and I have already asked a friend to assist with some of the writing. I’ll be paying him a portion of the fee that I’m getting for the work. Are there other areas I could farm out to others? I think there are and I will be vigilant for those opportunities as they arise.
|
||
When I was naming the company, I seriously struggled about whether to call it Ned in the Cloud or something else. Ned in the Cloud was already my website and I owned the domain. Two points in favor of using it. But the name implies that it’s just me, and not a larger organization. Right now that is true; however, there may come a day that I hire others on full-time to assist. When I made the decision, I couldn’t imagine hiring others to work for me. It just didn’t seem like a likely outcome. I was becoming a freelancer, and freelancers don’t have other full-time employees. It’s like a lone wolf pack. One year in, I’m looking at my current workload and the opportunities coming down the pike. The idea of adding some full-time or even part-time employees is looking a lot less ridiculous. That might be 3-5 years down the road, but I should start thinking about it now.
|
||
Summary Deciding to become an independent consultant and content creator was the best career move I could have made at the time. I was ready for the adventure and I had done the necessary planning to ensure a successful transition. This is exactly where I want to be at this point in my career and in my personal life. I’m sure circumstances will change, and if there’s anything the last few months have taught us, we have no idea what the future looks like. For now, I happy to report that my first year of Ned in the Cloud LLC was a resounding success and I am looking forward to see what the next year brings!
|
||
`,summary:"One year ago I decided to quit my consulting job at a respectable company and go rogue as the Founder and sole employee of Ned in the Cloud LLC. I thought that now might be a good time to reflect on how I feel about that decision a year in, how things are going personally and financially, and where I see the company going in the future. Come with me friend on this introspective journey to examine the real-world of an independent content producer.",date:"7 May, 2020",url:"https://nedinthecloud.com/2020/05/07/ned-in-the-cloud-first-year-review/",image:"NitC-First-Year.png",readingTime:"9"},"https://nedinthecloud.com/2020/05/06/mysterious-missing-region-argument-in-terraform/":{title:"Mysterious missing region argument in Terraform",tags:["hashicorp","terraform"],content:`I’m working on my next course for Pluralsight, Implementing Terraform on AWS. I probably don’t need to explain what the course is about. Anyhow, I was trying to show how you can create multiple instances of an AWS provider using the alias argument. Running through the initialization and validation process I ran into an error that was not very helpful.
|
||
Error: Missing required argument The argument "region" is required, but was not set. No mention of what line the error occurred on, or what resource in the configuration was throwing it. Just a missing region argument. Let’s see what’s going on here.
|
||
In the configuration I have two providers being defined:
|
||
provider "aws" { version = "~> 2.0" region = var.region alias = "infra" profile = "infra" } provider "aws" { version = "~> 2.0" region = var.region alias = "sec" profile = "sec" } It’s pretty clear that I am defining a region for each provider. So that’s not the issue. I thought maybe since I was using the profile argument, Terraform wanted to use the region set from the profile and not the provider. I tried omitting the region argument instead, and just ended up with more errors.
|
||
Error: Missing required argument The argument "region" is required, but was not set. Error: Missing required argument on main.tf line 18, in provider "aws": 18: provider "aws" { The argument "region" is required, but no definition was found. At this point it was time to turn up the logging and see what’s going on. I’m working in Windows, so I ran $env:TF_LOG="DEBUG" and reran the validation process. I won’t paste in all the output here, it’s amazingly verbose! But after looking for a bit, I found two sets of lines that gave me a clue as to what was going on.
|
||
2020/05/04 09:20:06 [TRACE] dag/walk: added new vertex: "provider.aws" 2020/05/04 09:20:06 [TRACE] dag/walk: added new vertex: "provider.aws.sec (close)" 2020/05/04 09:20:06 [TRACE] dag/walk: added new vertex: "provider.aws.infra" ... 2020/05/04 09:20:06 [TRACE] buildProviderConfig for provider.aws: no configuration at all Terraform was constructing three AWS providers even though I only defined two. And the third AWS provider had no configuration, which means that Terraform was implicitly creating it for a resource. I looked through the configuration and sure enough:
|
||
resource "aws_iam_group_policy" "peering-policy" { name = "peering-policy" group = aws_iam_group.peering.id policy = <<EOF { "Version": "2012-10-17", "Statement": { "Effect": "Allow", "Action": "sts:AssumeRole", "Resource": "\${aws_iam_role.peer_role.arn}" } } EOF } There is no provider argument in the configuration block. Since both of my explicitly defined AWS providers have an alias, Terraform was creating a third, non-alias provider for this resource. Unfortunately, the implicit provider had no configuration data and I had not set any environment variables that Terraform could use.
|
||
Terraform was correct, I was missing the region argument for the provider, and it could not provide a line reference because the provider was being created implicitly.
|
||
I updated the configuration with a provider argument and the error went away.
|
||
resource "aws_iam_group_policy" "peering-policy" { name = "peering-policy" group = aws_iam_group.peering.id provider = aws.infra policy = <<EOF { "Version": "2012-10-17", "Statement": { "Effect": "Allow", "Action": "sts:AssumeRole", "Resource": "\${aws_iam_role.peer_role.arn}" } } EOF } I’ve never been a fan of the implicit provider creation feature of Terraform. The first time I saw it in a demo, I thought, “Well that’s confusing and possibly problematic.” Turns out I was right, but I don’t expect HashiCorp to change it now.
|
||
The moral of the story is that the debug logs for Terraform are your friend. When faced with an inscrutable error, go ahead and turn them on and take your time parsing through. There’s a lot of information, and buried in there will be the answer to your problem.
|
||
`,summary:`I’m working on my next course for Pluralsight, Implementing Terraform on AWS. I probably don’t need to explain what the course is about. Anyhow, I was trying to show how you can create multiple instances of an AWS provider using the alias argument. Running through the initialization and validation process I ran into an error that was not very helpful.
|
||
Error: Missing required argument The argument "region" is required, but was not set.`,date:"6 May, 2020",url:"https://nedinthecloud.com/2020/05/06/mysterious-missing-region-argument-in-terraform/",image:"Terraform-missing-region.png",readingTime:"3"},"https://nedinthecloud.com/2020/05/04/the-terraform-certified-study-guide/":{title:"The Terraform Certified Study Guide",tags:["hashicorp","terraform","terraform-certified"],content:`As I mentioned in a previous post, HashiCorp has officially announced the availability of two certifications, Terraform Certified Associate and Vault Certified Associate. In that post I detailed a bunch of different resources to help you study for the Terraform exam. One of those resources was a study guide that Adin Ermie and I put together called the HashiCorp Terraform Certified Associate Preparation Guide, which does not lend itself well to an acronym - HTCAPG? I guess we could go with Hat Cap? Nah. Anyway, I thought I would give you an idea of what is in the guide, and a free sample of a few pages.
|
||
The Guide The certification is broken into a list of high-level objectives. Each of the high-level objectives is composed of enabling objectives. There are nine high-level objectives in total, and so Adin and I chose to dedicate a chapter to each objective. Within each chapter is an introduction, and then a section dedicated to each enabling objective. At the end of each section is a Key Takeaway for that enabling objective to help you focus on what was most important. At the end of each chapter is a list of the Key Takeaways for each section, so you can quickly reference them when preparing for the exam.
|
||
The whole point of the guide is to help prepare you to take the exam. It is not meant to teach you Terraform from the ground up. We assumed that the reader would already have a passing familiarity with Terraform, and was looking for a resource that would help them focus their studying on what would actually be in the exam. I think we were fortunate that HashiCorp chose topics that were practical and useful in day-to-day Terraform operations and not esoteric arcana pushed by a marketing group with an agenda. This was an exam clearly steered by SMEs and boots on the ground operations people. I know since I was a contributor to the exam.
|
||
Let’s take a closer look at a typical chapter, starting with a favorite of mine - chapter 5, which deals with the high-level objective: Understand Terraform Basics.
|
||
Introduction Here’s the intro to the chapter:
|
||
Before you can get started with Terraform, you’ll need to install it somewhere. You also can’t do a whole lot without using the various and sundry providers that enable Terraform to talk to cloud providers, platform services, datacenter applications, and more. We’ll also touch on provisioners, and when to use them. You might find the answer counterintuitive!
|
||
Solid. We get an overview of what’s in the chapter. If you’re feeling pretty competent in this topic, you could skip it. But what about that mention of provisioners? You might decide to jump to the section entitled: Explain When to Use and Not Use Provisioners and When to Use Local-Exec or Remote-Exec.
|
||
Enabling Objective Let’s see what’s in that provisioners section:
|
||
When a resource is created, you may have some scripts or operations you would like to be performed locally or on the remote resource. Terraform provisioners are used to accomplish this goal. To a certain degree, they break the declarative and idempotent model of Terraform. For this reason, HashiCorp recommends using provisioners as a last resort only.
|
||
That probably bears repeating, because HashiCorp has changed their stance over time on using provisioners. We can’t guarantee any particular question will be on the exam, but HashiCorp seems to feel passionately about this. To repeat:
|
||
Don’t use provisioners unless there is absolutely no other way to accomplish your goal.
|
||
Use cases What are the use cases for a provisioner anyway? There’s probably three that you’ve already encountered, or will in the near future.
|
||
Loading data into a virtual machine Bootstrapping a VM for a config manager Saving data locally on your system The remote-exec provisioner allows you to connect to a remote machine via WinRM or SSH and run a script remotely. Note that this requires the machine to accept a remote connection. Instead of using remote-exec to pass data to a VM, use the tools in your cloud provider of choice to pass data. That could be the user_data argument in AWS or custom_data in Azure. All of the public cloud support some type of data exchange that doesn’t require remote access to the machine.
|
||
In addition to the remote-exec provisioner, there are also provisioners for Chef, Puppet, Salt, and other configuration managers. They allow you to bootstrap the VM to use your config manager of choice. A better alternative is to create a custom image with the config manager software already installed and register it with your config management server at boot up using one of the data loading options mentioned in the previous paragraph.
|
||
You may wish to run a script locally as part of your Terraform configuration using the local-exec provisioner. In some cases there is a provider that already has the functionality you’re looking for. For instance the the Local provider can interact with files on your local system. Still there may be situations where a local script is the only option.
|
||
Provisioner details Despite your best efforts, you’ve decided that you need to use a provisioner after all. A provisioner is part of a resource configuration, and it can be fired off when a resource is created or destroyed. It cannot be fired off when a resource is altered.
|
||
When a creation-time provisioner fails, it sets the resource as tainted because it has no way of knowing how to remediate the issue outside of deleting the resource and trying again. This behavior can be altered using the on_failure argument. When a destroy-time provisioner fails, Terraform does not destroy the resource and tries to run the provisioner again during the next apply.
|
||
A single resource can have multiple provisioners described within its configuration block. The provisioners will be run in the order they appear.
|
||
Of course you can avoid all this nonsense by not using provisioners in the first place. We’re just sayin'.
|
||
Supplemental Information * https://www.terraform.io/docs/provisioners/index.html * https://www.terraform.io/docs/provisioners/local-exec.html * https://www.terraform.io/docs/provisioners/remote-exec.html
|
||
Key Takeaways: Provisioners are a measure of last resort. Remote-exec will run a script on the remote machine through WinRM or ssh, and local-exec will run a script on your local machine.
|
||
Wow. That’s good to know. If I see a question about provisioners on the exam, the answer is probably not to use them unless they match one of the use cases.
|
||
Summary I hope you’ve enjoyed this sneak peak into what is in the guide and how it is constructed. If you’re interested in picking up a copy, head on over to the Leanpub page. If you’re looking for a physical copy, let me know! Right now the guide is entirely digital, but if enough people ask I will get it set up with a physical publisher as well.
|
||
`,summary:"As I mentioned in a previous post, HashiCorp has officially announced the availability of two certifications, Terraform Certified Associate and Vault Certified Associate. In that post I detailed a bunch of different resources to help you study for the Terraform exam. One of those resources was a study guide that Adin Ermie and I put together called the HashiCorp Terraform Certified Associate Preparation Guide, which does not lend itself well to an acronym - HTCAPG?",date:"4 May, 2020",url:"https://nedinthecloud.com/2020/05/04/the-terraform-certified-study-guide/",image:"Terraform-certified.png",readingTime:"6"},"https://nedinthecloud.com/2020/04/27/preparing-for-the-hashicorp-terraform-certification/":{title:"Preparing for the HashiCorp Terraform Certification",tags:["hashicorp","terraform","terraform-certified"],content:`HashiCorp has recently announced the availability of the Terraform Certified Associate exam. This is an excellent way to assess your skills and demonstrate your competence with the Infrastructure as Code tool, Terraform. Those who have been following me for any period of time know that I am a pretty big fan of Terraform, and may have authored more than a few posts and courses on the topic. What you might not know is that I was actively involved in writing and reviewing the questions for the exam. In this post, I will give you an overview of what to expect in the exam, how I think you should study for it, and some materials to help you along the way.
|
||
The Certification The HashiCorp Certified Terraform Associate exam is meant to test that you have an associate level of experience and knowledge with Terraform. But what does an associate level really mean? It means that you are comfortable with the core concepts of Infrastructure as Code and the main underpinnings of Terraform. It means that you are comfortable using Terraform in a non-production environment and you have created enough configurations to understand how the tool works. You’ll also need some knowledge of what Terraform Enterprise and Terraform Cloud do, although you do not need practical experience with either.
|
||
HashiCorp has laid out the expectations for the certification with nine high-level objectives. Each objective has enabling objectives that further detail what you should know about a subset of the high-level objective. Let’s take a look at one of the high-level objectives:
|
||
Implement and maintain state
|
||
Implementing and maintaining state is a broad topic and could include a lot of basic and advanced information. The first enabling objective is:
|
||
Describe default local backend
|
||
Ah, okay. So we are dealing with the default local backend. That narrows things down considerably. What do I need to know about the default local backend? Well, it’s the default. If I don’t specify a backend, Terraform will use this one. It supports state locking. It uses the local filesystem, and creates the file terraform.tfstate in the directory of the configuration, unless I tell it otherwise. There’s a bit more to it, but I think that’s a pretty good start. You can always read the official docs for more information, or try out Terraform yourself to see how the local backend works.
|
||
I’ve taken a lot of certifications in my time as an IT practitioner, and I have to tell you that many vendors are vague and ambiguous about what is going to be on the exam. That is not the case with HashiCorp. It is very clear what will be included. I can tell you from experience, questions that strayed from the core objectives were rejected. This exam is not meant to be tricky or clever. It is meant to test your knowledge on the objectives defined by HashiCorp. I wish every exam was designed this way.
|
||
The Exam The exam itself is a combination of true/false, multiple choice, and multi-select questions. It is not a practical exam where you are presented with a command line and a task. The exam is administered with a remote proctor, which means you’ll need a webcam and a clear workspace. It is a pass/fail type exam, so you’ll need to get a certain amount of questions right to pass. Just like any modern exam, you get the results immediately after taking the exam.
|
||
If you want to hear more about my thoughts on taking a remote exam, I recorded a whole YouTube video just about that.
|
||
Studying for the Exam There are many ways to study for an exam. You can read a book, watch a video, or do some practice questions. I have used a combination of all of these and more to prepare for previous certification exams. Below are some recommended resources you can use to study and prepare. Note, I do not earn commissions from any of these links, although I may earn money if you purchase something I wrote or created.
|
||
Books - Books are a great way to study the core concepts of Terraform. Here are two books that I think do a great job of getting you ready: Terraform: Up & Running The Terraform Book Courses - Courses and walkthroughs are a great way to absorb the information and get some keyboard time: Terraform - Getting Started: My course on Pluralsight Learn Terraform: The official learning site by HashiCorp Study guides and practice exams - These are focused on the exam itself and should probably be used to augment other learning: HashiCorp Terraform Certified Associate Preparation Guide: A guide I wrote with Adin Ermie Study Guide - Terraform Associate Certification: Official guide from HashiCorp HashiCorp Certified: Terraform Associate Practice Exam: Practice Exam by Bryan Krausen All of these resource are great, and I have first-hand knowledge of using them or creating them. There are plenty of other resources out there, but I cannot speak directly to their content. In addition to using these study materials, I think it’s important to get actual keyboard time with Terraform. By tinkering around with the product, you will learn how the tool actually works in multiple contexts. The easiest way to get started is to simply download Terraform onto your workstation and create a configuration. Another quick way is to use the Azure Cloud Shell.
|
||
The Cloud Shell in Azure has many tools pre-installed when you launch it, including Terraform and Git. It also inherits the Azure AD credentials you used to launch the Cloud Shell, so you don’t have to worry about provider authentication for Azure resources. And the Cloud Shell has a built-in code editor, meaning you don’t even need to edit the config files locally. If you want to get started quickly with minimal friction, create an Azure account and launch Cloud Shell. No local installation, and no local dependencies.
|
||
Summary I can tell you from personal experience that Terraform’s popularity has skyrocketed. The launch of a formal certification for Terraform is a major milestone for HashiCorp, and I suspect the cert will be incredibly popular. The good news is that the exam was created with fairness and transparency in mind. What you see in the objectives is what you get in the exam. Between the resources above and practice time with the Terraform itself, I think you will be well positioned to pass the exam and achieve the Terraform Associate certification. Good luck, and let me know how you do!
|
||
`,summary:"HashiCorp has recently announced the availability of the Terraform Certified Associate exam. This is an excellent way to assess your skills and demonstrate your competence with the Infrastructure as Code tool, Terraform. Those who have been following me for any period of time know that I am a pretty big fan of Terraform, and may have authored more than a few posts and courses on the topic. What you might not know is that I was actively involved in writing and reviewing the questions for the exam.",date:"27 Apr, 2020",url:"https://nedinthecloud.com/2020/04/27/preparing-for-the-hashicorp-terraform-certification/",image:"Terraform-certified.png",readingTime:"6"},"https://nedinthecloud.com/2020/01/19/enabling-conditional-access-for-azure-active-directory-applications/":{title:"Enabling Conditional Access for Azure Active Directory Applications",tags:["azure","azure-ad"],content:`I’m in the process of updating my Managing Identities in Azure Active Directory course on Pluralsight. One of the demos in the course is configuring Conditional Access for an Azure Active Directory integrated application. The idea is that you can set up a Conditional Access policy that restricts users from logging into the application from outside the US. When I went to go record the updated demo, the application I had created in Azure AD was missing. What followed was a journey into the bowels of Azure AD to find what triggers the appearance of an app in Conditional Access.
|
||
Update: Microsoft has fixed the bug. Hurray!
|
||
Earlier in the course I create an application called Marketing App using PowerShell and the AzureAD module. The commands I run are as follows:
|
||
$app = New-AzureADApplication -DisplayName "Marketing App" -IdentifierUris "http://marketing.contosoned.xyz" New-AzureADServicePrincipal -AppId $app.AppId The commands will successfully create an application in Azure AD and give it a Service Principal to be used for SSO.
|
||
In a later module, I walk through the process of configuring a Conditional Access policy to control access to an Azure AD application. In the original demo I used the Marketing App, but this time when I got to the Cloud Apps section of the policy the Marketing App was not there. There is no indication of why some apps appear and others do not. Looking at the documentation on the Microsoft Docs site, the guides walk you through adding an application, but they presuppose that the application already appears in this Cloud Apps menu.
|
||
I sent out a Tweet to the Microsoft MVP peeps. I sent an email to the MVP distribution lists. There were some nice suggestions, but nothing panned out. I stepped away from the keyboard for a bit. It occurred to me that it might be a licensing issue. Conditional Access for Azure AD apps requires at least an Azure AD Premium 1 license. When I created the Marketing App, I had not yet purchased the Azure AD Premium license. That was something I did a day or two later. Maybe there was something in the application creation process that checks to see if you have an Azure AD Premium license in the tenant, and flags the application as licensed for Conditional Access? To test my theory, I deleted the existing app from the portal and created a new Marketing App from the portal.
|
||
Huzzah! Now my Marketing App appeared in the Cloud Apps list. It must have been the missing Azure AD Premium licensing when I created the first app. Well that’s just dumb, and there should be a way to fix it without deleting the app. But at least I had the solution right? Right?
|
||
No.
|
||
Eagle-eyed observers will notice that the first time I created the application it was using PowerShell, and the second time was through the portal. Suffice to say, those workflows do not produce the same results. This occurred to me later in the day, so I decided to test my theory. I created a new Marketing App v2 using PowerShell, and sure enough that application did not appear in Cloud Apps. I also noticed that the Marketing App I created with the portal had Conditional Access available when viewing it as an Enterprise Application, and v2 of the Marketing App did not.
|
||
Clearly there was some difference between how the portal was creating my application, and how the commands in PowerShell were doing it. Essentially, the portal was making some implicit assumptions, without telling me. The first one I discovered was in the configuration of the Service Principal. When you register an application in the portal, it automatically creates a Service Principal for you. When using PowerShell, you first create the application in one command and then the Service Principal in a separate command. I compared the properties of the Marketing App and Marketing App v2 Service Principals and determined that the former had the entry WindowsAzureActiveDirectoryIntegratedApp in the Tag property. The v2 Service Principal had no entry in its Tag property. Sure enough, looking at the docs for the New-AzureADServicePrincipal command it notes the following for the -Tags parameter:
|
||
Note that if you intend for this service principal to show up in the All Applications list in the admin portal, you need to set this value to {WindowsAzureActiveDirectoryIntegratedApp}
|
||
Wow, so that’s not obvious at ALL, but there it is in the docs. I recreated v2 of the application by running the following:
|
||
$app = New-AzureADApplication -DisplayName "Marketing App v2" -IdentifierUris "http://marketingv2.contosoned.xyz" New-AzureADServicePrincipal -AppId $app.AppId -Tags "WindowsAzureActiveDirectoryIntegratedApp" Going back to the portal, I now see both versions of the application in the Enterprise Applications - All Applications view.
|
||
The v2 of the application was not in this view before, so basically the WindowsAzureActiveDirectoryIntegratedApp tag is what marks an application as an Enterprise Application. Good to know.
|
||
Looking at the v2 entry details, I now can see the Conditional Access tab. Great!
|
||
Things aren’t great though. If I click through on that Conditional Access button and try to create a policy for the application, I see this weird error on the Cloud App selection screen.
|
||
Huh? I’m going to quote the error so people can find this through web searching.
|
||
1 cloud apps configured in this policy have been deleted from the directory, but this doesn’t affect the other apps in the policy. The next time you update the application section of the policy, the deleted apps will be automatically removed from it.
|
||
What the hell? The application definitely has not been deleted. I got here from its details page. Something is - to use the technical term - hinky with this error. I went back and compared the properties of the two Service Principals again, and went down a few dead ends having to do with Oauth2 permissions.
|
||
I won’t bore you with the details. Suffice to say, it was a waste of time, but it did point me back at the Application objects. After comparing the properties of the two application objects, I realized that my original Marketing App had a ReplyURL set, whereas my v2 app did not. It appears that the Conditional Access portal needs the app to have an entry in the ReplyURLs property, or it won’t list the application in the Cloud Apps.
|
||
Basically, the details page for the Enterprise Application will show the Conditional Access link if the Service Principal Tags property has the WindowsAzureActiveDirectoryIntegratedApp value in it. The Cloud Apps selector appears to filter out any applications that do not have an entry in the ReplyUrls property. Using the Conditional Access link on the details page of the v2 application circumvented that check, but once it got to the actual policy the check was made and the application couldn’t be found due to it lacking a ReplyUrl entry. The Cloud Apps selector makes the logical assumption that the application was selected at one point, but it is now deleted, and so it shows the error above.
|
||
This was a pretty easy fix. I deleted the v2 application and recreated it with the following commands:
|
||
$app = New-AzureADApplication -DisplayName "Marketing App v2" -IdentifierUris "http://marketingv2.contosoned.xyz" -ReplyUrls @("https://marketing.contoso-ned.xyz/.auth/login/aad/callback") New-AzureADServicePrincipal -AppId $app.AppId -Tags "WindowsAzureActiveDirectoryIntegratedApp" Now the application shows up in the Cloud Apps selector, and I can finish my demo.
|
||
To sum up. When you create an application and Service Principal using PowerShell, you must include at least one entry in the ReplyUrls property of the application and set the Tags property of the Service Principal to be WindowsAzureActiveDirectoryIntegratedApp. It’s weird, completely non-intuitive, and poorly documented. Hopefully, I’ve saved some others a couple hours of banging their head against a wall.
|
||
`,summary:"I’m in the process of updating my Managing Identities in Azure Active Directory course on Pluralsight. One of the demos in the course is configuring Conditional Access for an Azure Active Directory integrated application. The idea is that you can set up a Conditional Access policy that restricts users from logging into the application from outside the US. When I went to go record the updated demo, the application I had created in Azure AD was missing.",date:"19 Jan, 2020",url:"https://nedinthecloud.com/2020/01/19/enabling-conditional-access-for-azure-active-directory-applications/",image:"Azure-AD-SP.png",readingTime:"6"},"https://nedinthecloud.com/2020/01/01/a-vision-and-strategy-for-2020/":{title:"A Vision and Strategy for 2020",tags:[],content:`Although I don’t generally subscribe to New Year’s resolutions, I do like to review my professional goals on a regular basis and make sure they align with my overall strategy and vision for my career. I suppose that the beginning of a new year is a useful reminder to check in and see how things are going. In a previous post, I took a look at my goals for 2019 and how I did on achieving those goals. I also mentioned how I needed to revise those goals for 2020 based on my new circumstances, i.e. being self-employed.
|
||
Before I can even formulate new goals for the year, I think I need to spend a little time figuring out what the long term vision is for Ned in the Cloud LLC. The vision creates a strategy, the strategy determines goals. The fundamental question is, what is the vision for me?
|
||
The Vision When I think of a vision or a mission statement, I naturally think of Jerry McGuire. Most people think of “Show me the money!” or “You had me at hello,” but I think of the beginning of the movie where the titular character writes a vision statement/manifesto. It is… not well received. And it becomes the butt of many jokes. I also think of the Weird Al song Mission Statement that successfully lampoons the vapid corporate speak prevalent on most companies’ websites when you hit up the About Us page.
|
||
I don’t want to write a mission statement that is vapid and superficial to the point of meaninglessness. I also don’t want to write some earnest and heartfelt manifesto fueled by too much cough syrup and insomnia. A vision should be concise, pragmatic, and actionable. With that in mind, I have come up with this:
|
||
Ned in the Cloud’s vision and purpose is to create compelling technical content for IT professionals across multiple mediums based on impactful and emerging technologies.
|
||
At one sentence, the statement is certainly concise. I mean, you wouldn’t put it on a coffee mug or t-shirt, but it’s not War & Peace either. It’s also pragmatic. I’m not trying to change the entire world, unlike Microsoft’s vision of “Our mission is to empower every person and organization on the planet to achieve more.” Microsoft is a trillion dollar company, I’m one person. There’s an implied action in the statement too. I am going to create compelling technical content. The content could be video, audio, written, or some new medium that does not yet exist. The medium is less important than the goal of creating something that IT professionals will find compelling and informative.
|
||
The Strategy Now I’ve got a vision, so I can formulate a strategy to achieve that vision. There are three things to pick apart here.
|
||
Create compelling technical content Across multiple mediums Impactful and emerging technologies Create compelling technical content Compelling content doesn’t just happen, it needs to be crafted and tested. I’m sure some people are natural born communicators and entertainers. But the rest of us need to work at it. Creating compelling content means that I need to find a way to improve my skills across a range of skill-sets including communication, presentation, instructional design, narrative arcs, storytelling, etc. Since I am defining a strategy here, I am not so much concerned with how I am going to improve my skills, it’s more about what I need to improve. I think I can identify a few skills that will help me across all mediums:
|
||
Public speaking and presentation: Even if things are pre-recorded, becoming a better public speaker will improve my delivery of content in all formats. Instructional design: While not all of my content will be educational in nature, much of it will. Having a solid grounding in instructional design will help me be more effective when delivering training, but will also help me craft more compelling content in general by focusing my delivery. Creative writing: Making compelling content requires a certain level of creativity. The best delivery of well-crafted instructional design is likely to fall flat if there is no hook to draw in the consumer. Successful content will be creative content, and becoming more skilled in creative writing can only help. Across multiple mediums There are few people that only consume content through a single medium. Folks like to read, listen, and watch. Sometimes they are looking for something interactive, like a hands-on lab. Other times they want to learn about a new topic while commuting to work. Perhaps they want to read a physical book and make notes in the margins. Creating content in a single way, on a single medium is simply not going to work. But I cannot be master of all mediums, so I need to be strategic about what media channels I want to pursue. This one is fairly simple since I already create content, but of course there is room for improvement.
|
||
Video training courses: This is the bread and butter of Ned in the Cloud. I create courses for Pluralsight, and I plan to continue doing so for the foreseeable future. My primary focus here should be on improving the production quality of my content and finding ways to streamline the production process. Podcasting: I already run two podcasts, Day Two Cloud and Buffer Overflow. I have no plans to add another to the roster. Instead, just like the Pluralsight courses, I should endeavor to improve the production and content of both podcasts. Technical writing: There’s more than one format for writing. I have my blog, I write posts for technology vendors, I write analyst reports, and I write documentation for software. Oh and I have co-authored a book. The best approach is improving the writing quality and keep on doing what I am doing. Video: Just like writing, there are multiple video types. There are panel conversations, how-to walkthroughs, lightboarding, live-streaming one-on-one chats. Honestly, I haven’t done much with video beyond Pluralsight and the occasional panel chat here and there. Many people prefer video as their primary medium, and I need to embrace that in the next year. Impactful and emerging technologies The landscape of technology is ever evolving. New applications, services, and products are constantly being introduced. I cannot possibly cover all of the new developments, so I need to be strategic about which topics I focus on for my content. They need to be impactful, meaning that the topic/product/solution makes an actual difference in the life of the IT practitioner.
|
||
Impactful It’s hard to pin down exactly what makes a given technology impactful, but I guess I know it when I see it. I suppose I don’t want to focus on something that is a niche product with a limited audience. An edge application designed for the medical industry is probably too specific for me and not impactful for IT practitioners on the whole. The edge platform that hosts the medical application - whether its Kubernetes, VMware, or something else - is of more interest to me, and probably to most IT practitioners.
|
||
Emerging I got into this technology racket to play with cutting edge solutions, not to upgrade the same old application for the umpteenth time. I think people are going to find more value in content that addresses new and emerging solutions and examines how they might actually fit into the life of an IT practitioner. For that reason, I need to stay current with what is coming down the pike and take time to talk to industry leaders and startups who are producing the innovative technology today that will become the standard tomorrow.
|
||
Goals My goals for 2020 should be directly informed by the main points of my strategy. I need to find achievable goals that align with creating compelling content across multiple mediums on impactful and emerging technologies. This post is probably long enough already, and I need to have a good think about what my goals might be. In my next post, I will share what goals I have come up with and how I plan to achieve them.
|
||
`,summary:"Although I don’t generally subscribe to New Year’s resolutions, I do like to review my professional goals on a regular basis and make sure they align with my overall strategy and vision for my career. I suppose that the beginning of a new year is a useful reminder to check in and see how things are going. In a previous post, I took a look at my goals for 2019 and how I did on achieving those goals.",date:"1 Jan, 2020",url:"https://nedinthecloud.com/2020/01/01/a-vision-and-strategy-for-2020/",image:"featured-image-nedinthecloud.jpg",readingTime:"7"},"https://nedinthecloud.com/2019/12/31/the-2010s-a-decade-in-review/":{title:"The 2010s: A Decade in Review",tags:["aws","azure","buffer-overflow","gcp","kubernetes"],content:`In the most recent episode of Buffer Overflow, we talked about the biggest tech trends for the 2010s. I thought I would expand on my thoughts a little bit with this post. Check out the full episode below.
|
||
[embed]https://www.anexinet.com/wp-content/uploads/2019/12/BufferOverflow-Episode140.mp3[/embed]
|
||
Public cloud accelerates everything In the 2010s public cloud computing exploded and enabled massive technological innovation, particularly in the realm of Software as a Service (SaaS). In addition to providing a common platform for startups and enterprises alike to build on, it also served as a example to existing datacenter providers and in-house IT operations teams on how to effectively run a cloud service at scale using automation and standardization. We are now seeing the impact of cloud-native operations infiltrate the on-premises datacenter, a trend I think will further drive innovation in the 2020s.
|
||
Some background Public cloud didn’t just come out of nowhere, it began its life as a relatively simple set of services offered by companies like Amazon and Google. Here’s a few representative data points leading up to the 2010 explosion.
|
||
2004: SQS was the first service available through Amazon. This predates AWS. 2006: S3 and AWS are launched, EC2 comes later that year with SimpleDB in 2007. 2008: Google launches App Engine to compete (this still exists!) 2010: Microsoft launches Azure with basic cloud app hosting At the start of 2010 the three major cloud providers had started their roll. AWS was launched in 2006, Google made their foray in 2008 with App Engine, and Microsoft launched Azure with the Cloud Services offering in 2010. The primary focus for all of these clouds was providing a web application hosting platform on which to build applications.
|
||
Each of the public clouds realized that customers wanted a lower level offering in the stack, specifically Infrastructure as a Service (IaaS), which would provide the datacenter components people were used to - network, compute, and storage. There can be no doubt that AWS was first to the party and still holds the lion’s share of customers in 2019. Microsoft has been making significant headway, driven by the strategy adopted by Satya Nadella when he took the throne in 2014 with a cloud and services first attitude. GCP is a distant third, in large part due to the immaturity of their platform, their poor track record working with enterprises, and a general uncertainty that GCP will exist in five years time and not be unexpectedly cancelled during a Google Developers Keynote.
|
||
Public cloud making money With the big three ready to go in 2010, public cloud computing was still only a $15B industry. In 2019 it is expected to hit $228B. That’s 15x growth over ten years. And that doesn’t account for all of the the SaaS applications that are hosted using Platform as a Service (PaaS) and IaaS.
|
||
It’s interesting to look at the most recent yearly financial reports from Microsoft and Amazon to get a sense of just how big things have become with cloud. Microsoft is touting $38B in revenue for their Intelligent Cloud division for FY2019, and Amazon posted $25.66B for AWS in FY2018 with a projected total of $34B for 2019.
|
||
Please don’t email me about how Microsoft’s number is bigger than AWS. I know it is. And I also know that Microsoft includes more than just Azure in their Intelligent Cloud BD. I also don’t care. Any way you slice it, it’s a lot of money!
|
||
Ancillary effects It’s more than just money though. The public cloud made it trivial for startups to consume technology in a non-capital intensive way. And it gave growing companies a global reach without them having to build out datacenters across the world or negotiate with ISPs in far flung geographies. You need a presence in South Korea, boom you got it. This, along with mobile devices, changed the landscape of software consumption in the world.
|
||
I just can’t stress this enough. Without AWS, there is no Netflix, no Dropbox, no Uber. Thousands of SaaS companies, gaming companies, and websites would not and could not exist. I can run chrisistheworst.com out of S3 for basically nothing. The most expensive thing is not the storage or the web hosting. It was the $2 I paid for the domain. And that dumb website could easily serve hundreds of thousands of requests per minute without me having to do anything. That is immense power at my fingertips for essentially no capital investment.
|
||
Public cloud has become the platform on which other businesses can be built, and just like any good platform or utility, we only notice it when something goes wrong. When your power goes out, your plumbing fails, or the WiFi acts up then you notice. Otherwise, all of these services fade into the background and allow us to live our normal modern lives. While public cloud computing is not quite as pervasive and utilitarian as those other services, it is also not far off.
|
||
Standardize and devour There is one thing the public cloud doesn’t have today, and that’s standardization. Each public cloud does things in a slightly different way when it comes to IaaS. When it comes to PaaS and SaaS, the differences increase dramatically. The upside is that it enables incredible innovation and application development. The downside is that users are locked to that provider. OpenStack was intended to prevent this type of lock-in by creating a set of open standards for cloud providers. That, um, didn’t work out as planned.
|
||
Even though public cloud hasn’t standardized on a single open standard, the providers made the consumption of computing services friendly and simple. So simple that a whole new class of IT rose called shadow IT. People in your organization using cloud services without IT approval, and this problem will only grow. I would consider it the iPhone problem for the datacenter.
|
||
There was a time when we all had ugly Blackberries that were locked down by corporate policy and could barely browse the internet or install a useful app aside from Brickbreaker. Then the CEO got an iPhone and realized how pleasant and simple mobile computing could be. Then they got an iPad and the problem was further exacerbated. Consumer tech became the driving factor behind innovation, and enterprises lagged behind. Cloud computing is the same problem. And the solution is for the IT department to adopt a cloud-native operations philosophy or risk becoming irrelevant. The on-premises IT department and datacenter need to become consumer friendly, agile, and streamlined. That is going to require serious changes in culture and technology which are well beyond the scope of this simple essay.
|
||
The modern datacenter will need a cloud-native platform to support what consumers are demanding. In the second half of the 2010s, Kubernetes started to rise to prominence as a potential platform to provide a common abstraction layer across all the public cloud providers. One of the K8s luminaries, Kelsey Hightower, made the argument to me - rather convincingly - that Kubernetes will become that common platform and API for cloud, edge, and on-prem.
|
||
What’s coming? The 2020s are going to see the advent of Edge applications, which will cause an explosion of computing power that will make our current cloud computing efforts seem quaint. The lessons learned through cloud-native adoption will serve us well as we move to a massive, distributed computing model with thousands of mini-datacenters sprinkled across the globe. The big three public cloud providers are already making their bid to be the big three edge providers, for good or ill. If public cloud computing was one of the biggest technology shifts of the 2010s, edge computing will be its equivalent in the 2020s.
|
||
`,summary:`In the most recent episode of Buffer Overflow, we talked about the biggest tech trends for the 2010s. I thought I would expand on my thoughts a little bit with this post. Check out the full episode below.
|
||
[embed]https://www.anexinet.com/wp-content/uploads/2019/12/BufferOverflow-Episode140.mp3[/embed]
|
||
Public cloud accelerates everything In the 2010s public cloud computing exploded and enabled massive technological innovation, particularly in the realm of Software as a Service (SaaS). In addition to providing a common platform for startups and enterprises alike to build on, it also served as a example to existing datacenter providers and in-house IT operations teams on how to effectively run a cloud service at scale using automation and standardization.`,date:"31 Dec, 2019",url:"https://nedinthecloud.com/2019/12/31/the-2010s-a-decade-in-review/",image:"BufferOverflow191.png",readingTime:"6"},"https://nedinthecloud.com/2019/12/30/2019-year-in-review/":{title:"2019 Year in Review",tags:[],content:`In December of 2018, I wrote a post about my goals for 2019. I’d like to reflect on those goals and take a look at what I actually accomplished in 2019. My goals for 2020 will be part of a separate post coming next week.
|
||
2019 was a heck of a year for me professionally. I started my own business, and that had an outsize impact on the goals I had made in the beginning of the year. At the beginning of 2019 I already knew that I was planning to quit my current job as Director of Cloud Solutions at a local VAR and go into business for myself. But I wasn’t ready to make that information public, since I thought I wouldn’t be ready to quit until June at the earliest - it turned out to be May. The necessity of keeping such a large goal quiet meant that I couldn’t be entirely straightforward in my goals list, lest I tip my hand. With that in mind, let’s take a look at my goals for 2019:
|
||
Try and stay positive Be kinder Mentoring someone Produce five new courses for Pluralsight Start a new podcast Blog every week Get the Microsoft MVP Award again Speak at more events Learn Go for real That’s a solid list of goals, some of which are well-defined and others which are somewhat nebulous. Let’s see how I did.
|
||
Try and stay positive While this does seem like a poorly-defined goal, I did have a list of specific things I wanted to work on.
|
||
Whenever I want to make a negative comment, find a positive one as well When a new idea is posed, look for ways to improve the idea (don’t tear it down) When something goes wrong, look for ways to make it right (don’t just complain) Start with the premise that people are usually trying to do good, and wait for compelling evidence to the contrary Be encouraging to anyone who is starting a new project, business, etc. (They’ll be plenty of detractors, be a supporter instead) My major concern was that the world of Tech Twitter can be sarcastic and snarky to a point where it’s just a constant stream of negativity. I don’t want to be one of those voices. It’s okay to be snarky and skeptical, but there are limits. Corey Quinn has talked about this on his Screaming in the Cloud podcast, and essentially it comes down to this. When you are being snarky, it’s okay to punch up. Meaning that it’s okay to snark about big companies and extremely public figures that have far more money and influence than you. It is NOT okay to make it personal (except for Oracle apparently). The snark and commentary should directed at companies, ideas, and products not at individuals. It is also NOT okay to punch down, meaning snarking at those who have less influence or are not in a privileged position. That is the rule I have tried to follow, in addition to being supportive of people trying to build cool new things. I even championed a special episode of Buffer Overflow where each of the hosts brought their favorite, positive things to the table.
|
||
2019 was a year that needed a lot of positive people to drown out the torrent of bad news and bad actors. I think I helped in that regard and I would like to continue this stance moving into 2020. I’d give myself a B+ for this goal.
|
||
Be kinder Just like staying positive, this was also a nebulous goal that needed some well-defined metrics.
|
||
Be honest about what people have done, but don’t speak ill of them Try to help someone out everyday in some small way Look for opportunities to promote other people (at least once a week) No personal attacks (criticize the action, not the person) Practice focused acts of kindness (I don’t believe in the random thing, which is a whole other post on its own) Volunteer or raise money for at least one charity (in addition to making contributions) The impetus behind this goal was the fact that I had become increasingly negative in regards to other people. I think I can attribute this to staying at a job I no longer enjoyed for too long. I was dealing with people I didn’t like on a daily basis and watching individuals do things I found reprehensible. Quitting my job and working for myself has massively improved the caliber of people that I interact with on a daily basis. All of my coworkers are selected by me. If I encounter someone I don’t want to do business with, I don’t. It’s really that simple.
|
||
I’ve also been trying to promote the efforts of others when on Twitter or LinkedIn, as well as in real life. If someone is doing something cool, I try and let the world know. Not sure if I managed to do that every week, but I have been trying. I don’t think I’ve personally attacked anyone in private or public since leaving my old job. Once I was gone, all the petty nonsense became just that - petty nonsense. None of it was important to me, and whatever malice I felt towards certain individuals evaporated. I simply don’t need to talk shit about other people. If I don’t like them, I don’t interact with them. If someone else asks me about that person, I will be factual not malicious.
|
||
I did not raise money for a charity, so that needs to be a goal for next year. I’m a big fan of St. Jude’s, and I have raised money for them before. Maybe I will do a fundraiser through Day Two Cloud, we’ve got some sweet t-shirts coming soon.
|
||
I’d give myself a solid B for this goal.
|
||
Mentoring someone Part of the process of growing and maturing as a person is to share what you have learned with others. Someone did reach out to me at the end of 2018 looking for guidance. We had a number of lunches over the first half of 2019 as he searched for a direction about what to do next. There were also several people who I spent some time with after I quit my job, who wanted to know more about what I was doing and what direction they might want to head in their career. Hopefully I helped each of those people in some small way. I’d love to do something a bit more formal this coming year.
|
||
Even though I started strong in 2019, the mentoring thing fell by the wayside as my new business took off. I didn’t make mentoring a priority. I’d have to give myself a middling C for this goal.
|
||
Produce five new courses for Pluralsight Ha! I knew I was going to publish a lot of courses for Pluralsight in 2019. I ended up publishing eight, yes eight! One of them was a revision of my Terraform - Getting Started course. I would count that since it was a full rewrite with all new video and examples. So wow, eight courses in twelve months. I expect that next year will be more of the same. Pluralsight continues to be the lion’s share of my income, and making more courses is one of the best uses of my time. I’d have to give myself an A+ for this goal.
|
||
Start a new podcast In late 2018, Ethan Banks approached me about doing a podcast for Packet Pushers that was focused on the cloud, and thus Day Two Cloud was born. The first episode was published on January 23rd, and it was publishing on a fortnightly basis until November. At that time Ethan joined me as a cohost and we switched to a weekly production schedule. You can hear more about that on episode 22.
|
||
I’m happy to say that Day Two Cloud (D2C) has been a resounding success! Thanks in no small part to the existing audience at Packet Pushers. They were able to provide a solid group of initial subscribers by adding D2C to their Full Feed, and now it’s up to me and Ethan to attract additional subscribers by providing informative and entertaining content every week. If you’ve got a topic or story to tell, please reach out!
|
||
I’m happy to give myself an A on this goal.
|
||
Blog every week Er. So I was kinda busy in 2019 building a business and writing for others, which had the net effect of me writing less for myself. I did manage to crank out 40 posts, including this post, but that falls far short of the 52 posts that would be required. That’s not to say I wasn’t busy creating content. In the last year I created eight courses for Pluralsight, wrote original content for four different vendors, wrote a book on the Azure Kubernetes Service, published a weekly podcast with Buffer Overflow, published 29 episodes of Day Two Cloud, and started writing technical documentation two days a week for a startup. I am writing a lot, just not on this site.
|
||
I’d like to say I’ll do better in 2020, but I don’t know that I will. I’ve already got six courses planned for Pluralsight, another book in the works, D2C and Buffer Overflow publishing weekly, and my writing gig for the startup. That just doesn’t leave a whole lot of time. Probably the best thing I could do is create a blog post based off the D2C or Buffer Overflow episode that publishes each week.
|
||
The impetus behind this goal was to become a better writer by writing more. While I accomplished that goal, I did not meet the metric I laid out for myself. Based on my complete failure to publish weekly, I’m afraid I must give myself an F.
|
||
Get the Microsoft MVP Award again Mission accomplished! I was re-awarded the Microsoft MVP award for Azure/Azure Stack in July of this year. This was a pass/fail type grading, and I passed!
|
||
Speak at more events Here’s the quote from my original post:
|
||
In 2019, I am going to submit more, submit better, and get rejected more as well.
|
||
And boy did I get rejected! I submitted talks for about eight conferences and had two accepted. I got to go speak at the MMS Mall of America conference, and it was a great event with ton of awesome speakers and attendees. None of the sessions are recorded at the event, so I’m afraid you can’t go watch any of them. The demos for the Terraform talk are available on my GitHub. I also submitted a talk for the HashiTalks event, which was 24 hours of talks about HashiCorp products. That talk is available on their website.
|
||
I participated in Cloud Field Day 5 and 6 as a delegate, which is not a speaking gig per se, but you do end up doing a bit of talking.
|
||
I presented at the Azure Philly group on using AKS, and then was invited to present remotely to the Austin Azure meetup for the same topic.
|
||
I also participated in a bunch of webinars, panel discussions, and podcasts. It was a pretty good year for me and my public persona.
|
||
When I decided to start my own business, I wasn’t exactly sure where the money would come from, so I had a lot of irons in the fire. One of those irons was becoming a public speaker who is paid to speak at conferences and events. Ultimately, I ended up not pursuing that possibility. Being a professional speaker at events requires a lot of travel. Those who do it are constantly going to new conferences across the globe, and traveling 50% or more of the time. That was not for me. I max out at about four events per year, in part because I have three small children that I like seeing, and in part because I am a bit of a homebody.
|
||
I have to give myself a B- for this goal, which is a grade I am completely okay with. My priorities shifted through the year and this became less of an important goal for me.
|
||
Learn Go for real Nope. Wanted to do this in 2018 and I didn’t. Wanted to do this in 2019 and I didn’t. Learning Go will probably remain one of those low priority goals that keeps getting overridden by more important things. The only way I will ever learn Go is if there is a financial driver behind it.
|
||
When you are self-employed, every project you undertake needs to be evaluated on it’s monetary implications. Learning to write better has a direct impact on how often I am paid to write, so that is a worthwhile endeavor. Learning about Kubernetes administration makes me a better host for Day Two Cloud, where so many of the conversations include something about K8s. Likewise, learning about GCP or other cloud related technologies. And being a better host of D2C means more subscribers and downloads, which means more interest from sponsors, which means more money in my pocket. But learning Go? So far I don’t see a direct path from learning Go to making more money as Ned in the Cloud LLC. Until I see that path, this will remain an item of vague interest for me.
|
||
I give myself an F for this goal, and that is fine by me.
|
||
That’s it for my goals in review. As I alluded to earlier, my goals and priorities shifted over the course of the year. There was no way to predict that D2C would take off like it did, that I would end up writing a book, or that I’d get a gig writing technical docs part-time. The thing about goals is they need to align with your overall strategy, and each time you revise your strategy you also have to revise your goals. As 2020 starts, I need to make sure I have a viable strategy for Ned in the Cloud, and make sure my personal and business goals align with that strategy. That is the topic of my next post.
|
||
`,summary:`In December of 2018, I wrote a post about my goals for 2019. I’d like to reflect on those goals and take a look at what I actually accomplished in 2019. My goals for 2020 will be part of a separate post coming next week.
|
||
2019 was a heck of a year for me professionally. I started my own business, and that had an outsize impact on the goals I had made in the beginning of the year.`,date:"30 Dec, 2019",url:"https://nedinthecloud.com/2019/12/30/2019-year-in-review/",image:"featured-image-nedinthecloud.jpg",readingTime:"12"},"https://nedinthecloud.com/2019/12/03/azure-advent-calendar-2019-aks-and-pod-identity/":{title:"Azure Advent Calendar 2019 - AKS and Pod Identity",tags:["aks","azure-kubernetes-service","azureadventcalendar"],content:`The Azure Advent Calendar is a fantastic idea from Gregor Suttie and Richard Hooper. Everyday, starting on December 1st and going until Christmas, they are posting three new Azure videos on their YouTube channel. They were originally planning to post a single video each day, but the response was so overwhelming they were able to schedule 75 total videos! I guess they shouldn’t be surprised, this is the Microsoft MVP community we are talking about. When I asked for a few guests on my fledgling Day Two Cloud podcast, the response was similarly overwhelming.
|
||
My humble entry is being published today, December 3rd, and the topic is running the Pod Identity solution on Azure Kubernetes Service.
|
||
The code for the demo can be found on my GitHub in this repository. As I mentioned in the demo, there is a bug in the azure-identity library for Python that is preventing my Flask application from working properly. I’m submitting a bug report, and if it gets fixed I’ll post an update here.
|
||
I’m excited about all the great content that will be published in the run up to Christmas. You can see the complete calendar at their website. Thanks again to Gregor and Richard for inviting me to be a part of this event.
|
||
Happy viewing and happy holidays!
|
||
-Ned
|
||
`,summary:"The Azure Advent Calendar is a fantastic idea from Gregor Suttie and Richard Hooper. Everyday, starting on December 1st and going until Christmas, they are posting three new Azure videos on their YouTube channel. They were originally planning to post a single video each day, but the response was so overwhelming they were able to schedule 75 total videos! I guess they shouldn’t be surprised, this is the Microsoft MVP community we are talking about.",date:"3 Dec, 2019",url:"https://nedinthecloud.com/2019/12/03/azure-advent-calendar-2019-aks-and-pod-identity/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2019/11/16/ned-in-the-cloud-llc-six-month-check-in/":{title:"Ned in the Cloud LLC - Six Month Check-in",tags:[],content:`Six Month Check-In: Cruisin' In May of 2018, I decided to start my own company (Ned in the Cloud LLC), quit my lucrative consulting job, and hang my own shingle for content creation, education, and technical writing. Well, that’s not entirely accurate. I actually decided to strike out on my own in September of 2018, it just took me until May of 2019 to do it. I won’t rehash the entire thing now. If you want to hear the whole story, check out my detailed post or this episode of The Full Stack Journey podcast.
|
||
So how are things going six months later? In a word. Awesome. That’s all you wanted to know? Sweet, thanks for reading.
|
||
.
|
||
..
|
||
…
|
||
OH. You’re still here? Neat! Here’s what’s been going on since my post back in August.
|
||
Day Two Cloud Podcast I started the Day Two Cloud (D2C) podcast back in beginning of 2019 in the Packet Pushers Community channel for podcasts. In May of 2019, after about eight episodes, D2C graduated to its own channel and received its own webpage. At that point, I was still publishing on a fortnightly cadence and handling the hosting duties solo. Starting in November, D2C is moving to a weekly cadence and I am joined by Ethan Banks as my cohost. The podcast now has its own website, Twitter handle, and email address. It also has its own mascot, the mischievous Nimbie.
|
||
On the business side of things, D2C is now going to have sponsored episodes and ad-spots. I suppose this makes me a professional podcaster? I know some people bristle at the idea of sponsorships and ad-spots, and I get that. I bristle at the idea of not being able to feed my family. Since content creation is now my full-time gig, I need to be able to monetize the content I am creating and that includes podcasts. That being said, I am going to make every effort to ensure that all sponsored podcasts are delivering interesting and informative content. If vendors are just there to make a hard pitch for their product, then I am not interested. We have the luxury of choosing our sponsors, and they are briefed ahead of time what we are looking for in an episode.
|
||
Don’t think I’m living up to my end of the bargain? Let me know!
|
||
Cloud Field Day 6 Cloud Field Day is a source of joy and opportunity for me. Joy because I get to spend three days with the other delegates, who are - to a person - amazing individuals. Cloud Field Day 6 was no exception. In addition to finally meeting Nate Avery, I also was introduced to awesome individuals like Larry Smith Jr., Chris Williams, and Rita Younger. Seriously, go check out the CFD6 page and follow everyone on it. You’ll be glad you did.
|
||
Many of the delegates at CFD6 are also full-time influencers and content creators. Being together gives us a chance to bounce around ideas and band together for new projects. This has been especially helpful for me, as people like Keith Townsend are full of incredibly useful and practical advice that has helped me grow my business in a healthy way.
|
||
Time Management Sucks Time management continues to be a major challenge for me in two important ways. Actually managing my time, and accepting too much work.
|
||
High Friction = Low Adoption I had such bold dreams of having a highly organized and efficient time management system that would track all of my projects and make me a more efficient revenue generating machine. As usual, the best laid plans… are exactly that. In my three month post, I mentioned using Trello and Toggl to track my projects and manage my time. Three months later, I use Trello for the Day Two Cloud podcast only and I use Toggl to track my time for one client. This is not the time management utopia I dreamed of, and it’s likely that it never will be.
|
||
The fact of the matter is that I have no appetite for complicated project management tools or activity tracking. The biggest challenge is one of friction. The more effort it takes me to engage with a system, the lower the chance I will continue to use that system unless there is some outside motivating force.
|
||
I use Trello b/c Packet Pushers uses Trello and I need to coordinate with them and my D2C cohost. I stopped using Trello for everything else b/c no one else was collaborating on those projects. I was the only one involved, and I know what I need to do.
|
||
It’s the same with Toggl. I was trying to track how long it takes me to create and produce courses for Pluralsight, but ultimately I decided it wasn’t worth the effort. I have a rough idea of how long it takes me to produce a given course, and remembering to open Toggl each time I was doing Pluralsight work was too much of a mental effort, i.e. the friction was too high. Plus, it didn’t capture the times I was working on the course, but not at my computer. When I am trying to think of a good demo or how to organize a module, I go out for a run or think about it in the shower. Toggl isn’t capturing that information, and I don’t really care.
|
||
I am using Toggl for one client that I am billing by the hour. I have to track my time to invoice properly, and thus there is an external force to overcome the friction of the tool.
|
||
Why do I do this to myself? The other aspect of time management that I am still struggling with is being an accurate judge of my bandwidth. There’s a few factors at play here.
|
||
Worry about income Saying “no” Fear of missing out When I quit my full-time job and went independent, it was hella scary. Before quitting, I was not responsible for finding my work. That job belonged to someone else. Now I am responsible for sourcing all of my work and getting it done. I have this irrational fear that suddenly I won’t be able to find more work. If something comes along, I better agree to do it, because who knows when the next opportunity will come along. As I spend more time as an independent, I am realizing that work for competent people is never in short supply, and turning down work is the only way to stay sane.
|
||
Which leads me to the problem of saying “no.” I don’t like to say no to people. I like to be helpful, and if someone is reaching out for assistance, I want to be able to jump in and get the job done. That’s not realistic, and is 100% the road to burnout. Saying “yes” is also the path of least resistance. Saying “no” requires confrontation, and I do not like confrontation. I’m happy to say that I have turned down four major opportunities in the last three months. I’m sad to say that I probably should have turned down more. What can I say? I’m a work in progress.
|
||
There is also the matter of FOMO. Technology moves at a ridiculous pace, and there are always more things that I want to dive into than there are hours in the day. For instance, I want to dive deeper into Morpheus Data, NetApp NKS, VMware Project Pacific, GoLang programming, Azure Stack updates, and using Git more effectively. And that’s just a few off the top of my head. When there is a paid opportunity to learn more about a topic I am already interested in, it’s that much harder to say “no.” I worry that I am going to miss the boat on a major technology trend, instead of being front and center for the next big thing. And I need to get over that.
|
||
I NEED to get out MORE One of my main concerns with going independent and working from home 100% of the time is that I wouldn’t be getting enough social time with other human beings. My initial idea was to go out to lunch or breakfast with someone every week. I was studious about this in the first quarter, but since then I have become much more slack about it. I can feel myself making excuses for not leaving the house. I’m too busy. There’s stuff I need to do around the house. It’s raining.
|
||
There’s a slippery slope here that I need to avoid. What I really need to do is stop accepting so much work and start having more lunch dates. It’s not as if there’s no one nearby to have lunch with. I just need to go through the effort of actually doing it.
|
||
Conclusion In summary, business at Ned in the Cloud is good. People are mostly awesome. And the future looks bright. My goals for the next six months are to accept less work, get out more, and keep growing the Day Two Cloud podcast. The next check-in post will be at the one year mark. Let’s see how I’m doing then.
|
||
`,summary:"Six Month Check-In: Cruisin' In May of 2018, I decided to start my own company (Ned in the Cloud LLC), quit my lucrative consulting job, and hang my own shingle for content creation, education, and technical writing. Well, that’s not entirely accurate. I actually decided to strike out on my own in September of 2018, it just took me until May of 2019 to do it. I won’t rehash the entire thing now.",date:"16 Nov, 2019",url:"https://nedinthecloud.com/2019/11/16/ned-in-the-cloud-llc-six-month-check-in/",image:"tutorials.png",readingTime:"8"},"https://nedinthecloud.com/2019/11/05/microsoft-ignite-2019-the-great-marketing-shift/":{title:"Microsoft Ignite 2019 - The Great Marketing Shift",tags:["azure","microsoft","microsoft-ignite"],content:`The Ennui of Success In 2018, I attended Microsoft Ignite and one of the things I noticed during the keynote was how stale everything seemed. Satya appeared to not only be lacking in excitement, but also clearly overcompensating by trying to be super excited about boring things. It was painful to watch and I thought perhaps we were witnessing the beginning of the end for Satya’s tenure as the CEO of Microsoft. His mission in many ways had been accomplished. The culture at Microsoft had been irrevocably altered, the product direction and priorities shifted to accommodate our service based economy, and their market cap was set to crest $1 trillion. If Satya was planning to leave on a high note, this would definitely qualify.
|
||
SPOILER ALERT: Satya did not leave. The stock price had not peaked, in fact since September of 2018 it has grown 31%. While the keynote may have been a weak showing, Microsoft itself had never been healthier and more profitable. And since stock price is an indicator of future potential, all signs point to a bright outlook for Redmond.
|
||
I think what happened last year was simply feature exhaustion. By which I mean, the bulk of the announcements were simply additional features on existing products or simple iterations. Server 2019 is great I’m sure, but it isn’t the quantum leap from 2003 to 2008. There were a ton of improvements in Azure services, but these were mostly cool features. There was nothing bombastic to get people excited.
|
||
Rename all the things! The solution this year was to rename everything. No better way to make something seem new than to rename it and pretend it’s a brand new service with all these cool features, rather than a newer version of an existing product. And so we have before us the great rebranding of 2019.
|
||
Azure Stack -> Azure Stack Hub Databox Edge -> Azure Stack Edge Config Mgr + Intune -> Microsoft Endpoint Management Azure Data Warehouse -> Project Synapse Microsoft Flow -> Power Automate
|
||
While Edge did not get a new name, it did get a new soul (Chromium) and a new logo. There are probably some more that I am forgetting at the moment. My point is that these are not new products in any realistic sense. They are all rebrands of existing products. But during the keynote, Satya and team made no attempt to link the old product name to the new. To an outside observer the effect was that Microsoft had just launched a ton of brand new services.
|
||
There is one exception to point out and that is Azure Arc. Details are still incredibly scarce on exactly how Azure Arc works and what the requirements and limitations are for the hybrid cloud management platform. Is it the answer to Google Anthos? Is it an attempt to compete against services like AWS RDS on VMware? And what does this do to the Azure Stack family? Time will tell, and I am excited to know more about this product. Much moreso than all the other announcements made during the keynote.
|
||
But my PowerPoint Decks! The massive renaming of products is going to cause more than a little consternation for those in the training and marketing world who have already created content and curriculum based on the current product names. Unfortunately, that is the nature of working in technology. I have a book on Azure Kubernetes Service coming out on December 28th, and I’d be pretty miffed if they had changed the name of that product so close to release.
|
||
Some of the renaming makes sense from a product rationalization perspective, but the bulk of it was done for marketing buzz and PR. After all, a rose by any other name would smell even sweeter, right?
|
||
`,summary:"The Ennui of Success In 2018, I attended Microsoft Ignite and one of the things I noticed during the keynote was how stale everything seemed. Satya appeared to not only be lacking in excitement, but also clearly overcompensating by trying to be super excited about boring things. It was painful to watch and I thought perhaps we were witnessing the beginning of the end for Satya’s tenure as the CEO of Microsoft.",date:"5 Nov, 2019",url:"https://nedinthecloud.com/2019/11/05/microsoft-ignite-2019-the-great-marketing-shift/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2019/10/16/netapp-and-the-cloud-conundrum/":{title:"NetApp and the Cloud Conundrum",tags:["cfd6","cloud-field-day","netapp"],content:`There are few technology changes that have been as disruptive as the public cloud. All aspects of the tech industry have been impacted in some way, but the hardware industry in particular has seen a major disruption. The advent of cloud hyperscalers has skewed the server, network, and storage markets. Suppliers can have an entire quarter made or ruined by the decision of a single hyperscaler to purchase new gear for a datacenter. By the same token, the cloud hyperscalers now have outsized influence over pricing, and can negotiate heavy discounts by purchasing at massive scale. Add in the fact that enterprise IT is shrinking, and traditional hardware companies are in a bit of a pickle.
|
||
Vendors like NetApp need to address the massive shifts in markets by developing fresh products and services that will appeal to enterprise customers in a post-cloud environment. Let’s look at some of the biggest shifts that have occurred due to the prevalence of public cloud.
|
||
Low friction consumption of services Consumption based spending Capacity on demand Management of hardware and base systems These shifts are all part of what public cloud offers to the consumer, and it is one reason that organizations have flocked to adopt it. While cost is often cited as a reason to move, the reality is that convenience, scale, and outsourced management are probably the most compelling features. How does a traditional hardware vendor become competitive in a landscape ruled by the new public cloud principles?
|
||
That is the cloud conundrum. NetApp has a possible solution.
|
||
NetApp Kubernetes Service When you think of NetApp, you probably think of storage. That is what they have been known for, and in fact they are a leader for primary storage in Gartner’s Magic Quadrant for 2019.
|
||
You’d be forgiven for wondering why on earth NetApp would create its own Kubernetes service if it is primarily a storage company. But let’s remember that NetApp is trying to bring the principles of convenience, scale, and outsourced management to a hybrid cloud world. NetApp has been expanding its solutions into the major cloud providers, with first party support in Azure and GCP and third party support in AWS. To make these solutions work in the cloud, NetApp had to create a deployment and provisioning process that was automated and scalable. Based on this post by Jonsi Stefansson, the acquisition of Greenqloud brought with it a Service Delivery Engine (SDE) that enabled the rollout of Cloud Volumes. It appears that the SDE leverages Kubernetes to run cloud native applications, and so NetApp ended up inadvertently building a Kubernetes management platform for their own internal use.
|
||
While I’m not convinced of the long-term viability of NKS, I am certain there is a need for Kubernetes management software on the market. Rancher Labs has seen great success providing a platform that can manage multiple Kubernetes installations across a plethora of platforms. I’m not going to compare and contrast the two offerings - I’ll leave that as an exercise to you, but suffice to say there is a significant amount of overlap. NKS is able to provision and manage K8s clusters in the major cloud providers, with support coming for on-premises installations. It’s the on-premises installations that throws me for a moment, because while the cloud providers are highly standardized and well documented, on-premises deployments are heterogeneous, wildly differentiated, and a support nightmare. How is NetApp going to offer NKS on-premises? How is NetApp going to offer the same convenience, scale, and outsourced management in the private cloud sphere?
|
||
NetApp HCI NetApp HCI is a great way to break up the heterogeneity of the on-premises datacenter, and make it possible for them to offer NKS on a standardized platform. By controlling the hardware and software layers of the solution, NetApp can offer a consistent, managed experience to end users. Imagine that you are already using NetApp’s Cloud Volumes in Azure and GCP, NKS to manage your K8s clusters, and Trident to manage your K8s storage. Adding NetApp HCI gives you that same deployment and management model on-premises. Except… it’s on-premises. Which means you have to work with your local VAR, buy the capacity you “think” you need, install the hardware in a rack, and get it configured. Then you have to manage the hardware going forward. That doesn’t sound very convenient or simplified for management. It’s a vastly different experience than using the public cloud, and it includes several points of friction that both infrastructure folks and developers find aggravating.
|
||
As Neil explains in the clip below. Enterprises are looking to move to the cloud for the lowered friction of getting started and the movement from CapEx to OpEx type spending.
|
||
NetApp’s goal with NetApp HCI and it’s Data Fabric approach is to abstract away the hardware on the private cloud side and simplify adoption by providing a managed offering. I couldn’t find a clip that talked about it, but Neil indicated that NetApp HCI will be made available as something approaching hardware-as-a-service, that would truly be an OpEx cost to the finance side. NetApp would be on the hook for managing the hardware, providing service, and adding capacity in what Neil called “packs”.
|
||
The idea of a managed, on-premises private cloud is not new. There have been several attempts from other vendors to sell a private cloud solution that matches the needs of an enterprise buyer. The cloud hyperscalers are also trying to get into the market with solutions like Azure Stack, AWS Outposts, and Google Anthos. I believe this validates NetApp’s approach of creating a product offering that enterprises can use in their datacenter. I think it also has an advantage in being hybrid solution not tied to a specific public cloud vendor, unlike Azure Stack and Outposts. In addition, NetApp already has a foothold in many enterprise datacenters, which gives them a leg up on something like Outposts, which is an AWS OEM hardware solution. There are a few considerations for execution.
|
||
NetApp needs to engage with decision makers outside of the storage silo to build interest NetApp has to tread lightly with their existing VAR partners to strike a balance between convenience and the channel NetApp has to prove they have the services and support chops for the entire solution None of these are simple tasks to execute, and I would be curious to see what NetApp’s strategy is to conquer each one. If nothing else, NetApp’s CFD6 presentation and new line of products and services has given me much food for thought.
|
||
`,summary:"There are few technology changes that have been as disruptive as the public cloud. All aspects of the tech industry have been impacted in some way, but the hardware industry in particular has seen a major disruption. The advent of cloud hyperscalers has skewed the server, network, and storage markets. Suppliers can have an entire quarter made or ruined by the decision of a single hyperscaler to purchase new gear for a datacenter.",date:"16 Oct, 2019",url:"https://nedinthecloud.com/2019/10/16/netapp-and-the-cloud-conundrum/",image:"featured-image-nedinthecloud.jpg",readingTime:"6"},"https://nedinthecloud.com/2019/09/30/solo.io-stole-the-show-at-cloud-field-day-6/":{title:"Solo.io stole the show at Cloud Field Day 6",tags:["cfd6","cloud-field-day","solo-io"],content:`I attended Cloud Field Day 6 this September and my favorite presentation hands-down was from Solo.io. There were other strong contenders, but Idit and Christian stole the show with their laser-precise product focus, infectious enthusiasm, and high-quality demonstrations. I overhead someone say that it was a masterclass in how to do a Tech Field Day presentation, which is especially impressive for a first-time presenter. I wrote a post about what I thought of Solo.io’s various projects before I went to CFD, here are my updated thoughts on their offering and the team behind it.
|
||
The YouTube videos for Solo.io can be found on their Cloud Field Day page, and I have embedded the introduction video below.
|
||
Idit Levine, the Founder and CEO is a bit nervous, as are many presenters, but she is also clearly excited about the project. Words seems inadequate for the deluge of information she is trying to communicate, and I feel like she would rather open our heads and pour knowledge directly on our brains. Sadly that doesn’t appear to be possible today. Maybe that will be her next startup.
|
||
Idit and her Field CTO, Christian Posta, are both clearly excited with the solutions they have to offer. The larger question is whether others will find it equally compelling. What are their solutions and is there a product fit? The presentation focused on three main offerings:
|
||
Gloo SuperGloo Dr. Gloo Gloo is an API gateway using the Envoy proxy. It functions as a control plane layer on top of Envoy, to provide a more user friendly experience and integration with non-Envoy endpoints. There were two things that gelled with me after I noodled on it a bit. The first is that they are not trying to build a competitor to the proxy solutions already out there. The primary target is Envoy - you’ll get the best performance from containers leveraging it, but other targets are fine including Serverless and legacy applications. The demo showed an application that was composed of a legacy web server, a microservice for displaying table data, and an AWS lambda app providing a contact page for the website. All three were stitched together using Gloo. The process was simple, secure, and could be driven via UI or with YAML configs. The project itself is open-core with certain additional functions only available in the paid version.
|
||
SuperGloo is a service mesh orchestrator. Today many service meshes are not particularly simple to manage, and if you have multiple service meshes across your organization your problems increase exponentially. You may also have different teams choosing different service mesh solutions, with one using Linkerd and another using Istio. Each one has a different API and complicated settings. SuperGloo is meant to be the control plane of service meshes, abstracting their APIs and settings to a common framework in order to simplify operations. I wasn’t entirely sure how many companies are actually deploying a single service mesh, let alone multiple service meshes necessitating a service mesh hub. However, I was assured by my new friend cxi, for those companies that are using microservices and have not yet deployed a service mesh, they will very soon.
|
||
“Service Mesh is a hot topic, but not too many people are actually doing it.” #CFD6 @soloio_inc @christianposta - So true, Service Mesh is widely not used, but widely needed!
|
||
— Christopher Kusek (@cxi) September 27, 2019 SuperGloo seems to simplify the management of a even a single service mesh, and will help organizations as they grow to multiple service meshes in their environment.
|
||
Dr. Gloo is a solution meant to assist with the health of your application in a number of ways. It’s an extensible solution, with a number of extensions already provided by Solo.io. They chose to highlight two of those, Squash and Loop. Squash provides the ability to debug distributed applications across multiple services in order to determine what is causing an outage. Distributed apps are notorious hard to debug, since there is no longer a monolithic application on one or two servers to debug against. Squash is trying to provide that monolithic debugging experience on a distributed application. In a lower environment you might be able to run the debugging directly on the impacted systems, but in production that might not be feasible. Loop allows you to collect the state of a distributed application when an error occurs, and then reproduce that error and debug the state of the application with Squash on a non-production system.
|
||
Many of the solutions shown by Solo.io lean towards the developer side of the experience, but as with many things in the cloud, that line is becoming increasingly blurred. Who is responsible for setting up and managing an API gateway? Who should be making sure that the proper toolset is in place for deploying service meshes? While many of these objects are software-defined, I think at least some of these job functions fall squarely in the Ops category or at the very least sit on the fence between Dev and Ops.
|
||
There were two overriding concepts that Solo.io kept talking about in their presentation. The first is that they are customer driven. Product roadmaps and features are based off of real customer requests from large companies, and not the fanciful ideas dreamt up in a windowless room. I don’t mean that Solo.io is just giving customers a better buggy whip, I mean they are looking at the real problems customers are having and developing a novel solution to meet the customer and their challenges where they are today. The second big point was that they are developing solutions that are complimentary to existing solutions. SuperGloo is not trying to be a service mesh - an already crowded space - they are providing a service mesh agnostic solution to manage the service mesh sprawl that is surely coming.
|
||
Whether Solo.io gets picked up by a large company or continues to develop their solutions on their own, I have no doubt that they are a name to remember in the cloud-native market segment.
|
||
`,summary:"I attended Cloud Field Day 6 this September and my favorite presentation hands-down was from Solo.io. There were other strong contenders, but Idit and Christian stole the show with their laser-precise product focus, infectious enthusiasm, and high-quality demonstrations. I overhead someone say that it was a masterclass in how to do a Tech Field Day presentation, which is especially impressive for a first-time presenter. I wrote a post about what I thought of Solo.",date:"30 Sep, 2019",url:"https://nedinthecloud.com/2019/09/30/solo.io-stole-the-show-at-cloud-field-day-6/",image:"analysis.png",readingTime:"5"},"https://nedinthecloud.com/2019/09/27/cloud-field-day-first-impressions-on-hammerspace/":{title:"Cloud Field Day - First Impressions on Hammerspace",tags:["cfd6","cloud-field-day","hammerspace"],content:`I am currently attending Cloud Field Day 6 in Silicon Valley. We are in the thick of it, just about to start the third day of presentations. Before I wander into the maelstrom of presentations and conversations that will leave my brain feeling like it’s been through a Vitamix, I thought I would jot down some initial thoughts about Hammerspace, one of the presenters from Wednesday.
|
||
In the lead up to CFD6, I wrote about Hammerspace, and I had some real questions about where their product fit within the marketplace. If you’ll forgive me for quoting myself from that post.
|
||
During their CFD presentation, I would like to focus on the use cases for the technology and why I would choose to use their solution over something like Azure Files or Amazon S3, both of which have a global namespace concept.
|
||
Did Hammerspace address these concerns during the course of their presentation. Here is the first video of their session if you’d like to see for yourself.
|
||
You can find the rest of the videos on their Tech Field Day page.
|
||
In a nutshell, I don’t think they provided ample differentiation to make me choose their solution over options that exist in the marketplace. Before I dive into the reasons why, I want to say something about the technology.
|
||
Much of the presentation was spent talking about how their solution works from a technical perspective, and I have to say that this is some pretty impressive technology. There is no doubt that David and Douglas are whip-smart dudes who have some serious technical chops. They have built a platform that sits in front of multiple NFS systems and decouples the metadata layer and the data layer from each other. The Anvil and DSX components seem to do an admirable job of providing an abstraction layer between the storage backend and the clients that consume that storage. The real questions is this:
|
||
Is there sufficient value being delivered by their solution to justify the additional layer of abstraction and administrative overhead?
|
||
We are talking about adding a layer of abstraction to existing storage, which in most of their use cases was an NFS filer. Applications that are already happily using the NFS mounts would need to be updated. New infrastructure would need to be provisioned, managed, and monitored. The storage team would need to learn a new product. And the IT Ops team would need to support it. So again, is the benefit derived from their solution worth all that additional fuss?
|
||
Let me make a comparison to another abstraction layer that adds complexity, but also provides significant benefits. Virtualization and vCenter. vCenter takes a collection of virtualization hosts and provides a common abstraction layer to simplify the consumption and management of virtualized resources. Managing vCenter and the ESXi is a non-trivial amount of work - as many a VI admin would tell you, but the benefits of virtualization and vCenter cannot be denied.
|
||
Hammerspace adds the capability to have a global namespace for NFS and they are pinning their solution to Linux and the use of pNFS. I am not convinced that a global namespace is a massive benefit. They also mentioned the ability to apply security policies, deduplication, and replication. Those are nice features to have, but many solutions you might consume in the cloud or on-premises already have these features and it appears in many cases that Hammerspace is leveraging those native capabilities.
|
||
They also had a portion of the presentation dedicated to tapping into Kubernetes. While that is an impressive feat of technical engineering, again I am not convinced that their CSI plug-in provides a net benefit over other off-tree solutions from native storage in the cloud or on-prem.
|
||
There is a lot of potential in the technology, but I am not sure that Hammerspace has found the proper market fit for what they do. I think there are a few different paths for them:
|
||
Evolve the product to be a more general storage abstraction layer focusing on policy and security management supporting multiple backend and frontend protocols. Be acquired by an existing storage vendor that is heavy in the NFS space, and layer their IP into the existing vendor’s solution to provide additional functionality. Pivot to providing a storage backend on commodity hardware with a better feature set that competes with other NAS providers Time will tell which direction Hammerspace goes, but I have no doubt that the folks at Hammerspace will be successful.
|
||
`,summary:`I am currently attending Cloud Field Day 6 in Silicon Valley. We are in the thick of it, just about to start the third day of presentations. Before I wander into the maelstrom of presentations and conversations that will leave my brain feeling like it’s been through a Vitamix, I thought I would jot down some initial thoughts about Hammerspace, one of the presenters from Wednesday.
|
||
In the lead up to CFD6, I wrote about Hammerspace, and I had some real questions about where their product fit within the marketplace.`,date:"27 Sep, 2019",url:"https://nedinthecloud.com/2019/09/27/cloud-field-day-first-impressions-on-hammerspace/",image:"analysis.png",readingTime:"4"},"https://nedinthecloud.com/2019/09/23/cloud-field-day-6-prep-hammerspace-and-solo.io/":{title:"Cloud Field Day 6 Prep - Hammerspace and Solo.io",tags:["cfd6","cloud-field-day","hammerspace","solo-io"],content:`This is part of a series of posts I’m writing as I prepare to attend Cloud Field Day 6. There are a total of eight presenters planned for CFD6, and I am going to cover two vendors per post. My goal is to have a basic understanding of each vendor’s product portfolio with a focus on cloud related products. Some of these vendors I am already familiar with, and others are new to me. In this post we are going to look Hammerspace and Solo.io.
|
||
Hammerspace Hammerspace is not a company I had heard of before CFD6. Fortunately, they held a pre-event briefing for CFD delegates to give an overview of the company. Giving the overview ahead of time will give them more time to get down into the weeds, instead of spending time giving us a bunch of company background. Time is limited, and while that stuff is interesting, it doesn’t necessarily enhance the presentation.
|
||
One thing they didn’t address was the name of the company, which I thought they just grabbed out of thin air. Turns out I was kind of right. There is this concept in cartoons called Hammerspace, and it refers to a “fan-envisioned extradimensional, instantly accessible storage area in fiction” where characters can find whatever items they need for a given situation. In early cartoons, a character might pull a hammer or mallet out of thin air to whack another character, hence the name. How does this line up with the company Hammerspace? Good question!
|
||
Hammerspace is a storage company that specializes in provide a data-as-a-service middle tier to your existing storage backends. The DaaS component creates global namespaces for your storage that are accessible from anywhere. The naming therefore is quite apropos. In many ways I would compare what Hammerspace is doing to Lucidlink, another company that is presenting at CFD6. Both of these companies are trying to find a way to make the consumption of storage simpler and ubiquitous across multiple datacenters and clouds. Their approach differs in several ways.
|
||
The Hammerspace solution is composed of a metadata service and a data connection service, respectively called Anvil and DSX. Neither of these components actually hosts the storage, instead they function as an abstraction layer for various pools of storage you may have. That includes existing storage in your datacenter and cloud based services. They didn’t get very specific on what your backend needs to look like to hook into their solution, so I don’t know if they are looking for raw pools of disk, NFS shares, SMB shares, or something else. Regardless, the idea is that you don’t have to buy a whole bunch of new storage, you can migrate your existing storage over to this service.
|
||
Anvil, the metadata service and global namespace provider, keeps track of all the metadata and can be deployed in multiple datacenters with bi-directional replication. When it comes to storage metadata is king, providing front-systems with a list of files, status, and location. This allows Hammerspace to point an application at the closest copy of a file or object, and decouples the replication requirements of data from metadata. It’s an approach that is becoming more common as storage systems are re-imagined.
|
||
Even though storage systems are evolving to glom onto new concepts and structures, most applications are still reliant on traditional file system protocols like SMB and NFS. The DSX component of the Hammerspace product provides those services to Windows and Linux servers alike, although operating systems that support pNFS can connect directly to the backend storage. There is also a Kubernetes CSI for applications running in a K8s cluster. Object based storage services are on the roadmap, but will not be available until later this year.
|
||
Cloud Field Day Presentation The way the product presentation is written, it appears that they are primarily going after storage presentation to the application layer. Unlike Lucidlink, which has an agent that can run on workstations and mobile devices, Hammerspace still uses a traditional network file system approach for delivering the data. I’d say the main strength here is their universal global namespace and integration with existing storage. During their CFD presentation, I would like to focus on the use cases for the technology and why I would choose to use their solution over something like Azure Files or Amazon S3, both of which have a global namespace concept.
|
||
Solo.io I had seen some signage about Solo.io prior to CFD6 and I knew they had something to do with Kubernetes, but I wasn’t sure exactly what. Their website makes it pretty easy to understand where they sit in the K8s stack, they are all about the network including being an API gateway and a service mesh manager. They have a bunch of open-source projects that provide different types of functionality, and two paid offerings called Gloo and Service Mesh Hub.
|
||
Gloo Gloo is an open-source project from Solo.io that provides an Envoy based API gateway for Kubernetes. There’s the free version that contains a lot of bells and whistles, but the paid-for, enterprise version adds some more goodness on top. Most of those additional features are around security functions like authentication, WAF, and policy. However, the free version supports Let’s Encrypt and HashiCorp Vault, which is pretty cool. I also noticed that the project supports Kubernetes, Consul, and serverless functions, which really covers the gamut when it comes to networking for microservice architectures and cloud-native applications. It’s really cool that the solution is not Kubernetes only, since we are not yet in a K8s-only world and probably won’t be any time in the near future.
|
||
I would like to take the tech for a test drive against a few different managed K8s solutions - AKS, EKS, GKE - and see how it stacks up to other API gateway solutions. I’d also really like to see some authentication mechanisms available on the free tier, or an additional tier that is not enterprise level pricing. HashiCorp just did something similar for Terraform, and I think it is going to generate a lot of sales for HashiCorp from smaller teams that couldn’t afford the price tag on Terraform Enterprise.
|
||
Service Mesh Hub Service Mesh Hub appears to be an aggregation point for multiple service meshes that can be managed through a central console. I’m not sure how many organizations have reached the required scale to need such a management solution, though I am sure they are out there. But for the majority off SMBs the need for even a single service mesh solution, let alone multiple solutions that need a management overlay is incredibly small. I expect that might change. The degree of the change will depend on whether multiple service mesh solutions are necessary. If I may be permitted an analogy that might resonate with IT Ops folks.
|
||
A cluster of ESXi servers without vCenter would be akin to a cluster of linked services without a service mesh. With a small number of ESXi servers, you don’t really NEED vCenter. It’s nice to have, but it also introduces additional administrative overhead and licensing costs. Adding in vCenter server is a bit like adding a service mesh. Now you can manage all your ESXi servers through a centralized console and perform orchestrated operations that weren’t possible before. A service mesh hub would be a bit like adding vCloud Director to the mix. If you have to manage a fleet of vCenter servers, and provide multitenancy, and add additional visibility, then you might need vCloud Director. But you need a LOT of vCenters and a lot of clusters to require this level of management, and it also comes with a hefty price tag and a bunch of administrative overhead. Likewise, I have to imagine a Service Mesh Hub solution would also require a proliferation of services and service meshes that most people don’t have to necessitate standing up the solution and administering it.
|
||
The SMH complies with the Service Mesh Interface specification, so I imagine it can work with any service mesh technology that is also SMI-compliant. For those that are not SMI-compliant today, I suspect that their days are numbered. As more service mesh solutions throw their weight behind SMI it is going to become a requirement and not a nice-to-have.
|
||
Cloud Field Day Presentation At their CFD6 presentation I would like Solo.io to concentrate on the value-proposition of their enterprise Gloo solution and whether they have plans to introduce a lower, paid-for tier of the product for small teams. I’d also like to know how they integrate with other solutions out in the market like Sysdig. And I would like to know if they have plans to open-source the Service Mesh Hub. It didn’t seem like that was available on the website.
|
||
Cloud Field Day 6 is happening September 25-27. There will be live-streamed presentations from both of these vendors. If you’d like to join in with the conversation just use the #CFD6 hashtag on Twitter. All of the delegates will be watching that tag and asking questions on your behalf!
|
||
`,summary:"This is part of a series of posts I’m writing as I prepare to attend Cloud Field Day 6. There are a total of eight presenters planned for CFD6, and I am going to cover two vendors per post. My goal is to have a basic understanding of each vendor’s product portfolio with a focus on cloud related products. Some of these vendors I am already familiar with, and others are new to me.",date:"23 Sep, 2019",url:"https://nedinthecloud.com/2019/09/23/cloud-field-day-6-prep-hammerspace-and-solo.io/",image:"analysis.png",readingTime:"8"},"https://nedinthecloud.com/2019/09/18/cloud-field-day-6-prep-lucidlink-and-extrahop/":{title:"Cloud Field Day 6 Prep - Lucidlink and ExtraHop",tags:["cfd6","cloud-field-day","extrahop","lucidlink"],content:`This is part of a series of posts I’m writing as I prepare to attend Cloud Field Day 6. There are a total of eight presenters planned for CFD6, and I am going to cover two vendors per post. My goal is to have a basic understanding of each vendor’s product portfolio with a focus on cloud related products. Some of these vendors I am already familiar with, and others are new to me. In this post we are going to turn to the networking side of things with a closer look at Lucidlink and ExtraHop.
|
||
Lucidlink I have never heard the name Lucidlink prior to seeing their name on the CFD6 page. That’s not a dig. There’s a TON of companies out there, and I tend to know the ones I worked with through consulting, or have giant signs at conferences, or have presented at a TFD event. It is thus with fresh eyes that I come to the world of Lucidlink, a company that will continually flummox my capitalization instincts by choosing not to capitalize the Link half of their name. This is entirely my fault and not theirs for choosing a sane way to write their name.
|
||
You would think that Lucidlink was a networking company, in fact that’s exactly what I thought when I wrote the introduction to this post. Then I went and read the documentation on their site and realized that they are a storage company. They are a storage company that relies heavily on networks and the cloud, so I suppose it makes sense.
|
||
Lucidlink has a single product, a distributed file system using object-based storage like Amazon S3. On the client side, there is a Lucidlink client that runs and mounts to the local file system. Their file system is log centric, like zfs, and separates the data and metadata planes from each other. For the data plane, Lucidlink is using object-based storage, but it is not writing each file to an object in an S3 compliant bucket. Instead, they have created a data overlay that breaks files into uniform chunks, and each of those chunks are stored as an object. This allows Lucidlink’s file system to treat the objects more like blocks on a traditional drive.
|
||
Since each file is scattered across multiple objects, Lucidlink is heavily reliant on its metadata system to describe the status and location of each chunk that makes up a file. The client software keeps a full copy of the metadata for the file system in an eventually consistent model. Data is retrieved directly from the each bucket location to the client, removing the need for a centralized file server. The client also maintains a local cache of files and can be configured to a maximum size. The presumption here is that each client will have constant connectivity to the metadata service and the bucket locations storing its data.
|
||
The solution supports all the usual suspects like encryption, compression, and caching. Being based on object storage for the data backend, the solution should be able to scale to whatever the limitation is for their metadata service.
|
||
There are many questions I have for Lucidlink. Let’s start with the other competitors in the field. What about existing solutions like Dropbox, Box, and OneDrive? If the primary play here is for end users, then I am not sure there is sufficient differentiation. All of those solutions have a mature client, syncing capabilities, and the ability to selectively determine client-side caching; not to mention their collaboration and sharing capabilities. If their target is more for servers and applications, then why would I want to use this over native solutions in a given cloud? Especially for container based workloads that will need to warm up the storage cache each time a new container spawns.
|
||
Cloud Field Day Presentation I assume that the CFD6 presentation is primarily going to focus on explaining their solution to us. It is an elegant solution, but I really want to know where they feel the business fit is for their product. Who is their target market? How are existing clients using the solution today? I’m sure we could easily get buried in the technical weeds, but I’d like to take a more pragmatic view of the solution.
|
||
ExtraHop Here’s the actual networking company! And it is one that I already know a bit about. About four years ago an ExtraHop rep came to the consulting company I was working for. They were interested in partnering with us, and I got to test drive their software in our lab. While I really liked the software, I wasn’t in a position to directly influence partnerships. On the bright side, their analysis product helped me troubleshoot some thorny issues we were having around DNS and Active Directory. Thanks for that ExtraHop!
|
||
That was for their Application Performance product, but I am noticing on their website that they have a product category specifically for cloud. What’s going on over there? Let’s have a little looksie.
|
||
The product is called Reveal(x) Cloud, and can we just stop here for a moment? I took a fair amount of math(s) in high school and college. Is ExtraHop trying to make a reference to a function, as in f(x) = y? Or are they trying to reference programming languages where you called a function called Reveal and pass it an argument (x)? I honestly don’t know, but I immediately dislike the product. You’ve confused me and made me think I might be on the outside of a clever joke. Don’t do that. Your product reveals what’s going on in the network of a cloud? Call it Cloud Reveal and be done with it.
|
||
[ANYWAY]
|
||
Cloud Reveal - as it will now be known on my site - is basically a SaaS version of their existing ExtraHop analysis engine that ingests packets from your public cloud provider of choice. Provided of course that your public cloud provider is AWS or Azure, which based off current stats is a fairly strong guess. AWS has VPC Traffic Mirroring and Azure has Network vTap, both of which are able to provide mirrored packet flows to the specified destination - Cloud Reveal - for collection and analysis. Cloud Reveal uses machine learning and standard detection methods to determine if there is something in the packets that warrants your attention. ExtraHop does this sort of thing really well, and it makes sense that they would plug into the public clouds now that the capability is there to do extensive, agentless packet capture.
|
||
Does this really check the box as a cloud service? It’s SaaS to be sure, and I like that it is agentless. Of course there is an associated cost with running something like Azure vTap for every virtual machine in your environment. Current preview pricing for vTap is about $9 a month per VM, with the pricing set to double once the product goes GA. AWS VPC Traffic Mirroring is about $11 a moth per VM. Layer on the cost of the ExtraHop solution, and things could get pricey for a large organization with 100s or 1000s of virtual machines running in the cloud.
|
||
I’m also curious how this handles Kubernetes managed services like AKS and EKS. Both run inside a Vnet and VPC respectively, but intra-node traffic wouldn’t leave the virtual machine, so it would not be captured. And what about non-VM services like AWS Lambda, Azure App Service, AWS Network Load Balancer, Azure Application Gateway, and others? Are all of those included in the ExtraHop offering? Cloud is more than just VMs, and the network goes well beyond simple IaaS components. Tell me how you are protecting and monitoring all of those assets and I’m all ears. Otherwise you’re just updating your product - a very good product mind you - to work in someone else’s datacenter.
|
||
Cloud Field Day Presentation I’m hoping that ExtraHop talks more about their Cloud Reveal product and how it goes beyond typical VM monitoring. Show me that this is a solution designed for cloud-native workloads, not just workloads that have been moved to the cloud.
|
||
Cloud Field Day 6 is happening September 25-27. There will be live-streamed presentations from both of these vendors. If you’d like to join in with the conversation just use the #CFD6 hashtag on Twitter. All of the delegates will be watching that tag and asking questions on your behalf!
|
||
`,summary:"This is part of a series of posts I’m writing as I prepare to attend Cloud Field Day 6. There are a total of eight presenters planned for CFD6, and I am going to cover two vendors per post. My goal is to have a basic understanding of each vendor’s product portfolio with a focus on cloud related products. Some of these vendors I am already familiar with, and others are new to me.",date:"18 Sep, 2019",url:"https://nedinthecloud.com/2019/09/18/cloud-field-day-6-prep-lucidlink-and-extrahop/",image:"analysis.png",readingTime:"7"},"https://nedinthecloud.com/2019/09/14/cloud-field-day-6-prep-netapp-and-dell-technologies/":{title:"Cloud Field Day 6 Prep - NetApp and Dell Technologies",tags:["cfd6","cloud-field-day","dell-emc","netapp"],content:`This is part of a series of posts I’m writing as I prepare to attend Cloud Field Day 6. There are a total of eight presenters planned for CFD6, and I am going to cover two vendors per post. My goal is to have a basic understanding of each vendor’s product portfolio with a focus on cloud related products. Some of these vendors I am already familiar with, and others are new to me. In this post I am going to take a look at NetApp and Dell Technologies.
|
||
NetApp NetApp was kind enough to present at Cloud Field Day 3, where they had Eiki Hranfnsson present about where NetApp was going. Before joining NetApp, Eiki was CEO of Greenqloud, a public cloud service acquired by NetApp in 2017. I have detailed my thoughts on the CFD3 presentation in a previous post. And I also wrote about NetApp in general for Gestalt IT, so feel free to read those posts if you are curious. To summarize, I went in thinking of NetApp as that company that makes NFS filers that are sometimes used for VMware. I left the presentation thinking, NetApp is trying to build something completely different than their core product.
|
||
Why would a company look to pivot so hard into the cloud, DevOps, and HCI? I think that NetApp saw the writing on the wall when it comes to being a traditional storage vendor. HCI and software defined storage are two technologies that threaten the core ideas of dedicated SAN arrays and filers. Why pay the high sticker price for branded hardware, when you have companies like Hedvig, and technologies like vSAN and Storage Spaces Direct, that allow you to leverage commodity hardware at a fraction of the price for the same or better performance. Add in the threat of people moving all their storage needs to the cloud, and it becomes obvious that a storage company needs to find a way to stay relevant. The market seemed to have bought into NetApp’s messaging in 2018, where their stock hit an all time high of $87.65. While the vision is there, investors might be feeling a little less confident now, with the stock price dipping fiercely after the last few earnings reports. They are currently trading at $56.21. Some of that mirrors overall fluctuation in the market, but the signal is clear, investors are not as bullish as they once were about NetApp.
|
||
NetApp has made some interesting announcements recently that I believe are a step in the right direction. They have launched a Kubernetes management service, as have many other companies in the last year, that is designed to work on all the major public clouds with on-premises “coming soon.” Why would NetApp create a Kubernetes management service? It seems like a weird thing for a storage, I’m sorry, Data Management company to do. The answer probably lies behind two prevailing forces.
|
||
Engaging with developers is the best way to become sticky in the enterprise. Kubernetes is hard to deploy and manage, especially on-prem. In regard to point one, you may have noticed that more and more companies are appealing directly to developers. That is because in many ways developers are helping to drive purchasing decisions at companies. With the realization that software development is a key indicator of company performance, many organizations are ready to heed the needs of their senior developers. It’s not just large enterprises though. Five person startups of today can become the tech giants of tomorrow. And of those five people, four are probably developers. Engaging with developers has been an extremely successful strategy for AWS, Azure, GCP, and others. Why do you think there are so many Developer Relationship Advocates? And I want to be clear, there is nothing wrong with this approach. As a non-developer, I do feel a bit left out in the cold, but that is my cross to bear.
|
||
Point two brings me to the next thing, the announcement of NetApp’s disaggregated HCI offering. The NetApp HCI is a hardware product offering that enables the independent scaling of compute and storage in an HCI cluster, while still maintaining the same management plane. You know what you might want to put on NetApp HIC? The forthcoming NetApp Kubernetes Service for on-premises. NetApp is certainly not the only company to throw their hat in the ring for managed K8s. VMware bought Heptio for exactly that reason.
|
||
Cloud Field Day Presentation The promise of managed K8s is a promise of hybrid cloud, where you can deploy and run your application anywhere, and Ops can manage all of the K8s clusters from a single UI. That is the vision NetApp is trying to sell. Can they execute? We’ll see what they have to say at Cloud Field Day 6.
|
||
Dell Technologies Dude, it’s a Dell! My first Windows laptop was a Dell. It ran Windows NT 3.5 and weighed about 10 pounds. That was back in 1998. I don’t think that is the Dell we are talking about here. I suspect that the Dell coming to CFD6 will be more from the Dell EMC side of the house. Although, and don’t take this the wrong way EMC, please leave your salespeople outside of the room. I’ve dealt with several sales reps from EMC over the years, and what I can say is the vast majority of them are aggressive, over-confident blowhards that love to trash talk on other vendor’s tech and make bold claims about their own tech without a scrap of whitepaper to back it up. You might think that salesperson archetype had waned post-merger, but I attended a VMUG lately with an EMC presentation, and that archetype is alive and well! I can tell you right now that the Tech Field Day crew is not going to put up with those shenanigans. It happened with a presenter at CFD3 - not going to name names - and the delegates picked him apart until he was visibly shaking with rage. Don’t bring that guy (or gal) to a CFD presentation, just don’t.
|
||
[ANYWAY]
|
||
Dell Technologies has so many products, I won’t even try to count them. It’s a pretty safe assumption that whatever Dell presents at Cloud Field Day will be at least tangentially related to cloud, but that’s about all I can say with any certainty. Dell has a majority stake in VMware, so they could talk about that. But I think they would rather have VMware present directly to the CFD crew, as they did in CFD5.
|
||
Dell makes a lot of hardware that could be used to build a private cloud, and that’s another direction they could steer the presentation. A few years ago, HPE took up the banner of composable infrastructure with their Synergy line of hardware. It’s a jazzed up blade chassis, although HPE was extremely reluctant to call it that. At least, they were reluctant until they realized that no one in enterprise IT had ANY idea what the heck composable was, and selling them on the concept was going to be a significant investment in marketing and sales “education.” Now they still talk about composable, and that nomenclature has become more prevalent, but they have also started pitching Synergy as the successor to the C7000 as well. Dell for their part, has introduced the PowerEdge MX blade chassis that also claims the same level of composability.
|
||
For me, the hardware war is kind of interesting from a purely technical standpoint, but from a business value standpoint it is… pointless. The real value of any of these platforms is either improving software deployment velocity or lowering administrative overhead. And those features are delivered through the management plane. An amazing hardware package with a terrible interface and management platform is a net negative for the Ops team. Maybe Dell is going to show all the private cloud options they have that are simple to manage and deploy.
|
||
Speaking of management, another major challenge for many organizations is the management of multiple clouds. There has a been a bit of a race by MSPs and vendors to create an all encompassing management platform that provides a marketplace, cost management, and security controls all from one console. Is that a market that Dell wants to get into? To a certain degree, the VMware products like Tanzu Mission Control and Cloud Foundation are meant to address these problems. But again that is a VMware product, and Dell is the one presenting.
|
||
There’s also a great team at Dell around the Azure Stack solution. I happen to know some of the people on the team, and they are a bunch of really sharp people that know their tech. Perhaps Azure Stack will be the big push.
|
||
Cloud Field Day Presentation Honestly, there are so many products at Dell that could fill a niche in the world of cloud, I have no idea where they are going to go with their presentation. But I am excited to find out!
|
||
Cloud Field Day 6 is happening September 25-27. There will be live-streamed presentations from both of these vendors. If you’d like to join in with the conversation just use the #CFD6 hashtag on Twitter. All of the delegates will be watching that tag and asking questions on your behalf!
|
||
`,summary:"This is part of a series of posts I’m writing as I prepare to attend Cloud Field Day 6. There are a total of eight presenters planned for CFD6, and I am going to cover two vendors per post. My goal is to have a basic understanding of each vendor’s product portfolio with a focus on cloud related products. Some of these vendors I am already familiar with, and others are new to me.",date:"14 Sep, 2019",url:"https://nedinthecloud.com/2019/09/14/cloud-field-day-6-prep-netapp-and-dell-technologies/",image:"analysis.png",readingTime:"8"},"https://nedinthecloud.com/2019/09/07/cloud-field-day-6-prep-morpheus-data-and-hashicorp/":{title:"Cloud Field Day 6 Prep - Morpheus Data and HashiCorp",tags:["cfd6","cloud-field-day","hashicorp","morpheus-data"],content:`This is part of a series of posts I’m writing as I prepare to attend Cloud Field Day 6. There are a total of eight presenters planned for CFD6, and I am going to cover two vendors per post. My goal is to have a basic understanding of each vendor’s product portfolio with a focus on cloud related products. Some of these vendors I am already familiar with, and others are new to me. In this post I am going to take a look at Morpheus Data and HashiCorp.
|
||
Morpheus Data Morpheus Data presented at my first Cloud Field Day, CFD3. That was actually my first Tech Field Day event, and Morpheus was the first presenter on the first day. I think I spent half the presentation trying to figure out the logistics of being a CFD delegate, and missed half of what they said.
|
||
If I remember correctly, Morpheus Data was trying to become an orchestrator of automators. The concept is that you already have a bunch of automation and cloud platforms, and you are unlikely to standardize on a single one. Morpheus is proposing that you use their platform to act as an orchestrator of all of these other platforms. At the time I felt that was a tall order, given the number of automation tools out there and the need for abstraction.
|
||
What I mean by abstraction is that an application that tries to provide an overlay for a bunch of other platforms has to make a fundamental decision. Do you abstract away the differences in the platforms to provide a consistent set of constructs, at the risk of losing the specialization of each platform? Or do you embrace the specialization of each platform, but attempt to build a unified interface for administrators?
|
||
I think the second approach makes the most sense, it means that you aren’t reducing the platforms down to the lowest common feature set, but it also requires more of your development team to integrate with each platform and keep those integrations fresh. Terraform has adopted this model for infrastructure automation, but they have enlisted the help of the vendors to write their own providers for the platform. They are able to do this because they have achieved critical mass with practitioners, so vendors are willing to devote development cycles to assist.
|
||
Cloud Management Platform Based on my reading of the website, Morpheus has chosen to pivot a bit and talk about being a Cloud Management Platform. They are trying to appeal to DevOps (read Developers), IT Ops, and Business Analysts. They are focusing on on their 100% agnostic platform as a way to appeal to all persons within the stack. By being agnostic, they are claiming you can managed 20+ cloud providers and avoid lock-in. Aside from the lock-in with Morpheus, since if you lean hard into their product you are basically locked into it.
|
||
I am not a fan of the lock-in argument. Any decision that is made at the technology level implies some amount of lock-in. Wrote your application in Java? You’re “locked-in” to Java. Deployed your application with AWS Lambda? You’re “locked-in” to Lambda. Lock-in is really just a measure of how difficult it would be to move some portion of the application to a different platform. Locking into a technology allows you to exploit its unique features to the fullest - remember the abstraction argument I made before, the trade-off being that it makes it more difficult to replace that technology with a different one.
|
||
This is more of a philosophical debate than a comment on Morpheus’ approach, but I think its a valuable perspective as I go into their presentation at CFD6.
|
||
Cloud Field Day Presentation Which approach has Morpheus taken? And have vendors begun to devote resources to assist with development? Those are some of the questions I have in my brain. When I saw them present in spring of 2018 they were still evolving their approach, and I’m curious to see if they have pivoted the product in any way based on market feedback.
|
||
HashiCorp If there is one vendor on the list that I am intimately familiar with, it’s HashiCorp. I have been using HashiCorp products since 2013, starting with Vargrant and then moving into Packer, Terraform, and Vault. In fact, I have four courses on Pluralsight around HashiCorp products, two on Terraform and two on Vault. When it comes to core products, HashiCorp has four that they focus on for their product messaging: Terraform, Vault, Consul, and Nomad. These four products correspond to their vision on of building, secure, connect, and run.
|
||
Build Terraform is acloud agnostic infrastructure automation platform. It is meant to build and maintain the underlying infrastructure that supports an application. When HashiCorp is talking about build they are talking about building up infrastructure to support applications.
|
||
Secure Vault is a secrets management platform. When you think about all the sensitive information that is needed to build infrastructure and deploy an application - things like access keys, API keys, database passwords, certificates, etc - that information needs to be stored securely somewhere. Vault provides a place to not only store those secrets, but also manage their lifecycle from creation to retirement.
|
||
Connect Consul is a distributed key value store and service discovery platform. The key value store can be used as a storage backend for Vault, or leveraged by other products. It can also act as a service discovery and service mesh type of platform, providing secure by default communications between applications using mutual TLS. As you can tell by the connect title, HashiCorp is leaning heavily into the network and service discovery portions of Consul.
|
||
Run Nomad is a platform used to deploy and manage the lifecycle of an application. The format of the application can take many forms, but ultimately needs to be describe in a Nomad deployment file. Nomad is responsible for taking the described application, deploying it to a target environment, and then maintaining the application through updates and alterations.
|
||
Cloud Field Day Presentation HashiCorp has given no hint what their presentation is going to be about at CFD6. If I had to guess, I think they are going to focus on their Nomad and Consul. Terraform and Vault and hugely popular and well known in the industry. But Consul and especially Nomad are less well known. In many respects, Consul and Nomad could be considered competitors to the Kubernetes ecosystem. HashiCorp has been trying to reposition both applications as complimentary or supplementary to Kubernetes. I think they are probably right, and Nomad in particular addresses applications that do not run entirely in Kubernetes. Nomad supports the Cloud Native Application Bundle (CNAB) format, which is meant for applications that are cloud-native, but not fully K8s.
|
||
Cloud Field Day 6 is happening September 25-27. There will be live-streamed presentations from both of these vendors. If you’d like to join in with the conversation just use the #CFD6 hashtag on Twitter. All of the delegates will be watching that tag and asking questions on your behalf!
|
||
`,summary:"This is part of a series of posts I’m writing as I prepare to attend Cloud Field Day 6. There are a total of eight presenters planned for CFD6, and I am going to cover two vendors per post. My goal is to have a basic understanding of each vendor’s product portfolio with a focus on cloud related products. Some of these vendors I am already familiar with, and others are new to me.",date:"7 Sep, 2019",url:"https://nedinthecloud.com/2019/09/07/cloud-field-day-6-prep-morpheus-data-and-hashicorp/",image:"analysis.png",readingTime:"6"},"https://nedinthecloud.com/2019/09/02/using-ultra-ssd-storage-with-azure-kubernetes-service/":{title:"Using Ultra SSD Storage with Azure Kubernetes Service",tags:["aks","azure","azure-kubernetes-service"],content:`In a previous post, I performed a storage performance benchmark of Azure Managed Disks and Azure Files for Azure Kubernetes Service. The testing included the now generally available Ultra SSD class of Managed Disk. The process for using Ultra SSD with AKS was fraught with peril, caveats, and an assist from the AKS product group to get it all working. I thought I would detail how I went about enabling Ultra SSDs with AKS in case someone else was struggling with the same.
|
||
Ultra SSD The Ultra SSD class of storage from Microsoft Azure is their highest performance tier of managed disks. This is the first time Microsoft has added provisioned IOPS and throughput to a storage offering, something that AWS has had for quite some time. The disks can serve up to 160,000 IOPS and 2000 MBps throughput to an Azure VM. The solution was in preview until August 15th of this year, when it was released to general availability. Even though it is GA, the Ultra SSD class is not available in all regions and the feature is not enable by default on all subscriptions. It’s also worth noting that there is no Azure VM that is capable of pushing 160k IOPS or 2000 MBps. The highest published numbers are 80k IOPS and 1200 MBps.
|
||
Requirements There are several requirements and prerequisites that need to be met before using the Ultra SSD. At a high level they are:
|
||
Enabled Ultra SSD for subscription Pick a supported region Deploy a cluster using availability zones Enable the Ultra SSD feature on the VMSS In order to get your subscription enabled for Ultra SSD, you will need to fill out a form and wait. I wish there was a super sexy Azure CLI command that would register the feature, you know like az feature register --namespace Microsoft.Compute --name UltraSSD. But that will not work. Ultra SSDs are available by request only even though the feature is no longer in preview. Go ahead and fill out the form and wait for your subscription to be enabled. I’ll wait…
|
||
All good? Great. Let’s proceed.
|
||
Now that you have Ultra SSDs enabled, it’s time to pick a region. The supported regions as of this post are East US 2, North Europe, and Southeast Asia. You can always check the most recent list on the Ultra Disk section of the FAQs. For my testing I chose to go with East US 2.
|
||
Preparing the cluster There are a few other things to know about Ultra SSDs. The disks are provisioned in an availability zone (AZ) and will only attach to a VM in the same AZ. It makes sense that only regions that have AZs could support Ultra SSDs.
|
||
The reason to force availability zone matching between disk and VM also makes a lot of sense. The storage backend supporting the Ultra SSD disk feature needs to be in the same data center as the VM attached to the disk in order to meet the target IOPS and bandwidth. Because of this requirement, an AKS cluster that want to use Ultra SSD disks will need to be provisioned with AZ support, which also requires the use of Virtual Machine Scale Sets. These are preview features that must be enabled within your subscription:
|
||
AvailabilityZonePreview AKSAzureStandardLoadBalancer VMSSPreview You are also going to need to install the aks-preview extension for your Azure CLI instance. I recommend doing all of this in Cloud Shell to simplify matters.
|
||
First let’s install the aks-preview extension.
|
||
az extension add --name aks-preview az extension update --name aks-preview Great, now register the preview features.
|
||
az feature register --name AvailabilityZonePreview --namespace Microsoft.ContainerService az feature register --name AKSAzureStandardLoadBalancer --namespace Microsoft.ContainerService az feature register --name VMSSPreview --namespace Microsoft.ContainerService The features may take up to ten minutes to register. You can check on the features by running the following command.
|
||
az feature list -o table --query "[?contains(name, 'Microsoft.ContainerService/')].{Name:name,State:properties.state}" Then refresh the registration by running.
|
||
az provider register --namespace Microsoft.ContainerService You are now ready to deploy your AKS cluster. Run the following command, substituting the proper values for the placeholders in ALL_CAPS.
|
||
#Change these rg="RESOURCE_GROUP_NAME" loc="LOCATION" clus="CLUSTER_NAME" #Leave this one alone mcrg="MC_\${rg}_\${clus}_\${loc}" az group create --name $rg --location $loc az aks create --resource-group $rg --name $clus --generate-ssh-keys --enable-vmss --load-balancer-sku standard --node-count 1 --node-zones 1 --node-vm-size Standard_D64s_v3 az aks get-credentials --resource-group $rg --name $clus Only the VM families DS and ES V3 support Ultra SSDs. The node pool used with the Ultra SSD must be of that family. Also, to maximize potential performance, I went with the D64s size that has the highest disk throughput available on an Azure VM - 80,000 IOPS and 1,200 MBps throughput for uncached disks. Ultra SSDs only support a setting of None for caching, so the uncached performance is what we’re looking at here. The Ultra SSDs are also deployed in a specific availability zone, so the node pool only needs to have one node in a single AZ. When we create the Ultra SSD, we will create it in the same zone as the single node in the node pool.
|
||
Updating the VMSS The node pool is created using a VMSS, but it isn’t ready to support Ultra SSD disks just yet. There is an ultraSSDEnabled property setting that needs to be configured on all VMs and VMSS.
|
||
{ "additionalCapabilities": { "ultraSSDEnabled": true } } The Azure CLI provides a switch for adding this property when creating a VM or VMSS directly. Since the node pool creation process abstracts the underlying VMSS creation, there is no opportunity to set this property. The property must be added after creation by deallocating the VMSS, updating the setting, and starting the VMSS back up.
|
||
There is a resource group which is created when the AKS cluster is generated with the naming standard MC_resourcegroup_clustername_location. The VMSS is in that resource group and is named something like aks-nodepool1-#######-vmss, where the ####### is some set of integers. Since there is only a single VMSS in the resource group - assuming you only have one node pool - then we can simply show all VMSSs and query the name.
|
||
vmss=$(az vmss list --resource-group $mcrg --query [].name -o tsv) az vmss deallocate -g $mcrg -n $vmss az vmss update -g $mcrg -n $vmss --set additionalCapabilities.ultraSSDEnabled=true az vmss start -g $mcrg -n $vmss That az vmss update command is not well documented, or at least I had trouble understanding exactly how the command wanted me to structure the set parameter, or if I should use the add parameter instead. Big thanks to the AKS product team for helping out there!
|
||
Storage Classes My original plan was to create an Ultra SSD storage class and use that to provision volumes for the pods. While I was able to create a storage class, there is no setting in the Azure disk provider to specify the AZ targeted for creation. It appears that the provider will simply create an Ultra SSD disk in a random zone that is included in the cluster configuration. That will work if the node pool supporting the Ultra SSDs covers all of the AZs set in the cluster configuration. If that is not the case, you either have to roll the dice or manually create the disk and attach it.
|
||
All Azure managed disks are subject to this limitation with availability zones, and Microsoft makes note of that in the AKS documentation. Other managed disk classes do not have to be deployed in an availability zone, and thus the cluster doesn’t have to use AZs. Due to the unique nature of Ultra SSDs, the cluster must use AZs and thus the problem comes to the forefront.
|
||
Here is the storage class I created for testing.
|
||
kind: StorageClass apiVersion: storage.k8s.io/v1 metadata: name: managed-ultra provisioner: kubernetes.io/azure-disk parameters: storageaccounttype: UltraSSD_LRS kind: Managed cachingMode: None DiskIOPSReadWrite: "160000" DiskMBpsReadWrite: "2000" You can save that to a file and run kubectl apply -f azure-ultra-sc.yaml to create the storage class. If you’d like to see the storage class in action, simply create a new file with following contents.
|
||
kind: PersistentVolumeClaim apiVersion: v1 metadata: name: dbench-pv-claim spec: storageClassName: managed-ultra accessModes: - ReadWriteOnce resources: requests: storage: 1024Gi Then run kubectl apply -f azure-ultra-pvc.yaml. That will generate an Ultra SSD Disk in the cluster’s resource group. You can view the disk by running az disk list --resource-group $mcrg
|
||
We can destroy the disk by running kubectl delete -f azure-ultra-pvc.yaml. These disks are pretty expensive, so I recommend deleting it as soon as possible.
|
||
Manual Disk Creation If you are in a situation where the storage class will not work because the zonality is random, then it is pretty easy to create an Ultra SSD disk and use it within a configuration. Create the disk by running the following.
|
||
az disk create -g $mcrg --name ultraSSD --size-gb 1024 --zone 1 --sku UltraSSD_LRS --disk-iops-read-write 160000 --disk-mbps-read-write 2000 az disk show -g $mcrg -n ultraSSD --query id -o tsv You are going to need the disk URI in order to attach it to a pod. Here is an example configuration that I used for storage testing.
|
||
apiVersion: batch/v1 kind: Job metadata: name: dbench spec: template: spec: containers: - name: dbench image: ndrpnt/dbench:1.0.0 imagePullPolicy: Always env: - name: DBENCH_MOUNTPOINT value: /data - name: DBENCH_QUICK value: "no" - name: FIO_SIZE value: 1G - name: FIO_OFFSET_INCREMENT value: 256M - name: FIO_DIRECT value: "1" volumeMounts: - name: dbench-pv mountPath: /data restartPolicy: Never volumes: - name: dbench-pv azureDisk: kind: Managed diskName: ultraSSD diskURI: DISK_URI cachingMode: None backoffLimit: 4 Simply update the DISK_URI placeholder with the correct value and save it. Then run kubectl apply to create it. Don’t forget to run az disk delete -g $mcrg --name ultraSSD at the end to clean up the Ultra SSD disk!
|
||
Conclusion The process of using Ultra SSDs with AKS is a bit more challenging than other types of managed disk. While all of these hurdles can be overcome, the biggest remaining challenge is dealing with the zonality of Ultra SSD disks. Ideally the Azure disk provider should be able to specify an AZ, rather than picking one at random. I’m sure that will be a feature added as availability zones in AKS move towards GA.
|
||
I hope this was helpful to some people out there! Let me know if you have additional questions or run into trouble.
|
||
`,summary:"In a previous post, I performed a storage performance benchmark of Azure Managed Disks and Azure Files for Azure Kubernetes Service. The testing included the now generally available Ultra SSD class of Managed Disk. The process for using Ultra SSD with AKS was fraught with peril, caveats, and an assist from the AKS product group to get it all working. I thought I would detail how I went about enabling Ultra SSDs with AKS in case someone else was struggling with the same.",date:"2 Sep, 2019",url:"https://nedinthecloud.com/2019/09/02/using-ultra-ssd-storage-with-azure-kubernetes-service/",image:"tutorials.png",readingTime:"9"},"https://nedinthecloud.com/2019/08/29/storage-showdown-at-the-aks-corral/":{title:"Storage Showdown at the AKS Corral",tags:["aks","azure","azure-kubernetes-service","kubernetes"],content:`This is a follow-up post to my analysis of using Azure NetApp Files for AKS storage versus the native solutions. After I wrote the post, with some surprising findings about Azure File performance, a number of people from Microsoft reached out to bring up a few key facts. In this post I will review the points that they brought up and include an updated analysis of the native Azure storage solutions for the Azure Kubernetes Service. Hold on to yer butts everyone!
|
||
There are basically two native solutions on Azure to provide persistent storage to AKS, Azure Managed Disks and Azure Files. In my original assessment, I used the Standard and Premium tiers of Managed Disk, but only the Standard tier for Azure Files. Either I misread the documentation or it was updated after my initial reading - this is the cloud after all and stuff is changing by the moment - and I thought only the Standard tier was supported for Azure Files. The Premium tier is definitely supported, and that is one of the things that the folks from Microsoft wrote in to let me know.
|
||
They also wanted to bring it to my attention that performance on premium storage is highly dependent on the amount of storage presented. That goes for both Managed Disk and Azure Files. Basically, the more storage you allocate, the more IOPS and throughput you get. I can only assumed that in both cases you are getting the aggregated throughput of multiple disks on the back end. By only using a single disk size of 500GB for my testing, I was restricting the max throughput of the Premium Managed Disk to 2300 IOPS and 150MBps. Standard tier for Managed Disks and Azure Files have a performance level that is unaffected by capacity, except in the case of very large disk sizes for Standard HDD disk. Standard HDD disks are pegged at 500 IOPS and 60MBps, which matches my findings of 555 IOPS and ~62MBps. Standard Azure Files are pegged at 1000 IOPS regardless of size, although apparently a 10K IOPS option is coming in the near future for larger file shares.
|
||
Another thing I discovered is that the default storage class for Azure Managed disks uses the Standard HDD class of disk and not Standard SSD. I was moderately surprised by this discovery, as I would expect Standard SSD to be the default for any new storage out there. The performance levels for IOPS and throughput are roughly the same, but Standard SSDs deliver lower latency. Standard SSDs do cost about 30% more, so maybe they were trying to save money for users? Who knows.
|
||
I also needed to consider the IO limitations of the Azure VM I was using for my AKS cluster. In my original test I was using the DS2_v2 class of Azure VM. The DS2_v2 has a max of 6400 IOPS and 96 MBps throughput for uncached workloads. The max bandwidth on the NICs is 1500Mbps, which is roughly 187.5MBps. The way I understand it, Azure Managed Disks are directly attached to the VM, so they are beholden to the storage restrictions of 6400 IOPS and 96MBps. Azure Files and the Azure NetApp Files are using the NICs to mount the storage, so their restriction would be the NIC bandwidth of 187.5MBps and whatever IOPS the storage solution can push. Looking at the data from my previous testing, the max throughput was Azure Files at 193MBps for a Read workload. Slightly above the stated limit, but not way outside of tolerances. The max IOPS for Managed Disks was 8159 for a Read workload. Again, above the limit of 6400, but not wildly out of spec. There could definitely be some caching going on as well.
|
||
Lastly, the Ultra tier of Azure Managed Disks is now GA. The performance of Ultra disks scales with capacity and has a theoretical max of 160,000 IOPS and 2,000MBps throughput for a 1TB disk. They are only supported on the ES and DS v3 Azure VMs, in select Azure regions, and must be deployed in an availability zone.
|
||
Test Planning In lieu of all the feedback I received, and things I discovered after the fact, it would appear that I need to update my testing methodology. First of all, I need to use a different family of VM for testing. The maximum IOPS for all the ES and DS v3 VMs is 128,000 cached or 80,000 uncached. The max throughput is 1024MBps cached or 1200MBps uncached. Based on the specs for the Ultra disk tier, I cannot select a VM that would actually achieve the max performance for a 1TB disk. To achieve the maximum available throughput, I will need to use a D64s_v3 VM, which will run me about $3.07 per hour. The price you pay for glory I suppose! If I want to test the Ultra tier, the AKS cluster will need to be created using the availability zone option in the East US 2 region.
|
||
I will test all six options for storage, with different sizes for each. Here is the rough testing matrix:
|
||
Storage Type Storage Class Storage Size Managed Disk Standard HDD 1TB Managed Disk Standard HDD 10TB Managed Disk Standard HDD 20TB Managed Disk Standard SSD 1TB Managed Disk Standard SSD 10TB Managed Disk Standard SSD 20TB Managed Disk Premium SSD 1TB Managed Disk Premium SSD 10TB Managed Disk Premium SSD 20TB Managed Disk Ultra SSD 1TB Azure Files Standard 1TB Azure Files Standard 5TB Azure Files Premium 1TB Azure Files Premium 10TB Azure Files Premium 20TB You’ll note that the Ultra Disk is only using a 1TB size. According to the documentation, any Ultra disk above 1TB has the maximum available performance, and therefore I see no reason to spend more money on a larger disk. The max size for an Azure Files Standard share is 5TB. Since the performance on the Standard Files is capped anyway, it probably doesn’t matter.
|
||
For the testing software, I am going to be using the same FIO project, with the same environment variables as last time.
|
||
env: - name: DBENCH_MOUNTPOINT value: /data - name: DBENCH_QUICK value: "no" - name: FIO_SIZE value: 1G - name: FIO_OFFSET_INCREMENT value: 256M - name: FIO_DIRECT value: "1" Performance Results Before I get to the results, I discovered several things during my testing that bear mentioning. First, there are two storage classes in AKS by default, default and managed-premium. As I mentioned before, the default class uses Standard HDD. What I didn’t realize until the test is that it has cachingMode set to ReadOnly. After the first test run, I saw that my Random IOPS Read test was getting 180k IOPS on the Standard HDD. That seemed… wrong. I ended up creating a new storage class called managed-standard-hdd with cachingMode set to None. Same thing with the managed-premium storage class.
|
||
I also learned that managed disks are provisioned in an availability zone, but not in a deterministic way. The managed disk would be provisioned in any zone that was part of the cluster definition, and if I didn’t have a node in that zone, the disk would never attach. To repeat that, managed disks in an availability zone can only attach to VMs in the same availability zone. If you are like me and trying to keep costs down by having a single node running in the cluster, then there’s a 66% chance the disk will be created in a zone where you don’t have a node. You can either running N number of nodes, where N is equal to the number of availability zones in your AKS cluster, or manually create the managed disk and set it to the same zone as your one node.
|
||
Configuring Ultra SSD was also challenging. So challenging, in fact, that I ended up reaching out to the AKS product team for help. Big shout out to Justin Luk and Jorge Palma for the assistance. I will cover the process of getting Ultra SSD working in a separate post, since this post is already fairly large.
|
||
Finally! To the numbers!
|
||
As in my initial post, we are first going to look at bandwidth for both random and sequential workloads.
|
||
Random Bandwidth It’s pretty obvious from the chart that the Ultra SSD crushed the random write test. It maxed out at basically 1GBps of bandwidth for random writes. I provisioned the Ultra SSD disk with 2000MBps of bandwidth, but as I mentioned earlier the max for a D64s_v3 VM is 1200MBps for uncached. I was hitting the limitation of the VM, not the Ultra SSD. It was a little surprising that the random read test came back under the Premium SSD numbers. I attribute this to the smaller volume size. Even though I am supposed to be able to get 2000MBps of throughput on a 1TB volume, there may be an advantage to provisioning more storage.
|
||
Aside from the Ultra SSD, it’s notable that Standard SSD drives at 10TB and 20TB are not that far behind the Premium SSD, and they also blow the Standard HDD out of the water for bandwidth. The key takeaway here is, if you want more performance from the Standard and Premium SSD disks, then you need to provision more storage. With Azure Files that does not appear to be the case. Even with 20TB of Azure Premium Files, the throughput was still about the same as 1TB.
|
||
Sequential Bandwidth Once again, the Ultra SSD crushed it for sequential write, but not for sequential read. At least it’s a consistent result. The Standard HDD disks did a lot better for sequential work, and that’s not surprising for a spinning medium. 20TB of Standard HDD actually did better than the 1TB of Premium SSD and matched the performance of Azure Files.
|
||
Now let’s move on to IOPS.
|
||
Mixed IOPS We can’t be terribly surprised that Ultra SSD killed it for IOPS can we? Remember I provisioned 160,000 IOPS for the Ultra SSD disk and I only hit 31,400. The max for the VM type is 80,000 IOPS, so I also wasn’t maxing out the VM. I’m not sure whether the IOPS limitation was on the CPU side or something to do with the container? Regardless, the max I got for NetApp’s Ultra tier was 32K IOPS, so the Ultra SSD disks are comparable. It’s also worth noting that Premium SSD and Premium Files were pretty similar in IOPS performance. Doing a quick cost exercise, 10TB of Premuim SSD is $1,638 and 10TB of Premium Files is $1,536. Azure Files charges for write, list, and read operations, so it would be a wash cost-wise. The nice thing is that Azure Files are network based and thus you don’t have to worry about the zonality of storage and the storage can be shared across multiple pods if desired.
|
||
I’d also like to point out that Standard SSDs really start performing at a capacity larger than 1TB. The performance between a 1TB SSD and 1TB HDD is very similar. But a 10TB SSD blows the 10TB HDD out of the water. Standard SSDs are more expensive than Standard HDD, but despite what the official numbers say on Microsoft’s docs, there is a significant performance difference in IOPS.
|
||
Random IOPS Here’s where the Ultra SSD really shined. Peaking at 48,900 IOPS for random writes, we still haven’t hit the VM or Ultra SSD maximums. There should still be more performance to be eked out if the proper tweaking was in place. Just as in the Mixed IOPS, the Premium SSD and Premium Files were fairly close in performance. The Standard SSD continued its pattern of outperforming the Standard HDD at larger sizes.
|
||
Average Latency I am not going to include average latency here. The results I got out of FIO were inconsistent and sometimes missing entirely. The Standard HDD was showing lower read latency than the Ultra SSD, for example, and I am pretty certain that isn’t correct.
|
||
Conclusion There’s probably a few key takeaways here. First, the amount of storage provisioned matters. Whether it’s standard or premium storage, more storage tends to results in better performance. It’s not a surprising result, but I am glad to see it proved out. Second, the VM type matters as well. My initial tests were performed on a VM that was undersized to the task. Third, the default storage classes may not match what you want in your AKS cluster. Definitely consider creating custom storage classes that define exactly what you are looking for. Finally, the Azure managed disks can only be attached to a single VM in a single zone. For the moment, that prevents them from being particularly useful to stateful workloads on an AKS cluster using availability zones. You’re going to need to go with Azure Files or some other network based solution.
|
||
Going forward I would like to tweak my testing to run each test multiple times and aggregate the data across runs to even out anomalies. I’m thinking I would need to have the container push the logs to an Azure Files share in JSON format, and then have something that ingests the logs and does the number crunching. But that is a story for another day.
|
||
`,summary:"This is a follow-up post to my analysis of using Azure NetApp Files for AKS storage versus the native solutions. After I wrote the post, with some surprising findings about Azure File performance, a number of people from Microsoft reached out to bring up a few key facts. In this post I will review the points that they brought up and include an updated analysis of the native Azure storage solutions for the Azure Kubernetes Service.",date:"29 Aug, 2019",url:"https://nedinthecloud.com/2019/08/29/storage-showdown-at-the-aks-corral/",image:"analysis.png",readingTime:"11"},"https://nedinthecloud.com/2019/08/09/ned-in-the-cloud-llc-three-month-update/":{title:"Ned in the Cloud LLC - Three Month Update",tags:[],content:`Three months ago I decided to leave the world of VAR consulting and try my hand at a new venture. That new venture is Ned in the Cloud LLC. I wrote a long post about the events that led up to my decision and I encourage you to go check that out if you have questions. The focus of Ned in the Cloud is to create technical content that is educational in nature. That could be courses on Pluralsight, sponsored blog posts about vendor technology, webinars about a technical topic, podcasts about the cloud, or even a book about the Azure Kubernetes Service. The unifying thread is a desire to learn about technology and share that knowledge with others. Now that I have been doing this for a full quarter, I thought it might be nice to post an update about how things are going so far.
|
||
When I started this voyage three months ago, there were several items of concern that fell into three basic categories: Financial, Business, Personal. Let’s dive into those categories a bit and see what I was worried about, and how I’ve fared in the last 90 days.
|
||
Financial Ned in the Cloud LLC is not a publicly traded company or anything, so I won’ be divulging any specific numbers. But I did have some concerns on the financial end of things. I was worried about building and maintaining a pipeline of work, effectively tracking my financials, dealing with accounting, and properly structuring for taxes.
|
||
When it comes to building and maintaining a pipeline, I needn’t have worried. In fact, I overstuffed the pipeline with work, and as a result the last two weeks of May and all of June were completely packed with work. I was working eight hour days and sometimes nights to hit all my deadlines. This was not my intention when I decided to strike out on my own. I’ll talk about that more in the personal section, but suffice to say that working less was one of my goals when I quit my previous job. In my worried state over pipeline, I made the mistake of saying “yes” to any opportunity that presented itself. My key takeaway is that it’s okay to say “no” to opportunities when you don’t have sufficient bandwidth to get them done. Saying “no” doesn’t slam a door on future opportunity, in fact it communicates to the prospect that you are in high demand! I have also created an Excel spreadsheet to track my pipeline of opportunities to help make financial projections. It’s very much a work in progress, but it’s nice to have some visibility into future earnings.
|
||
In terms of tracking my financials and dealing with accounting, I got a free 12-month membership to Quickbooks Self-Employed because I did my taxes through TurboTax at the small business tier for 2018. I’ve been using Quickbooks to track my income and spending. One of the things I wish I had done sooner was to get a dedicated checking account and business credit card to completely separate business and personal finances. I had the checking account in place three months ago, but I didn’t have the credit card until mid-June. Once I had both, I went through all the business related accounts, like AWS and my Wordpress hosting, and moved them to the business credit card. That makes tracking finances way easier. I can also use Quickbooks to do invoicing for clients. I won’t be using TurboTax for my taxes next year, more on that in a moment, but I will probably pay up for another year of Quickbooks. The key takeaway I have here is to separate my finances as much as possible, and avail myself of software that makes the whole process simpler.
|
||
There’s only two things that are certain in life, death and taxes. And I’m not so sure of the death thing. Back in April, I engaged with an accountant to figure out how to structure my company for tax purposes. Based on the expected income, he recommended that I structure as an S-corp. Of course, being an S-corp is much more complicated than continuing to file as an individual. But there are lots of potential savings when it comes to taxes. If I filed as a self-employed person, then all of my income would be subject to standard income tax rates, whereas as an S-corp only the salary I pay myself is subject to said taxes. It also means that I need to pay unemployment and worker’s compensation taxes. Things get complicated quickly, so I hired the accountant I mentioned earlier to handle all of my tax related affairs. Unless you really like digging into the fascinating world of corporate taxes, I highly recommend getting a good CPA and letting them handle it.
|
||
Business A business is more than just financials; it’s marketing, IT, customer service, strategy, etc. When I was working for someone else, I could leave many of the business related activities to them and focus on my job responsibilities. What I quickly realized in starting my own business is that I am now responsible for all of these activities. No one is going to take care of my marketing, set the strategy for the next three years, or manage my ongoing projects.
|
||
To that end, I have started using several applications to try and manage my workflow and marketing. I was already using Buffer and Zapier to automate posts across social media, and I have continued to lean in to those platforms for marketing. I also had to go through all of profiles on multiple websites and update my status to include the company website and contact information. Here is the short list of apps that I’ve started using:
|
||
Office 365: mostly for email and OneDrive Slack: multiple groups I work with use Slack for comms Trello: for managing the various projects I’m doing Toggl: for tracking time spent on projects in Trello Ringr: for recording podcasts for Day Two Cloud I am also forced to use some other solutions as part of working with partners. For example, I have to use Google Docs, Drive, and Sheets with a few partners because that is their workflow. Another one uses Asana, which I’ve never touched before. I am the small fish in all of these situations, so I need to adapt to whatever each partner prefers to use for their workflow.
|
||
My larger strategic vision for the company has not changed. I don’t have a mission statement as such, but it really boils down to creating technical content for the purpose of educating others. That’s the high level vision, and it’s up to me to transform that vision into a strategy and a tactical plan. Currently my largest partner is Pluralsight, from both a workload and revenue perspective. It is a fantastic platform to create educational content of a technical nature, but I want to diversify and add new clients to the roster. Generally you don’t want a majority of your work and income coming from a single source. That constitute a significant business risk if something happens to the relationship with that client. Fortunately, I have been building relationships with others in the technical community, and that has resulted in several opportunities to write, podcast, and host webinars. Over the next few years I am planning to continue to cultivate those relationships while also looking to create some original content that is more direct to consumer, so I am not entirely reliant on clients and partners for my income.
|
||
Personal There were two things I was concerned with when it came to my personal life: maintaining work/life balance and working from home.
|
||
I’ve heard some serious horror stories about people starting their own business and suddenly working 80 hour weeks and questioning why they ever did this foolish thing to start with. I did not want to become one of those people. Prior to quitting my previous job, I was already working way too much. I had my regular 40-hour a week job, and then I was doing a bunch of things to build up my side-hustle, which would eventually become my full-time job. Authoring Pluralsight courses, writing sponsored blog posts, recording podcasts, all of these things were adding at least another 20 hours of work to my week. The whole point of doing all this side-hustle work was to build up a business that was ready to launch without the 80 hour work weeks. I am happy to report that my plan worked. I’ve been able to maintain a maximum of 40 hours a work, with the exception of the end of June. In my worry over having enough work in the pipeline, I over-committed myself in June, and as a result I had to work a few nights in the final two weeks to hit my deadlines. Then I took the first two weeks of July off, restoring balance to my life.
|
||
I am a creature of routine, as I think many of us are. Part of maintaining a good balance and my own mental health was establishing a new routine around the house. Even though I set my own hours, I still need some semblance of structure. Previously I would get all the kids up by 6:30AM, get them fed and dressed, and be out the door by 7:15AM or so. I would drop the kids off and be at work by 8:00AM and work until 4:30PM. Even when I worked from home, that was my general schedule. Now that I am able to make my own schedule, it didn’t make sense to drop the kids off so early, and I could save money by not sending my oldest to before-care at the elementary school. Instead, I settled into a routine of getting them up by 7:30AM - if they didn’t wake themselves up - and heading out by 8:15AM for drop-off in the car line. By the time I dropped them all off and returned home, I was able to eat breakfast and start work by 9AM. I would work until 4:30PM or 5PM at the latest, with a 30 minute break for lunch. Even once school ended and summer started, I have kept this schedule. My pro-tip here is to create a schedule for yourself and stick to it as much as possible. It provides a solid separation of work and personal life, which could otherwise become a bit blurred.
|
||
Speaking of work and life being blurred, another thing I did before leaving my old job was to get a dedicated office space in the basement of my house. We had wanted to get the basement finished since we moved into the house, and having a dedicated office served as a catalyst to finally doing it. Now I have a space carved off that is specifically for doing work, with a door that shuts and locks. When I am in my office, I am at work. Sometimes I do some work sitting in the sunporch or in front of the TV, as I did before I left my old job, but the majority of my work is happening in my office. Also important, since it is summer and the kids are home, they know that when I am in my office, I am not accessible. Go ask mom, unless it’s an emergency. And no, being unable to find your missing Shopkins does not constitute an emergency.
|
||
I did have two more concerns of a personal nature that my prior work addressed. First, I was worried about not getting enough social contact. Like many IT professionals, I am not the most social person. I do not need to spend all day around other people, and in many cases I find that I am more productive when left alone. But there is a lot to be said for the casual interaction that happens around an office. You have random conversations that spark a new idea, or a hallway chat helps solve an issue you were kicking around. Getting up from your desk, and getting out of your own head-space is important. By working 100% remote and without any actual co-workers, I was giving up that aspect of social interaction. My solution, so far, has been to make sure and go out to lunch with someone at least once a week. That gets me out of the house, and it gives me a reason to reach out and connect with people I might not have seen in a while. I have also been making ample use of online social interaction with things like Slack, Zoom, WhatsApp, and even Twitter. Finally, I go back to my old work every Friday to record Buffer Overflow, providing yet another avenue of social interaction.
|
||
The second concern I had was one of travel anxiety. I’ll be blunt here, travel makes me anxious and my anxiety expresses itself through intestinal discomfort. When I say travel, I could mean driving to a client site, flying to a conference, or taking the train up to NYC. It’s not the mode of transport necessarily, although driving myself makes me the least anxious. It is more the uncertainty and loss of control that is inherent in travel. You are no longer in a space that you control, i.e. home. And unless you are driving a car by yourself, you no longer control your mode of transport either. What I have discovered is that the more I travel, the less anxious travel makes me. My brain gets conditioned to frequent travel, and doesn’t have time to fixate on a particular trip. Working as a consultant for six years meant a lot of travel - mostly local, but still a lot of travel - to unfamiliar places to meet new people. The repetition of doing so has built up a level of tolerance to travel that I didn’t used to have. My fear was that I would lose that level of tolerance if I were no longer traveling as much with the new job. At this point, I don’t know if my tolerance is changing, since I haven’t been traveling for work at all since May. I do have a number of trips coming up in late summer and early fall that will test my tolerance.
|
||
I don’t really have a way to address this issue. As I said, traveling for consulting was mostly day trips to regional locations, so I was always home at night with my family. There were occasional conferences that had me going away for multiple days, but the majority of my travel was to client sites. My new job doesn’t require me going to client sites on a regular basis, so that option is out. I could go to more conferences, but I don’t want to. I value my time with my family too much to spend a week away every month at a different conference. For now, this concern appears to be a wait and see type situation. All I can do is be mindful of it.
|
||
Conclusion Whew! I guess I had a lot to say after 90 days! If you’re thinking about going independent, I hope you’ve found some part of this helpful. If you’ve got advice for me moving forward, I’d love to hear it. Overall, I have to say that I LOVE my new job. It has improved my mental state, my home life, and my enjoyment of work. Financially it was a pay-cut, but personally it was a huge raise. I know which one I value more.
|
||
`,summary:"Three months ago I decided to leave the world of VAR consulting and try my hand at a new venture. That new venture is Ned in the Cloud LLC. I wrote a long post about the events that led up to my decision and I encourage you to go check that out if you have questions. The focus of Ned in the Cloud is to create technical content that is educational in nature.",date:"9 Aug, 2019",url:"https://nedinthecloud.com/2019/08/09/ned-in-the-cloud-llc-three-month-update/",image:"featured-image-nedinthecloud.jpg",readingTime:"13"},"https://nedinthecloud.com/2019/07/23/azure-netapp-files-performance-with-azure-kubernetes-service/":{title:"Azure NetApp Files Performance with Azure Kubernetes Service",tags:["aks","azure","cfd3","netapp"],content:`In April of 2018, I was delegate for Cloud Field Day 3. One of the presenters was NetApp, and they showed off a few different services they had under development in the cloud space. In a previous post I went over the services in some detail, so I won’t regurgitate all that now. One of the services that was still in private preview at the time was NetApp Files for Azure. The idea was relatively simple, NetApp would place their hardware in Azure datacenters and configure the hardware to support multi-tenancy and provisioning through the Azure Resource Manager. That solution is now generally available, and I was curious how it would perform in comparison with the other storage options for the Azure Kubernetes Service (AKS). In this post I will detail out my testing methodology, the performance results, and some thoughts on which storage makes the most sense for different workload types.
|
||
NetApp Files The process for getting access to Azure NetApp Files is a little more complicated than I would have liked. Most services on Azure are enabled by default, and you simply have to go through the resource creation process. In order to use NetApp Files, you first have to fill out a Microsoft Form, and wait for someone at NetApp to get back to you. They will want to know which subscriptions to enable the service for, and will also want to have a phone call to discuss your use case. I was… not interested in talking to a NetApp sales person. They did enable my subscription for the service, and then I had to go into Azure PowerShell and enable the Resource Provider for Microsoft.NetApp. The entire process took a couple of days, longer than I would have liked for a cloud service. I understand that they have limited capacity on the hardware in the various datacenters, so they don’t want people just spinning up NetApp capacity pools willy-nilly. At the same time, this is the cloud and I am impatient.
|
||
Regardless, I gained access to the service. There are three tiers of performance for NetApp Files: Standard, Premium, and Ultra. That follows the same categories as Azure Storage, which also exists as Standard, Premium, and Ultra. The process for consuming NetApp Files is to create a NetApp Files account, add a Capacity pool of the performance level you need, and then carve that capacity pool up into volumes. The volumes are exported to a dedicated subnet on an Azure Vnet. The process is very straightforward.
|
||
For my testing, I created a capacity pool for each of the three performance levels, and then exported a 1TB volume to an Azure Vnet that also had my running AKS cluster. My AKS cluster was a single worker node of size DS2_v2. There was nothing running on the cluster aside from the default pods and a tiller pod for Helm. I also wanted to compare the performance to the other storage options on AKS:
|
||
Azure Managed Disks Standard Azure Managed Disks Premium Azure Files Azure Files uses the standard tier for performance. There is a Premium tier of Azure Files that is not yet available for consumption by AKS. There is also the aforementioned Ultra tier for Azure Disk, which also is not yet supported. I’m sure both are forthcoming, and perhaps I will update my post when that happens.
|
||
**Update:**As a couple commentators point out, Azure Files at the Premium SKU is supported with dynamic provisioning for AKS clusters running K8s version 1.13.0 and higher. I have run the performance tests against the Premium tier, and I will document the results in a separate post dealing exclusively with Azure File.
|
||
The performance testing is using the FIO project packaged into a container. I am following the guidance of this GitHub repository, except the container image is no longer available on Docker Hub. Someone else created an identical image based on the DockerFile, so I decided to use that. My tests are stored on my own GitHub repository for your reference if you wanted to recreate the performance testing for yourself. I tried presenting the NFS storage from NetApp using both the NFS-Client provisioner deployed using Helm and also the more direct process of creating a persistent volume with an NFS server and mount specified. From a performance standpoint, they both appeared to be identical.
|
||
For the dbench tests I used the following settings:
|
||
env: - name: DBENCH_MOUNTPOINT value: /data - name: DBENCH_QUICK value: "no" - name: FIO_SIZE value: 1G - name: FIO_OFFSET_INCREMENT value: 256M - name: FIO_DIRECT value: "1" There are many possible variations for these settings, and I would encourage you to test based on an actual workload you plan to use. For me, I didn’t have a specific workload in mind, so I used these settings for all the different storage types. I ran the tests one at a time without anything else running on the cluster to prevent any potential contention for resources. Let’s check out the performance metrics for each test.
|
||
Performance Benchmarks First let’s look at bandwidth for random and sequential workloads.
|
||
Random Bandwidth Storage Type Random Reads Random Writes Azure Standard Disk 63.9 63.1 Azure Premium Disk 60.4 57.2 Azure Files 125 96.3 NetApp Standard 16.1 16 NetApp Premium 64.3 64.6 NetApp Ultra 128 128 There’s a couple surprising findings here. For starters, Azure Standard and Premium disk are essentially the same for both read and write bandwidth. The second big surprise is that Azure Files is as performant or close to the Ultra class of the NetApp files. If you’ve got a bandwidth hungry application on AKS, Azure Files appears to be the way to go.
|
||
Sequential MBps Storage Type Sequential Reads Sequential Writes Azure Standard Disk 63.8 63.1 Azure Premium Disk 63.2 62.5 Azure Files 193 109 NetApp Standard 16.2 18.3 NetApp Premium 63.9 64.7 NetApp Ultra 128 128 Once again we’re seeing that Azure Standard and Premium disk have roughly the same performance levels, and Azure Files is outperforming both. The real kicker here is that Azure Files outperforms NetApp Files Ultra by a decent margin, 193MBps vs 109MBps. Whether your workload is random or sequential, Azure Files seems to be the way to go for bandwidth.
|
||
Random IOPS Storage Type Random Reads Random Writes Azure Standard Disk 8159 555 Azure Premium Disk 7641 2306 Azure Files 735 1010 NetApp Standard 4094 4093 NetApp Premium 16400 16400 NetApp Ultra 32800 32900 While Azure Files may rule the roost for bandwidth, it doesn’t hold a candle to all the other storage types for IOPS. Azure Standard and Premium Disks seem to have roughly the same read IOPS, but Premium starts edging out Standard for writes, 2306 vs 555. The real standout in this case is NetApp Files. Standard (4096), Premium (16400), and Ultra (32900) are consistent on both Read and Write. While Standard is slightly below Azure Disk, the Premium and Ultra tiers dominate all other options. If you’ve got a workload that needs a ton of random IOPS, NetApp Files is the way to go.
|
||
Mixed IOPS Storage Type Random Reads Random Writes Azure Standard Disk 1670 558 Azure Premium Disk 5723 1915 Azure Files 770 257 NetApp Standard 3072 1018 NetApp Premium 12300 4100 NetApp Ultra 24600 8186 The picture for mixed IOPS is roughly the same, although in this case the Azure Premium Disk outperformed the Standard Disk by at least a factor of 3x. Azure Files was the worst with a meager 257 IOPS for mixed writes. The NetApp tiers once again shined, although this time the write IOPS were about a third of the read IOPS numbers. Still, if you need IOPS, NetApp Files is the way to go.
|
||
You might be curious what the difference is between the random and mixed IOPS tests. Looking at the docker-entrypoint.sh file in the dbench repo, the Random IOPS is made up of two separate runs - one for Read and one for Write. The Mixed IOPS is a single run that has a mix of read and write operations. The relevant setting in FIO is readwrite, which is set to randread, randwrite, and randrw respectively. The Random IOPS setting of randrw includes additional switch options to configure the mix of operations called rwmixread and rwmixwrite. The dbench setup uses a value of 75 for rwmixread which means that 75% of the operations will be read. You can check out the official docs to learn more about FIO.
|
||
In a nutshell, the Random IOPS tests are run separately, so each set of results is for read or write operations only. The Mixed IOPS test runs a balance of 75% read and 25% write in the same test run, giving you a nice blend.
|
||
Average Latency Storage Type Avg Latency Reads Avg Latency Writes Azure Standard Disk 489 7.3 Azure Premium Disk 518.9 3511.7 Azure Files 22.5 5.1 NetApp Standard 973.4 973.5 NetApp Premium 416 463.4 NetApp Ultra 413.1 465 As opposed to the other metrics - where higher is better - latency is generally considered a bad thing. Lower latency, therefore, is better. And you can’t get much lower than the latency on Azure Files! The measurement here is in microseconds, abbreviated as usec. The latency for Azure Files is so low as to be suspicious, as if the information is somehow being cached locally and not actually reaching out to Azure Files at all. The other big standout here is that the average latency for Azure Premium Disk is so much higher than all the other storage types. There’s definetely something weird going on here, and it would bear further investigation, possibly as a separate blog post.
|
||
NetApp Files is what you might expect, not as good as Azure Disk at the Standard level and better than Azure Disk at the Premium and Ultra levels. Based on the graph, I would recommend Azure Files for latency, but with a grain of salt. There something that doesn’t quite sit right with all these numbers, and I need to find out what is really going on.
|
||
Conclusion The point of this process was to compare some different storage options for AKS, and the results were somewhat surprising. If pure bandwidth is what your application craves, then Azure Files is probably the way to go. On the other hand, if pushing IOPS is your goal, then NetApp’s Premium and Ultra tiers have you well covered. In the world of latency it appears that Azure Files is the champion, but again I caution you to try it out for yourself. Another key finding is that there is not a ton of difference between Azure Standard and Premium disk. The IOPS for mixed workloads was better, but other than that it doesn’t make sense to pay the extra for Premium storage.
|
||
`,summary:"In April of 2018, I was delegate for Cloud Field Day 3. One of the presenters was NetApp, and they showed off a few different services they had under development in the cloud space. In a previous post I went over the services in some detail, so I won’t regurgitate all that now. One of the services that was still in private preview at the time was NetApp Files for Azure.",date:"23 Jul, 2019",url:"https://nedinthecloud.com/2019/07/23/azure-netapp-files-performance-with-azure-kubernetes-service/",image:"tutorials.png",readingTime:"9"},"https://nedinthecloud.com/2019/07/16/demystifying-azure-ad-service-principals/":{title:"Demystifying Azure AD Service Principals",tags:["azure","azure-ad","powershell"],content:`Anyone who’s worked with Azure for a bit has encountered the need to create a service principal. If you are an IT Ops person, you probably equate an SP with a service account in local Active Directory. If you’re more of an application developer, then you may have created an SP as part of your application in Azure, because you want to give that application permissions to Azure resources. The purpose of this post is to tease apart what service principals are, how they interact with application objects, and all the myriad ways to create an SP on Azure.
|
||
Azure Active Directory Applications The first thing you need to understand when it comes to service principals is that they cannot exist without an application object. If that sounds totally odd, you aren’t wrong. The service principal construct came from a need to grant an Azure based application permissions in Azure Active Directory. Since access to resources in Azure is governed by Azure Active Directory, creating an SP for an application in Azure also enabled the scenario where the application was granted access to Azure resources at the management level. That is all very useful in the context of creating applications in Azure. If you wanted an account in Azure AD to use for automation or to power a service, then you could use the same construct. Instead of creating a separate object type in Azure AD, Microsoft decided to roll forward with an application object that has a service principal. For the purposes of using an SP like a service account, the application it creates as part of the process sits unused and misunderstood.
|
||
To make things even more confusing, a single application object can have multiple service principals across different Azure AD tenants. When an application object is registered with the home tenant, an SP is also created in that Azure AD tenant. If the application being developed is a single-tenant application, that’s the only SP needed. But if the application is meant to be multi-tenant - by which I mean multiple Azure AD tenants - then it will have an SP created in each tenant that uses it. The consent process of enabling an application for your Azure AD tenant includes creating and granting permissions to that application object in the form of an SP in your tenant.
|
||
Of course, if your whole goal was to use a service principal to do some automation, then you don’t care about any of this nonsense. You just want to create an SP and be done with it.
|
||
Service Principals As an IT Ops person trying to get some work done, you don’t care about the application object. You probably don’t want to deal with the application object. You just want to create an SP. There is NO way to do this without also creating an application object. That’s the decision that Microsoft made, and it seems to be sticking with it. What that means is that depending on which tool you use to create a service principal, you may need to create an application object first. For instance, the portal requires that you create the application object first and doesn’t even mention the service principal as a construct. If you are using a different tool, it may automatically create that application object for you. For instance, the Azure CLI allows you to directly create an SP, and it will take care of creating that application object for you in the background. How helpful! The downside is that there are so many different tools to use with Azure, and they ALL seem to have a different workflow. You can create an SP by using:
|
||
The Azure Portal The AzureAD PowerShell Module The Az PowerShell Module The AzureRM PowerShell Module The Azure CLI Other Azure SDKs The Azure AD API The Microsoft Graph API Holy cow! What’s a poor IT Ops person to do? Let’s break it down with what will likely be the most common ways you will create a Service Principal.
|
||
Creating Service Principals Microsoft Graph API The one thing that all of these tools have in common is that they are all referencing one of two APIs to perform their operations. Microsoft is in the process of deprecating the Azure AD API in favor of the Microsoft Graph API. All the other methods are using some kind of SDK to interact with one of these two APIs. If you’re curious about the Azure AD API, the relevant sections for the application and service principal objects can be found in the entity and complex types area of the docs. The Microsoft Graph API docs seem to be a little better organized, and you can find information on applications and service principals. Funny thing that I noticed, there is no create function for the service principal object. The resource appears to be implicitly created when an application is registered with a tenant.
|
||
The Azure Portal The experience for registering an application and creating a service principal has changed recently. I’d like to say it makes more sense now, but I would be lying. As far as I can tell it’s more confusing with check boxes that don’t fully explain what they want you to do. Regardless, if you want to create a service principal through the portal, just follow these directions.
|
||
The AzureAD PowerShell Module If you don’t already have the AzureAD PowerShell module, you can install it by running Install-Module AzureAD -Force. Then run the following commands:
|
||
Connect-AzureAD -TenantId "YOUR_TENANT_ID" $myApp = New-AzureADApplication -DisplayName "AzureAD Module App" -IdentifierUris "https://azureadmoduleapp" $mySP = New-AzureADServicePrincipal -AppId $myApp.AppId Obviusly, the AzureAD module does not take care of creating the application object for you. You have to do that first and then create the SP. The commands above will get you a service principal, but without any type of credentials to login. If you want a password associated with the service principal, then you can run the following:
|
||
$spCredParameters = @{ StartDate = [DateTime]::UtcNow EndDate = [DateTime]::UtcNow.AddYears(1) Value = 'MySuperAwesomePasswordIs3373' ObjectId = $mySP.ObjectID } New-AzureADServicePrincipalPasswordCredential @spCredParameters Now you have a service principal that you can assign roles and permissions to.
|
||
The Az PowerShell Module There’s a new Azure PowerShell module on the block. The Az module is replacing the original AzureRM module. The reason? Partly, Microsoft just wanted to shorten the commands by five letters. They also wanted to rewrite the module to take advantage of new functionality in PowerShell and in Azure and get rid of some of the old commands that maybe weren’t following best practices. Regardless, this is the module you’ll be using to do things with Azure going forward. If you’re currently running AzureRM, beware here there be dragons. You need to completely remove AzureRM first, or install PowerShell 6 and run the Az module in PowerShell 6 context instead. To create a service principal with the Az module, run the following commands:
|
||
Connect-AzAccount $mySP = New-AzADServicePrincipal That’s it. The good news is that the command creates the application in the background for you. And that is pretty much where the good news ends. In this case, the command creates a service principal with a display name that starts azure-powershell- and appends the current date and time. It also gives it a secret of the type System.Security.SecureString which is not particularly useful. It is possible to decrypt it, but I would recommend setting a password credential manually like we did in the AzureAD module example. You might think that there is a command like New-AzureADServicePrincipalPasswordCredential in the Az module, and you would be partly correct. There is the New-AzADSpCredential command, but that only allows you to add a certificate type and not a password. No idea why that choice was made. Maybe because Microsoft hates passwords?
|
||
Before we get into the process for creating a password based credential, which I assure you is non-intuitive and annoying, I would first like to point out something that really annoys me. The service principal object from the AzureAD module isn’t the same type as the service principal object from the Az module. If you run Get-Member on the SP object from the AzureAD module you get the TypeName Microsoft.Open.AzureAD.Model.ServicePrincipal, whereas with the Az module you get the TypeName Microsoft.Azure.Commands.Resources.Models.Authorization.PSADServicePrincipalWrapper. The properties exposed in each object type also differ. The AzureAD module exposes 25 different properties, and the Az module exposes only 7. If I might present a table for comparison:
|
||
Az AzureAD ApplicationId AppId DisplayName DisplayName Id ObjectId ObjectType ObjectType Secret N/A ServicePrincipalNames ServicePrincipalNames Type N/A Right off the bat, the ApplicationId is named differently across the two objects. So is the ObjectId. And it’s not even consistent in its inconsistency. The Az modules uses the longer ApplicationId property and the shorter Id property. Then there is the Secret property, which is really just the value stored in one of the keys in the PasswordCredential property. A service principal can have multiple passwords - aka secrets - which are held in an array in the PasswordCredential property. The PasswordCredential property is an object type of Microsoft.Open.AzureAD.Model.PasswordCredential. There is a separate KeyCredentials property and object type that houses certificate based authentication. The New-AzADSpCredential command takes a cert value and adds it to the service principal in a property that is not even exposed by the Az implementation of the service principal type. That means you need to run the Get-AzADSpCredential command to get the value back. In order to create a password based credential, you have to create PasswordCredential object directly like this:
|
||
$credProps = @{ StartDate = Get-Date EndDate = (Get-Date -Year 2024) KeyId = (New-Guid).ToString() Value = 'MySuperAwesomePasswordIs3373' } $credentials = New-Object Microsoft.Azure.Graph.RBAC.Models.PasswordCredential -Property $credProps Set-AzADServicePrincipal -ObjectId $mysp.Id -PasswordCredential $credentials That’s if you want to configure a password after creation of the SP. If you wanted to set the password while creating the service principal, you have to create a completely different object type. When you run New-AzAdServicePrincipal with the PasswordCredential parameter, the command is expecting an object of type Microsoft.Azure.Commands.ActiveDirectory.PSADPasswordCredential. When you run Set-AzAdServicePrincipal with the PasswordCredential parameter, the command is expecting an object of type Microsoft.Azure.Graph.RBAC.Models.PasswordCredential. And it will not do an implicit conversion for you! Instead you have to do the following:
|
||
$credProps = @{ StartDate = Get-Date EndDate = (Get-Date -Year 2024) Password = 'MySuperAwesomePasswordIs3373' } $credentials = New-Object Microsoft.Azure.Commands.ActiveDirectory.PSADPasswordCredential -Property $credProps $sp = New-AzAdServicePrincipal -DisplayName "WhateverItDoesntMatter" -PasswordCredential $credentials You may have also noticed that the two different object types have different arguments. One expects a KeyId and Value and the other expects a Password argument only.
|
||
My advice is this. Don’t use the Az module for managing Azure AD resources. It’s a hot mess.
|
||
The Azure CLI Unlike the PowerShell modules, the Azure CLI is written in Python. The process for creating a service principal is simple. Run the following command:
|
||
az ad sp create-for-rbac -n "MySpCLI" The command will create the application object in the background for you. And the output will include all the information you need to use the service principal, including the password in clear text.
|
||
{ "appId": "d9bdde72-73eb-4bc2-8dbf-0855c0d19820", "displayName": "MySpCLI", "name": "http://MySpCLI", "password": "7778ef85-2c5a-4f8d-81d5-88389c9a4feb", "tenant": "f06624a8-558d-45ab-8a87-a88094a3995d" } That’s all there is to it. Super easy and simple.
|
||
Conclusion I started this post hoping to demystify the application and service principal relationship and shed some light on how to use different tools to accomplish the same goal. At the end, I may have made things a little more confusing. But that simply reflects the confusing nature of service principal kludge. My advice would be to use the Azure CLI to create a service principal. The command is simple. You can run it from the Cloud Shell if you don’t have the Azure CLI locally. It is faster than using the portal, and easier than using PowerShell. Additionally, many resources in Azure now have the ability to use Managed Service Idenities (MSIs) to access other Azure resources. You still need service principals for some use cases, but I would highly recommend checking to see if an MSI can meet your requirements.
|
||
`,summary:"Anyone who’s worked with Azure for a bit has encountered the need to create a service principal. If you are an IT Ops person, you probably equate an SP with a service account in local Active Directory. If you’re more of an application developer, then you may have created an SP as part of your application in Azure, because you want to give that application permissions to Azure resources. The purpose of this post is to tease apart what service principals are, how they interact with application objects, and all the myriad ways to create an SP on Azure.",date:"16 Jul, 2019",url:"https://nedinthecloud.com/2019/07/16/demystifying-azure-ad-service-principals/",image:"tutorials.png",readingTime:"10"},"https://nedinthecloud.com/2019/07/06/privilege-and-perspective/":{title:"Privilege and perspective",tags:[],content:`Let me say this upfront. I am going to talk about privilege. I am not an expert on privilege. I am not a poly-sci major with dual minors in social justice and psychology. These are ideas that have been kicking around in my head for a couple weeks, and I thought they might resonate with others or spark conversation. If that’s not the sort of thing you’re interested in, probably best to skip to the next post, where you can read about HashiCorp Vault or Azure or something similar. Okay? Still here? Great. Let’s talk.
|
||
I am INCREDIBLY lucky. I was born at the right time, in the right place, to the right parents. I had access to a great education, solid role models, and a supportive network of people. That’s not something I worked hard to get. That’s something I had from day one, because I am incredibly lucky. As a teenager, I would classify myself as an obnoxious little shit who took all he had for granted, and still managed to complain about how the universe was unfair. I was self-centered and ego-centric in the way that only a teenage boy in the suburbs with no real problems can be. I intentionally hung with the wrong people and got in with the wrong crowd so I could have “real” problems. But that was always a facade, a vacation into some else’s crappy reality, because at the end of the day I knew that I could get my act together, ask my parents for help, and I would be fine. That’s privilege. The knowledge that no matter how bad you mess up - within limits, of course - you’re probably going to be okay.
|
||
The problem I had was one of perspective. I absolutely did not understand how other people struggle to scrape by. Didn’t they know how easy it was to get a good job and move up through the ladder? One of my jobs in high school was a stock boy at Wawa - it’s a convenience store. The day I turned 18 they made me a Shift Manager. I thought to myself, “This is so simple. Anybody can get what they want if they work hard.” What I didn’t appreciate at the time, was the fact that I was a well-educated, semi-responsible person, in large part due to my upbringing. I knew how to act appropriately at a job, and how to speak to my superiors. It’s not that I didn’t have to work hard, but I knew how to work in a way that would end up with a promotion. That was lucky. I was lucky. I might also mention that I had a family that could afford to lend me a car to drive to work. And a stable enough household that I could reliably show up for work. Not everyone has that. Again, very lucky.
|
||
When I decided to move into the world of IT, again my privilege helped tremendously. I had access to computers from a young age. We had a personal computer at home in the early 80s. My school had an enrichment program that included some basic programming on a PC. I grew up immersed in technology, from the constant upgrade cycle of our home computer to the introduction of game systems, and the presence of technology at school. When I went to college, I majored in Computer Engineering for a year and a half before I dropped out. All costs covered by my parents of course. And then finished a two-year degree in Computer Science, also covered by my parents. When I started interviewing for tech jobs a couple years later, I already had a solid grounding in technology. I got hired after my first interview. That’s ridiculous. I know lots of people that had to go through 10 or more interviews to break into tech. Me? I went on one. I’m sure I did well in the interview, but it didn’t hurt that I had years of privilege lending a helping hand.
|
||
It was not a glamorous position. It was level one helpdesk. My time in the service industry served me well here, since a lot of being good at helpdesk is being polite and providing good customer service. By the way, those are almost always good things to be. A year and a half later, I had moved up to third-level support, and then the company declared bankruptcy. During that time I had two amazing bosses that gave me the opportunity to learn and grow, and were super flexible about my schedule. I had gone back to school to get my bachelor’s degree. This time paying for it with my money.
|
||
That was sixteen years ago, and since that time my career has done pretty well. I’ve had to work hard, but I’ve also been lucky to have a string of really good bosses and coworkers. People who mentored me, and helped me develop both my technical and soft skills. I was given the tools to do well at an early age, and despite a few missteps, I’ve managed to use those tools to my advantage. But I can’t ignore the fact that I have been given tremendous privileges that many people do not get.
|
||
So what is the point? There’s got to be a point. The reason I started thinking about all this was a few posts I saw on LinkedIn and Twitter that really got my blood boiling. I’m not going to name names or anything, but the sentiment of it was something like this:
|
||
If a meeting doesn’t bring you joy, you should refuse to accept it. Find a job that brings you happiness. Time is the only finite resource, spend it wisely. I know these are somewhat innocuous koans that people like to toss out from time to time. But they are also tremendously pretentious, privileged bullshit. You know what doesn’t bring me joy? A weekly sales meeting. You know what happens if I don’t attend my weekly sales meeting? I would probably get told that I have to, and eventually fired. If you’re the CEO of a company, then by all means turn down whatever meeting doesn’t bring you joy. For the rest of us? You should probably go to that mandatory meeting so you can keep your job. You know, so you can feed your children?
|
||
Find a job that brings you happiness? No. Find a job that pays you a living wage, hopefully. When you work in an industry that has near negative unemployment, it’s hard to remember that the rest of the world isn’t like that. The IT industry is incredibly privileged at the moment, even outside the bubble of Silicon Valley. Most people outside of tech don’t get to choose between six different, six-figure salary packages. A lot of people have a shitty job, and they work that shitty job because if they don’t, they will have no money and eventually be homeless. And by eventually, I mean very quickly. Not everyone has a nest egg or severance pay floating them to their next dream job. The conceit of this particular phrase is that you have the means and luxury of being discriminating when it comes to your next job. Trust me when I say that no one working in the mall picked that job because it would bring them happiness, they chose it because it pays money and they need money to live.
|
||
Time is a finite resource. That’s true. And spending it wisely, also true. Again, the conceit here is that you are in charge of the vast majority of your time. I guess technically you are. You can quit your shitty job, and go spend every day at the beach. Unless of course you want to, I don’t know, eat? Or have a place to sleep. And god forbid you have a family that you are responsible for. The excuse, “I’m abruptly leaving my shift to spend quality time with my child.” might work in a RomCom movie. In real life? You boss will tell you you’re fired, and you and your kids will be homeless on the street. Now you can spend all the quality time you want with them, until Child Protective Services takes them all away.
|
||
All of these vapid, pointless sayings assume a level of privilege that I did when I was 13. They lack any perspective that isn’t the single person, working in IT, with no responsibilities, forging the way to a grander tomorrow. Years of experience have shown me just how far up their own asses some people are in the tech industry. And none quite so far as self-help gurus. Do yourself a favor and take their self-centered, self-help with a massive grain of salt. Don’t share their crap with your friends. Maybe instead try these out:
|
||
If you’re lucky enough to be doing well, help someone who isn’t If you hate your job, ask for help, but don’t quit on a lark Enjoy the time you have in the way that makes you happy, but don’t get fired Okay, rant over. Now I’m going to go do work on a weekend, not because I want to, but because I have to. You know, so I can eat.
|
||
`,summary:"Let me say this upfront. I am going to talk about privilege. I am not an expert on privilege. I am not a poly-sci major with dual minors in social justice and psychology. These are ideas that have been kicking around in my head for a couple weeks, and I thought they might resonate with others or spark conversation. If that’s not the sort of thing you’re interested in, probably best to skip to the next post, where you can read about HashiCorp Vault or Azure or something similar.",date:"6 Jul, 2019",url:"https://nedinthecloud.com/2019/07/06/privilege-and-perspective/",image:"featured-image-nedinthecloud.jpg",readingTime:"8"},"https://nedinthecloud.com/2019/06/04/microsoft-windows-is-dead-long-live-lite-os/":{title:"Microsoft Windows is dead, long live Lite OS",tags:["computex","lite-os","microsoft"],content:`On last week’s Buffer Overflow we were talking about this very strange blog post from Microsoft about a “Modern OS” and what it should include. It’s obvious that Microsoft has something big brewing over in Redmond. In the build-up to \\build, rumors were flying fast and free that Microsoft was going to unveil their new Lite OS and provide a roadmap for development. That did not happen, and the blog post from Computex seems to indicate that while Microsoft is definitely developing something that is not Windows, they aren’t quite ready to share with the world what that “Modern OS” actually is. The more I think about it, I believe that Microsoft might be poised to come barreling back into the mobile market from a totally unexpected direction.
|
||
Let’s review the past few years shall we? Microsoft has increasingly become a services company. That is the entirety of their focus. Every move they make is about increasing customer consumption of their services, including Azure, Office 365, Xbox Live, GitHub and more. The main reason they killed off Windows Phone was the simple fact that there was no financial incentive to compete against iOS and Android. Microsoft wants people who are using iOS and Android to consume their services on those platforms, and therefore their focus turned towards making their services integrate better with those devices. It also means that Microsoft doesn’t especially care about Windows on the consumer side. Sure they make some money from selling Windows licenses to consumers, but the lion’s share of that revenue comes from corporations. I think you’ll see all Windows 10 development focus squarely on enterprise friendly features, while the consumer side becomes more and more anemic. Consumers for their part are flocking to mobile devices and Chromebooks. That’s where we are today when it comes to client operating systems. Microsoft just isn’t that interested in developing Windows 10 for the consumer. And why would they be? Windows 10 has 20 years of cruft tied to its neck, making it a fragile beast at the best of times. Sure, Microsoft has done a lot to stabilize the platform, but given the chance to design a new operating system from the ground up, Microsoft would never build something that functioned like Windows 10. And that is especially true from an internal structure side.
|
||
So let’s assume that Windows 10 development is all about the enterprise customer at this point. Microsoft is left with a hole where their consumer OS could be. Given what I just said about Microsoft just wanting to make services better on iOS and Android, why would Microsoft develop a competitor to those operating systems? I’ve got three reasons:
|
||
Microsoft is trying to drive hardware innovation Controlling the stack leads to better outcomes There is a long game, and Microsoft is playing it Let me expand on these a little bit.
|
||
Microsoft is trying to drive hardware innovation One of the key themes of Computex was the need for better devices and more innovation. Intel came out with Project Athena to try and drive hardware manufactures to develop new and better hardware. Some of the OEMs, like Asus and Lenovo, had some pretty odd designs with different screen types and placement. We’re in a bit of a slump when it comes to consumers purchasing new mobile devices and laptops, and innovation is one way to drive sales and demand. Microsoft has been trying to do a similar thing with its Surface line of devices. In the same vein as Google’s Pixel line of devices, Microsoft is trying to get OEMs to up their game by showing a path forward. It seems to have worked, but now there is a new problem. These innovative devices need an operating system and applications that can take full advantage of unique form factors. Foldable screens, heads-up displays, ancillary screens in strange locations, and even stranger form factors require an operating system that is just as flexible with it’s shell and UI. Is that OS Windows? I’ll pause for a moment while you mop up whatever just shot out of your nose.
|
||
If Microsoft wants to stay relevant on new platforms and form factors, it has to be able to support them. I think that support includes building an operating systems because of point two.
|
||
Controlling the stack leads to better outcomes Microsoft could certainly wait for iOS and Android to get in the game and support all these wacky new devices. Then they can rewrite their applications to try and take advantage of the features exposed on iOS and Android, and hope that things work really well. That’s not a great bet, and certainly not one I would take. Despite their best efforts, many Microsoft products still deliver a subpar experience on iOS and Android and even MacOS. The best experience is usually on a Windows device, because Microsoft controls the whole stack. It’s the same reason that Apple builds the hardware, operating system, and apps for many of its services. Controlling the entire stack leads to a better experience. Better experiences lead to happy consumers that want to use more of your services. And what is Microsoft all about? Ding, ding, ding! Selling services.
|
||
Windows is not going to be the platform to make all this magic happen. Instead, Microsoft has been drastically evolving their approach. For instance, Edge has adopted Chromium for its rendering engine, and that has made it orders of magnitude faster. Most of our applications run in a browser these days, and having a first-class browser that you own and manage should be a priority for Microsoft. So should first class support of WebAssembly to power native applications running in a browser.
|
||
Microsoft is ditching the current Windows-only .Net framework for a cross-platform implementation called .Net Core. The current version of .Net Core is 3.0, and going forward Microsoft is going to drop the Core label and make .Net Core the only .Net Framework for the future. Microsoft has also announced that the next version of Windows will ship with a Linux kernel that can run the Windows Subsystem for Linux v2 in a super-lightweight VM that doesn’t require Hyper-V. The future of applications is containers, and that’s not just for server side deployments. In the Computex blog post they talked about isolating the hardware and operating system from the applications that run on it, in a similar vein to how iOS handles things. Containers can do that handily, and you can package up a Windows legacy app in a container if you need to by using the technology being developed by Droplet Computing.
|
||
Most of the device innovation is going to happen on the consumer side of the house, enterprises tend to move more slowly. But Microsoft needs to get ahead of this trend now, which brings me to point three.
|
||
There is a long game, and Microsoft is playing it When it comes to making money off of an operating system, enterprise money is where it’s at. Microsoft already has Windows 10 as a subscription as part of the Microsoft 365 package. But if you look to the horizon, there is trouble brewing. Chromebooks, tablets, and mobile phones are steadily being adopted by organizations as a primary device. As new form factors and device types are adopted by the consumer space, they will also slowly trickle into the enterprise. While Windows 10 may not die off entirely, a steady decline seems inevitable, and Microsoft needs to head that off at the pass. Hence they are putting money into develop a completely new operating system that meets the current and future needs of consumers and enterprises alike.
|
||
Building a successful and fully feature operating system isn’t going to happen overnight. Adoption of it won’t happen overnight either. But if Microsoft can gain critical mass with a new operating system in the consumer space over the next five years, they will be poised for the adoption of their new operating system into the enterprise.
|
||
I think it’s also clear that we are in the infancy of device form factors. The smartphone as we know it is less than 15 years old. We’ve pretty much maxed out what a slab of flat glass can do for us at this point, and I think that the next killer device will come along in the next five years. As tied as we are to our phones, they still deliver a sub-optimal experience when it comes to interacting with the rest of the world. As technology, hardware, and software progress we are going to reach a tipping point where new and better devices are possible and we have the software to make them viable. Don’t forget, the Newton existed back in the 80s, but the technology and software wasn’t ready yet to support the vision. Windows Phone may forever be dead, but Microsoft is betting that the next big device (whether it’s an AR headset, overlay contacts, or a neural implant) isn’t all that far away, and this time they don’t intend to miss it.
|
||
`,summary:"On last week’s Buffer Overflow we were talking about this very strange blog post from Microsoft about a “Modern OS” and what it should include. It’s obvious that Microsoft has something big brewing over in Redmond. In the build-up to \\build, rumors were flying fast and free that Microsoft was going to unveil their new Lite OS and provide a roadmap for development. That did not happen, and the blog post from Computex seems to indicate that while Microsoft is definitely developing something that is not Windows, they aren’t quite ready to share with the world what that “Modern OS” actually is.",date:"4 Jun, 2019",url:"https://nedinthecloud.com/2019/06/04/microsoft-windows-is-dead-long-live-lite-os/",image:"featured-image-nedinthecloud.jpg",readingTime:"8"},"https://nedinthecloud.com/2019/05/17/going-cloud-native-with-pure-storage-purity/":{title:"Going Cloud Native with Pure Storage Purity",tags:["cfd5","cloud-field-day","pure-storage"],content:`As a Cloud Field Day 5 delegate, I attended a presentation from Pure Storage regarding how their products are embracing the cloud. Whenever a storage vendor starts talking about embracing the cloud, I start to get a bit wary. In general, what they actually mean is they have created a virtualized version of their array, running on IaaS, using the public cloud. Or it could mean that they now have the ability to send data up to the public cloud from a local datacenter. Is that what Pure came to the table with? Yes. At least that was part of the presentation, but they actually had additional product details that I felt moved the needle and showed some real innovation.
|
||
Product Vision The initial product direction and vision for the company had some interesting points. The central theme was the divide between what storage looks like in a traditional datacenter versus the cloud. When you think of cloud storage, you probably think of a few key factors:
|
||
Pay as you consume pricing Elastic capacity and performance API driven interaction Traditional storage is very much the opposite of how cloud approaches these three categories. You almost always pay up front for capacity, with the exception of some larger programs from storage providers where you pay as you go (PAYG). But even those programs are limited in the amount of capacity you can consume in the PAYG model. Traditional storage is also limited in capacity and performance. A storage array has whatever capacity it has. You can expand that capacity up to a point, but eventually you will run out of room on the frame and have to buy another frame. That second frame is not part of the same storage service, unless you have invested in something like a VPLEX, but even then there are limits. And you can’t shrink down when you don’t need the storage anymore. Releasing capacity back into a storage pool is fraught with peril, and even if you successfully accomplish this feat you can’t exactly return the disks to the vendor for a refund. Lastly, the interaction with traditional frames is usually through some command line tool, or a janky UI that uses some outdated and insecure version of Java that only works on one workstation in the company, and you all RDP into that workstation to make changes to the frame. Not that I have any experience with something like that…
|
||
ANYHOW.
|
||
Pure’s main point was that traditional storage needs to take its cues from cloud storage and find a way to embrace those three principles, as well as support object-based storage, instead of just file and block.
|
||
There were two products showcased by Pure during the demos.
|
||
Cloud Snap Remember that thing I said about storage vendors sending data up to the cloud? Yeah, that’s basically what Cloud Snap is. You can watch the whole presentation if you like, but I can distill it down to the essential components. Cloud Snap replicates a snapshot from a Pure array to either NFS or AWS S3. After the first snapshot, each update only sends the incremental differences, which makes the storage consumption and network consumption fairly efficient. The snapshots include the metadata about the snapshot, so the snapshot can be recovered to any Pure array and not just the original array. Cloud Snap doesn’t use a virtual appliance in the public cloud destination, it manages the storage from the array itself. The presenter talked about using the replicated storage in tandem with backup software to efficiently move data to a cloud target. While that seems like a potential application, most of the backup providers are already using their own deduplication and compression algorithms to efficiently move data between targets. Instead, this seems more like a way to reduce space consumption on the array and provide offsite data protection for a crash consistent DR recovery.
|
||
Not a terribly innovative solution overall, but it seems to be table stakes for storage vendors these days.
|
||
Cloud Block Store Remember that thing I said about running a virtual frame on IaaS in the public cloud? For starters, that’s not a great idea. You’re going to end up spending a lot of money on the IaaS, and in all likelihood the virtual version of the frame has the same limitations that the original frame had. It probably makes more sense to replicate data using some other third party tool. The Cloud Block Store is a straight port of their on-premises frame - more on this later. It’s using a combination of S3 and EC2, and each instance is deployed in a single availability zone. You can set up a synchronous cluster of instances across AZs, but you cannot stretch a single instance across more than one AZ.
|
||
I thought that was a strange choice, and I asked about the decision to have the instance run in a single AZ instead of multiple AZs. The presenter talked about the mental model of what they were trying to create, which I took to mean that they are focused on building something that looks like a traditional datacenter. In the traditional world, this would be the equivalent of creating a metro-cluster with the VMAX. But I would remind you that this is a cloud product, and apparently it is using S3, which is a multi-AZ service. There is no reason to mimic the traditional datacenter approach when it brings no tangible benefit. It seemed clear from the remainder of the conversation that the team from Pure knows this, and once Cloud Block Store is a out of beta and a few revs into its production lifecycle, the instance will be multi-AZ capable.
|
||
The front-end of the Cloud Block Store uses the same APIs as a Pure Storage array, so any application that works with the on-premises array will work with the Cloud Block Store. The presenter showed an example where both versions were using the Pure Service Orchestrator to provision storage on the fly for a Kubernetes cluster. It was an interesting concept, but K8s already has native cloud integrations for storage. I don’t necessarily know why I would want to shove the Cloud Block Store into the mix. I suppose you could be replicating data from an on-premises array to Cloud Block Store and running a application on K8s that is running in both environments and consuming the replicated data. Maybe for a DR site? I guess what I’m saying is that it’s possible to do. I just don’t know if this is the killer use case for most people.
|
||
I like that Pure is at least thinking of ways to integrate with modern development methods. It’s far and away better than some storage vendors out there who can’t really think outside the LUN. Does that make them the Taco Bell of storage vendors? Maybe, and I mean that as a compliment.
|
||
Towards the end of the presentation, Pure mentioned that they had rewritten the Purity software that runs their array for Cloud Block Store. I want to emphasize that. They rewrote the Purity software to take advantage of the native constructs within AWS. This is not just a straight port of their array software running on a virtual machine and backed by traditional data disks. The API front-end is the same as their on-premises product, but the backend is something completely different. That’s HARD. And when companies do something hard, it shows a level of commitment to this cloud thing that I don’t always see with traditional vendors.
|
||
Conclusion Pure Storage is certainly doing some interesting things with their cloud offerings. They appear to have the basics covered, and are doing their best to bring some real innovation to their cloud products. If you are already all in with Pure on-premises, then expanding out to the cloud with the same management tools and APIs might make a lot of sense. There could be the fear of vendor lock-in, but I take those fears with a grain of salt. As soon as you pick any technology, you are “locked in.” I’m excited to see where Pure goes from here, especially once Cloud Block Store exits the beta stage and customers start putting Production workloads on it. Once a product is in customers hands, they often think of use cases that the product team never considered.
|
||
`,summary:"As a Cloud Field Day 5 delegate, I attended a presentation from Pure Storage regarding how their products are embracing the cloud. Whenever a storage vendor starts talking about embracing the cloud, I start to get a bit wary. In general, what they actually mean is they have created a virtualized version of their array, running on IaaS, using the public cloud. Or it could mean that they now have the ability to send data up to the public cloud from a local datacenter.",date:"17 May, 2019",url:"https://nedinthecloud.com/2019/05/17/going-cloud-native-with-pure-storage-purity/",image:"analysis.png",readingTime:"7"},"https://nedinthecloud.com/2019/05/08/rubriks-build-is-all-about-education/":{title:"Rubrik's Build is all about education",tags:["cfd5","cloud-field-day","rubrik"],content:`I was a delegate for Cloud Field Day 5 back in April. One of the companies presenting was Rubrik. They also presented at CFD 3, where I was a delegate for the first time. Their presentation for CFD 3 knocked it out of the park! Watching Chris Wahl and Rebecca Fitzhugh present was a delight. Needless to say I came into the CFD 5 session with high hopes. The results this time around were a bit mixed, and it was only after reflection that I understood what was going on.
|
||
At CFD 5, Rubrik had a very different presentation than what I was expecting. The introductory speaker, Mike Nelson, did a great job of summarizing things and introducing the rest of the presentation. He has also been a guest on the Day Two Cloud podcast, and I highly recommend checking that episode out. In the presentation he also referenced a tweet I made, which really encapsulates the reason I quit my job and started my own company.
|
||
You know what? It's New Music Friday. I've got my soft pretzel. My wine of the month box came. It's going to be OK.
|
||
— Ned Bellavance [MVP] (@Ned1313) March 29, 2019 After Mike finished, David Terei talked about the reference architecture for Polaris and Polaris Radar. At first it was very esoteric, and he decided to use a whiteboard app on his tablet. I think the information would have been better communicated with pre-drawn diagrams. Whiteboarding is great for a free flowing exchange of ideas, but less suited to a presentation about a pre-existing architecture. After he got through the initial explanation, we got down into the nitty-gritty details of how Polaris is deployed and consumes information. The Polaris service is running in its own Azure tenant, but it is talking to a deployment in your Azure account. And that deployment is using Azure Kubernetes Service, Azure Key Vault, and Storage Accounts. We asked a lot of questions about how they are keeping customer data safe and complying with security regulations while training their models. I encourage you to watch the whole video.
|
||
Then the conversation shifted into the Polaris Radar product and how they can use it to prevent ransomware attacks. Well, not prevent actually, but rather it could be a way to recover from a ransomware attack and detect that it is happening by analyzing the change rate and patterns of files that are being backed up. Which, I guess that’s cool? I don’t know. They kept harping on the use case of detecting ransomware attacks, but I think that’s a red herring. Rubrik is grabbing ALL of the data and metadata and throwing it at a massive training model for machine learning. Yes, you can detect some interesting anomalies regarding malware. But there’s got to be more interesting applications for their Polaris solution. I asked as much and got a knowing smile and a wink from the Rubrik team. They’re brewing something bigger and more impressive here, and they are just not ready to talk about it yet. Keep an eye on them and their product announcements though. Data is the new oil, and oil needs to be refined. Who would have thought that the same chemical cocktail could be refined to give us fuel for our vehicles, plastic for playground equipment, and fabric for our clothing? To stretch out that analogy, who knows what backup data could be refined into for use at your company?
|
||
The last presentation was from Matt Elliott, and here’s where things got strange. So far the whole presentation had been about the Rubrik product line-up, the nuts and bolts behind it, and a little on the future roadmap. Not terribly exciting - to be honest - but also not unexpected. With a lead-in from David Terei, they started going into the open-source credentials of Rubrik and the Build project. Then Matt Elliott took over the presentation and started talking about his Unix/Linux bonafides and long history in open-source. Matt seemed a little nervous, which I don’t blame him. Presenting in front of an audience is nerve-racking, and he knows better than many how difficult the Tech Field Day crew can be.
|
||
The core of his talk was around the Build project at Rubrik, and honestly I didn’t get it. Yeah, open source is pretty cool. You’ve got this weird little project on GitHub that teaches people how to interact with your software. How is this cloud? Why is this in the presentation? There are clearly some very cool things coming down the pike that Rubrik is not ready to talk about. Are they just bringing up the Build project because they needed to fill 45 minutes? I really did not get it.
|
||
Until I did about four hours later at the happy hour.
|
||
A recurring theme of all my CFD conversations was the need to educate IT Ops people about new ways to approach their job. The cloud brought with it the introduction of APIs for everyone. What was once the province of the developer now becomes something that all IT practitioners can take advantage of. The fact that some cloud services - Rubrik especially - have an API-first philosophy means that IT Ops have the opportunity to create new and incredibly useful improvements to their existing workflows. The drudgery of manually entering data across multiple systems because they “don’t talk”, becomes a thing of the past. The need to write convoluted, imperative scripts that screen-scrape and fail to adapt to interface changes also goes away. An IT Ops person has the chance to make their job better and more satisfying. At the same time I would argue that refusing to change will negatively impact their employment prospects in the next five years, even more so in the next ten. One is the carrot and the other is the stick. This seems obvious to those of us who live in the cloud world. We take these things for granted. But then you go and talk to a sysadmin who has been banging out manual labor in their datacenter for the last fifteen years and it occurs to you, they don’t even know to stop and look around.
|
||
This is essentially the conversation I started having with Chris Wahl at the happy hour that evening. Chris then started talking about how they are doing these small, in-person educational seminars in a bunch of cities across the country. The seminars aren’t sales pitches for Rubrik. They are an opportunity to connect with IT Pros who have never touched an API, never run a git commit, never even written a single line of code. Rubrik and the Build program are reaching out to those people and giving them a glimpse of another way. A better way. There’s only so much you can do in a single day, but by giving these IT practitioners a little taste of success, by giving them a tiny preview of the vast world beyond, by giving them a path forward, you are educating. And education is the thing that seems most lacking at the moment. Not education about the latest product features, or marketing hype, or vendor specific jargon. This isn’t self-serving education to help IT Pros master your UI or CLI. It’s about educating IT Pros to make them better at their job, and that is reward enough.
|
||
I’m not so blind as to ignore the fact that the Build project is focused on teaching people to be better with Rubrik APIs and learning to automate Rubrik related software. The fundamentals are still there, and it is hard to teach in a vacuum. There’s a reason why Microsoft has invented several fictional companies for their training. There’s a reason why every course I do for Pluralsight has a real-world scenario behind it. Learning to work on a real-world API and create real projects in GitHub is super important. I’m glad that Rubrik is doing this. I’m also glad that companies like Juniper have the NRE labs to help network engineers on their journey. In an industry that experiences as much change as ours, education is paramount. I’m glad that some vendors see it that way too.
|
||
When Matt Elliott got up to talk about his personal journey through IT, he was telling it as one possible path among many. That path led him to Rubrik and to the Build project. I didn’t get it at first, but I do now. Thank you to Matt, Chris, and the rest of the team at Rubrik for your efforts!
|
||
`,summary:"I was a delegate for Cloud Field Day 5 back in April. One of the companies presenting was Rubrik. They also presented at CFD 3, where I was a delegate for the first time. Their presentation for CFD 3 knocked it out of the park! Watching Chris Wahl and Rebecca Fitzhugh present was a delight. Needless to say I came into the CFD 5 session with high hopes. The results this time around were a bit mixed, and it was only after reflection that I understood what was going on.",date:"8 May, 2019",url:"https://nedinthecloud.com/2019/05/08/rubriks-build-is-all-about-education/",image:"analysis.png",readingTime:"7"},"https://nedinthecloud.com/2019/04/28/time-for-a-new-adventure/":{title:"Time for a new adventure",tags:[],content:`This post is about my recent decision to leave my current employer for a new opportunity to launch Ned in the Cloud LLC. I want to be very clear up front that I hold no ill will towards Anexinet or anyone who works there. Over the past six and a half years I’ve been provided with tremendous opportunity to grow and expand, and try out new things.
|
||
My search for a new job started about eight months ago. I told my parents that I was looking for a new job, and my Mom kind of freaked out. “Why would you leave a perfectly good job?”, was her main question. And it’s a fair one. She also worked at the same school for her entire career as a librarian, so she’s got some preconceived notions on the topic. Still though, it’s a fair question on why I would leave a well-paying job, at a stable company, that treats me well. There are reasons! And here they are:
|
||
I accepted a position two years ago that I do not really like. I am kind of done with IT consulting. I am ready for a new company. Let me expand on these a bit.
|
||
The Position About two years ago I was offered a position running the Cloud Solutions area. While the organization has been heavily involved in cloud, it was time to double down and create a dedicated role that would focus on expanding our cloud offerings and bringing them to market. The position sounded interesting, and I was done with going to clients and delivering solutions. More on that in the next section. I thought I understood what the position entailed. I thought it was something I really wanted to do. I was only about 50% correct.
|
||
See, I thought that the position would focus heavily on developing new solutions and working with cutting edge technologies. But I didn’t realize that it would also include training sales people, going on sales calls, training pre-sales, developing content, speaking at events, and more. There was a lot more to this job than just the product development side of things. Some of those things were actually a lot of fun. It turns out that I really like public speaking and creating new content for the marketing group. Some of those things were not a lot of fun.
|
||
Sales is hard. And I am not a sales person. There’s a certain personality that not only excels at selling, but also actually enjoys it. That is not my personality type. I will say that the entire experience made me appreciate the grind that is being a sales person. I’ve been on a dozen or so cold calls, and they are awful. There’s a dance to the process, and a certain level of palaver that is required. I felt like I came to a dance that I didn’t want to go to, wearing the wrong clothes, and not knowing the dance steps. That’s not anyone’s fault but mine. I could read-up on how to be better at sales calls and try to find a mentor. I could. But I won’t, because I hate it. Not in the way that you hate coffee at first, but it grows on you. More like the way that no matter how many times you get nauseous, you never grow to like it.
|
||
Consulting When I decided to get into consulting back in 2012, I thought to myself, “Let’s do this for five years, and then if you’re done, you will be able to write your own ticket to whatever you want to do next.” I knew that consulting would be challenging in many ways. For starters, I knew that I was going to be constantly challenged to learn new technologies and implement them. I also knew I was going to be shoved out of my comfort zone when it came to traveling and social interaction. I do not travel well. Traveling and going to unfamiliar places ties my stomach in knots and gives me all kinds of anxiety. Or at least it did back in 2012. The idea of going to a different and often unknown place on a regular basis seemed terrifying. In addition, I was going to be meeting new people and having to prove myself over and over again. Interacting with strangers is also a source of anxiety for me, and now I was going to be in a position where I was going to an unfamiliar place with unfamiliar people all the time. Seven years later, I can’t say that either of those things are easy for me, but they are much easier than they once were. The prospect of traveling up to NYC to meet with new clients and talk to them about AWS for two and a half hours doesn’t fill me with deep foreboding. That’s just Wednesday. Consulting taught me to think on my feet, form strong opinions that are loosely held, and enter an unfamiliar setting with confidence.
|
||
Those are the good things about consulting, but there are some down sides as well. Consulting is stressful. You have to balance the client’s needs with what is possible. You’re often thrown into new technologies and asked to be the expert in front of clients. Sometimes you are literally one chapter ahead of what you’re telling them. And the better you get at doing this, the more often you’ll be asked to do it. New project, with new technology that no one has experience with in the company? Give it to Ned, he’ll figure it out. Eventually, you get tired of that sort of thing. Well, maybe you don’t. But I did. There’s a lot of other unpleasantness with consulting, like the rat-race of certification, keeping yourself billable to earn your bonus, and trying to deliver the solution promised by presales, even though the product can’t really do it. If you’ve spent time in consulting, you know exactly what I mean. If you haven’t, trust me, this stuff happens across the entire industry. If you’re looking for more context, this episode of Datanauts does a great job.
|
||
This is why I decided to take a new position where I wasn’t delivering projects anymore. But I was still working at a consulting company, and to be frank, I was kind of done with consulting.
|
||
New Company At first, I thought I wanted to go work for a vendor. I had been creating courses for Pluralsight, and I thought that I might like doing that as a full-time job for a vendor. In July of 2018, I started looking for vendors interested in someone to create training or educational content. There were a few close opportunities. I got to advanced stages of the interviewing process with two open-source companies. One had a travel requirement I couldn’t meet, and the other found someone with more experience than me in the world of learning development. On the hardware vendor side, I came very close to getting a position evangelizing Azure Stack. Again, the primary issue was the amount of travel involved. Ultimately, none of these opportunities panned out, and I was feeling disheartened. And then I had a conversation with Stephen Foskett, and my attitude changed.
|
||
When you find yourself in a dilemma, there are few things more important than a friendly sounding board. By September, I was feeling despondent about my job search after going through eleven video interviews for two positions at the same company. Yes, you read that correctly. Eleven video interviews with different people throughout the company, and then they went dark for three weeks. The interview process had already stretched over two months, and I thought I was definitely going to get the job for one of the two positions. At the beginning of September, I finally heard back from their HR person that I didn’t get one position because they decided it had to be in California and the other had been filled by someone with more experience than me. That’s fine, but I wish they had communicated what was going, rather than me reaching out to their HR person every week for an update. All the feedback I had received up until then had been extremely positive. In my mind, I had already built up this idea of getting the job, quitting my current one, working from home 100%, etc. I was making mental plans for all the different ways my life was going to be better. To be completely honest, I actually started this blog post back in late August of 2018. That’s how confident I felt about getting the job! When I didn’t get the job, it popped the balloon on my inflated dreams. Almost the same day I got rejected, I received a message from Stephen Foskett about going to a Tech Field Day event. I couldn’t go due to other obligations, but I mentioned that I was looking for a new job. That led to a phone call, where Stephen was kind enough to let me vent, and act as a sounding board for my thought process. He offered advice when asked, and really helped me think about my options and what was available out there. By the end of the conversation, I was seriously starting to consider going independent and doing freelance work.
|
||
The idea of quitting a stable job and freelancing on my own was daunting. I’ve worked for a company, collecting an hourly wage or a salary since 1997. That’s 22 years of not having to worry about where my next paycheck was coming from. As long as I did my work in a competent way, and didn’t do anything too stupid, I would receive pay for that work every two weeks. By going independent, I was no longer going to have that steady, fortnightly deposit in my bank account. Work would not be assigned to me. I would have to go find new work. It took another few weeks to convince myself I could do it. I also had to make sure my wife was on board. This was going to impact our financial future, and I needed to make sure she was comfortable with my idea. To my surprise, she didn’t even blink, instead she went straight into planning mode. She knew I wasn’t happy at my job, and the idea of me making my own schedule meant that I would be able to spend more time with the family. Happy Ned and more family time? That sounds like a win-win.
|
||
Ned in the Cloud LLC With the support of my wife, family, and friends, I went forward with my plan. I formally registered Ned in the Cloud LLC, and got a Tax ID. Starting May 6th, this will be my one and only job. In preparation for the transition, I have started a new podcast through Packet Pushers. I’ve agreed to write a book on the Azure Kubernetes Service with some other awesome people. And I’ve stepped up my course production with Pluralsight. Soon I will be launching a new version of the website, with a listing of services and probably a better look and feel. Graphic design is NOT my strong suit, and I think the simplicity of this site demonstrates that.
|
||
In terms of services, I am available for webinars, speaking engagements, writing content, podcasting, and more! If you’re interested, please reach out to me. I will have a contact form on the site soon, but until then you can reach me ned-at-nedinthecloud.com, through Twitter, or LinkedIn. I am incredibly excited with this new venture and can’t wait to see where the journey leads me next!
|
||
`,summary:`This post is about my recent decision to leave my current employer for a new opportunity to launch Ned in the Cloud LLC. I want to be very clear up front that I hold no ill will towards Anexinet or anyone who works there. Over the past six and a half years I’ve been provided with tremendous opportunity to grow and expand, and try out new things.
|
||
My search for a new job started about eight months ago.`,date:"28 Apr, 2019",url:"https://nedinthecloud.com/2019/04/28/time-for-a-new-adventure/",image:"featured-image-nedinthecloud.jpg",readingTime:"10"},"https://nedinthecloud.com/2019/04/24/sysdig-monitoring-via-ebpf/":{title:"Sysdig - Monitoring via eBPF",tags:["cfd5","cloud-field-day","sysdig"],content:`A few weeks ago I attended a presentation by Sysdig as part of Cloud Field Day 5. Prior to attending CFD5 I did a little research about the company and their products and wrote up a quick post where I posed a few questions. I think some of those questions were answered by the presenters. The questions were:
|
||
Do you currently support or plan to support container deployments using AWS Fargate or Azure Container Instances? Is Sysdig a marketplace item in AWS or Azure today to simplify deployment? How are you handling balancing open-source and paid products? Are there plans to open-source the whole solution like Chef just did? What are you doing with all the aggregated monitoring data you are getting from clients? What are the major security concerns with your solution and how are you addressing them? You can watch all three videos on YouTube. Below are a few thoughts and a bit of a deep dive on how Sysdig is capturing and storing their information.
|
||
Sysdig was founded by Loris Degioanni, who co-authored Wireshark. As someone who has used Wireshark on more than one occasion, I know the value that tool has to a systems administrator. There was a time when I needed to troubleshoot an Active Directory Cross Forest Trust authentication issue between two forests with Selective Authentication enabled. The only way I ever figured out the problem was by firing up Wireshark and looking at the packet captures coming out of the domain controllers on both sides and the client computer. By analyzing the authentication handshakes between multiple clients, we were able to track down a DNS issue, a FSMO role issue, and a poorly configured DC in a child domain. That’s just one instance where Wireshark helped me crack a particularly tricky situation. That’s the old world though, the legacy world of VMs and physical machines where the packet was king and packet capture put you in the center of everything. The reason Loris founded Sysdig, and the guiding principle behind their solution, is that packet capture on the wire is dead. It just doesn’t work for a cloud-native world with containers being spawned and thrown away constantly. The solution was to find a way to capture all traffic information from containers, look for anomalous behavior, and send that information up to an aggregation point for further analysis and alerting.
|
||
Sysdig is doing this through a light-weight container that lives on each host and has access to eBPF (extended Berkeley Packet Filter) running on the kernel of the host. If you aren’t familiar with eBPF, don’t worry! Neither was I. Here’s a really good article introducing the eBPF and what it can do. At a basic level, eBPF gets attached to a code path in the kernel and it allows verified programs to interact with particular interfaces through the bpf() function. If you want to filter, monitor, and classify network traffic in a performant way, then eBPF is your friend. The Sysdig container running on each host is basically sniffing the traffic of all other containers running on that hosts. I don’t know the exact details, but I assume when a new container is spun up, the Sysdig program is attached to its code path. Sysdig can see not only network traffic, but also IPC calls. The information is processed locally by the Sysdig agent, aggregated and sent up to the Sysdig analysis platform for additional processing. There is also a circular buffer of about 100MB that captures every packet and stores it to see if it is anomalous or interesting behavior. If it is, the full packet capture is uploaded so that you can do a playback of exactly what was happening at that time. Even if the container that caused it has ceased to exist, you can still replay the entire packet flow and find out what was going on. That is pretty freakin’ awesome for a developer trying to debug code or for a security person trying to figure out how a breach occurred.
|
||
There is still a question of security here, in fact I asked Loris that question directly! You are instilling a lot of trust in that Sysdig container. It gets to inspect all the traffic on your other containers, which makes it a goldmine of information for a potential attacker. In that regard, I would hope that Sysdig is constantly patching and improving their solution to keep their container images secure. Loris made the point that otherwise you would need to instrument every container individually or run a sidecar with every container. A vulnerability in that situation would be much harder to mitigate. If a vulnerability is discovered with Sysdig, you can stop the containers on all the hosts until a patch is available. In the meantime, the rest of your workloads continue to function uninterrupted. The aggregated data being uploaded to their service is being stored per client, but they did say that they anonymize the data and use it for model training. You can opt out of that, or you can choose to run the Sysdig software locally and not use their cloud service at all.
|
||
I did not get to ask them about Fargate or ACI. My guess is that neither is currently supported because of the way that Sysdig works. It’s using the eBPF to collect data, so it needs a container running on every host in the environment. If you’re running regular K8s or something like AKS, that’s fine. Each worker node is dedicated to you, and you can run that additional container on each host. Something like Fargate or ACI does not allocate a whole host to you, and there is no way that Microsoft or AWS is going to let you collect the network data on all containers running on each host you use in ACI or Fargate. Talk about a security nightmare! There are a few other ways to support it.
|
||
Embed the Sysdig agent in any container image you deploy to ACI or Fargate Run a sidecar proxy with each container image you deploy to ACI or Fargate Use native tooling in Azure or AWS to stream telemetry to Sysdig I did ask about converting from Open Core to fully Open Source in the same way that Chef did. For the time being, Sysdig is going to continue down the path with Open Core. I did not get to ask them about publishing something to the Azure or AWS marketplace, but having actually run a demo myself I can tell you that it is very easy to deploy.
|
||
While I think that the death knell of packet capture is a big premature, it helped Sysdig tell a compelling story about the unique challenges that cloud-native applications pose for monitoring, performance, and security. Their tool approaches solving that problem in a unique and sophisticated way.
|
||
`,summary:`A few weeks ago I attended a presentation by Sysdig as part of Cloud Field Day 5. Prior to attending CFD5 I did a little research about the company and their products and wrote up a quick post where I posed a few questions. I think some of those questions were answered by the presenters. The questions were:
|
||
Do you currently support or plan to support container deployments using AWS Fargate or Azure Container Instances?`,date:"24 Apr, 2019",url:"https://nedinthecloud.com/2019/04/24/sysdig-monitoring-via-ebpf/",image:"analysis.png",readingTime:"6"},"https://nedinthecloud.com/2019/04/20/day-two-cloud-podcast-update/":{title:"Day Two Cloud Podcast update",tags:["day-two-cloud","packet-pushers"],content:`I am super excited to announce that the Day Two Cloud podcast has officially graduated from the Packet Pushers Community Channel to its own dedicated channel and feed! That means you can directly subscribe to the podcast and listen to informative and educational episodes about operating a cloud on day two and beyond. The guests so far have been fantastic and I already have more episodes queued up for your enjoyment.
|
||
You know when I started a music podcast in 2006, I never thought it would go anywhere. And I was absolutely right! I recorded about 13 episodes and then pod-faded into that good night. Then in 2016 I tried my hand at this podcast thing again, except this time it was a technical podcast for my employer. There have been some fits and starts with that one, but it’s got 41 episodes published. Since I didn’t have enough to do at work - heavy sarcasm here - I decided to start a second podcast about tech news on a fortnightly schedule. That podcast became Buffer Overflow, shifted to a weekly schedule, and now is over 100 episodes.
|
||
When Ethan from Packet Pushers initially suggested doing a podcast for them, I was surprised to say the least. If I had to guess, I started listening to Packet Pushers podcast - mostly the main feed and Datanauts for about three and a half years. While I had a solid background in most things infrastructure, I hadn’t really ventured into the networking space too much. I had my CCNA, and I sat the first CCNP exam, and I failed. Let’s say I knew enough about networking to be dangerous and talk coherently with the networking guys. One of those networking guys at work knew I listened to podcasts, and suggest I might like the Packet Pushers stuff. That was the beginning of it all. Over the next few years I had both Greg Ferro and Chris Wahl on the AnexiPod to talk about networking and Rubrik respectively. I also was a guest on Datanauts twice, once for Azure Stack and again for Terraform. I began to interact with their team on Twitter, and build a bit of rapport. So when Ethan did reach out about doing a podcast, it wasn’t totally out of left field, but it was a surprise and a great opportunity.
|
||
I jumped at the chance! As someone who was already running a couple podcasts, I knew I could do it. The real struggle is getting great guests and earning a great audience. And I do mean earning! Building and maintaining an audience in the world of podcasts is hard, especially with the medium experiencing unprecedented growth. The Packet Pushers have worked tirelessly to earn an audience of loyal listeners, and I was being presented a chance to leverage some of that goodwill for my podcast. But like I said, I have to earn the loyalty of my audience. You might discover Day Two Cloud through Packet Pushers, but will you subscribe and listen? That part is on me. Now that I am out of the community channel and on my own, it’s time to really ramp things up. And that, dear potential listener, is where you come in.
|
||
Do you have a topic you’d like to talk about on the show? Are there guests you would suggest? Do you have some feedback on the format of the show? I want to hear from you! Please hit me up on Twitter, my DMs are open. Regardless, I hope that you will give the podcast a chance, and that you find it helpful and informative.
|
||
`,summary:"I am super excited to announce that the Day Two Cloud podcast has officially graduated from the Packet Pushers Community Channel to its own dedicated channel and feed! That means you can directly subscribe to the podcast and listen to informative and educational episodes about operating a cloud on day two and beyond. The guests so far have been fantastic and I already have more episodes queued up for your enjoyment.",date:"20 Apr, 2019",url:"https://nedinthecloud.com/2019/04/20/day-two-cloud-podcast-update/",image:"graduation-995042_1920.jpg",readingTime:"3"},"https://nedinthecloud.com/2019/04/07/cloud-field-day-vmware/":{title:"Cloud Field Day – VMware",tags:["aws","cfd5","cloud-field-day","vmware"],content:`I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up. As a delegate, I am meant to represent the larger IT community, so I want to know what you think! In this post I am going to consider VMware and what they’re doing with AWS.
|
||
Almost two years ago I wrote a post called “VMware on AWS - You’re doing it wrong.” If you don’t want to read the whole thing, then I would boil the central thesis down to this:
|
||
If you’re going to make a move to AWS, you’re better off learning and adopting their native service than running a more expensive service that is less flexible.
|
||
I stand behind that statement. In an ideal world, IT Ops would adopt the new tools and services in AWS and developers would migrate their applications to AWS by leveraging cloud-native services. When I wrote the post, I was deep in the midst of coming up with new solutions for my company to offer. I had read all about 12-Factor Apps and really believed that refusing to adopt the cloud was patently ridiculous. When you first learn about something, it can take a while to synthesize that information into a cogent philosophy. And there’s also a tendency to try and apply your new knowledge to every solution. Here I was, diving into the awesomeness of cloud, becoming an evangelist, and not taking a hard look at the actual world we live in.
|
||
Since that time, my views have mellowed. Part of that is due to the failure of the first few product offerings I came up with. Here I was, coming up with these fantastic, strategic solutions that would empower my clients to move forward and advanced their cloud adoption process. It was manifestly obvious why they should engage with my group immediately. Our solution was just that good. And then… we consistently failed to sell that solution. In a vacuum, on paper, in a perfect world, my idea made total sense. Once that idea had to interact with the real world things rapidly deteriorated.
|
||
Back to VMware on AWS (VMC). In an ideal world, this solution would not exist. It doesn’t make sense from a financial perspective. It’s really no different than a hosted colo. Your team is going to have to work with AWS anyway. There are so many reasons that VMC didn’t make sense. Yet 18 months later, it is still chugging along. Growing in fact. So what gives?
|
||
VMC is not an ideal solution, but it is a viable solution for many organizations. Because those organizations don’t live in an ideal world. In the real world you have a procurement department that is used to dealing with VMware as an approved vendor. Updating your enterprise agreement to include VMC could be an easier solution than trying to get a new vendor added to the approved list. It’s not efficient or elegant, or even cost effective, but it IS politically expedient. And that will trump all the rest of those reasons if you’re working for a large corporation where politics are the motivating factor for most of middle management, instead of saving money or cutting costs.
|
||
Then there’s the training issue. Your IT Ops folks have been honing their skills on VMware for the last decade. In an ideal world they would rapidly learn how to manage AWS resources. But in the real world, IT Ops folks are overworked and might balk at the idea of learning yet another technology when there is an alternative that allows them to keep using the skills and processes that already exists. There’s also an added benefit here, since VMC is a managed solution, the admin burden for IT Ops might lessen post-adoption. Thus, those overworked IT Ops folks might actually have time to learn about AWS.
|
||
Migration is also a potential friction point. In an ideal world, if you had to do a lift-and-shift migration, you would use something like Cloudamize to migrate your VM workloads to EC2. Lift-and-shift by itself is already a sub-optimal approach, but it also can be a necessary step in your cloud journey. In the real world, moving VMs to EC2 can be tricky. There are unsupported configurations (clusters come to mind), VMs using FT, and a need to re-IP Address the workloads moving. Seriously, I have seem migrations grind to a halt because some application’s IP addressing could not change. That’s reality. And VMC removes a lot of those concerns by enabling a straight-up VMotion to the cloud without changing network addressing.
|
||
I guess what I am saying is that when it comes to using VMware on AWS, you’re still doing it wrong, and that might be alright.
|
||
For Cloud Field Day, most of the questions I have for VMware revolve around their plans to support hybrid cloud workloads. Here’s a few things I want to know:
|
||
What is the roadmap for NSX-V and NSX-T? Are they going to be merged? What’s going on with VMware on AWS Outposts? How will this integrate with VMC, the AWS Console, and existing on-premises VMware installations? What is VMware’s container strategy going forward? It’s been a little spray-and-pray up to this point. How is VMC going to continue to integrate AWS services? Whatever happened to the rumors about VMware on Azure? Do you questions for VMware you’d like me to ask? LMK!
|
||
`,summary:"I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up.",date:"7 Apr, 2019",url:"https://nedinthecloud.com/2019/04/07/cloud-field-day-vmware/",image:"analysis.png",readingTime:"5"},"https://nedinthecloud.com/2019/04/03/cloud-field-day-sysdig/":{title:"Cloud Field Day – Sysdig",tags:["cfd5","cloud-field-day","sysdig"],content:`I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up. As a delegate, I am meant to represent the larger IT community, so I want to know what you think! In this post I am going to consider Sysdig and what they have to offer in the cloud world.
|
||
Much like my previous Datrium post, I have heard the name Sysdig before and even seen them at a conference, but I’d be hard-pressed to tell you exactly what they do. Based on the name, I would assume that they specialize in data and analytics, or logging of some kind. The name is evocative of syslog, that venerable logging solution, and digging into it in some way. Fortunately, their website makes it really clear what they do. The main page probably says it best:
|
||
Sysdig is the first unified approach to monitor and secure containers across the entire software lifecycle.
|
||
Okay. I get it. They monitor and secure containers. I’m sure there’s a lot more to it than that. But I want to pause. Take a moment. And really appreciate the fact that they were able to sum up what they do in a single sentence. As someone who is exposed to enough marketing slicks to choke a wildebeest, it is refreshing to have something so concise and direct.
|
||
[Brief period of reflection over, we now return to the regularly scheduled blog post]
|
||
Cool. Now how do they accomplish this modern miracle of monitoring acumen? From what I can tell on their site, and here’s where things get a little hazy, it looks like they run a container on each host in your container cluster. The assumption is that you are probably already using an orchestrator, like Kubernetes, and you’ll use that same orchestrator to deploy the sysdig agent on each host. There is also a manual install process if you aren’t using an orchestrator. There’s also a support group out there for you while you come to grips with the choices you’ve made in life that left you without an orchestrator to lean on.
|
||
Once you have your agents provisioned and configured, they will start reporting information back to Sysdig. Now I’m not sure if Sysdig is a SaaS platform, or if each agent is part of a larger Sysdig deployment that collects and analyzes the information from your containers and services. The architecture here is not exactly clear. A cursory glance at the docs seems to point to a SaaS offering if you want it, or an on-prem deployment if you need it. Maybe on-prem isn’t the right word, more like you can host the backend if you want, or have Sysdig host it for you.
|
||
It’s also not entirely clear how Sysdig agents are pulling information from all the other containers running on the host. That seems like a security issue. Obviously you need to be careful about exactly what Sysdig is collecting and whether it passes muster with your security and compliance folks.
|
||
Once the Sysdig agents are pulling the info, they appear to put it to work in multiple contexts. If you’re worried about metrics and monitoring, then you have Sysdig Monitor. If your concern is with security, then you have Sysdig Secure. There are also some open-source projects with Sysdig Inspect, Falco, and Sysdig + Prometheus.
|
||
All in all it seems like a good concept. I’d like to dig deeper into how the solution works and integrates with the various cloud platforms out there. Here are the questions I have:
|
||
Do you currently support or plan to support container deployments using AWS Fargate or Azure Container Instances? Is Sysdig a marketplace item in AWS or Azure today to simplify deployment? How are you handling balancing open-source and paid products? Are there plans to open-source the whole solution like Chef just did? What are you doing with all the aggregated monitoring data you are getting from clients? What are the major security concerns with your solution and how are you addressing them? Do you have questions for Sysdig? LMK and I’ll be happy to ask them too.
|
||
`,summary:"I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up.",date:"3 Apr, 2019",url:"https://nedinthecloud.com/2019/04/03/cloud-field-day-sysdig/",image:"analysis.png",readingTime:"4"},"https://nedinthecloud.com/2019/03/30/cloud-field-day-datrium/":{title:"Cloud Field Day – Datrium",tags:["cfd5","cloud-field-day","datrium"],content:`I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up. As a delegate, I am meant to represent the larger IT community, so I want to know what you think! In this post I am going to consider Datrium and what they have to offer in the cloud world.
|
||
Aside from a passing familiarity with the name - I’ve seen their booths at conferences and stickers on things - I don’t have any idea what Datrium does. Naturally, I went to their website to find out! Like most vendors these days, their website has Products and Solutions. I generally find the solutions area more interesting, because it tells me more about their aspirations and market fit than reading a product slick with speeds and feeds. Let’s do a fun exercise, I’ll list out the solutions and then try and figure out what the product is without reading the product page. Here we go!
|
||
Datrium Solutions Private Cloud Consolidation Microsoft SQL Server Virtual Desktops DevOps Cloud Data Management Oracle Database Healthcare That is a strange mix of solutions. Obviously they are doing something with storage and virtualization. They threw DevOps in there, so I am guessing that their product can fit into a CI/CD pipeline and has some type of API. The cloud data management makes me think that they are managing metadata to a certain degree. That also lines up with healthcare, which is a vertical very concerned with data governance and tracking. Okay, here’s my big guess. Datrium is software defined storage that assists in the management, consolidation, and categorization of data in disparate environments. I bet they do this through three products: OEM appliances, virtual appliances, and a SaaS offering housed in one of the major public clouds.
|
||
OK, how did I do? Off to the Products page!
|
||
Datrium Products On-Prem DVX Cloud DVX CloudShift Cool names, but what do they do?
|
||
On-Prem DVX The On-Prem DVX appears to be a combination of compute and disk nodes. The compute nodes are either Datrium branded or bring-your-own running the Datrium software. Datrium splits their solution into Performance and Protection. The performance nodes are running the actual hypervisor and are using Datrium to provide the storage abstraction layer. Protection nodes have a bunch of disk capacity and appear to be focused on efficient data storage and encryption.
|
||
Alright, so I am wrong on the virtual appliance product idea. This is a physcial device running a hypervisor and providing software-defined storage for compute to consume. Datrium can be installed on third-party hardware for the compute layer, but it looks like the storage layer is all Datrium hardware and software. Unsurprisingly, the special sauce appears to be how they handle data requests and the lifecycle management of virtual machines, databases, and containers.
|
||
Cloud DVX The Cloud DVX component is a SaaS offering running in AWS that allows the On-Prem DVX to send backups up to the cloud. Datrium uses global dedupe across their systems, so the data being written up to your Cloud DVX instance should use less storage than just shipping your uncompressed data up to AWS S3. Since it’s globally deduped, pulling data back down from the cloud should be faster as well since it can be rehydrated once it has been pulled down. This reminds me a lot of what Druva does for backups.
|
||
I was pretty on point with my Cloud DVX guess. It’s a SaaS offering sitting in AWS.
|
||
CloudShift This one is kind of interesting. It’s basically DR as a Service. You are still using DVX to manage all of your data. CloudShift is the management and control plane for the DR. Think of it as VMware SRM, if you’re familiar with that product set. The data plane can either be On-Prem DVX to another instance of On-Prem DVX or you could be syncing up to Cloud DVX. Since your Cloud DVX doesn’t have anywhere to recover your VMs, Datrium’s solution is to allow you to recover to VMware on AWS (VMC). Why? You have to figure that the On-Prem DVX is providing data services to VMware. So everything is already in that format - VMDKs and the like. In order to recover to anything non-VMware, such as EC2, Datrium would have to convert the format and do some other technical magic. Rather than doing that, they’ve opted to recover to VMC, which doesn’t require a conversion of formats. As someone who is not really into VMC, that doesn’t seem compelling to me. I’d rather see the ability to recover to AWS EC2, Azure VMs, or GCP.
|
||
Now that I’ve got a solid background in what Datrium is doing, I’ve got some questions:
|
||
Does your SaaS solution run entirely in AWS? Has that been an issue for anyone? Are you planning to support other recovery options outside of VMC? Are you planning to bring your storage tech to the cloud in either a virtual appliance or hardware form factor (think NetApp Cloud Volumes)? How are you supporting container based workloads? Is there a native storage driver/plug-in? Those are the questions I have for Datrium so far.
|
||
Do you have questions for Datrium ? LMK and I’ll be happy to ask them too.
|
||
`,summary:"I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up.",date:"30 Mar, 2019",url:"https://nedinthecloud.com/2019/03/30/cloud-field-day-datrium/",image:"analysis.png",readingTime:"5"},"https://nedinthecloud.com/2019/03/29/cloud-field-day-kemp/":{title:"Cloud Field Day – Kemp",tags:["aws","azure","azure-stack","cfd5","cloud-field-day","kemp"],content:`I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up. As a delegate, I am meant to represent the larger IT community, so I want to know what you think! In this post I am going to consider Kemp and what a load balancer company can do in the cloud better than the native tooling.
|
||
The first time I ever encountered Kemp was at a client site. They were using the virtual LoadMaster to load balance their Exchange 2010 environment. If you are unlucky enough to have lived through Exchange 2010 load balancing, then you know that it wasn’t exactly an optimal experience. You needed to enable sticky sessions, and add a whole host of different listeners. The setup could get rather complex. But I digress, the point is that they were using Kemp and I had to jump into using it as well to update their configuration to support Exchange 2013. Kemp’s UI was simple. They had templates for Exchange 2013 that worked almost out of the box. And I was told that the virtual appliance was very affordable. In my mind, I put Kemp in the category of low cost, simple to use, and probably not enterprise grade. Maybe that’s unfair, but it was my first impression.
|
||
The next time I saw Kemp was in the context of Azure Stack. Around the introduction of Azure Stack in general availability - I can’t remember if it was at GA, or just after - Microsoft added the Azure Stack Syndicated Marketplace. 3rd party vendors could make their Azure Marketplace solutions available for Azure Stack, and you could download the marketplace items and make them available to tenants of your Azure Stack deployment. One of the first vendors I saw on the list was Kemp! That was surprising. I had assumed I would find Citrix’s NetScaler or F5’s Big-IP. F5 is there now, Citrix is notably absent. But at launch, Kemp was the only option, which I found impressive.
|
||
Kemp obviously has an investment in cloud, even hybrid cloud. A quick look at their solutions areas lists the following:
|
||
Load balancing (LoadMaster) Multi-cloud (Kemp 360 Central) App optimization (LoadMaster) Security (Kemp 360 Vision) Basically they have their flagship product, the LoadMaster. And the LoadMaster has a bunch of additional features that help it support multiple clouds, application optimization, and security. This is common across all the major load balancers, who now are mostly calling themselves application delivery controllers. Load balancer just isn’t fancy enough.
|
||
Since they are embracing the cloud heavily, I have some questions about the next generation of applications and how they are handling it.
|
||
Supporting cloud native applications - A lot of the documentation is looking like it is aimed at traditional IaaS. Does the LoadMaster handle cloud native contructs like a Azure VM Scale Sets or AWS AutoScaling Group? Acting as an API gateway - Can the LoadMaster also act as an API gateway, providing protection, control, and throttling for external facing APIs? Integration with Kubernetes - Can the LoadMaster automatically front-end a service from Kubernetes? Container based deployment - Can the LoadMaster be deployed in a container? Could it work as a side car? Feature richness over native load balancers - Why would someone choose the LoadMaster over the AWS Application LB or Azure Standard LB? Metrics collection - Do the metrics, logging, and telemetry stream to native cloud services like AWS CloudWatch or Azure Monitor? Those are the questions I have for Kemp right now. I am sure they will be plenty more as they take us on a journey down their cloud roadmap.
|
||
Do you have questions for Kemp? LMK and I’ll be happy to ask them too.
|
||
`,summary:"I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up.",date:"29 Mar, 2019",url:"https://nedinthecloud.com/2019/03/29/cloud-field-day-kemp/",image:"analysis.png",readingTime:"4"},"https://nedinthecloud.com/2019/03/26/cloud-field-day-cohesity/":{title:"Cloud Field Day – Cohesity",tags:["cfd5","cloud-field-day","cohesity"],content:`I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up. As a delegate, I am meant to represent the larger IT community, so I want to know what you think! In this post I am going to consider Cohesity and how a backup company can become a data aggregator.
|
||
Cohesity at first glance appears to be a backup product vendor hocking a hardware appliance. In fact, I first heard of them in direct comparison with Rubrik. Whether or not that’s a fair comparison, it is fair to say that both vendors have a hardware appliance that performs backup and recovery of data.
|
||
And just like Rubrik, Cohesity has set its sights on loftier goals than being a backup company. The reason why has to do with aggregation theory and the relationship of backup products to their consumers and the data being backed up. I am largely basing this on the aggregation theory ideas that Ben Thompson has laid out on his blog Stratechery. When you think about a value chain, there are three players. The supplier, the distributor, and the consumer. It’s a reductive approach, and I realize there’s more to it than that. But the simple idea stands. The supplier is the one or more organizations that produce a good or service that they want to sell to a consumer. The distributor brokers the buying and selling of goods and services between suppliers and consumers. And the consumer purchases the goods or services.
|
||
Let’s apply that model to a backup appliance and the software it is running. The thing being supplied is data. Every device that stores data is a potential supplier. Backup software serves to collect this data in a central repository and then make that collected data available to a consumer. The consumer in this case is the target for a recovery of the data. Backup software is serving as a distribution channel by aggregating the information stored in disparate systems and making that information available when and where it is needed.
|
||
One of the original focuses of aggregation theory was Google. Google aggregates information from disparate sources and surfaces that data up to consumers. Suppliers who are interested in improving their contact with consumers can pay Google to surface results in a way that is favorable. Each individual supplier has very limited power, and the consumers also have very little power. By aggregating search results, Google has put itself in a place of power and their earnings have born this out. In the last 10 years, Google’s ad business has accumulated over $620B in revenue. Being able to aggregate the world’s data and serve it up to hungry consumers was key. The other major key was simplicity and performant user interaction. Yahoo was once considered a rival to Google, but they chose to take a curated portal approach. That made their website slow, clunky, and harder to use. Google kept things spartan and simple, and also managed create some amazing search algorithms that blew the competition out of the water.
|
||
What does all this have to do with a backup company? Backups take all your important information and keep a historical record for however long you need it. That sounds a bit like Google doesn’t it? Since it has all that information, it would probably make sense to index and catalog it. Maybe it also makes sense to run some data analysis tools on it. In fact, you could run data analysis tools across multiple datasets that aren’t traditionally stored together, since the backup software is aggregating all the data in your organization. What can you do with all that information?
|
||
Data analysis and trend recognition Machine learning model training Compliance discovery and auditing Document search and retrieval Those are some quick and easy ones, but I’m certain there are plenty more. The main point here is that the backup software in your organization has aggregated all your important data in one location. Why aren’t we doing more with it?
|
||
In the past there were a lot of technical limitations on the amount of data you could keep, and how rapid the retrieval of that data was. Storing backups on tape and shipping them to an offsite location was fine for disaster recovery and restoring missing data, but not practical for quick indexing and retrieval. Modern storage solutions and the public cloud have removed many of those blockers. Instead of doing incremental backups throughout the week and a full every weekend, now almost all solutions perform some kind of incremental backup that gets turned into a synthetic full programmatically. Data deduplication is basically a given for these solutions, and the only questions is how much data is included in each data dedupe library. Physical disk storage has been plummeting in price, with the average cost per GB now in the neighborhood of 2.5 cents. And with the advent of public cloud storage, the costs have dropped even further with support to backup locally and archive to the cloud for longer term storage. In most ways we’ve solved the data storage problem.
|
||
The other problem, frankly, is the backup software itself. As someone who has worked with multiple vendors, I can say that almost all of them have a terrible interface that is mired in the old way of doing things. With very few exceptions the backup UI is ugly and confusing. The backup policies and process are arcane and needlessly complex. And the backup clients are unreliable and constantly in need of tending. If backup vendors want to embrace the potential of being an aggregator, they have to do what other aggregators have done; build a compelling and easy user interface. That innovation is highly unlikely to come from one of the incumbents. They are deathlocked with their existing user base, being forced to continue supporting their terrible user interface and processes for those who have spent enough time with them to be trapped in a Stockholm syndrome type arrangement. I’ve seen it at clients who have these bizarre and incredibly inefficient processes they have built on top of their backup solution, and any shift in the solution - even a positive one - will result in great wailing and gnashing of teeth.
|
||
The innovation must come from a new vendor. One who can pioneer a better interface and user experience, leading the way to embrace all of the amazing possibilities a data aggregator makes available. Is that Cohesity’s vision? I have no idea, but I intend to find out at Cloud Field Day 5.
|
||
Do you have questions for Cohesity? LMK and I’ll be happy to ask them too.
|
||
`,summary:"I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up.",date:"26 Mar, 2019",url:"https://nedinthecloud.com/2019/03/26/cloud-field-day-cohesity/",image:"analysis.png",readingTime:"6"},"https://nedinthecloud.com/2019/03/21/cloud-field-day-nginx/":{title:"Cloud Field Day - NGINX",tags:["cfd5","cloud-field-day","f5","nginx"],content:`I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up. As a delegate, I am meant to represent the larger IT community, so I want to know what you think! In this post I am going to consider NGINX and where it fits in the cloudy, cloudy world.
|
||
I am most familiar with NGINX as a product that I use for demos and examples instead of Apache. When I need to give an example of how to spin up a new server with Terraform or deploy code through a CI/CD pipeline with Azure DevOps, I am probably going to install NGINX and run a simple website. It’s a super easy to do, and very lightweight. But that is not how I think most of the world uses NGINX. Looking at their website they have several products:
|
||
NGINX Plus: Basically the open-source version of NGINX with additional features. It is meant to replace your hardware load balancers. NGINX Controller: The controller component appears to be the centralized management plane for multiple NGINX Plus instances. NGINX Unit: Unit is an application server that sits behind your NGINX Plus load balancer and helps run your application written in Python, Go, Ruby, etc. NGINX WAF: WAF stands for Web Application Firewall NGINX Amplify: The monitoring solution for NGINX Plus including application and OS layer NGINX fits neatly into the world for application hosting and load balancing. All their other ancillary services appear to be expanding on that basic notion. If you have an application you want to run on the web, then NGINX can help you with the front end portion of that application. Most of what I read leads me to believe that NGINX is for more traditional style application deployment and management. You deploy a VM with an OS, install NGINX, and lay down your application. Then you join it to the NGINX Controller and monitor it with Amplify. That’s not exactly a unique proposition, but I appreciate that it removes the need for a physical load balancer.
|
||
On March 11th, NGINX announced that it was being purchased by F5. I’d heard that NGINX was shopping around for an acquisition prior to this announcement, so it wasn’t terribly surprising. It does raise some questions for me. F5 is very much a traditional load balancing company. The bulk of their portfolio is all about selling enterprise grade load balancers, and then keeping customers on the hook for software and support after the fact. The idea of big honking, hardware load balancers sitting in front of your applications doesn’t translate well for cloud native applications. The way in which applications are structured and presented is moving to micro-services with service discovery and service mesh. Purchasing NGINX is a way for F5 to modernize their portfolio and stay relevant for the next five years. With that perspective, I have some questions I’d like to ask NGINX at CFD 5.
|
||
What is the future of your open-source version of the product in lieu of the F5 acquisition? What are you doing to embrace and integrate with micro-services and Kubernetes based deployments? How do you plan to keep a culture of innovation after the merger with F5? Do you have questions for NGINX? If so, hit me up on Twitter or drop a comment below and I will make sure to include it during the session!
|
||
`,summary:"I will be a delegate for Cloud Field Day 5 on April 10-12. During the event we will be attending presentations from several vendors, which will be livestreamed. Before I leave on this grand adventure, I wanted to familiarize myself with each of the presenters and consider how their product/solution integrates with cloud computing. I’m also interested to hear from you about what questions you might have for each vendor, or topics you’d like me to bring up.",date:"21 Mar, 2019",url:"https://nedinthecloud.com/2019/03/21/cloud-field-day-nginx/",image:"featured-image-nedinthecloud.jpg",readingTime:"3"},"https://nedinthecloud.com/2019/03/03/scaling-the-kubernetes-cluster-in-azure-stack/":{title:"Scaling the Kubernetes Cluster in Azure Stack",tags:["acs","aks","azure","azure-stack","kubernetes"],content:`This is a follow-up to my post about the Kubernetes Cluster running on Azure Stack. In that post, I asked myself how to scale a deployed cluster and how to update the cluster. Since that post went live, I’ve done experimentation on my own, and also learned a few things about the deployment toolset being used for the Kubernetes Cluster Template.
|
||
If you remember from the previous post, the deployment process creates a virtual machine and clones the azsmaster branch of the Azure Cluster Service engine (acs-engine) from GitHub. The repo is a fork of the main Azure/acs-engine repo. But as Kenny Lowe helpfully pointed out to me, the acs-engine is being deprecated in favor of the aks-engine.
|
||
So ASKT* is not AKS, but AKS uses ACS Engine and so does ASKT, and ACS Engine is part of ACS but ACS is being deprecated, and ACS Engine will become AKS Engine, which will power AKS and ASKT (which isn't AKS). ACI is an independent nation state.
|
||
*Azure Stack Kubernetes Template
|
||
— Kenny Lowe (@KennyLowe) February 20, 2019 Which is also clearly echoed on the GitHub repo’s readme.md:
|
||
If you read through the notes on the repo, it becomes readily apparent that the code for the Kubernetes component of the acs-engine was moved over as is, and the rest of the codebase has been left to languish. The acs-engine that the Kubernetes Cluster on Azure Stack uses is from a deprecated source repo. One could easily infer from that information, that the next version of the K8s Cluster template will be using the aks-engine instead. Sure enough, if you look at the forks for the Azure/aks-engine repo, you’ll find there is a fork made by msazurestackworkloads.
|
||
Before I learned about all this, I did some reading about the commands that are available in acs-engine and determined that there was a scale command. It seemed to me that you could log into the dvm virtual machine that is created from the K8s Cluster template and manually run the commands. The scale command needs the following information:
|
||
Subscription ID of the existing cluster Resource Group of the existing cluster Name of the node pool Type of Azure environment (Azure Stack in this case) Authentication method type and credentials Deployment directory from original cluster creation FQDN of the master node(s) New number of worker nodes All of the necessary information is stored in the files from the original deployment, you just need to pull it from the files: apimodel.json, azuredeploy.parameters.json, and acsengine-kubernetes-dvm.log.
|
||
Here’s the full script:
|
||
Unfortunately, it turns out that the scale and upgrade commands are not supported on Azure Stack yet. You can see more about the error here on the closed issue for the GitHub repo. Since the repo is being deprecated in favor of the aks-engine, the issue was closed with a note that they are working on getting these commands supported with the aks-engine.
|
||
For the time being, the answer to scaling and updating the Kubernetes Cluster on Azure Stack is that you can’t. At least not with the toolset used to deploy it. You could go through the process of adding more nodes manually. You might even be able to clone one of the existing nodes and run some config scripts to add it as a new node.
|
||
My next project is to try out the aks-engine fork on Azure Stack and see if it is deploying properly. I’ll let everyone know how that goes in a future post.
|
||
`,summary:`This is a follow-up to my post about the Kubernetes Cluster running on Azure Stack. In that post, I asked myself how to scale a deployed cluster and how to update the cluster. Since that post went live, I’ve done experimentation on my own, and also learned a few things about the deployment toolset being used for the Kubernetes Cluster Template.
|
||
If you remember from the previous post, the deployment process creates a virtual machine and clones the azsmaster branch of the Azure Cluster Service engine (acs-engine) from GitHub.`,date:"3 Mar, 2019",url:"https://nedinthecloud.com/2019/03/03/scaling-the-kubernetes-cluster-in-azure-stack/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2019/02/28/using-azure-speech-service-with-powershell/":{title:"Using Azure Speech Service with PowerShell",tags:["azure","powershell"],content:`The 100th episode of Buffer Overflow - a weekly tech news podcast I host - is steadily approaching. As I write this, we are getting ready to record episode 98. In preparation for the 100th episode, I thought it might be nice to look over past episodes and find some common themes, running gags, and anything else that caught my eye. At an average of 35 minutes, that’s roughly 57 hours of combined audio. There’s no way I could listen to the entirety of the episodes, and so I started thinking. What if I could transcribe the audio to text, and then search through the text to find all the times we talked about Derrick and Miranda, how we’re all doomed, or smiling poop? The Azure Speech to Text API can be used to convert speech to text of audio files. Why not start there?
|
||
When I decided to try and use the Azure Speech Service, I initially assumed that I would be able to upload the files to Azure Storage and point the service at all my MP3s. Then it would dutifully transcribe all the files and dump out the transcriptions in some file repository. That is not exactly what the Speech Service does. The service itself can be used in concert with the Speech SDK to convert snippets about 15 seconds long. The point of that is to integrate speech to text in your applications. That’s not what I was looking to do. There is also a Rest API that supports batch processing, allowing you to send a request to the transcription endpoint with a file stored in Azure Storage. The batch process will transcribe the file, and then you can retrieve the results using another call to the API. That is more like it! Let’s do that.
|
||
Except, the Rest API is just an API, not a GUI or a menu. And there’s no PowerShell module for it. If you want to use the Speech Service SDK, there are examples on GitHub. The batch service examples in particular are only available in C#. I haven’t used C# in about six years. But I use PowerShell a lot! And PowerShell can talk to a Rest API using Invoke-WebRequest or Invoke-RestMethod. Why not just write the whole thing with PowerShell functions? And that is exactly what I did.
|
||
There are six functions that compose what I needed.
|
||
Get-AzSSBatchStatus: retrieves the current status of a particular transcript request ID Get-AzSSBatchResults: retrieves the JSON result files from a successful transcription New-AzSSBatchRequest: creates a new transcription request, and returns the URI of the request Remove-AzSSBatchRequest: removes a completed transcription request, whether is was successful or not New-AzStorageSASTokenAllBlobs: creates a SAS token for every blob item in a container on an Azure Storage Account, required for the Speech Service to access the blob New-AzSSMultiBatchRequest: creates a transcription request for every blob in a container, waits for each to complete, and saves result files to a directory The New-AzSSMultiBatchRequest is the function that leverages most of the other functions to get its work done. Essentially, you upload all your audio files to a blob storage container. Then run the New-AzSSMultiBatchRequest function and pass it the storage account information, the speech service information, and the destination directory for the results. Each blob is submitted to the speech service and then tracked to see if it is successful or fails. The successful items will have their JSON results files written out to the destination directory. The failed items will be reported in the final counts.
|
||
I chose to use the abbreviation AzSS for the Azure Speech Service. I thought maybe I would turn this into a full blown PowerShell module, with support for each function in the Rest API as defined in the Swagger document. Then I thought to myself, well if there’s a Swagger document wouldn’t it be possible to automatically develop a PowerShell module based on a Swagger spec? And then I checked and discovered that people have already written a project that does exactly that. At that point I realized that I probably could have saved a lot of time and used that to create four out of the six functions in my PowerShell project.
|
||
But you know what? I had fun writing the functions, as weird as that might sound. As someone who spends entirely too much time in PowerPoint instead of PowerShell and Microsoft Word instead of VS Code, it was really nice to just get into the zone and write some half-decent PowerShell scripts. And I learned some more about reading Swagger docs and using the Invoke-WebRequest and Invoke-RestMethod cmdlets. Ultimately, it was a worthwhile exercise. And if you have a need similar to mine, you are welcome to take my scripts and use them to your heart’s delight.
|
||
Of course now that I’ve run some transcription jobs, I have quickly realized how bad the default language model is. That’s especially true in a tech podcast that is full of jargon, acronyms, and strange proper nouns. Now I’m thinking I need to train a model with some pre-transcribed audio files. I’ll report back later on how that effort is going.
|
||
`,summary:"The 100th episode of Buffer Overflow - a weekly tech news podcast I host - is steadily approaching. As I write this, we are getting ready to record episode 98. In preparation for the 100th episode, I thought it might be nice to look over past episodes and find some common themes, running gags, and anything else that caught my eye. At an average of 35 minutes, that’s roughly 57 hours of combined audio.",date:"28 Feb, 2019",url:"https://nedinthecloud.com/2019/02/28/using-azure-speech-service-with-powershell/",image:"tutorials.png",readingTime:"5"},"https://nedinthecloud.com/2019/02/20/azure-stack-kubernetes-cluster-is-not-aks/":{title:"Azure Stack Kubernetes Cluster is NOT AKS",tags:["azure","azure-stack","kubernetes"],content:`If you’ve been test driving Azure Stack with a full stamp or just the ASDK, you may have decided to try out the Kubernetes Cluster template that is available in the marketplace syndication. This post is meant to walk through what the K8s template is, what it isn’t, and how it works.
|
||
The first thing to understand is that the Kubernetes Cluster template - herein KCT - is NOT the Azure Kubernetes Service (AKS). It is more akin to the Azure Container Service (ACS) that preceded the AKS. With ACS, Microsoft had developed a series of templates to roll out a container deployment using the orchestrator of your choosing. You could pick Mesosphere DC/OS, Docker Swarm, or Kubernetes. The template would help you roll out the deployment, but it wasn’t a managed service. It did not apply updates or automatically scale your cluster of nodes. That was up to you. Incidentally, ACS has been deprecated as an offering, replaced by AKS. Azure Kubernetes Service is a managed K8s cluster that allows you to easily scale the number of worker nodes and update the version of K8s in a managed fashion. The VMs running the master nodes in an AKS cluster are not even accessible to you.
|
||
Last year, Microsoft introduced a preview version of the Kubernetes Cluster on Azure Stack. It functions in a similar way as ACS, in that it deploys your K8s cluster through a set of ARM templates. In fact, as we work through the contents of the templates, we’ll see that it uses the same open-source acs-engine that the original Azure Container Service did. But it does not provide any kind of ongoing management of the cluster. It also does not have the same programmatic hooks as AKS. If you try to do a deployment from Azure DevOps using the AKS provider as a task, the deployment will fail with an error letting you know that the Microsoft.Containers resource provider was not available. That resource provider is what AKS uses, and it does not exist on Azure Stack. You can use a regular K8s deployment instead and provide the cluster dns name and kube config contents in the Pipeline.
|
||
That still leaves questions about how the KCT deploys the cluster and how you can manage it. How do you scale the workers? How do you update the version of K8s? How do you monitor the cluster health? Let’s take a look at the template deployment to get a sense of what is going on here.
|
||
The deployment is actually a nested deployment process. There are three deployments in total. The first two are executed directly, and the third uses an ARM template generated by the acs-engine running on a VM created for that purpose. The first template takes all of the parameters, creates some variables, and invokes a new deployment called the dvmdeployment. The dvmdeployment creates the following resources:
|
||
Storage account Network security group allowing port 22 access Public IP address with a DNS label Virtual network with a single subnet Network interface with the virtual network, NSG, and PIP associated to it VM using the storage account and network interface already created Virtual machine extension using the Linux custom script extension The important things here are the customData being passed to the VM and the Linux custom script being invoked. Let’s dissect those two items to see what’s going on with this dvm machine.
|
||
The customData is as follows
|
||
[base64(concat('#cloud-config\\n\\nwrite_files:\\n- path: \\"/opt/azure/containers/script.sh\\"\\n permissions: \\"0744\\"\\n encoding: gzip\\n owner: \\"root\\"\\n content: !!binary | [binary blob] I’ve omitted the binary blob, since it is super long and not helpful. Basically this is writing a file to /opt/azure/containers/script.sh. The custom script extension is running the following command:
|
||
[concat(variables('scriptParameters'), ' PUBLICIP_FQDN=', '\\"', reference(resourceId('Microsoft.Network/publicIPAddresses',variables('publicIPAddressName')),'2015-06-15').dnsSettings.fqdn,'\\"',' /bin/bash /opt/azure/containers/script.sh >> /var/log/azure/acsengine-kubernetes-dvm.log 2>&1')] Which is the JSON way of saying, run the script.sh bash script created in the customData field, pass it the PUBLICIPFQDN parameter and the variable _scriptParameters, and log output to acsengine-kubernetes-dvm.log. The variable scriptParameters has all of the deployment information that was part of the master template, including information like the subscription, tenant, service principal, and K8s cluster specifics. Which begs the question, what is in that script file? The script itself is 263 lines long, so I created a public gist. Feel free to peruse. I’ll hit the highlights.
|
||
First it copies the Azure Stack root certificate to /usr/local/share/ca-certificates/azsCertificate.crt. Then it installs pax, jq, and curl. Then it clones the azsmaster branch of the GitHub repo located here. Once it is done cloning, the script expands the archive examples/azurestack/acs-engine.tgz into a bin directory. Now it gets the json file necessary for deployment from the directory examples/azurestack/azurestack-kubernetes[version_number].json. Once it has that, it injects the storage account information into a temporary json file based on the primary json file. The temporary file is validated with a custom function and then copied back to the primary json file. Lastly, the credentials from the service account are added to the json file depending on whether the system is using ADFS or Azure AD for authentication. If everything is ready to go, the script finally runs the ./bin/acs-engine deploy command with all the necessary parameters passed in from the customScript invocation.
|
||
Got all that? The TL;DR is that we have created a Virtual machine to run an acs-engine process, and passed it all the settings from the original template filled out during deployment. Once that deployment is complete, this VM and all the accompanying resources could probably be deleted.
|
||
What is the acs-engine doing? Glad you asked. That is described is pretty good detail on the GitHub docs, so I won’t rehash it all. Suffice to say that it creates yet another ARM deployment, with its own associated template file. That final template does all the heavy lifting of creating the actual K8s cluster. The first 2098 lines of the template are simply the parameters and variable definitions. I’m not even going to post a gist of that. It’s a bit crazy pants. The only saving grace is that the template was programmatically generated. In terms of actual resource creation, the template creates the following:
|
||
Route table for networking Network Security Group for ssh and kubectl Virtual network Public IP Address for master nodes Load balancer for master nodes Internal load balancer for master nodes API Network interfaces, one per each worker node Storage accounts, one per the variable linuxpool2StorageAccountsCount An availability set for the worker nodes A VM for each worker node An availability set for the master nodes Storage account for the master nodes Network interfaces, one for each master node A VM for each master node And that is it. Each virtual machine has an associated customData and custom script extension command. The customData for the worker node is 8074 characters long. I’m not going to even try and parse all the stuff it is doing. Suffice to say that it is putting in place all the necessary components for a worker node in a K8s cluster. The custom script extension invokes the /opt/azure/containers/provision.sh script and logs to /var/log/azure/cluster-provision.log. The master node customData is 37,044 characters long. That is b/c of binary data being included for the kube-dns-deplyment.yaml file. Again, I am not going to try and parse through all of that. The master node custom script extension also invokes /opt/azure/containers/provision.sh script and logs to /var/log/azure/cluster-provision.log. If something goes wrong with your K8s cluster deployment, those are some of the log files you would want to investigate.
|
||
To get back to the original questions. How would you add another worker node to the cluster? Well, the short answer is that ;the acs-engine has a scale command that you could run from the dvmdeployment VM, in the same way that the deploy command was run. More information on the acs-engine scale command can be found in the docs. I haven’t tested it yet, but I plan to in the next few days. How would you update the version of Kubernetes? Acs-engine does not have an upgrade command. You might be able to use something like kubeadm, but I’m not sure. Since the cluster wasn’t built with kubeadm, I’m not sure if it can be used to do the upgrade. Again, more investigation is needed. Lastly, how do you monitor the cluster? That’s a bit more straight forward. Each VM has the Azure VM agent installed, and you can deploy the OMS agent on top of that. Then it’s a simple matter of getting the cluster plugged into Azure Monitor. Otherwise you can use the monitoring solution of your choice.
|
||
The offering is still evolving. And I expect that Microsoft will eventually replace this template with a more managed version. But until then, if you are trying to follow a hybrid application deployment model with Azure, Azure Stack, and Kubernetes then hopefully this helped you out.
|
||
`,summary:`If you’ve been test driving Azure Stack with a full stamp or just the ASDK, you may have decided to try out the Kubernetes Cluster template that is available in the marketplace syndication. This post is meant to walk through what the K8s template is, what it isn’t, and how it works.
|
||
The first thing to understand is that the Kubernetes Cluster template - herein KCT - is NOT the Azure Kubernetes Service (AKS).`,date:"20 Feb, 2019",url:"https://nedinthecloud.com/2019/02/20/azure-stack-kubernetes-cluster-is-not-aks/",image:"tutorials.png",readingTime:"7"},"https://nedinthecloud.com/2019/02/11/deploying-a-kubernetes-cluster-on-azure-stack-fails/":{title:"Deploying a Kubernetes Cluster on Azure Stack fails",tags:["asdk","azure-stack","kubernetes"],content:`I just finished updating my Azure Stack ASDK to the latest 1901 version. Before the upgrade I was messing around with the Kubernetes cluster offering, and I wanted to get that added back to my ASDK now that I’ve performed the update. I rushed through the process, and of course got an error. And that error was not very helpful. Just in case you’re like me, and missed a step in the setup for K8s on Azure Stack, here are the various error messages and the solution.
|
||
The TL;DR? Add the Service Principal you created for K8s as a Contributor on the subscription the cluster will be running in.
|
||
You’re welcome.
|
||
After attempting the deployment I got this.
|
||
Which is not terribly helpful. What does Conflict even mean? If you drill down to the Operation details, you get the following status message:
|
||
{ "status": "Failed", "error": { "code": "ResourceDeploymentFailure", "message": "The resource operation completed with terminal provisioning state 'Failed'.", "details": [ { "code": "VMExtensionProvisioningError", "message": "VM has reported a failure when processing extension 'LinuxCustomScriptExtension'. Error message: Enable failed: failed to execute command: command terminated with exit status=1\\n[stdout]\\n\\n[stderr]\\n" } ] } } Which basically means that the Custom Script Extension for Linux exited with an Error Code of 1. Again, this does not tell me anything that is helpful. Since the deployment leaves the virtual machine up and running, I connected via SSH and started to poke around. According to the troubleshooting docs on Microsoft’s site, I can check the log file located in /var/log/waagent.log for more info. I did that and got this helpful error message:
|
||
2019/02/11 16:25:26.423933 WARNING ExtHandler Unknown property: extensionHandlers.properties.upgradePolicy Repeated about a hundred times. That didn’t really point me in the right direction either. Nevertheless, I am persistent. After a bit more poking around I found another log I needed to look at /var/log/azure/custom-script/handler.log. That is the log which contains the error:
|
||
0 event="failed to handle" error="failed to execute command: command terminated with exit status=1" That’s the error I was seeing in the Azure Stack portal. Clearly something was amiss with the commands being run by the Custom Script extension. Now I needed to dig a little into the commands being handed to the Custom Script Extension. I looked at the actual template being submitted for this deployment. In that template, under the Resources category I selected the Virtual Machine.
|
||
In the JSON there is a nested virtual machine extension resource with the following command to execute:
|
||
"[concat(variables('scriptParameters'), ' PUBLICIP_FQDN=', '\\"', reference(resourceId('Microsoft.Network/publicIPAddresses',variables('publicIPAddressName')),'2015-06-15').dnsSettings.fqdn,'\\"',' /bin/bash /opt/azure/containers/script.sh >> /var/log/azure/acsengine-kubernetes-dvm.log 2>&1')]" I admit that isn’t easy to look at. The key portion of the script is that the logging is going to /var/log/azure/acsengine-kubernetes-dvm.log. Looking through that log I finally found the following error message:
|
||
failed to load apimodel: failed to get client: resources.ProvidersClient#List: Failure responding to request: StatusCode=403 -- Original Error: autorest/azure: Service returned an error. Status=403 Code="AuthorizationFailed" Message="The client 'b6d59fc6-e848-4612-aec0-817918cf6b8a' with object id 'b6d59fc6-e848-4612-aec0-817918cf6b8a' does not have authorization to perform action 'Microsoft.Resources/subscriptions/providers/read' over scope '/subscriptions/39c86ab5-8ac8-422b-b570-5fc6092c8702'." Again, not the prettiest thing to look at, but the main thing to get out of this is that the client does not have authorization to perform an action on the subscription. That triggered in my mind the fact that when you fill out the deployment fields for the Kubernetes template you need to supply a Service Principal that has Contribute permissions on the subscription. I had reused the Service Principal I created last time, but I had not granted it permissions on the new subscription that was created after updating the ASDK.
|
||
via GIPHY
|
||
I guess that means that next time I should RTFM, even if I have already read it once before. That being said, the troubleshooting process could have been easier if the logging messages had been surfaced up through the Azure portal. When a script fails, it should pass a helpful message through stderr and the invoking script should grab that error and surface it up through the layers. That way, instead of getting the incredibly unhelpful “command terminated with exit status=1”, I would instead have received the actual error message and would not have to dig through configs and logs to find it.
|
||
`,summary:"I just finished updating my Azure Stack ASDK to the latest 1901 version. Before the upgrade I was messing around with the Kubernetes cluster offering, and I wanted to get that added back to my ASDK now that I’ve performed the update. I rushed through the process, and of course got an error. And that error was not very helpful. Just in case you’re like me, and missed a step in the setup for K8s on Azure Stack, here are the various error messages and the solution.",date:"11 Feb, 2019",url:"https://nedinthecloud.com/2019/02/11/deploying-a-kubernetes-cluster-on-azure-stack-fails/",image:"tutorials.png",readingTime:"4"},"https://nedinthecloud.com/2019/02/09/day-two-cloud-podcast-launched/":{title:"Day Two Cloud Podcast Launched",tags:["aws","azure","day-two-cloud","packet-pushers"],content:`In case you didn’t notice, the Day Two Cloud podcast has officially launched! Big thanks to Tim Warner and Kenny Lowe for being the guests in the first two episodes! There is a lot more great content coming. I’ve got ten more episodes already recorded, and two more scheduled. If I stick to a fortnightly schedule for publishing, that should take me through July. That is pretty ridiculous!!! Needless to say that I am already considering moving to a weekly schedule.
|
||
I’ve had a few people ask me about where the podcast is hosted, what topics I might be interested in, and what my process is for publishing. The process for recording and publishing is a whole post unto itself, but I can address the other two topics here.
|
||
The seed for the Day Two Cloud podcast was planted by Ethan Banks. He hosts the Datanauts podcast with Chris Wahl, and one of their most popular episodes for 2018 was called Getting to Day 2 Cloud. It’s a seriously good episode, go check it out. I have been on Datanauts a couple times myself. Once for Azure Stack and again for Terraform. And then most recently I was a presenter and panelist for their Virtual Design Clinic series of virtual conferences. Basically, Ethan knew that cloud was a popular topic, and could potentially be its own podcast. He also knew that I have more than a passing interest in cloud and have a lot of experience in running podcasts. So he approached me about developing a cloud specific podcast for Packet Pushers, and I happily agreed! The process for getting a podcast going with Packet Pushers is to first publish at least five episodes in the Community Feed. This shows that you aren’t going to podfade and lets you work out the format and details of the podcast. Once things seem to have stabilized, the podcast graduates to its own channel. Obviously I’ve already got enough raw material recorded to make it well past the five episode mark, so now it is a matter of time. Once the podcast has graduated to its own channel, I think I will probably bump it up to a weekly podcast. Time will tell.
|
||
In terms of topics for Day Two Cloud, I definitely have some broad categories as well as specific interests. Broadly the topics break down into the following categories:
|
||
Cloud Migrations Cloud Native Development Monitoring of Cloud Disaster Recovery and Resiliency Security in the Cloud Using PaaS Cloud Solutions Cost Management in the Cloud Application Deployment in the Cloud Skills Gaps and Training Those are some of the broad categories, and by no means exhaustive. The guests I have had on so far have been leaning towards the Azure side of things. That’s in part because when I lit the Bat-Signal for podcast guests, the Microsoft MVP community was made aware. And they came flocking in droves. Seriously, Microsoft MVPs are the freakin’ best. It also means that things feel a little lop-sided. I’d like to rectify that by getting more people from the AWS and GCP communities involved. I’d also love to hear from someone using a smaller player in the cloud space, like Digital Ocean or Oracle Cloud. Why are they choosing that platform? What are the benefits? I’d also like to talk about some specific products and services and get a feel for how they work in the real world. That includes stuff like:
|
||
Kubernetes Fargate or ACI No Code Automation AWS Landing Zones or Azure Blueprints Terraform/Ansible/Salt VMware on AWS Snowball Edge and GreenGrass If you you have experience with any of these topics and want to be a guest, please let me know! Hit me up on Twitter, LinkedIn, or email Ned-@-Ned-in-the-Cloud-com. I’m considering adding a form to the website for people who are interested in being a guest to make things easier. In the meantime, look for the next few episodes where I’ll be covering things like Immutable Infrastructure with Rob Hirschfeld and a Failed Azure Deployment with Iris Classon.
|
||
`,summary:"In case you didn’t notice, the Day Two Cloud podcast has officially launched! Big thanks to Tim Warner and Kenny Lowe for being the guests in the first two episodes! There is a lot more great content coming. I’ve got ten more episodes already recorded, and two more scheduled. If I stick to a fortnightly schedule for publishing, that should take me through July. That is pretty ridiculous!!! Needless to say that I am already considering moving to a weekly schedule.",date:"9 Feb, 2019",url:"https://nedinthecloud.com/2019/02/09/day-two-cloud-podcast-launched/",image:"nimbieserver.png",readingTime:"4"},"https://nedinthecloud.com/2019/01/29/using-azure-active-directory-authentication-with-hashicorp-vault-part-2/":{title:"Using Azure Active Directory Authentication with HashiCorp Vault – Part 2",tags:["azure","azure-ad","hashicorp-vault"],content:"This is the second and probably final post in this series. If you haven’t read the first post I would highly recommend it. When we last left our erstwhile heroes, they had successfully setup the Azure authentication method on a Vault server and created a policy associated with a role in the Azure auth method. The policy grants access to a key-value store called webkv. Now comes the fun part, how does an Azure VM go about using the Azure auth method to access the secrets stored in webkv? So glad you asked!\nEverything in Vault is accessible via the API. Even when running commands through the Vault CLI, it is interacting with the HTTP/S front end. Hence the reason you must set the VAULT_ADDR environment variable, or the address of the Vault server with each CLI command. And that address is something like http://yourserver:8200. Using the Azure VM to interact with Vault is going to rely on using this same HTTP API front end.\nFirst things first, we are going to spin up a new Azure VM running Ubuntu 18.04 in the same Vnet as the Vault server we set up previously. This VM will be our web server that needs access to a secret in Vault to connect to a backend DB or something similar. In a production scenario, your Vault server could be sitting wherever you want. It just needs to be accessible by resources in Azure. For demonstration purposes, it’s easier to spin up a dev instance in the same Vnet as the Azure VMs that will be using it.\nWhen you are setting up the new Azure VM that will be your web server, make sure that you set this toggle to “On”.\nThat will create the managed security identity that the VM can used to authenticate with Azure AD. Once the web server is provisioned, you can SSH into the box and start configuring things. But first, let me describe the process flow.\nThe Azure VM is going to access its local metadata store and get information about itself. Then it is going to request an access token from Azure AD. Now that it has all that information, it is going to put that info in a JSON document and POST that doc to the Vault server using the Azure auth method. Vault will verify the parameters in the JSON doc, and if all goes well, it will include a Vault token in the response that has the web policy assigned to it. The web server can now use that Vault token to request secret data from the Vault server. That’s the workflow. We’re going to be running a bunch of manual commands to do all this, but it would be just as simple to create a simple function that would run during boot-up of your web server instances.\nWith all that said, let’s dive in. First we are going to install jq on the web server so that we can parse the JSON in the responses from Azure AD and Vault.\nsudo apt update sudo apt install jq -y Now let’s get some metadata from Azure using the special 169.254.169.254 address.\nmetadata=$(curl -H Metadata:true "http://169.254.169.254/metadata/instance?api-version=2017-08-01") That will get the metadata information about the instance, including important things, like the VM Name, Resource Group, and Subscription ID. If you want to know what is in the response just run:\necho $metadata | jq And you’ll get something fairly readable. Now let’s use jq to extract the information we need for the Vault token request.\nsubscription_id=$(echo $metadata | jq -r .compute.subscriptionId) vm_name=$(echo $metadata | jq -r .compute.name) resource_group_name=$(echo $metadata | jq -r .compute.resourceGroupName) Great! Now we need an access token from Azure AD. We will use the same magic 169.254.169.254 address, but with a different path.\nresponse=$(curl 'http://169.254.169.254/metadata/identity/oauth2/token?api-version=2018-02-01&resource=https%3A%2F%2Fmanagement.azure.com%2F' -H Metadata:true -s) In the response is the access token value. Let’s grab that value using jq.\njwt=$(echo $response | jq -r .access_token) Now we need to take all of this information and turn it into a JSON doc that we can send to Vault. I’ve got a template JSON doc that we can use.\n{ "role": "ROLE_NAME_STRING", "jwt": "JWT_STRING", "subscription_id": "SUBSCRIPTION_ID_STRING", "resource_group_name": "RESOURCE_GROUP_NAME_STRING", "vm_name": "VM_NAME_STRING" } Save that in a file named authdata_complete.json. Now we can use _sed to replace the values with what we got from the metadata.\nsed -i "s/ROLE_NAME_STRING/web-role/g" auth_payload_complete.json sed -i "s/JWT_STRING/$jwt/g" auth_payload_complete.json sed -i "s/SUBSCRIPTION_ID_STRING/$subscription_id/g" auth_payload_complete.json sed -i "s/RESOURCE_GROUP_NAME_STRING/$resource_group_name/g" auth_payload_complete.json sed -i "s/VM_NAME_STRING/$vm_name/g" auth_payload_complete.json Cool, now we have a JSON doc with all the info that the Azure auth method is going to want. For more detail you can check out the documentation page. There’s a lot of info you can submit, but as far as I can tell the minimum is the jwt and role name. So let’s go ahead and send the login request to our Vault server.\nlogin=$(curl --request POST --data @auth_payload_complete.json $VAULT_ADDR/v1/auth/azure/login) So we’re sending the JSON doc in a POST request to the login method of Azure auth. In response we will get a token from Vault with the web policy applied.\nNow we can extract the client_token value and use that to make a request against the Vault server for a secret in the webkv secret store.\nexport VAULT_TOKEN=$(echo $login | jq -r .auth.client_token) curl --header "X-Vault-Token: $VAULT_TOKEN" $VAULT_ADDR/v1/webkv/webpass | jq And in return we get this response:\nThat’s the ticket! This process could be applied to anything that is capable of getting an access token from Azure AD and sending a web request to Vault. Obviously this is not a production worthy example of how to do it. I would probably write something in Python or C# as part of a larger web application.\nNow you know how to use the Azure AD authentication method with Vault. Share and Enjoy!\n",summary:"This is the second and probably final post in this series. If you haven’t read the first post I would highly recommend it. When we last left our erstwhile heroes, they had successfully setup the Azure authentication method on a Vault server and created a policy associated with a role in the Azure auth method. The policy grants access to a key-value store called webkv. Now comes the fun part, how does an Azure VM go about using the Azure auth method to access the secrets stored in webkv?",date:"29 Jan, 2019",url:"https://nedinthecloud.com/2019/01/29/using-azure-active-directory-authentication-with-hashicorp-vault-part-2/",image:"tutorials.png",readingTime:"5"},"https://nedinthecloud.com/2019/01/23/using-azure-active-directory-authentication-with-hashicorp-vault-part-1/":{title:"Using Azure Active Directory Authentication with HashiCorp Vault - Part 1",tags:["azure","azure-ad","hashicorp-vault"],content:`I am currently working on a Getting Started course for HashiCorp’s Vault product. There was a pretty cool demo I put together for using Azure AD as an authentication source for Vault, but unfortunately I had to cut it for sake of time. I didn’t want it to go to waste though; so I figured I’d write about it here instead. Here’s what we’re going to do. Use the Managed Service Identity feature in Azure to give an Azure VM permissions to access secrets in Vault. This is the sort of thing that could be applied to anything that can receive an MSI in Azure, including App Service, Functions, VMSS, and more!
|
||
Now I am not going to get into what HashiCorp Vault is or why you might use it over Azure Key Vault or AWS KMS. There are def reasons, but I don’t want to get into those reasons right now. Instead let’s dive right into the tech. First you are going to need to have Vault installed on your local machine. Just follow this link, and get the executable. You are also going to need an Azure subscription, because of course you are silly. In order to get Vault talking to Azure, we are going to need to know a few things. The Subscription ID your are going to spin resources up in, and the Tenant ID where the MSI will be created. That’s pretty easy to come by. From the Azure Portal, spin up the Cloud Shell in Bash mode and enter the following:
|
||
az account list You’ll get output that looks roughly like this:
|
||
{ "cloudName": "AzureCloud", "id": "XXXXX-XXXX-XXXXX-XXXXX-XXXXX", "isDefault": true, "name": "MAS", "state": "Enabled", "tenantId": "XXXXX-XXXX-XXXXX-XXXXX-XXXXX", "user": { "cloudShellID": true, "name": "live.com#ned.bellavance@gmail.com", "type": "user" } } The id field is your Subscription ID and the tenantID, well… I think you can figure that one out on your own. If you have multiple subscriptions, you’re going to want to have the proper one selected. Go ahead and run:
|
||
az account set -s [SubscriptionName] The next thing we need to do is create an account for Vault to use when attempting to access Azure AD. That account is called a Service Principal. And since we’re already in Cloud Shell, we might as well do that here.
|
||
az ad sp create-for-rbac --name http://vaultsp --role reader --scopes /subscriptions/AZURE_SUBSCRIPTION_ID You will get something along these lines:
|
||
{ "appId": "d2b5e840-5554-4a35-8e21-b9a189e6c5ff", "displayName": "vaultsp", "name": "http://vaultsp", "password": "328da32c-d16e-4680-9f5e-9955dd85ddaf", "tenant": "XXXXX-XXXX-XXXXX-XXXXX-XXXXX" } Don’t worry, I’ve already deleted the account since this post went live. In order to configure Vault, we will need the appId and password values. Now let’s get to configuring Vault!
|
||
First we are going to start up an dev instance of the Vault server. This instance needs to be running either from a publicly accessible endpoint, or in a Vnet that an Azure VM can talk to. Once the Azure VM is authenticated by Azure AD, it is going to want to talk to the Vault server. I recommend spinning up an Ubuntu 18.04 instance for this in Azure. As long as the new Azure VMs will be running in the same Vnet, you won’t need to open any additional ports. Once you are logged in using SSH, you’ll need to install Vault. The easier way is to run the following:
|
||
#Install unzip sudo apt update sudo apt install unzip -y #Set the version of Vault to download VAULT_VERSION="1.0.1" wget https://releases.hashicorp.com/vault/\${VAULT_VERSION}/vault_\${VAULT_VERSION}_linux_amd64.zip unzip vault_\${VAULT_VERSION}_linux_amd64.zip sudo chown root:root vault sudo mv vault /usr/local/bin/ Now that we have Vault installed and ready to go, we can get an instance of that Dev server running. By default the dev instance is configured to run on the loopback adapter, which doesn’t work so well for external resources. So we’re going to need the IP address of the VM running Vault. Once you have that, go ahead and run:
|
||
vault server -dev -dev-listen-address=X.X.X.X:8200 Replacing X.X.X.X with your IP address. You’ll get a running instance of Vault, along with the root token. You should grab that info to login.
|
||
Now start a separate terminal session and run:
|
||
export VAULT_ADDR='http://X.X.X.X:8200' vault login Paste in the root token and you are now logged into Vault!
|
||
Now let’s enable the Azure authentication method in Vault:
|
||
vault auth enable azure Once we’ve enabled it, then we need to configure it.
|
||
vault write auth/azure/config \\ tenant_id=AZURE_TENANT_ID \\ resource=https://management.azure.com/ \\ client_id=AZURE_CLIENT_ID \\ client_secret=AZURE_CLIENT_SECRET Go ahead and plug in the information you gathered earlier. Now in order for an identity in Azure to get access to resources in Vault, we need to create a role in the Azure auth method. The role will be composed of policies within Vault that will be assigned to the role, and then a number of other properties. Some of the properties are prefixed with bound, and they serve to restrict which resources in Azure can request a token using this role. For instance, if you wanted to restrict the role to a specific location and resource group, you could use the bound_resource_groups and bound_location properties. The full doc is here if you are curious about other properties that can be used. Here’s what I used for this example:
|
||
vault write auth/azure/role/web-role \\ policies="web" \\ bound_subscription_ids=AZURE_SUBSCRIPTION_ID \\ bound_resource_groups=vault This command assumes that we have a policy named web and a resource group named vault. Go ahead and change the resource group to one that makes sense for you. Now we need to create the web policy. In order to create the policy, we are going to need a policy file formatted in either HashiCorp Configuration Language (HCL) or JSON. And since only a masochist would choose JSON, let’s try HCL instead:
|
||
path "webkv/*" { capabilities = ["read", "list"] } Take that text and save it as webpol.hcl. In a nutshell, we are granting the policy holder read and list rights to the path “webkv”. Of course the path “webkv” doesn’t yet exist, so let’s go ahead and create it:
|
||
vault secrets enable -path=webkv kv vault kv put webkv/webpass password=marvin And I dropped a secret in there for good measure. Now we can create the policy:
|
||
vault policy write web webpol.hcl That’s it on the Vault side. We’ve configured the Azure AD connection. We’ve created a role for an authenticated resource to assume. We’ve created a policy for that role, and created an instance of the key value store to hold secrets. Now we need to test this all out. That will be in the next post.
|
||
`,summary:"I am currently working on a Getting Started course for HashiCorp’s Vault product. There was a pretty cool demo I put together for using Azure AD as an authentication source for Vault, but unfortunately I had to cut it for sake of time. I didn’t want it to go to waste though; so I figured I’d write about it here instead. Here’s what we’re going to do. Use the Managed Service Identity feature in Azure to give an Azure VM permissions to access secrets in Vault.",date:"23 Jan, 2019",url:"https://nedinthecloud.com/2019/01/23/using-azure-active-directory-authentication-with-hashicorp-vault-part-1/",image:"tutorials.png",readingTime:"6"},"https://nedinthecloud.com/2019/01/10/some-helper-scripts-for-azure-stack-development-kit/":{title:"Some helper scripts for Azure Stack Development Kit",tags:["asdk","azure-stack"],content:`Not too long ago, I got a DL380 Gen10 from HPE to deploy the Azure Stack Development Kit. I had been limping along with a couple Frankstein systems running on Gen8 and Gen9 hardware. They had slow disks, not enough storage, and not enough RAM. This new beast has 384GB of RAM, 20 cores, and SSDs for the OS disk. Basically it’s awesome, and I am a very happy nerd. Since the early days of the ASDK, when it was just a little Technical Preview, there have appeared a growing library of scripts to help with the deployment of the ASDK. Since I am deploying the latest version today (1811), I thought it might be a good idea to share some helper scripts I put together to make the process a bit faster.
|
||
Below are two scripts that will help grease the deployment wheels. The first is meant to simplify the download of the ASDK files. If you are deploying for the first time, you should fill out the form on the Microsoft website and agree to whatever the EULA is for the ASDK. That being said, if this is the 20th time you’re deploying the ASDK, filling out the form and getting the downloader executable is a bit of a hassle. This script replaces that downloader executable and just pulls down the bin files and extraction executable to the location of your choice. It also grabs the askd-installer.ps1 file for good measure. After all, once you have the VHDX to deploy the ASDK, the next natural step is to actually deploy the dang thing.
|
||
The second script helps out with the process of pulling and running the ConfigASDK.ps1 script written by the excellent Matt McSpirit. I based it off this post by Kristopher Turner. I added a little piece to download the Server 2016 ISO. That’s an eval copy, so don’t be surprised if it expires after 120 days. Of course, after 120 days you’re probably going to redeploy the ASDK to get the latest version anyhow.
|
||
I hope this helps out someone besides me. I would also recommend checking out Kenny Lowe’s post about exposing your ASDK to your network.
|
||
`,summary:"Not too long ago, I got a DL380 Gen10 from HPE to deploy the Azure Stack Development Kit. I had been limping along with a couple Frankstein systems running on Gen8 and Gen9 hardware. They had slow disks, not enough storage, and not enough RAM. This new beast has 384GB of RAM, 20 cores, and SSDs for the OS disk. Basically it’s awesome, and I am a very happy nerd. Since the early days of the ASDK, when it was just a little Technical Preview, there have appeared a growing library of scripts to help with the deployment of the ASDK.",date:"10 Jan, 2019",url:"https://nedinthecloud.com/2019/01/10/some-helper-scripts-for-azure-stack-development-kit/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2019/01/08/day-two-cloud-podcast/":{title:"Day Two Cloud Podcast",tags:["buffer-overflow","day-two-cloud","podcasting"],content:`One of my goals for 2019 was to launch a new podcast. That process has officially started. The podcast is going to be called Day Two Cloud. I sent a tweet last week about the podcast to see if anyone would be interested in being a guest. The reaction was overwhelming. I was hoping to get two or three people to be guests. Instead I now have 14 interviews booked, and more people who are interested. I thought I would take to time to lay out what the podcast is meant to be, along with answers to FAQs that I have received from potential guests and listeners.
|
||
What is it? The Day Two Podcast is intended to explore topics related to real world cloud deployments. We’ve all been told how magically delicious the cloud is. It is all bacon, rainbows, and unicorns. The buzzword salad is hyperbolic and overwrought:
|
||
Deploy a million times a minute! Infinite scale on-demand!! Real-time metrics powered by ML and AI!!! Unlimited storage for free!!!!!!!!!!! The reality could not possibly live up to the marketing hype. And it doesn’t. Now bear in mind, I’m not crapping on the public cloud. It has brought some significant advances and helped me forge another chapter in my career. At the same time, it has given me more than my fair share of indigestion. [One cannot live on bacon and unicorns alone.] So I thought it would be great to sit down with some folks who have done real world deployments, who have felt the elation of success and the scourge of failure, and have a frank discussion of what worked, what didn’t, and what some lesson learned were. There are many topics I want to delve into including:
|
||
Hybrid Cloud / Multi Cloud Machine learning / AI High Performance Computing Cost optimization Automation and Orchestration CI/CD for apps Securing the cloud Cloud networking Skills gaps and training Monitoring Disaster Recovery Migration That’s just the first dozen I came up with. Each of those topics has the potential to be multiple shows, and broken out into more fine grained topics. There’s a lot of meat there, and I am starving!
|
||
Because the show is meant to reflect the reality of cloud deployments and not just the hype, I would like the guests primarily to be practitioners and not vendors. That’s not to say that vendors can’t tell a compelling story, it’s just that they are beholden to their employer. And while they won’t outright lie, they may sugarcoat or soften the edges on topics where their solution runs into some trouble.
|
||
I also want people to talk about their failures as well as successes. A story where you deployed a massively complex application across three different public cloud providers without a single hiccup is okay. But it isn’t really educational, it’s more like humble-bragging. I think listeners will learn a lot more from failure than success. It also makes for a better story telling arc. We all love the hero’s journey. A big part of that journey is the initial failure and learning, where our protagonist has a major setback that they now need to overcome. How they overcome and save the world is the truly interesting part.
|
||
I want to know how you saved the world despite insurmountable odds.
|
||
If there are no setbacks. If everything goes perfectly. If you never fail. You might just be living in a marketing slick. And those are NOT compelling stories.
|
||
The format of the podcast will be me interviewing one or more guests about a particular topic. Right now I am sticking with a single guest, but I am definitely considering multiple guests, or a round table format for special episodes. The run time will be around 30-40 minutes. That seems to be the ideal length, although I welcome feedback about that as well.
|
||
How often will it publish? I plan to publish it on a fortnightly basis. That means every two weeks. I would say bi-weekly, but that term is ambiguous. The cadence may change to a weekly podcast over time, but I didn’t want to commit to such an aggressive schedule out of the gate. I already run a weekly podcast, so I know how much work is involved. On the other hand, based on the number of people who have already raised their hand to be a guest, I think I could easily fill 52 slots in a year. As you can see, I am still very much on the fence about this.
|
||
Where and when can I get it? Right now? You can’t. I have only recorded one interview and I still need to edit it. I want to have at least five episodes ready before this thing goes live. Once publishing starts, I will make sure to update this post and publish it all over social media. If you’re really worried about missing out, send me a DM on Twitter. I will personally message you when the first episode drops.
|
||
How can I get involved? I am always going to be looking for new guests and topics. If you think you’ve got a great story about Day Two Cloud, let me know! You can reach me on Twitter or hit me up on email. All my contact details are on the About Me page.
|
||
I am super excited to get this project moving. More details will follow in the next couple weeks. Until then, may the sun shine brightly on your cloud and may Day Two be merry and bright.
|
||
UPDATE: You can now get the episodes from the Packet Pushers Community Feed. If you are already subscribe to the Packet Pushers Full Feed, then you will see the episodes appear auto-magically. Eventually, the podcast will graduate from the Community Feed to its own separate feed, and I will join the ranks of an official Packet Pushers podcast.
|
||
`,summary:"One of my goals for 2019 was to launch a new podcast. That process has officially started. The podcast is going to be called Day Two Cloud. I sent a tweet last week about the podcast to see if anyone would be interested in being a guest. The reaction was overwhelming. I was hoping to get two or three people to be guests. Instead I now have 14 interviews booked, and more people who are interested.",date:"8 Jan, 2019",url:"https://nedinthecloud.com/2019/01/08/day-two-cloud-podcast/",image:"nimbiewink.png",readingTime:"5"},"https://nedinthecloud.com/2018/12/31/goals-for-2019/":{title:"Goals for 2019",tags:["aws","azure","cfd3","hashicorp"],content:`I don’t believe in making New Year’s Resolutions. Or at least, I don’t believe in making the type of New Year’s Resolutions that you might typically think of. A grandiose resolution to achieve an overly ambitious goal in an unrealistic time-frame. Whether it’s resolving to start working out five days a week when you don’t work out at all, or losing 100 lbs. and keeping it off, or finally reading War & Peace. Those are all laudable goals, but setting your sights too high tends to end in failure. As in all things, moderation is key. I think it’s important to have a high-level goal, along with smaller milestones, and achievable tasks.
|
||
Let’s take running a marathon as an example. The high-level goal is to run a marathon. But if you just leave your house and try to run without any kind of plan or milestones, you’re probably going to stick with that plan for about a week. You have to set milestones, like being able to run a 5k in one month, a 10k in three months, a half-marathon in six months, etc. Then break those milestones into smaller goals, like run three times a week for the first month. Each of the activities, each run per se, is a task that has a purpose. In week 1 you might set a goal of running for 30 minutes each day, regardless of distance or speed. Breaking a monumental goal, like running a marathon, into something simple - running for 30 minutes - makes the entire process feel realistic. And each time you achieve your tiny goal, you get a sense of accomplishment. And if you track those accomplishments over the course of the high-level goal, you’ll be able to see real progress. Seeing that progress is a true motivator! How do I know? In 2012 I ran my first marathon, and this is exactly how I did it.
|
||
All of this is a VERY long-winded way of saying that I don’t believe in typical New Year’s Resolutions. I believe in setting goals, no matter what time of year it is, and creating a realistic plan to achieve those goals. That being said, the end of the year is an especially good time to reflect on what you accomplished in the previous year, and what goals you have in-flight for the next year. Having a well-defined moment in time to pursue internal reflection is necessary to staying on track or updating your plans to accommodate changes to your situation, and I don’t see any reason not to use the changing of the calendar year as such. The following items are goals that I have for 2019. Most of these goals are based on something that is already in-flight - remember, I don’t wait until January 1 to start a new project. I am going to try to provide some actionable tasks for each goal as well as metrics for success. Away we go!
|
||
Try and stay positive
|
||
That’s a pretty nebulous goal, I have to admit. Why did I even put it here? How does one stay positive? One thing I’ve noticed over the last couple years is that I am growing increasingly negative about the tech world. Being in consulting means that I usually don’t get involved until something is broken or a new product needs to be rolled out that no one understands. New products tend to be buggy, and often the marketing is a little ahead of what the product can actually do. The result is that I get a warped view of technology in production and in theory. Being a podcaster with a weekly tech news podcast also means that I tend to read a lot of stories about tech gone wrong, people in tech behaving badly, or the constant dumpster fire that is security on the web. Once again, I am getting a skewed version of the world, distorted through the lens of what reporters tend to focus on. The end result is that I have found myself becoming increasingly sarcastic, dismissive, and cynical about technology, people, and the world.
|
||
There’s a certain level of snark and skepticism that is healthy. And my concern is that I may have overindulged and crossed the line into unhealthy territory. Now I have to work my way back, and I’ve got a plan to do so.
|
||
Whenever I want to make a negative comment, find a positive one as well When a new idea is posed, look for ways to improve the idea (don’t tear it down) When something goes wrong, look for ways to make it right (don’t just complain) Start with the premise that people are usually trying to do good, and wait for compelling evidence to the contrary Be encouraging to anyone who is starting a new project, business, etc. (They’ll be plenty of detractors, be a supporter instead) There’s no easy way to measure this, but I’d like to pose one for myself. Put these ideas into practice when writing and podcasting. As I look back on each post, I need to keep a scorecard of how I am doing. That’s not to say that I can’t be critical, snarky, or skeptical. All those things are fine, in moderation. I need to maintain perspective and make sure I am not tearing anyone else down to build myself up. Which leads me to my next goal.
|
||
Be kinder
|
||
Here’s another goal that is completely vague and hard to measure. The reason it is on the list dovetails nicely with my previous goal. My increased negativity has made me meaner to individuals. And that is in direct violation of one of my three pillars. It doesn’t help that I had some rocky times with coworkers in the past two years. Those bad interactions have led to a certain amount of distrust and negativity about people in general. As a hypothetical example, let’s say that you had a terrible interaction with a used car salesperson. You might end up with a general dislike of all salespeople. That’s not fair to all the good salespeople out there who are trying to do their jobs. I am not saying that I want to be naive or trusting to a fault. But I do need to give people the benefit of the doubt and not cast judgment based on their occupation, company affiliation, or preference for spaces over tabs. Here are a few rules that I need to try and follow in the new year:
|
||
Be honest about what people have done, but don’t speak ill of them Try to help someone out everyday in some small way Look for opportunities to promote other people (at least once a week) No personal attacks (criticize the action, not the person) Practice focused acts of kindness (I don’t believe in the random thing, which is a whole other post on its own) Volunteer or raise money for at least one charity (in addition to making contributions) My default setting needs to be one of hopeful optimism and mutual respect. If one person lets me down, that doesn’t justify being mean to them. I don’t have to trust them, but I should still be kind about it. Rather than focus my energy on people who I dislike or distrust, instead I need to focus my energy on people who make the world better and brighter. The world is what we make of it, and I’d like to make it better. To that end I would like to help by…
|
||
Mentoring someone
|
||
Don Jones wrote an intriguing book called Be the Master that I read this summer. I don’t necessarily agree with everything in the book, but there was one thing that made an impression on me.
|
||
There’s a perception that you’re not “good enough” to teach until you know everything… you don’t need to be an “expert” in order to share knowledge. -Don Jones
|
||
I was invited to be part of a program at work where I would be assigned a mentor, and that mentor would rotate every six months. The idea is that each mentor would be able to assist me in moving forward on my career path, as well as impart wisdom to me on a topic they had expertise in. My mentors are not going to be infallible, omniscient gods. They are going to be mere mortals like me that have more experience in a particular area than I do, and have agreed to take the time to share that experience with me. That’s huge. It’s a commitment of their time above and beyond their regular job responsibilities, all with the focus of improving someone else’s life. I don’t know if they are getting some form of remuneration as part of the program, but I strongly suspect they are all motivated more by paying it forward.
|
||
Over my years in the tech industry, I have been immensely lucky to have bosses and coworkers who were willing to take the time and provide guidance and mentorship to me. It’s never been an official program, just awesome people who were willing to spend some of their precious time giving me a boost. Now it’s time to pay it forward. Someone reached out to me on LinkedIn to ask if I would be his mentor. I have to admit that I feel some trepidation about my level of expertise in career development and technology. While reflecting on the possibility, I found Don’s words to be comforting, and help give me the confidence to say ‘yes’. Even if this mentorship doesn’t work out, I would like to formally become a mentor for someone either at work or at large. I expect that the process will be deeply rewarding, and it would help me achieve my goal of being kinder and staying positive.
|
||
Produce five new courses
|
||
I don’t have a great segue from the last item into this one, so please forgive me. The last three items where externally focused. Goals that I hope will improve the world as a whole. The next few are focused on growing my career. Although, they do have an ancillary effect of helping others. The first item is to produce five new courses, ostensibly for Pluralsight. I have five courses on their site now, of which three were produced entirely in 2018. I really like creating learning content. It’s not all that surprising. My father, sister, wife, and mother-in-law are all teachers. My mother is a librarian. Learning and teaching is in my blood. Choosing to go work in technology made me a bit of an anomaly at family gatherings. No more! In the next year I want to produce five courses for Pluralsight, along with updating the content in some of the existing courses. My first course of the year, Getting Started with HashiCorp Vault, has already been approved and I am currently creating it. Now I just need four more! I’m open to suggestions if anyone has them.
|
||
Start a new podcast
|
||
I know, you’re thinking, “Ned, you already have two podcasts. Isn’t that enough?!” No, no it isn’t. Here’s the deal. I’ve been approached by Packet Pushers to create a new podcast about the reality of operating IT in the cloud, called Day Two Cloud. This is an opportunity for me to create a podcast not linked to my current employer, have really interesting conversations about cloud technology - my blog is called Ned in the Cloud after all, and make a little pocket change on the side. I’m still working out the details on format, length, and production. I am planning to start recording interviews with guests in January, and launching the podcast sometime in February. The basic premise is around the actual operation of products and services in the cloud. I want to have frank discussions about the original plan for implementing something in the cloud, and then what actually ended up getting rolled out. I want to hear about the challenges, successes, and most importantly the failures. I want to know what were the lessons learned, and things that could have been done differently. The cloud is a rapidly evolving ecosystem, with best practices being changed almost daily. I’d like to separate out the marketecture from the architecture, and see what happens when the cloud stops being polite and starts acting real. If you’re interested in being a guest, hit me up! You can be anonymous if you want. We can change the names of people to protect the innocent or guilty, if you think it might be an issue with your current employer or client. I do want to talk to people who are running real workloads in the cloud, and not the marketing folks from vendor X. They are fine people, but they also have an agenda and a paycheck to worry about.
|
||
Blog every week
|
||
I’d like to become a better writer. We can consider that the high-level goal behind this one. As part of accomplishing that goal, I want to make sure that I am blogging on a regular basis. If you want to be a better runner, you run more. If you want to be a better musician, you play more. And if you want to be a better writer, you write more. I think it is just that simple. Of course, I also should try and write posts in a way that is more elegant, clear, and concise. I can’t just write crap once a week and expect to get better. I also can’t expect to get better if I don’t write at all. Thus, the goal of this year is a minimum of one post per week.
|
||
Get Microsoft MVP award again
|
||
I’ve received the Microsoft MVP award for two years running. I’d like to keep up the streak. Access to product teams at Microsoft, Azure credits for testing, and the MSDN subscription are some of the biggest and most tangible benefits. But I also love meeting fellow MVPs and exchanging ideas. My award category is in Azure Stack, and with any luck I’ve participated enough in the Azure Stack community this year that I am awarded again for 2019.
|
||
Speak at more events
|
||
I’d like to become a better public speaker. We can consider that the high-level goal behind this one. In the same way that blogging every week is going to help make me a better writer, speaking at more events will help make me a better public speaker. In 2018, I spoke at a lot more events than I had in any previous year. That wasn’t an explicit goal for the year, rather it was more of something that I fell into. I started attending the Philly Azure group in the beginning of the year, and soon I ended up being a presenter. My employer started hosting the local AWS user group, and I volunteered to present for that. Then I was asked to be part of a panel on multi-cloud for a joint Philly Azure, Philly AWS, and Philly Google Cloud meetup. At the same time, I was invited to be part of Cloud Field Day three, and the Packet Pushers Virtual Design Clinic. Basically, a lot of things fell into my lap or required minimal effort. I also submitted talks to conferences, so it wasn’t all handed to me. I managed to get talks accepted for SQL Saturday, Microsoft Ignite, and NYC CloudExpo. There were some rejections too, I submitted talks to eight conferences and was only picked to present at three of them. I don’t know if that is high or low, but it did show me that I am not going to be selected every time. I don’t take that as a dis on me. The conference selection committees have to sort through all the submissions and find the talks they think will resonate best with their audience. In 2019, I am going to submit more, submit better, and get rejected more as well.
|
||
Learn Go for real
|
||
This might seem completely out of left field given all the other goals. Basically it comes down to two key observations. The first is that Go is a super powerful and popular language, and it is used on many of the open source projects I am most interested in. The second is that DevOps and cloud have a lot of cross over, and as someone who has primarily lived on the Ops side of the house I want to expand my horizons into Development. The larger goal is to start contributing to open source projects, and learning a programming language is one way to accomplish that. I attempted to learn Go last year, as I documented in my 2018 review post. That effort was stymied by other projects demanding my attention. This year, I am going to carve out some dedicated time to learn Go. I don’t think it will happen until June. In the meantime, I need to devise some smaller achievable tasks to help me stay motivated about learning Go. In training videos I watched this year, the examples were interesting enough to teach a concept, but I won’t really learn the language until I have a project with tasks. That might end up being contributions to Terraform, or I might try to create my own video game in Go. It has to be something that interests me, and something I will actually be able to do. Once again, I am open to suggestions.
|
||
And that’s it. That’s the list for 2019. I do have one other goal for 2019 that I am not ready to share at the moment, but I will when the time is right. In the meantime, I wish everyone a Happy 2019! Go make the world a little better than you found it.
|
||
`,summary:"I don’t believe in making New Year’s Resolutions. Or at least, I don’t believe in making the type of New Year’s Resolutions that you might typically think of. A grandiose resolution to achieve an overly ambitious goal in an unrealistic time-frame. Whether it’s resolving to start working out five days a week when you don’t work out at all, or losing 100 lbs. and keeping it off, or finally reading War & Peace.",date:"31 Dec, 2018",url:"https://nedinthecloud.com/2018/12/31/goals-for-2019/",image:"featured-image-nedinthecloud.jpg",readingTime:"14"},"https://nedinthecloud.com/2018/12/24/2018-year-in-review/":{title:"2018 Year in Review",tags:[],content:`This was a busy year for me personally and professionally. I started the beginning of 2018 feeling a little unsure about what was going to happen. There were a lot of changes happening at work. I had accepted a new position in June 2017, and that role was still evolving and refining. I was slogging through the development of my second course on Pluralsight. And I was trying to get back in the habit of blogging on a semi-regular basis. Now that the year is winding down, I thought it might be nice to take a look back at what happened, and how I fared on things I wanted to accomplish. The esteemed Scott Lowe does something similar each year, although he is far more organized about it. My plans for 2019 will be a separate post.
|
||
Here are the goals I had for 2018, or at least the ones I knew about at the beginning of the year:
|
||
Build a successful cloud practice Learn to program in Go or Python (or both!) Keep creating new Pluralsight courses Blog on a weekly basis Podcast on a weekly basis Speak at more user groups and conferences Build a successful cloud practice
|
||
In the summer of 2017, I accepted a new position at work to develop our Cloud Solutions offerings. It was a shift from doing consulting work at a client, which was my job for the previous five years. That’s the kind of work that is rewarding, but also can burn you out after a while. In mid-2017 I knew I was ready for a change, and this shift was something new and interesting. Director of Cloud Solutions became my somewhat nebulous title. I don’t manage anyone directly, so I’m not directing people. I’m no longer delivering on projects for clients. I’m also not a sales person, and I am not directly responsible for a P&L center. The closest thing I can equate it to is a combination of a program manager and technical evangelist. It took me the better half of 2018 to figure out exactly how I fit into the organization in my new role, and what I needed to do to be successful. I’m still figuring that out, but things have shifted internally so that I now feel I have the required support and guidance to be successful. That’s the abbreviated version, and probably the one that is appropriate to tell publicly.
|
||
Learning a new role in a new group at a company that is shifting its focus is perilous ground to dance on. I think I could have done a better job asking for help, taking ownership of the new practice group, pushing back on unrealistic expectations, and setting proper expectations early on. People usually don’t like to talk about their failures, and neither do I. But I want to be honest with myself about this one. I failed to achieve my goal here, and the best I can do is to try and learn lessons that I can apply in 2019. Overall I would give myself a D on this one.
|
||
Learn to program in Go or Python
|
||
“I am not a developer.” That’s a mantra I hear from Ops and Infrastructure people a lot, myself included. As I started 2018, I was working on my second course about Terraform for Pluralsight. I was spending a tremendous amount of time in VS Code creating Terraform configurations, ARM templates, and PowerShell scripts. For someone who is not a developer, I was spending a lot of time in development tools. I thought that perhaps I should start learning a new language outside of PowerShell. Terraform is written in Go, and that seemed like it might be a good place to start. I also started to do some work with Ansible, which relies heavily on Python, so that also seemed like a logical place to start.
|
||
How did I do? I worked through half a course on Python, and then got distracted. The main problem was that I didn’t have an actual project or goal to apply my knowledge to. I’m a practical learner, and I really need a project to drive my learning process. I actually managed to learn a decent amount of Go. I did a series of blog posts about Terraform functions. When a particular function was acting weird, I went and looked at the source code on Github. As a result I ended up learning how to read Go code, even if I wasn’t actively programming. There was one function, timeadd, that didn’t handle dates. I decided that I wanted to try and write a a new function, dateadd, that would introduce this functionality. I issued a pull request into Terraform main, but that’s where things got bogged down. One of the maintainers wanted me to incorporate the functionality in dateadd as an extension of timeadd with additional arguments. They also wanted me to create a struct for the data type and do some other stuff. That was beyond the little I had learned in Go, and unfortunately I didn’t have time to dedicate to learning it. Overall I would give myself a C on this one. I learned some Go and applied it to a project. I learned some Python and applied it to nothing.
|
||
Keep creating new Pluralsight courses
|
||
In 2017 I created my first Pluralsight course. In 2018 I created four more. The reason I didn’t have enough time to keep learning Go? It was because the bulk of my free time was dedicated to creating these courses. The acceleration of course development was fueled by a new partnership between Microsoft and Pluralsight, which created a demand for Azure based courses as a ramp up to Microsoft Ignite. My last three courses have all been about topics in Azure, and now you know why. Creating these additional courses on a compressed timeline forced me to streamline the process by which I created the slide decks, recorded the content, and edited my recordings. The last course I did was eight modules over four weeks. That’s two modules a week! Overall I would give myself an A on this one.
|
||
Blog on a weekly basis
|
||
Based on my count, I published 81 posts in 2018, not including this one. On average that is more than one a week, but the posts were not evenly distributed. The series of posts on Terraform functions published daily. The series kept me blogging consistently from June through September. It also burned me out for a bit. During the rest of the year I was able to post at least once a month. Overall I would give myself a B+ on this one.
|
||
Podcast on a weekly basis
|
||
One of the podcasts I do for work, Buffer Overflow, is a weekly tech news podcast. Knowing that we had committed to a weekly cadence made it easier to be consistent in delivery. My cohost Chris and I realized pretty early on that we needed some help generating the necessary content for each week. To that end, we started asking other people around the office to be a guest on the show. Guests would add more perspectives to our podcast, provide coverage when Chris or I weren’t available, and help source topics for us to talk about it. It was honestly the best thing we could have done for the show! Kim, Brenda, Jen, and Hank have helped the podcast become more entertaining and kept Chris and I grounded in reality. I cannot thank all of them enough for being willing to dedicate their time and effort on the show in addition to their daily job responsibilities. Overall I would give myself an A on this one.
|
||
Speak at more user groups and conferences
|
||
Part of building up my personal brand and developing my career was doing more public speaking. This year I did a lot of that:
|
||
Philly Azure user group (twice) Philadelphia AWS user group SQL Saturday in Philly Microsoft Ignite CloudExpo in NYC Anexinet Annual Golf Outing Emerging Technology Council (twice) Virtual Design Clinic by the Packet Pushers (panelist and presenter) Cloud Field Day 3 (delegate) On-Premise Podcast by Gestalt IT Datanauts Podcast by The Packet Pushers You know when I put it all in a single list, it seems like a lot. And it was! It’s not the crazy speaking schedule that professional presenters might have, but it was still a lot more than the previous year. Over the course of the year, I could tell me presentation skills were getting better. My PowerPoint decks were more engaging and my demeanor was more relaxed. Overall I would give myself an A on this one.
|
||
It was a busy and productive year. I met some of my goals, and fell short on others. That’s okay. If I was successful in every endeavor, it would probably mean that I wasn’t challenging myself enough. If you’ll forgive the indulgence, my three pillars are:
|
||
Make yourself uncomfortable Plan to fail Be nice There’s no doubt that I made myself uncomfortable this year, and there’s also no doubt that I failed on some of my goals. I tried to be nice this year as well, though I think I could have been kinder to some people. How do I plan on pursuing my pillars in 2019? That will be the topic of my next post.
|
||
`,summary:"This was a busy year for me personally and professionally. I started the beginning of 2018 feeling a little unsure about what was going to happen. There were a lot of changes happening at work. I had accepted a new position in June 2017, and that role was still evolving and refining. I was slogging through the development of my second course on Pluralsight. And I was trying to get back in the habit of blogging on a semi-regular basis.",date:"24 Dec, 2018",url:"https://nedinthecloud.com/2018/12/24/2018-year-in-review/",image:"featured-image-nedinthecloud.jpg",readingTime:"8"},"https://nedinthecloud.com/2018/12/16/hybrid-cloud-is-on-target-for-2019/":{title:"Hybrid Cloud is On Target for 2019",tags:["aws","aws-outposts","azure","azure-stack","hybrid-cloud","kubernetes","vmware"],content:`If you were going to build a brand new application today, your approach would probably be fundamentally different than five or ten years ago. And I do mean fundamentally, as in the fundaments of the architecture would be different. In the last ten years we have moved rapidly from traditional three-tier applications to 12-factor apps using microservices, and now things are shifting again to serverless. That’s all well and good for any business looking to build a new application, but what about organizations that have traditional applications? I’ve also heard them called legacy or heritage applications. These applications are deeply ingrained in the business and are often what is actually generating the bulk of a company’s revenue. The company cannot survive without these applications, and modernizing them will be costly and fraught with risk. Due to the inherent risk, most companies opt to either keep these applications running on-premises or move them as-is to the public cloud, aka life and shift. That’s the reality we’re living with today, but tomorrow is knocking on the door and promising hybrid cloud to fix all this. What’s the reality and what’s the hype? And what is the most likely journey for most companies?
|
||
First we have to define some terms. The original meaning of hybrid cloud referred to running an application with some components running in the public cloud, and some running in a private cloud. What does that look like? Well, let’s say that you have a company called OnTarget that makes paper targets for archery and firing ranges. You have a B2C website that runs in your on-premises datacenter with a web front-end, an application middle tier, and a database backend. Your busiest time is the ramp up to summer camp in the late Spring, and your website gets slammed with requests. Rather than buy the necessary capacity to handle your peak load, you’ve decided to try and burst to AWS with EC2 instances running your website. That was the original vision behind hybrid cloud, stretching the application across public and private cloud. However, you run some load testing on your website with the EC2 instances in place, and discover that the latency between the EC2 instances in AWS and the application and DB tiers on-prem is just too high. Looks like bursting to the cloud isn’t going to work. And you’re not alone in this discovery. The idea of stretching an application may look attractive on a balance sheet, but in practice it usually doesn’t work well.
|
||
The newer version of hybrid cloud doesn’t require this stretching of an application. Instead, the idea is to have a consistent deployment methodology and set of services being utilized in both the public and private cloud. Running with the OnTarget example, let’s say you’ve decided to move your Production systems up to AWS to take advantage of the scalability and elasticity of the public cloud. You can run your website at a nominal capacity most of the year, and then ramp it up for peak season without having to by enough hardware to cover the maximum demand. The other environments - development, QA, staging - will continue to run on-premises using your existing hardware investment. Your lead software developer wants to start using some of the cloud-native services in AWS, such as Elastic Beanstalk and RDS, but you’re hesitant to do so. Those services don’t exist in your on-prem datacenter, and you want a consistent deployment model across all your application environments. After all, QA is validating the application based on what is running on-prem. If you have to make code changes to deploy the Production version of the application in AWS, then you are invalidating the work that QA has done.
|
||
Now you’re faced with a choice. You wanted to keep using the hardware you had on-premises, at least until it has depreciated off your balance sheet. But you also want your application to evolve and improve, by using the native services in AWS. What to do? In the recent past, most companies would end up moving all the application environments to the cloud, and begin the work of transforming their application. There is another choice on the horizon, and it takes the form of the new hybrid cloud. I think there are three options emerging:
|
||
Azure Stack and AWS Outposts offer an on-premises version of the public cloud in your datacenter Managed Kubernetes are creating a container-centric hybrid deployment model VMware on AWS creates the ability to change nothing and still have hybrid applications The first option doesn’t allow you to keep using your existing hardware. Both Azure Stack and AWS Outposts require pre-configured hardware from an OEM with Azure Stack or from the vendor directly with AWS Outposts. Either way, you cannot just deploy one of these solutions on your existing hardware. If that is your goal, then option one is not really an option. What both Azure Stack and Outposts bring is the ability to start re-architecting your application to use cloud-native constructs in Azure or AWS, and still be able to run that application on local hardware as needed. You can also use the same deployment and management tool chains for public or private cloud. This option paves the way for the use of serverless in your application. Azure Stack already supports Azure Functions today, and I am 100% certain that Outposts will support Lambda very soon after launch. Taking the broader definition of serverless being any service that you can allocate on demand and do not have to manage the underlying service, then other “serverless” offerings, like DynamoDB and CosmosDB, won’t be far behind either.
|
||
The second option would allow you to keep using your existing hardware, assuming it meets the requirements of the managed Kubernetes solution. You could manage it yourself, but that is probably a bad idea. I’m not going to even try and get into Geoffrey Moore’s guidance around Core vs. Context, but suffice to say that OnTarget’s key market differentiator is not managing their own Kubernetes. The great thing about using managed K8s, is that you can re-architect your application in a container-centric way. Once you have packaged up your application with something like Helm, you can deploy it to any Kubernetes solution, including public and private cloud. In the case of Red Hat’s OpenShift, you actually deploy the management components in the public clouds as well (OpenShift on Azure), and now you’ve got the same management and deployment toolkit across all your environments. That certainly hits the mark with hybrid cloud. This option won’t let you start using native cloud services in AWS and Azure in your on-premises environment. There are several projects to use Kubernetes as a base layer for serverless - functions as a service really, so this may become a viable option if you are trying to add FaaS to your application.
|
||
The third option allows you to keep the exact same deployment model you have today, assuming you are a VMware shop. You can replicate and migrate existing virtual machines to VMware Cloud on AWS (VMC). As your older hardware ages out, you can move workloads to VMC. AWS Outposts is also offering a flavor of VMC on the Outposts hardware, so you could replace your on-prem hardware with Outposts as it ages out. The bad news is that you won’t be taking advantage of any cloud native services or serverless technologies. VMware does have plans to pursue a managed version of Kubernetes called PKS and PKS Cloud. That begs the question, why wouldn’t you just go with managed Kubernetes and skip VMC altogether?
|
||
Which option would you choose at OnTarget? That all depends on your business drivers and IT strategy. Let’s say that OnTarget is looking to move to a container based deployment model for its application. And you’d like to keep using your on-prem hardware for lower environments. Then managed Kubernetes is probably the thing that makes the most sense for you. Since you would also like a consistent deployment model, you might consider deploying Red Hat OpenShift on-prem and in AWS.
|
||
That is just one possible scenario, and I think all three of the above options will continue to be viable for the next five years. Beyond that, I suspect we’re going to see Azure Stack and AWS Outposts slowly begin to edge out most other on-premises offerings. Organizations that had performed an early lift and shift to the cloud will choose to repatriate some of their applications on-prem using one of these hybrid offerings. I also expect some kind of partnership to happen between Nutanix and Google Cloud. Nutanix is blowing up, and right now GCP doesn’t have a dance partner at the hybrid mixer.
|
||
The second generation of hybrid cloud is well underway, and I think 2019 is going to bring an acceleration of this trend.
|
||
`,summary:"If you were going to build a brand new application today, your approach would probably be fundamentally different than five or ten years ago. And I do mean fundamentally, as in the fundaments of the architecture would be different. In the last ten years we have moved rapidly from traditional three-tier applications to 12-factor apps using microservices, and now things are shifting again to serverless. That’s all well and good for any business looking to build a new application, but what about organizations that have traditional applications?",date:"16 Dec, 2018",url:"https://nedinthecloud.com/2018/12/16/hybrid-cloud-is-on-target-for-2019/",image:"publiccloud.png",readingTime:"7"},"https://nedinthecloud.com/2018/12/07/is-aws-going-to-destroy-your-business/":{title:"Is AWS going to destroy your business?",tags:["amazon","aws","azure"],content:`There was a recent article on CNBC talking about how AWS is creating new solutions that could potentially put existing companies out of business. That prompted a tweet from Matthew Prince about how if you are running your company on AWS, you are feeding them data about how to beat you. Prince is certainly not the first to make this leap of logic, and I am certain he won’t be the last. During the Microsoft Ignite Keynote this year, Satya Nadella said something similar about choosing a public cloud partner. Here’s the direct quote:
|
||
“If you’re dependent on a provider, who through some game theory construct, is providing you a commodity on one end, only to compete with you on another end. Then you could be making another strategic mistake.”
|
||
Is any of this true? Does is matter if you run your company on AWS or Azure? Is AWS or Amazon going to destroy your business? I have some thoughts.
|
||
First let’s define what is actually being argued here. As a hypothetical, let’s say you have a business called LogSlap! that provides a new log analytics capability to clients. You’re a modern hip company, and running your own infrastructure is so passe. Naturally, you’re going to run your service out of a public cloud. But which one?! What if you put your service in AWS, and they harvest analytical data about LogSnap! 1.0 and recreate your service? Then you’re out of business! Well, better not use AWS then.
|
||
There’s some many problems with this argument, I feel a list coming on!
|
||
Your spend in AWS is irrelevant AWS is not spying on you AWS doesn’t need you to run in AWS to destroy you Azure makes services too Competition is normal Your spend in AWS is irrelevant
|
||
One thing I keep hearing is that you are paying AWS to compete with you. That’s malarkey. You are paying AWS for services that they are rendering to you. One of the core tenets of AWS’ Well Architected Framework is cost optimization. One of the core tenets of running a successful business is managing costs. When you are deciding which public cloud provider to use, one of the factors should be how much it will cost to run your solution in that environment. It may not be the sole determinant, but it will definitely be in the top five. All other things being equal, if AWS is the cheapest place to run your service, then you’d be a fool not to run it there. There’s no reason to pay more for the same service, unless things aren’t equal. For instance, what if AWS is spying on me?
|
||
AWS is not spying on you
|
||
I should probably qualify this. When you deploy a application in AWS, they definitely know what services you are using and how much of those services you are consuming. Otherwise they wouldn’t be able to bill you. Beyond that, they don’t have access to anything. They aren’t looking inside your EC2 instances, or S3 buckets, or Lambda code. They aren’t pulling the source code for your solution our of CodeCommit. To do so would violate all of their security attestations, and land them in some significant legal trouble. Plus, it would completely destroy people’s trust in AWS as a company. Is AWS willing to risk their $40B annual run rate on potentially gaining some insight on your service? I’m going with ’no’ on that one.
|
||
AWS doesn’t need you to run in AWS to destroy you
|
||
There is no reason to believe that AWS isn’t going to eventually destroy you. If you have a successful business model and product, and AWS/Amazon decides they can do better, then they are probably right. And it doesn’t matter if your service was running on AWS or Azure or Bob’s Crab Shack and Public Cloud. AWS has a lot of very smart engineers that can figure out how your service works, and maybe even design one that works better. Your best bet in that situation is to hope they acquire you, rather than build it in-house. But, just because AWS puts out a similar offering doesn’t mean that your company is de facto done for. We live in a competitive, capitalistic marketplace. Even though AWS has announced a new Log Analytics platform, it doesn’t mean that LogSlap! is suddenly worthless. The clients you have today chose your service for reason. It’s up to you to make sure you retain a competitive advantage over the offering from AWS. That advantage doesn’t have to be price, in fact you’re going to lose on price. AWS and Amazon can have multi-year loss leaders that put other brands out of business. Just ask Toys’R’Us… oh wait, you can’t. Any service from AWS is not going to meet the needs of every customer, and they aren’t striving to do that. There will almost always be market segments that are under-served, and you can be there to help. Should you move your service to Azure though? I mean AWS is now a competitor. Not so fast Charlie.
|
||
Azure makes services too
|
||
Despite Satya’s attempts to quell fears about Microsoft being a competitor, the fact remains that they are. Microsoft makes services; these days they are primarily a services company. Everything else they make is dedicated to increasing consumption of their services. That means that Microsoft may very well create a service in Azure that directly competes with your fancy Log Analytics software. What are you going to do then? Move to Google Cloud? Same problem. If you get big enough, and it makes financial sense, you can move things in-house. Dropbox did that. Dropbox also still uses AWS for some portions of their service, because it makes good financial sense to do so.
|
||
Competition is normal
|
||
It would be a fair point that all the examples I gave are around SaaS offerings. Those are the things that are brought up most often in these types of arguments. But the counter example would be something like Walmart. They are not a SaaS company, and they have decided not to run in AWS because Amazon is a major competitor. And I say, that’s fair. Walmart’s annual spend on Azure is probably in the hundreds of millions. At that size and scale, Walmart’s spend would have a material impact on Amazon’s bottom line. That is an extreme example though, and I still say, if running your business is more cost effective in AWS, and you aren’t Walmart, then don’t worry about it. Amazon looks for new businesses to break into. There’s a chance that they could look at spend in AWS and see which business appears to be spending the most money, and then leap to the conclusion that maybe they should get into that business. Of course, Amazon doesn’t really need that information. Businesses that are successful tend to have investors and produce financial reports. Even privately held companies give indications about how much revenue they are bringing in. If you are a successful business, you’re usually not quiet about the fact. If Amazon catches wind and turns their Gorgon gaze upon you, running in AWS is going to be the least of your worries.
|
||
`,summary:"There was a recent article on CNBC talking about how AWS is creating new solutions that could potentially put existing companies out of business. That prompted a tweet from Matthew Prince about how if you are running your company on AWS, you are feeding them data about how to beat you. Prince is certainly not the first to make this leap of logic, and I am certain he won’t be the last.",date:"7 Dec, 2018",url:"https://nedinthecloud.com/2018/12/07/is-aws-going-to-destroy-your-business/",image:"publiccloud.png",readingTime:"6"},"https://nedinthecloud.com/2018/11/30/aws-outposts-and-azure-stack/":{title:"AWS Outposts and Azure Stack",tags:["aws","aws-outposts","azure","azure-stack","vmware"],content:`This week was AWS re:Invent, and I watched the keynote live-on Wednesday.
|
||
The three. hour. keynote.
|
||
During which Andy Jassy announced new features at a pace that is frankly astounding. Three hours should be too long for a keynote, and if I wasn’t watching from the comfort of my office, it would have been. Not only did the announcements keep unfolding for the full 270 minutes, but some didn’t even make it into the keynote. Running a three hour keynote is tough, creating enough new services and features to overflow a three hour keynote is amazing. My hat goes off to the engineering teams at AWS. It is truly staggering what you manage to accomplish each year.
|
||
There was one announcement that struck close to home. AWS announced Outposts towards the end of the keynote. In case you don’t want to read the whole marketing blurb on their website, let me give you the TL;DR. Outposts is an on-premises, hardware solution that you will be able to purchase via the AWS console and install in your datacenter. Andy Jassy said that the hardware in question is the same hardware they use in their AWS datacenters. The software running on this hardware? It will come in two flavors:
|
||
Running the same VMware on AWS software that they use to deliver VMware on AWS in the cloud. Running the same AWS software with a subset of the AWS services. The first one is like an olive branch to VMware. Andy Jassy is saying, “Yes Pat, we used to be mortal enemies. And now we’re best friends and we’re definitely not out to destroy you.”
|
||
No, definitely not. They are besties for life…
|
||
Except for that second option. That’s the takeover bid. The soldiers in the Trojan Horse. When VMware ruled the datacenter and AWS ruled the cloud, they could be friends. We are living in the Hybrid Cloud world after all. They needed to unite. Microsoft had a story that spanned the Hybrid Cloud, and AWS and VMware created one that did the same. The second option disrupts that uneasy alliance. You wouldn’t know it by all the smiles and back slapping we saw on stage, but if you checked the next day, there would be bruises from those slaps.
|
||
On your Outposts deployment you will be able to run a slice of AWS in your datacenter with EC2 and EBS to start. Naturally, that comes with the VPC construct to wrap around it. You can also start with as little as one node and scale out. The networking can use NSX to integrate with your Layer 2 network. The management of the solution will be done through the existing AWS console, or the standard APIs that are available in AWS today. AWS is fully managing the hardware and software updates and support of the box.
|
||
This is strikingly familiar to a product I’ve been working on for the last few years, Microsoft’s Azure Stack. Microsoft and analysts have predictably crowed about how Outposts is a validation of Azure Stack, and since Microsoft was there first, they’re going to win. The former is true, the latter is… less true. Just because you have the idea first, doesn’t mean that you win. Go ahead and talk to pets.com about that. Oh wait, you can’t. Right idea, wrong time and bad execution. (Pets.com I mean.)
|
||
There are three things that immediately jump out to me about the Outposts product:
|
||
You can start at one node. Azure Stack requires a minimum of four nodes (with the exception of the ASDK, which doesn’t count.) Of course, how much can you actually run on a single node? You know there won’t be any redundancy if the node fails. But I will point out that you can pack 96 cores, 2TB of RAM, and 1PB of storage into a single server these days. I have no idea what the specs are going to be on the nodes, but what is possible is quite staggering. AWS is owning the entire hardware and software stack. Azure Stack works with multiple OEMs. AWS is taking the Apple approach. There is definitely an advantage in owning the entire stack when it comes to consistency, predictability, and simplified testing. When a new Azure Stack update is released from Microsoft, the OEMs then need to verify the update on all of their system variants. This is trivial when the change is purely internal, like an updated API. This is incredibly complex when you want to introduce new networking capabilities that rely on hardware acceleration. AWS may be able to move faster since they don’t have the built-in lag of third party OEMs and validation. You can order this thing in the console. This might be the killer app. You want an Azure Stack? Well you’re going to need to work with one of the OEMs, and deal with a VAR, and go through procurement, and… I’m already exhausted. Outposts? Click, click BOOM. Outposts is the Saliva of on-prem, hybrid cloud. If there’s one thing that Amazon perfected, it’s low friction customer interaction. That’s great for when I need more cheezy-poofs. Maybe not so much for a rack of gear in my datacenter. There are also three issues that immediately jump out at me:
|
||
Integrating any solution into a datacenter is hard. There’s a reason Systems Integrators exist. I’ve added hardware to a lot of datacenters. In the best scenario, the customer has properly filled out the pre-flight checklist, and they have actually followed through on the to-do items. The more common scenario is that they don’t have the proper cables, rack space, power connectors, etc. And that’s just for the physical installation. Once you get to the software configuration, now you are dealing with their network, Active Directory, DNS, and time servers. Some or all of which may be incorrect on the checklist and not functioning properly. This is not unpacking an Amazon Echo and plugging it in. AWS is used to doing things at scale, where the logistics are totally different. Now they are directly shipping and supporting hardware for customers, and customers are sort of awful. That’s probably not fair. Some customers are great! They listen, they provide prompt responses, and they know their environment. It’s really the 80/20 rule. But that 20%? They are going to be the thing that makes you want to burn down the whole datacenter and walk away. My point though, is that AWS deals with servers in the thousands, not the ones. And like Andre the Giant said, “You have to use different moves when you are fighting half a dozen people, than when you have to worry about fighting only one.” That’s a long quote, but it was worth it. As Microsoft has discovered, the caveats of this type of solution are numerous and difficult. There’s a reason Azure Stack took so long to ship, and even now is still evolving. AWS does have the benefit of seeing what issues the Azure Stack team had to deal with, but they’ll still have to clear those same hurdles with their own solution. Azure Stack was in gestation for about three years with multiple technical previews. Scaling down Azure for a small number of nodes is hard. And when it first came out, it really only had IaaS solutions available. It wasn’t until this past year that the App Service and DB as a Service started being truly viable. AWS said that EC2 and EBS will be available at launch. Don’t expect anything else for a while, and don’t expect it to work well. Also dealing with code updates for remote boxes is difficult, and the API on your Outposts is going to be a different version than the public AWS regions. How are you going to deal with that? Will Outposts be successful? Yes. I have no doubt that this is the wave of the future. Microsoft has a head start on that wave, and they need to keep riding it. What about other competitors? I have a sneaking suspicion that Google Cloud will have something similar soon, and I’m not just talking about GKE on Nutanix. In fact, if I were Google Cloud, I would just buy Nutanix and have done with it.
|
||
Lastly, yes you can run VMware on this thing. No, don’t do that. I suspect 95% of Outposts that ship will be running the AWS variant. AWS has basically just eaten VMware’s breakfast, lunch, and dinner, and now it’s eyeing up the Baked Alaska that Pat ordered for dessert.
|
||
`,summary:`This week was AWS re:Invent, and I watched the keynote live-on Wednesday.
|
||
The three. hour. keynote.
|
||
During which Andy Jassy announced new features at a pace that is frankly astounding. Three hours should be too long for a keynote, and if I wasn’t watching from the comfort of my office, it would have been. Not only did the announcements keep unfolding for the full 270 minutes, but some didn’t even make it into the keynote.`,date:"30 Nov, 2018",url:"https://nedinthecloud.com/2018/11/30/aws-outposts-and-azure-stack/",image:"publiccloud.png",readingTime:"7"},"https://nedinthecloud.com/2018/11/23/what-happens-after-public-cloud/":{title:"What happens after public cloud?",tags:["aws","azure","cloud"],content:`AWS and Azure have won the public cloud race. Some might think that it’s too early to call it. But some might also be wrong. The fact is, AWS is the 10k lbs. gorilla in the market, and Azure is the alternative for those who don’t want to use AWS. We can argue about the potential technological superiority of Oracle Cloud’s claimed L2 networking, or the Machine Learning capabilities of Google Cloud. This is not an argument about technological superiority. We all know that the best product doesn’t always win. There are a lot of other factors involved, including things like: getting to the market first, removing friction for adoption, and being good enough. My point being, before you start telling me that the Cloud Spanner DB is better than Cosmos DB - just know that while you may be right, it doesn’t actually matter that you’re right.
|
||
Now that AWS and Azure have won the top seats in public cloud, I’ve started thinking about what is next. Is it possible for a competitor to come in and supplant one of these behemoths? Could someone just graduating from MIT be preparing to create a company that will displace AWS as king of the public cloud? Or will it be some other, hitherto, unknown technology that will bring a sea change to the tech industry - in the same way that public cloud is slowly replacing traditional datacenter architecture and approaches? I think we might be able to look back at the past to get a sense of what is lurking in our future.
|
||
A Little History
|
||
I’ve been doing this tech thing for a little while now. Not nearly as long as some, but long enough to see some massive shifts over the last 20-30 years. The first major shift was the move from mainframes to x86 servers. My first job had me supporting Windows NT boxes and an AS/400. The company was slowly shifting all of the software from the AS/400 to Windows NT 4.0, and the main drivers where obvious. The hardware was cheaper, the operating system was simpler to administrate, and it made applications more modular. It was a process of decoupling and dissaggregating the hardware from the software. Instead of specialized hardware, running specialized operating systems, and highly customized applications - there was now a way to purchase commodity hardware from one vendor, an operating system from another vendor, and an off-the-shelf application from another. This was revolutionary.
|
||
Six years later I watched the same process begin with VMware and virtualization. The main driver here was two-fold: a massive improvement in hardware utilization and a further dissaggregation of hardware from software. Now an operating system was no longer tightly coupled to the underlying hardware. The hardware had been abstracted by a hypervisor, and so the operating system could be shifted around in the datacenter, deployed as a template, and use specialized backup tools like snapshots. VMware’s introduction of VMotion changed the way that datacenters would function forever. This was revolutionary as well.
|
||
While VMware was strutting its stuff all around town, AWS released their first public cloud offering, the Simple Queue Service. And… most people didn’t really care. Some developers did though, and if you were looking for a public service to provide queuing for your applications, SQS was there to take your money. Things really started moving with the introduction of the Simple Storage Service (S3), which provided almost infinite storage of items in the cloud and the ability to host a static website. Some of the final pieces of the puzzle were Elastic Compute Cloud (EC2) and the Virtual Private Cloud (VPC). It was now possible to use someone else’s datacenter to host your workloads, without drastic modifications to your application. The management and monitoring of the underlying resources was no longer your problem, and the capital expenditure required to start a new tech company became effectively zero.
|
||
And that leads us to where we are today. Public cloud is all the rage. AWS was first to market with viable services that were easy to consume. Microsoft eventually came around to the idea with Azure, and then went ALL IN on being a cloud company. They are the one and two, and at this point everyone else is an also-ran.
|
||
King of the Hill
|
||
Being the biggest and baddest in the industry only works so long as the industry stays where you are. Ask IBM about AIX and their big iron. They still have revenue coming in from those business units, but it’s a far cry from when they were the King of the Hill. They ceded their crown to the x86 platform and Windows and Linux operating systems. An abstraction was created between the hardware and the operating system, and while IBM continued to produce hardware, they no longer owned the whole stack. For a long time Microsoft was high on its horse, but eventually VMware came along and was able to create a unique position in the market. Instead of trying to fight Microsoft for dominance of the operating system, instead they created a new abstraction layer in the stack and staked their indelible claim on it. There’s also KVM, and Xen, and Hyper-V, but let’s face it. VMware is the king of the virtualization hill. And then along came public cloud, and they did exactly what VMware had done. Instead of fighting VMware on their turf, they created a new battleground and staked their claim.
|
||
The next company to displace AWS as King of the Hill will need to find a way to switch up the game once again. They will need to find a way to change the value proposition to favor their solution either as a replacement to public cloud - in the same way that x86 replaced big iron - or as a new layer of abstraction to ride the public cloud - in the way that VMware created a new layer with virtualization. What could that look like?
|
||
Kubernetes?
|
||
When containers started gaining prominence with Docker, I thought to myself that we were witnessing VMware all over again. This was another abstraction to further divorce applications from the underlying hardware and operating systems they run on. Still, while containers were great for developers, they were and are a nightmare for operations folks. Enter Kubernetes and a host of other technologies aimed at operationalizing the deployment and management of containers. These technologies create a new layer in the stack, rather that replacing a component of the stack with a new technology. Containers and Kubernetes are available in all the major public clouds and in your datacenter. They aren’t going to replace operating systems, virtualization, or physical hardware platforms. It’s another layer, and in that regard they will not be what displaces AWS and public cloud.
|
||
The Value
|
||
What is the value that public cloud brings to an organization? There are several, but I think it mostly comes back to the idea that running a datacenter and services to support your software is usually not a key differentiator or core competency for most companies. That’s not to say that they cannot do these things, it simply means that it does not create significant competitive advantages for the company or provide some essential distinction that makes them more attractive than other companies in their market. Let’s say you are running a real estate agency, and your office workers and clients rely on a web application to list and review properties. Does it create significant value if you host that application in your own private datacenter or in AWS? Chances are the answer is no. Is your IT staff more efficient than a public cloud datacenter? Again, the answer is likely no. Is you datacenter infrastructure more reliable or secure than the public cloud? No. I’m just going to flat out say, it is not.
|
||
So what is the key differentiator about your web application? It’s the application itself. It makes more sense to focus your time and money on creating the best possible application for your clients and employees, than spending time running a datacenter. Build what makes you different, buy what doesn’t. That’s some solid advice from Geoffrey Moore, and it’s the main reason that AWS has become such a Goliath in the industry.
|
||
How can a new company change the value proposition so that public cloud isn’t the answer? By creating a platform that makes it even easier to buy those things that don’t provide a competitive advantage for a company. And by solving a problem that public clouds are struggling with today. What is that problem? One is data locality.
|
||
A Thought Experiment
|
||
I think I’ve spent enough words framing the situation. If a company wants to displace AWS, they are going to need to provide a better value proposition, make adoption seamless, and solve an issue companies are facing today that AWS cannot solve. What could such a solution be. Well here’s one.
|
||
Imagine that there’s a company called Quantum Entanglement Data (QED) that was just started by some grad students from China. They have been working on a technology to create quantum entangled particles that can stay in synchronization and transmit data instantaneously over large distances. Let’s suspend disbelief for a moment, after all this is something that should be theoretically possible. The group has found a way to not only get the technology working, but also a way to produce it reliably at scale for a relatively low cost. The bandwidth of each QED-bit is only 10Mb/s, but they are working on running multiple QED-bits in parallel to reach bandwidth speeds of up to 10Gb/s in two years. This is a technology that is going to revolutionize everything.
|
||
Why would this be such a massive breakthrough? Because of data gravity. Moving data is hard, and moving lots of data is harder. As a result, applications tend to move towards where the data is being stored, as if the data had a gravitational pull. Why is it hard to move data? Mostly because the lanes in which data can travel slow down significantly as they change mediums. The fastest data exchange happens between the processor and the L1 cache, where things can hit over 1TB/s with nanosecond latency. Zoom out to the SAS bus where most SSDs live and now you’re looking at something closer 12GB/s and 10x the latency. A network link in a datacenter is probably hitting 100Gb/s (which is roughly 12GB/s) with latency measured in the low millisecond range. Move outside of the datacenter and now you’re look at lines running at 10Gb/s if you’re lucky, and more likely in the 100Mb/s range. Latency now becomes a function of distance - you can’t make light travel faster - and how many devices have to process the signal. Over a hundred milliseconds wouldn’t be unreasonable.
|
||
The technology that QED is talking about would remove the latency - remember, these are quantum entangled particles. And the more QED-bits that run in parallel, the more bandwidth is available to move data. This would lead to a situation where distance is no longer a major barrier to moving data and accessing data. An application can be massively distributed, with no concern on data access latency to a back-end database. Likewise, a back-end database could replicate itself synchronously between multiple locations, which if you didn’t already know, is a helluva a challenge. If your computer had a QED-bit in its storage sub-system, you could have a corresponding bit in multiple datacenters and have instant access to any of those data sources with no latency.
|
||
QED would have invented a technology that completely changes the technology landscape. They could patent their technology, build datacenters, and start providing service. The whole paradigm of public cloud would be shifted to the edge overnight. Decentralized applications - already an important movement today - would explode along with the integration of QED-bits into devices. The need for reliable connectivity through cellular lines, cable internet, fiber optics, etc. would be removed. The need for massive datacenters full of storage and compute would seem like putting all one’s eggs in a single basket creating unnecessary risks. The change wouldn’t happen overnight, but it would completely change the tech industry over the course of just a few years.
|
||
All About the Benjamins That’s just a quick example. I have no special crystal ball that tells me what major tech innovation is coming next. If I did, I would be a rich investor and not a humble blogger. What I do know is that whatever the next sea change is, it will be driven by changing the value proposition for technology. If I were trying to move the needle, I would look for problems that the public cloud is struggling to solve today, and develop an innovative solution that creates more value for public cloud consumers than what they are doing today.
|
||
I don’t think that change is too far away either. The public cloud has risen to prominence in the last two or three years, and in the next five years there will be some type of technology that changes the game. It will probably start slow, like SQS or GSX, and then - like the proverbial snowball - grow exponentially until it displaces public cloud or subsumes it. I kind of hope it will be something like the QED-bit, but I doubt we’ll be that lucky.
|
||
`,summary:"AWS and Azure have won the public cloud race. Some might think that it’s too early to call it. But some might also be wrong. The fact is, AWS is the 10k lbs. gorilla in the market, and Azure is the alternative for those who don’t want to use AWS. We can argue about the potential technological superiority of Oracle Cloud’s claimed L2 networking, or the Machine Learning capabilities of Google Cloud.",date:"23 Nov, 2018",url:"https://nedinthecloud.com/2018/11/23/what-happens-after-public-cloud/",image:"publiccloud.png",readingTime:"11"},"https://nedinthecloud.com/2018/10/05/terraform-fotd-wrap-up/":{title:"Terraform - FotD - Wrap Up",tags:["fotd","hashicorp","terraform"],content:`This is the final post of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I would like to take a moment to review the project as a whole, mung together some semi-coherent thoughts about the functions, and maybe even make a plan for the future.
|
||
Thoughts on the project In June of this year, I undertook a project to examine each built-in interpolation function within Terraform. My plan was to examine a function each day - weekends excluded - and document how the function works, why it might be useful, and any interesting tidbits I discovered along the way. There were a few drivers behind the project:
|
||
I was not consistently blogging, and taking on a daily blogging project would force me to do so. There were several functions that I did not understand in Terraform, and this seemed like a good way to learn. I’ve got a couple courses on Pluralsight dealing with Terraform and I thought this would help when I revise them. I’d like to say that I had thought this whole process through and started my first post with the whole series and format laid out. I’d like to say that, but it wouldn’t be the truth. The first few posts had me working through the format of the post and the way in which I wanted to create examples. It also took me a few iterations to figure out how I wanted to structure the GitHub repo, and how to effectively create new examples with minimum effort. After the first few posts I got into a groove. I also realized that this was going to be a significant undertaking in terms of time. When I started the posts, there were about 63 functions, and now there are 65. Working at five functions a week, the whole thing should have taken about 13 weeks to complete. That’s three months of posting, five days a week. Considering I started in June and finished in September, three months was exactly right. It wasn’t till the end of the first week that I realized that I had made a three month commitment, and once I did realize that, I started to look for ways to make the process sustainable. Using a basic template for posting and examples helped a lot. I also didn’t get too ambitious with the number of examples for each function. I wouldn’t say that the process was easy, and I don’t know if I would do it again. But I will say that I learned a lot about Terraform and blogging.
|
||
Thoughts on Terraform functions If I could start the whole project over again, I would have grouped the functions by what they do, and not alphabetically. It makes sense to put all the functions dealing with strings together, and same with lists and maps. Then there are the functions that take no arguments. Basically, if you are looking for a function to assist with string manipulation, then it would be much more helpful if all of those types of functions were grouped together. So here’s my attempt at doing exactly that:
|
||
Terraform Interpolation Function Cheat Sheet String functions chomp coalesce format indent lower replace split substr title trimspace upper Number and date functions abs ceil floor log max min pow signum timestamp timeadd List functions chunklist coalescelist compact concat contains distinct element flatten formatlist index join length list matchkeys slice sort Map functions keys lookup map merge transpose values zipmap File and path functions basename dirname file pathexpand Network functions cidrhost cidrnetmask cidrsubnet Security and web functions base64decode base64encode base64gzip base64sha256 base64sha512 bcrypt jsonencode md5 rsadecrypt sha1 sha256 sha512 urlencode uuid Thoughts on Terraform and the future While writing all of these posts, I did discover some things I would like to see changed and improved. For instance, the documentation page for the built-in functions is listed alphabetically for the most part. But there are a few examples of it not being entirely alphabetical. That’s a minor thing to most. As the son of a librarian, that’s the type of thing that really gets stuck in my craw. More importantly, I think the functions should be grouped by loose affiliation; the way I just listed them out. There were also a few functions that I thought were either missing functionality, or could be improved in some way. I don’t want to be the person just complaining about things, I want to be part of the solution as well. The whole project is open source, including the docs. The website is written in markdown, and I can easily update the page with the built-in functions and submit a pull request. In fact, I plan to do that. The functions are written in Golang, and that makes things a little trickier since I am not a developer by trade. To that end, I have started learning Golang in order to assist with improving the functions. My first attempt was adding a function called dateadd that would add the ability to manipulate dates in the same way that timeadd deals with time. After submitting my pull request, I had some good conversations with one of the maintainers about whether that should be a separate function or added to the timeadd function. We agreed on adding it to the timeadd function and I am working on learning enough Go to make that a reality.
|
||
Terraform v0.12 will be dropping in the near future. It is going to bring with it a lot of enhancements to the Hashicorp Configuration Language, and the way that Terraform functions overall. There will be breaking changes, but it will all be in the name of creating a better product. Terraform is becoming more widely adopted by the moment, and it’s probably best to make the necessary changes now before this juggernaut picks up any more steam. I’m sure once the new version drops I will need to revise my courses on Pluralsight and maybe review some new functions here. I’m excited for the future and what it might bring!
|
||
`,summary:`This is the final post of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I would like to take a moment to review the project as a whole, mung together some semi-coherent thoughts about the functions, and maybe even make a plan for the future.
|
||
Thoughts on the project In June of this year, I undertook a project to examine each built-in interpolation function within Terraform.`,date:"5 Oct, 2018",url:"https://nedinthecloud.com/2018/10/05/terraform-fotd-wrap-up/",image:"tutorials.png",readingTime:"5"},"https://nedinthecloud.com/2018/09/21/terraform-fotd-zipmap/":{title:"Terraform - FotD - zipmap()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the zipmap() function. The example file is on GitHub here.
|
||
What is it? Function name: zipmap(list, list)
|
||
Returns: Takes two lists of equal length and returns a map with the first list as keys and the second list as values.
|
||
Example:
|
||
# Returns { "k1":"v1" "k2":"v2" } output "zipmap_output" { value = "\${zipmap(list("k1","k2"), list("v1","v2"))}" } Example file: ############################################## # Function: zipmap ############################################## ############################################## # Variables ############################################## variable "list_1" { type = "list" default = [["n1v1","n1v2"],["n2v1","n2v2"]] } variable "map_1" { type = "map" default = { "k1" = "v1" "k2" = "v2" } } variable "list_2" { default = [[], [], []] } variable "list_3" { default = [false, false, false] } variable "list_4" { default = ["So", "Long", "and", "Thanks"] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_standard_list" { value = "\${zipmap(list("k1","k2","k3"),list("v1","v2","v3"))}" } output "2_empty_values" { value = "\${zipmap(var.list_3, var.list_2)}" } output "3_nest_list_values" { value = "\${zipmap(list("k1","k2"), var.list_1)}" } output "4_nest_map_values" { value = "\${zipmap(list("k1","k2"), list(var.map_1,var.map_1))}" } output "5_nest_mixed_lengths" { value = "\${zipmap(list("k1","k2"), list(var.list_4,var.list_3))}" } Run the following from the zipmap folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? You’ve got a list and you want a map. Or more likely, you have strings that represent lists and maps from the output of a module. Since modules currently only support a string output, you might use the split function to get the list of keys and list of values. Then use zipmap to create the map from that. You could also dynamically create maps from separate data sources that return lists.
|
||
Lessons Learned The keys and values lists have to be the same length. The function will allow both nested lists and nested maps for the values, but only a list of strings for the keys. Not a surprise there. There can be a different number of elements in the nested lists. You cannot mix and match maps, lists, and strings for the values.
|
||
Coming up next is no function! That’s it folks. My next post will be a wrap up of this whole FotD series.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the zipmap() function. The example file is on GitHub here.
|
||
What is it? Function name: zipmap(list, list)
|
||
Returns: Takes two lists of equal length and returns a map with the first list as keys and the second list as values.`,date:"21 Sep, 2018",url:"https://nedinthecloud.com/2018/09/21/terraform-fotd-zipmap/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/09/20/terraform-fotd-values/":{title:"Terraform - FotD - values()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the values() function. The example file is on GitHub here.
|
||
What is it? Function name: values(map)
|
||
Returns: Takes a map and returns the values in the order the keys appear. The values must be strings, not lists or maps.
|
||
Example:
|
||
variable "map" { type = "map" default = { "one" = "1" "two" = "2" } } # Returns ["1","2"] output "value_output" { value = "\${values(var.map)}" } Example file: ############################################## # Function: values ############################################## ############################################## # Variables ############################################## variable "map_value" { type = "map" default = { "life" = "42" "universe" = "6" "everything" = "7" } } variable "empty_map" { type = "map" default = {} } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_map_value_output" { value = "\${values(var.map_value)}" } #String values only #output "2_map_function_output" { # value = "\${values(map("life",list("42"),"the universe",list("six","times","seven")))}" #} output "3_empty_map_output" { value = "\${values(var.empty_map)}" } Run the following from the values folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? Usually, you would be more interested in knowing the keys in a map, and then getting the respective values. But there are times you might just want all the values. The returned list may have some duplicates in it, so using the distinct function can help remove those. If you are trying to flip the values and keys, then you can use this function along with keys and zipmap to create a reversed map. You might think the transpose function would do that. But you’d be wrong.
|
||
Lessons Learned The function does exactly as advertised. It doesn’t mind an empty map, but all the values have to be strings. If the value is a map or list, then the function will throw an error.
|
||
Coming up next is the zipmap() function. And that will be the final function of this series. Wow!
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the values() function. The example file is on GitHub here.
|
||
What is it? Function name: values(map)
|
||
Returns: Takes a map and returns the values in the order the keys appear. The values must be strings, not lists or maps.`,date:"20 Sep, 2018",url:"https://nedinthecloud.com/2018/09/20/terraform-fotd-values/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/09/19/terraform-fotd-uuid/":{title:"Terraform - FotD - uuid()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the uuid() function. The example file is on GitHub here.
|
||
What is it? Function name: uuid()
|
||
Returns: Takes no arguments. Returns a string with a UUID conforming to RFC 4122 v4.
|
||
Example:
|
||
# Returns something like XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX output "uuid_output" { value = "\${uuid()}" } Example file: ############################################## # Function: uuid ############################################## ############################################## # Variables ############################################## ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_uuid_basic" { value = "\${uuid()}" } Run the following from the uuid folder to get example output for a number of different cases:
|
||
#No arguments for function terraform apply Why use it? Having a UUID is going to be pretty useful when generating new resources. Especially if the provider for those resources doesn’t have a UUID generation mechanism. AWS resources all have an identifier associated with them, but that might not be the case everywhere. The important thing to remember is that this function will generate a new, unique ID each time it is run. And that means that if you run it against the same resource it would change the ID each time. That’s why you’ll need to add the ignore_changes lifecycle attribute to prevent that behavior.
|
||
Lessons Learned The official docs don’t actually explain what RFC 4122 is, or how the UUID is constructed. For that, you can turn to the RFC itself. The UUID is a 128-bit value generated by an algorithm in order to be unique across time and space. It also doesn’t need a central authority to register each ID with, which simplifies the creation of UUIDs by a lot. 2 ^ 128 potential values works out to about 3.4 X 10^38. That’s the same address space being used by IPv6. So like, that’s a lot of values. The output from Terraform formats the value as:
|
||
32 bits - 16 bits - 16 bits - 16 bits - 48 bits
|
||
I won’t go into too much more detail, but basically the fields correspond to the following:
|
||
time values - time values - time values - clock sequence - node value
|
||
The actual calculation of the values depends on the version in use. Version 4 is meant for creating unique values from psuedo-random numbers. The whole RFC was pretty fascinating, I recommend reading it!
|
||
Coming up next is the values() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the uuid() function. The example file is on GitHub here.
|
||
What is it? Function name: uuid()
|
||
Returns: Takes no arguments. Returns a string with a UUID conforming to RFC 4122 v4.
|
||
Example:
|
||
# Returns something like XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX output "uuid_output" { value = "\${uuid()}" } Example file: ############################################## # Function: uuid ############################################## ############################################## # Variables ############################################## ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_uuid_basic" { value = "\${uuid()}" } Run the following from the uuid folder to get example output for a number of different cases:`,date:"19 Sep, 2018",url:"https://nedinthecloud.com/2018/09/19/terraform-fotd-uuid/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/09/18/terraform-fotd-urlencode/":{title:"Terraform - FotD - urlencode()",tags:["fotd","hashicorp","terraform"],content:"This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the urlencode() function. The example file is on GitHub here.\nWhat is it? Function name: urlencode(string)\nReturns: Takes a string and returns the a URL-safe copy of the string.\nExample:\nvariable "string" { default = "My Text" } # Returns "My+Text" output "urlencode_output" { value = "${urlencode(var.string)}" } Example file: ############################################## # Function: urlencode ############################################## ############################################## # Variables ############################################## variable "urlencode" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "urlencode_output" { value = "${urlencode(var.urlencode)}" } Run the following from the urlencode folder to get example output for a number of different cases:\n#We have to get something to encode first $fileContent = Get-Content .\\textFile.txt terraform apply -var "urlencode=$fileContent" #Test out some special characters terraform apply -var "urlencode=Here's a string: Bistromatically speaking I have a $40 check, with 20% gratuity, (don't forget the -2^e for service) & #1 for good measure." #More fun $testString = 'HI!@#$%^&*()_+-=;,./?\\|`~BYE' terraform apply -var "urlencode=$testString" #Empty string test terraform apply -var "urlencode=" Why use it? Better safe than sorry I always say. If you’re trying to hit up a URL using the HTTP or External providers, it’s probably best to sanitize that input first. The urlencode function will do exactly that.\nLessons Learned My first question when I read the official docs was, “What is a URL-safe string?” As always the answer is to check Wikipedia. The actual standard that appears to be in use is the application/x-www-form-urlencoded standard. You can tell since the space character is encoded as a + instead of the more common %20. I suppose this makes it safe for use in form data submissions as well.\nComing up next is the uuid() function.\n",summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the urlencode() function. The example file is on GitHub here.
|
||
What is it? Function name: urlencode(string)
|
||
Returns: Takes a string and returns the a URL-safe copy of the string.
|
||
Example:
|
||
variable "string" { default = "My Text" } # Returns "My+Text" output "urlencode_output" { value = "\${urlencode(var.`,date:"18 Sep, 2018",url:"https://nedinthecloud.com/2018/09/18/terraform-fotd-urlencode/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/09/17/terraform-fotd-upper/":{title:"Terraform - FotD -upper()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the upper() function. The example file is on GitHub here.
|
||
What is it? Function name: upper(string)
|
||
Returns: Takes a string and returns the uppercase version of all characters in the string as a string.
|
||
Example:
|
||
variable "string" { default = "all lower" } # Returns "ALL LOWER" output "upper_output" { value = "\${upper(var.string)}" } Example file: ############################################## # Function: upper ############################################## ############################################## # Variables ############################################## variable "upper" { default = "This Line Is Capitalized." } variable "sourcefile" { default = "textFile.txt" } ############################################## # Resources ############################################## data "local_file" "source" { filename = "\${var.sourcefile}" } ############################################## # Outputs ############################################## output "upper_output" { value = "\${upper(var.upper)}" } output "file_output" { value = "\${upper(data.local_file.source.content)}" } Run the following from the upper folder to get example output for a number of different cases:
|
||
#We have to get the local resource first terraform init #Default values terraform apply #All caps terraform apply -var "upper=THIS IS IN ALL CAPITALS" #All lowercase terraform apply -var "upper=this is in all lowercase" #Empty string terraform apply -var "upper=" #All standard US-EN characters terraform apply -var 'upper=qwertyuiopasdfghjklzxcvbnm1234567890,./;[]\\<>?:"{}|~!@#$%^&*()_+-=\`' Why use it? I had some really good ideas on why to use the lower function. There are many cases of resources that will not accept uppercase characters as valid input, so lower makes sense. I can’t think of a single service that uses all uppercase. Maybe you use this for metadata tagging to really bring attention to something. Or for your SQL statement, which seem to always want to be in uppercase. Or this was just included for completeness.
|
||
Lessons Learned The function does precisely what you would expect. I only tested it with standard EN-US characters, so your mileage could vary. It should be able to do uppercase translation for any character in the Unicode set. There are 1,781 uppercase characters in Unicode version 11. If you’d like to test all of them, have fun! Personally I will trust the Go function that is being called here to do the right thing.
|
||
Coming up next is the urlencode() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the upper() function. The example file is on GitHub here.
|
||
What is it? Function name: upper(string)
|
||
Returns: Takes a string and returns the uppercase version of all characters in the string as a string.
|
||
Example:
|
||
variable "string" { default = "all lower" } # Returns "ALL LOWER" output "upper_output" { value = "\${upper(var.`,date:"17 Sep, 2018",url:"https://nedinthecloud.com/2018/09/17/terraform-fotd-upper/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/09/14/terraform-fotd-trimspace/":{title:"Terraform - FotD - trimspace()",tags:["fotd","hashicorp","terraform"],content:"This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the trimspace() function. The example file is on GitHub here.\nWhat is it? Function name: trimspace(string)\nReturns: Takes a string and returns the string with any trailing whitespace removed\nExample:\n# Returns "String" output "trimspace_output" { value = "${trimspace("String ")}" } Example file: ############################################## # Function: trimspace ############################################## ############################################## # Variables ############################################## variable "trimspace" { default = "A string with tab \\t" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_trimspace_output" { value = "${trimspace(var.trimspace)}" } output "2_length" { value = "${length(var.trimspace)}" } output "3_length_trimspace" { value = "${length(trimspace(var.trimspace))}" } Run the following from the trimspace folder to get example output for a number of different cases:\n#Test with newline $mystring = "String with one newline `n" terraform apply -var "trimspace=$mystring" #Just spaces $mystring = " " terraform apply -var "trimspace=$mystring" #Empty String terraform apply -var "trimspace=" #Spaces after newline $mystring = "String with one newline `nNext line " terraform apply -var "trimspace=$mystring" Why use it? Sanitized data is happy data! The same reasons you might want to use the chomp function apply to this function as well. Trailing spaces can reek havoc when that value is passed to another resource for use. Generally speaking, don’t trust anything from an end user or an outside data source.\nLessons Learned It’s pretty hard to tell if the trimspace function has done anything, since the string looks the same in the output. I ended up using the length function to measure the string before and after trimspace was applied to make sure it was doing something. Fun fact, length says that it is only for lists, but it will give you the length of a string as well. I guess a string is just a list of characters after all. A tab and newline are also considered white space characters, but only if the newline is at the end of the string. Otherwise you’re not really at the end of the string are you? I tried passing the newline and tab characters from the command line with no success. That might be a consequence of using Windows. Instead I ended up storing the value in a PowerShell variable and then passing it to Terraform.\nComing up next is the upper() function.\n",summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the trimspace() function. The example file is on GitHub here.
|
||
What is it? Function name: trimspace(string)
|
||
Returns: Takes a string and returns the string with any trailing whitespace removed
|
||
Example:
|
||
# Returns "String" output "trimspace_output" { value = "\${trimspace("String ")}" } Example file: ############################################## # Function: trimspace ############################################## ############################################## # Variables ############################################## variable "trimspace" { default = "A string with tab \\t" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_trimspace_output" { value = "\${trimspace(var.`,date:"14 Sep, 2018",url:"https://nedinthecloud.com/2018/09/14/terraform-fotd-trimspace/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/09/13/terraform-fotd-transpose/":{title:"Terraform - FotD - transpose()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the transpose() function. The example file is on GitHub here.
|
||
What is it? Function name: transpose(map)
|
||
Returns: Takes a map with lists of strings as the values. Each unique string in the lists is turned into a key. Each key has a list of strings as a value. The list of strings are any keys where the string was present in the original map. Returns a map with lists of strings as values. That is confusing, and I acknowledge it. See below in the examples for some clarification.
|
||
Example:
|
||
# Returns {Blue = [Green Purple], Red = [Orange Purple], Yellow = [Green Orange]} output "transpose_output" { value = "\${transpose(map("Purple",list("Red","Blue"),"Orange",list("Yellow","Red"),"Green",list("Yellow","Blue")))}" } Example file: ############################################## # Function: transpose ############################################## ############################################## # Variables ############################################## variable "map_value" { type = "map" default = { "app_servers" = ["Ford","Arthur","Zaphod"] "db_servers" = ["Marvin","Trillian"] "web_servers" = ["Ford","Marvin"] } } variable "empty_map" { type = "map" default = {} } variable "second_map" { type = "map" default = { "four" = [1,2,4] "six" = [1,2,3,6] "eight" = [1,2,4,8] "ten" = [1,2,5,10] } } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_map_value_output" { value = "\${transpose(var.map_value)}" } output "2_map_function_output" { value = "\${transpose(map("Purple",list("Red","Blue"),"Orange",list("Yellow","Red"),"Green",list("Yellow","Blue")))}" } output "3_second_map_value_output" { value = "\${transpose(var.second_map)}" } output "4_empty_map" { value = "\${transpose(var.empty_map)}" } Run the following from the transpose folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? I struggled with this one initially, and then I started to see how this could be useful with tags or groupings. The example I used above with the various types of servers shows how if I wanted to know what roles a particular server had, I could transpose the map of roles to servers. Now I know that Marvin is both a web server and a DB server. That’s a simple example, but I imagine there are lots of times where flipping the grouping could be helpful.
|
||
Lessons Learned At first I was so confused by the documentation, I thought that the function was totally useless. It wasn’t till I tried to actually use it that things clicked. The map has to have a list type as values. It’s also okay with an empty map, since there’s nothing to do. I thought this function would just swap the keys and values in a map, but it doesn’t do that. If you were looking to take a map of keys and values that are just strings, then you could use the keys and values functions and then zipmap to put them back together in a transposed way. This transpose function is more subtle that that.
|
||
Coming up next is the trimspace() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the transpose() function. The example file is on GitHub here.
|
||
What is it? Function name: transpose(map)
|
||
Returns: Takes a map with lists of strings as the values. Each unique string in the lists is turned into a key. Each key has a list of strings as a value.`,date:"13 Sep, 2018",url:"https://nedinthecloud.com/2018/09/13/terraform-fotd-transpose/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/09/11/terraform-fotd-timeadd/":{title:"Terraform - FotD - timeadd()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the timeadd() function. The example file is on GitHub here.
|
||
What is it? Function name: timeadd(time, duration)
|
||
Returns: Takes a date time value and a duration value. Returns the date value plus the duration value. All values should conform to the RFC 3339 format.
|
||
Example:
|
||
# Returns 2018-09-10T21:00:00Z output "timeadd_output" { value = "\${timeadd("2018-09-10T20:00:00Z","1h")}" } Example file: ############################################## # Function: timeadd ############################################## ############################################## # Variables ############################################## variable "date" { default = "1970-01-01T00:00:00Z" } variable "add" { default = "1h" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_timeadd_basic" { value = "\${timeadd(var.date, var.add)}" } Run the following from the timeadd folder to get example output for a number of different cases:
|
||
# For your reference # ns = nanosecond # ms = millisecond # s = second # m = minute # h = hour # There is nothing bigger than an hour #Defaults terraform apply #Negative duration terraform apply -var="add=-1h" #Zero duration terraform apply -var="add=0h" #Multiple durations terraform apply -var="add=1h10m" #My favorite day, adding a year terraform apply -var="add=8760h" -var="date=2018-03-14T01:59:27Z" Why use it? I imagine using this to schedule a task in the future or retrieve a log from an hour ago. I’ve never had an occasion to use this function personally, so I don’t know what the most common use cases are. Basic date and time manipulation does seem like a fairly common task though.
|
||
Lessons Learned One thing that bothered me is that the documentation references RFC 3339 for its formatting and provides a link to the format, but it doesn’t provide a link or reference what values are accepted for the duration argument. Through a bit of trial and error I found the accepted values in the example above: ns, ms, s, m, h. I assumed that there would be something to alter the day, month, and year, but there is not. Why? Well, the actual function in Terraform makes use of the Golang time package, specifically the Add function for time type variables. Digging into the Golang docs, I found that the Add function uses a Duration value, which is a type with a constructor that takes the following values: nano-second, micro-second, millisecond, second, minute, and hour. And that is because the Add function is only dealing with Time, not Dates. There is a separate function called AddDate that handles manipulating Dates. There should probably be a corresponding dateadd function in Terraform to support that.
|
||
Coming up next is the title() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the timeadd() function. The example file is on GitHub here.
|
||
What is it? Function name: timeadd(time, duration)
|
||
Returns: Takes a date time value and a duration value. Returns the date value plus the duration value. All values should conform to the RFC 3339 format.`,date:"11 Sep, 2018",url:"https://nedinthecloud.com/2018/09/11/terraform-fotd-timeadd/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/09/11/terraform-fotd-title/":{title:"Terraform - FotD - title()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the title() function. The example file is on GitHub here.
|
||
What is it? Function name: title(string)
|
||
Returns: Takes a string and returns that string with the first letter of every word capitalized
|
||
Example:
|
||
# Vogon Jeltz Is My Hero output "title_output" { value = "\${title("Vogon Jeltz is my hero")}" } Example file: ############################################## # Function: title ############################################## ############################################## # Variables ############################################## variable "source" { default = "don't panic" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_title_basic" { value = "\${title(var.source)}" } Run the following from the title folder to get example output for a number of different cases:
|
||
#Default terraform apply #Inverse Caps, capitalizes everything terraform apply -var "source=oNCE uPON a tIME" #all lower case terraform apply -var "source=the restaurant at the end of the universe" #empty string terraform apply -var "source=" Why use it? I guess you would use this to make things look nice for cases where putting something in Titlecase matters. Honestly, there are some major issues with this function.
|
||
Lessons Learned The function does exactly what you would expect, mostly. It counts an apostrophe as the beginning of a new word. If you run it against don’t panic, you get back Don’T Panic. I checked a few others and a dash ‘-’ is also consider a word separator, so state-of-the-art becomes State-Of-The-Art. Those are probably both wrong. A larger issue is that Titlecase is a real thing that has rules. And this function does not follow those rules. If a developer is going to add a function that does Titlecase, then it should do it reasonably well. Or don’t add it. In all fairness, this function uses the Golang title function in the strings package to do the work. And that function has a bug:
|
||
BUG(rsc): The rule Title uses for word boundaries does not handle Unicode punctuation properly.
|
||
Still, I think if you have a function that is trying to provide Titlecase, it should do that. Or at least know that an apostrophe is not a word separator.
|
||
Coming up next is the transpose() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the title() function. The example file is on GitHub here.
|
||
What is it? Function name: title(string)
|
||
Returns: Takes a string and returns that string with the first letter of every word capitalized
|
||
Example:
|
||
# Vogon Jeltz Is My Hero output "title_output" { value = "\${title("Vogon Jeltz is my hero")}" } Example file: ############################################## # Function: title ############################################## ############################################## # Variables ############################################## variable "source" { default = "don't panic" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_title_basic" { value = "\${title(var.`,date:"11 Sep, 2018",url:"https://nedinthecloud.com/2018/09/11/terraform-fotd-title/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/09/07/terraform-fotd-substr/":{title:"Terraform - FotD - substr()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the substr() function. The example file is on GitHub here.
|
||
What is it? Function name: substr(string, offset, length)
|
||
Returns: Takes a string, offset, and length. Returns a string as long as the length value and starting at the offset value.
|
||
Example:
|
||
# Returns abc output "substr_output" { value = "\${substr("abcdefg",0,3)}" } Example file: ############################################## # Function: substr ############################################## ############################################## # Variables ############################################## variable "source" { default = "Don't Panic" } variable "offset" {} variable "length" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_substr_basic" { value = "\${substr(var.source,var.offset,var.length)}" } Run the following from the substr folder to get example output for a number of different cases:
|
||
#Standard substr terraform apply -var "offset=0" -var "length=5" #Both 0s terraform apply -var "offset=0" -var "length=0" #Negative offset terraform apply -var "offset=-5" -var "length=5" #Negative length terraform apply -var "offset=5" -var "length=-1" #Both Negative terraform apply -var "offset=-5" -var "length=-1" Why use it? There are plenty of reasons why you might want a portion of a string. Take a hostname like prod-web-001. Maybe you need the server type, or the environment, or the number of the server. Assuming your name formatting is consistent, you can always use substr to extract the information you are looking for. There’s probably a ton of other examples. This is just the first that came to mind.
|
||
Lessons Learned As the documentation explains, a negative offset starts at the last character and moves to the left. If a string is 10 characters long, then a -2 would start at the 9th character. This is great if you don’t know the length ahead of time. There is no length function for a string in Terraform, so the only way to know is either turn the string into a list and use the actual length function, or use a negative offset. A length of -1 corresponds to the end of the string, regardless of the length of the actual string. Don’t try and use a -2 for length, as it will not work. As in the previous example of a hostname, if you know the last three characters are the server number, then you can use an offset of -3 and a length of -1 to get the numbers consistently regardless of the length of the string. I have to say I really like the way this function works. You can tell someone had been frustrated by other substr functions and looked for a way to make things easier for the user.
|
||
Coming up next is the timestamp() function. It takes no arguments at all, because it is timeless.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the substr() function. The example file is on GitHub here.
|
||
What is it? Function name: substr(string, offset, length)
|
||
Returns: Takes a string, offset, and length. Returns a string as long as the length value and starting at the offset value.`,date:"7 Sep, 2018",url:"https://nedinthecloud.com/2018/09/07/terraform-fotd-substr/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/09/07/terraform-fotd-timestamp/":{title:"Terraform - FotD - timestamp()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the timestamp() function. The example file is on GitHub here.
|
||
What is it? Function name: timestamp()
|
||
Returns: Takes no arguments. Returns the current time in UTC using the RFC 3339 standard.
|
||
Example:
|
||
# Returns something like 2018-09-05T00:00:00Z output "timestamp_output" { value = "\${timestamp()}" } Example file: ############################################## # Function: timestamp ############################################## ############################################## # Variables ############################################## ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_timestamp_basic" { value = "\${timestamp()}" } Run the following from the timestamp folder to get example output for a number of different cases:
|
||
#No arguments for function terraform apply Why use it? What time is it? 4:30, it’s not late, nah, it’s just early. The Spin Doctors really said it best. Seriously though, grabbing the current time is great for logging, file naming, and a host of other things. This function also works hand-in-hand with the timeadd function to get the date in the future or the past.
|
||
Lessons Learned Not gonna lie. There wasn’t much to learn. The most important bits of this function are what is in the official documentation. Every time you run terraform apply, the value is going to change. So if a portion of the resources is using that value, then Terraform will see it as a change. It may be necessary to add ignore_changes as a lifecycle property to avoid the diff.
|
||
Coming up next is the timeadd() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the timestamp() function. The example file is on GitHub here.
|
||
What is it? Function name: timestamp()
|
||
Returns: Takes no arguments. Returns the current time in UTC using the RFC 3339 standard.
|
||
Example:
|
||
# Returns something like 2018-09-05T00:00:00Z output "timestamp_output" { value = "\${timestamp()}" } Example file: ############################################## # Function: timestamp ############################################## ############################################## # Variables ############################################## ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_timestamp_basic" { value = "\${timestamp()}" } Run the following from the timestamp folder to get example output for a number of different cases:`,date:"7 Sep, 2018",url:"https://nedinthecloud.com/2018/09/07/terraform-fotd-timestamp/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/09/05/terraform-fotd-split/":{title:"Terraform - FotD - split()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the split() function. The example file is on GitHub here.
|
||
What is it? Function name: split(delimiter, string)
|
||
Returns: Takes a string and a delimiter. The string is split by the supplied delimiter and returned as a list of values.
|
||
Example:
|
||
# Returns [a, b, c] output "split_output" { value = "\${split(":","a:b:c")}" } Example file: ############################################## # Function: split ############################################## ############################################## # Variables ############################################## variable "source" { default = "Arthur scratched his head and discovered an Arthur doppelganger. \\nFord screamed and threw a towel over his head." } variable "delimiter" { default = "\\n" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_split" { value = "\${split(var.delimiter,var.source)}" } output "original" { value = "\${var.source}" } Run the following from the split folder to get example output for a number of different cases:
|
||
#Default values terraform apply #Comma delimited terraform apply -var "source=One,Two,Three,Four" -var "delimiter=," #Double delimited terraform apply -var "source=One::Two::Three::Four" -var "delimiter=::" #Whitespace delimited - regular expressions not supported terraform apply -var "source=One Two Three Four" -var "delimiter=\\s" Why use it? The most common reason to use this in Terraform is to split a string that has been produced as output by a module. In their current form, modules can only supply output in the form of a string. So if you wanted to have a list as output, you would use the join function to turn it into a string, and then use the split function in the configuration calling the module to turn the output back into a list. With Terraform version 0.12 coming, that restriction should be lifted. There will still be external data sources that supply strings you’ll want to split, so this function isn’t going to become useless. But it probably won’t be as commonly used as it once was.
|
||
Lessons Learned The function does exactly what it advertised. It only returns a flat list. There’s no fancy nested lists or map types available. You could create a map by using the map function once you have your list. That’s how you can get a map out of a module. It appears that escape sequences are supported, so I was able to use a tab-delimited string and a new-line delimited string. Multi-character delimiters are also supported, so you could use :: as the delimiter if you’re worried that your source data will match on single characters that are not a delimiter. Regular expressions for the delimiter are not supported, so you can’t just split on any whitespace character, or something funky like that. It probably would be nice to add regular expression capabilities - like what is in the replace function - but that does seem like overkill for most use cases.
|
||
Coming up next is the substr() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the split() function. The example file is on GitHub here.
|
||
What is it? Function name: split(delimiter, string)
|
||
Returns: Takes a string and a delimiter. The string is split by the supplied delimiter and returned as a list of values.`,date:"5 Sep, 2018",url:"https://nedinthecloud.com/2018/09/05/terraform-fotd-split/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/09/04/terraform-fotd-sort/":{title:"Terraform - FotD - sort()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the sort() function. The example file is on GitHub here.
|
||
What is it? Function name: sort(list)
|
||
Returns: Takes a list and returns the same list with the values sorted in lexicographical order.
|
||
Example:
|
||
# Returns [a, b, c] output "sort_output" { value = "\${sort(list("c","b","a"))}" } Example file: ############################################## # Function: sort ############################################## ############################################## # Variables ############################################## variable "string_list" { default = ["One", "Two", "Three", "Four"] } variable "string_list_2" { default = ["b","c","a",""] } variable "int_list" { default = [400, 42, 41] } variable "empty_string_list" { default = ["", "", ""] } variable "empty_list" { default = [] } variable "bool_list" { default = [true, false] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## #Test a standard list output "1_sort_list" { value = "\${sort(var.string_list)}" } output "1b_sort_list" { value = "\${sort(var.string_list_2)}" } #Test a standard int list output "2_int_list" { value = "\${sort(var.int_list)}" } #Test list of empty strings output "3_empty_string_list" { value = "\${sort(var.empty_string_list)}" } #Test what will happen with a list boolean values output "4_bool_list" { value = "\${sort(var.bool_list)}" } #This will fail. Nested lists are not supported #output "5_messing_around" { # value = "\${sort(list(list("l1v1","l2v2"),list("l2v1","l2v2")))}" #} Run the following from the sort folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? When you’re working on a list, there’s a good chance you are going to end up sorting it. This is barebones sorting though. If you need a flat list sorted in ascending lexicographical order, then this is the function for you. Anything more advanced, and you’ll need to look elsewhere. It would be really nice if this function took a second, optional argument to reverse the sort order. And maybe a number sort function as well? It could be called numsort to avoid confusion.
|
||
Lessons Learned I did learn a few interesting things about the sort function that I thought I would share. Per the docs, the function only takes flat lists. Don’t try to use a nested list or a list with a map in it. Errors will ensue. As noted, the sorting is based on a lexicographical order, which is a fancier way of saying alphabetical. Actually, that’s a bit flippant. Lexicographical is a more generalized version of alphabetical that is meant to deal with ordering any set of characters or other combination of data types. That allows it to deal with non-alphabetical characters and Unicode. It also means that if you hand the function a list of integers, it will not sort them by value as you might expect. My example above uses 400, 42, and 43. By value it should be sorted as (42,43,400). By lexicographical order, the 400 stays first since the first character of all the values is the same, and the second character of 0 in 400 comes before 2 or 3 in the other two values. I also discovered that an empty string will be sorted as the first value. That seems like an important thing to note if you are seeing strange behavior in your configs. I recommend using compact to remove those empty strings. Sanitized data is happy data.
|
||
Coming up next is the split() function. So let’s make like a banana…
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the sort() function. The example file is on GitHub here.
|
||
What is it? Function name: sort(list)
|
||
Returns: Takes a list and returns the same list with the values sorted in lexicographical order.
|
||
Example:
|
||
# Returns [a, b, c] output "sort_output" { value = "\${sort(list("c","b","a"))}" } Example file: ############################################## # Function: sort ############################################## ############################################## # Variables ############################################## variable "string_list" { default = ["One", "Two", "Three", "Four"] } variable "string_list_2" { default = ["b","c","a",""] } variable "int_list" { default = [400, 42, 41] } variable "empty_string_list" { default = ["", "", ""] } variable "empty_list" { default = [] } variable "bool_list" { default = [true, false] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## #Test a standard list output "1_sort_list" { value = "\${sort(var.`,date:"4 Sep, 2018",url:"https://nedinthecloud.com/2018/09/04/terraform-fotd-sort/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/08/31/terraform-fotd-slice/":{title:"Terraform - FotD - slice()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the slice() function. The example file is on GitHub here.
|
||
What is it? Function name: slice(list, start, end)
|
||
Returns: Takes a list, a starting index, and an ending index for the list. Returns a list that starts with the element found at the start value index and ends with the last element found in the list preceding the end value index.
|
||
Example:
|
||
# Returns [two, three] output "slice_output" { value = "\${slice(list("one","two","three","four"),1,3)}" } Example file: ############################################## # Function: slice ############################################## ############################################## # Variables ############################################## variable "string_list" { default = ["One", "Two", "Three", "Four"] } variable "int_list" { default = [41, 42, 43] } variable "empty_string_list" { default = ["", "", ""] } variable "empty_list" { default = [] } variable "bool_list" { default = [false, true] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## #Test a standard list output "1_slice_list" { value = "\${slice(var.string_list,0,1)}" } #Test length output "1b_slice_list" { value = "\${slice(var.string_list,1,length(var.string_list)-1)}" } #Test no items output "1c_slice_list" { value = "\${slice(var.string_list,0,0)}" } #Test a standard int list output "2_int_list" { value = "\${slice(var.int_list,1,3)}" } #Test list of empty strings output "3_empty_string_list" { value = "\${slice(var.empty_string_list,0,2)}" } #Test what will happen with a list boolean values output "4_bool_list" { value = "\${slice(var.bool_list,0,2)}" } Run the following from the slice folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? Slice is something I’ve used when I only want a subset of values from a list. A good example is when provisioning resources in AWS AZs. You can get the full list of AZs, which might have six entries, but you only want to deploy in two of those. You can use the slice function to get a list of only the first two AZs and use those as arguments for other resources.
|
||
Lessons Learned The slice function includes the value found at the starting index, but not the value found at the ending index. That can lead to an odd scenario where you have to give slice an index that doesn’t exist in the list. For instance, a list with three elements [1,2,3] has an maximum index of 2. If I want a slice to include the last element, I would need to use the following: slice(list(1,2,3),1,3)) And I would get the list [2,3] back. The index 3 doesn’t exist, but I need to submit it as a value so I get back the last element in the list. It actually makes things easier if you don’t know the list length. You can do something like this: slice(var.list,1,length(var.list)) To get the elements starting at index 1 through the remainder of the list. This was a design decision that someone at HashiCorp made, and while I don’t agree with it, I understand the choice. I think the start and end should both be inclusive, as that is what I would assume. The substr function works the same way. If you feed slice the same start and end value, then you get back nothing.
|
||
Coming up next is the sort() function. Sadly, there is no sprite function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the slice() function. The example file is on GitHub here.
|
||
What is it? Function name: slice(list, start, end)
|
||
Returns: Takes a list, a starting index, and an ending index for the list. Returns a list that starts with the element found at the start value index and ends with the last element found in the list preceding the end value index.`,date:"31 Aug, 2018",url:"https://nedinthecloud.com/2018/08/31/terraform-fotd-slice/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/08/30/terraform-fotd-signum/":{title:"Terraform - FotD - signum()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the signum() function. The example file is on GitHub here.
|
||
What is it? Function name: signum(int)
|
||
Returns: Takes an int and returns -1 if negative, 0 if 0, and 1 if positive
|
||
Example:
|
||
# Returns 0 output "signum_output" { value = "\${signum("0")}" } Example file: ############################################## # Function: signum ############################################## ############################################## # Variables ############################################## variable "signum" { default = 0 } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "signum_output" { value = "\${signum(var.signum)}" } Run the following from the signum folder to get example output for a number of different cases:
|
||
#base case returns 0 terraform apply #negative 0 is fine terraform apply -var "signum=-0" #returns 1 terraform apply -var "signum=1" #returns -1 terraform apply -var "signum=-1" #returns 1 terraform apply -var "signum=42" #returns -1 terraform apply -var "signum=-42" #errors, the max positive value is 1e+18 terraform apply -var "signum=10000000000000000000" #errors, the max negative value is -1e+18 terraform apply -var "signum=-10000000000000000000" Why use it? This function is dead simple. And I really don’t know why you would use it. It provides a way to determine if an int is negative, positive, or 0. The suggested use case in the official doc is to have one value for the first item in a list, and another value for the rest. If that’s not clear, let’s say you have a list of resources [main, other1, other2, other3] and another list of settings [main_settings, other_settings]. Then you could use signum to assign the proper settings to each resource in the list based on a count loop.
|
||
element(list(main_settings,other_settings),signum(count.index))
|
||
For the first iteration of the loop your count is 0, and the element function returns main_settings. The rest of the loop counts resolve to 1, and the element function returns other_settings. I thought for sure that this function was built into Golang’s math package, but no! Someone actually took the time to write a function to natively handle the operation. Here’s the code in case you were curious.
|
||
func interpolationFuncSignum() ast.Function { return ast.Function{ ArgTypes: []ast.Type{ast.TypeInt}, ReturnType: ast.TypeInt, Variadic: false, Callback: func(args []interface{}) (interface{}, error) { num := args[0].(int) switch { case num < 0: return -1, nil case num > 0: return +1, nil default: return 0, nil } }, } } So I guess someone felt strongly enough that this should be part of the native functions in Terraform. And to them I say, “You go, Glen Coco!”
|
||
Lessons Learned Well, I learned that the max value for the int type is 1 quintillion. That’s probably big enough. I also learned that -0 evaluates without an issue. That’s pretty much it. To be honest, I think I said it all in the Why use it section.
|
||
Coming up next is the slice() function. Sadly, there is no sprite function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the signum() function. The example file is on GitHub here.
|
||
What is it? Function name: signum(int)
|
||
Returns: Takes an int and returns -1 if negative, 0 if 0, and 1 if positive
|
||
Example:
|
||
# Returns 0 output "signum_output" { value = "\${signum("0")}" } Example file: ############################################## # Function: signum ############################################## ############################################## # Variables ############################################## variable "signum" { default = 0 } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "signum_output" { value = "\${signum(var.`,date:"30 Aug, 2018",url:"https://nedinthecloud.com/2018/08/30/terraform-fotd-signum/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/08/29/terraform-fotd-sha1-and-sha256-and-sha512/":{title:"Terraform - FotD - sha1() and sha256() and sha512()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the sha1(), sha256(), and sha512() functions. The example file is on GitHub here.
|
||
What is it? Function names: sha1(string) sha256(string) sha512(string)
|
||
Returns: Takes a string and returns a hash in hexadecimal form of the string based on which function is used.
|
||
Example:
|
||
# Returns 92cfceb39d57d914ed8b14d0e37643de0797ae56 output "sha1_output" { value = "\${sha1("42")}" } Example file: ############################################## # Function: sha1, sha256, sha512 ############################################## ############################################## # Variables ############################################## variable "string" { default = "So long, and thanks for all the fish!" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "sha1_output" { value = "\${sha1(var.string)}" } output "sha256_output" { value = "\${sha256(var.string)}" } output "sha512_output" { value = "\${sha512(var.string)}" } Run the following from the sha folder to get example output for a number of different cases:
|
||
#Start with the default variable terraform apply #Try submitting a string terraform apply -var 'string="Oh freddled gruntbuggly, Thy micturations are to me, As plurdled gabbleblotchits on a lurgid bee."' #Empty string test terraform apply -var "string=" Why use it? You use it to generate a SHA hash of some kind. There are some resources that require a SHA-1 or SHA-2 hash of a string or file. That hash would be compared to a value held by the resource to confirm validity. Usually this is for submitting a secret or password that the other side knows the hash of. That way you don’t have to send the sensitive data in the clear.
|
||
Lessons Learned Don’t use sha1 unless you have to. It’s not considered secure. I realize the function had to be in there for backwards compatibility, but if your resource or data source is using SHA-1, then it’s probably time to question whether you really want to use that resource. All three functions work the same, no surprise there. Sha256 used to choke on an empty string, as I found out when doing the base64sha256 post. But it would appear that bug has been fixed in the most recent versions of Terraform. Speaking of those functions, it’s important to remember that the sha functions all return a hexadecimal representation of the hash, not its raw form. The base64sha256 and base64sha512 functions take the raw output of the sha hash and base64 encode it. You will not get the same output if you run base64encode on the output of one of the sha functions. Terraform’s docs have that in bold, so I figured I’d mention it too.
|
||
Coming up next is the signum() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the sha1(), sha256(), and sha512() functions. The example file is on GitHub here.
|
||
What is it? Function names: sha1(string) sha256(string) sha512(string)
|
||
Returns: Takes a string and returns a hash in hexadecimal form of the string based on which function is used.`,date:"29 Aug, 2018",url:"https://nedinthecloud.com/2018/08/29/terraform-fotd-sha1-and-sha256-and-sha512/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/08/28/terraform-fotd-rsadecrypt/":{title:"Terraform - FotD - rsadecrypt()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the rsadecrypt() function. The example file is on GitHub here.
|
||
What is it? Function name: rsadecrypt(string, key)
|
||
Returns: Takes a base64 encoded string that has been encrypted with an RSA key and decrypts it using the supplied RSA key. Returns the decrypted string.
|
||
Example:
|
||
# Returns decrypted string output "rsadecrypt_output" { value = "\${rsadecrypt(file("path_to_encrypted_file"),file("path_to_rsa_key"))}" } Example file: ############################################## # Function: rsadecrypt ############################################## ############################################## # Variables ############################################## variable "encrypted_string_path" {} variable "private_key_path" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "rsadecrypt_output" { value = "\${rsadecrypt(file(var.encrypted_string_path),file(var.private_key_path))}" } Run the following from the rsadecrypt folder to get example output for a number of different cases:
|
||
#We have to generate a key pair first #You will need openssl to do this openssl genrsa -out private.key 2048 #Then we have to encrypt a string openssl rsautl -encrypt -inkey .\\private.key -in .\\cleartext.txt -out encrypted.txt #And base64 encode it openssl enc -a -in .\\encrypted.txt -out .\\encrypted64.txt -none #Now we can use Terraform to decrypt terraform apply -var "encrypted_string_path=encrypted64.txt" -var "private_key_path=private.key" Why use it? I am guessing that you are getting the encrypted string from some parameter store and then getting the RSA key from some secure location and decrypting the value. This could probably be used to store passwords in a key value store, and decrypt them at runtime with an RSA key that was stored somewhere else. I’d rather use Vault for storing sensitive data, or Azure Key Vault, or AWS Secrets Manager. What I am saying, is that there are options, ones that are likely better than this. This function was added in version 11.2 with the reasoning: “This is particularly useful for decrypting the password for a Windows instance on AWS EC2, but is generic and may find other uses too.” That is a legitimate use case, but a better plan would be to store the password on one of the secure options I listed above.
|
||
Lessons Learned Getting this to work as an example was a bit of a challenge. I don’t exactly use openssl everyday. But after about 20 minutes of furious Google-Fu and general messing around, I was able to work out some simple commands to see this bad boy in action. You will need openssl installed on your local machine to work through the example. If you don’t already have it installed, you probably should. I used Chocolately to install it: choco install OpenSSL.Light -y
|
||
If you’re on Linux, just hit up your package manager.
|
||
Coming up next is a triple threat, I’m going to combine the sha1(), sha256(), and sha512() functions into a single post. Because I am crazy efficient.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the rsadecrypt() function. The example file is on GitHub here.
|
||
What is it? Function name: rsadecrypt(string, key)
|
||
Returns: Takes a base64 encoded string that has been encrypted with an RSA key and decrypts it using the supplied RSA key.`,date:"28 Aug, 2018",url:"https://nedinthecloud.com/2018/08/28/terraform-fotd-rsadecrypt/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/08/27/terraform-fotd-replace/":{title:"Terraform - FotD - replace()",tags:["fotd","hashicorp","terraform"],content:"This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the replace() function. The example file is on GitHub here.\nWhat is it? Function name: replace(string, search, replace)\nReturns: Takes a string and matches instances in the search value and replaces them with the replace value. Returns the updated string. The search value supports regular expressions if the value is bounded by the forward slash ‘/’.\nExample:\n# Returns bar output "replace_output" { value = "${replace("foo","foo","bar")}" } Example file: ############################################## # Function: replace ############################################## ############################################## # Variables ############################################## variable "source" { default = "Arthur scratched his head and discovered an Arthur doppelganger. Ford screamed and threw a towel over his head." } variable "search" {} variable "replace" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_replace_basic" { value = "${replace(var.source,var.search,var.replace)}" } output "2_replace_long_hostname" { value = "${replace("heart-of-gold-42","/^(.)[^-]+-(.)[^-]+-(.)[^-]+-(\\\\d+)$/","$1$2$3$4")}" } output "3_replace_arthur" { value = "${replace(var.source,"/^(A\\\\S+)\\\\s.*$/","$1 stands alone.")}" } Run the following from the replace folder to get example output for a number of different cases:\n#Standard replace terraform apply -var "search=Arthur" -var "replace=Zaphod" #Regex replace with or terraform apply -var "search=/Arthur|Ford/" -var "replace=Zaphod" #Empty replace is fine terraform apply -var "search=/Arthur|Ford/" -var "replace=" #Emtpy search matches all characters except whitespace terraform apply -var "search=" -var "replace=42" Why use it? Replace is a pretty critical function if you’re planning to do any text string manipulation. You might use it to replace region names, hostnames, domain names, heck anything you want really. This does not work on lists or maps, so you would need to iterate through a list using the element function and a count loop. Same thing for map values, except first you’d need to use the values function to get a list of values.\nLessons Learned The regex functionality is not well explained by the documentation. I guess either you’re super familiar with regex or you’re not, and the Terraform docs are not there to teach regex anyhow. For starters, you need to add the forward slash to the search string, e.g. /foo/ will match foo. You also need to double escape any backslashes as usual with Terraform, so the whitespace special character is \\\\s. The use of capture groups is also supported. Let’s say I have the string "foo man chu" and I use the regular expression /(foo).*/. That matches foo followed by 0 or more characters. The captured value of foo will be captured and stored in the placeholder $1. In the replace field you can put $1 bar and since the regular expression matches the entire source string, the replace value will replace all of the string. Since foo is captured in $1, the resulting string will be foo bar. Why would you use this? A quick Google search pointed out a situation where a person was trying to shorten the location names in AWS to be prepended to a hostname. I am sure there are other occasions where this would be useful. The syntax for the regular expressions is based on this GitHub project. You can also test your regular expressions, including capture groups, with this handy website. I also learned that if you submit an empty search string, the function will prepend the replace value to each non-whitespace character in the source string. That might actually be useful in some universe. Otherwise it just leads to really wacky behavior that will be “fun” to debug.\nComing up next is the rsadecrypt() function.\n",summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the replace() function. The example file is on GitHub here.
|
||
What is it? Function name: replace(string, search, replace)
|
||
Returns: Takes a string and matches instances in the search value and replaces them with the replace value. Returns the updated string.`,date:"27 Aug, 2018",url:"https://nedinthecloud.com/2018/08/27/terraform-fotd-replace/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/08/25/terraform-fotd-pow/":{title:"Terraform - FotD - pow()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the pow() function. The example file is on GitHub here.
|
||
What is it? Function name: pow(base,exponent)
|
||
Returns: Takes a base value and raises it by the exponent
|
||
Example:
|
||
# Returns 8 output "pow_output" { value = "\${pow("2","3")}" } Example file: ############################################## # Function: pow ############################################## ############################################## # Variables ############################################## variable "base" {} variable "exp" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "pow_output" { value = "\${pow(var.base, var.expZ)}" } Run the following from the pow folder to get example output for a number of different cases:
|
||
#Basic functionality terraform apply -var 'base=2' -var 'exp=3' terraform apply -var 'base=-1' -var 'exp=3' terraform apply -var 'base=3.1' -var 'exp=-2' #Gives NaN, weird, guess it doesn't like decimal exponent terraform apply -var 'base=-3.1' -var 'exp=2.3' #Checking maximums, 308 seems to be the max terraform apply -var 'base=10' -var 'exp=309' terraform apply -var 'base=10' -var 'exp=308' Why use it? This is a bit like the log function, both from a mathematical and usefulness perspective. Mathematically, this is the inverse operation of a log function. How useful this function is to deploying automated infrastructure is a bit more debatable. It is pretty easy to implement though, since the pow function is a part of the math package in Golang.
|
||
Lessons Learned The function has an upper limit on value, so anything more than 10^308 is rendered as infinity. Since there are only about 10^82 atoms in the universe, I guess that’s more than sufficient. The function also does not like a negative base value with a decimal for the exponent. The resulting value is NaN. I checked and that is how Golang renders it as well, so that’s not an issue of implementation on Terraform’s part.
|
||
Coming up next is the replace() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the pow() function. The example file is on GitHub here.
|
||
What is it? Function name: pow(base,exponent)
|
||
Returns: Takes a base value and raises it by the exponent
|
||
Example:
|
||
# Returns 8 output "pow_output" { value = "\${pow("2","3")}" } Example file: ############################################## # Function: pow ############################################## ############################################## # Variables ############################################## variable "base" {} variable "exp" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "pow_output" { value = "\${pow(var.`,date:"25 Aug, 2018",url:"https://nedinthecloud.com/2018/08/25/terraform-fotd-pow/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/08/23/terraform-fotd-pathexpand/":{title:"Terraform - FotD - pathexpand()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the pathexpand() function. The example file is on GitHub here.
|
||
What is it? Function name: pathexpand(string)
|
||
Returns: Takes a string that should be a file path, and expands the ~ to the current user’s home directory
|
||
Example:
|
||
# Returns /home/ned/test.txt output "path_expand_output" { value = "\${path_expand("~/test.txt")}" } Example file: ############################################## # Function: pathexpand ############################################## ############################################## # Variables ############################################## variable "pathexpand" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "pathexpand_output" { value = "\${pathexpand(var.pathexpand)}" } Run the following from the pathexpand folder to get example output for a number of different cases:
|
||
#Windows tests #Try the home directory terraform apply -var 'pathexpand=~\\' #Try with a file extension terraform apply -var 'pathexpand=~\\test.txt' #Try non home path terraform apply -var 'pathexpand=C:\\test.txt' #Linux tests #Try with home directory terraform apply -var 'pathexpand=~/' #Try with file terraform apply -var 'pathexpand=~/test.txt' #Try empty value terraform apply -var 'pathexpand=' Why use it? I would assume that this is mostly used by provisioning scripts to return the full path to a file or directory that is sitting in the user’s home directory. Since different operating systems might be using different usernames for connection or deployment, it’s probably important to get back the actual path and not some type of shorthand.
|
||
Lessons Learned The function will expand a path that starts with the ‘~’ character to the current user’s home directory path. If you submit a path that does not begin with a tilda, then it just returns back the same path. The function also handles an empty string by returning the empty string with nothing additional.
|
||
Coming up next is the pow() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the pathexpand() function. The example file is on GitHub here.
|
||
What is it? Function name: pathexpand(string)
|
||
Returns: Takes a string that should be a file path, and expands the ~ to the current user’s home directory
|
||
Example:
|
||
# Returns /home/ned/test.`,date:"23 Aug, 2018",url:"https://nedinthecloud.com/2018/08/23/terraform-fotd-pathexpand/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/08/22/terraform-fotd-md5/":{title:"Terraform - FotD - md5()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the md5() function. The example file is on GitHub here.
|
||
What is it? Function name: md5(string)
|
||
Returns: Takes a string and returns the md5 hash of that string
|
||
Example:
|
||
# Returns 827ccb0eea8a706c4c34a16891f84e7b output "md5_output" { value = "\${md5("12345")}" } Example file: ############################################## # Function: md5 ############################################## ############################################## # Variables ############################################## variable "md5" {} variable "fileText" { default = "textFile.txt" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "md5_output" { value = "\${md5(var.md5)}" } output "md5_file" { value = "\${md5(file(var.fileText))}" } Run the following from the md5 folder to get example output for a number of different cases:
|
||
#Should return a26d69fa3acf4022a4505a0c823af393 terraform apply -var "md5=amanaplanacanalpanama" #Should output 8D93AD635E2C8C822D796BD8726EEF6B terraform apply -var "md5=ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789" #Empty string test give d41d8cd98f00b204e9800998ecf8427e (whut?) terraform apply -var "md5=" Why use it? There’s probably two good reasons to use this. If you’re downloading a file and want to verify its integrity, then you could calculate the md5 hash of the file and compare it to a hash supplied by the source. The other reason is that some resources require an md5 hash to be submitted along with a source file in order to validate a successful upload. An example would be the etag optional attribute on the aws_s3_bucket_object resource.
|
||
Lessons Learned The function works as advertised. I provided an example where I pulled a file’s contents and calculated the md5 hash along with just doing it to a standard string. I think the file hash calculation is probably the more common of the two options. Funny thing is that running the md5 calculation with an “empty” string resulted in d41d8cd98f00b204e9800998ecf8427e. I don’t know how or why, but that’s cool I guess.
|
||
Coming up next is the pathexpand() function. You might wonder what happened to ’n’ and ‘o’. I guess there needs to be more functions!
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the md5() function. The example file is on GitHub here.
|
||
What is it? Function name: md5(string)
|
||
Returns: Takes a string and returns the md5 hash of that string
|
||
Example:
|
||
# Returns 827ccb0eea8a706c4c34a16891f84e7b output "md5_output" { value = "\${md5("12345")}" } Example file: ############################################## # Function: md5 ############################################## ############################################## # Variables ############################################## variable "md5" {} variable "fileText" { default = "textFile.`,date:"22 Aug, 2018",url:"https://nedinthecloud.com/2018/08/22/terraform-fotd-md5/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/08/21/terraform-fotd-min/":{title:"Terraform - FotD - min()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the min() function. The example file is on GitHub here.
|
||
What is it? Function name: min(float1, float2, float3,…)
|
||
Returns: Takes a list of numeric values and returns the lowest value of the set.
|
||
Example:
|
||
# Returns 0 output "min_output" { value = "\${min("0","2","3.14159")}" } Example file: ############################################## # Function: min ############################################## ############################################## # Variables ############################################## variable "min1" { default = false } variable "min2" { default = false } variable "min3" { default = false } variable "min4" { default = false } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "min_output" { value = "\${min(var.min1,var.min2,var.min3,var.min4)}" } Run the following from the min folder to get example output for a number of different cases:
|
||
#Regular values terraform apply -var 'min1=5.9' -var 'min2=4.9' -var 'min3=6.9' -var 'min4=5.4' #Negative values terraform apply -var 'min1=-5.9' -var 'min2=5.9' -var 'min3=3.9' -var 'min4=-3.9' #Same values terraform apply -var 'min1=5.9' -var 'min2=5.9' -var 'min3=5.9' -var 'min4=5.9' #Longer decimal point max out at 15 terraform apply -var 'min1=5.1123456789123456' -var 'min2=5.1123456789123455' -var 'min3=5.1123456789123454' -var 'min4=5.1123456789123453' #Ints terraform apply -var 'min1=5' -var 'min2=6' -var 'min3=7' -var 'min4=8' Why use it? I guess you’ve got a bunch of floats and need to know which one is the smallest. Really this is just a pass through of a basic function in the Go Math package. I’ve never had to use it, but that doesn’t mean others won’t. The functionality is a little weird to me, as I’ll explain in a moment.
|
||
Lessons Learned There wasn’t much to get excited about here. I tried using some really long floats and it appears that the float type is a double, meaning it can go up to about 16 decimal places of precision. That should be more than enough for anything you are doing in Terraform. The thing I don’t like about the function is that it won’t take numbers in a list. You have to submit individual float values. You could use the sort function on a list and take the first or last element, but that is for lexographical sorting, not for numerical sorting. I’m guessing the two are a bit different. I’ll test that out when I get to sort.
|
||
Coming up next is the md5() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the min() function. The example file is on GitHub here.
|
||
What is it? Function name: min(float1, float2, float3,…)
|
||
Returns: Takes a list of numeric values and returns the lowest value of the set.
|
||
Example:
|
||
# Returns 0 output "min_output" { value = "\${min("0","2","3.`,date:"21 Aug, 2018",url:"https://nedinthecloud.com/2018/08/21/terraform-fotd-min/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/08/20/terraform-fotd-merge/":{title:"Terraform - FotD - merge()",tags:["fotd","hashicorp","terraform"],content:"This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the merge() function. The example file is on GitHub here.\nWhat is it? Function name: merge(map1, map2, map3, …)\nReturns: Takes one or more maps and merges them into a single map by performing a union operation. Duplicate keys are overwritten.\nExample:\n# Returns {m1k1:m1v1 m2k1:m2v1} output "max_output" { value = "${merge(map("m1k1","m1v1"),map("m2k1","m2v1"))}" } Example file: ############################################## # Function: merge ############################################## ############################################## # Variables ############################################## variable "map_1_value" { type = "map" default = { "m1k1" = "m1v1" "m1k2" = "m1v2" "m1k3" = "m1v3" } } variable "map_2_value" { type = "map" default = { "m2k1" = "m2v1" "m2k2" = "m2v2" "m2k3" = "m2v3" } } variable "empty_map" { type = "map" default = {} } variable "nested_map_1" { type = "map" default = { "n1m1" = { "n1m1k1" = "n1m1v1" "n1m1k2" = "n1m1v2" "n1m1k3" = "n1m1v3" } "n1m2" = { "n1m2k1" = "n1m2v1" "n1m2k2" = "n1m2v2" "n1m2k3" = "n1m2v3" } } } variable "nested_map_2" { type = "map" default = { "n2m1" = { "n2m1k1" = "n2m1v1" "n2m1k2" = "n2m1v2" "n2m1k3" = "n2m1v3" } "n2m2" = { "n2m2k1" = "n2m2v1" "n2m2k2" = "n2m2v2" "n2m2k3" = "n2m2v3" } } } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_merge_basic" { value = "${merge(var.map_1_value,var.map_2_value)}" } output "2_merge_empty" { value = "${merge(var.map_1_value,var.empty_map)}" } output "3_merge_duplicates" { value = "${merge(var.map_1_value,var.map_2_value,map("m1k1","m1v1-alt"))}" } output "4_merge_nested" { value = "${merge(var.nested_map_1,var.nested_map_2)}" } output "5_merge_mixed" { value = "${merge(var.map_1_value,var.nested_map_2)}" } output "5_merge_mega" { value = "${merge(var.map_1_value,var.map_2_value,var.nested_map_1,var.nested_map_2)}" } output "6_merge_null" { value = "${merge(var.empty_map)}" } Run the following from the merge folder to get example output for a number of different cases:\n#All examples are in variables terraform apply Why use it? Makes sense that if you have maps from disparate sources and you want to merge them together that you would have a function to do that. I could see using this with tagging. Maybe you’ve got a set of standard tags and then some custom tags for a specific project. Both tag sets are stored in maps, so you use the merge function to stitch them together and apply to all the resources in the project. Now that I think about it, I’m pretty sure that is exactly what I used this for in the past. Huzzah!\nLessons Learned I learned a couple of fun things. First, you can provide a single map value and the function will still run. Not that it has anything to do, but it won’t throw an error either. The function doesn’t care about nested maps. It just merges the top level keys and maintains the nested maps as values. The one important thing I learned is the order of overwrite for duplicate keys. If you have duplicate keys in your maps, the last map wins. In the example file above, I have a map with key m1k1 and value m1v1. I also have a map with key m1k1 and value m1v1-alt. Since the map with the m1v1-alt value is last in the merge order, that value is what ends up in the key m1k1. That could be hell to debug if you don’t know that’s how the function merges keys.\nComing up next is the min() function.\n",summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the merge() function. The example file is on GitHub here.
|
||
What is it? Function name: merge(map1, map2, map3, …)
|
||
Returns: Takes one or more maps and merges them into a single map by performing a union operation. Duplicate keys are overwritten.`,date:"20 Aug, 2018",url:"https://nedinthecloud.com/2018/08/20/terraform-fotd-merge/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/08/17/terraform-fotd-max/":{title:"Terraform - FotD - max()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the max() function. The example file is on GitHub here.
|
||
What is it? Function name: max(float1, float2, float3,…)
|
||
Returns: Takes a list of numeric values and returns the largest value of the set.
|
||
Example:
|
||
# Returns 3.14159 output "max_output" { value = "\${max("0","2","3.14159")}" } Example file: ############################################## # Function: max ############################################## ############################################## # Variables ############################################## variable "max1" { default = false } variable "max2" { default = false } variable "max3" { default = false } variable "max4" { default = false } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "max_output" { value = "\${max(var.max1,var.max2,var.max3,var.max4)}" } Run the following from the max folder to get example output for a number of different cases:
|
||
#Regular values terraform apply -var 'max1=5.9' -var 'max2=4.9' -var 'max3=6.9' -var 'max4=5.4' #Negative values terraform apply -var 'max1=-5.9' -var 'max2=5.9' -var 'max3=3.9' -var 'max4=-3.9' #Same values terraform apply -var 'max1=5.9' -var 'max2=5.9' -var 'max3=5.9' -var 'max4=5.9' #Longer decimal point max out at 15 terraform apply -var 'max1=5.12345678912345' -var 'max2=5.123456789123456' -var 'max3=5.1234567891234567' -var 'max4=5.12345678912345678' #Ints terraform apply -var 'max1=5' -var 'max2=6' -var 'max3=7' -var 'max4=8' Why use it? I guess you’ve got a bunch of floats and need to know which one is the biggest. Really this is just a pass through of a basic function in the Go Math package. I’ve never had to use it, but that doesn’t mean others won’t. The functionality is a little weird to me, as I’ll explain in a moment.
|
||
Lessons Learned There wasn’t much to get excited about here. I tried using some really long floats and it appears that the float type is a double, meaning it can go up to about 16 decimal places of precision. That should be more than enough for anything you are doing in Terraform. The thing I don’t like about the function is that it won’t take numbers in a list. You have to submit individual float values. You could use the sort function on a list and take the first or last element, but that is for lexographical sorting, not for numerical sorting. I’m guessing the two are a bit different. I’ll test that out when I get to sort.
|
||
Coming up next is the merge() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the max() function. The example file is on GitHub here.
|
||
What is it? Function name: max(float1, float2, float3,…)
|
||
Returns: Takes a list of numeric values and returns the largest value of the set.
|
||
Example:
|
||
# Returns 3.14159 output "max_output" { value = "\${max("0","2","3.`,date:"17 Aug, 2018",url:"https://nedinthecloud.com/2018/08/17/terraform-fotd-max/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/08/16/terraform-fotd-matchkeys/":{title:"Terraform - FotD - matchkeys()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the matchkeys() function. The example file is on GitHub here.
|
||
What is it? Function name: matchkeys(values, keys, search_set)
|
||
Returns: Takes a list of values and a list of keys that are of the same length and a list of keys to search for. If a matching key from the search_set is found in the keys list, then the function returns the corresponding value from the value list. The return type is always a list. If no matching terms are found, an empty list is returned.
|
||
Example:
|
||
variable "value" { type = "list" default = "value" } variable "key" { type = "list" default = "key" } # Returns [value] output "matchkey_output" { value = "\${matchkey(var.value,var.key,list("key"))}" } Example file: ############################################## # Function: matchkeys ############################################## ############################################## # Variables ############################################## variable "value_list" { type = "list" default = ["v1", "v2", "v3", "v4"] } variable "key_list" { type = "list" default = ["k1", "k2", "k3", "k4"] } variable "search_list" { type = "list" default = ["k1","k2"] } variable "map_test" { type = "map" default = { "k1" = "v1" "k2" = "v2" "k3" = "v3" } } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## #Basic query output "1_matchkeys_output" { value = "\${matchkeys(var.value_list,var.key_list,var.search_list)}" } #No match query output "2_matchkeys_output" { value = "\${matchkeys(var.value_list,var.key_list,list("k5"))}" } #Test with a map output "3_matchkeys_output" { value = "\${matchkeys(values(var.map_test),keys(var.map_test),var.search_list)}" } Run the following from the matchkeys folder to get example output for a number of different cases:
|
||
#All examples are in output terraform apply Why use it? The main thing for me is the ability to search a map for multiple values and get them back as a list. There’s already a lookup function to check a map for a particular key value, but that only handles a single lookup operation. If you wanted to search multiple keys and get back multiple values, then lookup won’t do that for you. Instead you can take a map object and use the values and keys functions to get lists and drop them into the matchkeys function. I don’t have a real world example of where I would use this, but I can see the need for such a function to exist.
|
||
Lessons Learned I had to read the doc for this function about five times before I thought I understood what it did. And the example given did not really help clear that up for me at all. That’s part of the reason I started this series of posts in the first place. Hopefully, if you are also confused by this function, this post helped clear up some of that confusion. It’s also important to note that the search set must also be a list type and not a string. If you are only looking for a single value in a map, just use the lookup function. If you have two lists, you can always use the zipmap function to create a map first, then use the lookup function.
|
||
Coming up next is the max() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the matchkeys() function. The example file is on GitHub here.
|
||
What is it? Function name: matchkeys(values, keys, search_set)
|
||
Returns: Takes a list of values and a list of keys that are of the same length and a list of keys to search for.`,date:"16 Aug, 2018",url:"https://nedinthecloud.com/2018/08/16/terraform-fotd-matchkeys/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/08/15/terraform-fotd-map/":{title:"Terraform - FotD - map()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the map() function. The example file is on GitHub here.
|
||
What is it? Function name: map(key, value,…)
|
||
Returns: Takes a set of keys and values, alternating, and returns a map based off the key/value pairs. All values must be of the same type. No duplicate keys are allowed.
|
||
Example:
|
||
variable "value" { default = "value" } # Returns { key = value } output "map_output" { value = "\${map("key",var.value)}" } Example file: ############################################## # Function: map ############################################## ############################################## # Variables ############################################## ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_basic_map" { value = "\${map("key1","value1","key2","value2")}" } output "2_map_lists" { value = "\${map("life",list("42"),"the universe",list("six","times","seven"))}" } output "3_empty_map" { value = "\${map()}" } output "4_nested_map" { value = "\${ map( "nested1", map("n1k1","n1v1","n1k2","n1v2"), "nested2", map("n2k1","n2v1","n2k2","n2v2") )}" } output "5_triple_nested" { value = "\${ map( "level1", map( "level2", map( "level3","infinity" ) ) ) }" } Run the following from the map folder to get example output for a number of different cases:
|
||
#All examples are in output terraform apply Why use it? Sometimes you just need a map. I’ve been using this function to generate maps for other functions when creating these posts. But this is also a convenient way to package up information from something that maybe doesn’t output the map data type you wanted. If you’re reading in data using an external data source, then you could use map to construct a map from that external information and then parse that out for something like a firewall rule set.
|
||
Lessons Learned The function does precisely what you would expect. I was pleasantly surprised to see that it let me create nested maps and triple nested maps. I don’t know how far down you could go with the nesting, but it works to three levels. Beyond that you are probably just probing the limit for the fun of it, rather than having some practical application. One of the limitations here is that the values of a map need to be the same data type, so I can’t mix strings, lists, and maps for the value. Not sure why I would want to, but also not sure why the limitation exists.
|
||
Coming up next is the matchkeys() function, which appears to be hella confusing in the docs.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the map() function. The example file is on GitHub here.
|
||
What is it? Function name: map(key, value,…)
|
||
Returns: Takes a set of keys and values, alternating, and returns a map based off the key/value pairs. All values must be of the same type.`,date:"15 Aug, 2018",url:"https://nedinthecloud.com/2018/08/15/terraform-fotd-map/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/08/14/terraform-fotd-lower/":{title:"Terraform - FotD - lower()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the lower() function. The example file is on GitHub here.
|
||
What is it? Function name: lower(string)
|
||
Returns: Takes a string and returns the lowercase version of all characters in the string as a string.
|
||
Example:
|
||
variable "string" { default = "ALL CAPS" } # Returns "all caps" output "lower_output" { value = "\${lower(var.string)}" } Example file: ############################################## # Function: lower ############################################## ############################################## # Variables ############################################## variable "lower" { default = "This Line Is Capitalized." } variable "sourcefile" { default = "textFile.txt" } ############################################## # Resources ############################################## data "local_file" "source" { filename = "\${var.sourcefile}" } ############################################## # Outputs ############################################## output "lower_output" { value = "\${lower(var.lower)}" } output "file_output" { value = "\${lower(data.local_file.source.content)}" } Run the following from the lower folder to get example output for a number of different cases:
|
||
#We have to get the local resource first terraform init #Default values terraform apply #All caps terraform apply -var "lower=THIS IS IN ALL CAPITALS" #Already lowercase terraform apply -var "lower=this is in all lowercase" #Empty string terraform apply -var "lower=" #All standard US-EN characters terraform apply -var 'lower=QWERTYUIOPASDFGHJKLZXCVBNM1234567890,./;[]\\<>?:"{}|~!@#$%^&*()_+-=\`' Why use it? There are a number of resources that will not accept uppercase characters for some of their values. For instance, storage accounts in Microsoft Azure need to be all lowercase for… well, for reasons obviously. Yeah, not sure why Microsoft decided to start caring about case for this one service when they have been case insensitive forever. But my point stands regardless, if you have a resource that won’t accept uppercase characters, then don’t trust the end user to know that, just make it lowercase for them.
|
||
Lessons Learned The function does precisely what you would expect. I only tested it with standard EN-US characters, so your mileage could vary. It should be able to do lowercase translation for any character in the Unicode set. There are 1,781 uppercase characters in Unicode version 11. If you’d like to test all of them, have fun! Personally I will trust the Go function that is being called here to do the right thing.
|
||
Coming up next is the map() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the lower() function. The example file is on GitHub here.
|
||
What is it? Function name: lower(string)
|
||
Returns: Takes a string and returns the lowercase version of all characters in the string as a string.
|
||
Example:
|
||
variable "string" { default = "ALL CAPS" } # Returns "all caps" output "lower_output" { value = "\${lower(var.`,date:"14 Aug, 2018",url:"https://nedinthecloud.com/2018/08/14/terraform-fotd-lower/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/08/13/terraform-fotd-lookup/":{title:"Terraform - FotD - lookup()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the lookup() function. The example file is on GitHub here.
|
||
What is it? Function name: lookup(map, key, [default])
|
||
Returns: Takes a map, a key, and an optional default value. Returns the value in the map that corresponds to the key if found. If not found, returns the optional default value or throws an error.
|
||
Example:
|
||
variable "map" { type = map default = { "one" = "1" } } # Returns 1 output "lookup_output" { value = "\${lookup(var.map, "one")}" } Example file: ############################################## # Function: lookup ############################################## ############################################## # Variables ############################################## variable "map_value" { type = "map" default = { "life" = "42" "universe" = "6" "everything" = "7" } } variable "empty_map" { type = "map" default = {} } variable "nested_map" { type = "map" default = { "Zaphod" = { "First" = "Zaphod" "Last" = "Beeblebrox" } "Arthur" = { "First" = "Arthur" "Last" = "Dent" } } } variable "lookup_key" { default = "life" } variable "default_key" { default = "Trillian" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_map_value_output" { value = "\${lookup(var.map_value,var.lookup_key,var.default_key)}" } output "2_empty_map_output" { value = "\${lookup(var.empty_map,var.lookup_key,var.default_key)}" } #Nested maps not allowed. #Will work if you specify a non-existent key and a default value #output "3_nested_map_output" { # value = "\${lookup(var.nested_map,var.lookup_key,var.default_key)}" #} Run the following from the lookup folder to get example output for a number of different cases:
|
||
#Basic use terraform apply #Alternate key terraform apply -var "lookup_key=everything" #Non-existent key terraform apply -var "lookup_key=missing" Why use it? This has got to be one of the most used functions in Terraform. There are a ton of data sources that return a map of values you might want to parse through. AWS AMIs, Azure VMs, firewall rules, metadata tags. That’s just a few off the top of my head. If you’ve got a map of key pairs, lookup is your pal.
|
||
Lessons Learned The documentation on the Terraform site states that the function will not work with nested maps, i.e. maps with maps as the value associated to a key. I tried using a nested map for funsies, and it didn’t throw an error. What I learned is that the function evaluates the presence of the key first. So if you pass a key that isn’t in the map, then the function checks for a default value. If you specified a default value, then that will be returned, even though the values in the map are not valid. If you pass a key that is in the map, then you’ll get the error that nested maps are not allowed. The key takeaway (<–see what I did there?) is that you shouldn’t use a nested map. Empty maps are fine, provided you specified a default value. Again, the function checks for the presence of a key before checking the values for validity or existence. Good times.
|
||
Coming up next is the lower() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the lookup() function. The example file is on GitHub here.
|
||
What is it? Function name: lookup(map, key, [default])
|
||
Returns: Takes a map, a key, and an optional default value. Returns the value in the map that corresponds to the key if found.`,date:"13 Aug, 2018",url:"https://nedinthecloud.com/2018/08/13/terraform-fotd-lookup/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/08/06/terraform-fotd-log/":{title:"Terraform - FotD - log()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the log() function. The example file is on GitHub here.
|
||
What is it? Function name: log(value, base)
|
||
Returns: Takes a numeric value and a base numeric value and returns the logarithm of the value in the base.
|
||
Example:
|
||
variable "value" { default = "100" } # Returns 2 output "list_output" { value = "\${log(var.value, 10)}" } Example file: ############################################## # Function: log ############################################## ############################################## # Variables ############################################## variable "log" { default = 10 } variable "base" { default = 10 } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "log_output" { value = "\${log(var.log,var.base)}" } Run the following from the log folder to get example output for a number of different cases:
|
||
terraform apply -var 'log=5' -var 'base=10' terraform apply -var 'log=-5' -var 'base=2' terraform apply -var 'log=5' -var 'base=-10' terraform apply -var 'log=-5' -var 'base=-10' terraform apply -var 'log=0' -var 'base=10' terraform apply -var 'log=5' -var 'base=0' terraform apply Why use it? You know, I had to think pretty hard to remember what a logarithm actually is. If you’re like me, here’s a quick refresher. The logarithm is the exponent that would have to be applied to the base number to get the desired value. So if you want to know the log of 100 with base 10, you would need to know what exponent to apply to 10 to get 100. That one is easy of course, you would just square 10, thus the log is 2.
|
||
Now, why would you use this in a Terraform configuration? I have no idea. I suppose if you are already using the math package in Golang for more common functions like floor or ceil, then this is super easy to add in.
|
||
Lessons Learned I didn’t really learn anything new on this one. The function does exactly what you would expect it to do.
|
||
Coming up next is the lookup() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the log() function. The example file is on GitHub here.
|
||
What is it? Function name: log(value, base)
|
||
Returns: Takes a numeric value and a base numeric value and returns the logarithm of the value in the base.
|
||
Example:`,date:"6 Aug, 2018",url:"https://nedinthecloud.com/2018/08/06/terraform-fotd-log/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/27/terraform-fotd-list/":{title:"Terraform - FotD - list()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the list() function. The example file is on GitHub here.
|
||
What is it? Function name: list(value, value, value,…)
|
||
Returns: Takes one or more values of string, list, or map and returns a list containing those values as elements.
|
||
Example:
|
||
variable "string" { type = "string" default = "Four" } # Returns ["Four"] output "list_output" { value = "\${list(var.string)}" } Example file: ############################################## # Function: list ############################################## ############################################## # Variables ############################################## variable "value_1" { default = "Ford" } variable "value_2" { default = "Prefect" } variable "value_3" { default = "Arthur" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_basic_list_output" { value = "\${list(var.value_1,var.value_2,var.value_3)}" } output "2_nested_list_output" { value = "\${list(list(var.value_1),list(var.value_2,var.value_3))}" } output "3_map_list_output" { value = "\${list(map("life","42","universe","6"),map("everything","7"))}" } output "4_empty_list" { value = "\${list()}" } Run the following from the list folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? If you need to pack stuff up as a list to pass to another resource or module, then this is the way to do it. This one is pretty straightforward. I’ve also used this function in a bunch of the other examples to test a list against another function.
|
||
Lessons Learned One thing I learned, which is missing from the documentation, is that this function will happily accept a bunch of string, list, or map values and return a list containing them. What it will not do is allow you to mix value types in the list. If you try something like list("a",list("b")) then it will throw an error saying that element one (1) is a list and it was expecting a string. No big deal, but something to keep in mind. Otherwise the function works as advertised.
|
||
Coming up next is the log() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the list() function. The example file is on GitHub here.
|
||
What is it? Function name: list(value, value, value,…)
|
||
Returns: Takes one or more values of string, list, or map and returns a list containing those values as elements.`,date:"27 Jul, 2018",url:"https://nedinthecloud.com/2018/07/27/terraform-fotd-list/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/26/terraform-fotd-keys/":{title:"Terraform - FotD - keys()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the keys() function. The example file is on GitHub here.
|
||
What is it? Function name: keys(map)
|
||
Returns: Takes a map and returns the keys of that map in a list that has been lexically sorted.
|
||
Example:
|
||
variable "map" { type = "map" default = { "one" = "1" "two" = "2" } } # Returns ["one","two"] output "key_output" { value = "\${keys(var.map)}" } Example file: ############################################## # Function: keys ############################################## ############################################## # Variables ############################################## variable "map_value" { type = "map" default = { "life" = "42" "universe" = "6" "everything" = "7" } } variable "empty_map" { type = "map" default = {} } variable "nested_map" { type = "map" default = { "map_one" = { "key_1-1" = "value_1-1" "key_1-2" = "value_1-2" } "map_two" = { "key_2-1" = "value_2-1" "key_2-2" = "value_2-2" } } } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_map_value_output" { value = "\${keys(var.map_value)}" } output "2_map_function_output" { value = "\${keys(map("life",list("42"),"the universe",list("six","times","seven")))}" } output "3_empty_map_output" { value = "\${keys(var.empty_map)}" } output "4_nested_map_output" { value = "\${keys(var.nested_map)}" } Run the following from the keys folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? There’s a good chance you are going to need the keys in a map to help look up values. Especially if you don’t know what is going to be in that map to begin with. Once you have extracted the keys into a list, now you can use something like contains or element to parse through that list in a conditional logic statement or in a count loop. Lots of fun to be had there.
|
||
Lessons Learned The function does exactly as advertised. It doesn’t mind an empty map or a nested map. With the nested map, you are only going to get the keys of the top level map, and not the keys of the nested maps. Not surprising, but good to know.
|
||
Coming up next is the length() function. That might end up being a long post, har-har-har.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the keys() function. The example file is on GitHub here.
|
||
What is it? Function name: keys(map)
|
||
Returns: Takes a map and returns the keys of that map in a list that has been lexically sorted.
|
||
Example:
|
||
variable "map" { type = "map" default = { "one" = "1" "two" = "2" } } # Returns ["one","two"] output "key_output" { value = "\${keys(var.`,date:"26 Jul, 2018",url:"https://nedinthecloud.com/2018/07/26/terraform-fotd-keys/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/26/terraform-fotd-length/":{title:"Terraform - FotD - length()",tags:["fotd","hashicorp","terraform"],content:"This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the length() function. The example file is on GitHub here.\nWhat is it? Function name: length(value)\nReturns: Takes a value of string, list, or map and returns the number of elements in the value. For a string it is the number of characters, a list is the number of first level elements, and a map is the number of first level key-value pairs.\nExample:\nvariable "string" { type = "string" default = "Four" } # Returns 4 output "length_output" { value = "${length(var.string)}" } Example file: ############################################## # Function: length ############################################## ############################################## # Variables ############################################## variable "simple_value" { default = "Ford" } variable "simple_list" { type = "list" default = ["So", "long", "and", "thanks"] } variable "nested_list" { type = "list" default = [ ["So"], ["Long"], ["and"], ["thanks!"], ] } variable "mixed_list" { type = "list" default = [ [ ["3-1"], "2-1", ], [ ["3-2"], "2-2", ], "1-1", ] } variable "bool_value" { default = true } variable "map_value" { type = "map" default = { "life" = "42" "universe" = "6" "everything" = "7" } } variable "empty_map" { type = "map" default = {} } variable "nested_map" { type = "map" default = { "map_one" = { "key_1-1" = "value_1-1" "key_1-2" = "value_1-2" } "map_two" = { "key_2-1" = "value_2-1" "key_2-2" = "value_2-2" } } } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_simple_value_output" { value = "${length(var.simple_value)}" } output "2_simple_list_output" { value = "${length(var.simple_list)}" } output "3_nested_list_output" { value = "${length(var.nested_list)}" } output "4_mixed_list_output" { value = "${length(var.mixed_list)}" } output "5_bool_value_output" { value = "${length(var.bool_value)}" } output "6_map_output" { value = "${length(var.map_value)}" } output "7_empty_map_output" { value = "${length(var.empty_map)}" } output "8_map_output" { value = "${length(var.nested_map)}" } Run the following from the length folder to get example output for a number of different cases:\n#All examples are in variables terraform apply Why use it? Length is one of those fundamental functions that you just assume will be there. There’s not a whole lot more to say about that. If you use Terraform for anything beyond the simplest of configs, you are going to end up using the length function for something. I guarantee it.\nLessons Learned The nice thing is that there is a single function for all three primitive types. I don’t have to use lengthlist and lengthmap or something similar. Regardless of the value type being passed, you will get the appropriate count of elements. This is also a potential danger for debugging, since Terraform is not going to throw an error if the wrong type is passed. That can lead to potential unexpected behavior. Probably important to keep in mind if you are applying length to data coming from a data source or third-party module.\nComing up next is the list() function.\n",summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the length() function. The example file is on GitHub here.
|
||
What is it? Function name: length(value)
|
||
Returns: Takes a value of string, list, or map and returns the number of elements in the value. For a string it is the number of characters, a list is the number of first level elements, and a map is the number of first level key-value pairs.`,date:"26 Jul, 2018",url:"https://nedinthecloud.com/2018/07/26/terraform-fotd-length/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/07/25/terraform-fotd-jsonencode/":{title:"Terraform - FotD - jsonencode()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the jsonencode() function. The example file is on GitHub here.
|
||
What is it? Function name: jsonencode(value)
|
||
Returns: Takes a value and returns a properly JSON encoded version of that value. The value could be a string, list, map, or combination of primitives.
|
||
Example:
|
||
variable "list" { default = ["one","two","three"] } # Returns ["one","two","three"] output "json_output" { value = "\${jsonencoded(var.list)}" } Example file: ############################################## # Function: jsonencode ############################################## ############################################## # Variables ############################################## variable "simple_value" { default = "Ford" } variable "simple_list" { default = ["So", "long", "and", "thanks"] } variable "nested_list" { default = [ ["So"], ["Long"], ["and"], ["thanks!"], ] } variable "mixed_list" { default = [ [ ["3-1"], "2-1", ], [ ["3-2"], "2-2", ], "1-1", ] } variable "bool_value" { default = true } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_simple_value_output" { value = "\${jsonencode(var.simple_value)}" } output "2_simple_list_output" { value = "\${jsonencode(var.simple_list)}" } output "3_nested_list_output" { value = "\${jsonencode(var.nested_list)}" } output "4_mixed_list_output" { value = "\${jsonencode(var.mixed_list)}" } output "5_bool_value_output" { value = "\${jsonencode(var.bool_value)}" } output "map_output" { value = "\${jsonencode(map("life",list("42"),"the universe",list("six","times","seven")))}" } Run the following from the jsonencode folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? If the resource you are sending information to needs to be formatted correctly in JSON then I guess this is a good function to have in your back pocket. I would envision using this with the HTTP resource or an external resource that uses a custom script.
|
||
Lessons Learned The function appears to be pretty boring. It’s basically outputting the value I know is in the variable in the format I used. But the important thing here is you may need that value as a string in JSON to pass along. If it’s a list or a map, you may need the whole list or map as a JSON encoded string, not the individual values in the primitive. You know who loves JSON? Lamba functions. Love it, live it, script it.
|
||
Coming up next is the keys() function. Yeah, yeah, yeah. We’re getting into maps.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the jsonencode() function. The example file is on GitHub here.
|
||
What is it? Function name: jsonencode(value)
|
||
Returns: Takes a value and returns a properly JSON encoded version of that value. The value could be a string, list, map, or combination of primitives.`,date:"25 Jul, 2018",url:"https://nedinthecloud.com/2018/07/25/terraform-fotd-jsonencode/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/23/terraform-fotd-join/":{title:"Terraform - FotD - join()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the join() function. The example file is on GitHub here.
|
||
What is it? Function name: join(delimiter, list)
|
||
Returns: Takes a list and joins all the elements of the list using the delimiter. Returns the resultant string.
|
||
Example:
|
||
variable "list" { default = ["one","two","three"] } # Returns "one-two-three" output "join_output" { value = "\${join("-",var.list)}" } Example file: ############################################## # Function: join ############################################## ############################################## # Variables ############################################## variable "string_list" { default = ["So", "long", "and", "thanks"] } variable "int_list" { default = [41, 42, 43] } variable "empty_string_list" { default = ["", "", ""] } variable "empty_list" { default = [] } variable "bool_list" { default = [false, true] } variable "nested_list" { default = [[41], [42]] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## #Test a standard list output "1_join_list" { value = "\${join(" ",var.string_list)}" } #Test length output "1b_join_list" { value = "\${join(",",var.string_list)}" } #Test wrap output "1c_join_list" { value = "\${join("+",var.string_list)}" } #Test a standard int list output "2_int_list" { value = "\${join("\\n",var.int_list)}" } #Test list of empty strings output "3_empty_string_list" { value = "\${join(" ",var.empty_string_list)}" } #Test what will happen with a list boolean values output "4_bool_list" { value = "\${join("\\\\",var.bool_list)}" } output "5_nested_list_flatten" { value = "\${join("-",flatten(var.nested_list))}" } Run the following from the join folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? This function is useful if you need to pass a list to another module or an outside source that doesn’t natively support a list object. Once the constructed string gets to the other side, a mechanism can pull it apart. There is a split function as well that can re-create the list. The output of a module doesn’t currently support a list, so you might join the list for the output and split it apart on the other side.
|
||
Lessons Learned The function explicitly doesn’t deal with nested lists, so you’ll need to use flatten first. An example is provided. I also checked to see what escape sequences I could use as a delimiter. The only ones that seemed to work are newline (’\\n’) and backslash (’\\’). A sequence like (’\\t’) resulted in an error that the escape sequence is invalid. The delimiter can be multiple characters, so if you suspect that your list might have a common character, like a comma, already in it, then you could use a combination of characters instead.
|
||
Coming up next is the jsonencode() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the join() function. The example file is on GitHub here.
|
||
What is it? Function name: join(delimiter, list)
|
||
Returns: Takes a list and joins all the elements of the list using the delimiter. Returns the resultant string.
|
||
Example:
|
||
variable "list" { default = ["one","two","three"] } # Returns "one-two-three" output "join_output" { value = "\${join("-",var.`,date:"23 Jul, 2018",url:"https://nedinthecloud.com/2018/07/23/terraform-fotd-join/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/07/20/terraform-fotd-indent/":{title:"Terraform - FotD - indent()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the indent() function. The example file is on GitHub here.
|
||
What is it? Function name: indent(spaces, string)
|
||
Returns: Takes a string and indents all lines after the first line by the number of spaces specified. Returns the string.
|
||
Example:
|
||
variable "string" { default = "My\\nmulti-line\\nstring" } variable "spaces" { default = 2 } # Returns "My\\n multi-line\\n string" output "indent_output" { value = "\${indent(var.spaces,var.string)}" } Example file: ############################################## # Function: indent ############################################## ############################################## # Variables ############################################## variable "indent" { default = "This is line one\\nThis is line two\\nThis is line three" } variable "spaces" { default = 4 } variable "sourcefile" { default = "textFile.txt" } ############################################## # Resources ############################################## data "local_file" "source" { filename = "\${var.sourcefile}" } ############################################## # Outputs ############################################## output "indent_output" { value = "\${indent(var.spaces,var.indent)}" } output "file_output" { value = "\${indent(var.spaces, data.local_file.source.content)}" } Run the following from the indent folder to get example output for a number of different cases:
|
||
#We have to get the local resource first terraform init #Default values terraform apply #negative spaces will crash Terraform terraform apply -var "spaces=-1" #Empty string test works just fine terraform apply -var "indent=" Why use it? This function is just completely baffling to me. I guess it has something to do with YAML? I don’t know. There are so many possibly useful functions out there, and this seems to be frivolous at best. How about a function hard types a variable? That would be nice. PLEASE let me know if you use this function and why.
|
||
Lessons Learned There’s not much to learn here. Well, I did learn that if you give the function a negative value for the spaces it will crash Terraform. So that’s bug number three for me! Beyond that, the function does exactly what you would expect.
|
||
Coming up next is the index() function. That’s one I’ve used before. Very useful!
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the indent() function. The example file is on GitHub here.
|
||
What is it? Function name: indent(spaces, string)
|
||
Returns: Takes a string and indents all lines after the first line by the number of spaces specified. Returns the string.`,date:"20 Jul, 2018",url:"https://nedinthecloud.com/2018/07/20/terraform-fotd-indent/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/20/terraform-fotd-index/":{title:"Terraform - FotD - index()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the index() function. The example file is on GitHub here.
|
||
What is it? Function name: index(list, value)
|
||
Returns: Takes a list and looks for a value in the list matching the submitted value. Returns the index number of the first item found in the list with a matching value.
|
||
Example:
|
||
variable "list" { default = ["one","two","three"] } # Returns 1 output "index_output" { value = "\${index(var.list,"two")}" } Example file: ############################################## # Function: index ############################################## ############################################## # Variables ############################################## variable "string_list" { default = ["So", "long", "and", "thanks", "So"] } variable "int_list" { default = [41, 42, 43, 42] } variable "empty_string_list" { default = ["", "", ""] } variable "empty_list" { default = [] } variable "bool_list" { default = [false, true] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## #Test a standard list with more than one matching element output "1a_index_list" { value = "\${index(var.string_list,"So")}" } #Test a standard int list output "2_int_list" { value = "\${index(var.int_list,42)}" } #Test list of empty strings output "3_empty_string_list" { value = "\${index(var.empty_string_list,"")}" } #Test what will happen with boolean values output "4_bool_list" { value = "\${index(var.bool_list,1)}" } Run the following from the index folder to get example output for a number of different cases:
|
||
#Default examples terraform apply #Cannot find list element terraform apply -var "string_list=[]" #int or string doesn't matter terraform apply -var 'int_list=["41","42"]' Why use it? This function is sort of the inverse of the element function. Rather than supplying an index and getting the element, we are supplying the element and getting the index. I can only assume you would want to then slice up the list in some way based off where an element is in the list. To that end, you would probably want to do a little list hygiene beforehand. I imagine you would use things like sort, chomp, and distinct to do that.
|
||
Lessons Learned The function explicitly doesn’t deal with nested lists, so you’ll need to use flatten first. Boolean values are converted to 0 and 1, so you will need to search the list for 0 or 1. The function doesn’t really differentiate between ints and strings, so it works just as well to search for a string in a list of ints. The function only returns the first index of the value, even if the value appears multiple times. You could probably use that to your advantage if you have combined multiple lists that have a common ending element. Now you can get the index of where the original first list ended, and slice it off.
|
||
Coming up next is the join() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the index() function. The example file is on GitHub here.
|
||
What is it? Function name: index(list, value)
|
||
Returns: Takes a list and looks for a value in the list matching the submitted value. Returns the index number of the first item found in the list with a matching value.`,date:"20 Jul, 2018",url:"https://nedinthecloud.com/2018/07/20/terraform-fotd-index/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/07/19/terraform-fotd-formatlist/":{title:"Terraform - FotD - formatlist()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the formatlist() function. The example file is on GitHub here.
|
||
What is it? Function name: formatlist(format, list, args…)
|
||
Returns: Takes a format value and a list, followed by additional arguments. The additional arguments must be either single values or a list of the same length as the original list. Returns a list with values formatted based on the format value.
|
||
Example:
|
||
variable "string_list" { default = ["one","two","three"] } variable "int_list" { default = [1,2,3] } # Returns ["one-01,"two-02","three-03"] output "formatlist_output" { value = "\${formatlist("%v-%02v",var.string_list,var.int_list)}" } Example file: ############################################## # Function: formatlist ############################################## locals { int_list_local = "\${list("1","2","3")}" } ############################################## # Variables ############################################## variable "string_1" { default = "Beeblebrox" } variable "string_list" { default = ["Ford", "Arthur", "Trillian"] } variable "int_1" { default = 42 } variable "int_list" { default = [1, 2, 3] } variable "float_1" { default = 3.14159 } variable "float_list" { default = [3.14159, 1.21, 42.42] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "formatlist_one_item" { value = "\${formatlist("%v-%v",var.string_list,var.string_1)}" } output "formatlist_same_items" { value = "\${formatlist("%v-%v",var.string_list,var.int_list)}" } output "formatlist_three_items" { value = "\${formatlist("%v is the answer to %v times %v",var.string_list,var.int_list,var.float_list)}" } Run the following from the formatlist folder to get example output for a number of different cases:
|
||
#Using defaults terraform apply #Try using a string instead of list terraform apply -var "string_list='cat'" #Try using an empty list, may crash Terraform terraform apply -var "string_list=[]" #Try two different size lists, will fail on a mismatch terraform apply -var "int_list=[1,2]" Why use it? It’s pretty likely you’re going to have a list of values that you would like to apply some formatting to before passing it along to another resource, module, or output. Maybe you want to take all the names of load balancers you created and create a link for each one mapped to a port in another list. This function let’s you create those links and forward them along. I could also see some use cases for taking an environment name and appending it to a bunch of list values, and then adding it to a metadata tag on an object.
|
||
Lessons Learned Okay, for starters, learning to use the Sprintf formatting is a whole blog post in and of itself. Possibly many blog posts! If you find the documentation confusing, since you are probably not a Go developer by trade, then check out my post on the format function, where I dive into a few things that were mysterious to me.
|
||
Much like the format function, Terraform will submit all the values in variables as a string which makes using int or float specific formatting a bit challenging. I also discovered that you cannot use two different size lists, which is pretty obvious, but not necessarily intuitive at first. Even more exciting, I was able to make Terraform crash by giving it an empty list to process. Second bug found! I’ll be submitting that one shortly.
|
||
Coming up next is the indent() function. Probably has something to do with YAML. Yuck.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the formatlist() function. The example file is on GitHub here.
|
||
What is it? Function name: formatlist(format, list, args…)
|
||
Returns: Takes a format value and a list, followed by additional arguments. The additional arguments must be either single values or a list of the same length as the original list.`,date:"19 Jul, 2018",url:"https://nedinthecloud.com/2018/07/19/terraform-fotd-formatlist/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/07/17/terraform-fotd-format/":{title:"Terraform - FotD - format()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the format() function. The example file is on GitHub here.
|
||
What is it? Function name: format(format, args,…)
|
||
Returns: Takes a format value and one or more arguments that are part of the format. Returns a value formatted as specified in the format value. Uses the format style from Sprintf in Golang.
|
||
Example:
|
||
variable "format" { default = "string" } # Returns 737472696e67 which is base 16 lowercase of the string output "format_output" { value = "\${format("%x",var.format)}" } Example file: ############################################## # Function: format ############################################## ############################################## # Variables ############################################## variable "string_1" { default = "example string" } variable "int_1" { default = 42 } variable "float_1" { default = 3.14159 } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "format_string" { value = "\${format("%q",var.string_1)}" } output "format_string_16byte" { value = "\${format("%.2X",var.string_1)}" } #Addition forces int value type output "format_int_base2" { value = "\${format("%b",var.int_1 + 0)}" } #Addition forces float value type output "format_float_scientific" { value = "\${format("%E",var.float_1 + 0.0)}" } output "format_float_precision" { value = "\${format("%+.3f",var.float_1 + 0.0)}" } output "string_int_combo" { value = "\${format("%v-%03d",var.string_1,var.int_1 + 0)}" } output "format_bool" { value = "\${format("%t",true)}" } Run the following from the format folder to get example output for a number of different cases:
|
||
#Using defaults terraform apply #Try a negative float terraform apply -var "float_1=-3.14159" #Try a negative int terraform apply -var "int_1=-42" #Try a different string terraform apply -var "string_1=Trillian" Why use it? Formatting probably works best when you have multiple arguments you want to munge together and apply a little formatting to. A good example might be naming a resource, where there are specific requirements around how the name must be formatted. You may also need to manipulate floating values to lower the precision or be converted to scientific notation. I’m not sure exactly when you would need this, but there’s no doubt that it will come up eventually.
|
||
Lessons Learned Okay, for starters, learning to use the Sprintf formatting is a whole blog post in and of itself. Possibly many blog posts! If you find the documentation confusing, since you are probably not a Go developer by trade, then let me explain a couple things that were mysterious to me. The formatting language uses a percentage sign to start a formatting block. The value being formatted is the argument being passed to the format function. Unless you explicitly specify an evaluation order, it will use each argument in the order they were given. For example, something like this:
|
||
format("%s%d","string",42) The %s will format the first argument - string in this case - using the ’s’ format which explicitly deals with strings. The %d will format the second argument - 42 in this case - using the ’d’ format which formats an integer as base 10.
|
||
You can add additional formatting to each block. For instance adding a ‘+’ to %+d tells the formatter to always display a ‘+’ sign for positive values. That might be important if spacing of the value regardless of sign is consistent. There’s a ton more, but hopefully that clears up the initial haze a bit.
|
||
I also learned in the process that Terraform will submit all the values in variables as a string, unless you force an implicit conversion. The way I found to do this is to perform a math operation on the variable. By doing something like:
|
||
var.int + 0 Terraform converts the string to an int, even though the value is not altered. Resources and data sources natively return a correct data type, so this is really only useful for variables.
|
||
Coming up next is the formatlist() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the format() function. The example file is on GitHub here.
|
||
What is it? Function name: format(format, args,…)
|
||
Returns: Takes a format value and one or more arguments that are part of the format. Returns a value formatted as specified in the format value.`,date:"17 Jul, 2018",url:"https://nedinthecloud.com/2018/07/17/terraform-fotd-format/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/07/16/terraform-fotd-flatten/":{title:"Terraform - FotD - flatten()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the flatten() function. The example file is on GitHub here.
|
||
What is it? Function name: flatten(list)
|
||
Returns: Takes a list with nested values and returns a flat list with all values removing nested lists.
|
||
Example:
|
||
variable "flatten" { default = ["one",["two",3]] } # Returns ["one","two",3] output "flatten_output" { value = "\${flatten(var.flatten)}" } Example file: ############################################## # Function: flatten ############################################## ############################################## # Variables ############################################## variable "list_1" { default = [] } variable "list_2" { default = [[], [], []] } variable "list_3" { default = [0, [1, 2, 3], [4, 5, 6], [7, 8], 9, 0] } variable "list_4" { default = ["So", "Long", "and", "Thanks"] } variable "list_5" { default = [[[1, 2, 3], [4, 5, 6]], [[7, 8], [9, 10]], 11, 12] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_flatten_empty_list" { value = "\${flatten(var.list_1)}" } output "2_flatten_list_of_empty_lists" { value = "\${flatten(var.list_2)}" } output "3_flatten_mixed_list" { value = "\${flatten(var.list_3)}" } output "4_flatten_flat_list" { value = "\${flatten(var.list_4)}" } output "5_flatten_double_nested_list" { value = "\${flatten(var.list_5)}" } Run the following from the flatten folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? Depending on the data source, or resource, you might get a messy nested list back. Say you have a list of security groups, and each security group has rules, and each rule has a set of allowed ports. If you want to know all the allowed ports in all the rules, then you could flatten the list of a list of a list to get that. Fun times really.
|
||
Lessons Learned I would have thought that the function would take multiple lists and flatten them all into one. But that is the job of the concat function. The concat function, if you’ll remember, doesn’t handle nested lists and flat lists in the same function call. So you may need to use flatten first before concatenating the lists together. And you might also use distinct and chomp to clean up your data. Basically, there’s a lot of good list functions.
|
||
Coming up next is the format() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the flatten() function. The example file is on GitHub here.
|
||
What is it? Function name: flatten(list)
|
||
Returns: Takes a list with nested values and returns a flat list with all values removing nested lists.
|
||
Example:
|
||
variable "flatten" { default = ["one",["two",3]] } # Returns ["one","two",3] output "flatten_output" { value = "\${flatten(var.`,date:"16 Jul, 2018",url:"https://nedinthecloud.com/2018/07/16/terraform-fotd-flatten/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/13/terraform-fotd-file/":{title:"Terraform - FotD - file()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the file() function. The example file is on GitHub here.
|
||
What is it? Function name: file(path)
|
||
Returns: Takes a path to a file and reads the content of the file into a string variable. Returns the string value.
|
||
Example:
|
||
variable "file" { default = "/path/to/file.txt" } # Returns string of file contents output "file_output" { value = "\${file(var.file)}" } Example file: ############################################## # Function: file ############################################## ############################################## # Variables ############################################## variable "file1" {} variable "file2" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "file_output" { value = "\${file(var.file1)}" } output "path_file_output" { value = "\${file("\${path.cwd}\\\\\${var.file2}")}" } Run the following from the file folder to get example output for a number of different cases:
|
||
terraform apply -var 'file1=input.txt' -var "file2=input.txt" terraform apply -var 'file1=.\\\\input.txt' -var "file2=input.txt" terraform apply -var 'file1=.\\\\module\\\\mod_input.txt' -var "file2=input.txt" terraform apply -var "file1=.\\\\module\\\\input-noext" -var "file2=input.txt" Why use it? I’ve definitely had to pull file contents into Terraform to use with other resources. Sometimes you have to copy a file to a destination. Other times the file contains a template that you will use with the template resource. You might also pull in a CSV or something similar to work as a data source.
|
||
Lessons Learned The function will read in the file as a single string and will not perform any interpolation. If you want to do something like that, you will need to use the template resource to leverage those functions. You can also use the path.PATH built-in variable to reference the path in the Terraform configuration. The way to call it is a little weird, so I included an example in the output values. The function doesn’t care if the file has an extension on it. You cannot give it a directory, as it will throw an error.
|
||
Coming up next is the floor() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the file() function. The example file is on GitHub here.
|
||
What is it? Function name: file(path)
|
||
Returns: Takes a path to a file and reads the content of the file into a string variable. Returns the string value.`,date:"13 Jul, 2018",url:"https://nedinthecloud.com/2018/07/13/terraform-fotd-file/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/13/terraform-fotd-floor/":{title:"Terraform - FotD - floor()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the file() function. The example file is on GitHub here.
|
||
What is it? Function name: floor(float)
|
||
Returns: Rounds down to the closest integer, unless the float is already an integer value. Will attempt to perform implicit conversion of other data types to float prior to conversion.
|
||
Example:
|
||
variable "floor" { default = "3.14159" } # Returns 3 output "floor_output" { value = "\${floor(var.floor)}" } Example file: ############################################## # Function: floor ############################################## ############################################## # Variables ############################################## variable "floor" { default = true } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "floor_output" { value = "\${floor(var.floor)}" } Run the following from the floor folder to get example output for a number of different cases:
|
||
terraform apply -var 'floor=5.9' terraform apply -var 'floor=-5.9' terraform apply -var 'floor=3.1' terraform apply -var 'floor=-3.1' terraform apply -var 'floor="-1"' terraform apply -var 'floor="String"' terraform apply Why use it? Floor is a fairly common mathematical device. If you are going to be manipulating float values and need an integer, floor might suit your needs.
|
||
Lessons Learned I learned in the process that there is an implicit conversion of boolean true to float as 1 and false to float as 0. This conversion only happens when using a default value or passing the value in a tfvars file. True or false passed in the -var argument will be treated as a string without conversion. Strings will be converted to float if it’s a number, but not if it’s regular text. Negative numbers are rounded down away from 0, so -5.9 is rounded down to -6.
|
||
Coming up next is the flatten() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the file() function. The example file is on GitHub here.
|
||
What is it? Function name: floor(float)
|
||
Returns: Rounds down to the closest integer, unless the float is already an integer value. Will attempt to perform implicit conversion of other data types to float prior to conversion.`,date:"13 Jul, 2018",url:"https://nedinthecloud.com/2018/07/13/terraform-fotd-floor/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/12/terraform-fotd-element/":{title:"Terraform - FotD - element()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the element() function. The example file is on GitHub here.
|
||
What is it? Function name: element(list, index)
|
||
Returns: Takes a list and an index value. Returns the element from the list with the submitted index value.
|
||
Example:
|
||
variable "element" { default = ["one","two","three"] } # Returns "one" output "element_output" { value = "\${element(var.element,0)}" } Example file: ############################################## # Function: element ############################################## ############################################## # Variables ############################################## variable "string_list" { default = ["So", "long", "and", "thanks"] } variable "int_list" { default = [41, 42, 43] } variable "empty_string_list" { default = ["", "", ""] } variable "empty_list" { default = [] } variable "bool_list" { default = [false, true] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## #Test a standard list output "1_element_list" { value = "\${element(var.string_list,1)}" } #Test length output "1b_element_list" { value = "\${element(var.string_list,length(var.string_list)-1)}" } #Test wrap output "1c_element_list" { value = "\${element(var.string_list,10)}" } #Test a standard int list output "2_int_list" { value = "\${element(var.int_list,2)}" } #Test list of empty strings output "3_empty_string_list" { value = "\${element(var.empty_string_list,2)}" } #Test what will happen with a list boolean values output "4_bool_list" { value = "\${element(var.bool_list,2)}" } Run the following from the element folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? There are a ton of cases I can think of where I want an element from a list. Many of them involve looping through the list. Let’s say you grab all the availability zones for an AWS region. You can alternate the placement of resources by using element in a count loop. The function will do a standard mod function on the index, so if you have 3 elements and the index value is 4, then you will get the 4 mod 3 element or the element at index 1. Remember that index values start at 0.
|
||
Lessons Learned The function explicitly doesn’t deal with nested lists, so you’ll need to use flatten first. Boolean values are converted to 0 and 1. While the function is happy to take a positive number greater than the number of elements, it will not accept a negative number. It will also not process an empty list. You may want to clean up your list ahead of time using functions like sort and compact.
|
||
Coming up next is the file() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the element() function. The example file is on GitHub here.
|
||
What is it? Function name: element(list, index)
|
||
Returns: Takes a list and an index value. Returns the element from the list with the submitted index value.
|
||
Example:
|
||
variable "element" { default = ["one","two","three"] } # Returns "one" output "element_output" { value = "\${element(var.`,date:"12 Jul, 2018",url:"https://nedinthecloud.com/2018/07/12/terraform-fotd-element/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/10/terraform-fotd-distinct/":{title:"Terraform - FotD - distinct()",tags:["fotd","hashicorp","terraform"],content:"This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the distinct() function. The example file is on GitHub here.\nWhat is it? Function name: distinct(list)\nReturns: Takes a list and removes any duplicate values, keeping the first of each duplicate value. Returns the updated list.\nExample:\nvariable "distinct" { default = ["one","two","three","one"] } # Returns ["one","two","three"] output "distinct_output" { value = "${distinct(var.distinct)}" } Example file: ############################################## # Function: distinct ############################################## ############################################## # Variables ############################################## variable "string_list" { default = ["So", "so", "long", "long", "so"] } variable "int_duplicates" { default = [42, 42, 42] } variable "string_and_int_list" { default = [42, "42", 42] } variable "empty_string_list" { default = ["", "", ""] } variable "empty_list" { default = [] } variable "bool_list" { default = [false, true, false, true] } variable "nested_list" { default = [[1, 2, 3], [1, 2, 3]] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## #Test a standard list output "1_distinct_list" { value = "${distinct(var.string_list)}" } #Test a standard int list output "2_int_duplicates" { value = "${distinct(var.int_duplicates)}" } #Test int and strings with same value output "3_string_and_int_list" { value = "${distinct(var.string_and_int_list)}" } #Test list of empty strings output "4_empty_string_list" { value = "${distinct(var.empty_string_list)}" } #Test what it will do with an empty list output "5_empty_list" { value = "${distinct(var.empty_list)}" } #Test what will happen with a list boolean values output "6_bool_list" { value = "${distinct(var.bool_list)}" } #Flatten a nested list and then distinct output "7_nested_list" { value = "${distinct(flatten(var.nested_list))}" } Run the following from the distinct folder to get example output for a number of different cases:\n#All examples are in variables terraform apply Why use it? This would be immensely useful when combining a bunch of lists from data sources and trying to extract only the unique records for other processes. I would use this in concert with other list functions, such as compact, concat, and flatten to clean up nested lists and empty values. In fact, I would probably run compact prior to distinct, so I can remove all empty strings, and then remove the duplicates. That seems more efficient from a computation perspective. The compact function only has to check if the value is empty, whereas the distinct has to check if that value exists in the list already.\nLessons Learned The function explicitly doesn’t deal with nested lists, so you’ll need to use flatten first. Boolean values are converted to 0 and 1. Strings and ints are treated as the same value, so “42” and 42 are considered the same. The function is case sensitive, so “Fish” and “fish” are not the same value.\nComing up next is the element() function. I’m partial to the element(var.movies,5) myself.\n",summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the distinct() function. The example file is on GitHub here.
|
||
What is it? Function name: distinct(list)
|
||
Returns: Takes a list and removes any duplicate values, keeping the first of each duplicate value. Returns the updated list.
|
||
Example:
|
||
variable "distinct" { default = ["one","two","three","one"] } # Returns ["one","two","three"] output "distinct_output" { value = "\${distinct(var.`,date:"10 Jul, 2018",url:"https://nedinthecloud.com/2018/07/10/terraform-fotd-distinct/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/07/09/terraform-fotd-dirname/":{title:"Terraform - FotD - dirname()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the dirname() function. The example file is on GitHub here.
|
||
What is it? Function name: dirname(path)
|
||
Returns: Takes a file path and returns the directory path containing the file, without the file name.
|
||
Example:
|
||
variable "dirname" { default = "/bin/bash" } # Returns /bin output "dirname_output" { value = "\${dirname(var.dirname)}" } Example file: ############################################## # Function: dirname ############################################## ############################################## # Variables ############################################## variable "dirname" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "dirname_output" { value = "\${dirname(var.dirname)}" } Run the following from the dirname folder to get example output for a number of different cases:
|
||
#Windows tests #Try the root directory terraform apply -var 'dirname=C:\\\\' #Try without a file extension terraform apply -var 'dirname=C:\\\\level0\\\\level1\\\\level2' #Try with trailing Windows backslash terraform apply -var 'dirname=C:\\\\level0\\\\level1\\\\level2\\\\' #Try with file terraform apply -var 'dirname=C:\\\\level0\\\\level1\\\\level2\\\\file1.txt' #Linux tests #Try the root directory terraform apply -var 'dirname=/' #Try without a file extension terraform apply -var 'dirname=/level0/level1/level2' #Try with trailing forward slash terraform apply -var 'dirname=/level0/level1/level2/' #Try with file terraform apply -var 'dirname=/level0/level1/level2/file1.txt' #Try empty value terraform apply -var 'dirname=' Why use it? Sometimes you need the directory path to put some other files in it. Other times you want to grab all the files from a directory given a particular file’s path. There are a ton of potential uses here. You could even combine this with basename to get the name of the directory holding a file.
|
||
Lessons Learned The function is happy working with either Windows or Linux file paths. For Windows you still need to double-escape the backslash (\\) since the backslash is the escape character for HCL. Whether the file has an extension or not doesn’t matter, the function is just looking for the last path separator and returning everything to the left of it. If you leave the slash on the end of the path, the function gives you everything to the left of the slash, as you might expect. If you give it an empty path, then it will give you a “.” representing the current directory, which I find clever.
|
||
Coming up next is the distinct() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the dirname() function. The example file is on GitHub here.
|
||
What is it? Function name: dirname(path)
|
||
Returns: Takes a file path and returns the directory path containing the file, without the file name.
|
||
Example:
|
||
variable "dirname" { default = "/bin/bash" } # Returns /bin output "dirname_output" { value = "\${dirname(var.`,date:"9 Jul, 2018",url:"https://nedinthecloud.com/2018/07/09/terraform-fotd-dirname/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/06/terraform-fotd-contains/":{title:"Terraform - FotD - contains()",tags:["fotd","hashicorp","terraform"],content:"This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the contains() function. The example file is on GitHub here.\nWhat is it? Function name: contains(list, element)\nReturns: Takes a list and an element and tests if the element is in the list. Returns true if the element is present, false if it is not.\nExample:\nvariable "list" { default = ["My","list","of","stuff"] } # Returns true output "compact_list" { value = "${compact(var.list, "list")}" } Example file: ############################################## # Function: contains ############################################## ############################################## # Variables ############################################## variable "string_list" { default = ["So", "long", "and", "thanks"] } variable "string_duplicates" { default = [42, 42, 42] } variable "empty_list" { default = [] } variable "empty_string_list" { default = ["", "", ""] } variable "bool_list" { default = [false, true] } variable "nested_list" { default = [["one"], ["two"]] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## #Test a standard list output "1_contains_list" { value = "${contains(var.string_list,"So")}" } #Test a standard list to false output "2_contains_list_false" { value = "${contains(var.string_list,"fish")}" } #Test case sensitivity output "3_string_list_case" { value = "${contains(var.string_list,"so")}" } #Test duplicates output "4_string_list_duplicate" { value = "${contains(var.string_duplicates,42)}" } #Test what it will do with an empty list output "5_contains_empty_list" { value = "${contains(var.empty_list,"")}" } #Test what will happen with a list full of empty strings output "6_contains_empty_string_list" { value = "${contains(var.empty_string_list,"")}" } #Test how the function interprets boolean values output "7_contains_bool_list" { value = "${contains(var.bool_list, true)}" } output "8_contains_bool_list_2" { value = "${contains(var.bool_list, 1)}" } output "9_contains_nested_list" { value = "${contains(var.nested_list,"one")}" } Run the following from the contains folder to get example output for a number of different cases:\n#All examples are in variables terraform apply Why use it? This seems like a pretty fundamental conditional test. If the element is present, like a firewall rule, don’t create a new firewall rule. I’m sure there are a ton of other examples.\nLessons Learned The function will accept nested or flat lists, but if you try to pass it a list as the element value, then it throws an error. So although you can pass a nested list, you cannot test for the presence of a specific list in the nested list. It will always return false. I also learned that boolean values in the original list get converted to 0 and 1, so if you test for boolean true or false, it will return false since it only has 0s and 1s in the processed list. An empty list will always return false as well. The function is case sensitive, so fish and Fish are not the same thing. Duplicate values don’t seem to effect the function. If a value is present, then it is returns true, regardless of how many times that element appears.\nComing up next is the dirname() function.\n",summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the contains() function. The example file is on GitHub here.
|
||
What is it? Function name: contains(list, element)
|
||
Returns: Takes a list and an element and tests if the element is in the list. Returns true if the element is present, false if it is not.`,date:"6 Jul, 2018",url:"https://nedinthecloud.com/2018/07/06/terraform-fotd-contains/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/07/05/terraform-fotd-concat/":{title:"Terraform - FotD - concat()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the concat() function. The example file is on GitHub here.
|
||
What is it? Function name: concat(list1, list2, …)
|
||
Returns: Takes two or more lists and combines them into a single list. Returns the combined list.
|
||
Example:
|
||
variable "list" { default = ["My","list","of","stuff"] } variable "list2" { default = ["is","very","long"] } # Returns ["My","list","of","stuff","is","very","long"] output "concat_list" { value = "\${concat(var.list, var.list2)}" } Example file: ############################################## # Function: concat ############################################## ############################################## # Variables ############################################## variable "list_1" { default = [] } variable "list_2" { default = [6, 7, 42] } variable "list_3" { default = [false, false, false] } variable "list_4" { default = ["So", "Long", "and", "Thanks"] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "concat_output" { value = "\${concat(var.list_1,var.list_2,var.list_3,var.list_4)}" } Run the following from the concat folder to get example output for a number of different cases:
|
||
#List of different value types is fine terraform apply #Cannot combine nested lists and regular lists terraform apply -var-file="example1.tfvars" #All nested lists works terraform apply -var-file="example1.tfvars" -var "list_4=[]" Why use it? If you are working with multiple resources, you might get back a list from each one. You can use this function to combine all those lists, and then use something like the distinct function to remove any duplicates. An example that comes to mind is a list of instance sizes available in different Azure regions. I might want some common denominator size, in case there are special sizes in each region.
|
||
Lessons Learned The function will accept lists with values or nested lists. If you want to combine nested lists, you cannot have a flat list as well. There is a flatten function you could use to flatten nested lists before concatenating them. Also, I had assumed that concat was going to be a function for strings and not lists, and I’m not the first one. The function should probably be called concatlist() instead, though I supposed it’s too late for that now.
|
||
Coming up next is the contains() function, I can hardly contain my excitement!
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the concat() function. The example file is on GitHub here.
|
||
What is it? Function name: concat(list1, list2, …)
|
||
Returns: Takes two or more lists and combines them into a single list. Returns the combined list.
|
||
Example:
|
||
variable "list" { default = ["My","list","of","stuff"] } variable "list2" { default = ["is","very","long"] } # Returns ["My","list","of","stuff","is","very","long"] output "concat_list" { value = "\${concat(var.`,date:"5 Jul, 2018",url:"https://nedinthecloud.com/2018/07/05/terraform-fotd-concat/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/04/terraform-fotd-compact/":{title:"Terraform - FotD - compact()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the compact() function. The example file is on GitHub here.
|
||
What is it? Function name: compact(list)
|
||
Returns: Takes a list and removes any empty string values in the list.
|
||
Example:
|
||
variable "list" { default = ["My","list","","of","","stuff"] } # Returns ["My","list","of","stuff"] output "compact_list" { value = "\${compact(var.list)}" } Example file: ############################################## # Function: compact ############################################## ############################################## # Variables ############################################## variable "list" { default = ["", "so", "long", "", "and", " ", "thanks", ""] } variable "empty_list" { default = [] } variable "empty_string_list" { default = ["", "", ""] } variable "bool_list" { default = [false, "zaphod", "", true, 0, 1] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## #Test a standard list output "compact_list" { value = "\${compact(var.list)}" } #Test what it will do with an empty list output "compact_empty_list" { value = "\${compact(var.empty_list)}" } #Test what will happen with a list full of empty strings output "compact_empty_string_list" { value = "\${compact(var.empty_string_list)}" } #Test how the function interprets boolean values output "compact_bool_list" { value = "\${compact(var.bool_list)}" } Run the following from the compact folder to get example output for a number of different cases:
|
||
#All examples are in variables terraform apply Why use it? One thing I learned early on, is to sanitize your input. Assume that any module or user is going to give you crappy input, so try and clean it up. Compact could be used with other functions to clean up list data, making sure there are no empty values. You might also run something like distinct() to remove duplicate values.
|
||
Lessons Learned The function will only accept a flat list, no nested lists allowed. If you suspect you’re going to get a nested list, use the flatten function first to un-nest those lists. As usual, boolean false evaluates to 0. Whitespace does not count as an empty string, so a list with the string " “, will not remove that element. The function is fine with an empty list or a list of just empty strings.
|
||
Coming up next is the concat() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the compact() function. The example file is on GitHub here.
|
||
What is it? Function name: compact(list)
|
||
Returns: Takes a list and removes any empty string values in the list.
|
||
Example:
|
||
variable "list" { default = ["My","list","","of","","stuff"] } # Returns ["My","list","of","stuff"] output "compact_list" { value = "\${compact(var.`,date:"4 Jul, 2018",url:"https://nedinthecloud.com/2018/07/04/terraform-fotd-compact/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/07/03/terraform-fotd-coalescelist/":{title:"Terraform - FotD - coalescelist()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the coalescelist() function. The example file is on GitHub here.
|
||
What is it? Function name: coalescelist(list,list,…)
|
||
Returns: Takes two or more lists and returns the first non-empty list from the arguments. Must include at least two arguments to evaluate. Does not accept a list of lists as multiple arguments.
|
||
Example:
|
||
variable "list" { default = ["My","list","of","stuff"] } variable "empty_list" { default = [] } # Returns ["My","list","of","stuff"] output "coalescelist" { value = "\${coalescelist(var.empty_list,var.list)}" } Example file: ############################################## # Function: coalescelist ############################################## ############################################## # Variables ############################################## variable "list_1" { default = [] } variable "list_2" { default = [[], [], []] } variable "list_3" { default = [false, false, false] } variable "list_4" { default = ["So", "Long", "and", "Thanks"] } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "coalescelist_output" { value = "\${coalescelist(var.list_1,var.list_2,var.list_3,var.list_4)}" } Run the following from the coalescelist folder to get example output for a number of different cases:
|
||
#Evaluate to "one", "two", "three"] terraform apply -var-file="example1.tfvars" #Is an list of empty lists empty? Nope, list_2 will be returned terraform apply #What about a list of boolean false values? Evaluates to all 0s terraform apply -var "list_2=[]" #All empty, evaluate to empty terraform apply -var "list_2=[]" -var "list_3=[]" -var "list_4=[]" Why use it? I have to admit I was a little stumped on when you would actually use this. It seems like you would need a situation where you would check the output of multiple resources, and then be okay with selecting the first non-empty value as the value you would want to use. I am sure there is a situation for this, but I have no idea what it might be.
|
||
Lessons Learned The function will take a list of lists as a value, but it doesn’t expand the list of lists. The function will evaluate the boolean value false in a list as an int value of 0. If you have a list of boolean false values or a list of empty lists [] those both count as values in a list. I also discovered that passing a list of strings via the var command line argument doesn’t work so well in Windows.
|
||
terraform apply -var 'list_1=["one","two","three"]' That command errors out for me when it tries to parse the list. So I had to create a tfvars file to test that combination. Other than that, it does precisely what it advertises. I do find the naming a little confusing, since coalesce means to combine multiple things into a single entity. But then I found that there is a Coalesce function in T-SQL that also returns the first non-null value in a list. I suppose in the world of programming, coalesce has this other established definition.
|
||
Coming up next is the compact() function.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the coalescelist() function. The example file is on GitHub here.
|
||
What is it? Function name: coalescelist(list,list,…)
|
||
Returns: Takes two or more lists and returns the first non-empty list from the arguments. Must include at least two arguments to evaluate.`,date:"3 Jul, 2018",url:"https://nedinthecloud.com/2018/07/03/terraform-fotd-coalescelist/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/07/02/terraform-fotd-coalesce/":{title:"Terraform - FotD - coalesce()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the coalesce() function. The example file is on GitHub here.
|
||
What is it? Function name: coalesce(string,string,…)
|
||
Returns: Takes two or more strings and returns the first non-empty string in the list. Must include at least two arguments to evaluate. Does not accept a list of strings.
|
||
Example:
|
||
variable "string" { default = "some_string" } variable "empty_string" { default = "" } # Returns some_string output "coalesce" { value = "\${coalesce(var.empty_string,var.string)}" } Example file: ############################################## # Function: coalesce ############################################## ############################################## # Variables ############################################## variable "1_string" { default = false } variable "2_string" { default = 0 } variable "3_string" { default = "three" } variable "4_string" { default = "four" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "coalesce_output" { value = "\${coalesce(var.1_string,var.2_string,var.3_string,var.4_string)}" } Run the following from the coalesce folder to get example output for a number of different cases:
|
||
#Evaluate to one terraform apply -var "1_string=one" -var "2_string=two" -var "3_string=three" -var "4_string=four" #First string is empty, evaluate to second string terraform apply -var "1_string=" -var "2_string=two" -var "3_string=three" -var "4_string=four" #What will a default value of bool false do? Evaluate as string 0 terraform apply -var "2_string=two" -var "3_string=three" -var "4_string=four" #What will a default value of int 0 do? Evaluate as string 0 terraform apply -var "1_string=" -var "3_string=three" -var "4_string=four" #All empty, evaluate to empty terraform apply -var "1_string=" -var "2_string=" -var "3_string=" -var "4_string=" Why use it? I have to admit I was a little stumped on when you would actually use this. It seems like you would need a situation where you would check the output of multiple resources, and then be okay with selecting the first non-empty value as the value you would want to use. I am sure there is a situation for this, but I have no idea what it might be.
|
||
Lessons Learned The function will not take a list as a value, only multiple strings, which I found really odd. The function will evaluate the boolean value false as an empty string. It evaluates the int value of 0 as a string value of 0. Other than that, it does precisely what it advertises. I do find the naming a little confusing, since coalesce means to combine multiple things into a single entity. But then I found that there is a Coalesce function in T-SQL that also returns the first non-null value in a list. I suppose in the world of programming, coalesce has this other established definition.
|
||
Coming up next is the coalescelist() function, you’ll never guess what it does.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the coalesce() function. The example file is on GitHub here.
|
||
What is it? Function name: coalesce(string,string,…)
|
||
Returns: Takes two or more strings and returns the first non-empty string in the list. Must include at least two arguments to evaluate.`,date:"2 Jul, 2018",url:"https://nedinthecloud.com/2018/07/02/terraform-fotd-coalesce/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/06/29/terraform-fotd-cidrsubnet/":{title:"Terraform - FotD - cidrsubnet()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the cidrsubnet() function. The example file is on GitHub here.
|
||
What is it? Function name: cidrsubnet(iprange,newbits,netnum)
|
||
Returns: Takes an IP address range in CIDR notation and adds the newbits to the the subnet mask. Then it finds the network in the original CIDR range with the network number in netnum. It returns the new subnet in CIDR notation. This is confusing at first, but it’s really meant to divide up a larger subnet into smaller subnets.
|
||
Example:
|
||
variable "iprange" { default = "10.0.0.0/16" } # Returns 10.0.1.0/24 output "cidrsubnet" { value = "\${cidrsubnet(var.iprange,8,1)}" } Example file: ############################################## # Function: cidrsubnet ############################################## ############################################## # Variables ############################################## variable "iprange" { default = "10.0.0.0/16" } variable "newbits" { default = 4 } variable "netnum" { default = 0 } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_iprange" { value = "\${var.iprange}" } output "2_newbits" { value = "\${var.newbits}" } output "3_netnum" { value = "\${var.netnum}" } output "4_cidrsubnet_output" { value = "\${cidrsubnet(var.iprange,var.newbits,var.netnum)}" } Run the following from the cidrsubnet folder to get example output for a number of different cases:
|
||
terraform apply -var "iprange=172.16.0.0/24" -var "newbits=2" -var "netnum=0" terraform apply -var "iprange=172.16.0.0/24" -var "newbits=4" -var "netnum=2" terraform apply -var "iprange=172.16.0.0/24" -var "newbits=-8" -var "netnum=20" terraform apply -var "iprange=10.0.0.0/8" -var "newbits=16" -var "netnum=40" #Fails with negative number for netnum terraform apply -var "iprange=172.16.0.0/24" -var "newbits=2" -var "netnum=-1" #Fails with invalid subnet net number terraform apply -var "iprange=172.16.0.0/24" -var "newbits=2" -var "netnum=255" #Results in the 0.0.0.0 address range terraform apply -var "iprange=172.16.0.0/0" -var "newbits=2" -var "netnum=0" Why use it? Whenever I am creating a new network in AWS or Azure and need to parcel out subnets, this function is my best friend. My most common use case is using the cidrsubnet function with the count.index to split a large subnet range into private and public subnets in a VPC.
|
||
Lessons Learned The function does not like getting a negative number for the netnum. I thought it would count backwards from the highest subnet, like cidrhost does, but it errors out instead. The interesting thing I found was submitting a CIDR network with 0 as the subnet mask results in the function returning 0.0.0.0 for the network. It doesn’t handle that value well, so maybe don’t do that. It is also possible to submit a negative number for the newbits, though I have no idea why you would want to do that.
|
||
Coming up next is the coalesce() function, which is one of my favorite words and bands.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the cidrsubnet() function. The example file is on GitHub here.
|
||
What is it? Function name: cidrsubnet(iprange,newbits,netnum)
|
||
Returns: Takes an IP address range in CIDR notation and adds the newbits to the the subnet mask. Then it finds the network in the original CIDR range with the network number in netnum.`,date:"29 Jun, 2018",url:"https://nedinthecloud.com/2018/06/29/terraform-fotd-cidrsubnet/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/06/27/running-ansible-on-windows-10/":{title:"Running Ansible on Windows 10",tags:["ansible","docker","hpe","terraform"],content:`I’ve been going through Ansible training for work, and I want to be able to easily try out some of the playbooks and modules on my Windows 10 box. There’s lots of ways I could go about this, but for me the thing that makes the most sense is to leverage Docker for Windows (DfW) and the magic of containers to be able to get up and running quickly. In this post I will walk through what I did to get started, and why I landed on containers instead of another solution.
|
||
At my day job, we are working on building demos for our HPE Synergy system using Ansible for automation. Obviously, in order to run the demos I need something that can run Ansible playbooks and has access to the API endpoint for the OneView Composer that manages the Synergy frame. I’ve used Ansible before for some small projects and in my Deep Dive Terraform course on Pluralsight. In those cases, I was creating a Linux host and having it run the Ansible playbooks on itself, so I didn’t need a traditional controller node. In the case of OneView, I now need a controller node that can reach out to a system and configure it via API calls.
|
||
I am also going through Ansible training, and while Red Hat does provide a lab environment, I also want a local instance I can try things out on.
|
||
I reviewed my options for a local instance and here’s what I came up with:
|
||
Run a VM locally on Hyper-V: I don’t really want to run a full-blown VM. Plus I would have to download an ISO for a Linux OS and go through the install. No thanks. Run a VM in AWS and install Ansible: This could potentially cost some small amount of money, and it also requires me to maintain a remote system, and make sure I have an internet connection whenever I want to try something out. Use the Azure Cloud Shell: I LUV the Azure Cloud Shell, like a lot. It comes with all kinds of goodies pre-installed, including Ansible. But again this assumes that I will have an internet connection, and it won’t have access to local resources in my work lab. Use a Vagrant Box: This one really appealed to me and I was about to start down this road. But I’ve also been trying to get more comfortable with containers, so if I’m going to use a Vagrantfile to customize a CentOS box, then I might as well use a container instead. Use Docker for Windows to run Ansible in a container: This was ultimately the winner for me. I already had DfW installed on my Win 10 laptop. I’m trying to get more into the world of containers. The container will have access to my work lab. And I can share the image with coworkers easily if they want to repeat the demo. And so DfW FTW as it were!
|
||
If you’d like to do the same, first you need to install DfW. You will also need to install the Hyper-V and Containers features in Windows, which DfW will do for you if you haven’t already. Once DfW is installed, go ahead and clone my GitHub repo for this project. You can build the docker image using:
|
||
docker build -t centos7-ansible-oneview .
|
||
And then spin up a container using:
|
||
docker run -it centos7-ansible-oneview /bin/bash
|
||
If you happen to be trying to use this for your own OneView and Synergy install, then definitely take a look at the additional requirements HPE has listed on their GitHub repo for Ansible and Terraform. As I work through the demo process, I will try to update this post and the Dockerfile with additional configuration to make the process as seamless as possible.
|
||
`,summary:"I’ve been going through Ansible training for work, and I want to be able to easily try out some of the playbooks and modules on my Windows 10 box. There’s lots of ways I could go about this, but for me the thing that makes the most sense is to leverage Docker for Windows (DfW) and the magic of containers to be able to get up and running quickly. In this post I will walk through what I did to get started, and why I landed on containers instead of another solution.",date:"27 Jun, 2018",url:"https://nedinthecloud.com/2018/06/27/running-ansible-on-windows-10/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/06/27/terraform-fotd-cidrnetmask/":{title:"Terraform - FotD - cidrnetmask()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the cidrnetmask() function. The example file is on GitHub here.
|
||
What is it? Function name: cidrnetmask(iprange)
|
||
Returns: Takes an IP address range in CIDR notation and returns the netmask value in X.X.X.X type notation.
|
||
Example:
|
||
variable "cidrnetmask" { default = "10.0.1.0/16" } # Returns 255.255.0.0 output "cidrnetmask" { value = "\${cidrnetmask(var.cidrnetmask)}" } Example file: ############################################## # Function: cidrnetmask ############################################## ############################################## # Variables ############################################## variable "iprange" { default = "10.0.0.0/16" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_iprange" { value = "\${var.iprange}" } output "2_cidrnetmask_output" { value = "\${cidrnetmask(var.iprange)}" } Run the following from the cidrnetmask folder to get example output for a number of different cases:
|
||
terraform apply -var "iprange=172.16.0.0/8" terraform apply -var "iprange=172.16.0.0/16" terraform apply -var "iprange=1.1.1.1/24" terraform apply -var "iprange=156.25.42.0/0" terraform apply -var "iprange=172.16.2.128/32" #Fails with invalid CIDR expression terraform apply -var "iprange=172.16.2.128/33" #Fails with missing mask terraform apply -var "iprange=172.16.0.0" Why use it? There are many systems out there that expect their IP address and subnet mask to be in dot notation instead of CIDR. In my experience most of the cloud-native stuff would prefer CIDR notation, and so there needs to be someway of converting one to the other. This is the way, and the truth, and the… well, it’s the way at least. Combine this with the cidrhost function and you’re off to the IP addressing races.
|
||
Lessons Learned You have to give the function a valid mask value (0-32). Anything else will throw an error. Omitting the mask value also throws an error. The function tests the IP address itself too, so if you give it something like 500.500.1.1/24 it will also throw an error.
|
||
Coming up next is the cidrsubnet() the last of what might be my favorite series of functions thus far.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the cidrnetmask() function. The example file is on GitHub here.
|
||
What is it? Function name: cidrnetmask(iprange)
|
||
Returns: Takes an IP address range in CIDR notation and returns the netmask value in X.X.X.X type notation.
|
||
Example:
|
||
variable "cidrnetmask" { default = "10.`,date:"27 Jun, 2018",url:"https://nedinthecloud.com/2018/06/27/terraform-fotd-cidrnetmask/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/06/26/terraform-fotd-cidrhost/":{title:"Terraform - FotD - cidrhost()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the cidrhost() function. The example file is on GitHub here.
|
||
What is it? Function name: cidrhost(iprange, hostnumber)
|
||
Returns: Takes an IP address range in CIDR notation and finds the specified hostnumber in the range. The hostnumber is an int value, and it accepts positive or negative values. The negative values cycle down from the top of the range. The return value will be an IP address.
|
||
Example:
|
||
variable "cidrhost" { default = "10.0.1.0/16" } # Returns 10.0.1.2 output "chomp" { value = "\${cidrhost(var.cidrhost,2)}" } Example file: ############################################## # Function: cidrhost ############################################## ############################################## # Variables ############################################## variable "iprange" { default = "10.0.0.0/16" } variable "hostnum" { default = 2 } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "1_iprange" { value = "\${var.iprange}" } output "2_hostnum" { value = "\${var.hostnum}" } output "3_cidrhost_output" { value = "\${cidrhost(var.iprange,var.hostnum)}" } Run the following from the cidrhost folder to get example output for a number of different cases:
|
||
terraform apply -var "iprange=172.16.0.0/24" -var "hostnum=2" terraform apply -var "iprange=172.16.0.0/19" -var "hostnum=12" terraform apply -var "iprange=192.168.0.1/25" -var "hostnum=17" terraform apply -var "iprange=156.25.42.0/19" -var "hostnum=100" terraform apply -var "iprange=172.16.2.128/27" -var "hostnum=-5" #Fails with missing mask terraform apply -var "iprange=172.16.0.0" -var "hostnum=2" #Fails with no available ip address terraform apply -var "iprange=172.16.0.0/32" -var "hostnum=2" Why use it? Dealing with IP Addresses is notoriously difficult. I think I only truly understood all the subnet math when I was studying for the CCNA, and then promptly found the IP Subnet Calculator and discarded most of what I knew to the dustbin of history. If you’ve ever tried to work with IP Addresses in CloudFormation or ARM Templates, you will immediately understand what a godsend this function is. Put together with a count loop and you can assign static IP Addresses like a champ.
|
||
Lessons Learned Even if you feed the function the wrong starting value for a CIDR range, it still figures it out for you. The broadcast address (255) is included in the count, so just remember that while the function will return it, you shouldn’t use it.
|
||
Coming up next is the cidrnetmask() which does not include any type of cider, hard or otherwise.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the cidrhost() function. The example file is on GitHub here.
|
||
What is it? Function name: cidrhost(iprange, hostnumber)
|
||
Returns: Takes an IP address range in CIDR notation and finds the specified hostnumber in the range. The hostnumber is an int value, and it accepts positive or negative values.`,date:"26 Jun, 2018",url:"https://nedinthecloud.com/2018/06/26/terraform-fotd-cidrhost/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/06/25/terraform-fotd-chunklist/":{title:"Terraform - FotD - chunklist()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the chunklist() function. The example file is on GitHub here.
|
||
What is it? Function name: chunklist(list, size)
|
||
Returns: Takes the list and splits it into several lists. Each list will have the number of elements specified in the size property. Any remainder will be in the final list. The function returns the chunked lists as a list value.
|
||
Example:
|
||
variable "chunklist" { default = ["A","list","of","elements"] } # Returns [["A","list"],["of","elements"]] output "chomp" { value = "\${chunklist(var.chunklist,2)}" } Example file: ############################################## # Function: chunklist ############################################## ############################################## # Variables ############################################## variable "chunklist" { default = ["one", "two", "three", "four"] } variable "chunklist_size" { default = 2 } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "list_output" { value = "\${var.chunklist}" } output "chunklist_output" { value = "\${chunklist(var.chunklist,var.chunklist_size)}" } Run the following from the chunklist folder to get example output for a number of different cases:
|
||
terraform apply -var "chunklist_size=1" terraform apply -var "chunklist_size=2" terraform apply -var "chunklist_size=3" terraform apply -var "chunklist_size=4" terraform apply -var "chunklist_size=5" #Empty list terraform apply -var "chunklist=[]" Why use it? There’s a lot of data sources that will return a list. Whether it’s a subnets in VPC, datastores in VMware, network security groups in Azure. Now I’m not sure exactly why I would need to chunk that list into smaller lists, and hope that the data source returned the correct number of items so things chunk evenly. I’m sure there is a use for it, since it exists. If you’re using it, let me know!
|
||
Coming up next is cidrhost(), the first of several functions dealing specifically with IP Addressing. For which I am eternally grateful.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the chunklist() function. The example file is on GitHub here.
|
||
What is it? Function name: chunklist(list, size)
|
||
Returns: Takes the list and splits it into several lists. Each list will have the number of elements specified in the size property.`,date:"25 Jun, 2018",url:"https://nedinthecloud.com/2018/06/25/terraform-fotd-chunklist/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/06/22/terraform-fotd-chomp/":{title:"Terraform - FotD - chomp()",tags:["fotd","hashicorp","terraform"],content:"This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the chomp() function. The example file is on GitHub here.\nWhat is it? Function name: chomp(string)\nReturns: Takes the string and removes any trailing newlines. Trailing in this case are any newline characters that are at the end of the string.\nExample:\nvariable "chomp" { default = "A string with newlines \\n\\n\\n\\n" } # Returns "A string with newlines " output "chomp" { value = "${chomp(var.chomp)}" } Example file: ############################################## # Function: chomp ############################################## ############################################## # Variables ############################################## variable "chomp" { default = "A string with newlines \\n\\n\\n\\n\\n" } ############################################## # Resources ############################################## resource "local_file" "chomp_file" { content = "${chomp(var.chomp)}" filename = "output.txt" } ############################################## # Outputs ############################################## output "chomp_output" { value = "${chomp(var.chomp)}" } Run the following from the chomp folder to get example output for a number of different cases:\n#Initialize for local_file resource terraform init #Create strings in PowerShell to test $something = "String with one newline`n" terraform apply -var "chomp=$something" -auto-approve $something = "String with two newlines`n`n" terraform apply -var "chomp=$something" -auto-approve $something = "String with two lines`n Second line" terraform apply -var "chomp=$something" -auto-approve $something = "String with two lines`n Second line with newline `n" terraform apply -var "chomp=$something" -auto-approve #Just newlines $something = "`n`n`n`n`n" terraform apply -var "chomp=$something" -auto-approve #Empty string terraform apply -var "chomp=" Why use it? As someone who has been bit before with trailing newlines in text files, the chomp function is great for cleaning up data pulled from potentially messy sources. If you know your text absolutely should not end in a newline, then chomp is your new best friend. Garbage in, garbage out as the old adage goes. While you’re at it, I would recommend cleaning up whitespace at the end of string, which is something chomp does not do.\nLessons learned I found a fun bug! Let’s say you create a file using the local_file resource and populate it with a test string from a variable. Then you run the same config with an empty string as the value for the file. Terraform tries to run a diff to see if the contents of the file are different, and then it crashes. I submitted an issue so we’ll see what happens.\nComing up next is chunklist() which I hear is a favorite of sloth.\n",summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the chomp() function. The example file is on GitHub here.
|
||
What is it? Function name: chomp(string)
|
||
Returns: Takes the string and removes any trailing newlines. Trailing in this case are any newline characters that are at the end of the string.`,date:"22 Jun, 2018",url:"https://nedinthecloud.com/2018/06/22/terraform-fotd-chomp/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/06/21/terraform-fotd-bcrypt/":{title:"Terraform - FotD - bcrypt()",tags:["fotd","hashicorp","terraform"],content:"This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the bcrypt() function. The example files are on GitHub here.\nWhat is it? Function name: bcrypt(string, count)\nReturns: The bcrypt function takes a string and performs the Blowfish based encryption algorithm on it with the specified number of passes, returning a hash. Defaults to 10 passes.\nExample:\nvariable "bcrypt" { default = "1234" } # Returns varies depending on run # I got back $2a$05$0uOnQG9a8YjBRCBM4blwfOejPZ9RAfFVK2dxdVAsh2ovkzt0ZkBCO output "bcrypt" { value = "${bcrypt(var.bcrypt,5)}" } Example file: ############################################## # Function: bcrypt ############################################## ############################################## # Variables ############################################## variable "bcrypt" { default = "So long, and thanks for all the fish!" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## #Cost defaults to 10 output "bcrypt_no_cost" { value = "${bcrypt(var.bcrypt)}" } output "bcrypt_5_cost" { value = "${bcrypt(var.bcrypt, 5)}" } output "bcrypt_12_cost" { value = "${bcrypt(var.bcrypt, 12)}" } Run the following from the bcrypt folder to get example output for a number of different cases:\n#Start with the default variable terraform apply #Try submitting a string terraform apply -var 'bcrypt="Oh freddled gruntbuggly, Thy micturations are to me, As plurdled gabbleblotchits on a lurgid bee."' #Empty string test terraform apply -var "bcrypt=" Why use it? Bcrypt is a pretty secure way to create a hash or a string. And you can amplify the hash based on the count parameter. The higher the count, the better the security, or at least that’s the overall idea. Again, this is the sort of thing that will probably be used by the External provider or maybe when using the remote_exec provisioner.\nLessons learned I have to be honest here, I didn’t really know anything about the bcrypt process before I started learning about this function. In that regard, it was an excellent learning experience. If you want to know more about the bcrypt hash, I definitely recommend reading through the Wikipedia article. The format of the output is especially interesting. Take this example:\n$2a$12$hUy7uT6.K8B2cPNCdErpCOpx0.fXdeg1CWgHWULEw0Z4Vx2AINh02\nThe $2a refers to the version of bcrypt being used. The $12 means that 2^12 (4096) passes were performed. The next 22 characters are a 128-bit salt for the hash (hUy7uT6.K8B2cPNCdErpCO) and the remainder is the 184-bit hash (px0.fXdeg1CWgHWULEw0Z4Vx2AINh02). Another thing I realized was that setting the count to something high, like 20, means that your computer will be busy for a while. I recommend exercising caution.\nComing up next is ceil() which I assume has little to do with our mammalian, aquatic friends.\n",summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the bcrypt() function. The example files are on GitHub here.
|
||
What is it? Function name: bcrypt(string, count)
|
||
Returns: The bcrypt function takes a string and performs the Blowfish based encryption algorithm on it with the specified number of passes, returning a hash.`,date:"21 Jun, 2018",url:"https://nedinthecloud.com/2018/06/21/terraform-fotd-bcrypt/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/06/21/terraform-fotd-ceil/":{title:"Terraform - FotD - ceil()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the ceil() function. The example file is on GitHub here.
|
||
What is it? Function name: ceil(float)
|
||
Returns: Rounds up to the closest integer, unless the float is already an integer value. Will attempt to perform implicit conversion of other data types to float prior to conversion.
|
||
Example:
|
||
variable "float" { default = 41.66666666666 } # Returns 42 output "float" { value = "\${ceil(var.float)}" } Example file: ############################################## # Function: ceil ############################################## ############################################## # Variables ############################################## variable "ceil" { default = false } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "ceil_output" { value = "\${ceil(var.ceil)}" } Run the following from the ceil folder to get example output for a number of different cases:
|
||
terraform apply -var 'ceil=5.9' terraform apply -var 'ceil=-5.9' terraform apply -var 'ceil=3.1' terraform apply -var 'ceil=-3.1' terraform apply -var 'ceil="-1"' terraform apply -var 'ceil="String"' terraform apply Why use it? Ceil or ceiling is a fairly common mathematical device. If you are going to be manipulating float values and need an integer, ceil might suit your needs.
|
||
Lessons learned I learned in the process that there is an implicit conversion of boolean true to float as 1 and false to float as 0. This conversion only happens when using a default value or passing the value in a tfvars file. True or false passed in the -var agrument will be treated as a string without conversion. Strings will be converted to float if it’s a number, but not if it’s regular text. Negative numbers are rounded up towards 0, so -5.9 is rounded up to -5. Fun stuff!
|
||
Coming up next is chomp() which is one of my favorite Mario characters.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the ceil() function. The example file is on GitHub here.
|
||
What is it? Function name: ceil(float)
|
||
Returns: Rounds up to the closest integer, unless the float is already an integer value. Will attempt to perform implicit conversion of other data types to float prior to conversion.`,date:"21 Jun, 2018",url:"https://nedinthecloud.com/2018/06/21/terraform-fotd-ceil/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/06/20/terraform-fotd-base64sha256-and-base64sha512/":{title:"Terraform - FotD - base64sha256() and base64sha512()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the base64sha256() function and the base64sha512() function. Double feature! The example files are on GitHub for base64sha256 and base64sha512.
|
||
What is it? Function name: base64sha256(string)
|
||
Returns: The base64sha256 function takes a string and creates a sha-256 hash from it in raw byte form. The raw byte form is then encoded using base64 and returned as a string value.
|
||
Example:
|
||
variable "base64sha256" { default = "1234" } # Returns A6xnQhbz4Vx2HuGl4lXwZ5U2I8iziLRFnhP5eNfIRvQ= output "base64sha256" { value = "\${base64sha256(var.base64sha256)}" } Example file: ############################################## # Function: base64sha256 ############################################## ############################################## # Variables ############################################## variable "base64sha256" { default = "So long, and thanks for all the fish!" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "base64sha256_output" { value = "\${base64sha256(var.base64sha256)}" } output "sha256_output" { value = "\${sha256(var.base64sha256)}" } output "decoded_output" { value = "\${base64decode(base64sha256(var.base64sha256))}" } Run the following from the base64sha256 folder to get example output for a number of different cases:
|
||
#Start with the default variable terraform apply #Try submitting a string terraform apply -var 'base64sha256="Oh freddled gruntbuggly, Thy micturations are to me, As plurdled gabbleblotchits on a lurgid bee."' #Empty string test - sha256(string) does NOT like this! terraform apply -var "base64sha256=" Function name: base64sha512(string)
|
||
Returns: The base64sha512 function takes a string and creates a sha-512 hash from it in raw byte form. The raw byte form is then encoded using base64 and returned as a string value.
|
||
Example:
|
||
variable "base64sha512" { default = "1234" } # Returns 1ARVn2Auq2/WAqx2gNrL+q3RNjAzXpUfCXrzkA6d4Xa22yhRLy4AC50E+6UTPoscbo31nbOoq51gvkuXzJ6B2w== output "base64sha512" { value = "\${base64sha512(var.base64sha512)}" } Example file: ############################################## # Function: base64sha512 ############################################## ############################################## # Variables ############################################## variable "base64sha512" { default = "So long, and thanks for all the fish!" } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "base64sha512_output" { value = "\${base64sha512(var.base64sha512)}" } output "sha512_output" { value = "\${sha512(var.base64sha512)}" } output "decoded_output" { value = "\${base64decode(base64sha512(var.base64sha512))}" } Run the following from the base64sha512 folder to get example output for a number of different cases:
|
||
#Start with the default variable terraform apply #Try submitting a string terraform apply -var 'base64sha512="Oh freddled gruntbuggly, Thy micturations are to me, As plurdled gabbleblotchits on a lurgid bee."' #Empty string test terraform apply -var "base64sha512=" Why use it? Sha-256 and sha-512 hashes are often used to validate passwords or ensure that files haven’t been tampered with. I couldn’t find a provider that would require either of these functions explicitly, so if you know of one please let me know and I’ll add it to the examples. I can envision using this with the HTTP or External provider to get information and validate it.
|
||
Lessons learned According to the documentation, the base64sha256 and base64sha512 functions are not equivalent to running base64encode(sha256(string)) or base64encode(sha512(string)) since both sha256() and sha512() return a hexadecimal string and not the raw bytes that are encoded in the these functions. I added an output that decodes the string to show what the bytes look like. It’s, well, it’s not pretty. I also learned that the base64sha256 function does not like being fed an empty string. The configuration will hang on that output till you cancel it.
|
||
Finally we are done with the base64 functions. Hooray! Coming up next is bcrypt(password, cost). I have no idea what bcrypt does, so exciting times are ahead.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the base64sha256() function and the base64sha512() function. Double feature! The example files are on GitHub for base64sha256 and base64sha512.
|
||
What is it? Function name: base64sha256(string)
|
||
Returns: The base64sha256 function takes a string and creates a sha-256 hash from it in raw byte form.`,date:"20 Jun, 2018",url:"https://nedinthecloud.com/2018/06/20/terraform-fotd-base64sha256-and-base64sha512/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/06/19/terraform-fotd-base64gzip/":{title:"Terraform - FotD - base64gzip()",tags:["fotd","hashicorp","terraform"],content:"This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the base64gzip() function. The example files are on GitHub here.\nWhat is it? Function name: base64gzip(string)\nReturns: The base64gzip function first takes a string and applies gzip compression, then encodes the compressed data using base64.\nExample:\nvariable "base64gzip" { default = "1234" } # Returns H4sIAAAAAAAA/zI0MjYBAAAA//8BAAD//6Pg45sEAAAA output "base64gzip " { value = "${base64gzip (var.base64gzip )}" } Example file: ############################################## # Function: base64gzip ############################################## provider "aws" { access_key = "${var.aws_access_key}" secret_key = "${var.aws_secret_key}" region = "us-east-1" } ############################################## # Variables ############################################## variable "base64gzip" { default = "1234" } variable "sourcefile" { default = "input.txt" } variable "aws_access_key" {} variable "aws_secret_key" {} ############################################## # Resources ############################################## data "local_file" "source" { filename = "${var.sourcefile}" } resource "local_file" "base64gzip" { content = "${base64gzip(data.local_file.source.content)}" filename = "output.txt" } resource "aws_s3_bucket" "bucket" { bucket_prefix = "base64gzip" acl = "public-read" } resource "aws_s3_bucket_object" "object" { bucket = "${aws_s3_bucket.bucket.id}" key = "output.txt" content = "${base64gzip(data.local_file.source.content)}" content_encoding = "identity" acl = "public-read" } ############################################## # Outputs ############################################## output "file" { value = "https://${aws_s3_bucket.bucket.bucket_domain_name}/${aws_s3_bucket_object.object.id}" } output "base64gzip " { value = "${base64gzip (var.base64gzip )}" } Run the following from the base64gzip folder to get example output for a number of different cases:\n#Since we are using the aws and local provider we need to initialize terraform init -var "aws_access_key=YOURACCESSKEY" -var "aws_secret_key=YOURSECRETKEY" #Run a plan to get prepped terraform plan -var "aws_access_key=YOURACCESSKEY" -var "aws_secret_key=YOURSECRETKEY" -out terraform.tfplan #Run apply to get the output terraform apply "terraform.tfplan" #Apply it to any file you'd like! terraform plan -var "sourcefile=\\Path\\to\\a\\source\\text\\file.txt" -var "aws_access_key=YOURACCESSKEY" -var "aws_secret_key=YOURSECRETKEY" -out terraform.tfplan terraform apply "terraform.tfplan" Why use it? If you have a string that needs to be base64 encoded for a web service or some other resource and it also need to be compressed, this is how you can do that. I would think primarily of sending GET requests where you are trying to embed some information in the header. If you are using the HTTP Provider to send information to a site, then you might use this function.\nLessons learned I was combing through resources in Terraform that might use base64 gzipped files and I found that the S3 bucket object type let’s you specify the encoding of an object being added. So I created this example that will upload the string as a file. The string is encoded as base64, but once in the S3 bucket, it is decoded and stored as compressed in gzip. You can download the file using the link in the output. Copy and paste that text into a gzip decompression engine, like this one to get back the original text. The example also writes out an identical file locally, so you can remove the AWS pieces if you don’t have an account.\nI thought I was done with the base64 functions, but it looks like a few new ones were added. Coming up next is base64sha256() and base64sha512(). They are basically the same function, so I figured I’d do two in a day.\n",summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the base64gzip() function. The example files are on GitHub here.
|
||
What is it? Function name: base64gzip(string)
|
||
Returns: The base64gzip function first takes a string and applies gzip compression, then encodes the compressed data using base64.
|
||
Example:
|
||
variable "base64gzip" { default = "1234" } # Returns H4sIAAAAAAAA/zI0MjYBAAAA//8BAAD//6Pg45sEAAAA output "base64gzip " { value = "\${base64gzip (var.`,date:"19 Jun, 2018",url:"https://nedinthecloud.com/2018/06/19/terraform-fotd-base64gzip/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2018/06/15/terraform-fotd-base64encode/":{title:"Terraform – FotD – base64encode()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the base64encode() function. The example files are on GitHub here.
|
||
What is it? Function name: base64encode(string)
|
||
Returns: The base64encode function returns a base64 encoded value of string.
|
||
Example:
|
||
variable "base64encode" { default = "1234" } # Returns MTIzNA== output "base64encode" { value = "\${base64encode(var.base64encode)}" } Example file: ############################################## # Function: base64encode ############################################## ############################################## # Variables ############################################## variable "base64encode" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "base64encode_output" { value = "\${base64encode(var.base64encode)}" } Run the following from the base64encode folder to get example output for a number of different cases:
|
||
#We have to get something to encode first $fileContent = Get-Content .\\textFile.txt terraform apply -var "base64encode=$fileContent" #Should output QUJDREVGR0hJSktMTU5PUFFSU1RVVldYWVowMTIzNDU2Nzg5 terraform apply -var "base64encode=ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789" #Empty string test terraform apply -var "base64encode=" Why use it? If you have a string that needs to be base64 encoded for a web service or some other resource, this is how you can transform it. I would think primarily of creating URL strings that are base64 encoded to assist with data transfer. If you are using the HTTP Provider to send information to a site, then you might use this function.
|
||
Lessons learned I have to be honest here, I didn’t really know anything about base64 encoding before I started learning about this function. In that regard, it was an excellent learning experience. If you want to know more about base64 encoding, I definitely recommend reading through the Wikipedia article.
|
||
Coming up next is base64gzip() which will round out the base64 functions.
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the base64encode() function. The example files are on GitHub here.
|
||
What is it? Function name: base64encode(string)
|
||
Returns: The base64encode function returns a base64 encoded value of string.
|
||
Example:
|
||
variable "base64encode" { default = "1234" } # Returns MTIzNA== output "base64encode" { value = "\${base64encode(var.`,date:"15 Jun, 2018",url:"https://nedinthecloud.com/2018/06/15/terraform-fotd-base64encode/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/06/14/terraform-fotd-base64decode/":{title:"Terraform - FotD - base64decode()",tags:["fotd","hashicorp","terraform"],content:"This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the base64decode() function. The example files are on GitHub here.\nWhat is it? Function name: base64decode(string)\nReturns: The base64decode function returns a decoded value of a base64 encoded string.\nExample:\nvariable "base64decode" { default = "MTIzNA==" } # Returns 1234 output "base64decode" { value = "${base64decode(var.base64decode)}" } Example file: ############################################## # Function: base64decode ############################################## ############################################## # Variables ############################################## variable "base64decode" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "base64decode_output" { value = "${base64decode(var.base64decode)}" } Run the following from the base64decode folder to get example output for a number of different cases:\n#We have to encode something first $fileContent = get-content .\\terraform_image.png $fileContentBytes = [System.Text.Encoding]::UTF8.GetBytes($fileContent) $fileContentEncoded = [System.Convert]::ToBase64String($fileContentBytes) terraform apply -var "base64decode=$fileContentEncoded" #Should output ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789 terraform apply -var "base64decode=QUJDREVGR0hJSktMTU5PUFFSU1RVVldYWVowMTIzNDU2Nzg5" #Empty string test terraform apply -var "base64decode=" Why use it? If you are being handed a string that is base64 encoded by a web service or some other data source, this is how you can transform it back into a usable string for the rest of the application. I would think primarily of parsing URL strings that are base64 encoded to assist with data transfer. If you are using the HTTP Provider to query information from a site, then you might use this function.\nLessons learned I have to be honest here, I didn’t really know anything about base64 encoding before I started learning about this function. In that regard, it was an excellent learning experience. If you want to know more about base64 encoding, I definitely recommend reading through the Wikipedia article.\nComing up next is base64encode().\n",summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the base64decode() function. The example files are on GitHub here.
|
||
What is it? Function name: base64decode(string)
|
||
Returns: The base64decode function returns a decoded value of a base64 encoded string.
|
||
Example:
|
||
variable "base64decode" { default = "MTIzNA==" } # Returns 1234 output "base64decode" { value = "\${base64decode(var.`,date:"14 Jun, 2018",url:"https://nedinthecloud.com/2018/06/14/terraform-fotd-base64decode/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/06/13/terraform-fotd-basename/":{title:"Terraform - FotD - basename()",tags:["fotd","hashicorp","terraform"],content:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the basename() function. The example file is on GitHub here.
|
||
What is it? Function name: basename(string) Returns: The basename function returns the last portion of a file path. Example:
|
||
variable "basic_path" { default = "/usr/bin" } # Returns bin output "basic_path" { value = "\${basename(var.basic_path)}" } Example file: ############################################## # Function: basename ############################################## ############################################## # Variables ############################################## variable "basename" {} ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "basename_output" { value = "\${basename(var.basename)}" } Run the following from the basename folder to get example output for a number of different cases:
|
||
terraform apply -var 'basename=C:\\\\' terraform apply -var 'basename=C:\\\\level0\\\\level1\\\\level2' terraform apply -var 'basename=C:\\\\level0\\\\level1\\\\level2\\\\' terraform apply -var 'basename=C:\\\\level0\\\\level1\\\\level2\\\\file1.txt' terraform apply -var 'basename=/' terraform apply -var 'basename=/level0/level1/level2' terraform apply -var 'basename=/level0/level1/level2/' terraform apply -var 'basename=/level0/level1/level2/file1.txt' terraform apply -var 'basename=' Why use it? If you really need to extract a file name from a long path, or just the directory name at the end of the path, this is really helpful.
|
||
Lessons learned From my testing I determined a few things. You need to double escape any backslash characters for Windows files since ‘\\’ is the escape character that Terraform uses for adding newlines and other special characters to a string. The function will return the last file name or folder name on either Linux or Windows file system paths. If the path is just the root of the file system then you’ll get back the root ‘/’. If you leave a trailing slash on the value, such as /level0/level1/ it will be ignored and level1 will be returned. I did try a 256 character directory with a 256 character filename to see if basename would choke, but it just shrugged like Whatevs. I also tried an empty string which surprisingly gave me a ‘.’ as output. Which I guess makes sense since the ‘.’ refers to the current directory.
|
||
Coming up next is base64decode().
|
||
`,summary:`This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the basename() function. The example file is on GitHub here.
|
||
What is it? Function name: basename(string) Returns: The basename function returns the last portion of a file path. Example:
|
||
variable "basic_path" { default = "/usr/bin" } # Returns bin output "basic_path" { value = "\${basename(var.`,date:"13 Jun, 2018",url:"https://nedinthecloud.com/2018/06/13/terraform-fotd-basename/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/06/12/terraform-fotd-abs/":{title:"Terraform - FotD - abs()",tags:["fotd","hashicorp","terraform"],content:`This post references Terraform version 0.11 and older. Check out the official docs for 0.12 and newer here
|
||
This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the abs() function. The example file is on GitHub here.
|
||
Function name: abs(float) Returns: The absolute value of the float value passed to the function. Will attempt to perform implicit conversion of other data types to float prior to conversion. Example:
|
||
variable "negative_float" { default = -3.14159 } # Returns 3.14159 output "plus_float" { value = "\${abs(var.plus_float)}" } Here’s a generalized example:
|
||
############################################## # Function: abs ############################################## ############################################## # Variables ############################################## variable "abs" { default = false } ############################################## # Resources ############################################## ############################################## # Outputs ############################################## output "abs_output" { value = "\${abs(var.abs)}" } Run the following from the abs folder to get example output for a number of different cases:
|
||
terraform apply -var 'abs=5' terraform apply -var 'abs=-5' terraform apply -var 'abs=3.14159' terraform apply -var 'abs=3.14149' terraform apply -var 'abs="-1"' terraform apply -var 'abs="String"' terraform apply I learned in the process that there is an implicit conversion of boolean true to float as 1 and false to float as 0. This conversion only happens when using a default value or passing the value in a tfvars file. True or false passed in the -var agrument will be treated as a string without conversion. Strings will be converted to float if it’s a number, but not if it’s regular text. Fun stuff!
|
||
Coming up next is basename().
|
||
`,summary:`This post references Terraform version 0.11 and older. Check out the official docs for 0.12 and newer here
|
||
This is part of an ongoing series of posts documenting the built-in interpolation functions in Terraform. For more information, check out the beginning post. In this post I am going to cover the abs() function. The example file is on GitHub here.
|
||
Function name: abs(float) Returns: The absolute value of the float value passed to the function.`,date:"12 Jun, 2018",url:"https://nedinthecloud.com/2018/06/12/terraform-fotd-abs/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2018/06/11/terraform-function-of-the-day-lets-go/":{title:"Terraform - Function of the Day - Let's Go!",tags:["fotd","hashicorp","terraform"],content:`As a frequent user of Terraform, I’ve found that the documentation is good, but I feel that it could do a better job of explaining how to use the various functions included in the interpolation engine. In this series of posts I am going to document a different Terraform function each day and how you might use it. This won’t require using a provider per se. I’m just going to use variables, the null resource, and outputs to work through these functions. During the course of these posts I am also going to include the examples on GitHub, so you can pull the code whenever you want. So without further ado, let’s get started!
|
||
Before we dive into functions here’s my basic setup. I am running Terraform 0.11.7 on a Windows 10 (1804) laptop. The GitHub repo I’ll be referencing is
|
||
here. The functions I am going to run through are on the Terraform docs site here. I’ll be going through in alphabetical order, so abs() will be the first entry. That’s really it. I don’t want to over-complicate this, so let’s go!
|
||
`,summary:"As a frequent user of Terraform, I’ve found that the documentation is good, but I feel that it could do a better job of explaining how to use the various functions included in the interpolation engine. In this series of posts I am going to document a different Terraform function each day and how you might use it. This won’t require using a provider per se. I’m just going to use variables, the null resource, and outputs to work through these functions.",date:"11 Jun, 2018",url:"https://nedinthecloud.com/2018/06/11/terraform-function-of-the-day-lets-go/",image:"tutorials.png",readingTime:"1"},"https://nedinthecloud.com/2018/05/30/i-hate-microsoft-except-i-dont/":{title:"I HATE MICROSOFT - except I don't",tags:["apple","linux","microsoft"],content:`Recently at two separate client meetings, I heard two people say that they hate Microsoft. As you may have guessed, both people were happily using some type of Apple device, and were complaining about having to use some form of Microsoft technology. In one case it was Office 365, and the other was Active Directory. And listen everyone, I get it. Sometimes I hate Microsoft too. Except not really.
|
||
What is going on with the level of ire being directed at Microsoft? I feel like there’s more to unpack here. In my experience it seems to come from one of three basic situations:
|
||
The person is a Linux admin who still uses the abbreviation M$ for Microsoft and considers anything not open source a complete waste of time The person is an Apple user who believes in the one true way of Steve Jobs, and relishes the fact that their Mac is inherently more secure, faster, and 3x as expensive The person is a Windows user who has been repeatedly burned by the bad choices Microsoft has made with some of their products, and honestly just can’t forgive the whole clippy thing All of these people have valid points about Microsoft - and I’d like to address some of the issues - but first a little background on me, so you know where I am coming from. I grew up using an Apple IIgs, then a Macintosh, then a PowerPC Mac. I was all in for the MacOS through OS9, which you might remember is the reason that Apple had to completely rewrite their operating system for OSX. I even went to a university that only used Macs (Drexel). I was firmly entrenched in the Apple world and I loved it. Then I transferred to a technical college, and was thrown unceremoniously into the deep end of PCs. Suddenly, I needed a laptop running Windows, and I was learning about the 8088 and 8086 processor families. I had to learn all about ISA, and interrupts, and low-level formatting for hard drives. For years I had been completely sheltered from the world of PCs, and now I needed to learn about them if I wanted a snowball’s chance in hell of getting an entry-level tech job.
|
||
So I learned. And I was somewhat miserable for a while. I missed the slick graphics and perceived simplicity of the Mac. I still thought of myself as an Apple person who had to use Microsoft for work. But they were the evil empire, and Apple was the scrappy underdog. Apple had cachet, cool factor, a je ne sais quoi that Microsoft lacked. Sure Microsoft had to buy Apple stock to prop them up through the lean years, but whatevs, they were still the edgy outsider, and Microsoft was the corporate tool. Which is to say I still thought in those terms 20 years ago.
|
||
Once you spend a little time in the industry, you come to realize it is all corporations vying for your attention with shiny gadgets. Apple is not a scrappy underdog, they are a multi-billion dollar company. Google isn’t your hip cousin who listens to Fugazi, they are the corporate record machine that presses out Fugazi CDs for mass consumption. And to all the open source elitists, you know who the biggest contributors to open source projects are? Google, Facebook, and Microsoft to name a few. Open source would not be nearly as vibrant or well written if it weren’t for the very real and expensive efforts of large corporations to subsidize it. Don’t get me wrong, they benefit from it too. This isn’t some charity project. When Microsoft contributes to the Kubernetes project, they are doing so because they want a stable product to be available to run on Azure.
|
||
That’s not to say that Microsoft is all sunshine and rainbows. I think that SCCM is a convoluted mess of a product that needs a serious rewrite from scratch using modern programming methodologies. I think that SCOM should be taken out behind the barn and shot. SCVMM is such a difficult product to use, I would rather build an OpenStack cluster from scratch, from the source code. Don’t even get me started on Windows Phone, or Windows Mobile, or Windows CE, or all the other failed attempts at making a viable mobile OS. Also, Group Policy is just the worst. I could write a novel about how painful Group Policy is. Microsoft has some serious failings, but they also have some seriously good tech. I’d rather debate the merits of a specific technology than the perceived reputation of the company that created it. Unless of course they are killing black rhinos for fun, because guys? That’s just not OK.
|
||
Coming back to my original trio of people, let’s address some of this hate:
|
||
Person number one who is a Linux admin and hates good ole M$ because they are capitalist pigs.
|
||
Bad news buddy. The open source tech you are using probably has tons of code commits from Microsofties who want to improve the open source community for their own nefarious ends. Also, you work at a business to make money. Unless you are some kind of non-capitalistic, non-profit, magical unicorn you are probably throwing stones at glass houses. Linux is awesome. Open source software is awesome. No one is even debating that at this point, least of all Microsoft. sudo get a life.
|
||
Person number two who is an Apple user and better than everyone else. It’s not elite if everyone has one. It’s not VIP if everyone is a VIP. You’ve not the tragically-hip, counter-culture, free-thinking, box-breaker you thought you were. You can have any Mac you like as long as it’s black or white. I get that Microsoft isn’t especially cool or hip. And you know what? Neither are you. Apple makes some really excellent consumer tech. We’re talking about business products here, so why don’t we put on our big-boy pants and use products that do the best job, not that have the best semi-annual infomercial.
|
||
Person number three who is a once-bitten, twice-shy Microsoft user. You and I have a lot in common. Microsoft has done me wrong. But remember, it’s not really one monolithic company. It’s a tangled mess of warring product line fiefdoms, some of which are well run and some of which are not. Some of the divisions are crushed under the weight of legacy code and others are so bloated with features that they regularly end up beached on the Seattle shoreline. You’ll find that newer product lines, having shed the sorrows imposed by the elder gods, are nimble and and lithe. Focus on the future where Office 365 replaces managing Exchange, Azure AD slowly phases out Active Directory, and Azure rolls out a new feature on the daily. And by all means, enjoy it on a Mac.
|
||
When all is said and done, Microsoft is a flawed company, just like all companies. To say you hate Microsoft is a pointless statement, edging up against inanity. You might as well join a Nu-metal band and rage on about the MAN and your corporate shackles. Or you can grow up, realize that we’re all mostly trying to make the best of the mess we were handed, and stop saying ridiculous things like, “I HATE MICRO$OFT.” You’re not impressing anyone.
|
||
`,summary:`Recently at two separate client meetings, I heard two people say that they hate Microsoft. As you may have guessed, both people were happily using some type of Apple device, and were complaining about having to use some form of Microsoft technology. In one case it was Office 365, and the other was Active Directory. And listen everyone, I get it. Sometimes I hate Microsoft too. Except not really.
|
||
What is going on with the level of ire being directed at Microsoft?`,date:"30 May, 2018",url:"https://nedinthecloud.com/2018/05/30/i-hate-microsoft-except-i-dont/",image:"publiccloud.png",readingTime:"6"},"https://nedinthecloud.com/2018/05/17/using-azure-functions-and-buffer-api-for-automation/":{title:"Using Azure Functions and Buffer API for Automation",tags:["automation","azure-functions","buffer"],content:"In a previous post I mentioned how I am using Buffer, Feedly, and Zapier to automate parts of my online persona. In this post I will talk about how I found a workflow that wasn’t support by Zapier, and how I used the API from Buffer and Azure Functions to automate post generation.\nI run two podcasts for work, Buffer Overflow and AnexiPod. Both of them have been running for over a year, and I thought it would be cool to repost the previous year’s episode on Thursday for Throwback Thursdays (TBT). I figured that I should be able to set up a workflow in Zapier that would check daily to see if there was a post from a year ago, and then repost that for the upcoming Thursday slot. Sadly, that action did not exist in Zapier. How could I fix this? Well I know how to interact with websites and APIs via PowerShell due to my work in Azure and Azure Stack. So I figured I could write a script that runs as a scheduled task. The script would have to do the following:\nPoll the source feed to see if there was a post a year ago that day If no, do nothing If yes, proceed Create a new Buffer item for the upcoming Thursday including the previous post and the #TBT hashtag The next decision I had to make was where to run the script. I could have it running on my desktop at home, but then I would need to make sure that my desktop was always on and transfer the task when I replace or rebuild it. Instead I thought of Azure Functions, which is something I’ve been meaning to dig into anyway. I know that you can write functions in PowerShell, so that meant I could develop the script locally and then adapt it for Azure Functions.\nStep one was to poll the source feed. Since this is a podcast, there is already an XML feed available. So I chose to write a function that takes a URL and a date. The function pulls the feed in as an XML object, and then uses the date to filter the posts for a specific publish date. These fields are pretty standardized for podcasts, so this should be pretty portable. The function returns one or more post objects, or null if no posts for that date are found.\n###Function Get blog post from a year ago for AnexiPod ###Takes returns blog post URLs if any found. Returns null otherwise function Get-BlogPosts { Param( $date = (Get-Date).AddYears(-1), $url = "https://www.anexinet.com/blog/category/anexipod/feed/" ) $response = [xml](Invoke-WebRequest -Uri $url -UseBasicParsing) $items = $response.rss.channel.item $returnItems = $items | Where-Object{([DateTime]$_.pubDate).Date -eq $date.Date} if($returnItems){ return $returnItems }else{ return $null } } The next step was to add interaction with the Buffer API. In order to do that, I had to create an application on the Buffer side. The application information is not particularly important for my purposes. Here is the application information I used. The key is that this is a native application, and not a web application, meaning that there is no sign-in page that Buffer has to redirect the authentication request back to. Once I created the application, I had to figure out how to interact with the API. The authentication piece uses an Access Token that is generated when the application is created. All requests made to the API should include the Access Token in the header or in the URI string. That access token is linked to your account on Buffer, so if you are only using this application with your personal Buffer account then you are good to go. If you want to access other accounts in Buffer, then you will need to generate an access token for each through an OAuth process. I worked all that out after I figured out how to generate the access token for other accounts, so maybe that little tidbit will save you some trouble.\nArmed with the access token, I first had to get the profiles associated with my Buffer account. Each profile corresponds to a social media account of some kind, for instance I have three profiles for Twitter, LinkedIn, and Facebook respectively. Here is the bit of script that retrieves the profile information.\n#Build request header $requestheader = @{Authorization="Bearer $access_token"} #Get profiles $profiles = Invoke-RestMethod -Method Get -Uri "$url/profiles.json" -Headers $requestheader -UseBasicParsing Then I had to create the buffer item for a post that was found. When I tried to do that the first time I got an error, [Disallowed Key Characters]. Unfortunately, I had no idea what that meant and the Google machine wasn’t very helpful. I turned to Twitter and hit up Buffer. They got back to me, and after a little back and forth I discovered that I was sending the POST data in JSON form, which they don’t support. They just want an ‘&’ separated string of key value pairs. Once we got all that worked out, I was able to get the POST portion of my script working. Big props to the Buffer team for being helpful and responsive.\n#Create post $body = "" $text = "#TBT Check out last year's #AnexiPod episode $($postInfo.title) $($postInfo.link)" foreach($profile in $profiles){ $body += "profile_ids[]=$($profile._id)&" } $body += "text=$text" $body += "&scheduled_at=$($thursday.GetDateTimeFormats("s"))" Write-Output "Body is $body"; Write-output "URL is $url"; Invoke-RestMethod -Method Post -Uri "$url/updates/create.json?access_token=$access_token" -Body $body -ContentType $content_type -UseBasicParsing -Verbose Finally I adapted the script for Azure Functions. The main thing I had to figure out there was how to store my access token in Azure Key Vault, so it wasn’t sitting in plain text in my Function code. Thank goodness someone else has already figured that out, and I just followed the steps here. At the end of the day I now have an Azure Function that runs on a daily trigger and does exactly what I want. This was a great exercise because now I know more about interacting with APIs, creating Azure Functions, and I have a script I can adapt for other automation tasks. Hooray!\nHere’s the full Azure Function for those that are interested:\nWrite-Output "PowerShell Timer trigger function executed at:$(get-date)"; ###Function Get blog post from a year ago for AnexiPod ###Takes returns blog post URLs if any found. Returns null otherwise function Get-BlogPosts { Param( $date = (Get-Date).AddYears(-1), $url = "https://www.anexinet.com/blog/category/anexipod/feed/" ) $response = [xml](Invoke-WebRequest -Uri $url -UseBasicParsing) $items = $response.rss.channel.item $returnItems = $items | Where-Object{([DateTime]$_.pubDate).Date -eq $date.Date} if($returnItems){ return $returnItems }else{ return $null } } function Add-BufferPost { Param( $url = "https://api.bufferapp.com/1", $postInfo, $access_token, $content_type = "application/x-www-form-urlencoded" ) [Net.ServicePointManager]::SecurityProtocol = [Net.SecurityProtocolType]::Tls12 #Build request header $requestheader = @{Authorization="Bearer $access_token"} #Get profiles $profiles = Invoke-RestMethod -Method Get -Uri "$url/profiles.json" -Headers $requestheader -UseBasicParsing #Find the next Thursday to post $thursday = get-date while($thursday.DayOfWeek -ne "Thursday"){ $thursday = $thursday.AddDays(1) } $thursday = $thursday.AddHours(4) #Create post $body = "" $text = "#TBT Check out last year's #AnexiPod episode $($postInfo.title) $($postInfo.link)" foreach($profile in $profiles){ $body += "profile_ids[]=$($profile._id)&" } $body += "text=$text" $body += "&scheduled_at=$($thursday.GetDateTimeFormats("s"))" Write-Output "Body is $body"; Write-output "URL is $url"; Invoke-RestMethod -Method Post -Uri "$url/updates/create.json?access_token=$access_token" -Body $body -ContentType $content_type -UseBasicParsing -Verbose } #Get endpoint and password for MSI $endpoint = $env:MSI_ENDPOINT; $secret = $env:MSI_SECRET; #Vault URI to get AuthN token $vaultTokenURI = 'https://vault.azure.net&api-version=2017-09-01'; #Key vault secret to retrieve $vaultSecretURI = 'https://functions.vault.azure.net/secrets/buffer-access-key/9aa4e760ae4240f1a3976b2ee379cff6/?api-version=2015-06-01'; $header = @{'Secret' = $secret}; # Get Key Vault AuthN Token $authenticationResult = Invoke-RestMethod -Method Get -Headers $header -Uri ($endpoint +'?resource=' +$vaultTokenURI); # Use Key Vault AuthN Token to create Request Header $requestHeader = @{ Authorization = "Bearer $($authenticationResult.access_token)" } # Call the Vault and Retrieve Creds $buffer_secret = Invoke-RestMethod -Method GET -Uri $vaultSecretURI -ContentType 'application/json' -Headers $requestHeader $access_token = $buffer_secret.value; Write-Output "Retrieved access token $access_token"; $date = (Get-Date).AddYears(-1); $posts = Get-BlogPosts -date $date; Write-Output "Retrieved blog posts"; Write-Output $posts; if($posts){ Write-Output "Found $($posts.count) for $date"; foreach($post in $posts){ Write-Output "Writing post"; Write-Output $post.title; Add-BufferPost -postInfo $post -access_token $access_token; } } else{ Write-Output "No posts found for $date"; } ",summary:`In a previous post I mentioned how I am using Buffer, Feedly, and Zapier to automate parts of my online persona. In this post I will talk about how I found a workflow that wasn’t support by Zapier, and how I used the API from Buffer and Azure Functions to automate post generation.
|
||
I run two podcasts for work, Buffer Overflow and AnexiPod. Both of them have been running for over a year, and I thought it would be cool to repost the previous year’s episode on Thursday for Throwback Thursdays (TBT).`,date:"17 May, 2018",url:"https://nedinthecloud.com/2018/05/17/using-azure-functions-and-buffer-api-for-automation/",image:"tutorials.png",readingTime:"7"},"https://nedinthecloud.com/2018/05/14/automating-my-life-with-feedly-zapier-and-buffer/":{title:"Automating my Life with Feedly, Zapier, and Buffer",tags:["automation","buffer","feedly","zapier"],content:`This is about building my personal brand. If you don’t care about any of that, then you can safely skip this post. If you’re looking for ways to automate your brand building, then this will probably resonate with you. There was a period of time when I thought that marketing myself and having a personal brand was kind of gross. Now I realize that personal branding is pretty important if you want to build a career in the public side of IT. I am not making the argument that the only way to progress in your IT career is by being active on social media, community meetups, and actively blogging. But for a certain type of career path, the one that I seem to be on, those things definitely help.
|
||
At a client meeting a few weeks ago, one of the people I was meeting with commented on how they were enjoying my posts on LinkedIn. He asked how I had time to post on LinkedIn throughout the week, even on weekends, with all the other things I was doing. The truth is, I don’t have time to do that. But I understand that building up a following is about posting consistently. I tend to consume my news in large chunks at a time. Sitting on the couch at night I might spend an hour reading, and during that period I may want to share five or six articles of interest. But I don’t want to share them all in one big block. Those posts need to trickle out on a regular basis, to increase the likelihood that one or more of them will be seen. I also want to be able to share my content on more than one platform, and spread out posts about the same topic, so that my posts are more likely to be seen.
|
||
To that end I signed up for Buffer. Buffer allows you to schedule posts on various social media platforms well in advance. It also uses a link shortener that allows the platform to track the click through rate on your posts, so you know what people are interested in and what didn’t resonate. In theory I could analyze trends and try to make a determination about what types of posts are likely to generate interest. There’s probably a case here to be made for using AI and ML to devise a posting schedule and strategy that guarantees an improvement in click-through rates, but I’m not trying to make money on ads or anything. Click through does nothing if I’m alienating people with salacious posts or improperly polarizing statements. Sure that stuff probably generates interest in the short term, but it will backfire in the long term.
|
||
Wow, that was a bit of a digression now wasn’t it? Where was I? Oh yes, Buffer. So I signed up for the free version of Buffer initially, but eventually I started to bump up against the limits of the free account. There were certain posts I might want to repeat on a regular basis, called Re-buffer. There was also a limit of 10 posts in my queue, and if I wanted to schedule stuff out in the future, then I couldn’t do that. For instance, I might want to buffer up a bunch of posts about my Pluralsight courses. I couldn’t really dive into the analytics side of things, and there was a limit on how far ahead I could schedule posts. All in all, I need to go Pro. The cost was reasonable, and the features were enough to push me from Free to Paid.
|
||
Do you remember Google Reader? If you’re the type of person who does, then you may have had to take a moment and collect yourself. Google Reader was the best product that Google ever let die. It was a solid RSS reader and aggregator of content. No frills, no unnecessary graphics or weird UI quirks. It stripped out the crap from sites and let you just read the damn news. Obviously it was doomed from the start, but it was how I chose to consume the news. When the door shut forever on the best RSS reader - not bitter - I turned to Feedly to fill the hole left in my heart. It wasn’t a perfect fit, but it got the job done. And that continued to be the case for several years. I used the free version of Feedly, and it did what I wanted. There were a few things that pushed me over the edge to get the paid version. First, I started to hit the upper limit on how many things I can subscribe to - the free version is 100 sources. I have started subscribe to a lot of smaller blogs that may not post a lot, but tend to be high quality when they do. Scott Lowe and Ivan Pepelnjak are great examples. I also wanted to automate the process of sharing posts that I find interesting. Why? Let me explain, no no there is too much, let me sum up.
|
||
I have a podcast called Buffer Overflow. We have a lightning round at the end of each episode. I flag articles in Feedly as part of the lightning round Feed, and then reference that when preparing the episode. I want anything that goes into that Feed to automatically be buffered in my Twitter and LinkedIn feeds. This is pretty easy to do with Zapier or IFTTT, but Feedly only supports that integration if you are running the Pro version. That was the thing that really pushed me over to the Pro version.
|
||
So that’s how I am automating my brand building. There are some other integrations I want to pursue, like automating a repost of old articles that got more than a certain number of views or clicks, or if I add a new post on my Wordpress site, automatically creating 4 posts to be published over the next four weeks. There’s probably some other things that I haven’t even thought of yet, but I feel confident that between Feedly, Buffer, and Zapier I should be able to automate whatever I need to. Which gives me more time to be productive (read: lazy) and removes some of the drudgery of building a personal brand.
|
||
`,summary:"This is about building my personal brand. If you don’t care about any of that, then you can safely skip this post. If you’re looking for ways to automate your brand building, then this will probably resonate with you. There was a period of time when I thought that marketing myself and having a personal brand was kind of gross. Now I realize that personal branding is pretty important if you want to build a career in the public side of IT.",date:"14 May, 2018",url:"https://nedinthecloud.com/2018/05/14/automating-my-life-with-feedly-zapier-and-buffer/",image:"tutorials.png",readingTime:"5"},"https://nedinthecloud.com/2018/05/06/veritas-360-cloudpoint/":{title:"Veritas 360 - CloudPoint",tags:[],content:`At Cloud Field Day 3 we visited the Veritas office and they presented their cloud vision to us. It took a little while to ramp up, as I detailed in my post here. My fellow CFD delegate Martez Reed has already put an excellent post together detailing the high-level view of what Veritas had on display, so I won’t rehash that now. Instead, I would like to focus down on their CloudPoint offering and actually try and take it for a spin. Let’s see how far I get.
|
||
I went to the main Veritas site and I saw the Multi-Cloud offering on the front page. That’s a good start, so I clicked the link to “Discover” it. Now I feel like an explorer riding in my dirigible searching for the mythical cloud city of Tepulteca.
|
||
Alright, with my aviator goggles tightly fastened, I navigated the next WordPress page. Aha, I have struck gold! There is a section for CloudPoint. Two clicks and I am already on the product page. Sadly, I consider that a success in the world of modern web design.
|
||
And now a button advertising that I can “Get it free.” Pardon me for be slightly sceptical, but there’s still no such thing as a free lunch. Nevertheless, into the maw of the beast I must go. For science!
|
||
Ack! The dread form of incessant calling, emails, and other unwanted communication.
|
||
Oooh, and a T&C link. Pardon me while I read 64-pages of legal jargon…..
|
||
…reading….
|
||
Well that’s interesting: “You do not intend to use or access, nor will allow any other person to use or access, the software for any purpose prohibited by United States law, including without limitation, for the development, design, manufacture or production of nuclear, chemical or biological weapons of mass destruction.”
|
||
Not sure how I would use Cloud Backup Software to accomplish that, but important nonetheless I suppose. Can we also put a provision in there about not eating unicorns when using the software? It’s an objectionable practice and Veritas needs to take a stand!
|
||
…reading…
|
||
Then there’s this: “Licensee shall keep accurate business records relating to its use of the Licensed Software for a period of three (3) years following termination of this Agreement. Upon request from Veritas, Licensee shall provide Veritas with a report certifying the destruction of Licensed Software pursuant to Section 3. The provisions regarding license restrictions, confidentiality, audit, exclusion of warranty, and the general provisions in Section 8 will survive expiration of the evaluation Term or termination of this Agreement.”
|
||
I need to keep business records for three years relating to trial software? And I need to be able to provide a report detailing the software’s destruction? Wow, um maybe I don’t want to try it that badly. No, I must. For science!
|
||
…reading…
|
||
Ah, now there’s this gem: “Licensee shall not disclose the results of any benchmark tests run on the Licensed Software without Veritas’ prior written consent”
|
||
It was snuck into a section about Confidentiality. I can try this thing, but I can’t really tell people how well it works unless I clear it with Veritas first. And I’m totally sure they will be okay with numbers that aren’t flattering right. Right?
|
||
…reading…
|
||
Lastly, this thing is tied to their general license agreements page which includes two more four-page documents that continue in this vein. I will spare the reader any more “excitement.” Suffice to say that it is largely more of the same.
|
||
Okay, so here we go. I am filling out the form and getting the software.
|
||
Whew, that was easy. And gave me a download that is taking rather a long time. In the meantime, I have a couple comments. One, I immediately got this email, which is perfect:
|
||
Except I don’t see a link to CloudPoint documentation or a getting started video. That should definitely be in there. Secondly, the download screen is this.
|
||
That screen should have the same content as the email, and a link to a Getting Started document. Make this easy for me Veritas!
|
||
Well, I walked away to go get lunch, since the download was taking a bit of time. In total you are looking at a 1.6GB file. Don’t go downloading it on a hotspot at the airport. The file is Veritas_CloudPoint_2.0.1_IE.img.gz. And since I didn’t get directed to any kind of getting started docs, I have no idea what to do with this monster. Crack it open with 7-zip? Sure, why not.
|
||
Well, that’s not helpful now is it? Based on the contents of the json files, I gather that this is a container-based application. To the docs then! In case you are trying to find the docs, I’ll help. There were NO links to the docs on the main product site or on the main community page. But if you go into the Pinned Post “Welcome to the CloudPoint Community” and scroll down to Helpful Resources, then you will find the link. Or use this one, except don’t for reasons that becomes apparent a few paragraphs down. *[Foreshadowing]*
|
||
By the way, Amit was one of the presenters at CFD. He’s a good dude, highly recommend chatting with him if you get the chance.
|
||
I’ve got the docs and I’m reading through. The CloudPoint install is in fact a Docker image based on Ubuntu LTS 16.04. You will need a suitable host to run it. That host can be in the cloud or on-premises. They recommend a physical host, but I don’t see any reason you couldn’t use a virtualized host with sufficient resources. I’m going to choose to use Microsoft Azure for my host, since I have credits there.
|
||
Wait a second…
|
||
I could have saved myself a LOT of trouble and just used the Azure Marketplace image. The question is do I deploy my own Ubuntu host, or use the marketplace image. The answer of course is use the marketplace image, since I only have so much time here. I created the marketplace image in a Resource Group I am called CloudPointMarket. I ended using the D2S_v3 instance size and added a 64GB Premium disk for data. That’s the recommended sizing from the docs. The deployment was successful, so I am going to try and connect to the box. Based on the NSG rules created by the template, it looks like I can connect via SSH and HTTPS.
|
||
Or not.
|
||
SSH it is then.
|
||
So, erm. Now whut? Haha! Jokes on me. I’ve been looking at the CloudPoint 1.0 manual. And I should be looking at the CloudPoint 2.0 manual, which is here. Fortunately, the Microsoft Azure Marketplace item had a link to the CloudPoint support site. Which again, SHOULD BE IN THE INTRO EMAIL AND DOWNLOAD PAGE.
|
||
Since finding that document took me so long, I decided to check back on the website. Hey, turns out it was just taking a while for the container to initialize.
|
||
Filled out the form, and I’m all a-twitter with excitement.
|
||
Heh, sweet. Let’s back up some cloud.
|
||
I want to try a few things:
|
||
Backup an Azure VM Backup an EC2 Instance Restore an Azure VM Restore an EC2 Instance Let’s start with Azure VM. I click on Manage cloud and arrays and it brings me to a menu.
|
||
I’m guessing I need to click on the Microsoft Azure cloud to add one. Voila!
|
||
Click on Add configuration and we’re off to the races.
|
||
Of course, it doesn’t really guide me through the process beyond asking for this information. There should at least be a (?) button to help me out here. If I click through the documentation that launches from the main page.
|
||
I found in the unhelpfully titled Configuring off-host plug-in, and below that the Microsoft Azure plug-in configuration notes with this handy table.
|
||
The link will take you to directions for creating a service principal using the portal, but I highly recommend using that Azure CLI in the Azure CloudShell instead. You can find directions here for that. I have a VM named BackupTest in a resource group called BackupTest, so I want to create a service principal with the necessary rights to backup VMs in BackupTest.
|
||
Here’s the CLI commands, be sure to swap the AppID out for the actual AppID created by the first command. And make note of the password and tenant ID as well. You’re going to need those.
|
||
az ad sp create-for-rbac --name CloudPointSPN az role assignment delete --assignee "AppID" --role Contributor az role assignment create --assignee "AppID" --role Reader az role assignment create --assignee "AppID" --role Contributor --resource-group BackupTest The documentation doesn’t actually specify what role is required on what resources, so I am going to try giving the SPN Contributor to the resource group and Reader to the subscription. We’ll see if that works okay.
|
||
There we go. Now that I’ve added a data source I get a different dashboard page.
|
||
Let’s protect some assets then. Here’s my VM.
|
||
And if I click on the pane, I get this sub-pane.
|
||
Clicking on Create Snapshot, I get this dialog box.
|
||
Looking in the Job Log, I can see that the snapshot appears to be successful.
|
||
That was a manual snapshot. If I want to automatically protect the VM, then I need to create a protection policy from the Dashboard view.
|
||
The policy tab is a little weird.
|
||
I can choose to retain a certain number of copies, days, weeks, months, or years. This is not at all clear. I am going to say 1 Month of retention with Daily backups at 12:00AM. I am guessing that means I will have a month of daily backups. The interface is awkward, but I’ve accepted that is just par for the course with backup software.
|
||
Now I can go back to the Polices section of the protected VM and assign it a policy for backup.
|
||
Wait, whut? Why?
|
||
Ah, see I created a Disk policy not a Host policy. I am trying to protect a Host, so I cannot assign a Disk policy. Maybe filter the policies to only objects that it can apply to?
|
||
I created an AzureGoldHost policy and all appears well.
|
||
The AWS side of things doesn’t use an IAM role or anything fancy like that. Instead it just wants the Access and Secret keys for your account. I created a user account with the AmazonEC2FullAccess and AmazonRDSFullAccess managed policies. In the future I would prefer that CloudPoint provide a predefined role you can copy and paste in.
|
||
I added the us-east-1 region to the configuration and created a new protection policy called AWSGoldHost. I created a new instance in my AWS account, and after 10 minutes it still hadn’t shown up as a protectable asset. There’s no refresh button, so I guess I have to wait it out. It would be nice to be able to automatically assign protection based on a metadata tag, and that goes for Azure, AWS, and any other target that supports tagging. Going in and manually assigning a protection policy doesn’t scale.
|
||
And after another few minutes the instance appeared, and I was able to add a protection policy and initiate a snapshot. Looks like all is well with the AWS side too.
|
||
At this point I decided to let things run for a little bit and verify that the policies I had created were actually working. I came back a week later and I could see that the AWS instance was grabbing a snapshot every day, just like I asked it to.
|
||
If I select a specific snapshot I get a side pane item that gives me a few action items.
|
||
In this case I am going to try and restore the instance to a “New Location”, which in the world of AWS and CloudPoint means that I am restoring it to a different destination subnet.
|
||
After kicking off the job, I can see in the AWS console that a new instance has been created and is in the process of initializing.
|
||
CloudPoint confirms that the job ran successfully.
|
||
Going back to the AWS console, I checked to see if CloudPoint had added any type of metadata tagging to the restored instance. Sadly, it had not.
|
||
I would really like the option to add some default tags to instances that are recovered. I would also like CloudPoint to add some tags to protected instances, so it’s easy to check from the AWS side if something has been protected and what policy has been applied.
|
||
Logging into the instance, I check to make sure that some files I had dropped on the original instance were there.
|
||
And they were. So I would call this a success. I also tested the restore process sticking with the original location.
|
||
The restore will create the instance in the same subnet, but it will not overwrite the existing instance.
|
||
Next I attempted to restore an Azure VM. I select a snapshot and clicked restore, just like the previous attempt.
|
||
I chose to restore it to the original location, which in the world of Azure means the same Resource Group.
|
||
The job was, um, not successful.
|
||
Sadly, this is ALL the information the CloudPoint interface gives you. Examining the job log from the GUI gives no additional clue as to what happened. My first thing to check was maybe I was too restrictive on the permissions I gave to the CloudPointSPN account. I granted it Contributor access to the entire Azure subscription.
|
||
The job failed again. I dug into the documentation a little bit and found this page detailing the log files and their contents, as well as location.
|
||
Based on this document, I should be looking for the flexsnap-agent-offhost.log file. So I SSH’d into the CloudPoint box and tried to find that file. It did not exist. But I was able to find the log files in /cloudpoint/logs. The flexsnap-agent.log turned out to have the information I needed. The last error for restore was the following:
|
||
OperationFailed: Failed to restore snapshot: Source snapshot does not exist I reviewed the accompanying json further up in the log, and I was able to verify the snapshot name and resource ID.
|
||
Azure: Restore instance snapshot azure-snapvm-4d8e572a321440e9a26f8f71ecd24e0d:backuptest:azuregoldhost20180505000000-1024 I check in the Azure portal and sure enough, the snapshot does exist.
|
||
It was at that point that I gave up. The product should be able to accomplish this simple task without me having to SSH into the box and go through log files to troubleshoot. I am having NetBackup flashbacks at this point, and I am not interested in going down that long, tortuous path to ruin.
|
||
As a side note, as I was letting these policies cook for a week, I got a call from someone at Veritas asking how things were going. That’s nice, I guess, but personally I don’t want anyone calling me. If I like the product and want to purchase it, I will reach out.
|
||
Based on my interaction with the CloudPoint product here are my thoughts.
|
||
This is a fairly immature product, I was surprised to see that it was version 2.0. The product team needs to be MUCH more helpful to those diving into the product The product interface need to be a bit more intuitive, with some guided wizards and inline assistance Integration into the cloud providers should be well documented and automated where possible a. I would really like for CloudPoint to generate the necessary Azure CLI or AWS CLI commands to create the necessary service account in each Tag early, tag often, tag everything – that is how things work in the cloud I think this should probably be a SaaS solution instead of something I have to spin up and manage myself Is this product ready to protect Production-level and scale workloads in the cloud? Not really. There’s a lot of good stuff here, but the product needs more time to bake in features and functionality. I think Veritas has a good solution here, and it’s going in the right direction. My impression is that the team behind CloudPoint is fairly small. If Veritas was waiting to get to MVP before heavily investing, then I would say that threshold has been crossed. Time to open the floodgates and get this product to a maturity level that I would feel comfortable recommending to clients.
|
||
`,summary:"At Cloud Field Day 3 we visited the Veritas office and they presented their cloud vision to us. It took a little while to ramp up, as I detailed in my post here. My fellow CFD delegate Martez Reed has already put an excellent post together detailing the high-level view of what Veritas had on display, so I won’t rehash that now. Instead, I would like to focus down on their CloudPoint offering and actually try and take it for a spin.",date:"6 May, 2018",url:"https://nedinthecloud.com/2018/05/06/veritas-360-cloudpoint/",image:"analysis.png",readingTime:"13"},"https://nedinthecloud.com/2018/04/11/innovate-or-die-three-ways-to-cloud-enablement/":{title:"Innovate or Die? Three Ways to Cloud Enablement",tags:["cfd3","cloud-field-day"],content:`Last week I participated in Cloud Field Day 3. If you’re not familiar with Cloud Field Day, then I would highly recommend checking out the post from my fellow delegate Nick Janetakis detailing his experience. It’s a thorough and well thought-out post describing what Cloud Field Day is, and why you might be interested.
|
||
As I watched each vendor present, I kept coming back to the same set of questions. The focus of this post is one of those. How can companies use the cloud to innovate their current product portfolio? There was a stark difference between those organizations that had embraced the cloud as an enabler of new solutions and those who had instead approached cloud like a check box on a list of things organizations should be doing. Broadly, I think the innovative companies - or the innovative branch of the company - fell into three broad categories.
|
||
Those that were “born in the cloud” Those that purchased an innovative startup Those that created a center of excellence to embrace innovation In the first category there were companies like Druva and Morpheus. These vendors got their start in tandem with cloud technologies, and so they naturally glommed onto cloud native approaches to technology. They might not be in a greenfield situation where they can pick the best of breed cloud tech for everything, but they also aren’t saddled with years of accumulated technical debt anchoring them to a legacy mindset and approach. Druva went to great lengths to explain how their SaaS data protection solution was a cloud native offering using AWS technologies to achieve their goals. That presentation actually led to a heated discussion with the delegates over whether you should care about where your SaaS solution is actually running - hint: it depends, but I would say not really. They were trying to show how their solution was architected to be massively scalable, provide unparalleled security, and drive cost optimizations that they pass on to the client.
|
||
The second category included a vendor that was actually quite a surprise to me, NetApp. Nothing against storage vendors, but I don’t look to them for any kind of real innovation in the cloud. Too many of the storage behemoths are busy burning marketing cash to explain why Hyperconverged and cloud are not the solution, and how their monolithic storage arrays are the one true and right way to storage nirvana. Simultaneously, they are developing virtualized versions of their software, and branding everything as Software Defined - a term which at this point carries the same weight as “new and improved” on paper towel packaging. So when I arrived at NetApp, I assumed that the presentation would be a conga line of “new and improved” storage products that were “Software Defined” and “Built for the Cloud” whilst all evidence spoke to the contrary. I was happily incorrect! The presentation was from Eiki Hrafnsson who was the CEO of GreenQloud, an Icelandic company purchased by NetApp last year. GreenQloud was focused on developing public cloud platforms, and NetApp purchased them in order to create a whole new solution. It leverages the ONTAP file system from NetApp, but that is where the similarities end. Eiki treated us to demos showing how their Cloud Volumes were able to outperform EBS volumes in AWS, and a full orchestrated container deployment using Kubernetes with Cloud Volumes providing native persistent container storage.
|
||
NetApp purchased an innovative company and left them alone to keep doing awesome things. If they can continue with that strategy and let it infect some of the more established portions of the organization, then I see a bright future for NetApp. Whether or not Eiki is still there in 12 months should be a solid indicator of whether NetApp’s old guard has allowed this project to thrive or crushed it with internal polictical machinations.
|
||
The third category is exemplified by another unexpected contender, Veritas. Yes, that Veritas. The presentation got off to a rocky start, with Veritas doing exactly what I had expected NetApp to do. They were talking about how their backup products could ship information up to the cloud, which is okay, I guess. Seriously, your backup product being able to use cloud storage is table stakes at best, along the same lines as being able to use VM snapshots. They even started bragging about their new backup appliances and how many petabytes it can hold. The delegates were getting restless, and finally Tim Crawford stopped the presenter and asked him to start over with some additional guidance. It was a tough love moment, but it paid off! Out of nowhere, Veritas brings out two new presenters that jump almost immediately into a demo of a cloud native backup solution that they developed from scratch. This solution, CloudPoint, was capable of backing up cloud workloads like Azure VMs, AWS RDS databases, and more. It was lightweight, and leveraged the built-in backup capabilities of cloud services, while also providing indexing, metadata management, and scheduling. The application ran in a container and had a lightweight UI. The product was clearly unpolished and rough around the edges, but it was a breath of fresh air to what had been a stifling atmosphere of superiority emanating from the Veritas folks. Be humble, stay humble, would be my advice.
|
||
What was particularly noteworthy about the Veritas offering is that the team responsible for developing it were pretty new to the organization. It felt as if someone at Veritas had hired the team, and told them to go off and make something cool with the cloud. And they did! I’ve heard of centers of excellence or centers of innovation at companies, and while no one said that’s what they were doing at Veritas, that’s exactly how it felt.
|
||
So there you go, three paths to cloud innovation. Be it, buy it, or add it. Just don’t let the haters squash your dreams.
|
||
`,summary:`Last week I participated in Cloud Field Day 3. If you’re not familiar with Cloud Field Day, then I would highly recommend checking out the post from my fellow delegate Nick Janetakis detailing his experience. It’s a thorough and well thought-out post describing what Cloud Field Day is, and why you might be interested.
|
||
As I watched each vendor present, I kept coming back to the same set of questions. The focus of this post is one of those.`,date:"11 Apr, 2018",url:"https://nedinthecloud.com/2018/04/11/innovate-or-die-three-ways-to-cloud-enablement/",image:"analysis.png",readingTime:"5"},"https://nedinthecloud.com/2018/04/03/my-life-as-a-tech-impostor/":{title:"My Life as a Tech Impostor",tags:[],content:`Before I even knew there was a term, I thought I was an impostor. Not in tech mind you, but in life. I was 11 years old. Over the summer I had visited my first Head Shop, ushered in by my older and ostensibly wiser cousin. I didn’t know what a bong was, or what all these dancing bears were about. The whole place stank of some unknown odor, which I would later be able to identify as a mélange of patchouli, sandalwood, and pot. Mostly pot. What I could dimly sense - as a budding, rebellious teenager - was that this place was cool, and I wanted to be cool. My cousin explained that the dancing bear was in fact a totem of The Grateful Dead, and I vaguely recognized the name from MTV. That would, of course, be the Touch of Grey single that served as an introduction of The Dead to many of my peer group.
|
||
The name was good, The Grateful Dead. The bear was cool. And the baja poncho bearing the emblem was also cool. Therefore - my eleven-year-old brain sang with glee - purchasing the baja and wearing it make me cool. QED. To show how cool I was, I wore that poncho unchallenged for the remainder of the summer. Claiming to be a Dead Head, and humming Touch of Grey, mostly because I didn’t know any other songs from my theoretical, favorite band. It wasn’t till the first day of school arrived that I was called on my bullshit. A schoolmate who was an occasional friend - more often a nemesis - challenged me to name another Grateful Dead song. If you’ve followed along so far, you won’t be surprised to discover that I couldn’t meet his challenge. And he called me what could easily be the most damning name an aspiring cool kid could hear, “Poser.”
|
||
That label, however accurate, stung deep. That very night I disposed of my poncho and the glamour I thought it had worked upon my status. I had been found out! I was a poser, a fake, a wannabe. And I was never going to do that again.
|
||
That single incident created within me a constant fear; a nagging voice in the back of my head that says, “You’re a fake and a liar, and sooner or later someone is going to call you on your bullshit.”
|
||
Fast forward several years, and several increasingly more awkward haircuts later, and I found myself in the tech industry. Getting started was easy, but gaining confidence was hard. Each technology I learned simply showed me how little I knew about the other technologies that radiate away from it. I thought I was a pretty sharp Active Directory admin, and then I dug a little deeper and realized that I maybe knew 5% of what makes up Active Directory. I thought I knew a lot about VMware, and then I started talking to people managing thousands of VMs, and I realized I maybe knew 30% of what makes up VMware vSphere. Same thing with networking, storage, servers, cloud, etc. Each time I gained a certain level of confidence in a topic, I would feel like someone else was calling bullshit. My old fear of being an impostor would rear its familiar, ugly head, and I would scamper away to my corner and lick my wounds.
|
||
But here’s the thing. The vast majority of people I met weren’t calling me an impostor. They didn’t think less of me for not knowing everything. Those sages of technology had advanced far enough in their career that they learned a central truth. No one knows everything about a topic. No one. If you think you do, I suggest that you talk to some other people about the topic, or recognize that you might be suffering from the Dunning-Kruger effect. So these bastions of tech knowledge knew that they didn’t know everything, knew that no one knew everything, and by extension knew that I didn’t know everything. That wasn’t a point of shame. It was a simple, subtle truth. All I needed to do was accept it, and ask for help. I didn’t need to fear being an impostor, as long as I didn’t claim to know everything. And it only took me ten years to figure that out!
|
||
What really opened my eyes was accepting a job as an IT consultant. There were some essential truths I had to acknowledge in order to be a good consultant:
|
||
You are surrounded by incredibly talented people who have deep knowledge on specific topics You are going to develop deep knowledge on specific topics, share accordingly Clients have an expectation that you are the expert, and you need to prepare for that There’s no shame in not knowing something, provided you do something about it You are going to be bombarded with a constant stream of new technologies, prepare for that too Accepting these truths has been difficult at times. I still struggle with the feeling that I am in some way an impostor. Each new area I engage in brings that old familiar feeling back. It doesn’t matter what level of accomplishment I have achieved in a given area, that nagging voice is never entirely quelled.
|
||
To a certain degree, that nagging voice is a driver; a source of motivation pushing me to learn more, do more, achieve more. It’s part of the reason that I got my Microsoft MCSE back in 2007. I didn’t need to for work. There was no significant financial remuneration for it. My hope was that in achieving the MCSE, I would no longer feel like an impostor. After all, Expert is in the certification for goodness sake! This is the same force that pushed me to try and become a Microsoft MVP, and a Pluralsight author, and probably some other nonsense in the future. It’s not the only reason, I don’t mean to imply that. But it is a contributing factor that helps drive me forward.
|
||
When I received the invitation to join Cloud Field Day, I was surprised. I had heard of Tech Field Day, and I knew some of the impressive people who were previous delegates. I considered many of these people to be celebrities of the tech industry, or at least my corner of it. People that I looked up to, and I thought had things “figured out”. How do I qualify to be invited to such a gathering? My friend the impostor syndrome raged forth in full force. Once again, I had to remind myself that I may not have everything “figured out”, but neither does anyone else. Those who claim to are probably operating under a delusion that reality will only be too happy to dispel them of. When I take a step back and look more objectively, it turns out that I’ve been hanging out in the tech industry for the last 16 years, and in the full course of time, I might actually have a bit of knowledge to share.
|
||
So if I’m an impostor, then we all are. And if we are all impostors, then nobody is. Know what you know. Learn what you can. Be generous those willing to learn. Learn from those willing to share. We’re all in this together.
|
||
`,summary:"Before I even knew there was a term, I thought I was an impostor. Not in tech mind you, but in life. I was 11 years old. Over the summer I had visited my first Head Shop, ushered in by my older and ostensibly wiser cousin. I didn’t know what a bong was, or what all these dancing bears were about. The whole place stank of some unknown odor, which I would later be able to identify as a mélange of patchouli, sandalwood, and pot.",date:"3 Apr, 2018",url:"https://nedinthecloud.com/2018/04/03/my-life-as-a-tech-impostor/",image:"featured-image-nedinthecloud.jpg",readingTime:"6"},"https://nedinthecloud.com/2018/03/06/azurestack-on-azure-part-3/":{title:"AzureStack on Azure - Part 3",tags:["azure","azure-stack","azurestack"],content:`In the last two parts we deployed an Azure Stack Development Kit on an Azure VM and got it registered with Azure. Then we created an Offer and Plan for the default user and started the download of marketplace items for use on Azure Stack. Now that those items have completed their download, we can move on to the process of installing the Resource Providers (RPs) for Microsoft SQL Server (MSSQL), MySQL Server, and the App Service. In this post I will cover the process and scripts you can use to get the MSSQL and MySQL RPs running. The App Service will be a separate post, due to the additional complexity involved.
|
||
The MSSQL and MySQL RPs consist of a VM running Windows Server 2016 Core that acts as a broker for the Resource Provider. It doesn’t actually provide the databases services itself, so you will need something to provide those services outside of the standard deployment. The parlance used by Azure Stack is a hosting server, which could actually be a cluster of servers - as in a SQL Server AlwaysOn Availability Group. In a production deployment you could choose to run the databases services as VMs on Azure Stack in a dedicated subscription, or you could provide the services from database servers that live outside the Azure Stack hardware. For the purposes of the ASDK, we are going to create a VM running Microsoft SQL Server 2017 on Windows Server 2016, at least that’s what I have specified in the template. Any Microsoft SQL installation 2014 and newer should work. You could even get SQL on Linux working if you really wanted to.
|
||
Within the directories that are pulled from GitHub is the C:\\AzureStackonAzureVM\\RPs\\MSSQL directory. That contains the deployment template for the VM that will supply the MSSQL database services. In order to deploy the template and Resource Provider, you will run the MSSQLDeploy.ps1 script. The script requires two parameters. First run the following to get values for those parameters.
|
||
$AzureStackAdmin = Get-Credential -Message "Enter Azure AD Admin credentials" $cloudAdminCreds = Get-Credential -Message "Enter Azure cloudAdmin credentials" If you are running the script from a non-default path, then make sure to also supply a value for the $defaultLocalPath parameter. The AzureStackAdmin should be the administrative account that you have been using to deploy Azure Stack resources.
|
||
The cloudAdminCreds will be the AzureStack\\cloudadmin account, which has the same password as the one used for the local admin account AzureStackAdmin.
|
||
Go ahead and run the script as shown below:
|
||
.\\MSSQLDeploy.ps1 -AzureStackAdmin $AzureStackAdmin -cloudAdminCreds $cloudAdminCreds The script will prompt you for a password to use for the local Administrator (mssqladmin) on the SQL Server, which will also serve as the sa password.
|
||
Then it will deploy the SQL 2017 server on AzureStack in the default provider subscription, creating a new resource group called MSSQL.
|
||
The template exposes port 1433 on its public IP address and uses a public DNS name of mssql01.local.cloudapp.azurestack.external. That value will be used to configure database services for the resource provider after it has been deployed.
|
||
After a successful deployment of MSSQL01, the script downloads and expands the latest MSSQL RP files from the internet and configures some of the parameters for the RP deployment. The VM running the services will be using the super secure password of “P@ssw0rd1”, though you can change that to something else in the script if you would like. Finally the script calls the DeploySQLProvider.ps1 script with the necessary parameters to deploy the VM and the resource provider components. The whole process will take a while, so I recommend going to get a caffeinated beverage or something similar.
|
||
Once the resource provider is done its work, you will need to set up a database server connection, and create a plan for the MSSQL provider. If you are already logged into the Admin portal, I recommend logging out and back in. The new provider may not show as available otherwise. Under all services in the Administrative Resources section, you will now see a SQL Hosting Servers section.
|
||
Clicking on that will take you to a blade where you can add, edit, and remove hosting servers that are providing MSSQL database services.
|
||
There doesn’t appear to be any PowerShell cmdlets for this yet… yet. It’s one more thing on my list of stuff to do. Click Add and go through the workflow to add a SQL hosting server. You will need the public IP address or public DNS name of the MSSQL server the script created earlier. Assuming you haven’t altered the script, the public DNS should be mssql01.local.cloudapp.azurestack.external.
|
||
You will also need to create a SKU that users can select when they are adding databases. The SKU is meant to define the capabilities being offered to the user, so things like version, edition, and HA status of the MSSQL server would all be applicable.
|
||
The SKU can take up to an hour to generate, so once you add the MSSQL RP to a plan it may still take an hour for the user to create a database.
|
||
Once the hosting server is successfully added, we can go ahead and create a Plan and Quota for the service and add it to the existing default offer.
|
||
This makes the MSSQL plan an optional add-on plan for subscribers using the default offer.
|
||
Once that is all saved, you can log in to the user portal and add a plan to your user subscription.
|
||
Once it’s added to the subscription, you will be able to deploy a SQL DB, assuming that the SKU has had time to be added to the service.
|
||
The process for the MySQL RP is very similar. The VM running the MySQL service is actually a Windows Server 2016 instance with MySQL installed on it. The script is sitting in the C:\\AzureStackonAzureVM\\RPs\\MySQL directory, and I creatively named it DeployMySQL.ps1. The parameters you need to supply are the same, so if you still have the PowerShell session open from running the MSSQL RP, you can go ahead and run:
|
||
.\\MySQLDeploy.ps1 -AzureStackAdmin $AzureStackAdmin -cloudAdminCreds $cloudAdminCreds The VM providing the service in this case will be MySQL01 with a public DNS of mysql01.local.cloudapp.azurestack.external. You will be prompted for the mysqladmin credentials, which will be used for both the local Administrator account on the VM and the sysadmin for the MySQL installation.
|
||
The resource provider VM again is using the super secret password of “P@ssw0rd1”. This script will take about as long as the MSSQL script to complete.
|
||
Now that it has been deployed successfully, the steps are the same as what we ran through to set up the MSSQL RP.
|
||
Add a hosting server with a SKU Add a Plan and Quota to an Offer Wait a suitable amount of time for the SKU to deploy Add the plan to the user subscription Deploy a MySQL DB Obviously capacity is somewhat limited on these VMs, but this is a PoC after all. You could instead choose to provide services from a VM running in Azure outside of the ASDK. I haven’t tried it, but I don’t see why it wouldn’t work. Maybe I can put that on the list of things to develop as well.
|
||
Next post we will be installing the App Service RP using a script, and I’ll talk a little about why you might not want to do that.
|
||
`,summary:"In the last two parts we deployed an Azure Stack Development Kit on an Azure VM and got it registered with Azure. Then we created an Offer and Plan for the default user and started the download of marketplace items for use on Azure Stack. Now that those items have completed their download, we can move on to the process of installing the Resource Providers (RPs) for Microsoft SQL Server (MSSQL), MySQL Server, and the App Service.",date:"6 Mar, 2018",url:"https://nedinthecloud.com/2018/03/06/azurestack-on-azure-part-3/",image:"tutorials.png",readingTime:"6"},"https://nedinthecloud.com/2018/02/28/azurestack-on-azure-part-2/":{title:"AzureStack on Azure – Part 2",tags:["azure","azure-stack"],content:`Let the registering begin! So you’ve got a working Azure Stack on Azure, which is like some kind of crazy inception thing. But let’s be honest with each other, there’s not a whole lot to do on Azure Stack at this point. You could like provision a VNet, maybe set up some sweet Resource Groups, but if you want to do something useful - like create some VMs or deploy services - you’re going to want to register your Azure Stack with Azure and syndicate with the marketplace. So let’s go ahead and do that using the Register-AzureStackLAB.ps1.
|
||
Azure Stack Registration The Register-AzureStackLAB.ps1 script is copied to the C:\\AzureStackonAzureVM folder during the creation of the Azure VM. For the script you will need to provide the CloudAdmin credentials
|
||
and the Azure Subscription ID
|
||
you’ll be using to register your Azure Stack.
|
||
You can go ahead and run the following two commands:
|
||
$cloudAdminCredential = Get-Credential -UserName "AzureStack\\CloudAdmin" -Message "CloudAdmin Credentials" $subscriptionID = "YourAzureSubscriptionID" Then execute the script:
|
||
.\\Register-AzureStackLAB.ps1 -CloudAdminCredential $cloudAdminCredential -AzureSubscriptionId $subscriptionID The script does a few useful things. It copies the GitHub files for the Resource Providers, which we’ll use in part three. It then removes the current Azure PowerShell modules and installs the correct version of the Azure and AzureStack PowerShell modules that will work with Azure Stack.
|
||
Then the script runs the Add-AzureRMAccount command, which replaces the Login-AzureRMAccount command, since Login is not an approved PowerShell verb. And someone in the PowerShell community got real angry about that. It was a whole thing, don’t ask. You’ll be prompted for an account that is an Owner of the Subscription.
|
||
Once that account is added, the script will register the AzureStack provider with your subscription, then download and import the Registration module from AzureStack-tools on GitHub. Maybe at some point the tools will make their way into the AzureStack module proper, but for now they’re still separate. The registration module is imported, and the Set-AzSRegistration command is run. That will invoke some remote commands on the AzS-ERCS01 privileged endpoint, which is why the CloudAdmin account is required. At the end of the script you should see something like this.
|
||
Azure Stack Offer and Plan Creation While we’re running scripts, let’s keep the party going. At this point your Azure Stack is syndicated with the Azure Marketplace, but there aren’t any user subscriptions on your Azure Stack to start consuming all this goodness. There is the Default Provider Subscription, which is used by the Admin console, but that is not how regular users are going to be consuming Azure Stack. Let’s go ahead and create a Plan and Offer, and then create a user subscription.
|
||
Plans are a collection of one or more services in Azure Stack and quotas for that service.
|
||
For instance, you might include the Compute service in a plan, but restrict the subscriber to 10 cores of compute. That’s pretty important on a limited system like the Azure Stack Dev Kit, or even on a multi-node deployment of Azure Stack.
|
||
An Offer is a collection of a Base Plan and some optional add-on plans.
|
||
So you could include the MSSQL resource provider as an add-on plan, which the subscriber can choose to activate. The Offers can be public or private, and you can control who can actually create a subscription based on an offer. Pretty heady stuff, I know. So we are going to create a basic plan and offer using the Create-OfferandPlan.ps1 script located in the same C:\\AzureStackonAzureVM folder. The only required argument is the identity of the Subscription Owner, which in my case I use the existing admin account:
|
||
You will be prompted to enter credentials, which the script uses to connect to the Default Provider Subscription.
|
||
The script retrieves the quotas stored in the Default Provider Subscription, creates a Resource Group for Plans and Offers - creatively called PlansAndOffers - and then creates a plan and offer with the standard Compute, Storage, Network, and Key Vault services. Lastly, the script creates a user subscription using the generated offer for the user specified.
|
||
Now we can log into the user portal and start consuming services.
|
||
But before we do that, let’s add some items from the marketplace so there are actually things to create!
|
||
Add Items from the Marketplace There’s not really a good way to add marketplace content from PowerShell yet… yet. Suffice to say I am working on it, and I’ll have something eventually. In the meantime, however, we’ll need to use the Admin portal. There should be a link on the desktop, or you can use the following address: https://adminportal.local.azurestack.external. Once logged in, you can click on the Marketplace Management item. The local Marketplace will be empty, so go ahead and click the Add from Azure button to be take to the full list of items. In order to add the MySQL, MSSQL, and AppService resource providers, you are going to need to download some specific images. The Windows Server 2016 images weigh in at a hefty 127GB per image! Suffice to say the download will take a while. In fact that’s probably a good place to stop the post. Next time we’ll be adding the MSSQL, MySQL, and AppService resource providers. See you then!
|
||
`,summary:"Let the registering begin! So you’ve got a working Azure Stack on Azure, which is like some kind of crazy inception thing. But let’s be honest with each other, there’s not a whole lot to do on Azure Stack at this point. You could like provision a VNet, maybe set up some sweet Resource Groups, but if you want to do something useful - like create some VMs or deploy services - you’re going to want to register your Azure Stack with Azure and syndicate with the marketplace.",date:"28 Feb, 2018",url:"https://nedinthecloud.com/2018/02/28/azurestack-on-azure-part-2/",image:"tutorials.png",readingTime:"5"},"https://nedinthecloud.com/2018/02/22/docker-swarm-on-azure-for-docker-deep-dive/":{title:"Docker Swarm on Azure for Docker Deep Dive",tags:["azure","docker"],content:`I’ve been working my way through the very excellent Docker Deep Dive book by Nigel Poulton. If you’ve been meaning to get into the whole Docker scene, this is the perfect book to get you started. Nigel doesn’t assume prior container knowledge, and he makes sure that the examples are cross platform and easily executed on your local laptop/desktop/whatever. That is, until you get to the section on Docker Swarm. Now instead of using a single Docker host, a la your local system, you now need six systems - three managers and three worker nodes. It’s entirely possible to spin those up on your local system - provided you have sufficient RAM, but I prefer to use the power of the cloud to get me there. See I might be working through the exercises on my laptop over lunch and then my desktop at night. I’d like to be able to access the cluster from whatever system I am working on, without deploying the cluster two or three times.
|
||
I decided to use Microsoft Azure and the Azure Cloud Shell to deploy the setup. I have an MSDN subscription with Azure credits, and these tiny VMs aren’t going to overrun my monthly allocation. I chose to use the Azure Cloud Shell so I could deploy the cluster using the Azure CLI, and interact with the cluster from any system that has a browser. If all that sounds pretty good to you, then’s here’s what you’ll need to do. First, you’ll need an Azure subscription, I’ll leave that as an exercise to the reader. Then you’ll need to login into the portal and launch the Azure Cloud Shell. That’s this icon
|
||
. That will open a window in the bottom half of the screen. If you’ve never used the Azure Cloud Shell before, then you will be prompted to select Bash or PowerShell and create a storage account to house your user home environment for the shell. The shell itself is a container spun up on demand, running either Bash (Linux) or PowerShell (Windows) and attached to the storage account to load your user directory. That storage account provides persistence across Cloud Shell sessions. The Cloud Shell is also already logged into Azure with the credentials you used to log into the portal, and it already has Azure PowerShell or the Azure CLI pre-installed. Lastly, it has the docker client bits installed including docker-machine, which is pretty critical to the next few items. First you’re going to want to select a Bash shell for this exercise. Then you need to select a subscription, if you have more than one.
|
||
az account list --query [].name That will give you the names of all your subscriptions. If you only have one, it should be selected by default. If you have a few, then pick the one you want to work with by running the following.
|
||
az account set -s [Subscription Name or ID] Now we’re ready to create some docker hosts! First we have to store the Subscription ID in a variable:
|
||
sub=$(az account show --query "id" -o tsv) Then we create some Docker hosts by using the docker-machine utility:
|
||
docker-machine create -d azure / --azure-subscription-id $sub / --azure-ssh-user azureuser / --azure-open-port 80 / --azure-size "Standard_A1_v2" / --azure-availability-set mgr_aset --azure-location eastus / mgr1 In the above command we are asking docker-machine to create an Azure VM using the supplied subscription. The username for ssh will be azureuser, and docker-machine will automatically generate the necessary certificate key-pair. We’re opening port 80, which is not strictly necessary. The VM size will be A1 v2 because it’s cheap and uses Standard LRS storage. We’re placing the VM in an availability set, since that is just good practice, and placing the VM in the East US region. What about the resource group and networking you might ask? If you don’t give docker-machine a resource group, it will use docker-machine by default, either creating that resource group or using an existing one. If you don’t supply it with a vnet or subnet, it will default to the docker-machine-vnet and docker-machine subnet. Again, it will create those networks if they don’t exist, or use an existing one if they do. The source image is Canonical Ubuntu 16.04 by default. The provisioning will look something like this:
|
||
Running pre-create checks... (mgr1) Completed machine pre-create checks. Creating machine... (mgr1) Querying existing resource group. name="docker-machine" (mgr1) Resource group "docker-machine" already exists. (mgr1) Configuring availability set. name="mgr_aset" (mgr1) Configuring network security group. name="mgr1-firewall" location="eastus" (mgr1) Querying if virtual network already exists. name="docker-machine-vnet" rg="docker-machine" location="eastus" (mgr1) Virtual network already exists. rg="docker-machine" location="eastus" name="docker-machine-vnet" (mgr1) Configuring subnet. name="docker-machine" vnet="docker-machine-vnet" cidr="192.168.0.0/16" (mgr1) Creating public IP address. name="mgr1-ip" static=false (mgr1) Creating network interface. name="mgr1-nic" (mgr1) Using existing storage account. name="vhds81zdtyaskjdhkad4l74bj" sku=Standard_LRS (mgr1) Creating virtual machine. name="mgr2" location="eastus" size="Standard_A1_v2" username="azureuser" osImage="canonical:UbuntuServer:16.04.0-LTS:latest" Waiting for machine to be running, this may take a few minutes... Detecting operating system of created instance... Waiting for SSH to be available... Detecting the provisioner... Provisioning with ubuntu(systemd)... Installing Docker... Copying certs to the local machine directory... Copying certs to the remote machine... Setting Docker configuration on the remote daemon... Checking connection to Docker... Docker is up and running! To see how to connect your Docker Client to the Docker Engine running on this virtual machine, run: docker-machine env mgr1 Now rinse and repeat the command for mgr2 and mgr3. For the worker nodes you can run the same command again except change the availability set to wrk_aset and the name to wrk1, wrk2, and wrk3. Once you’re finished, you can access the mgr1 node by running:
|
||
docker-machine ssh mgr1 Or you can run:
|
||
eval $(docker-machine env mgr1 --shell bash) And that will configure the local docker client to use mgr1 for docker commands. From here, you can follow along with Nigel’s Docker Swarm chapter, or mess around with Docker Swarm however it suits you. If you want to stop the VMs, go ahead and run:
|
||
docker-machine stop wrk1 wrk2 wrk3 docker-machine stop mgr1 mgr2 mgr3 vms_ids=$(az vm list -g docker-machine --query "[].id" -o tsv) az vm deallocate --ids $vms_ids Docker-machine will stop the VMs, but not deallocate them, so you’re still getting charged. The last two lines grab all the VMs in the docker-machine resource group and deallocate them. If you’re done with this experiment and want to get rid of them, you can run the following:
|
||
docker-machine rm wrk1 wrk2 wrk3 docker-machine rm mgr1 mgr2 mgr3 az group delete --name docker-machine --yes And you’re all cleaned up! Let me know if you find this helpful. And seriously, go buy Nigel’s book on Leanpub. And listen to a podcast I did with him if you’re feeling extra saucy.
|
||
`,summary:"I’ve been working my way through the very excellent Docker Deep Dive book by Nigel Poulton. If you’ve been meaning to get into the whole Docker scene, this is the perfect book to get you started. Nigel doesn’t assume prior container knowledge, and he makes sure that the examples are cross platform and easily executed on your local laptop/desktop/whatever. That is, until you get to the section on Docker Swarm. Now instead of using a single Docker host, a la your local system, you now need six systems - three managers and three worker nodes.",date:"22 Feb, 2018",url:"https://nedinthecloud.com/2018/02/22/docker-swarm-on-azure-for-docker-deep-dive/",image:"tutorials.png",readingTime:"6"},"https://nedinthecloud.com/2018/02/17/tortoise-and-hare-software-will-devour-us-all/":{title:"Tortoise and Hare, Software will devour us all!",tags:[],content:`Is Technology Moving Too Fast? Yes. Will It Slow Down? No.
|
||
As the spectre and meltdown drama continues to play itself out in the increasingly disdainful and disinterested public eye, a few things have come to my notice. The first is a post from The Math Citadel taking Intel and the IT field at large to task for having a fundamental design flaw in processor that is 20 years old. They posit that the ideas of “fail fast and fail often, the perfect is the enemy of the good, and just get something delivered” should be summarily rejected.
|
||
They rightly point out the following:
|
||
“Rushed thinking and a desperation to be seen as ‘first done’ with the most hype has led to complexity born of brute force solutions, with patches to fix holes discovered after release. When those patches inevitably break something else, more patches are applied to fix the first patches.”
|
||
They claim that in our rush to minimally viable product, and adopt Agile software development practices, we have set ourselves on a path of continual and ever increasing failure. Being mathematicians, their prescription is of course that IT needs to think more like them. Because as we all now know, mathematicians are rigorous and correct all of the time. Yes I linked a Cracked article, no I don’t see why that should matter.
|
||
The writer of the post clearly has an agenda, and a lens of perception coloring how they view the world. If I asked a baker to tell me how to fix what is wrong with IT, they might make allusions to careful measurement and have a recipe that is followed vigorously. It’s difficult to escape the trappings of your own occupation. But that doesn’t make the mathematicians wrong, just insufferably smug.
|
||
The other article is from Danny Crichton care of TechCrunch, and he points out a similar trend with an alarming number of examples of technology going wrong. I don’t really need to consult that since in the last week alone, my laptop has Green Screened on reboot, my phone had to be factory reset for “reasons”, Hulu live streaming was down for 3 hours during Friday primetime, and Skype’s login has been dodgier than usual for the last five days. Everything is broken, or in various states of broken, ever since Google started their forever Beta program. *Amen*
|
||
Crichton cites some interesting research about the need to maintain existing software rather than just speeding ahead to the next thing. He also points out that the massively complex systems we have put together are beyond our own ken, and small changes or failures can have a massive and totally unpredictable impact. Apparently Charles Perrow, a professor at Yale, calls these normal accidents. I prefer the more upbeat and Bob Ross inspired happy accidents. What are we to make steaming pile of technology that has been pooped out over the last ten years? Do we hold our nose and smile? Crichton brings up the fact that it is possible to make a highly available, and resilient complex system, such as modern US aviation. That makes sense when people’s lives are literally on the line, and not so much when the stakes are whether or not I can view pictures of tacocat whenever I want on my phone. What is the motivation for a company to create product that is incredibly stable, at triple the cost, for none of the return? Consumers have become completely numb to the constant broken state of their technology. The only time it becomes patently obvious is when you try to teach a loved one how to use some new piece of gadgetry, only to discover how broken and non-intuitive the product is. Apple’s no bastion of light these days, but I will hand it to the Apple of five years ago. They made a rock solid product, and ruled the ecosystem with an iron fist, much to the appreciation of their stockholders. That level of quality is, erm, lacking somewhat these days.
|
||
I don’t know of a simple way out it, but I think it will probably end up being AI. Our systems have become so ludicrously complex, and unfathomably labyrinthine that mere mortals stand a snowball’s chance in hell of taming the beast. I suspect that the only way that our foundations get shored up is if a neutral third-party with infinite patience, computational power, and memory is able to start fixing things for us. The proverbial mother cleaning up the children’s’ mess so that we may make another mess tomorrow.
|
||
Or we’ll just accidentally create our own oblivion. I feel like the odds are pretty even.
|
||
`,summary:`Is Technology Moving Too Fast? Yes. Will It Slow Down? No.
|
||
As the spectre and meltdown drama continues to play itself out in the increasingly disdainful and disinterested public eye, a few things have come to my notice. The first is a post from The Math Citadel taking Intel and the IT field at large to task for having a fundamental design flaw in processor that is 20 years old. They posit that the ideas of “fail fast and fail often, the perfect is the enemy of the good, and just get something delivered” should be summarily rejected.`,date:"17 Feb, 2018",url:"https://nedinthecloud.com/2018/02/17/tortoise-and-hare-software-will-devour-us-all/",image:"featured-image-nedinthecloud.jpg",readingTime:"4"},"https://nedinthecloud.com/2018/02/16/azurestack-on-azure-part-1/":{title:"AzureStack on Azure - Part 1",tags:["azure","azure-stack","azurestack"],content:`With the introduction of Dv3 and Ev3 VMs in Microsoft Azure, it became possible to run nested virtualization on Azure. Since I’ve got Azure Stack on the brain these days, my immediate thought was, “I wonder if I can run Azure Stack on Azure?” (cue Inception music). Not only was the answer yes, but others had already started the process for me. Following in the footsteps of Daniel Neumann and Florent Appointaire, I was able to bet the process running. One of the engineers at Microsoft took some of that work, added their special sauce and rolled out a GitHub repo that helps you through the process. I have forked that repo, and started adding some automation myself.
|
||
In this series of posts I will show you how to deploy the Azure Stack Development Kit (ASDK) on Azure, perform the initial configuration tasks, and add additional services using resource providers. Let’s begin!
|
||
First of all, before we go any further, please be aware that this will cost money. The hardware requirements for the ASDK are 12 cores, 96BG RAM, and four data drives, each at a minimum of 140GB. Part of the reason to run the ASDK in Azure is that many IT Pros and Developers will not have a machine hanging around their home or work lab that meets these requirements. The VM size I will be using is an E16S v3. It will be using premium storage with 4x 256GB data disks and a 256GB OS disk. I will also be using a Azure Hybrid Use license. Using the Azure cost calculator, I can see that running in the East US region is going to cost $778.85 per month for the VM and another $190.06 per month for the storage, bringing it to a grand total of $968.91 per month to run the whole thing. Of course, you can shut the VM down when you’re not working on it, bringing the price down to $1.06 per hour for the VM. The premium storage, however, is charged while the disk exists. If you plan to follow along and deploy this yourself, it is going to cost you money in Azure.
|
||
Second disclaimer - wow now I feel like a lawyer - this scenario is supported for the ASDK, but it is not meant for any type of production workload. That goes for the ASDK in general, but especially running the ASDK on Azure. The nested virtualization can get a little buggy sometimes, and the process has been known to BSOD during the network configuration.
|
||
OK, with all the hand waving out of the way, here’s the process. First deploy the template from my GitHub repo. There’s a nice shiny button, or you could use PowerShell. As much as I love PowerShell, I cannot resist a nice shiny button. Clicking said button takes you to the portal where you can fill out the required parameters. Auto Shutdown is enabled by default - trying to save you money - but I usually turn it off. Once everything looks good, accept the terms and conditions and then deploy the template. That is going to take a while, so let’s look at what it’s actually doing. To the JSON!
|
||
"resources": [ { "type": "extensions", "name": "CustomScriptExtension", "apiVersion": "2016-04-30-preview", "location": "[resourceGroup().location]", "dependsOn": [ "[parameters('virtualMachineName')]" ], "properties": { "publisher": "Microsoft.Compute", "type": "CustomScriptExtension", "typeHandlerVersion": "1.8", "autoUpgradeMinorVersion": true, "settings": { "fileUris": [ "[variables('fileUri')]" ], "commandToExecute": "[concat('powershell.exe -ExecutionPolicy Unrestricted -File ', variables('scriptFileName'), ' ', variables('adminUsername'))]" } } } ] In the resources section there is the customScriptExtension, which executes a PowerShell script defined in the variables.
|
||
"scriptFileName": "post-config.ps1", "scriptPath": "https://raw.githubusercontent.com/ned1313/AzureStack-VM-PoC/master/scripts/", "fileUri": "[concat(variables('scriptPath'), variables('scriptFileName'))]", The script in this case points to the GitHub repository, in the scripts folder, to a script called post-config.ps1. That scripts does some helpful things.
|
||
Create a local directory (AzureStackOnAzureVM) to place files in Disable IE’s Enhanced Security Configuration so that you can actually download files and visit websites Make a bunch of registry key changes to enable Credential Delegation and allow the ASDK VMs to interact with the host VM as needed Add some of the Windows features needed to function as the ASDK host Rename the local Administrator to Administrator Set the PowerShell execution policy to Unrestricted Download a PowerShell script for the ASDK installation process Download the ASDK downloader executable 9. Add the rest of the Windows features that require a restart, and then restart the VM Once all those steps are complete, we will need to log into the VM to continue the process. Now remember that the local admin account has been renamed to Administrator, so you will use .\\Administrator as the username to login. On the desktop you should see the next step of the process, a shortcut to the Install-ASDK.ps1 file that was downloaded in during the post-config.ps1 script. Rather than run it directly, let’s take a look at the contents. Open the file in an Administrative PowerShell ISE window. You can run the Install-ASDK.ps1 script now, and it will prompt you for all the necessary values. This script checks to see what versions of the ASDK are available, starting with version 1711. Then it will prompt you for the local Administrator password and verify it. Then it will prompt for a Tenant Admin to use with the ASDK. I created a Tenant Admin user for this purpose called AzsAdmin1@azslab.us. Now the scripts asks which version of the ASDK to install, if you’re not sure just pick the most recent. If everything looks good, press any key and the process will begin.
|
||
The script will now download the appropriate ASDK files and merge them into the CloudBuilder.vhdx. Then it will mount the vhdx file and extract the necessary components for the host VM to become the ASDK host. In a typical deployment of the ASDK, you would set the vhdx as a boot option and reboot directly into it. Since that is not an option when running in Azure, instead we pull out the deployment files from the vhdx. Then the script makes a few setting tweaks to allow the ASDK to run in a nested virtualization mode, rather than on a physical server. With that work out of the way, the script then downloads some additional scrips from the GitHub repo and schedules a job to run at next restart, invoking the ASDKCompanionService.ps1 script. This script is what configures the networking component for the BGPNAT VM, and it will kick off after the host VM is restarted. Finally, the script kicks off the ASDK installer PowerShell script, passing it a parameter block for configuration.
|
||
This process will take a while, like a really long time. My first run took about 3 hours. Now would be a good time to get a cup of coffee, run a half marathon, and write your mother back since she hasn’t heard from you in a while and you know how she worries.
|
||
During the network configuration, you will likely be disconnected from the RDP session. That is normal, and it should come back after about one minute. Later you will be logged out of the system, this is because the host VM will be joined to the local ASDK Active Directory. After the reboot, you can log back in as AzureStack\\AzureStackAdmin - using the same password as the local Administrator - to continue watching the process. It is at this point that the ASDKCompanionService starts checking to see if the networking configuration of the ASDK has created the internal switch and the BGPNAT VM. It will create an additional internal switch, change the BGPNAT VM adapter and fix the IP addressing. This is where the process typically throws a BSOD if it is going to. If you get past this point, the deployment will likely complete successfully. Now we’ve got ourselves a deployment of Azure Stack running in Azure. In the next post I will through the process of registering the Azure Stack with Azure for marketplace syndication, and creating some offers and plans for Azure Stack users.
|
||
`,summary:"With the introduction of Dv3 and Ev3 VMs in Microsoft Azure, it became possible to run nested virtualization on Azure. Since I’ve got Azure Stack on the brain these days, my immediate thought was, “I wonder if I can run Azure Stack on Azure?” (cue Inception music). Not only was the answer yes, but others had already started the process for me. Following in the footsteps of Daniel Neumann and Florent Appointaire, I was able to bet the process running.",date:"16 Feb, 2018",url:"https://nedinthecloud.com/2018/02/16/azurestack-on-azure-part-1/",image:"tutorials.png",readingTime:"7"},"https://nedinthecloud.com/2018/02/14/network-disaggregation-tsunamis/":{title:"Network Disaggregation Tsunamis",tags:[],content:`AT&T, a company that I generally unleash scorn upon for their cell phone service, has actually done something fairly interesting. On Jan 29th they announced that they would be releasing their dNOS (distributed network operating system) to the Linux Foundation. Now before you roll your eyes and quote Jessie Frazelle, who you should be following on Twitter and not one of the garbage Kardashians, I am aware that sometimes orgs donate their project to the Linux Foundation and leave itto languish and die in the hot and unforgiving light of the desert sun. But I don’t think dNOS falls under this particular category. AT&T has not only developed dNOS internally, they have a working prototype of it on production hardware possibly in actual production. I mean that’s the way the whitepaper reads.
|
||
So what is dNOS and why is AT&T so psyched about it? The concept behind dNOS is the development of an open source operating system for network hardware, that can run on commodity gear, so called whiteboxes, though why’s it gotta be white? What about pink boxes, or taupe? The reason AT&T is so jazzed about this idea is the rather high cost of the switches and routers they use to run their carrier grade networks. These boxes are vertically integrated using custom hardware, custom software, and proprietary everything. This is not only a large cost to AT&T, but it also slows their innovation cycle as they are at the mercy of the vendor when asking for new features.
|
||
I’ve mentioned network disaggregation before, going so far as to predict that we would see significant progress in 2017. That may have been a little too aggressive, but there were a lot of key components leading up to this. dNOS was announced in November of 2017. The P4 open source programming language also started gaining momentum in 2017. Barefoot Networks released their Tofino programmable ASIC, and Broadcom released their Tomahawk processor that is more than capable of handling the speeds and feeds of a carrier. Now in 2018 we have the introduction of the Linux Foundation Networking Fund, the release of an open-source SDK for the Broadcom Tomahawk chipset, and this announcement of dNOS being given to the Linux Foundation. Things may have gotten off to a slow start, but I feel confident that we are reaching critical mass. And I’m not even going to get into the new open-source, reduced cost optics that Facebook is pushing.
|
||
Basically the world of networking is in for a major shakeup, and the tide of open source and disaggregation is going to spur some incredible innovation. The major cloud players and the carriers will see the first fruits of their labor, but all that innovation is definitely going to trickle down to the Enterprise and SMB markets. With the coming Tsunami of IoT devices that will be thirsty for bandwidth and advanced networking solutions, this renaissance of networking cannot come soon enough.
|
||
`,summary:"AT&T, a company that I generally unleash scorn upon for their cell phone service, has actually done something fairly interesting. On Jan 29th they announced that they would be releasing their dNOS (distributed network operating system) to the Linux Foundation. Now before you roll your eyes and quote Jessie Frazelle, who you should be following on Twitter and not one of the garbage Kardashians, I am aware that sometimes orgs donate their project to the Linux Foundation and leave itto languish and die in the hot and unforgiving light of the desert sun.",date:"14 Feb, 2018",url:"https://nedinthecloud.com/2018/02/14/network-disaggregation-tsunamis/",image:"analysis.png",readingTime:"3"},"https://nedinthecloud.com/2017/11/30/vmware-on-azure-youre-still-doing-it-wrong/":{title:"VMware on Azure - You're still doing it wrong",tags:["aws","azure","cloud","microsoft","nutanix","vmware"],content:`Sigh. There’s an old adage that I always come back to. Just because you can do something, doesn’t mean that you should. In this case I am thinking about the recent announcement by Microsoft that Azure would be supporting bare metal deployments of VMware on Azure hardware. In case you’ve been living under a rock, AWS went GA with a very similar offering back in late August. Of course there are some specifics that differ, but the overall theme is the same. You can run your VMware workloads in their public cloud on bare metal, but still have close proximity to their respective public cloud services. Alas, just because it’s on Azure now, doesn’t make the idea any better, and I stand by my previous post.
|
||
The actual announcement and the subsequent gnashing of teeth is covered pretty well on the latest episode of Buffer Overflow. I’ll sum up here just in case. On November 21st Microsoft dropped the blog post called “Transforming your VMware environment with Microsoft Azure” which seems fairly innocuous. Most of the post is actually about the newly announced Azure Migration service. It’s not exactly a secret that Microsoft would very much like you to run everything on Azure if possible, on Hyper-V if you can’t, or on VMware if you must. The relationship between Microsoft and VMware has always been a bit strained, like they know that they need each other, but kinda wish they didn’t. Sort of like the UN and the US; shaking hands and smiling unconvincingly, all the while trying to crush the other’s hand to pulp.
|
||
The central issue is that Microsoft has a hypervisor, operating system, and successful public cloud. VMware has well, it’s a really nice hypervisor. And since System Center VMM is contortionists idea of easy to use software, VMware handily won the on-premises war of hypervisors and the management thereof (aka the SDDC). Sadly, things are getting cloudy for VMware, and so grasping at straws (can’t spell straws without AWS) they reached out to Amazon Web Services. After all, the enemy of my enemy, etc.
|
||
So, coming back around to the actual post, apparently some workloads are so special that they will only run in VMware, and so these magical prancing unicorns will be able to run on a full VMware stack running in Azure. At first I was thinking maybe it was nested virtualization, but no, it is a “bare-metal solution that runs the full VMware stack on Azure hardware, co-located with other Azure services.” It is going to be offered in partnership with a premier VMware certified partner and generally available next year.
|
||
A couple things:
|
||
Azure hardware is a tricky wording. Are they talking about Microsoft’s own Open Compute Project compliant hardware that runs other Azure workloads being repurposed to run VMware? Who are these mysterious VMware-certified partners? Why not name any of them, unless they aren’t ready to name one yet? And why not VMware itself? Well, someone apparently let VMware know about this whole thing. And they were, shall we say surprised and a little sour about it? Ajay Patel, SVP of Product Development in Cloud Services (so probably heavily involved in the whole VMware on AWS thing) wrote a by turns snarky, gloating, and somewhat petulant post on the VMware blog, saying such wonderful things as:
|
||
“This offering has been developed independent of VMware, and is neither certified nor supported by VMware.”
|
||
“VMware does not recommend and will not support customers running on the Azure announced partner offering.”
|
||
“Microsoft recognizing the leadership position of VMware’s offering and exploring support for VMware on Azure as a superior and necessary solution for customers over Hyper-V or native Azure Stack environments is understandable but, we do not believe this approach will offer customers a good solution to their hybrid or multi-cloud future.”
|
||
“VMware HCX technologies, a superior alternative to Azure Migration Service, enables organizations to migrate and manage hybrid cloud deployments.”
|
||
Ajay goes on to extoll the virtues of VMware’s partnership with AWS, OVH, and the VMware Cloud Foundation. Which I mean sure, you have to plant your flag somewhere, but considering that most analysts are unimpressed with VMware’s hybrid cloud motions, it takes a little bit of gall and a not insignificant amount to self-delusion to crow about your vast superiority to Azure’s offerings.
|
||
The speculation from The Register is that the mysterious partner that Microsoft is working with is in fact Nutanix, who is known for their own brash and brazen approach to the truth and marketing. Last year, Nutanix claimed that their Cisco based UCS deployment was validated and fully supported, to which Cisco replied, “um WHUT? No no, tis not true. Hyperflex is awesome, thank you for your time.” I find this line of reasoning possible, if not probable, although it is just as likely that Microsoft is building this with OCP servers and having some poor VMware partner try and engineer this thing together.
|
||
Microsoft subsequently announced a webinar in which inquiring minds could learn more about the solution, scheduled for 11/28. That webinar was then postponed until December 13th, for reasons that were not forthcoming. Maybe Microsoft did not anticipate the kerfuffle that would occur? Or maybe they have to do a little legal wrangling before they make any more announcements? Regardless, the whole solution is a bit silly.
|
||
What Microsoft is trying to do is make it easier to get VMware workloads into Azure. At least in the short term. In the long term, they want clients to migrate those workloads to Azure proper. This is the same thing that AWS is trying to do. AWS had the good grace to at least appear to partner with VMware whilst trying to destroy them, while Microsoft appears to be uninterested in offering up that poisoned apple at all.
|
||
Since I am a Microsoft MVP in Cloud and Datacenter, most people would reasonably assume that I am an Azure fanboy and that had colored my opinion on the whole VMware on AWS thing. Reasonable, but false. A bad idea on AWS is also a bad idea on Azure. I get why Azure and AWS are doing this. But I still say, if you’re running VMware in the cloud, you’re doing it wrong.
|
||
`,summary:"Sigh. There’s an old adage that I always come back to. Just because you can do something, doesn’t mean that you should. In this case I am thinking about the recent announcement by Microsoft that Azure would be supporting bare metal deployments of VMware on Azure hardware. In case you’ve been living under a rock, AWS went GA with a very similar offering back in late August. Of course there are some specifics that differ, but the overall theme is the same.",date:"30 Nov, 2017",url:"https://nedinthecloud.com/2017/11/30/vmware-on-azure-youre-still-doing-it-wrong/",image:"checkmark-circle.png",readingTime:"5"},"https://nedinthecloud.com/2017/09/27/vmware-on-aws-youre-doing-it-wrong/":{title:"VMware on AWS - You're doing it wrong",tags:["aws","cloud","vmware"],content:`This is going to be a controversial post I am almost certain. Basically, I am going to argue that the whole premise behind running VMware on AWS is fundamentally flawed and not a viable strategy for those who are currently running VMware or for VMware itself as a company. Get your angry comments ready, here we go!
|
||
The genesis behind this was an innocent LinkedIn post I made. Before we dive into the topic, I want to level set a little. First of all, I really like VMware’s virtualization product. I have been using ESX since the 3.5 days, and it has been nothing short of revolutionary in the datacenter. The first time I launched the vSphere client and realized what I could do with virtualization, I felt a wave of excitement wash over me. This was the future! And I was hitching my wagon to go along for the ride. That was in 2007. Checking my calendar, I see that it has been a decade, and sure enough virtualization has swept across the datacenter leaving a trail of positive destruction in its wake.
|
||
The point is that I really like VMware. This is coming from a place of love.
|
||
People have called VMware and virtualization disruptive, and it was to a certain degree; but in terms of application and operating system design, it wasn’t. It was simple, trivial even, to take an existing physical system and virtualize it. You didn’t need to rewrite the application or reinvent the operating system. Existing workloads chugged along, blissfully unaware that they were no longer sitting on bare metal. It was transformational for IT Ops, but not so much for application development.
|
||
Cloud native computing is not that.
|
||
It is entirely possible to lift and shift your existing workloads and the VMs they run on up to the cloud. However, that is certainly not the best way to take advantage of the public cloud. In my experience, you are going to end up paying more to run your VMs in Azure or AWS than you would in a local datacenter or a co-location facility. And you aren’t getting the true benefits of cloud native computing. I don’t want to wander into a discussion of the 12-factor application, suffice to say that a true cloud native application is very different than a traditional application you might have in your datacenter today. Platform as a Service enables those cloud native apps to run in the cloud effectively and efficiently.
|
||
So what is the point of the VMware on AWS offering?
|
||
Let’s look at the sales material:
|
||
Simple and Consistent Operations Flexibility to Suit Your Business Needs Enterprise-Grade Capabilities Delivered as a Service from VMware Simple and Consistent Operations Sounds to me like your sys admins don’t want to learn about this new-fangled cloud thing. Wrap them up in a comfy VMblanket and let them take a nap in the hot aisle of the datacenter. Seriously though, a simple consistent interface for your infrastructure is the holy grail, and no one has come even close to making it a reality. Why? Well you already have a single pane of glass, it’s called your monitor. Beyond that there are simply too many disparate systems to expect all of them to integrate nicely into a single management console and platform. Not having your admins learn how to use cloud services is simply kicking the can down the lane for another year of two. At which point your admins and possibly your company will be obsolete.
|
||
Flexibility to Suit Your Business Needs All long as VMware on AWS suits your business needs, then sure. Of course if you are going to branch out into AWS, then you’ve violated selling point one, since you now have to manage those resources through AWS. They also talk about flexible consumption models, which is a bit laughable when you actually look at the economics of the offering. You need a minimum 4-node deployment in AWS for VMware. That deployment will run you about $32 an hour at the list price. And it’s not like you can just spin these up and down on a whim, there’s a decent ramp up time to get it all running and you have to coordinate with VMware and set up the integration with your existing vSphere environment. Can you get some price breaks? Sure, but then you have to make an up-front commitment of 1-3 years. That doesn’t sound super flexible. You know what’s really flexible? EC2. I can spin up instances on demand in a VPC and only pay for them while they are running. If I need connectivity back to my environment, then I can use S2S VPN and use the new NSX-T to provide seamless connectivity. Although layer 2 across multiple datacenters is a terrible idea and is only used to help legacy applications limp along rather than modernizing them.
|
||
Enterprise Grade Capabilities Um, OK? The marketing here starts to get really fuzzy. Something about using NSX, vSAN, and vSphere on AWS bare-metal to achieve elastic, next generation… ugh I couldn’t even get through the last few terms. What are they really trying to say here? You should use vSAN in AWS instead of using their EBS or S3 storage? That’s patently ludicrous. If you want to talk about elastic, then it doesn’t get much more elastic than EBS and S3 (it’s literally in the name). Plus, vSAN is limited by the number of nodes in a cluster and some upper limits on the architecture itself. At release, the vSAN cluster will be limited to 16 hosts, so maybe you’ll be able to provision a VM with a 10TB disk? What’s the maximum volume size in EBS? 16TB and you can create multiple 16TB volumes and RAID them if you really needed to. S3 is essentially limitless in terms of storage. As I mentioned you could use NSX-T to get NSX in your VPC. For management, well like I just said you’re going to be using multiple management tools anyhow.
|
||
Delivered as a Service from VMware At least you don’t have to manage the underlying hypervisors or physical hardware, which is exactly what you would already get with any major public cloud vendor. Or in the case of a vCloud Air partner, this is what you can already get today. So again, what is the tremendous differentiator?
|
||
When you boil it all down, you basically have a managed colo offering sitting in AWS’ datacenters.
|
||
You’re Doing it Wrong. If what you want is to run workloads in AWS, why wouldn’t you just use AWS’ services?
|
||
If you need DRaaS, there are many much cheaper options to replicate your VMs to a public cloud without an entire vSphere cluster that is always running and costing you money.
|
||
If you need IaaS, then just use the public cloud IaaS services. They are easy to understand! And that’s the direction you’re likely heading anyhow.
|
||
You’re not going to VMotion workloads dynamically to VMware on AWS. Stop it. I know it’s cool. That doesn’t make it a good idea. You’re enabling bad application practices to perpetuate.
|
||
Cloud native computing is the future platform for modern applications. I think VMware knows this, and for the interim VMware on AWS was the closest they could introduce to keep their investors happy. But it’s the wrong approach and I think in five years time they will end up abandoning it in favor of something more modern.
|
||
`,summary:`This is going to be a controversial post I am almost certain. Basically, I am going to argue that the whole premise behind running VMware on AWS is fundamentally flawed and not a viable strategy for those who are currently running VMware or for VMware itself as a company. Get your angry comments ready, here we go!
|
||
The genesis behind this was an innocent LinkedIn post I made. Before we dive into the topic, I want to level set a little.`,date:"27 Sep, 2017",url:"https://nedinthecloud.com/2017/09/27/vmware-on-aws-youre-doing-it-wrong/",image:"analysis.png",readingTime:"6"},"https://nedinthecloud.com/2017/09/27/whats-in-a-name/":{title:"What's in a Name?",tags:[],content:`As the raging dumpster fire that is the Equifax breach continues to unfold, I find that I am thinking about identity and the way we use it in our modern life. Equifax was criminally negligent with information that was incredibly valuable to individuals. They should be penalized as an organization with fines and levies, and some of the individuals within the company who were responsible for the security of our data should face possible jail time. But when you step back for a moment, it becomes readily apparent that this is just the latest in a series of data breaches over the past decade, and despite fines, levies, and jail time; this is the sort of thing that is likely to happen again. Why? First, the monetary value of the information is high, meaning that criminal elements are willing to spend the resources to steal the information. Second, organizations are rarely incentivized to take the necessary precautions to secure data. As Greg Ferro likes to point out, as long as the cost of true security is higher than the cost of a breach, organizations are unlikely to adopt true security practices. Third, even if an organization tries to embrace true security, human beings are fallible. Applications have undiscovered exploits, misconfigurations happen, and hackers are always stepping up their game.
|
||
Simply saying that an organization needs to be more secure is not addressing the root of the problem, but rather a symptom of that problem. In my mind the primary issue is the intrinsic value of the social security number, date of birth, and other personally identifying information. How do we address this issue? The first thing I would say is to reduce the value of some of these numbers. The SSN is not valuable because of its original intended purpose. Originally instituted in 1936, it was used to keep track of US workers’ earnings histories in order to determine their social security eligibility and benefits. It was never intended to be used as a universal identifier for everyone in the country, and the just because it has become a de facto standard doesn’t mean that it should be. It is patently ridiculous that we expect a single number to uniquely identify everyone, and have everyone keep their number secret. Benjamin Franklin famously said, “Three can keep a secret, if two of them are dead.” There’s no way that your SSN can stay secret if you’re expected to give it during a credit application, your college application, to your doctor’s office, etc. If just one of them spills the beans, then the jig is up! And that’s exactly what happened with Equifax. Unique identifiers cannot also be secrets by their very nature. The SSN cannot be one (secret) and shouldn’t be the other (unique id). Those are two separate problems that require two different solutions.
|
||
I’d like to go on a slightly philosophical tangent here about identity. A more fundamental question to ask is, do we need a universal unique identifier for each person on the planet. The engineer in me says, yes of course. But the rest of me bristles at the idea. What is the true utility of a universal ID? Can I, as an individual, demand that I be forgotten? A UUID is a link between my past and future, and there may legitimately be times that I would like to sever that linkage. Think about how someone would be identified 200 years ago. It’s unlikely there would be a photograph. There were no fingerprint or DNA databases. Aside from relying on another person for identification, there was really no way to confirm someone’s identity. There’s a certain freedom in that, as well as the potential for abuse.
|
||
So do we need a UUID and who needs it? Let’s think about some of the institutions that use your SSN now. Your doctor asks for your SSN, but they don’t need it. They need a system to identify you for a couple reasons:
|
||
Your medical history Health insurance and billing Technically your medical history belongs to you. It’s useful for your doctor to have and be able to share with others. But there is no reason they need to use the SSN. What if instead you provided them a signed, access token to your medical records. When other institution needed access, you could grant them access with a unique token as well. In the event that you were done working with a particular doctor’s office, you could revoke the token and thereby their access. This is your information after all, and the doctors charge you plenty to add information to it.
|
||
Health insurance and billing is mostly handled today between your doctor’s office and the insurance company in a manner I can only describe as arcane and byzantine. There are entire industries that exist solely to navigate the twisted morass that is our healthcare debacle… I mean system. The point is, this is their problem, not yours. So it’s up to them to give you a unique identifier, which they do. It’s on your healthcare card. If they need to verify that the person making the claim is actually you, then they should set up a two-factor confirmation system. In that case, they could issue you a private key that you use to generate signed confirmation messages.
|
||
In both examples we are using a private/public key pair, wherein you never give anyone access to your private key. This enables you to both verify your identity and control access to sensitive information. This is security 101. Additionally, the unique identifier that your doctor and insurance company use for you is different and unrelated. If your insurance company is compromised, the information cannot then be used to access your doctor’s records, let alone your financial history and bank accounts.
|
||
Reduce the value of the information and limit the scope of usefulness for each ID.
|
||
What if your private key is compromised? Well the good news is that you would be to generate a new private key, and void all existing public keys and certificates tied to the original private key. Of course the burden of proof would be higher to issue the void command, but at least it isn’t a 9 digit number that you cannot change and that haunts you for your entire life.
|
||
Of course I realize that instituting such a program would be a massive undertaking both financially and logistically. Then again, how much is the current situation impacting us economically? The amount of money that credit card companies and financial institutions have to spend paying off fraud should easily balance out the cost of creating a better way to identify people. It’s in our best interest as the public, and it’s in the best interest of the banks, lenders, and credit agencies. Heck it’s even in the best interest of the government. I don’t expect immediate change since all of the systems tend to move slowly, but an important first step would be for the SSA (the Social Security Administration) to condemn and make a policy against the continued use of the SSN as a universal identifier in any system.
|
||
`,summary:"As the raging dumpster fire that is the Equifax breach continues to unfold, I find that I am thinking about identity and the way we use it in our modern life. Equifax was criminally negligent with information that was incredibly valuable to individuals. They should be penalized as an organization with fines and levies, and some of the individuals within the company who were responsible for the security of our data should face possible jail time.",date:"27 Sep, 2017",url:"https://nedinthecloud.com/2017/09/27/whats-in-a-name/",image:"featured-image-nedinthecloud.jpg",readingTime:"6"},"https://nedinthecloud.com/2017/08/24/windows-hosts-with-kubernetes-the-beginning/":{title:"Windows Hosts with Kubernetes - The Beginning",tags:["azure","containers","kubernetes","terraform"],content:`Well, it wasn’t even close. As mentioned in my previous post, I am moving to a less hands on role, and I want to keep close to the technology. The concept of running Windows container hosts in a Kubernetes cluster fascinates me and it appears that I wasn’t alone. With 82% of the votes on my Twitter poll, it was the clear winner. Now I guess I actually need to start diving in, and by diving in, I mean reading docs.
|
||
Based on what I have read so far, I plan to use Tectonic to deploy Kubernetes to Azure using Terraform. If I’m using Microsoft, might as well go whole hog right? The master nodes need to be Linux and running Kubernetes version 1.5 to support the Windows Containers. I will also need Windows Servers running Server 2016 RTM or newer with support for the Hyper-V and Containers roles. Microsoft recently announced support for nested virtualization within Azure for certain Azure VM families. That will give me the Windows container hosts I need.
|
||
My plan is to first build the Kubernetes cluster. I want to use Terraform because:
|
||
I am quite familiar with Terraform, which makes that part of things a bit easier I can build out the rest of the deployment in Terraform once I figure out exactly how it all works At the end of the project, I want to have a Terraform configuration that can deploy a full Kubernetes cluster and the Windows container hosts in Azure. My first attempts at adding a Windows container host will be manual though, and once I have a good grasp on what needs to happen, I can automate with Terraform and maybe DSC.
|
||
Once I actually have the Windows container hosts, I am going to need to find a good service to deploy. I am open to suggestions of any Windows based container service examples that exist out there. It doesn’thave to be super flashy, and I can figure out the service and pod configuration aspects of it. I just need a starting point. So if you already have a docker swarm service for Windows containers that you think would work for this project, please comment below or DM me on Twitter.
|
||
For now, I am off to read some exciting technical documentation!
|
||
`,summary:"Well, it wasn’t even close. As mentioned in my previous post, I am moving to a less hands on role, and I want to keep close to the technology. The concept of running Windows container hosts in a Kubernetes cluster fascinates me and it appears that I wasn’t alone. With 82% of the votes on my Twitter poll, it was the clear winner. Now I guess I actually need to start diving in, and by diving in, I mean reading docs.",date:"24 Aug, 2017",url:"https://nedinthecloud.com/2017/08/24/windows-hosts-with-kubernetes-the-beginning/",image:"tutorials.png",readingTime:"2"},"https://nedinthecloud.com/2017/08/19/welcome-back-kotter/":{title:"Welcome Back Kotter",tags:[],content:`You may have noticed a lapse in posts for that last few months. There’s a few reasons for that:
|
||
It’s summer, relax No don’t relax, cause you are writing a course for Pluralsight on Terraform And you got a promotion, which is a blessing and a curse Plus you’re now chasing around a 12 month old who is trying to chase your other two So yeah it’s been a little crazy and honestly blogging kinda fell off the table. The good news is that:
|
||
My course is effectively complete (more on that later) And the kids are going back to school in two weeks Plus I’ve shed my previous responsibilities at work, so I am only doing one job again My new role at work is not especially technical. Well, that’s not entirely true; I still need to know lots of technical things, but at a 10,000 foot level not at a keyboard-banging-things-out level. If I want to stay connected to the tech - and I do - then I need to dedicate some time to rolling up my sleeves and digging in. This blog will be the repository of that information as I learn it.
|
||
So what should be my next area of focus? Hard to say since there are so many new technologies coming out, and so little time to learn them all. A short list might look like this:
|
||
Kubernetes & Windows Digital Rebar on HPE Ansible & Terraform IoT using Azure AWS Lambda, Python, and Cloud Watch Hashicorp Nomad, Vault, Packer, and Consul That’s just a few off the top of my head. And since tech never stands still, there will be a few more by the end of the year. I’d like to turn one of those into a blog series, so I ask you, the reader, which would you most like to see? I’ve posted a Twitter poll about exactly that here. Happy Voting!
|
||
`,summary:`You may have noticed a lapse in posts for that last few months. There’s a few reasons for that:
|
||
It’s summer, relax No don’t relax, cause you are writing a course for Pluralsight on Terraform And you got a promotion, which is a blessing and a curse Plus you’re now chasing around a 12 month old who is trying to chase your other two So yeah it’s been a little crazy and honestly blogging kinda fell off the table.`,date:"19 Aug, 2017",url:"https://nedinthecloud.com/2017/08/19/welcome-back-kotter/",image:"analysis.png",readingTime:"2"},"https://nedinthecloud.com/2017/04/12/adding-windows-server-2012-r2-image-to-azure-stack/":{title:"Adding Windows Server 2012 R2 Image to Azure Stack",tags:["azure","azure-stack","microsoft","powershell","scripting"],content:`In a previous post I covered how to add a Linux image to Azure Stack. In this post I am going to detail a simple (if slow) way of adding a Server 2012 R2 image to Azure Stack as well. With Azure Stack TP3 (original and extra-crispy) there are no default VM Images included in the install. You are prompted to download a Windows Server 2016 ISO as part of the Azure Stack POC download, and there is a script in the AzureStack.ComputeAdmin module called New-Server2016Image that will take that ISO and turn it into a Core or Datacenter image. But what if you wanted that good ole Server 2012 R2 image?
|
||
For whatever reason, VM Images are not yet part of the marketplace syndication. So you are on your own when it comes to creating them. So I figured, why not go to the source? I can spin up a Windows Server 2012 R2 Datacenter VM in Azure, sysprep it, and then copy the VHD down locally. Then using the Add-VMImage cmdlet in AzureStack.ComputeAdmin I’ll add it to the Azure Stack Install. That is essentially what the script below does.
|
||
This set of commands is meant to be run from the MAS-CON01 VM, and it assumes that you have installed the proper version of AzureRM PowerShell and the AzureStack tools from Github. You also need to have a valid subscription in Azure, obviously, in which to spin up the VM. The commands assume that you have already logged into Azure Cloud and selected the proper subscription. There are a few fields to change here and there and I have added <> on either side of the text to indicate a necessary change. The ARM template being used is sitting on my GitHub, but feel free to use whatever template you like. The key is to allow WinRM traffic from the public subnet so that the script can remote in and sysprep the server. These commands could easily be repurposed to capture custom VM images using other operating systems too, assuming they support WinRM. I am looking into turning this into a proper script that can take some arguments and do the whole thing for you, but that’s not quite ready yet. Minimum Viable Product right? Lastly, the image is a whopping 128GB in size, so be congizant of that when you select a destination location for the VHD. I picked the SU1Fileserver share as a target since it should have ample space, but your mileage may vary.
|
||
Occasionally the token for uploading the VHD and creating the Gallery Item will time out. In that case, you can manually add the image and create the gallery item. I’m working with the Azure Stack team to fix that particular issue, possibly checking for token status after the initial VHD upload and refreshing it if expired.
|
||
`,summary:"In a previous post I covered how to add a Linux image to Azure Stack. In this post I am going to detail a simple (if slow) way of adding a Server 2012 R2 image to Azure Stack as well. With Azure Stack TP3 (original and extra-crispy) there are no default VM Images included in the install. You are prompted to download a Windows Server 2016 ISO as part of the Azure Stack POC download, and there is a script in the AzureStack.",date:"12 Apr, 2017",url:"https://nedinthecloud.com/2017/04/12/adding-windows-server-2012-r2-image-to-azure-stack/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2017/03/28/adding-linux-images-to-azure-stack/":{title:"Adding Linux Images to Azure Stack",tags:["azure","azure-stack","azurestack","powershell"],content:`If you are planning to add Linux Images to your Azure Stack deployment, first I would recommend reading through the documentation on the Azure Stack pages for Adding a VM Image and Using Custom Linux Images. From there you can get the base images and the process for adding the images to Azure Stack. What they don’t include is the Azure Cloud information for the various images, and if you would like to be able to use a JSON template against both Azure and Azure Stack without changing the image information, then you will want the publisher, offer, sku, and version to match. In this post I will walk through the basics of adding one Linux image, how to get the necessary information from Azure Cloud, and the current information for the images you may want to run.
|
||
All of these commands will be run from the Mas-Con01 VM in Azure Stack. You can run it from your local workstation if you set up a P2S VPN tunnel, but that’s a topic for another post. I am also running this on Azure Stack TP3, so if you are on a previous version your results may vary. Finally, I bump the Mas-Con01 VM up to 8GB of memory and turn off Windows Defender. That’s a personal choice; however, I think you’ll find that it speeds up anything you are planning to run on Mas-Con01 tremendously.
|
||
Step 1 - Install the Azure Stack tools and Azure Stack PowerShell From an Administrative PowerShell session run the following:
|
||
Step 2 - Retrive the VM Image information for Azure Cloud For this example I will be finding the VM Image information for Ubuntu Server 16.04 LTS. The first step is to log into Azure via PowerShell using Login-AzureRMAccount. Then the following commands will retrieve the information you would want.
|
||
Once the commands are finished running, you should get output that looks roughly like this:
|
||
Step 3 - Download and Add Image Now you have all the necessary values to add the Ubuntu image to your local Azure Stack installation. The next thing to do is download the images from the Deploy Linux virtual machines on Azure Stack page I linked before. Keeping with the theme of Ubuntu, I am downloading the 16.04 LTS image linked on the page. Right now that link is broken, so you can use this one instead if Microsoft hasn’t fixed it yet. The file will be compressed, and so you’ll need to extract it. The VHD is ~30GB in size, so I recommend copying it to the infrastructure file share created during the Azure Stack install, which should be \\\\SU1fileserver\\SU1_Infrastructure_1. That way you won’t run out of room on Mas-Con01 as you download and add more images.
|
||
Once the VHD file has been extracted, make a note of the location and run the following to add the VM Image.
|
||
You’ll be prompted for your Azure Stack credentials, and you’ll need to change the tenant to match whatever you used when installing Azure Stack. In the end you will now have a new VM Image in the marketplace.
|
||
Here is a table with all the current images linked on the Microsoft page:
|
||
PublisherName
|
||
Offer
|
||
Sku
|
||
Version
|
||
Title
|
||
Location
|
||
OpenLogic
|
||
CentOS
|
||
6.8
|
||
6.8.20170105
|
||
CentOS-based 6.8
|
||
Local
|
||
OpenLogic
|
||
CentOS
|
||
7.2
|
||
7.2.20170105
|
||
CentOS-based 7.2
|
||
Local
|
||
CoreOS
|
||
CoreOS
|
||
Stable
|
||
1298.6.0
|
||
CoreOS Linux (Stable)
|
||
Local
|
||
SuSE
|
||
SLES
|
||
12-SP1
|
||
10.21SLES 12 SP1
|
||
Local
|
||
Canonical
|
||
UbuntuServer
|
||
16.04-LTS
|
||
16.04.201703270
|
||
Ubuntu Server 16.04 LTS
|
||
Local
|
||
Canonical
|
||
UbuntuServer
|
||
14.04.5-LTS
|
||
14.04.201703230
|
||
Ubuntu Server 14.04 LTS
|
||
Local
|
||
Note that the version numbers are subject to change, so you may want to verify what version you have downloaded. Regardless, the latest version is automatically selected if you specify “latest” in your ARM template.
|
||
A word about the Sku property. The current version of the Add-VMImage cmdlet in the AzureStack.ComputeAdmin module does not allow for a “.” in the Sku parameter. You can either avoid using a “.” or alter the validation criteria for the parameter by changing the expression from “[a-zA-Z0-9-]{3,}” to “[a-zA-Z0-9-\\.]{3,}” in the Add-VMImage function. Doing so does not seem to affect the deployment of new images.
|
||
`,summary:"If you are planning to add Linux Images to your Azure Stack deployment, first I would recommend reading through the documentation on the Azure Stack pages for Adding a VM Image and Using Custom Linux Images. From there you can get the base images and the process for adding the images to Azure Stack. What they don’t include is the Azure Cloud information for the various images, and if you would like to be able to use a JSON template against both Azure and Azure Stack without changing the image information, then you will want the publisher, offer, sku, and version to match.",date:"28 Mar, 2017",url:"https://nedinthecloud.com/2017/03/28/adding-linux-images-to-azure-stack/",image:"tutorials.png",readingTime:"4"},"https://nedinthecloud.com/2017/03/17/oh-the-places-youll-go/":{title:"Oh, The Places You'll Go",tags:[],content:`You’ve got brains in your head… I think it’s fair to say that most of us have followed an interesting path to end up in IT. As you continue to progress, you may start considering what the next steps are in that path. Here are some of the questions to consider. FYI - this post was inspired by an excellent post that I absolutely cannot find and really wished I had bookmarked.
|
||
Do you stay technical or not? You probably got into IT because you love technology. I know I do. At home I like to tinker with electronics and dabble in some programming. Being in a technical role keeps me engaged and interested, but it also requires constantly keeping up with the industry. And I think we can all agree that this industry changes at a breakneck pace. The kind of pace that can burn someone out. If you decide to stay technical, then prepare yourself for the Sisyphean task of staying up to date on your skillset.
|
||
Do you specialize or not? If you do choose to stay technical, then the additional question is whether it makes sense to specialize in a particular discipline. There are those in the public sphere that advocate a generalist approach. You’ll hear terms like a Full Stack Admin or DevOps Engineer, which implies that you should know all the various components of the application stack from the node.js page that delivers the content to the fibre channel protocols that deliver the bits. Being a good generalist means knowing a little about a lot of things, knowing how they interact together, and knowing who to ask when you have a question. Who do you ask? A specialist of course! The specialist goes deep on a particular technical discipline, and I don’t just mean storage or network or compute. A specialist might drill down into a specific subset of one of those, like the aforementioned fibre channel, or wireless networking, or KVM virtualization. The problem with being a specialist is that the thing that you choose to specialize in may become obsolete, and consequently so do you.
|
||
Do you go into management? Some people decide to leave the technical rat race and try their hand at management. The skillset required to be an effective manager is vastly different than what is needed to be a good consultant, but there are still parallels. Being able to communicate effectively, listen to what people are saying, and be responsive to people’s needs are all things that will serve you well as both a consultant and a manager. If you think that you would enjoy the challenge and satisfaction of helping your coworkers grow and improve, then management could be right for you.
|
||
Do you go into presales? If you love the technology, but are getting tired of being in front of the keyboard, then presales could be a good fit for you. A presales architect needs to have a generalist’s understanding of technology, and the skills to explain it to a prospective client. Being in presales is about understanding what a client needs, and being able to design a solution based on those needs. If you enjoy presenting to people, talking technology, and designing solutions, then presales could be a good fit. You also need to enjoy, or at least be comfortable with the business side of things too. Presales needs to contend with margin, cost, budgets, deal registration, and a host of other non-technical items that might be unappealing to a technical person.
|
||
Do you stay put? This is an option that people don’t often talk about, but it’s perfectly legitimate. There is no reason your career path needs to be full of twists and turns, constant elevation, or veering angles. For the time being, it may make perfect sense to just keep heading straight down the path, enjoying the work that you do without trying to move up the ladder or spec out the next opportunity. If you like what you’re doing, and you are comfortable doing it, then I don’t see any reason to force yourself to change. Then again, I have mentioned something about making yourself uncomfortable before.
|
||
The road stands before you, and whatever you choose, so long as you’re happy, there’s no way you can lose.
|
||
`,summary:"You’ve got brains in your head… I think it’s fair to say that most of us have followed an interesting path to end up in IT. As you continue to progress, you may start considering what the next steps are in that path. Here are some of the questions to consider. FYI - this post was inspired by an excellent post that I absolutely cannot find and really wished I had bookmarked.",date:"17 Mar, 2017",url:"https://nedinthecloud.com/2017/03/17/oh-the-places-youll-go/",image:"featured-image-nedinthecloud.jpg",readingTime:"4"},"https://nedinthecloud.com/2017/03/05/cicd-pipeline-with-azure-stack-part-4/":{title:"CICD Pipeline with Azure Stack – Part 4",tags:["azure","azure-resource-manager","azure-stack","azurestack","microsoft","powershell","scripting","tfs"],content:`This is part 4 in a series. You can find part 1 here , part 2 here, and part 3 here.
|
||
First off, let me quell your anticipation. I got it working! It was not as straightforward as I might like, but it will work. If you haven’t already read post three, I would recommend doing so. The long and short of it is that the build task Azure Resource Group deployment in TFS doesn’t understand the Azure Stack environment. It doesn’t know how to talk to it, so any build task is going to fail. One of the engineers at Microsoft suggested I use a PowerShell task to deploy instead, which I did. That was not as simple as I would have liked, but here is what I had to do.
|
||
First the server running the build agent, in my case the TFS server, will need to have the Microsoft Azure PowerShell for Azure Stack TP2 - September 2016 installed. If you have a newer version installed, you will receive API version errors. Next, I created a very basic template for deployment in Visual Studio. I literally just used the single Windows VM Azure quickstart template from GitHub. When you create the solution in Visual Studio, it also creates a Deploy-AzureResourceGroup.ps1 PowerShell script, which can be used to deploy the template from Visual Studio directly. I created a copy of the script called Deploy-AzureResourceGroupTFS.ps1. In that script I added the following parameters: [string] $AzureStackUser,
|
||
[string] $AzureStackPwd,
|
||
[string] $AzureStackTenant,
|
||
[string] $AzureStackARMEndpoint = '[https://api.azurestack.local](https://api.azurestack.local)',
|
||
[string] $AzureStackName = "AzureStack",
|
||
[string] $AzureStackSubscriptionName = "Default Provider Subscription" Yes, I know there’s a password field. In order to store the string securely, I am using the variables in the Build configuration to store the password. I have also included the AzureStack-Tools\\Connect\\AzureStack.Connect.psm1 PowerShell module from this repo on GitHub. That module has a function I want to use to add the AzureStack environment from within the script. Now in the script I have added the following lines:
|
||
$AzureStackSecPwd = ConvertTo-SecureString $AzureStackPwd -AsPlainText -Force $AzureStackCredentials = New-Object System.Management.Automation.PSCredential($AzureStackUser,$AzureStackSecPwd) $ConnectModulePath = [System.IO.Path]::GetFullPath([System.IO.Path]::Combine($PSScriptRoot, "AzureStack.Connect.psm1")) Import-Module $ConnectModulePath
|
||
Write-Verbose "Add Azure Stack Environment" Add-AzureStackAzureRmEnvironment -AadTenant $AzureStackTenant -ArmEndpoint $AzureStackARMEndpoint -Name $AzureStackName
|
||
Write-Verbose "Add Azure Stack Account" Add-AzureRmAccount -EnvironmentName $AzureStackName -TenantId $AzureStackTenant -Credential $AzureStackCredentials
|
||
Write-Verbose "Login Azure Stack" Login-AzureRmAccount -EnvironmentName $AzureStackName -Credential $AzureStackCredentials
|
||
Write-Verbose "Select Azure Stack Sub" Select-AzureRmSubscription -SubscriptionName $AzureStackSubscriptionName
|
||
Note that I am using the newest version of the AzureStack-Tools. The Add-AzureStackRmEnvironment parameters have changed, adding ArmEndpoint and removing Domain. The PowerShell commands above are the process by which an Azure Stack environment is added and referenced by regular Azure PowerShell commands. This is the same process that is used by most of the PowerShell scripts I have seen when they are trying to perform work in Azure Stack, unless they are strictly using REST calls with an authentication header.
|
||
Now my deployment script has a connection to Azure Stack and it can start the deployment process. I check my project into TFS and create a Build sequence that includes a PowerShell Task.
|
||
In the Build definition, I will define some of the input parameters using the Variables tab. Specifically, I am defining the AzureStackUser, AzureStackPwd, and AzureStackTenant information.
|
||
Within the PowerShell task, the variables are referenced by adding a dollar sign and parentheses, like this $(AzureStackUser).
|
||
This allows the AzureStackPwd to stay encrypted on the Build server, rather than storing it in plain text. The parameters for the ARM template reside in the same project, and currently the administrator password for the VM I am creating is stored in plain text. I plan to refactor that to use a Key Vault secret. However, since the vault may be in Azure Cloud or Azure Stack, I need to alter the template to decide which key vault to access.
|
||
After running the Build definition, I can see that the template deploys successfully to my Azure Stack environment. Hooray!
|
||
Of course it failed more than a few times, but that’s the nature of running builds.
|
||
Now I will add a second task to the Build definition to deploy an Azure Resource Group deployment in Azure Cloud.
|
||
I will use the same templates and parameters.
|
||
If the first build task is successful in Azure Stack, then this task will begin and run in Azure. In a real world situation, I might have some tests that run after the first build task to verify the environment, then release it to Azure Cloud. My Azure Stack install might be a development environment and Azure Cloud is for production. I might also have slightly different parameters for the production deployment, like using larger VMs or a bigger VM Scale Set.
|
||
After adding the second task, I run the build process again. Voila, it is successful for both tasks, and now I have an identical build running in both Azure Stack and Azure Cloud.
|
||
Next I go into the properties of the Build definition and tell it to fire off a build each time I check in code.
|
||
Then I go back into the template and change the VM size from StandardA1 to Standard_A2. I commit the change and check it into source control. Sure enough, a build process kicks off and both my Azure Stack and my Azure Cloud VMs are brought up as Standard_A2 size VMs. I actually had to make a minor change in the Deploy-AzureResourceGroupTFS.ps1 script to get the continuous integration piece working with Azure Stack. For the New-AzureRmResourceDeployment command I added the Mode parameter and passed the value _Complete. I tried Incremental, but for whatever reason Azure Stack didn’t like that. It will probably be fixed with the next release. Azure Cloud was fine with the implied Incremental setting.
|
||
Is this a perfect CICD process? No, not by a long shot. I want to write tests to validate that the Azure Stack environment is running properly, and I can do that by adding a PowerShell task to run after the Azure Stack build. What I would like to do it try and leverage the Pester testing framework to validate my infrastructure. The inestimable Adam Bertram has some posts on this very topic, so I plan to read through those and come back to my pipeline to try adding tests.
|
||
Also, in case you didn’t hear, Azure Stack TP3 is out! All this development has been on Azure Stack TP2. I now have TP3 running on a separate server in the lab, and I plan to recreate the environment and see what’s different about TP3. Expect some posts on that topic very soon!
|
||
I’ve checked the whole project into my GitHub repo. You can check out the templates and deployment scripts there.
|
||
`,summary:`This is part 4 in a series. You can find part 1 here , part 2 here, and part 3 here.
|
||
First off, let me quell your anticipation. I got it working! It was not as straightforward as I might like, but it will work. If you haven’t already read post three, I would recommend doing so. The long and short of it is that the build task Azure Resource Group deployment in TFS doesn’t understand the Azure Stack environment.`,date:"5 Mar, 2017",url:"https://nedinthecloud.com/2017/03/05/cicd-pipeline-with-azure-stack-part-4/",image:"tutorials.png",readingTime:"6"},"https://nedinthecloud.com/2017/03/03/everything-is-broken...still/":{title:"Everything is Broken...Still",tags:["aws"],content:`Earlier this week Amazon Web Service’s Simple Storage Service, better known as S3, was experiencing higher than normal error rates in the us-east-1 region, i.e. S3 was down. To put that into perspective, it also means that several high-profile websites and applications were experiencing major issues. As you may know, S3 has a 99.9% SLA uptime, and it’s been down for a couple hours now. I’m no math genius, but that’s more than the 44 minutes a 99.9% uptime per month requires.
|
||
I don’t want to beat up on AWS though, and that’s not the intention of this article. But it is an excellent example of the fact that everything breaks. Everything. And as technologists we know this and must plan for the worst. A well-architected solution should be able to withstand outages anywhere in its infrastructure and keep running. We talk a lot about not leaving a single point of failure, and that includes a data center or a cloud service in a specific region. Not all of AWS S3 was down, just the us-east-1 region. There are still four other regions in North America alone that had a functioning S3 service. A well architected solution should be able to seamlessly failover to one of those other instances without interrupting the end user. Sadly, it appears that many services have not chosen to build their solution in that way.
|
||
According to TechCrunch about 0.8% of the top 1 million websites use S3 for storage, which is about 80,000 sites. Not an insignificant amount. There are also other services that appear to be impacted, such as IoT devices and the Amazon Fire TV. Even the AWS health dashboard was not functioning properly at first, showing an all green when things were clearly not. The irony here is that even AWS and Amazon didn’t follow their own best practice and make sure they were designing a fully resilient architecture. Turns out this stuff is hard, which is why we need to know how to do it and help our customers do the same.
|
||
Eventually AWS will recover from this snafu. Turns out that the whole issue was caused by a mistyped command, which is frankly amazing and not surprising at the same time. Some websites are likely to decry the cloud, and curse its name to the heavens. Cooler heads shall prevail, and in a week the churn cycle of the internet will make this barely a blip on the radar. For the rest of us, I think there are some lessons to be learned here:
|
||
Design your solution such that there are no single points of failure Make changes easy to roll-back Plan for outages, and test those outages in Production Try not to rely on a single solution or vendor for key components of your solution When I said plan for the worst, that’s a little inaccurate. We have to be prepared for the worst that we can reasonably afford to mitigate. When I am architecting a solution, sometimes I will start going down the rabbit-hole of “what if” until I’ve over-architected the hell out of something. Just as important as understanding the worst-case scenario, is understanding what a worst-case scenario will cost you and spending a commensurate amount of money to prevent it. If your Project Management website being down for a day costs you $1,000 in productivity, and you could prevent it for an extra $50 a day, then it’s probably worth it. If it’s an extra $1,000 a day, then it probably is not worth it. You must look at the likelihood of a failure, the cost of the failure, and the cost of preventing it. Run the numbers and you’ll know how much work to put into prevention versus setting a reasonable SLA.
|
||
The outage will cause AWS to credit 10% to all S3 customers impacted by the outages. Another hour or two and that would be 25%. If AWS is raking in $10 million from S3, they are about to cut a check for $1 million in refunds. That hurts, but probably doesn’t come close to the financial impact on those sites that were down.
|
||
`,summary:"Earlier this week Amazon Web Service’s Simple Storage Service, better known as S3, was experiencing higher than normal error rates in the us-east-1 region, i.e. S3 was down. To put that into perspective, it also means that several high-profile websites and applications were experiencing major issues. As you may know, S3 has a 99.9% SLA uptime, and it’s been down for a couple hours now. I’m no math genius, but that’s more than the 44 minutes a 99.",date:"3 Mar, 2017",url:"https://nedinthecloud.com/2017/03/03/everything-is-broken...still/",image:"featured-image-nedinthecloud.jpg",readingTime:"4"},"https://nedinthecloud.com/2017/02/17/cicd-pipeline-with-azure-stack-part-3/":{title:"CICD Pipeline with Azure Stack – Part 3",tags:["azure","azurestack","cicd","powershell","tfs"],content:`This is part 3 in a series. You can find part 1 here and part 2 here. Part 4 is now available here.
|
||
As I mentioned in my previous post, I was “ready” to deploy my TFS deploy template to Azure Stack. And as predicted, the universe laughed at my funny plans. The deployment failed due to a required Windows Update on the target image. I didn’t run into this on Azure b/c the Windows Server 2012R2 image on Azure is more up to date than the one that ships with Azure Stack. At this point I could have just installed TFS and Visual Studio manually, but no I refuse to give up my dreams of an automated future. I spent the next week creating a PowerShell script that will install all available, required Windows Updates, and then reboot and repeat until there are no updates left. Then I ran that script against a Windows Server 2012R2 VM in Azure Stack, and used that updated VM to create an updated VM Image. You can read all about that adventure here. Let’s just say that the yak is well and truly shorn.
|
||
Once I had the updated image, my TFS template deployed successfully, although the Output windows from Visual Studio claims otherwise.
|
||
What actually happened is that the access token that Visual Studio was using for the deployment expired during the deployment. The actual template was successfully deployed and my script ran without issue. When it was all done I had a A2 instance running TFS, Visual Studio, and a Build Agent as shown below.
|
||
Now I can finally get on with setting up my build pipeline. The first step is to create a new project. In the Control panel I select the DefaultCollection and View the collection administration page.
|
||
In the Overview tab I create a New team project.
|
||
The new project is my Azure Stack CICD Pipeline, and I am choosing to use Agile and Git for the process template and version control. So far I like Git over Team Foundation version control.
|
||
Gears turn, cranks whir, and a bell rings *ding*
|
||
Clicking to go to the project, I end up on the project Overview page, where an immaculately hip developer greets me with indifference and a cell phone.
|
||
I am going to create a build definition, so I click on the BUILD tab (all caps for some reason?) and select the All build definitions area.
|
||
I create a new build definition by clicking the green + sign. Since we are deploying an Azure Stack build, I select an Empty build definition.
|
||
For the build definition, I accept the defaults and click on Create.
|
||
Now I am going to add the task to deploy an ARM template to a Resource Group, so I select the Deploy task type and add the Azure Resource Group Deployment.
|
||
The task is added to the build definition, but I haven’t added any Azure RM subscriptions. I click on the Manage link to add one.
|
||
That takes me to the Services tab for the Project, and I choose to add a New Service Endpoint of the Azure Resource Manager Subscription type.
|
||
The service endpoint is going to need for me to create a Service Principal in the Azure AD environment with permissions to Contribute to the subscription. There is a nice PowerShell script out on Microsoft, along with documentation for the process here. The output of the script gives me all the info I need to fill out this form.
|
||
Since I plan to deploy to Azure and Azure Stack, I repeat the process for Azure. Now I have two Service Endpoints.
|
||
Back in the build task configuration, I select AzureStack from the Azure RM Subscription drop-down and… ah crap.
|
||
I guess I should have seen this coming. TFS has no idea how to find the local Azure Stack environment. I spent some time digging into files on the TFS server, and here’s what I found so far. There is a VSIX file in the path %InstallPath%\\Microsoft Team Foundation Server 14.0\\Tools\\Deploy\\TfsServicingFiles\\Extensions called Microsoft.TeamFoundation.Extension.AzureRM.vsix. That is probably what is actually doing the work. A VSIX file is really just an archive, so by renaming it with a zip extension I was able to open the archive. Inside there are two extension manifest files. The VSO manifest file is a JSON file with the Azure RM URL defined as https://management.core.windows.net/ along with a Computer and Storage endpoint URL with the same prefix. The Azure Stack API URL is https://api.azurestack.local, so it’s no wonder that the extension could not find it.
|
||
There’s more to it than that though. There are JavaScript files for the TFS controls, a PowerShell module for the deployment task, and a web.config file that does something.
|
||
I have to say this is all above my paygrade, of approximately $0. I could try to write a new extension and service endpoint for TFS to handle deployments to Azure Stack, but I don’t think I’m quite prepared to conquer that. Instead, I am going to reach out to the Azure Stack team at Microsoft and see if they already have something in the works for TFS that I can use. I’ll post an update once I know more about what my next step looks like.
|
||
Update: After conferring with the CICD Azure Stack team at Microsoft, they suggested that I try using a PowerShell script task in TFS instead of the Azure Resource Group task. They will eventually re-tool the native TFS task to support Azure Stack, but there is no hard date on when that will happen. I can’t believe that I didn’t think of this myself; I guess I was too tied up in trying to get the native tool to work. I’m going to get to work on using the PowerShell task, and be back with post 4 shortly!
|
||
`,summary:`This is part 3 in a series. You can find part 1 here and part 2 here. Part 4 is now available here.
|
||
As I mentioned in my previous post, I was “ready” to deploy my TFS deploy template to Azure Stack. And as predicted, the universe laughed at my funny plans. The deployment failed due to a required Windows Update on the target image. I didn’t run into this on Azure b/c the Windows Server 2012R2 image on Azure is more up to date than the one that ships with Azure Stack.`,date:"17 Feb, 2017",url:"https://nedinthecloud.com/2017/02/17/cicd-pipeline-with-azure-stack-part-3/",image:"tutorials.png",readingTime:"5"},"https://nedinthecloud.com/2017/02/14/dear-future-hci-partner/":{title:"Dear Future HCI Partner",tags:["hci","hyperconverged","marketing"],content:`If there’s one thing I wish HyperConverged Infrastructure (HCI) vendors would stop doing, it’s promising that the product will be up and running “in a matter of minutes”. First of all, it’s simply untrue. Second, it’s irresponsible and sets those of us deploying the hardware up for failure. When skewed perceptions intersect meaty reality, the deployment engineer is the first to be skewered. And you know what else?
|
||
I’m not speaking in a theoretical sense, but rather as someone who has deployed HCI solutions from four different vendors in the last six months. I’m not going to name names here, but the issues are so consistent amongst the four horsemen of my personal HCI apocalypse that I feel confident extending this to all solutions currently on the market.
|
||
First of all, the technical presales and their marketecture has got to go. Yes it is technically feasible to deploy your solution in 15 minutes, or rather it is possible to run the wizard in about 15 minutes. But that requires that:
|
||
The network is configured correctly based on whatever strange requirements your solution has (native VLAN, IPv6, IGMP, BGP with ISIS over PVXLAN) The technician has read your 600 page document, and the 75 page upgrade guide The pre-deployment worksheet has been filled out, reviewed, updated, rejected, and resubmitted You haven’t accidentally used a naming prefix that breaks the entire deployment and requires a full re-image (True Story!) The equipment has been unboxed, racked, cabled, powered, and labeled The 25GB of software updates have been downloaded It’s not a full moon or a day ending in ‘y’ Then SURE, YES, YOUR WIZARD is AMAZING!
|
||
But it’s not. The wizard is part of a UI that is explicitly designed to hide the complexity of the solution. The problem is not the complexity. I’m fine with a solution being complex, in fact I expect it. Complexity is okay, hiding complexity badly is not. The wizard and management GUI in most of these solutions is the front end to scripts that often break, fail, or give confusing feedback. HCI is still in infancy, and I know that making complexity simple is actually very difficult. If you want to see it done well, go log onto Azure or AWS. They take an incredibly complex set of technologies and wrap a streamlined GUI around it. It is in their best interest to make the consumption of their services as easy as possible, so they have every incentive to lower the barrier of entry on their platform. If you want to use the full power of the platform, you can of course use the CLI or the APIs, but you don’t have to. Azure, AWS, and other cloud platforms have had over a decade to streamline their interface and make sure it is intuitive and resilient. HCI, well not so much.
|
||
If HCI solutions were truly so simple deploy that anyone could do it, they wouldn’t require a 600 page installation manual. If the solution can truly be front-ended by a simple wizard and management GUI, then a manual that fells a small forest every time it is printed would not be necessary. If you try to hide complexity, and do it badly, you are simply inviting poor deployments and disillusioned customers.
|
||
And don’t tell me about your 200 node deployment that went flawlessly at client X. That deployment was designed and executed by a senior engineer from your company who probably wrote half the interface and scripts. Of course someone with intimate knowledge of the platform had no trouble deploying their own software. You know what a great test of software is? Hand it over to someone else and have them run it without reading the source code. They don’t know what assumptions you’ve made, or what you expect the user to do. They can and will find bugs. You know who is doing that testing? ME.
|
||
Also, let’s not talk about POCs. A Proof of Concept is usually executed in a standalone environment that doesn’t mirror actual Production. Most POCs are successful, because they only confirm that the solution works as already tested by the engineers at the vendor. Look, I know your technology works, mostly, and that’s not what I am annoyed with. The tech is not the problem, the marketing IS. Remember, every time you run a successful POC, a baby Arctic seal is clubbed by Katy Perry. And that makes her cry. You don’t want to make Katy Perry cry do you? You monster.
|
||
A few weeks ago I sat through yet another HCI presentation by a vendor wherein they started making claims about how easy it was to deploy their solution, and how it offered performance and features unparalleled by the competition. Naturally, my colleague and I expressed some cynicism about the marketing hype. The solution engineer tried to tell us that it wasn’t marketing hype, and things started to get a little heated. I told him that we hear a lot of great promises, and when the marketing meets the road things begin to fall apart. I’m not going to drink the kool-aid just cause you have some PowerPoint slides and a true believer on your side. The thing is, I know this is marketecture BS. The clients you are hocking your wares to don’t.
|
||
I’m sure your solution is a unique, precious snowflake, just like every other HCI solution put in front of me. And I’m sure it works, well probably, I wouldn’t know since I am still reading through the 350 page administration manual for a solution you describe as dead-simple to deploy in a matter of minutes.
|
||
So please just stop it. STOP.
|
||
`,summary:`If there’s one thing I wish HyperConverged Infrastructure (HCI) vendors would stop doing, it’s promising that the product will be up and running “in a matter of minutes”. First of all, it’s simply untrue. Second, it’s irresponsible and sets those of us deploying the hardware up for failure. When skewed perceptions intersect meaty reality, the deployment engineer is the first to be skewered. And you know what else?
|
||
I’m not speaking in a theoretical sense, but rather as someone who has deployed HCI solutions from four different vendors in the last six months.`,date:"14 Feb, 2017",url:"https://nedinthecloud.com/2017/02/14/dear-future-hci-partner/",image:"featured-image-nedinthecloud.jpg",readingTime:"5"},"https://nedinthecloud.com/2017/02/07/updating-the-azure-stack-2012r2-image/":{title:"Updating the Azure Stack 2012R2 Image",tags:["azure-stack","azurestack","microsoft","powershell","scripting"],content:`If you’ve started playing with Azure Stack, you might notice that the Windows Server 2012R2 image is a little behind on its Windows patches. Before you do any heavy duty testing, you’re going to want to update the image with the latest patches. This is a multi-step process:
|
||
Deploy an image to update Install all available Windows Updates (I’ve got a script for that!) Sysprep the machine to be a new image Locate the VHD file Update the image using the portal or PowerShell I’m not going to walk you through deploying a VM in Azure Stack, but I will recommend that you use the A2 size. Installing the update should go a little faster on a system with more horsepower.
|
||
Using PowerShell to run Windows Update You could manually update the VM with all the Windows Updates, but why do that when there’s PowerShell? I’m making use of the Windows Update PowerShell module available on the TechNet gallery. All you have to do is copy this script from my Gist to the target VM. Then run it. The script will download and import the module, install the available Windows Updates, and then create a scheduled task to run again on startup. It should keep running until there are no updates left. Fire away and come back in a few hours depending on your internet connection. It took an A2 VM about four hours to patch when I last ran this. Glad I didn’t have to babysit it!
|
||
Now that your VM is properly patched up, it needs to be prepared for use as an image. Fortunately, all the necessary settings and VM agent are already installed. From an administrative command prompt run sysprep:
|
||
The VM will shutdown when sysprep is complete. Make sure that you go into the portal and stop it from there, so it is deallocated properly.
|
||
The VHD location for the VM will vary depending on the storage account you used. From within the portal, go to the VM’s Disks
|
||
Select the OS disk and then copy the blob URI by clicking on the neat little clipboard icon. Paste that value into notepad or something similar.
|
||
Go into the storage account that was used to store the VHD. The blob properties of the VHD need to allow anonymous access. Select the Blob portion of the storage account, and then select the container which houses the vhd (usually vhds). Change the Access policy to Access type to Blob.
|
||
Now we’re going to add a new version of the Windows 2012R2 image. In the portal select Resource Providers
|
||
Select the Compute RP and then click on the VM Images on the far right
|
||
Click on the 2012-R2-Datacenter image and copy all the values to notepad
|
||
Now click on the Add button and use the previous values to fill out the fields. Be sure to increment the Version number in. In my case I went from 1.0.0 to 1.1.0.
|
||
Now click the Create button and wait. Once the creation process is complete, you will have a fully patched Windows Server 2012R2 image to use for your Azure Stack deployments. My creation time was about an hour, so don’t be surprised when it doesn’t create immediately.
|
||
You might wonder what happens with the existing Gallery Item that was using the version 1.0.0 template. Good question! The templates for the marketplace are unsurprisingly stored in a storage account in the System.Gallery Resource Group. If you dig down into the storage account you will find the blob container with the marketplace item here: dev20151001-microsoft-windowsazure-gallery/MicrosoftWindowsServer.WindowsServer-2012-R2-Datacenter.1.0.0
|
||
The template that controls deployment is called CreateUIDefinition.json. And that file doesn’t reference an actual version of the template. So despite the fact that the Gallery Item description claims that it uses the 1.0.0 version, it should use the latest version (1.1.0 in my case). I created a new VM to test, and as you can see, no Windows Updates were available.
|
||
PS - You can also add an image using PowerShell. If you’d like to know more about the process, then check here and use the same values you would have in the portal.
|
||
Here’s the full PowerShell script if you’re interested:
|
||
`,summary:`If you’ve started playing with Azure Stack, you might notice that the Windows Server 2012R2 image is a little behind on its Windows patches. Before you do any heavy duty testing, you’re going to want to update the image with the latest patches. This is a multi-step process:
|
||
Deploy an image to update Install all available Windows Updates (I’ve got a script for that!) Sysprep the machine to be a new image Locate the VHD file Update the image using the portal or PowerShell I’m not going to walk you through deploying a VM in Azure Stack, but I will recommend that you use the A2 size.`,date:"7 Feb, 2017",url:"https://nedinthecloud.com/2017/02/07/updating-the-azure-stack-2012r2-image/",image:"tutorials.png",readingTime:"4"},"https://nedinthecloud.com/2017/01/23/cicd-pipeline-with-azure-stack-part-2/":{title:"CICD Pipeline with Azure Stack – Part 2",tags:["azure","azurestack","cicd","powershell","tfs"],content:`This is part 2 of an ongoing series on building a CICD pipeline in Azure Stack. You can find part 1 here, part 3 here, and part 4 here.
|
||
When I last left things, I had successfully installed TFS on a virtual machine in Azure. And I wrote the template in such a way that it could be deployed to Azure Stack as well. After completing that process, I started working through deploying an ARM template through TFS using an automated build process. It turns out that the server running the build agent needs to have Visual Studio installed in order to deploy resources to Azure. I have since updated my ARM template and PowerShell script to automate the installation of Visual Studio Community 2015 and the TFS build agent. I also updated the template to take two new parameters: FileContainerURL and FileContainerSASToken. The former points to the blob container that holds the necessary installation files. The latter passes a SAS Token for read and list access to the blob container.
|
||
If you want to play along you will need to create a SAS Token thusly:
|
||
Login-AzureRmAccount
|
||
Get-AzureSubscription
|
||
Select-AzureSubscription -SubscriptionName _"YourSubscriptionName"_
|
||
$ctx = New-AzureStorageContext -_StorageAccountName yourstorageaccountname_ -StorageAccountKey "_YourStorageAccountKey_"
|
||
New-AzureStorageContainerSASToken -Name _"ContainerName_" -Permission rl -FullUri -Context $ctx -ExpryTime (get-date).AddDays(30)
|
||
That will return the FullURI for the token.
|
||
Here’s an example output:
|
||
https://nbinstallers.blob.core.windows.net/tfs?sv=2015-04-05&sr=c&sig=mVX%2FMu%2FcJRD%FIE3GgqOFAsB60dvWKYFcVTC0g7dNWvc%3D&se=2017-01-18T03%3A06%3A53Z&sp=rl
|
||
There are a number of fields in the URI, which you can read about in more detail here. The se field defines the expiration date and time, which in the case of the above token has already expired. The Full URI returned by the command can be used to list the contents of the container. If you want to grab a specific item from the container, then you insert the file name into the path and include the token afterwards like this:
|
||
“https://nbinstallers.blob.core.windows.net/tfs/AdminDeployment.xml?sv=2015-04-05&sr=c&sig=mVX%2FMu%2FcJRD%2FIE3GgqOFAsB60dvWKYFcVTC0g7dNWvc%3D&se=2017-01-18T03%3A06%3A53Z&sp=rl”
|
||
I am grabbing the AdminDeployment.xml file from the tfs container using the SAS token. The SAS token has special characters in it, so when it is passed to the PowerShell script in the ARM template, double quotes must be added like this:
|
||
"commandToExecute": "[concat('powershell -ExecutionPolicy Unrestricted -File ', variables('InstallTFSScriptScriptFolder'), '/', variables('InstallTFSScriptScriptFileName'),' -FileContainerURL ', parameters('FileContainerURL'),' -FileContainerSASToken \\"', parameters('FileContainerSASToken'),'\\"')]
|
||
I used \\" to escape the double-quote in the concatenated string.
|
||
I am currently using the Visual Studio Community Edition DVD to install VSC 2015, which weighs in at a heavy 7GB. Pulling that down to install VSC is time consuming and inefficient. I’d like to come back and revise the install process to use the Web Platform Installer, which has a command line component. That way I can install the light-weight WPI, and stream down the components of VSC that I need. If this becomes a presentation, or something others are truly interested in, then I will make that a priority. In the meantime I consider it a nice to have, but I’d like to actually move on to deploying IaC rather than tweaking the deployment platform. Speaking of which…
|
||
Since the universe laughs at our grand designs, it was inevitable that my Azure Stack box lab would be broken when I needed it for the next phase of the project. When I logged in, only the MAS-DC01 VM was running. I soon discovered that the Cluster Shared Volume had gone offline in Failover Cluster Manager due to failed writes to the disk. Since there is no redundancy in the Azure Stack POC (it’s a RAID 0 striped set), one drive having issues is all it takes to bring down the pool. In a real life deployment, you would have multiple nodes using Storage Spaces Direct (S2D) to replicate the volumes locally and across nodes.
|
||
Of course my hardware wasn’t reporting any issues with the drive, so I thought it might have been a fluke. I tried reinstalling the Azure Stack POC on the same hardware. I got as far as the deployment of the MAS-BGPNAT01 box and it failed. Looking in the Failover Cluster Manager, I could see that the CSV was offline again. The drives all looked okay individually, but the cache battery was dead on the RAID controller. I moved the drives to another lab server and tried deploying again. Sadly the same result. I can only conclude that one of the drives is bad, but not reporting as such.
|
||
In the meantime, I started installing the Azure Stack POC on the three nodes that make up the HPE 250-HC cluster. Each node has 256GB RAM, 2x 12-core processors, 4x 1.2TB HDDs, and 2x 400GB SSDs. More than enough horsepower, and I figured at least one of them would install properly. They all ended up installing, and now I have three Azure Stack boxes to share with the crew at work. As I mentioned in a post:
|
||
When life gives you lemons, install the #AzureStack POC on all three nodes of an #HPE 250-HC.
|
||
— Ned Bellavance (@Ned1313) January 17, 2017 So that’s where I am at now. Azure Stack is working again. My template is ready to deploy. And I have tested the build and deploy process using Azure. Now it’s time to deploy in Azure Stack and see what happens. In my next post I will show how to set up a project in Visual Studio using the deployed TFS server and how to create an automated build process to deploy the project to Azure Stack. At least, that’s what I plan to do. I can hear the universe chuckling as I type.
|
||
`,summary:`This is part 2 of an ongoing series on building a CICD pipeline in Azure Stack. You can find part 1 here, part 3 here, and part 4 here.
|
||
When I last left things, I had successfully installed TFS on a virtual machine in Azure. And I wrote the template in such a way that it could be deployed to Azure Stack as well. After completing that process, I started working through deploying an ARM template through TFS using an automated build process.`,date:"23 Jan, 2017",url:"https://nedinthecloud.com/2017/01/23/cicd-pipeline-with-azure-stack-part-2/",image:"tutorials.png",readingTime:"5"},"https://nedinthecloud.com/2017/01/14/cicd-pipeline-with-azure-stack-part-1/":{title:"CICD Pipeline with Azure Stack - Part 1",tags:["azure","azurestack","cicd","powershell","tfs"],content:`This is the first post in a series of getting a CICD pipeline working with Azure Stack. You can read part 2 here, part 3 here, and part 4 here.. I will add links to additional posts as they are created.
|
||
There are a few things that have been coming up a lot lately at work that I would like to dive into some more to get a better understanding. The first is Infrastructure as Code (IaC). I’ve started doing my fair share of this in both Azure and AWS, but I feel like I’m just starting to truly get my head around the best practices and patterns to use when deploying IaC. The next trend is a move towards continuous integration and continuous deployment, CICD. How can I take the principles of a CICD pipeline and apply them to the IaC work I’ve been doing? Finally, there is the hybrid cloud element that is coming with Azure Stack. Unless you’ve been living under a rock, you’ve probably heard about Azure Stack. If not, here are a couple resources to get you started. I wanted to put all of those items together and build out a project that uses them.
|
||
The goals of my project are straightforward.
|
||
Set up a CICD pipeline solution that can operate against both Azure and Azure Stack Create an IaC deployment that can be deployed on both Azure and Azure Stack Include built in unit tests and functional test for deployment of IaC Document the process and pitfalls The first step is to pick a code repository and build server that will work with Azure Stack. I’ve used Visual Studio Team Services to deploy from Visual Studio to Azure, so it seemed like a good fit for Azure Stack. However, the Azure Stack POC doesn’t have a publicly exposed front-end that VSTS could interact with, so I need to go with something on premise. I’ve decided on Team Foundation Server 2015 deployed on a VM in Azure Stack. TFS will be both my build sever and my code repository, Since I’m deploying to Azure Stack, I figured why not use an ARM template to automate as much of the deployment as possible? I’ve done just that by first developing the template against Azure and then adding Azure Stack compatibility based off the Azure Stack quickstart templates. I have created the template in such a way that it can be deployed to either environment.
|
||
Here is the template I developed. It deploys a single Windows 2012R2 VM in Standard_A2 size with a 100GB data disk. I expose RDP and TCP port 8080 for the TFS website. The custom script extension calls a PowerShell script that prepares the data disk, then it installs TFS 2015, Visual Studio Community 2015, and a Build agent. When the deployment is complete, I have a working TFS 2015 server with the ability to create a project that deploys to Azure.
|
||
Within the template I have created a parameter called DeployLocation as such:
|
||
"DeployLocation": { "type": "string", "defaultValue": "Azure", "allowedValues": [ "Azure", "AzureStack" ], "metadata": { "description": "Deploy to Azure or AzureStack" } }
|
||
The idea is that setting this parameter will automatically change components of the template to match what is available in Azure Stack. For instance, there is only Standard_LRS storage in Azure Stack, so I can lock the storage account type down to that. How do I do this? In the variables section I define two variables:
|
||
"DeployAzure": { "apiVersion": "2015-06-15" }, "DeployAzureStack": { "apiVersion": "2015-05-01-preview" }
|
||
Each variable defines a set of values to be used in the rest of the template. In this case I have only defined which API setting to use, but I can also define the VM sizes, storage account type, and location. Now I need to combine the DeployLocation parameter value with the variables above. That is done by defining a third variable.
|
||
"DeployLocation": "[variables(concat('Deploy',parameters('DeployLocation')))]"
|
||
This sets the DeployLocation variable to a concatenated value combining the string Deploy and the DeployLocation value to form either DeployAzure or DeployAzureStack. Within the resource area I access the values within the variable like this.
|
||
"apiVersion": "[variables('DeployLocation').apiVersion]"
|
||
If you’ve taken the time to read Patterns for designing Resource Manager templates, you might recognize this from the t-shirt size examples. So, I didn’t come up with this on my own. It’s a powerful way to dynamically set values for your template, or pass a set of values to a linked template.
|
||
In order to install TFS 2015 and Visual Studio Community, I had to store the ISOs in Azure blob storage. I had the blobs publicly accessible, but that’s probably not a good idea now that I am writing this post. Instead, I plan to add a Shared Access Signature parameter and storage account URL to the template parameters so that I restrict access to myself. I’ve set the blob container to Private for the moment. I’ll drop a new post once the SAS and storage account parameters are in.
|
||
I’ve successfully deployed the template in Azure, and now I need to deploy it in Azure Stack and validate it is working. Of course, my Azure Stack install died on me after the lab lost power unexpectedly. So I’ve got an exciting weekend of reinstalling it.
|
||
`,summary:`This is the first post in a series of getting a CICD pipeline working with Azure Stack. You can read part 2 here, part 3 here, and part 4 here.. I will add links to additional posts as they are created.
|
||
There are a few things that have been coming up a lot lately at work that I would like to dive into some more to get a better understanding. The first is Infrastructure as Code (IaC).`,date:"14 Jan, 2017",url:"https://nedinthecloud.com/2017/01/14/cicd-pipeline-with-azure-stack-part-1/",image:"tutorials.png",readingTime:"5"},"https://nedinthecloud.com/2017/01/07/installing-azure-powershell-on-server-2012-r2-using-powershell/":{title:"Installing Azure PowerShell on Server 2012 R2 using PowerShell",tags:["azure","azure-stack","powershell"],content:"Currently I am working on a project to create a CICD pipeline in Azure Stack. I am planning to use Team Foundation Server 2015, running on a virtual machine on Azure Stack. I want the installation of TFS to be automated using ARM templates and a CustomScript extension. (I would use DSC, but I feel I’ve shaved the yak enough already). The installer and configuration files are sitting in Azure Blob storage, and I want to be able to pull them from the Server 2012 R2 instance that will be running TFS. I can do it using Invoke-WebRequest or Start-BitsTransfer, but I wanted to try using the Azure PowerShell storage cmdlets instead. Of course the vanilla install of Server 2012 R2 does not have the Azure modules or PowerShellGet module to access the PowerShell Gallery and install them. So instead I am going to use Invoke-WebRequest and Github to grab it instead.\nFirst of all the most recent release of the Azure PowerShell module is on Github under: https://github.com/Azure/azure-powershell/releases/latest. So I grab the whole page and store it in a variable:\n$resp = Invoke-WebRequest -Uri "https://github.com/Azure/azure-powershell/releases/latest"\nNow if you look at the response members, there is a Links property:\nPS C:\\Windows\\system32> $resp | gm TypeName: Microsoft.PowerShell.Commands.HtmlWebResponseObject\nName MemberType Definition ---- ---------- ---------- Equals Method bool Equals(System.Object obj) GetHashCode Method int GetHashCode() GetType Method type GetType() ToString Method string ToString() AllElements Property Microsoft.PowerShell.Commands.WebCmdletElementCollection AllElements {get;} BaseResponse Property System.Net.WebResponse BaseResponse {get;set;} Content Property string Content {get;} Forms Property Microsoft.PowerShell.Commands.FormObjectCollection Forms {get;} Headers Property System.Collections.Generic.Dictionary[string,string] Headers {get;} Images Property Microsoft.PowerShell.Commands.WebCmdletElementCollection Images {get;} InputFields Property Microsoft.PowerShell.Commands.WebCmdletElementCollection InputFields {get;} Links Property Microsoft.PowerShell.Commands.WebCmdletElementCollection Links {get;} ParsedHtml Property mshtml.IHTMLDocument2 ParsedHtml {get;} RawContent Property string RawContent {get;} RawContentLength Property long RawContentLength {get;} RawContentStream Property System.IO.MemoryStream RawContentStream {get;} Scripts Property Microsoft.PowerShell.Commands.WebCmdletElementCollection Scripts {get;} StatusCode Property int StatusCode {get;} StatusDescription Property string StatusDescription {get;}\nAt the bottom of the page is a link to an msi file, which is the latest Azure PowerShell installer. I find that link and store it as an object.\n$link = ($resp.Links | ?{$_.href -like "https*.msi"}).href\nThere are actually two links that end in msi, one is a relative path and the other an absolute path. By prefixing https to the search string, I get the absolute path href. The $link variable has the string “https://github.com/Azure/azure-powershell/releases/download/v3.3.0-December2016/azure-powershell.3.3.0.msi" stored in it. Obviously your string will change depending on the latest release.\nI could do a one-liner to get the file and save it, then execute it, but I like to keep things readable. Here I am grabbing the filename and then downloading it to my selected location.\n$installer = $link.Split("/") | select -Last 1 Invoke-WebRequest -Uri $link -OutFile "F:\\$installer"\nNow that I have the msi file, I can use msiexec to install it. Only one problem, installing Azure PowerShell automatically closes all instances of PowerShell. Not a big deal if you are doing this interactively, or you are done your work in PowerShell. But I am assuming you are running a script next that needs Azure PowerShell, or at least that is what I am doing. What I want to do is set a scheduled task for 5 minutes in the future to run the next piece of the script.\n$A = New-ScheduledTaskAction -Execute "PowerShell.exe" -Argument "F:\\MyScript.ps1" $T = New-ScheduledTaskTrigger -Once -At (Get-Date).AddMinutes(5) $S = New-ScheduledTaskSettingsSet $D = New-ScheduledTask -Action $A -Settings $S -Trigger $T Register-ScheduledTask PS1 -InputObject $D\nIn five minutes time, the next script will run and have the Azure PowerShell cmdlets available. Now we install the msi, and since I have the filename in a variable it is easy to reference.\n& msiexec /i F:\\$installer /qn\nThe installer should only take a minute or so to complete. Hope this helps someone besides me! Now onto installing TFS 2015.\n",summary:"Currently I am working on a project to create a CICD pipeline in Azure Stack. I am planning to use Team Foundation Server 2015, running on a virtual machine on Azure Stack. I want the installation of TFS to be automated using ARM templates and a CustomScript extension. (I would use DSC, but I feel I’ve shaved the yak enough already). The installer and configuration files are sitting in Azure Blob storage, and I want to be able to pull them from the Server 2012 R2 instance that will be running TFS.",date:"7 Jan, 2017",url:"https://nedinthecloud.com/2017/01/07/installing-azure-powershell-on-server-2012-r2-using-powershell/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2017/01/03/trends-for-2017/":{title:"Trends for 2017",tags:["artificial-intelligence","aws","azure","hybrid-cloud","hyperconverged","machine-learning","predictions"],content:`The end of 2016 is here, and I think many of us are breathing a sigh of relief. The year has not been kind to some, and has been described as a “dumpster fire” by others. On the whole, I actually think that 2016 was a pretty decent year, or at least no worse than most previous years. But I am a bit biased since my second daughter was born in June, and she is awesome! That’ll tip the scales regardless of what else happened. Anyhow, I digress. The tech industry has seen a lot of change, with new technologies emerging and companies innovating at a rapid pace. I’d like to use this post to take a look at a few of those trends, and which ones I will be keeping an eye on in the coming year.
|
||
Machine Learning and Artificial Intelligence In 2016 there were a lot of exciting developments in the Machine Learning and Artificial Intelligence spaces. The big ones were the announcements from public cloud vendors Azure and AWS, who have released their ML and AI APIs. Azure now has the Cortana Cognitive Services suite, which includes APIs for Speech, Images, Search, and Knowledge. AWS has released Lex, Polly, Rekognition and their Machine Learning API. Both public cloud vendors have also added VMs to their IaaS offering that include an FPGA for those who need the hardware acceleration for high performance computing. In the past AI and ML were only available to the largest companies and universities who could afford the hardware and software required to run it. Innovation was slow because the barrier to entry was high. Now that barrier has tumbled down to fractions of a cent on the dollar, anyone can dive into the world of ML and AI without necessarily knowing how to code or run the hardware.
|
||
Given the rate at which data is being produced, there is no way for a human being to comb through it and glean relevant insight. At least not in a reasonable amount of time. AI and ML will be instrumental in consuming and processing the mountains of data and helping mere humans make sense of it all. I fully expect the market for AI and ML to explode in 2017, and if I were a young programmer I would focus my efforts learning those public cloud APIs and applying them to massive datasets.
|
||
Hybrid Cloud and Hyperconverged Infrastructure There was a mythical time when all enterprises were going to abandon their datacenters and migrate everything to the public cloud. There was no doubt in some analysts minds that the datacenter was dead, and public cloud was the only path to a bright future. And then reality intervened. Public cloud is still very much part of the modern enterprise, and some organizations have even taken a cloud-first approach. That doesn’t mean there isn’t still a need for on premise datacenters and applications. There are numerous reasons for this requirement, here are just a few:
|
||
Regulatory compliance Security concerns Legacy applications (especially mainframe) Data sovereignty Proximity to endpoint And thus the hybrid cloud was born. The big trend in 2016 was how to build a hybrid cloud, and while some big players like VMware have taken a crack at it, no one has a solid solution that ticks all the boxes. The public cloud has two amazing things that just don’t exist on premise, automation and orchestration. When an EC2 instance is launched in AWS, there is a lot happening in the background to get that process completed. But you don’t care about that and you don’t really need to worry about it. AWS takes care of the whole process for you by automating the steps and orchestrating the automation together using their control plane. That sort of automation barely exists on the private cloud side. Sure you could run vCloud Director or Microsoft CPS, but both of those are difficult to set up and maintain. Enterprises want a turnkey solution that will give them the convenience of public cloud in their on premise datacenter. Such a solution will need to run on net new hardware, so there needs to be a scalable hardware construct that can be expanded on without a lot of work by internal IT. That is where Hyperconverged Infrastructure comes in. Having the networking, compute, and storage packed into a single node allows for that kind of simple expansion and delivery.
|
||
All of the Hyperconverged vendors are trying to implement the automation and orchestration pieces with varying degrees of success. Companies, like Nutanix, that are heavily focused on software are likely to lead the charge for a better hybrid cloud experience. Or it could be a third-party application like Terraform or Kubernetes that provides the overlay for hybrid cloud. Finally, there is Azure Stack from Microsoft, which takes the code from Azure and shrinks it down for the datacenter. This is still an emerging field, and I think we’ll see some interesting competition in 2017. I also forsee several acquisitions as the HCI market starts to mature and consolidate.
|
||
Disaggregation of the Networking Stack Networking is about 20 years behind the it’s storage and compute brethren. There was a time when you purchased your compute and storage from a single vendor and they provided all components of the stack. The mainframe included the hardware, firmware, operating system and applications all from a single vendor. You couldn’t run one vendor’s application on another’s hardware. You were locked in. Then the idea of general purpose CPUs came along, and a general purpose operating system. It freed the consumer to purchase hardware from one company (or multiple), the operating system from another, and the applications from other third parties. If you look at what happened after the introduction of Linux and Windows server operating systems, there was an explosion of innovation from hardware and application vendors. Storage has followed a similar trajectory, albeit a decade behind the compute. Lots of people still buy their big storage frames from a single vendor and use their bundled software stack. But you don’t need to do that. Products like Ceph and Storage Spaces Direct remove the need to purchase a frame and the software from a single vendor. Just like the world of compute, you can choose the hardware, operating system, and applications. Networking has yet to embrace the idea of disaggregation, though the rise of SDN has pushed the issue somewhat. Today you are still mostly buying your networking gear with the OS baked in, and limited support for third party applications. That is changing though, with vendors like Cumulus Networks, Big Switch Networks, and even Dell’s OS10, the customer can choose the hardware and operating system independently. Even the lower layers of the networking gear can be manipulated using FPGA chips or the coming products from Barefoot networks.
|
||
In addition to the trend of disaggregation, the mega-scale operators like AWS, Facebook, and Microsoft have all started to design their own hardware and software to meet the needs that traditional networking vendors have not. Facebook has created FBoss to run their switching environment, and their own switching gear with the first product being their ToR switch Wedge. In a similar vein, Arista networks has been quickly replacing the incumbents in the data center space by offering a switch with less features and a much lower price point. Their EOS (extensible operating system) takes a modular approach to network operating systems, that makes it both more reliable and extensible to third parties.
|
||
Essentially, networking has been fairly stagnant in terms of innovation over the last 20 years. Get ready for that innovation to rev up in 2017 as disaggregated networking blends with SDN to create an explosion of possibilities.
|
||
Everything to the Edge For too long the focus of IT was concentrating everything in data centers, and forcing the client to come and get the information. That worked fairly well when most clients were located a few floors or across a decent WAN link from the data center holding the information, but that’s no longer the case. Clients are now operating in a mobile fashion from laptops, mobile devices, and across cellular data networks. They expect apps to be fast and responsive, and can’t wait for the signal to travel up to the cell tower, down to the NOC, across the ocean to a data center and back. By the time a page has loaded, the user has already moved on. Data access needs to be almost instantaneous, and content needs to live as close to the edge as possible. At the same time technologies like HTTP2 over UDP are designed specifically to speed up page loads and cache content closer to the edge. AWS has introduce Lamba@Edge to make dynamic web applications function at a CDN endpoint instead of going back to the data center. Anything that can be done to reduce roundtrips is a benefit to the consumer.
|
||
It’s not just content consumption though. It’s also data generation happening at the edge. IoT devices are generating heaps of data that needs to be collected and analyzed at the edge, especially for real-time applications. Products like Azure’s IoT gateway and AWS GreenGrass are intended to try and deal with the glut of data being created by IoT devices. I would expect that innovations in this area will keep accelerating, working in tandem with the hybrid cloud approach to put things in the public cloud when you can and keep them local when you must. So called “edge” cases if you will.
|
||
Those are the trends I will be tracking through 2017. Of course, who knows what the new year will bring. I can’t wait to review this list in a year and see what I got right and missed.
|
||
`,summary:"The end of 2016 is here, and I think many of us are breathing a sigh of relief. The year has not been kind to some, and has been described as a “dumpster fire” by others. On the whole, I actually think that 2016 was a pretty decent year, or at least no worse than most previous years. But I am a bit biased since my second daughter was born in June, and she is awesome!",date:"3 Jan, 2017",url:"https://nedinthecloud.com/2017/01/03/trends-for-2017/",image:"featured-image-nedinthecloud.jpg",readingTime:"8"},"https://nedinthecloud.com/2016/12/29/system-recovery-for-an-hpe-hyperconverged-250-running-vmware/":{title:"System Recovery for an HPE HyperConverged 250 running VMware",tags:["esxi","hc250","hci","hpe","hyperconverged","vmware"],content:`This is a technical post for someone trying to reset a node or an entire HC250 appliance running VMware. This is specific to the latest release of the recovery software for the HC250 running ESXi 6.0 update 2. If you have followed the directions for restoring the node which are included in the HPE Hyper Converged 250 System for VMware vSphere User Guide then you will have downloaded the necessary files and created a USB drive to perform the node reset. And that’s where things start to fall apart.
|
||
If you are following the guide, the first issue you’ll run into is the fact that they are asking you to use UNetbootin to create the USB stick. Unfortunately, UNetbootin doesn’t work on Windows 10, so don’t even bother. Instead use Rufus, because it is awesome and doesn’t require you to actually install anything in Windows. The second thing is that, assuming you are using iLO, you don’t even need to create a USB stick. iLO allows you to mount a local folder from your laptop as a USB stick. First unzip the ISO (using 7-zip of course) currently named HPE_HC250_VMware_ESXi_6.0_U2_K2Q48-10601.iso to a folder. Then from within an iLO Integrated Remote Console session click the Virtual Drives drop-down, select Folder, and navigate to the folder you just unzip to. But don’t do that yet, because there is a third issue.
|
||
The user guide directs you to unzip the HPE_HC250_USB_Recovery_Tools_6.0_K2Q48-10610.zip file, and copy the contents to the bootable USB drive, selecting to overwrite the boot.cfg and syslinux.cfg files. If you do that and boot your server, it will just boot into the vanilla ESXi installer. That’s not what you’re looking for, and obviously not very helpful. The recovery process works by changing the boot.cfg to specify a kickstart file. Here’s the default boot.cfg file:
|
||
bootstate=0 title=Loading ESXi installer timeout=5 kernel=/tboot.b00 kernelopt=runweasel modules=(…)
|
||
I have excluded the modules section for brevity. The recovery boot.cfg looks like this:
|
||
bootstate=0 title=Loading ESXi installer (USB Reset CS250 TD3.5) timeout=5 kernel=/tboot.b00 #kernelopt=runweasel kernelopt=ks=usb:/HPE-CS250-SV-v6-0_T3-5usb.cfg modules=(…)
|
||
As you can see, the kernelopt has a kickstart file specified and the runweasel kernel option has been commented out. The HPE-CS250-SV-v6-0_T3-5usb.cfg file contains the script to properly reset the node back to factory defaults. When I first tried the process, the installer took me directly to the standard ESXi installer. Having not performed a factory reset before, I thought this was normal, so I walked through the installer using defaults. Then I expected it would run some special post install process to do the rest. It did not. After a few hours, learning more that I ever wanted about kickstart scripting and automating the ESXi install process, I realized that the title of the install screen was not “Loading ESXi installer (USB Reset CS250 TD3.5)”. Which lead me to realize that the wrong boot.cfg was being loaded.
|
||
It turns out that in the installer ISO there are two boot.cfg files. The first is in the root of the ISO, and the second is in the efi\\boot subfolder. Since the Gen9 server is booting in UEFI mode, it will use the boot.cfg file found in that folder as opposed to the root folder. The simple solution is to copy the boot.cfg file to both folders and then create the USB drive or mount the folder in iLO IRC.
|
||
I’ve let HPE know about this, so hopefully they’ll update the documentation soon. If they do, I’ll update the post to reflect it. On the bright side, I know a lot more about scripting an ESXi install than I did 24 hours ago!
|
||
`,summary:"This is a technical post for someone trying to reset a node or an entire HC250 appliance running VMware. This is specific to the latest release of the recovery software for the HC250 running ESXi 6.0 update 2. If you have followed the directions for restoring the node which are included in the HPE Hyper Converged 250 System for VMware vSphere User Guide then you will have downloaded the necessary files and created a USB drive to perform the node reset.",date:"29 Dec, 2016",url:"https://nedinthecloud.com/2016/12/29/system-recovery-for-an-hpe-hyperconverged-250-running-vmware/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2016/10/30/everything-is-broken/":{title:"Everything is Broken",tags:["agile","cloud","microsoft"],content:`Everything is Broken…
|
||
There’s more than one occasion where I have uttered the phrase, “Why can’t this just work?” Usually after battling it out with some piece of software that the marketing fluff described as “simple” and “easy-to-use” and turns out to be more like incredibly complex and completely undocumented. I want my technology to just work, but I also want it to be cutting-edge, infinitely configurable, and fully documented. Those who are familiar with the Project Management Triangle may realize that having all three is impossible. To which I say, what about in n-th dimensions?
|
||
Seriously though, I have noticed that with the speed of innovation, especially in the cloud, most things that are released are at least partly broken. And that’s not just for beta or preview features, generally available features and functionality are buggy and partly undocumented. Major releases of software have always had some bugs, which is why “It ain’t done till SP1” was a mantra among the Microsoft cognoscenti.
|
||
An Agile Peacock Weren’t Agile software development and continuous integration supposed to rescue us from this tide of buggy flotsam? In the old bad days of waterfall software development, release cycles were measured in years. And the marketing team often had to advertise features that didn’t exist yet, or weren’t fully baked. Then it was up to the product team to make sure the feature was in the final release, which resulted in poor regression and unit testing and last minute code commits finding their way into the gold release. Teams would work in isolation for years, and then push out a half-baked product, which would ultimately fail to meet customer expectations, and we would all collectively hold our breath until the first service pack came out.
|
||
Agile and CICD were supposed to rescue us from this quagmire, and in some ways they did. But instead of putting us on solid ground, we are now firmly established on a swamp, where we need to keep building up just so we don’t sink into the muck. Cloud companies are in an arms race of features, and Agile software development allows them to pump those features out at an alarming rate. Every week I am reading about the latest features now available in Azure or AWS or Office 365 or in some other cloud service. These features are usually released in some kind of preview, which is good, and those previewing it help the development team find bugs. But that also means if you want to stay on the cutting edge then you are basically an unpaid member of the QA team for company X.
|
||
Forever Your Beta It seems that all this started back when Google Labs was pumping out new ideas every few months. Each new idea would be released in Beta onto an unsuspecting world. The Beta label gave that service special provenance, if a user complained that a feature was broken or unusable, then Google could shrug and say that it was still in beta. Gmail, the most widely used email platform in the world, was in beta for five years, of which two years were a beta open to the public. Just because it was in beta didn’t stop people from using it for business critical applications, but the beta label gave Google some type of plausible deniability. Gmail set the standard for leaving projects in beta for as long as you please, and we all got comfortable using a service that was in some ways inherently broken.
|
||
Welcome to the QA Team! As someone who had a Gmail account from early on, I benefited from the seemingly limitless storage and also suffered from the occasionally buggy interface and half-baked features. In essence I was an extension of the existing Gmail QA team, and I worked as a tester in exchange for early access to a beta product. That model has proliferated across the industry, with beta programs for VMware, Microsoft, Citrix, and more. You have the privilege of running buggy software and providing feedback and bug reports, all for no cost to the vendor. There are some tangible benefits to this approach, especially in regards to how a product matures, which features end up in the final release, and the overall stability of the system upon general availability release.
|
||
A perfect example of this approach is Microsoft’s most recent release of Windows 10 and Server 2016. In the past, Microsoft has developed their operating systems and applications in a closed room, with very little input and testing from the larger community. It allowed Microsoft to tightly control the flow of information and features being developed, so they could have a big bang release announcement. But it also lead to a mentality that Microsoft’s software wasn’t ready for production until the first service pack was released. That approach began to shift a little with Windows Vista and Server 2008 (aka Longhorn), where there was an early release program you could sign up for to test the operating system and provide some feedback. There was only one preview release, and when Vista finally came out… well I think we all remember how successful that was. Windows 7 and Windows 8 had an increased number of previews and additional feedback from customers, as did Server 2012. Those releases were moderately more successful, and while things didn’t really hit their stride till Windows 8.1 and Server 2012 R2, their predecessors were definitely more usable on Day 1. With Windows 10 and Server 2016, Microsoft started the early release program long before the official release. In fact, the first technical preview of Server 2016 dropped in May of 2015, a full 17 months before the product was released for general availability.
|
||
Open Source all the Things! What’s the next step? Open Source Software of course. Although we are accustomed to being QA testers for organizations, the code is still locked up in their purview. We can find the bugs, and report them, but we lack the ability to examine the root cause and suggest a fix. In the world of open source solutions, you could just file a bug report on Github. However, if you are feeling suitably motivated, you can fork the project and fix the bug, then submit a pull request to incorporate your fix back into the master branch. Lots of applications and operating systems already have an open source version, and they have been enjoying the benefits for years. Other organizations are coming around on the idea, even former bulwarks of conservatism (read: Microsoft) have started open sourcing select projects, such as PowerShell. Will they ever go full OSS? It seems unlikely. Microsoft has too many enterprise customers that would distrust an open source operating system, and Microsoft itself is probably not ready to let everyone see the Windows kernel laid bare in all its tangled glory. Nevertheless it sees the value in customer feedback and input, and the way it is embracing open source and technical previews helps demonstrate this.
|
||
Everything is Broken… and that’s OK! In the end, everything is broken, and it will never all be fixed. The rapid pace of innovation in the cloud and on premise has created a whack-a-mole situation where problems just keep popping up like vicious little groundhogs, and sometimes we all feel a little like Bill Murray in Caddyshack. Some days I just want to blow up the whole dang thing. But it’s OK. The rapid change and innovation has also brought with it opportunity. Those of us in IT consulting have a job because we understand technology, and can keep up with the rapid pace. It’s our job to stress about all the things that are broken, so to the end-users it all seems to “just work”. If technology is magic, then we are magicians, and I guess that’s OK too.
|
||
`,summary:`Everything is Broken…
|
||
There’s more than one occasion where I have uttered the phrase, “Why can’t this just work?” Usually after battling it out with some piece of software that the marketing fluff described as “simple” and “easy-to-use” and turns out to be more like incredibly complex and completely undocumented. I want my technology to just work, but I also want it to be cutting-edge, infinitely configurable, and fully documented. Those who are familiar with the Project Management Triangle may realize that having all three is impossible.`,date:"30 Oct, 2016",url:"https://nedinthecloud.com/2016/10/30/everything-is-broken/",image:"featured-image-nedinthecloud.jpg",readingTime:"7"},"https://nedinthecloud.com/2016/09/26/automating-snapshot-creation-and-copying-in-aws/":{title:"Automating Snapshot Creation and Copying in AWS",tags:["aws","disaster-recovery","powershell","scripting"],content:`Building IaaS in the cloud is becoming more popular. And part of building IaaS is providing some level of disaster recovery. After spending the last few weeks working in AWS, I realized that the toolsets I expected to exist just don’t. So what did I do? Scripted my own, or at least a start.
|
||
Coming from the world of Azure, I am used to the availability of what they call LRS (locally redundant storage) and GRS (globally redundant storage). When a virtual machine is backed by LRS, there are always three-local copies of the data in the data center which are in separate failure domains. GRS takes that a step further by creating three copies of the data in a companion datacenter. When I started deploying instances in AWS, I mistakenly assumed that a similar storage backing was available to provide disaster recovery to a companion region. Evidently, I was wrong. So what’s the suggested alternative? Create snapshots of the volumes backing your instance, and then copy those snapshots to another region for durability. So naturally I looked for the setting to auto-create snapshots, and… there isn’t one. So instead I rolled my own using AWS PowerShell (since I love PowerShell) and running the process as a scheduled task.
|
||
The first script in the series is below. It takes the following:
|
||
region: the region that has the source instance instanceID: the source instanceID The script does a few things. First it finds the existing volumes for the instance and tags them with new tags. I included the instance name, instance ID, the mount point of the volume, and what region it’s in. All this information can be used later when the snapshot it copied over to another region. If it’s the root volume of the instance, I also add in the VPC ID and the Subnet ID, and set an IsRootVolume tag to true. I’m doing that in order to automate the recovery of the instance in another region. That script is still in development, but I figure if I know the instance ID, VPC, and subnet, I can probably recover a bunch of instances in a mirrored config in another region. Now that the volumes are tagged appropriately, I’ll create a snapshot. The snapshot will copy all the tags on the volume, and add a Date and Name tag to the snapshot. The date is nice for deleting stale snapshots, or finding a consistent point in time for a bunch of them.
|
||
The script ends by outputting the snapshot IDs, which can be ingested by another script to copy them over to another region.
|
||
Here is the current version of the script:
|
||
In a follow up post I will show a script for copying the snapshots over, and an orchestration script to perform the operation for an entire set of instances based on tags or VPC ID.
|
||
`,summary:`Building IaaS in the cloud is becoming more popular. And part of building IaaS is providing some level of disaster recovery. After spending the last few weeks working in AWS, I realized that the toolsets I expected to exist just don’t. So what did I do? Scripted my own, or at least a start.
|
||
Coming from the world of Azure, I am used to the availability of what they call LRS (locally redundant storage) and GRS (globally redundant storage).`,date:"26 Sep, 2016",url:"https://nedinthecloud.com/2016/09/26/automating-snapshot-creation-and-copying-in-aws/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2016/09/07/autoscale-groups-with-domain-join/":{title:"AutoScale Groups with Domain Join",tags:["aws","cloudformation","powershell","windows"],content:`When it comes to AWS, it often feels like Windows is a second class citizen. I have been doing work in the Azure public cloud for a long time, where the situation is somewhat reversed. In AWS, the commands, assumptions, and example use cases are almost always in Linux with Windows as a bit of an after thought. I’ve been doing a lot of AWS work recently, and one of the things that came up was the ability to deploy AutoScale groups and Launch Configurations using CloudFormation. By itself, that process is relatively straightforward, especially if you are working in a Linux context. But what if you are deploying Windows boxes in a domain and want them to have a specific hostname structure? That’s a bit more tricky, but I got it figured out.
|
||
The first thing to know is that I am deploying this configuration in an existing domain, and I am using CloudFormation to do it. The AutoScale group is part of a larger environment that includes various SQL servers, web servers, and application servers. The requirement to have domain-joined servers stems from a need to centrally apply policies and deploy software. The members of the AutoScale group fall under that umbrella, so they need to be part of the domain. There are two parts to the deployment, the first is the CloudFormation resources that deploy the Launch Configuration and AutoScale group, and then there is the PowerShell script that helps perform the renaming of the servers. The full stack template and PowerShell script are included at the end of the post.
|
||
Here is the AutoScalingGroup resource:
|
||
[code] “MyServerAutoScalingGroup”: { “Type”: “AWS::AutoScaling::AutoScalingGroup”, “Properties”: { “LaunchConfigurationName”: { “Ref”: “MyServerLaunchConfiguration” }, “DesiredCapacity” : “2”, “MaxSize” : “5”, “MinSize” : “1”, “VPCZoneIdentifier” : [{“Ref” : “AutoScaleSubnetId”}] } } [/code]
|
||
The VPCZoneIdentifier refers to a subnet Id that I grab from the parameters in the template.
|
||
The LaunchConfiguration resource has a few things worth noting. In the Metadata section I have added an AWS::CloudFormation::Init section, which is how each instance gets configured. Most of it is straightforward, like creating the configuration file for the cfn-auto-reloader. There is a PowerShell script that gets loaded to handle the renaming of the server. The source is an S3 bucket that is defined in the template parameters.
|
||
[code] “c:\\\\cfn\\\\scripts\\\\Rename-LCComputer.ps1” : { “source” : { “Fn::Join” : ["", [ “https://”, { “Ref” : “BucketName” }, “.s3.amazonaws.com/ps-scripts/Rename-LCComputer.ps1” ]]} } [/code]
|
||
Within the rename configset the script is called.
|
||
[code] “a-execute-powershell-script-RenameComputer”: { “command”: { “Fn::Join” : ["", [ “powershell.exe -executionpolicy unrestricted -command “, “c:\\\\cfn\\\\scripts\\\\Rename-LCComputer.ps1 -ServerPrefix “, {“Ref” : “ServerPrefix”}, " -DCIPAddress \\””, { “Ref” : “ADServer1PrivateIP” }, “\\” -credentials “, “(New-Object System.Management.Automation.PSCredential(’”, { “Ref”: “DomainNetBIOSName” }, “\\\\”, { “Ref” : “DomainAdminUserName” }, “’,”, “(ConvertTo-SecureString ‘”, { “Ref”: “DomainAdminPassword” }, “’ -AsPlainText -Force)))” ]] }, “waitAfterCompletion”: “forever” } [/code]
|
||
The script takes a ServerPrefix value that is used to figure out the server name. It queries Active Directory for all servers that begin with the prefix and sorts them in descending order. If no servers are found, it tacks on an 01 to the prefix and renames the server. If existing servers are found, it takes the last one and adds one to it, e.g. is myservername03 is found, server is renamed myservername04. The script will work up to 99 servers. The rest of the Init changes the DNS servers for the network interface and joins the instance to the domain. In order to access the S3 bucket, I have included an AWS::CloudFormation::Authentication section.
|
||
[code] “AWS::CloudFormation::Authentication” : { “S3AccessCreds” : { “type” : “S3”, “accessKeyId” : { “Ref” : “CfnKeys” }, “secretKey” : {“Fn::GetAtt” : [“CfnKeys”, “SecretAccessKey”]}, “buckets” : [ { “Ref” : “BucketName” } ] } } [/code]
|
||
Here is the full stack template:
|
||
And here is the PowerShell script:
|
||
Share and enjoy!
|
||
`,summary:"When it comes to AWS, it often feels like Windows is a second class citizen. I have been doing work in the Azure public cloud for a long time, where the situation is somewhat reversed. In AWS, the commands, assumptions, and example use cases are almost always in Linux with Windows as a bit of an after thought. I’ve been doing a lot of AWS work recently, and one of the things that came up was the ability to deploy AutoScale groups and Launch Configurations using CloudFormation.",date:"7 Sep, 2016",url:"https://nedinthecloud.com/2016/09/07/autoscale-groups-with-domain-join/",image:"tutorials.png",readingTime:"3"},"https://nedinthecloud.com/2016/08/16/bridging-the-digital-divide/":{title:"Bridging the Digital Divide",tags:[],content:`There’s nothing I like more than talking tech with anyone who will listen, well except maybe actually doing the tech. But talking is pretty keen too. It’s a lot of fun to discuss the new technologies that are emerging, or listen to other’s experiences and issues. Talking tech is something that I think all technology professionals like to do, especially consultants. In the past, having an engrossing technical discussion could also lead to new business opportunities, as the technical leaders in the client’s organization were often the decision makers. While that is still true, the decision makers are increasingly from non-technical areas of the company. It could be someone in marketing who wants to adopt a new web platform for the next major marketing campaign. Or it could be the head of Finance, who is tired of waiting for the IT department to update their aging accounting software. These people have budget to spend, an interest in technology, and no one to talk to. That’s why as IT professionals, we need to bridge the gap.
|
||
Speaking in Tongues Have you ever heard an accounting person arguing with an IT person over a technical issue? It’s like the two of them are not even speaking the same language. In fact, even if they’re both using the English language, their words are meaningless to each other. The accounting person is talking about return on investment (ROI), general accounting principles (GAP), and finance controls. The IT person is talking about security patches, IP addressing issues, and firewall rules. It’s no wonder that these two parties walk away feeling misunderstood and frustrated. If someone were able to mediate the exchange and translate the jargon both ways, each party might feel that they have a common goal, and what the challenges are for each party to achieve those goals. Instead too often a request to IT feels like a battle, almost as if the IT department is actively trying to stymie the efforts of the business to improve their process and procedures. Suggestions are shot down, applications take months or years to roll out, and when they do finally launch the applications don’t even meet the business needs of the end users. The end result? Shadow IT and the diminished role of internal IT within the company.
|
||
Mind the Gap Non-technical folks speak in business terminology, and when they express a business need or requirement they expect that an IT team, internal or external, can interpret those requirements into a technical construct that meets the business need. Designing a solution is often as much about business requirements and goals as it is about technical architecture. After all, you can build an elegant technology solution, but if it doesn’t actually meet the needs of the business then it is effectively useless. Worse yet, it is a waste of resources which never sits well with business people. That’s why whenever I am working on a new solution, I want to make sure I am talking to the project sponsor and the business stakeholders. They know what they need from a solution, at least from a business perspective. Only after gathering the business drivers and requirements do I try to assemble a technical solution. Too often we approach from the opposite end where we let the technology we want to implement steer the conversation. There have been several planning meetings that have a pre-game meeting with just the IT folks and the conversation goes something like, “Alright, we need some additional capacity in our vSphere infrastructure, so let’s oversize the application and try to get some additional capacity and storage out of this project.” The assumption has already been made that the application will run in-house, virtualized, on some array-based storage. Those are some broad assumptions, and possibly totally wrong. We need to go into the planning meeting with an open mind, and walk away with a list of business requirements in a prioritized list. Once the requirements are known, then a solution can be pitched and validated by the business.
|
||
Be the Bridge Rather than a roadblock or a ferry - stretching the metaphor a bit here I know, we can be the bridge between business and IT. How do we do it? By learning some business fundamentals to be able to speak to non-technical and technical folks alike. I’m not saying you need to go out and get your MBA, but being able to understand concepts like business strategy, core competencies, ROI, and total-cost-of-ownership are key to engaging with the business people in the room. Being able to explain how an all-flash array enables a client to grow their business and improve their bottom line makes an ally out the business team, and helps present IT as a driving force for growth and not just a cost center. So while I would encourage you to go out and sharpen your technical skills, at the same time it couldn’t hurt to read a book on business fundamentals. There are also some great websites that track business trends along with changes in technology. Try occasionally looking at CRN, CIO.com, and BusinessWeek to keep abreast of the business side of the house, and next time you hear someone talking about CapEx and OpEx those won’t just be a marketing line from Office 365.
|
||
`,summary:"There’s nothing I like more than talking tech with anyone who will listen, well except maybe actually doing the tech. But talking is pretty keen too. It’s a lot of fun to discuss the new technologies that are emerging, or listen to other’s experiences and issues. Talking tech is something that I think all technology professionals like to do, especially consultants. In the past, having an engrossing technical discussion could also lead to new business opportunities, as the technical leaders in the client’s organization were often the decision makers.",date:"16 Aug, 2016",url:"https://nedinthecloud.com/2016/08/16/bridging-the-digital-divide/",image:"cloud-header-image2.jpg",readingTime:"5"},"https://nedinthecloud.com/2016/08/11/burned-bridges/":{title:"Burned Bridges",tags:[],content:`There’s a really good Coalesce song called Burned Bridges, which The Get Up Kids covered. Both versions are awesome so go take a listen, I’ll wait.
|
||
The reason I bring it up is not just because it’s a good song, which it is, but because some people prefer to flame out and burn bridges. Recently, I heard that someone I worked with had put in their notice, which was probably a good idea. He had been at the same place a long time, and had become enmeshed in the politics of the organization, and not in a positive way. This led him to be defensive, secretive, and at times cantankerous. A fresh start was definitely in order for both him, and the department he worked in. So, a win-win for everyone.
|
||
That should be the end of the story; an amicable split is not uncommon in our industry. People grow and change, and so do organizations. If the arrangement doesn’t fit anymore, then it’s time to move on. But of course that isn’t the end, because a few days before he left, he sent a series of scathing emails to his entire department criticizing the leadership, the group competency, and specific things he found objectionable about people he worked with. I like to call it a scorched earth policy.
|
||
And you know what? I get it. Not everyone you work with is a bastion of competence. You aren’t going to get along with everyone, and you may think that certain aspects of the company’s leadership are amiss or wrong-footed. That’s normal. That’s why you leave a company and work somewhere else. And you know what else? I bet writing those scathing emails felt good. It was probably cathartic, and possibly illuminating. But what will the recipients of those emails think?
|
||
“That guy was a huge jerk, I hope I never work with him again.”
|
||
Did you know that dissatisfied customers will tell between 9-15 people about their experience? If you think of fellow employees as customers, then being a total asshole to them means they will probably tell 9-15 people about what a Jerky McJerkface you were. Once word gets around that you are difficult to work with, or tend to flame out at a company, managers are not going to want to hire you. Even if you are extremely competent at your job, in all likelihood it is just not worth the headache for them.
|
||
The IT community is not that big. As a consultant, I end up at a lot of different places, and unsurprisingly I run into a lot of the same faces. The average tenure for a job in Information Technology is less than five years according to the 2014 Bureau of Labor Statistics. Combine that with LinkedIn and other networking technologies and your reputation can precede you big time. If you want to flame out, fine. But even if you already have a new job lined up, which I hope this guy did, that job will probably not last more than five years. You think in five years everyone will have forgotten? Get real. One Google search will bring it all right back.
|
||
When you are preparing to depart for a new job, be gracious. If you can’t be gracious, be courteous. If you can’t do that, then just be civil. Go ahead and write those nasty emails, then delete them. You still get the catharsis, without the penalty to your reputation.
|
||
It all goes back to my third pillar: Be Nice.
|
||
`,summary:`There’s a really good Coalesce song called Burned Bridges, which The Get Up Kids covered. Both versions are awesome so go take a listen, I’ll wait.
|
||
The reason I bring it up is not just because it’s a good song, which it is, but because some people prefer to flame out and burn bridges. Recently, I heard that someone I worked with had put in their notice, which was probably a good idea.`,date:"11 Aug, 2016",url:"https://nedinthecloud.com/2016/08/11/burned-bridges/",image:"featured-image-nedinthecloud.jpg",readingTime:"3"},"https://nedinthecloud.com/2016/08/04/hyperconverged-infrastructure-the-story-so-far/":{title:"Hyperconverged Infrastructure: The story so far",tags:["hci","hpe","hyperconverged","nutanix","virtualization","vmware","vsan","vxrail"],content:`In case you’ve been living under a rock, you might have noticed that HyperConverged Infrastructure (HCI) is a fast-growing segment of the hardware market. At least that’s what all the marketing campaigns and analysts are saying. Gartner said earlier this year that the HCI market will grow to consume 24% of the integrated systems market by 2019. According to IDC, HCI systems made up 11.4% of integrated systems sales in Q42015 with a YoY growth of 170.5%. Of course to a certain degree this is derived from vendors pushing HCI very hard at the consumer and in the channel. I can’t tell you how many briefings, webinars, and marketing campaigns I’ve been hit with about HCI in the last 12 months. Excessive is the word that comes most readily to mind. Nevertheless, where there’s smoke there is likely hyperconverged fire.
|
||
Working for a company with multiple hardware partners has its advantages, and this has allowed me to actually get my hands on four different HCI systems. Before I go any further, this is the part where I tell you that the opinions expressed in this post are mine and mine alone, and that this in no way reflects or is endorsed by my employer. There, now I feel better. Moving on…
|
||
So like I said, I’ve am lucky enough to have worked on four different HCI systems. They are as follows:
|
||
Nutanix HPE HC250 HPE HC380 VxRail Before I get into the details of the various systems, I want to pontificate a little about HCI and its place in the industry. The way that HCI is billed, you would think that it was a panacea for whatever ails your datacenter. Need more compute? Want to get rid of storage arrays? Need more automation? If you answered yes to any or all of these, then HCI is for you! The reality is that HCI is the right tool for some jobs, but by no means all. I see the main use cases for when:
|
||
You are considering purchasing new hardware for a specific application You are building a new datacenter and want a scalable solution You are placing hardware in a rented rack and want maximum density An HCI is not going to replace your storage arrays, at least not in the short term. Storage arrays are easily expandable, with tiered storage types that meet your needs. An HCI appliance has a pre-configured amount of storage in it, and you usually need to purchase identical nodes for the storage replication to work properly. So expanding the storage in a node is a non-trivial affair. For the time being, it’s likely you’ll be maintaining a SAN alongside your hyperconverged nodes. Maybe in the long term all lower tiers of storage can be shuffled up into S3 or StorSimple.
|
||
An HCI is also not going to replace all of your traditional dedicated servers, the ones that are not virtualized and aren’t going to be anytime soon. While it’s true that you can basically virtualize anything you want, that doesn’t necessarily mean that you should. HCI relies on a hypervisor, typically ESXi or Hyper-V, and virtualization to run all of its instances. If you have an Exchange buildout following the preferred architecture or an Oracle RAC on physical boxes, then HCI isn’t going to virtualize and condense those bad boys.
|
||
An HCI is not going to reduce your network footprint, at least not anymore than any other virtualization solution. And in some cases it will require a higher port density on your ToR switches. For instance, an HPE C7000 enclosure with two Virtual Connect FlexFabric 20/40 modules only really needs four 10Gb uplinks (two per module) to your ToR switches. It’s possible that you will need additional uplinks if things are getting saturated, so let’s say a maximum of eight uplinks (four per module), providing connectivity for 16 hosts in the enclosure. That’s a ratio of 1:2 per host, and I’m not even considering the 40Gb QSFP ports that would drop it down to 1:4 or 1:8. In HCI world, each host has two 10Gb ports, so the ratio gets flipped to 2:1 per host. What about rack density? C7000 enclosures are 10U a piece with 16 half-height hosts. So you can fit 4 in a rack with two 1U ToR switches, assuming you can serve enough power to the rack for all four enclosures. You would need 32 ports for the enclosures, and then sufficient ports to uplink your ToR to your spine switches (assuming leaf and spine). Two 24 port switches would probably be good enough, or you could go 32 or 48 ports to be safe. That would be 64 hosts per rack. An HC250 has four nodes in 2U, so each 2U requires eight 10Gb ports. Load up 40U and you’ve got 80 nodes requiring 160 ports! You’ve got an additional 16 hosts in your rack, but you need a port density that goes beyond two 48 port ToR switches. Now you’re looking at two 2U 96 port switches.
|
||
So what do you get with HCI that’s different from buying a blade chassis and storage array? Well first off you don’t need a storage array, this is converged infrastructure after all. Each node in the HCI has an identical amount of storage in it, and the contents of that storage is replicated to other nodes in the HCI environment. The value that an individual vendor adds to the equation is the software which handles that storage replication. In HPE based hardware, the storage replication is handled by StoreVirtual VSA (virtual storage appliance?) formerly known as Lefthand before they were acquired by HPE in 2008. Nutanix uses their own Acropolis data fabric. VxRail uses VSAN naturally. Each other HCI vendor either uses VSAN or some proprietary storage replication technology. For me, the different replication technologies can be a key differentiator since they determine how much usable storage you actually end up with, and how reliable/durable your data is at rest. Then there are questions of which ones can support native encryption, site-to-site replication, data deduplication, and more.
|
||
The second thing that HCI promises is the linear scalability of components. If you need additional capacity for a particular cluster, you just snap in an additional node and away you go. The management plane should make that process straightforward and quick. In theory you shouldn’t need to install a hypervisor, configure a bunch of settings, and all the other rigmarole associated with expanding your compute infrastructure. The key differentiator between the various HCI offerings is how truly simple it is to pop in another node, and what the caveats are surrounding node expansion. Do you need to add a certain number of nodes at a time? Does the expansion have to exactly match what you already have? What is the maximum nodes per cluster, and does performance suffer beyond a certain number of nodes? A lot of this comes down to the hypervisor being used, the management overlay, and the automation layer. Which brings me to my last point…
|
||
The final thing that HCI promises to bring is simplified automation and end user driven provisioning. That’s not unique to HCI, and you really could do it with any platform, but since an HCI vendor controls the whole stack of components, they should in theory be able to bring some powerful automation and a programmable API. This is chasing the dream that already exists in the public cloud sphere with AWS and Azure. Okay okay, Google Compute Engine too. All of those public cloud providers are offering Infrastructure as a Service (IaaS), the ability for users provision resources for themselves through a portal or an API. The ability to monitor consumption. The ability for admins to create a library of complex, pre-built application solutions and provide them to the end-user. And the ability to expand to web-scale capacity in a linear and predictable fashion.
|
||
That’s the promise of Hyperconverged Infrastructure. A nascent field with a bunch of offerings. How well do they live up to the promise? In my next post we’ll take a look at the HPE HC250 and HC380.
|
||
`,summary:"In case you’ve been living under a rock, you might have noticed that HyperConverged Infrastructure (HCI) is a fast-growing segment of the hardware market. At least that’s what all the marketing campaigns and analysts are saying. Gartner said earlier this year that the HCI market will grow to consume 24% of the integrated systems market by 2019. According to IDC, HCI systems made up 11.4% of integrated systems sales in Q42015 with a YoY growth of 170.",date:"4 Aug, 2016",url:"https://nedinthecloud.com/2016/08/04/hyperconverged-infrastructure-the-story-so-far/",image:"analysis.png",readingTime:"7"},"https://nedinthecloud.com/2016/07/28/my-three-pillars/":{title:"My Three Pillars",tags:[],content:`In the last six years I have been lucky enough in IT to be fairly successful and advance my career. Lately I’ve been reflecting on what I did right, and what I might change. In looking to the future and where my path is going, I find I have to look into the past and better understand how I got here. While there was no definitive five year plan, I think there were three pillars that served me faithfully to enable growth and advancement.
|
||
Make Yourself Uncomfortable In 2010 I had been at the same job for six and a half years. I had started as a desktop admin and progressed to be in charge of servers, networking, and storage. But I had definitely reached the zenith of where I could advance in the company. Short of taking my boss’ job, unlikely, there was really nowhere to go. The job was fairly easy, I knew everyone in the company, and I had a sweet three mile commute from my house. It was all very comfortable, and you know what? I could still be at that company doing the same work, with roughly the same people, and the same easy commute. It would be easy, not challenging, and completely boring.
|
||
If you want to advance in your career, you have to make yourself uncomfortable. You have to do things that make you uneasy or uncertain. Now I am not saying you have to do something that terrifies you. That seems counterproductive and likely to backfire. Instead, I think you have to make yourself a little uncomfortable in increments. For instance, I am not a huge fan of traveling to new places and working with people I don’t know. So naturally I went into consulting, wherein I meet new people and work in new places constantly. When I first started, it made me uncomfortable, and honestly sometimes it still does. But I know that each new place and person comes with new opportunities, so there is a benefit to all that discomfort.
|
||
Discomfort is usually a sign you’re doing something right. Plan to Fail After I left my comfortable job in 2010, I worked at a university for two years. And that was a great job. In two years I learned a ton about enterprise systems and working at scale that I never would have discovered had I stayed in my previous position. In 2012 opportunity knocked in the form of a move to consulting. There was a promise of working with cutting edge tech, and a healthy increase in salary to go with it. Even though the idea of consulting made me uncomfortable, I still went for it.
|
||
Four months in I knew I had made a bad decision.
|
||
There was a lot more travel than I had bargained for. I was being asked to implement tech that I wasn’t qualified to configure, and install it in a fraction of the time it should have been necessary to do properly. The team was me and two other people, who are both really nice guys, but it felt like we were totally disconnected from the parent company and pretty much making things up as we went. The phrase “fly by night” and “by the seat of one’s pants” come to mind. Smoke was being blown into orifices, and I was fairly miserable.
|
||
My first foray into consulting and I was failing. And that is okay.
|
||
Not that it felt okay. It felt pretty awful. Failure isn’t supposed to feel good. If I wanted to sit back and never fail, I would have never left the comfy couch.
|
||
Six months in I resigned. And that was okay too. I found a new job in consulting, and based on my prior experience I had a much clearer idea of what kind of company I wanted to work for and what type of consulting I wanted to do. I had two kids under the age of two, so travel was out. I wanted an established practice with subject matter experts I could rely on if I didn’t know the answer. And I wanted to be able to prepare myself before I was pushed out in the field to implement or design a solution. I ended up getting all of those things!
|
||
You’re going to fail. And that failure will help you succeed. Be Nice I had a boss who once introduced me as the nicest person he knew. I was a little embarrassed, and it made me sound a bit milquetoast. But he was totally right. I am nice. And I strive to be nice to other people. Being pleasant works for me, and I think it has greased the wheels a little when it comes to career advancement. Can you climb the ladder whilst being a total asshole? Sure you can, I’ve seen it. But those people tend to burn bridges, and unless you are especially smart, indispensable, or absurdly attractive, eventually you’re going to piss off the wrong person and find yourself in the unemployment line. I’ve seen that too.
|
||
I am not saying you need to be a doormat. You can be nice and still be assertive. You can be nice and still have an opinion. You can be nice until it is time to stop being nice. But the point here is the IT community is not that big, and people talk. I’d rather have a reputation as someone who is easy to work with, than a total dick.
|
||
I’ll let Dalton have the last say on that one:
|
||
All other things being equal, people would rather work with someone who is pleasant. `,summary:"In the last six years I have been lucky enough in IT to be fairly successful and advance my career. Lately I’ve been reflecting on what I did right, and what I might change. In looking to the future and where my path is going, I find I have to look into the past and better understand how I got here. While there was no definitive five year plan, I think there were three pillars that served me faithfully to enable growth and advancement.",date:"28 Jul, 2016",url:"https://nedinthecloud.com/2016/07/28/my-three-pillars/",image:"featured-image-nedinthecloud.jpg",readingTime:"5"},"https://nedinthecloud.com/2016/07/27/options-for-azure-migrations/":{title:"Options for Azure Migrations",tags:["azure","azure-classic","azure-resource-manager","microsoft"],content:`There are two deployment models in Azure, the older being the Service Management model, aka classic mode. The newer model is Azure Resource Manager (ARM). For reasons that extend beyond this post, Microsoft is moving away from the classic mode and adopting ARM wherever possible. Up until a few months ago, the two models had not yet reached feature parity and so classic was still required for some deployments. At this point the two models are at feature parity, and in fact ARM has pulled ahead. That gap is only going to widen as Microsoft continues to pour investment into ARM and leave classic to die on the vine.
|
||
If you are looking into migrating your Azure classic virtual machines to ARM, you might be wondering what your options are. There are several potential solutions, Microsoft supported and otherwise. Each has a set of limitations and gotchas, and in this post I intend to review them and provide a guide for using Azure Site Recovery to get the job done.
|
||
The solutions I have considered are the following:
|
||
Microsoft Supported Migration Process Community Supported Scripts Azure Site Recovery The Microsoft supported migration process allows you to migrate either an Azure VM that is not associated with a Vnet or an entire classic Vnet, including all the VMs on it. The nice part about their service is that it manipulates the management plane of Azure, without impacting the data plane. The data plane includes the storage, network, and actual running VMs, and since it is non-disruptive to the data plane, all of your VMs stay up during the migration process. Of course there are a bunch of caveats, for instance you cannot migrate your network gateway. So if you have site-to-site VPNs connecting to that Vnet, they are all going to be disconnected until you provision a new gateway in resource manager. Since the data plane is not changing, you cannot move those VMs to an existing Vnet or select a subset of VMs to move. This is an all or nothing migration. The first time I tried to use the service, the migration failed with a mysterious error. I posted in the forums, and Microsoft responded that the bug had been addressed. I retried and was successful, but it’s obvious this is very new code that still has some bugs to shake out. I think that this solution works best if you have not yet deployed anything into ARM, and you do not want to make any significant changes when moving over to ARM.
|
||
The community supported scripts break down in to a PowerShell module for migration (asm2arm), and a GUI tool called MigAz. Aidan Finn has done an excellent write up on MigAz, so please check that out that post, I won’t rehash it here. Either of the tools will clone the VM by creating a json deployment template and a script to copy the VHDs from a source storage account to a target storage account. In order to execute the full migration, the source VM will need to be stopped, then the VHDs are copied over, and the new VM is spun up. The VHD copy time is going to be pretty minimal, since you are copying within the same datacenter, but there will be downtime for the VM. The PowerShell module has some limitations as well. You cannot select a target virtual network, a target storage account, or place the storage account and network in separate resource groups. Again, that’s okay if you have nothing in ARM already, but if you are trying to migrate to an existing virtual network, then you’ve got some problems. I cloned the Github repo and updated the script to take arguments for all those items, so you can use my version of the script to migrate a classic VM to an existing virtual network and storage account.
|
||
Azure Site Recovery (ASR) is a service in Azure available in both classic and ARM mode. In classic, it targets Backup Vaults and in ARM it targets the Recovery Services vault. ASR uses scheduled replication to protect virtual and physical machines both on premise and in Azure. Machines have an agent installed which runs the replication process and sends the data to a process server, which in turn sends the rolled up replication to the vault. From the vault you can initiate an unplanned failover of the machine. In the plan for the unplanned failover, you specify the target virtual network, the target machine name, and target machine size. Before running the failover, the source machine needs to be powered down, so just like the community script, the migration process will require downtime. The ASR option natively supports more than just the basic classic to ARM migration. You can also do the following:
|
||
Migrate to new Azure subscription Migrate to a different Azure region Migrate within an Azure region Migrate from on premise to Azure In my opinion, if you are going to be migrating a significant number of VMs and some downtime is acceptable, then ASR is your best bet. The process is straightforward, supported by Microsoft, and has a lot more flexibility. If you would like to perform the migration, then the steps below will get you there.
|
||
Enabling Replication Prior to performing the migration, you will need to provision a Recovery Services vault in the region where the VMs will be migrating to. Then you will need to stand up a management server in the region and virtual network that has the source VMs to be migrated.
|
||
Once the basic ASR infrastructure is in place, you can follow the steps below to migrate a VM.
|
||
Azure virtual machines are treated as physical machines because there is no access to the hypervisor. Prior to enabling replication, the target machine should have the Windows firewall configured to allow the File and Printer Sharing and WMI management as shown below:
|
||
The applet can be found by opening the Control Panel and selecting Windows Firewall, not the Advanced Firewall. Make sure to select each feature for all three profiles: Domain, Private, and Public. Additionally, the account that was added to the local configuration server for client installation should be added as a local administrator. This will enable the automatic push of the replication agent. The target machine should also not have any pending reboots. If a pending reboot exists, the agent installation will fail. If a manual installation is desired, it should occur prior to attempting to enable replication. Instructions for manual installation can be found here.
|
||
In the Recovery Services vault under Settings select Protected Items -> Replicated Items. Then click on the + Replicate button. This will start the Enable replication wizard.
|
||
Select the configuration server, a Machine type of Physical Machines, and a Process server, typically collocated on the Management Server. If the site has more than one process server, select the one with the least load or the one on the same LAN of the physical machine.
|
||
Fill out the values for the target virtual machine including the subscription, storage account, network, and subnet. The deployment model should be Resource Manager.
|
||
Note: The storage account selected will be used when the virtual machine is restored. It must be a general storage account and not a Blob specific account, and it must be located in the same region as the vault.
|
||
Select an existing physical machine or click the + Physical Machine button to add a new one. When adding a new physical machine, provide the Name, IP Address and OS Type.
|
||
Select which account should be used to replicate the machine. The account should have local administrator permissions to install the replication agent and perform the replication back to the process server.
|
||
Select which replication policy should be used to protect the machine.
|
||
Once all settings are complete, click on Enable replication to start the replication process.
|
||
The machine will now begin the replication process. If there is a pending reboot on the target machine, the agent installation will fail. Once the replication process has completed, the next step is to perform the actual migration.
|
||
Machine Migration Shut down the source machine and note the time; it will be used to compare against the latest synchronization time in the vault. In the Recovery Services vault under Settings, select Replicated items in the Protected Items section.
|
||
Confirm that the last data sync corresponds with the time the VM was shut down.
|
||
Note: It may take up to 15 minutes for the last data sync time to change.
|
||
In the Computer and Network section of the replicated item Settings, verify that the Name, Size, and Network properties are correct.
|
||
On the replicated item blade select Unplanned failover to initiate the migration. Choose a Recovery Point and uncheck the Shut down machine option. Physical machines and Azure virtual machines cannot be shut down by the configuration server.
|
||
Monitor the status of the Unplanned failover task until it completes.
|
||
Verify that the target virtual machine is running.
|
||
If desired, allocate a public IP address for the target virtual machine. In the settings for the virtual machine navigate to the network interface blade and then the IP addresses blade. Enable the Public IP address setting and then allocate a public IP address either by selecting an existing public IP address object or by creating a new one. When finished click the Save button to complete the operation.
|
||
Associate a Network security group with the target virtual machine. NSG settings do not migrate with the virtual machine. In the settings for the network interface, click on Network security group and select an existing NSG or create a new one.
|
||
Log into the target virtual machine and validate that applications and services are working as expected.
|
||
Complete the migration. In the Recovery Services vault open the replicated item that is being migrated. Click on the Complete Migration button and then confirm on the pop-up blade. This will remove the item from replication protection and commit the point-in-time recovery point on the target virtual machine.
|
||
Once the failover is completed, the source Azure virtual machine is removed from protection. That virtual machine can be deleted once the target virtual machine is confirmed and validated. If backup protection was configured for the source virtual machine, it will need to be removed and re-enabled on the migrated VM. The migrated VM can also be enabled for protection for the purposes of disaster recovery.
|
||
Although this procedure only covers a single virtual machine, recovery plans can be created to orchestrate bigger migrations. Within the recovery plan it is possible to at pre and post script to perform actions like setting up availability sets, allocating public IP addresses, or configuring load balancing.
|
||
The Bottom Line Before taking my leave, a few words about cost. The cost of Azure Site Recovery by itself is based on the number of protected instances. For the first 31 days each protected instance is free. For the purposes of migrating a machine, the cost for ASR is negligible since the machine will not be protected long enough to incur charges. After the initial 31 days there is a standard cost of $54 per month for each protected instance. Actual pricing may vary.
|
||
In addition to the costs of protecting each instance in ASR, the cost of storage, storage transactions, and outbound data transfer must be taken into account. The vault storage is considered blob type and follows the pricing model depending on the storage replication type. The recommended replication type is LRS, which has the lowest cost per GB. The replication traffic from on-site to ASR would be considered all ingress traffic, and thus not charged. Replication traffic between Azure sites is charged based on the cost structure detailed here. Replication traffic within the same Vnet is not charged.
|
||
During disaster recovery testing, new virtual machines will be created. The charges associated with the test virtual machines follows the standard cost for Azure virtual machines. This includes the cost of additional storage for the test virtual machine VHD files. If the test virtual machines are not being used, they should be shut off to avoid incurring charges.
|
||
`,summary:"There are two deployment models in Azure, the older being the Service Management model, aka classic mode. The newer model is Azure Resource Manager (ARM). For reasons that extend beyond this post, Microsoft is moving away from the classic mode and adopting ARM wherever possible. Up until a few months ago, the two models had not yet reached feature parity and so classic was still required for some deployments. At this point the two models are at feature parity, and in fact ARM has pulled ahead.",date:"27 Jul, 2016",url:"https://nedinthecloud.com/2016/07/27/options-for-azure-migrations/",image:"tutorials.png",readingTime:"10"},"https://nedinthecloud.com/2016/07/20/write-it-right/":{title:"Write It Right",tags:[],content:`One of the things that I found most surprising about working in technology, and consulting in particular, is the sheer amount of writing we have to do. When I think about it, on a daily basis I am writing emails, IMs, text messages, documentation, memos, blog entries, or even newsletter articles. You would think that working in technology would be all plugging in cables, configuring software, and running scripts. But instead each day is full of reading and writing. That’s why it is so critical to be mindful of what you write and how you write it. In short, you need to Write it Right.
|
||
Basics Writing it Right means adhering to some fundamentals of writing, which include making sure your spelling and grammar are correct. That goes for any writing whether it is an IM, email, or formal document. Obviously, some spelling and grammar errors in chat and text are inevitable, and the medium expects a certain level of informality and brevity. In more formal mediums, such as email and documentation, there really is no excuse for misspelled words or incorrect grammar. The tools to prevent such errors are built right into the products we use to write! It may sound silly, but clients and coworkers can and will judge you by the quality of your writing, so make it easy on yourself and utilize the tools in Word and Outlook to prevent common errors.
|
||
Content Writing it Right also means being aware of the context and content of what you are writing. The language and style you use in an email to a coworker will probably be a little different than what you would use for client communication. Depending on the document type the content should have a sliding level of formality going from the most informal (text/IM) to the most formal (Statement of Work/Final Documentation). Just as you would not put legal verbiage in to a text message or tweet, you also should not put a humorous pun into an SOW. Clients and customer expect professionalism in communication, and may share or forward anything you write to other persons in their organization. When sending an email, think to yourself “Would I want everyone in the client’s company to read this?”. If the answer is “no”, then it’s probably best to reword or skip it. Don’t forget that anything you release out into wild is subject to sharing and legal discovery.
|
||
Permanence Writing it Right means writing less and being more concise. It also means creating templates and forms to replace manual creation of documents. This is important to create consistency in solution delivery and speed up the process of generating documentation. Writing it Right also means have a back catalog of documentation available. There are many times that I have been able to shorten delivery times or avoid common pitfalls by referring back to documentation on a previous project. Clients also appreciate solid documentation, especially if someone takes over the environment and never received knowledge transfer from the previous custodian. Having an excellent document to refer to can be a lifesaver for them, and a great advertisement for you. When the project is over and you are long gone, the only thing a client has to refer back to is the documentation they were left with. Their impression of the project and its execution is largely informed by the final documentation left behind. Every communication and document that you write is an advertisement for yourself and your competency. Don’t miss a chance to leave a lasting, positive impression.
|
||
`,summary:"One of the things that I found most surprising about working in technology, and consulting in particular, is the sheer amount of writing we have to do. When I think about it, on a daily basis I am writing emails, IMs, text messages, documentation, memos, blog entries, or even newsletter articles. You would think that working in technology would be all plugging in cables, configuring software, and running scripts. But instead each day is full of reading and writing.",date:"20 Jul, 2016",url:"https://nedinthecloud.com/2016/07/20/write-it-right/",image:"featured-image-nedinthecloud.jpg",readingTime:"3"}}</script><script src=/js/lunr.min.js></script><script src=/js/search.js></script></footer><script>var slideUp={distance:"100%",origin:"bottom",opacity:null};ScrollReveal().reveal(".slide-up",slideUp),ScrollReveal().reveal(".scrollreveal"),ScrollReveal().reveal(".seq",{interval:100}),ScrollReveal().reveal(".zoom",{scale:.85,duration:800,easing:"ease-in-out"}),ScrollReveal().reveal(".fade-in",{distance:"0px",opacity:0,duration:1500}),ScrollReveal().reveal(".spotlight",{distance:"0px",opacity:.8})</script><script>const swiper=new Swiper(".swiper",{direction:"horizontal",autoplay:!0,loop:!0,speed:1e3,effect:"fade",fadeEffect:{crossFade:!0},autoplay:{delay:8e3,disableOnInteraction:!1},pagination:{el:".swiper-pagination"},navigation:{nextEl:".swiper-button-next",prevEl:".swiper-button-prev"},scrollbar:{el:".swiper-scrollbar"}})</script></body></html> |