Files
nexus/sreweekly/markdown/413/05-sreday-london-uk-sep-19-20-2024.md
2026-09-12 17:23:01 +08:00

16 KiB
Raw Blame History

SREday – London, UK, Sep 19-20, 2024

简介

Check it out, it’s an entire SRE conference I was totally unaware of!

正文

2

Days

50+

Speakers

6

Tracks

250+

Attendees

Coralogix

        Let’s wander down memory lane, and revisit all of the strange things we used to do to keep systems alive. Remember the horrible things we used to do to the tail command? File roll-over errors? Staring in horror at failed SSH connections? Well, I do and now, so will you.



      Pulumi

        Feature flags allow you to enable and disable code without changing or deploying any source code, as well as letting you selectively route traffic to certain users or a percentage of certain users, along with other great tricks. It’s powerful stuff … but when you combine it with observability...



      Port

        While many SREs advocate for increased developer self-service to avoid becoming a bottleneck and to free up time for deeper, more impactful work, implementing effective self-service solutions—and establishing the right guardrails—can be challenging. In this talk, we will explore various...



      RiverSafe

        Navigating and managing an application estate can be overwhelming for Site Reliability Engineers (SREs) going into an organisation. What if you could instantly understand the codebase through visualisations, up-to-date technical document, and a chat experience, and pinpoint exactly where to focus...



      AWS

        This session covers the basics of how to use isolation boundaries to build resilient applications, and dives deep into using chaos engineering and AWS Fault Injection Service (AWS FIS) to test those boundaries.



      Elastic

        To understand how to retain, build and grow SREs, we need to know what not to do. In my time in tech I've been a contributor and manager and have good and bad experiences of managing and being managed. Join me as I share how not to manage SRE teams and the mistakes made in this ecosystem.



      OpenCredo

        Trying to mature your platform engineering team? Don't know where to start, or what pitfalls may be waiting for you. Then this talk is for you. Based on real world experience gained across a variety of clients, and using the CNCF platform maturity model as a guide, we explore & answer such questions



      Cloud Native Computing Foundation

        Discover the truth behind Platform Engineering! Join Horovits as he delves into its evolution, demystifies its purpose, and reveals how it maps to the DevOps maturity model. Gain insights from industry giants and community white papers, and learn when and why your organization needs it.



      GlobalLogic

        Imagine a system designed to process millions of events per day, but as use grows, data mysteriously seems to disappear, only to reappear later. We’ll talk about the what, why and how we investigated the symptoms, overturned some assumptions and finally delivered a resilient, serverless solution



      Container Solutions

        I spent years managing 3rd-line support for some of busiest transactional sites in the world: fixing issues in real-time; writing major incident reports; to managing processes for 50 engineers of varying levels. This talk is about what I learned in the process about what has come to be known as SRE.



      VMware / Pivotal

        We're developer productivity obsessed, worse, developer productivity METRICS obsessed. Deployment frequency! MTTR! eNPS! It's become a delightful distraction. Besides, who really benefits from this focus? This talk makes the case that we should focus more on actual craft and less on metrics.



      Google

        SRE are Google's specialists for designing, building, and running complex services that are reliable, scalable, efficient, and maintainable. The SRE Engagement Model describes how the collaboration between developers and SREs works and how reliability engineering can be applied early in development.



      Cognizant

        All of us have attended lots of talks with the premise “from Zero to Hero…”, but what does it take for a real Rock Star, a real Ninja, a magical Unicorn, a superhero to become a complete Zero? How can we destroy those heroes in a way that every supervillain on every film has dreamt about every day?



      RunWhen

        Inspired by work from hyperscale SRE teams, RunWhen is building the industry's largest open source library of troubleshooting automation. Contributing Authors receive royalties when the company's enterprise customers use their code to automate root cause analysis and remediation.



      Accenture

        Unlock the future of Kubernetes security with GenAI and eBPF. Learn to predict, detect, and neutralize threats in real-time. This session reveals how AI and monitoring transform security strategies for engineers and platforms in fast-paced environments.



      The Cloud Native Club

        Join me as I transition from football to tech, focusing on Kubernetes. Learn how teamwork, discipline, and strategy from sports influenced my tech career. Discover my journey into building developer platforms and the parallels between leading a sports team and orchestrating cloud-native environments



      NTT DATA

        If you need to define a road map to build more reliable products through SRE building ways of working adoption, you need to make an strategy to have a solid why and a road map to reach outcomes iteratively while you follow SRE road map adoption at your company, in this talk we share a real...



      StatusNeo

        The session would focus on the importance of Engineering Portals to accelerate Developer Experience by letting them "Create" Reusable Patterns, "Manage" their software catalog in a centralized inventory, and "Explore" the software ecosystem in a unified Single Pane of Glass.



      Kentik

        When people tell me "alerts suck", I tell them "No, YOUR alerts suck." Here's why: We suck at creating alerts that are useful, meaningful, and most of all actionable. But there's good news: We'll show you how your alerts can suck less, and even be more manageable, using a few easy techniques.



      Unlock the full potential of Kubernetes with Karpenter! Scale effortlessly, optimize efficiently, and upgrade seamlessly. Join my talk to revolutionize your cloud infrastructure journey in just minutes! Don't miss out on this game-changing solution!



      Agile Lab

        Imagine a product made of a number of microservices on Kubernetes and dozens of devs working on them. Now imagine multiple environments, where they must be deployed, tested and demoed at the speed of light. Let's put all of this together in a design that satisfies devs, customers and SRE's nights!



      Loft Labs

        Virtual Kubernetes cluster is a new approach to dealing with Kubernetes multi-tenancy pain. With virtual k8s clusters, each tenant has what feels like a full-blown, dedicated cluster inside a shared host cluster. Let's learn how you can use a virtual cluster to save costs and improve productivity.



      Riskified

        Adding new functionality to Terraform can be daunting: it’s written only in Go (which you may not know) and you have to understand the architecture and work through less than welcoming documentation. I’ll provide a walkthrough from my experience with it, going from zero to publishing a provider.



      Stenn International

        A comprehensive overview of various aspects of building a highly scalable fintech startup with a product-led strategy. In this talk, I will share insights on different problems and solutions for tackling the inevitable growing complexity during the scaling of a fintech startup. He will also...



      Varnish Software

        Learn how to accelerate web applications and APIs by caching the HTTP output in Varnish. Instead of focusing on basic use cases with static content that is easily cacheable, this presentation shows how to cache personalized content "on the edge", that is otherwise deemed uncacheable.



      Viam

        In the dynamic realm of modern infrastructure, challenges such as intricate security protocols and managing diverse environments are common across all technical teams. Infrastructure as Code (IaC) emerges as a transformative force, turning these challenges into opportunities for innovation. For...



      Sidero Labs

        We run declarative systems built on top of an imperative operating system. It’s time we simplified our interface for Linux with a declarative API.



      SRE Author

        Come and explore the landscape of SRE as it is in 2024, with the new trends, techniques and tools on the horizon.



      Splunk

        Are you struggling to make sense of your telemetry data or overwhelmed by your system's sheer volume and complexity? Learn how the OpenTelemetry Collector can help you streamline data collection, optimize performance, and gain unparalleled insights into your applications.



      Kubespaces

        Don’t panic! We will take head-on all the nitty gritty details and unknowns of what to do when things are going south, pods are on fire and nothing seems to work: we’ll give you hints, tools and techniques to face the most egregious of outages and the hidden dragons of creepy memory leaks, with...



      NGINX, part of F5

        Lies, damned lies and statistics. But statistics allow you to lie to yourself. Statistics can trick us into believing things that are less than true, though not on purpose. Learn how data choice, event focus and scale change perspective. See how graphs mislead and correlation can cause confusion.



      Ænix

        How to deploy a reproducible environment on bare metal with Talos Linux. Our experience in developing an open platform designed for on-premise. How we ensure component stability on any hardware.



      Fullstaq

        Perses is to Prometheus as what Grafana was for graphite. And even though Grafana has become so much more, Perses is an extremely interesting new entry into the dashboad/visualisation space. This talk will have an up to date demo and some of the unique advantages Perses offers over alternatives.



      Grafana Labs

        This talk is about Data analysts who mostly work as Data Scientists who need to visualize data mostly available in Google Excel data sets aka data sheets on the cloud and using Grafana can help to visualize and monitor it by using the Excel sheet plugin.



      OVHcloud

        The presentation explores how observability can improve service quality down to the database level. The speaker will discuss how OVHcloud refined the efficiency of their information system databases. The talk shares a recipe for getting the most out of your DBMS and how to get rid of slow queries.



      FanDuel / Blip.pt

        DevOps is not a trend that has come and gone. It's a cultural shift that has fundamentally changed how software is developed and deployed. While emerging practices like Platform Engineering may seem to be taking over, they are not meant to replace DevOps; instead, they are complementary to it.



      Netdata

        In SreCON19, Todd Underwood from Google gave a presentation with the title “All of Our ML Ideas Are Bad (and We Should Feel Bad)”. Let’s see a few ML ideas, implemented in the open-source Netdata, that may not actually be that bad.



      Fanatics

        Why don't you get dedicated product support from the company? Why is it so hard to get a headcount for your platform team? Why is it so hard to drive adoption of the products you create for the dev org? The truth is that Platform Engineering is a company within a company. Run it that way.



      TantusData

        Would you like to do something more than just create prompts for ChatGPT? How about building an application? During the presentation you will get an intuition about what it takes to build the LLM-base application. From writing the code to production deployment. We will cover tricky aspects like...



      ilert

        This talk will be particularly valuable for DevOps engineers looking to optimize their alert management systems and reduce the cognitive load caused by alert fatigue.



      This talk will explore the integration of Generative Artificial Intelligence (AI) to enhance decision-making and diagnosis in Site Reliability Engineering (SRE). Leveraging the capabilities of Generative AI, we propose an Omni-functional approach to synthesizing insights from diverse sources,...



      Zero Trust Security Framework deployed as guardrails to verify every request as if it originates from an open network and anticipates that threats can be both internal and external using Enterprise Solutions like Hashicorp Sentinel and Open-source tools like Kyverno and Open Policy Agent.



      Indeed Flex

        When delegating work to background jobs, developers often manage multiple queues, which can lead to unpredictable problems initially. In this talk, I propose a different approach based on latency tolerance to address common issues encountered with queues.



      Severalnines

        Learn how cloud-native technologies like Kubernetes are revolutionizing database software, enabling scalability, resilience, and agility. Join Divine a Data on Kubernetes Ambassador to explore real-world examples and best practices for building the next generation of database solutions.



      Modus Create

        You will see the advantages of using the k8s provider for terraform instead of helm to manage the k8s objects ("yamls"). It allows you to see and confirm exactly what changes will happen in your cluster. Later we'll see how to integrate that with the AWS ALB and forward traffic directly to pods.



      Have you encountered a project where every single app reads a config file on startup? Have you struggled to find those files or change them? Did you feel that there is something off with this approach? Look no further! We will discuss when flags can solve all of these issues (and when they can't).



      Rawkode Academy

        You don't learn until your back is up against the wall and the clock is ticking down. Join us as we explore 3 broken Kubernetes clusters in an attempt to fix each one within less than 30 minutes.



      Omnistrate

        Sometimes you delete production, you drop tables, you change a conf and everything breaks but what if you upgraded the wrong environment, to the wrong version, of the wrong customer (that also happens to be the bigger your company has)? And what if you did it right before going to a long lunch?...



      Red Hat

        Discover Debezium: Real-time database change streaming made seamless. Say goodbye to vendor lock-in with the innovative Debezium Server. Learn about parallelization and Kubernetes deployment using the Debezium Operator and supercharge your application performance.



      Tessl

        Kubernetes Dynamic Admission control can be used for advanced validation and error correction of workloads. If you've ever wanted to make sure workloads in your clusters are conforming to security best practices, this talk will show you how!



      This talk focuses on the cross-functional impacts and needs of software observability, including how it can benefit roles such as Product Management, Engineering Management, Customer Success, CTO's office, and development. Software observability is essential for quickly identifying faults in a...



      
        Crossrail Place,

        Canary Wharf,

        E14 5AR, London, UK

        Level -2

         

        Tube access

        Jubilee, Elizabeth and DLR lines: Canary Wharf station

Want to become a sponsor? Get in touch!