670 lines
32 KiB
HTML
670 lines
32 KiB
HTML
<!DOCTYPE html>
|
|
<html>
|
|
<head>
|
|
<meta charset="utf-8">
|
|
<meta http-equiv="X-UA-Compatible" content="IE=edge,chrome=1">
|
|
<title>A thorough introduction to bpftrace</title>
|
|
<meta name="viewport" content="width=device-width">
|
|
<meta name="description" content="A thorough introduction and tutorial to bpftrace: a high-level eBPF tracer">
|
|
<meta name="keywords" content="bpf,ebpf,bpftrace,linux,observability,blog">
|
|
|
|
<!-- syntax highlighting CSS -->
|
|
<link rel="stylesheet" href="/blog/css/syntax.css">
|
|
|
|
<!-- Custom CSS -->
|
|
<link rel="stylesheet" href="/blog/css/main.css">
|
|
|
|
<!-- Google tag (gtag.js) -->
|
|
<script async src="https://www.googletagmanager.com/gtag/js?id=G-SYD11KXST1"></script>
|
|
<script>
|
|
window.dataLayer = window.dataLayer || [];
|
|
function gtag(){dataLayer.push(arguments);}
|
|
gtag('js', new Date());
|
|
|
|
gtag('config', 'G-SYD11KXST1');
|
|
</script>
|
|
|
|
</head>
|
|
<body>
|
|
|
|
<div class="page">
|
|
<div class="nav">
|
|
<p class="navhdr">Brendan's site:</p>
|
|
<a href="/overview.html">Start Here</a><br>
|
|
<a href="/index.html">Homepage</a><br>
|
|
<a href="/blog/index.html">Blog</a><br>
|
|
<!-- <a href="/sitemap.html">Full Site Map</a><br> -->
|
|
<a href="/systems-performance-2nd-edition-book.html">Sys Perf book</a><br>
|
|
<a href="/bpf-performance-tools-book.html">BPF Perf book</a><br>
|
|
<a href="/linuxperf.html">Linux Perf</a><br>
|
|
<a href="/ebpf.html">eBPF Tools</a><br>
|
|
<a href="/perf.html">perf Examples</a><br>
|
|
<a href="/methodology.html">Perf Methods</a><br>
|
|
<a href="/usemethod.html">USE Method</a><br>
|
|
<a href="/tsamethod.html">TSA Method</a><br>
|
|
<a href="/offcpuanalysis.html">Off-CPU Analysis</a><br>
|
|
<a href="/activebenchmarking.html">Active Bench.</a><br>
|
|
<a href="/wss.html">WSS Estimation</a><br>
|
|
<a href="/flamegraphs.html">Flame Graphs</a><br>
|
|
<a href="/flamescope.html">Flame Scope</a><br>
|
|
<a href="/heatmaps.html">Heat Maps</a><br>
|
|
<a href="/frequencytrails.html">Frequency Trails</a><br>
|
|
<a href="/colonygraphs.html">Colony Graphs</a><br>
|
|
<a href="/dtrace.html">DTrace Tools</a><br>
|
|
<a href="/dtracetoolkit.html">DTraceToolkit</a><br>
|
|
<a href="/dtkshdemos.html">DtkshDemos</a><br>
|
|
<a href="/guessinggame.html">Guessing Game</a><br>
|
|
<a href="/specials.html">Specials</a><br>
|
|
<a href="/books.html">Books</a><br>
|
|
<a href="/sites.html">Other Sites</a><br>
|
|
|
|
</div>
|
|
|
|
<div class="recent">
|
|
<!-- (this is for the blog) recent books: -->
|
|
<!-- <center><a href="https://informit.com/sale/booksgiving"><img border=0 width=180 src="/Images/booksgiving2021.jpg"></center><br><b>Book sale until Dec 1, 2021: 55% off for 2 or more</b></a><br><br> -->
|
|
<center><a href="/systems-performance-2nd-edition-book.html"><img src="/Images/sysperf2nd_bookcover_360.jpg" width=180></a><br><font size=-2><i><a href="/systems-performance-2nd-edition-book.html">Systems Performance 2nd Ed.</a></i></font></center><br><br>
|
|
|
|
<center><a href="/bpf-performance-tools-book.html"><img src="/Images/bpfperftools_bookcover_360.jpg" width=180></a><br><font size=-2><i><a href="/bpf-performance-tools-book.html">BPF Performance Tools book</a></i></font></center>
|
|
<!--
|
|
<br><center><a href="https://www.portal.reinvent.awsevents.com/connect/search.ww?#loadSearch-searchPhrase=OPN303&searchType=session&tc=0&sortBy=abbreviationSort&p="><img src="/Images/Speaker/reInvent2019_200.jpg" width=180" border=0></a><br><font size=-2><i>I'm speaking at <a href="https://www.portal.reinvent.awsevents.com/connect/search.ww?#loadSearch-searchPhrase=OPN303&searchType=session&tc=0&sortBy=abbreviationSort&p=">AWS re:Invent 2019</a></i></font></center>
|
|
-->
|
|
<br>
|
|
Recent posts:<br>
|
|
<ul style="padding-left:18px">
|
|
|
|
<li>07 Feb 2026 »<br>
|
|
<a href="/blog/2026-02-07/why-i-joined-openai.html">
|
|
Why I joined OpenAI</a></li>
|
|
|
|
<li>05 Dec 2025 »<br>
|
|
<a href="/blog/2025-12-05/leaving-intel.html">
|
|
Leaving Intel</a></li>
|
|
|
|
<li>28 Nov 2025 »<br>
|
|
<a href="/blog/2025-11-28/ai-virtual-brendans.html">
|
|
On "AI Brendans" or "Virtual Brendans"</a></li>
|
|
|
|
<li>22 Nov 2025 »<br>
|
|
<a href="/blog/2025-11-22/intel-is-listening.html">
|
|
Intel is listening, don't waste your shot</a></li>
|
|
|
|
<li>17 Nov 2025 »<br>
|
|
<a href="/blog/2025-11-17/third-stage-engineering.html">
|
|
Third Stage Engineering</a></li>
|
|
|
|
<li>04 Aug 2025 »<br>
|
|
<a href="/blog/2025-08-04/when-to-hire-a-computer-performance-engineering-team-2025-part1.html">
|
|
When to Hire a Computer Performance Engineering Team (2025) part 1 of 2</a></li>
|
|
|
|
<li>22 May 2025 »<br>
|
|
<a href="/blog/2025-05-22/3-years-of-extremely-remote-work.html">
|
|
3 Years of Extremely Remote Work</a></li>
|
|
|
|
<li>01 May 2025 »<br>
|
|
<a href="/blog/2025-05-01/doom-gpu-flame-graphs.html">
|
|
Doom GPU Flame Graphs</a></li>
|
|
|
|
<li>29 Oct 2024 »<br>
|
|
<a href="/blog/2024-10-29/ai-flame-graphs.html">
|
|
AI Flame Graphs</a></li>
|
|
|
|
<li>22 Jul 2024 »<br>
|
|
<a href="/blog/2024-07-22/no-more-blue-fridays.html">
|
|
No More Blue Fridays</a></li>
|
|
|
|
<li>24 Mar 2024 »<br>
|
|
<a href="/blog/2024-03-24/linux-crisis-tools.html">
|
|
Linux Crisis Tools</a></li>
|
|
|
|
<li>17 Mar 2024 »<br>
|
|
<a href="/blog/2024-03-17/the-return-of-the-frame-pointers.html">
|
|
The Return of the Frame Pointers</a></li>
|
|
|
|
<li>10 Mar 2024 »<br>
|
|
<a href="/blog/2024-03-10/ebpf-documentary.html">
|
|
eBPF Documentary</a></li>
|
|
|
|
<li>28 Apr 2023 »<br>
|
|
<a href="/blog/2023-04-28/ebpf-security-issues.html">
|
|
eBPF Observability Tools Are Not Security Tools</a></li>
|
|
|
|
<li>01 Mar 2023 »<br>
|
|
<a href="/blog/2023-03-01/computer-performance-future-2022.html">
|
|
USENIX SREcon APAC 2022: Computing Performance: What's on the Horizon</a></li>
|
|
|
|
<li>17 Feb 2023 »<br>
|
|
<a href="/blog/2023-02-17/srecon-apac-2023.html">
|
|
USENIX SREcon APAC 2023: CFP</a></li>
|
|
|
|
<li>02 May 2022 »<br>
|
|
<a href="/blog/2022-05-02/brendan-at-intel.html">
|
|
Brendan@Intel.com</a></li>
|
|
|
|
<li>15 Apr 2022 »<br>
|
|
<a href="/blog/2022-04-15/netflix-farewell-1.html">
|
|
Netflix End of Series 1</a></li>
|
|
|
|
<li>09 Apr 2022 »<br>
|
|
<a href="/blog/2022-04-09/tensorflow-library-performance.html">
|
|
TensorFlow Library Performance</a></li>
|
|
|
|
<li>19 Mar 2022 »<br>
|
|
<a href="/blog/2022-03-19/why-dont-you-use.html">
|
|
Why Don't You Use ...</a></li>
|
|
|
|
</ul>
|
|
<a href="/blog/index.html">Blog index</a><br>
|
|
<a href="/blog/about.html">About</a><br>
|
|
<a href="/blog/rss.xml">RSS</a><br>
|
|
<!--
|
|
<br><center><a href="https://www.usenix.org/conference/lisa18"><img src="https://www.usenix.org/sites/default/files/lisa18_banner_join-me.png" width=180></a><br><font size=-2><i>I am program co-chair for LISA 2018</i></font></center>
|
|
-->
|
|
</div>
|
|
|
|
<div class="site">
|
|
<div class="header">
|
|
<h1 class="title"><a href="/blog/index.html">Brendan Gregg's Blog</a></h1>
|
|
<a class="extra" href="/blog/index.html">home</a>
|
|
</div>
|
|
|
|
<h2 class="big">A thorough introduction to bpftrace</h2>
|
|
<p class="meta">19 Aug 2019</p>
|
|
|
|
<div class="post">
|
|
<p><i>Originally posted at https:/opensource.com/article/19/8/introduction-bpftrace.</i></p>
|
|
|
|
<p>bpftrace is a new open source tracer for Linux for analyzing production performance problems and troubleshooting software. It is used by and has had contributions from many companies including Netfilx, Facebook, Red Hat, Shopify, and others. It was created by Alastair Robertson, a talented UK-based developer who has previously won various coding competitions.</p>
|
|
|
|
<p>Linux already has many performance tools, but these are often counter-based and have limited visibility. For example, iostat(1), or a monitoring agent, may tell you your average disk latency, but not the distribution of this latency. Distributions can reveal multiple modes, or outliers, either of which may be the real cause of your performance problems. bpftrace is suited for this kind of analysis: decomposing metrics into distributions or per-event logs, and creating new metrics for visibility into blind spots.</p>
|
|
|
|
<p>You can use bpftrace via one-liners or scripts, and it ships with many prewritten tools. Here is an example screenshot: tracing the distribution of read latency for PID 181, and showing it as a power-of-two histogram:</p>
|
|
|
|
<pre class="narrow">
|
|
# <b>bpftrace -e 'kprobe:vfs_read /pid == 30153/ { @start[tid] = nsecs; }
|
|
kretprobe:vfs_read /@start[tid]/ { @ns = hist(nsecs - @start[tid]); delete(@start[tid]); }'</b>
|
|
Attaching 2 probes...
|
|
^C
|
|
|
|
@ns:
|
|
[256, 512) 10900 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@ |
|
|
[512, 1k) 18291 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@|
|
|
[1k, 2k) 4998 |@@@@@@@@@@@@@@ |
|
|
[2k, 4k) 57 | |
|
|
[4k, 8k) 117 | |
|
|
[8k, 16k) 48 | |
|
|
[16k, 32k) 109 | |
|
|
[32k, 64k) 3 | |
|
|
</pre>
|
|
|
|
<p>This example instrumented one of many thousands of available events. If you have some weird performance problem, there's probably some bpftrace one-liner that can shed light on it. For large environments, this ability can help you save millions. For smaller environments, it can be of more use helping eliminate latency outliers.</p>
|
|
|
|
<p>I <a href="http://www.brendangregg.com/blog/2018-10-08/dtrace-for-linux-2018.html">previously</a> wrote about bpftrace vs other tracers, including <a href="https://github.com/iovisor/bcc">BCC</a> (BPF Compiler Collection). BCC is great for canned complex tools and agents. bpftrace is best for short scripts and ad hoc investigations. In this post I'll summarize the bpftrace language, variable types, probes, and tools.</p>
|
|
|
|
<p>bpftrace uses BPF (Berkeley Packet Filter), an in-kernel execution engine that processes a virtual instruction set. BPF has been extended (aka eBPF) in recent years for providing a safe way to extend kernel functionality, and has become a hot topic in systems engineering, with at least 24 talks on BPF at the last Linux Plumber's conference. BPF is in the Linux kernel, and bpftrace is the best way to get started using BPF for observability.</p>
|
|
|
|
<p>See the bpftrace <a href="https://github.com/iovisor/bpftrace/blob/master/INSTALL.md">INSTALL</a> guide for how to install it, and get the latest version: <a href="https://github.com/iovisor/bpftrace/releases/tag/v0.9.1">0.9.1</a> was just released. For Kubernetes clusters, there is also <a href="https://github.com/iovisor/kubectl-trace">kubectl-trace</a> for running it.</p>
|
|
|
|
<h2>Syntax</h2>
|
|
|
|
<pre class="narrow">probe[,probe,...] /filter/ { action }</pre>
|
|
|
|
<p>The probe specifies what events to instrument, the filter is optional and can filter down the events based on a boolean expression, and the action is the mini program that runs.</p>
|
|
|
|
<p>Here's hello world:</p>
|
|
|
|
<pre class="narrow">
|
|
# <b>bpftrace -e 'BEGIN { printf("Hello eBPF!\n"); }'</b>
|
|
</pre>
|
|
|
|
<p>The probe is <tt>BEGIN</tt>, a special probe that runs at the beginning of the program (like awk). There's no filter. The action is a <tt>printf()</tt> statement.</p>
|
|
|
|
<p>Now a real example:</p>
|
|
|
|
<pre class="narrow">
|
|
# <b>bpftrace -e 'kretprobe:sys_read /pid == 181/ { @bytes = hist(retval); }'</b>
|
|
</pre>
|
|
|
|
<p>This uses a kretprobe to instrument the return of the sys_read() kernel function. If the PID is 181, a special map variable <tt>@bytes</tt> is populated with a log2 histogram function with the return value <tt>retval</tt> of sys_read(). This produces a histogram of the returned read size for PID 181. Is your app doing lots of 1 byte reads? Maybe that can be optimized.</p>
|
|
|
|
<h2>Probe Types</h2>
|
|
|
|
<p>These are libraries of probes which are related. The currently supported types are (more will be added):</p>
|
|
|
|
<ul><table border=1>
|
|
<tr><th>Type</th><th>Description</th></tr>
|
|
<tr><td>tracepoint</td><td>Kernel static instrumentation points</td></tr>
|
|
<tr><td>usdt</td><td>User-level statically defined tracing</td></tr>
|
|
<tr><td>kprobe</td><td>Kernel dynamic function instrumentation</td></tr>
|
|
<tr><td>kretprobe</td><td>Kernel dynamic function return instrumentation</td></tr>
|
|
<tr><td>uprobe</td><td>User-level dynamic function instrumentation</td></tr>
|
|
<tr><td>uretprobe</td><td>User-level dynamic function return instrumentation</td></tr>
|
|
<tr><td>software</td><td>Kernel software-based events</td></tr>
|
|
<tr><td>hardware</td><td>Hardware counter-based instrumentation</td></tr>
|
|
<tr><td>watchpoint</td><td>Memory watchpoint events (in development)</td></tr>
|
|
<tr><td>profile</td><td>Timed sampling across all CPUs</td></tr>
|
|
<tr><td>interval</td><td>Timed reporting (from one CPU)</td></tr>
|
|
<tr><td>BEGIN</td><td>Start of bpftrace</td></tr>
|
|
<tr><td>END</td><td>End of bpftrace</td></tr>
|
|
</table></ul>
|
|
|
|
<p>If you're new to this terminology: dynamic instrumentation (aka dynamic tracing) is the superpower that lets you trace any software function in a running binary without restarting it. This lets you get to the bottom of just about any problem. However, the functions it exposes are not considered a stable API, as they can change from one software version to another. Hence static instrumentation, where event points are hard-coded and become a stable API. When you write bpftrace programs, try to use the static types first, before the dynamic ones, so that your programs are more stable.</p>
|
|
|
|
<h2>Variable Types</h2>
|
|
|
|
<ul><table border=1>
|
|
<tr><th>Variable</th><th>Description</th></tr>
|
|
<tr><td><tt>@name</tt></td><td>global</td></tr>
|
|
<tr><td><tt>@name[key]</tt></td><td>hash</td></tr>
|
|
<tr><td><tt>@name[tid]</tt></td><td>thread-local</td></tr>
|
|
<tr><td><tt>$name</tt></td><td>scratch</td></tr>
|
|
</table></ul>
|
|
|
|
<p>Variables with a '@' prefix use BPF maps, which can behave like associative arrays. They can be populated in one of two ways:</p>
|
|
|
|
<ul>
|
|
<li>variable assignment: <tt>@name = x;</tt></li>
|
|
<li>function assignment: <tt>@name = hist(x);</tt></li>
|
|
</ul>
|
|
|
|
<p>There are various map-populating functions as builtins that provide quick ways to summarize data.</p>
|
|
|
|
<h2>Builtins Variables and Functions</h2>
|
|
|
|
<p>I'll mention some here to help introduce bpftrace, but there are many more.</p>
|
|
|
|
<p>Builtin variables include:</p>
|
|
|
|
<ul><table border=1>
|
|
<tr><th>Variable</th><th>Description</th></tr>
|
|
<tr><td><tt>pid</tt></td><td>process ID</td></tr>
|
|
<tr><td><tt>comm</tt></td><td>Process or command name</td></tr>
|
|
<tr><td><tt>nsecs</tt></td><td>Current time in nanoseconds</td></tr>
|
|
<tr><td><tt>kstack</tt></td><td>Kernel stack trace</td></tr>
|
|
<tr><td><tt>ustack</tt></td><td>User-level stack trace</td></tr>
|
|
<tr><td><tt>arg0...argN</tt></td><td>Function arguments</td></tr>
|
|
<tr><td><tt>args</tt></td><td>Tracepoint arguments</td></tr>
|
|
<tr><td><tt>retval</tt></td><td>Function return value</td></tr>
|
|
<tr><td><tt>name</tt></td><td>Full probe name</td></tr>
|
|
</table></ul>
|
|
|
|
<p>Builtin functions include:</p>
|
|
|
|
<ul><table border=1>
|
|
<tr><th>Function</th><th>Description</th></tr>
|
|
<tr><td><tt>printf("...")</tt></td><td>Print formatted string</td></tr>
|
|
<tr><td><tt>time("...")</tt></td><td>Print formatted time</td></tr>
|
|
<tr><td><tt>system("...")</tt></td><td>Run shell command</td></tr>
|
|
<tr><td><tt>@ = count()</tt></td><td>Count events</td></tr>
|
|
<tr><td><tt>@ = hist(x)</tt></td><td>Power-of-2 histogram for x</td></tr>
|
|
<tr><td><tt>@ = lhist(x, min, max, step)</tt></td><td>Linear histogram for x</td></tr>
|
|
</table></ul>
|
|
|
|
<p>See the <a href="https://github.com/iovisor/bpftrace/blob/master/docs/reference_guide.md">reference guide</a> for everything.</p>
|
|
|
|
<h2>One-Liners Tutorial</h2>
|
|
|
|
<p>A great way to learn bpftrace is via one-liners, which I turned into the <a href="https://github.com/iovisor/bpftrace/blob/master/docs/tutorial_one_liners.md">one-liners tutorial</a>. That tutorial covers the following one-liners:</p>
|
|
|
|
<pre>
|
|
1. Listing probes
|
|
<b>bpftrace -l 'tracepoint:syscalls:sys_enter_*'</b>
|
|
|
|
2. Hello world
|
|
<b>bpftrace -e 'BEGIN { printf("hello world\n") }'</b>
|
|
|
|
3. File opens
|
|
<b>bpftrace -e 'tracepoint:syscalls:sys_enter_open { printf("%s %s\n", comm, str(args->filename)) }'</b>
|
|
|
|
4. Syscall counts by process
|
|
<b>bpftrace -e 'tracepoint:raw_syscalls:sys_enter { @[comm] = count() }'</b>
|
|
|
|
5. Distribution of read() bytes
|
|
<b>bpftrace -e 'tracepoint:syscalls:sys_exit_read /pid == 18644/ { @bytes = hist(args->retval) }'</b>
|
|
|
|
6. Kernel dynamic tracing of read() bytes
|
|
<b>bpftrace -e 'kretprobe:vfs_read { @bytes = lhist(retval, 0, 2000, 200) }'</b>
|
|
|
|
7. Timing read()s
|
|
<b>bpftrace -e 'kprobe:vfs_read { @start[tid] = nsecs }
|
|
kretprobe:vfs_read /@start[tid]/ { @ns[comm] = hist(nsecs - @start[tid]); delete(@start[tid]) }'</b>
|
|
|
|
8. Count process-level events
|
|
<b>bpftrace -e 'tracepoint:sched:sched* { @[name] = count() } interval:s:5 { exit() }'</b>
|
|
|
|
9. Profile on-CPU kernel stacks
|
|
<b>bpftrace -e 'profile:hz:99 { @[stack] = count() }'</b>
|
|
|
|
10. Scheduler tracing
|
|
<b>bpftrace -e 'tracepoint:sched:sched_switch { @[stack] = count() }'</b>
|
|
|
|
11. Block I/O tracing
|
|
<b>bpftrace -e 'tracepoint:block:block_rq_complete { @ = hist(args->nr_sector * 512) }'</b>
|
|
</pre>
|
|
|
|
<p>See the tutorial for an explanation of each of these.</p>
|
|
|
|
<h2>Provided Tools</h2>
|
|
|
|
<p>Apart from one-liners, bpftrace programs can be multi-line scripts. bpftrace ships with 28 of these as tools:</p>
|
|
|
|
<p><center><img src="/blog/images/2019/bpftrace_tools_early2019.png" width=640 border=0></center></p>
|
|
|
|
<p>These tools can be found in the /tools directory:</p>
|
|
|
|
<pre>
|
|
tools# <b>ls *.bt</b>
|
|
bashreadline.bt dcsnoop.bt oomkill.bt syncsnoop.bt vfscount.bt
|
|
biolatency.bt execsnoop.bt opensnoop.bt syscount.bt vfsstat.bt
|
|
biosnoop.bt gethostlatency.bt pidpersec.bt tcpaccept.bt writeback.bt
|
|
bitesize.bt killsnoop.bt runqlat.bt tcpconnect.bt xfsdist.bt
|
|
capable.bt loads.bt runqlen.bt tcpdrop.bt
|
|
cpuwalk.bt mdflush.bt statsnoop.bt tcpretrans.bt
|
|
</pre>
|
|
|
|
<p>Apart from their use with diagnosing performance issues and general troubleshooting, they also provide another way to learn bpftrace: by example.</p>
|
|
|
|
<h3>Source</h3>
|
|
|
|
<p>Here's the code to biolatency.bt:</p>
|
|
|
|
<pre>
|
|
tools# <b>cat -n biolatency.bt</b>
|
|
1 /*
|
|
2 * biolatency.bt Block I/O latency as a histogram.
|
|
3 * For Linux, uses bpftrace, eBPF.
|
|
4 *
|
|
5 * This is a bpftrace version of the bcc tool of the same name.
|
|
6 *
|
|
7 * Copyright 2018 Netflix, Inc.
|
|
8 * Licensed under the Apache License, Version 2.0 (the "License")
|
|
9 *
|
|
10 * 13-Sep-2018 Brendan Gregg Created this.
|
|
11 */
|
|
12
|
|
13 BEGIN
|
|
14 {
|
|
15 printf("Tracing block device I/O... Hit Ctrl-C to end.\n");
|
|
16 }
|
|
17
|
|
18 kprobe:blk_account_io_start
|
|
19 {
|
|
20 @start[arg0] = nsecs;
|
|
21 }
|
|
22
|
|
23 kprobe:blk_account_io_done
|
|
24 /@start[arg0]/
|
|
25
|
|
26 {
|
|
27 @usecs = hist((nsecs - @start[arg0]) / 1000);
|
|
28 delete(@start[arg0]);
|
|
29 }
|
|
30
|
|
31 END
|
|
32 {
|
|
33 clear(@start);
|
|
34 }
|
|
</pre>
|
|
|
|
<p>It's straightforward and easy to read, and short enough to include on a slide. This version is using kernel dynamic tracing to instrument the blk_account_io_start() and blk_account_io_done() functions, and passes a timestamp between them keyed on arg0 to each. arg0 on kprobe is the first argument to that function, which for these is the struct request *, and its memory address is used as a unique identifier.</p>
|
|
|
|
<h3>Examples</h3>
|
|
|
|
<p>Screenshots from these tools is also provided in *_example.txt files, where I explain exactly what we are seeing. For example:</p>
|
|
|
|
<pre>
|
|
tools# <b>more biolatency_example.txt</b>
|
|
Demonstrations of biolatency, the Linux BPF/bpftrace version.
|
|
|
|
|
|
This traces block I/O, and shows latency as a power-of-2 histogram. For example:
|
|
|
|
# biolatency.bt
|
|
Attaching 3 probes...
|
|
Tracing block device I/O... Hit Ctrl-C to end.
|
|
^C
|
|
|
|
@usecs:
|
|
[256, 512) 2 | |
|
|
[512, 1K) 10 |@ |
|
|
[1K, 2K) 426 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@|
|
|
[2K, 4K) 230 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@ |
|
|
[4K, 8K) 9 |@ |
|
|
[8K, 16K) 128 |@@@@@@@@@@@@@@@ |
|
|
[16K, 32K) 68 |@@@@@@@@ |
|
|
[32K, 64K) 0 | |
|
|
[64K, 128K) 0 | |
|
|
[128K, 256K) 10 |@ |
|
|
|
|
While tracing, this shows that 426 block I/O had a latency of between 1K and 2K
|
|
usecs (1024 and 2048 microseconds), which is between 1 and 2 milliseconds.
|
|
There are also two modes visible, one between 1 and 2 milliseconds, and another
|
|
between 8 and 16 milliseconds: this sounds like cache hits and cache misses.
|
|
There were also 10 I/O with latency 128 to 256 ms: outliers. Other tools and
|
|
instrumentation, like biosnoop.bt, can shed more light on those outliers.
|
|
[...]
|
|
</pre>
|
|
|
|
<p>Sometimes it can be most effective to switch straight to the example file when trying to understand these tools, since the output may be self evident (by design!).</p>
|
|
|
|
<h3>Man pages</h3>
|
|
|
|
<p>There are also man pages for every tool, under /man/man8. They include sections on the output fields, and expected overhead of the tool.</p>
|
|
|
|
<pre>
|
|
# <b>nroff -man man/man8/biolatency.8</b>
|
|
biolatency(8) System Manager's Manual biolatency(8)
|
|
|
|
|
|
|
|
NAME
|
|
biolatency.bt - Block I/O latency as a histogram. Uses bpftrace/eBPF.
|
|
|
|
SYNOPSIS
|
|
biolatency.bt
|
|
|
|
DESCRIPTION
|
|
This tool summarizes time (latency) spent in block device I/O (disk
|
|
I/O) as a power-of-2 histogram. This allows the distribution to be
|
|
studied, including modes and outliers. There are often two modes, one
|
|
for device cache hits and one for cache misses, which can be shown by
|
|
this tool. Latency outliers will also be shown.
|
|
[...]
|
|
</pre>
|
|
|
|
<p>Writing all these man pages was the least fun part of developing these tools, and in some cases tool longer to write than the tool took to develop, but it's nice to see the final result.</p>
|
|
|
|
<h2>bpftrace vs BCC</h2>
|
|
|
|
<p>Since eBPF has been merging in the kernel, most effort has been on the <a href="https://github.com/iovisor/bcc">BCC</a> front-end, which provides a BPF library and Python, C++, and lua interfaces for writing programs. I've developed a lot of <a href="https://github.com/iovisor/bcc#tools">tools</a> in BCC/python, and it works great, although coding in BCC is verbose. If you're hacking away at a performance issue, bpftrace is better for all the one-off custom queries you have. If you're writing a tool with many command line options, or an agent that uses Python libraries, you'll want to consider using BCC.</p>
|
|
|
|
<p>Here's how these will be used at Netflix: on the performance team, I use both: BCC for developing canned tools that others can easily use, and for developing agents; and bpftrace for ad hoc analysis. On, say, the network engineering team, they have been using BCC to develop an agent for their needs. The security team are most interested in bpftrace for quick ad hoc instrumentation for detecting zero day vulnerabilities. And for the developer teams: I expect they'll use both without knowing it via the self-service GUIs we are building (Vector), and occasionally may ssh onto an instance and run a canned tool or ad hoc bpftrace one-liner.</p>
|
|
|
|
<h2>More bpftrace reading</h2>
|
|
|
|
<ul>
|
|
<li>The <a href="https://github.com/iovisor/bpftrace">bpftrace</a> repository on github</li>
|
|
<li>The bpftrace <a href="https://github.com/iovisor/bpftrace/blob/master/docs/tutorial_one_liners.md">one-liners tutorial</a></li>
|
|
<li>The bpftrace <a href="https://github.com/iovisor/bpftrace/blob/master/docs/reference_guide.md">reference guide</a></li>
|
|
<li>The <a href="https://github.com/iovisor/bcc">BCC</a> repository for more complex BPF-based tools</li>
|
|
</ul>
|
|
|
|
<p>I also have a book coming out this year that covers bpftrace: <a href="http://www.brendangregg.com/blog/2019-07-15/bpf-performance-tools-book.html">BPF Performance Tools: Linux System and Application Observability</a>, to be published by Addison Wesley, and which contains many new bpftrace tools.</p>
|
|
|
|
<p><em>Thanks to Alastair Robertson for creating bpftrace, and the bpftrace, BCC, and BPF communities for all the work over the past five years.</em></p>
|
|
|
|
</div>
|
|
|
|
|
|
|
|
<br><hr>
|
|
<script type="text/javascript">
|
|
var disqus_shortname = 'brendangregg';
|
|
|
|
function loadshowdisqus(id) {
|
|
// from disqus:
|
|
var dsq = document.createElement('script'); dsq.type = 'text/javascript'; dsq.async = true;
|
|
dsq.src = '//' + disqus_shortname + '.disqus.com/embed.js';
|
|
(document.getElementsByTagName('head')[0] || document.getElementsByTagName('body')[0]).appendChild(dsq);
|
|
// mine:
|
|
var c = document.getElementById(id); c.style.display = (c.style.display == "none" ? "block" : "none");
|
|
}
|
|
</script>
|
|
<div id="comments_button" onclick='loadshowdisqus("comments_content")'><p style="color:blue"><u>Click here for Disqus comments (ad supported).</u></p></div>
|
|
<div id="comments_content" style="display:none">
|
|
<p><font size=-1><i>You are welcome to comment here, but I've been meaning to switch comment systems one day and I don't know yet if I can preserve existing comments (I'll try to find a way).</i></font></p>
|
|
<div id="disqus_thread"></div>
|
|
<noscript>Please enable JavaScript to view the <a href="http://disqus.com/?ref_noscript">comments powered by Disqus.</a></noscript>
|
|
<a href="http://disqus.com" class="dsq-brlink">comments powered by <span class="logo-disqus">Disqus</span></a>
|
|
</div>
|
|
|
|
|
|
<div class="recentmobile">
|
|
<hr><center>Site Navigation</center><br>
|
|
<!-- (this is for /index.html etc) recent books: -->
|
|
<!-- <center><a href="https://informit.com/sale/booksgiving"><img border=0 width=180 src="/Images/booksgiving2021.jpg"></center><br><b>Book sale until Dec 1, 2021: 55% off for 2 or more</b></a><br><br> -->
|
|
<center><a href="/systems-performance-2nd-edition-book.html"><img src="/Images/sysperf2nd_bookcover_360.jpg" width=180></a><br><font size=-2><i><a href="/systems-performance-2nd-edition-book.html">Systems Performance 2nd Ed.</a></i></font></center><br><br>
|
|
|
|
<center><a href="/bpf-performance-tools-book.html"><img src="/Images/bpfperftools_bookcover_360.jpg" width=180></a><br><font size=-2><i><a href="/bpf-performance-tools-book.html">BPF Performance Tools book</a></i></font></center>
|
|
<!--
|
|
<br><center><a href="https://www.portal.reinvent.awsevents.com/connect/search.ww?#loadSearch-searchPhrase=OPN303&searchType=session&tc=0&sortBy=abbreviationSort&p="><img src="/Images/Speaker/reInvent2019_200.jpg" width=180" border=0></a><br><font size=-2><i>I'm speaking at <a href="https://www.portal.reinvent.awsevents.com/connect/search.ww?#loadSearch-searchPhrase=OPN303&searchType=session&tc=0&sortBy=abbreviationSort&p=">AWS re:Invent 2019</a></i></font></center>
|
|
-->
|
|
<br>
|
|
|
|
Recent posts:<br>
|
|
<ul style="padding-left:18px">
|
|
|
|
<li class="recent">07 Feb 2026 »<br>
|
|
<a href="/blog/2026-02-07/why-i-joined-openai.html">
|
|
Why I joined OpenAI</a></li>
|
|
|
|
<li class="recent">05 Dec 2025 »<br>
|
|
<a href="/blog/2025-12-05/leaving-intel.html">
|
|
Leaving Intel</a></li>
|
|
|
|
<li class="recent">28 Nov 2025 »<br>
|
|
<a href="/blog/2025-11-28/ai-virtual-brendans.html">
|
|
On "AI Brendans" or "Virtual Brendans"</a></li>
|
|
|
|
<li class="recent">22 Nov 2025 »<br>
|
|
<a href="/blog/2025-11-22/intel-is-listening.html">
|
|
Intel is listening, don't waste your shot</a></li>
|
|
|
|
<li class="recent">17 Nov 2025 »<br>
|
|
<a href="/blog/2025-11-17/third-stage-engineering.html">
|
|
Third Stage Engineering</a></li>
|
|
|
|
<li class="recent">04 Aug 2025 »<br>
|
|
<a href="/blog/2025-08-04/when-to-hire-a-computer-performance-engineering-team-2025-part1.html">
|
|
When to Hire a Computer Performance Engineering Team (2025) part 1 of 2</a></li>
|
|
|
|
<li class="recent">22 May 2025 »<br>
|
|
<a href="/blog/2025-05-22/3-years-of-extremely-remote-work.html">
|
|
3 Years of Extremely Remote Work</a></li>
|
|
|
|
<li class="recent">01 May 2025 »<br>
|
|
<a href="/blog/2025-05-01/doom-gpu-flame-graphs.html">
|
|
Doom GPU Flame Graphs</a></li>
|
|
|
|
<li class="recent">29 Oct 2024 »<br>
|
|
<a href="/blog/2024-10-29/ai-flame-graphs.html">
|
|
AI Flame Graphs</a></li>
|
|
|
|
<li class="recent">22 Jul 2024 »<br>
|
|
<a href="/blog/2024-07-22/no-more-blue-fridays.html">
|
|
No More Blue Fridays</a></li>
|
|
|
|
<li class="recent">24 Mar 2024 »<br>
|
|
<a href="/blog/2024-03-24/linux-crisis-tools.html">
|
|
Linux Crisis Tools</a></li>
|
|
|
|
<li class="recent">17 Mar 2024 »<br>
|
|
<a href="/blog/2024-03-17/the-return-of-the-frame-pointers.html">
|
|
The Return of the Frame Pointers</a></li>
|
|
|
|
<li class="recent">10 Mar 2024 »<br>
|
|
<a href="/blog/2024-03-10/ebpf-documentary.html">
|
|
eBPF Documentary</a></li>
|
|
|
|
<li class="recent">28 Apr 2023 »<br>
|
|
<a href="/blog/2023-04-28/ebpf-security-issues.html">
|
|
eBPF Observability Tools Are Not Security Tools</a></li>
|
|
|
|
<li class="recent">01 Mar 2023 »<br>
|
|
<a href="/blog/2023-03-01/computer-performance-future-2022.html">
|
|
USENIX SREcon APAC 2022: Computing Performance: What's on the Horizon</a></li>
|
|
|
|
<li class="recent">17 Feb 2023 »<br>
|
|
<a href="/blog/2023-02-17/srecon-apac-2023.html">
|
|
USENIX SREcon APAC 2023: CFP</a></li>
|
|
|
|
<li class="recent">02 May 2022 »<br>
|
|
<a href="/blog/2022-05-02/brendan-at-intel.html">
|
|
Brendan@Intel.com</a></li>
|
|
|
|
<li class="recent">15 Apr 2022 »<br>
|
|
<a href="/blog/2022-04-15/netflix-farewell-1.html">
|
|
Netflix End of Series 1</a></li>
|
|
|
|
<li class="recent">09 Apr 2022 »<br>
|
|
<a href="/blog/2022-04-09/tensorflow-library-performance.html">
|
|
TensorFlow Library Performance</a></li>
|
|
|
|
<li class="recent">19 Mar 2022 »<br>
|
|
<a href="/blog/2022-03-19/why-dont-you-use.html">
|
|
Why Don't You Use ...</a></li>
|
|
|
|
</ul>
|
|
<a href="/blog/index.html">Blog index</a><br>
|
|
<a href="/blog/about.html">About</a><br>
|
|
<a href="/blog/rss.xml">RSS</a><br>
|
|
<!-- also edit _layouts/default.html -->
|
|
<!--
|
|
<br><a href="https://www.usenix.org/conference/lisa18"><img src="https://www.usenix.org/sites/default/files/lisa18_banner_join-me.png" width=180></a><br><font size=-2><i>I am program co-chair for LISA 2018</i></font>
|
|
-->
|
|
|
|
<hr>
|
|
</div>
|
|
<div class="navmobile">
|
|
<p class="navhdr">Brendan's site:</p>
|
|
<a href="/overview.html">Start Here</a><br>
|
|
<a href="/index.html">Homepage</a><br>
|
|
<a href="/blog/index.html">Blog</a><br>
|
|
<!-- <a href="/sitemap.html">Full Site Map</a><br> -->
|
|
<a href="/systems-performance-2nd-edition-book.html">Sys Perf book</a><br>
|
|
<a href="/bpf-performance-tools-book.html">BPF Perf book</a><br>
|
|
<a href="/linuxperf.html">Linux Perf</a><br>
|
|
<a href="/ebpf.html">eBPF Tools</a><br>
|
|
<a href="/perf.html">perf Examples</a><br>
|
|
<a href="/methodology.html">Perf Methods</a><br>
|
|
<a href="/usemethod.html">USE Method</a><br>
|
|
<a href="/tsamethod.html">TSA Method</a><br>
|
|
<a href="/offcpuanalysis.html">Off-CPU Analysis</a><br>
|
|
<a href="/activebenchmarking.html">Active Bench.</a><br>
|
|
<a href="/wss.html">WSS Estimation</a><br>
|
|
<a href="/flamegraphs.html">Flame Graphs</a><br>
|
|
<a href="/flamescope.html">Flame Scope</a><br>
|
|
<a href="/heatmaps.html">Heat Maps</a><br>
|
|
<a href="/frequencytrails.html">Frequency Trails</a><br>
|
|
<a href="/colonygraphs.html">Colony Graphs</a><br>
|
|
<a href="/dtrace.html">DTrace Tools</a><br>
|
|
<a href="/dtracetoolkit.html">DTraceToolkit</a><br>
|
|
<a href="/dtkshdemos.html">DtkshDemos</a><br>
|
|
<a href="/guessinggame.html">Guessing Game</a><br>
|
|
<a href="/specials.html">Specials</a><br>
|
|
<a href="/books.html">Books</a><br>
|
|
<a href="/sites.html">Other Sites</a><br>
|
|
|
|
</div>
|
|
|
|
<div class="footer">
|
|
<div class="contact">
|
|
Copyright 2025 Brendan Gregg.<br><a href="/blog/about.html">About this blog</a>
|
|
</div>
|
|
</div>
|
|
</div>
|
|
</div>
|
|
|
|
</body>
|
|
</html>
|