Files
nexus/sreweekly/markdown/384/01-scaling-merge-ort-across-github.md
2026-09-12 17:23:01 +08:00

126 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Scaling merge-ort across GitHub
- **期号**: SRE Weekly Issue #384(2023-08-06)
- **作者**: Jesse Toth — GitHub
- **链接**: https://github.blog/2023-07-27-scaling-merge-ort-across-github/
## 简介
They tested this new git merge strategy by using Scientist, a framework that runs both the old and new implementation and compares the results.
## 正文
# Scaling merge-ort across GitHub
GitHub switched to performing merges and rebases using merge-ort. Come behind the scenes to see why and how we made this change.
![](https://github.blog/wp-content/uploads/2023/07/Open-Source-Engineering@2x.png?resize=1600%2C850)
|
5 minutes
At GitHub, we perform a lot of merges and rebases in the background. For example, when you’re ready to merge your pull request, we already have the resulting merge assembled. Speeding up merge and rebase performance saves both user-visible time and backend resources. Git has recently [learned some new tricks](https://github.blog/2021-08-16-highlights-from-git-2-33/#merge-ort-a-new-merge-strategy) which we’re using at scale across GitHub. This post walks through what’s changed and how the experience has improved.
## [Our requirements for a merge strategy](https://github.blog#our-requirements-for-a-merge-strategy)
There are a few non-negotiable parts of any merge strategy we want to employ:
- **It has to be fast.** At GitHub’s scale, even a small slowdown is multiplied by the millions of activities going on in repositories we host each day.
- **It has to be correct.** For merge strategies, what’s “correct” is occasionally a matter of debate. In those cases, we try to match what users*expect* (which is often whatever the Git command line does).
- **It can’t check out the repository.** There are both scalability and security implications to having a working directory, so we simply don’t.
[Previously](https://github.blog/2015-12-15-move-fast/), we used `libgit2` to tick these boxes: it was faster than [Git’s default merge strategy](https://git-scm.com/docs/merge-strategies#Documentation/merge-strategies.txt-recursive) and it didn’t require a working directory. On the correctness front, we either performed the merge *or* reported a merge conflict and halted. However, because of additional code related to [merge base selection](https://public-inbox.org/git/539A25BF.4060501@alum.mit.edu/), sometimes a user’s local Git could easily merge what our implementation could not. This led to a steady stream of support tickets asking why the GitHub web UI couldn’t merge two files when the local command line could. We weren’t meeting those users’ expectations, so from their perspective, we weren’t correct.
## [A new strategy emerges](https://github.blog#a-new-strategy-emerges)
Two years ago, Git learned a new merge strategy, `merge-ort`. As the author [details on the mailing list](https://lore.kernel.org/git/4a0f088f3669a95c7f75e885d06c0a3bdaf31f42.1628055482.git.gitgitgadget@gmail.com/), `merge-ort` is fast, correct, and addresses many shortcomings of the older default strategy. Even better, unlike `merge-recursive`, it doesn’t need a working directory. `merge-ort` is much faster even than our optimized, `libgit2`-based strategy. What’s more, `merge-ort` has since become Git’s default. That meant our strategy would fall even further behind on correctness.
It was clear that GitHub needed to upgrade to `merge-ort`. We split this effort into two parts: first deploy `merge-ort` for merges, then deploy it for rebases.
## [`merge-ort` for merges](https://github.blog#merge-ort-for-merges)
`merge-ort` for merges
Last September, we [announced](https://github.blog/changelog/2022-09-12-merge-commits-now-created-using-the-merge-ort-strategy/) that we’re using `merge-ort` for merge commits. We used [Scientist](https://github.blog/2016-02-03-scientist/) to run *both* code paths in production so we can compare timing, correctness, etc. without risking much. The customer still gets the result of the old code path, while the GitHub feature team gets to compare and contrast the behavior of the new code path. Our process was:
1. Create and enable a Scientist experiment with the new code path.
2. Roll it out to a fraction of traffic. In our case, we started with some GitHub-internal repositories first before moving to a percentage-based rollout across all of production.
3. Measure gains, check correctness, and fix bugs iteratively.
We saw dramatic speedups across the board, especially on large, heavily-trafficked repositories. For our own `github/github` monolith, we saw a 10x speedup in both the average and P99 case. Across the entire experiment, our P50 saw the same 10x speedup and P99 case got nearly a 5x boost.
![Chart showing experimental candidate versus control at P50. The candidate implementation fairly consistently stays below 0.1 seconds.](https://github.blog/wp-content/uploads/2023/07/merge-ort-1.png?w=1024&resize=1024%2C459)
![Chart showing experimental candidate versus control at P99. The candidate implementation follows the same spiky pattern as the control, but its peaks are much lower.](https://github.blog/wp-content/uploads/2023/07/merge-ort-2.png?w=1024&resize=1024%2C458)
![Dashboard widgets showing P50 average times for experimental candidate versus control. The control averages 71.07 milliseconds while the candidate averages 7.74 milliseconds.](https://github.blog/wp-content/uploads/2023/07/merge-ort-3.png?w=1024&resize=1024%2C216)
![Dashboard widgets showing P99 average times for experimental candidate versus control. The control averages 1.63 seconds while the candidate averages 329.82 milliseconds.](https://github.blog/wp-content/uploads/2023/07/merge-ort-4.png?w=1024&resize=1024%2C215)
## [`merge-ort` for rebases](https://github.blog#merge-ort-for-rebases)
`merge-ort` for rebases
Like merges, we also do a huge number of rebases. Customers may choose [rebase workflows](https://docs.github.com/pull-requests/collaborating-with-pull-requests/proposing-changes-to-your-work-with-pull-requests/keeping-your-pull-request-in-sync-with-the-base-branch#updating-your-pull-request-branch) in their pull requests. We also perform test rebases and other “behind the scenes” operations, so we also [brought merge-ort to rebases](https://github.blog/changelog/2023-06-28-rebase-commits-now-created-using-the-merge-ort-strategy/).
This time around, we powered rebases using a new Git subcommand: `git-replay`. `git replay` was written by the original author of `merge-ort`, [Elijah Newren](https://github.com/newren) (a prolific Git contributor). With this tool, we could perform rebases using `merge-ort` and without needing a worktree. Once again, the path was pretty similar:
1. Merge `git-replay` into our fork of Git. (We were running the experiment with Git 2.39, which didn’t include the`git-replay` feature.)
2. Before shipping, leverage our test suite to detect discrepancies between the old and the new implementations.
3. Write automation to flush out bugs by performing test rebases of all open pull requests in `github/github` and comparing the results.
4. Set up a Scientist experiment to measure the performance delta between `libgit2` -powered rebases and monitor for unexpected mismatches in behavior.
5. Measure gains, check correctness, and fix bugs iteratively.
Once again, we were amazed at the results. The following is a great anecdote from testing, as relayed by [@wincent](https://github.com/wincent) (one of the GitHub engineers on this project):
Another way to think of this is in terms of resource usage. We ran the experiment over 730k times. In that interval, our computers spent 2.56 hours performing rebases with `libgit2`, but under 10 minutes doing the same work with `merge-ort`. And this was running the experiment for 0.5% of actors. Extrapolating those numbers out to 100%, if we had done all rebases during that interval with `merge-ort`, it would have taken us 2,000 minutes, or about 33 hours. That same work done with `libgit2` would have taken 512 hours!
## [What’s next](https://github.blog#whats-next)
While we’ve covered the most common uses, this is not the end of the story for `merge-ort` at GitHub. There are still other places in which we can leverage its superpowers to bring better performance, greater accuracy, and improved availability. Squashing and reverting are on our radar for the future, as well as considering what new product features it could unlock down the road.
### [Appreciation](https://github.blog#appreciation)
Many thanks to all the GitHub folks who worked on these two projects. Also, GitHub continues to be grateful for the hundreds of volunteer contributors to the Git open source project, including [Elijah Newren](https://github.com/newren) for designing, implementing, and continually improving `merge-ort`.
## Tags:
## Written by
## Related posts
![Copilot hovering above a mosaic of green squares in a decorative scene.](https://github.blog/wp-content/uploads/2026/01/generic-github-copilot-commit-logo.png?resize=400%2C212)
###
[How we make AI coding more cost efficient without sacrificing task quality](https://github.blog/ai-and-ml/github-copilot/how-we-make-ai-coding-more-cost-efficient-without-sacrificing-task-quality/)
Why shorter outputs can cost more, and how GitHub Copilot reduces wasted work across the complete coding task.
![Humorous header image with 'alt=IMG_2847.png' written across the top, and Mona's head looking up at a cube next to a GitHub logo on the bottom.](https://github.blog/wp-content/uploads/2026/08/header-1.png?resize=400%2C212)
###
[Your alt text passes automated checks. That doesn’t mean it’s any good.](https://github.blog/engineering/user-experience/your-alt-text-passes-automated-checks-that-doesnt-mean-its-any-good/)
We built a plugin for the GitHub Accessibility Scanner to make sure your alt text is actually accessible. Here’s how it works.
![Copilot hovering above geometric blocks featuring the GitHub invertocat logo in a decorative scene.](https://github.blog/wp-content/uploads/2026/01/generic-copilot-flying-invertocat-logo-github.png?resize=400%2C212)
###
[Using the GitHub Copilot SDK for Java](https://github.blog/engineering/using-the-github-copilot-sdk-for-java/)
Enterprise Java developers have a new superpower—drive GitHub Copilot from idiomatic Java code with annotations, virtual threads, and more.