Developer Productivity

Automate bisect with git bisect run

How to use git bisect run with a script to automatically find the commit that introduced a bug, including setup, exit codes, and common pitfalls.

Mohammed Saqib11 min read
Dark room setup with code displayed on PC monitors highlighting cybersecurity themes.
Photo by Tima Miroshnichenko on Pexels · Pexels License

You've got a regression that only shows up in production, or a test that started failing three weeks ago and nobody noticed until today. You know the bug exists in main, you know it didn't exist at some older tag, and now you're staring at a range of a few hundred commits wondering which one broke it. Running git bisect manually means checking out a commit, building, running the failing test, typing git bisect good or git bisect bad, and repeating that ten or twelve times. It works, but it's exactly the kind of repetitive, mechanical work that a script should do for you. git bisect run turns that whole loop into a single command, provided you give it a script that can decide "good" or "bad" without human input.

Why git bisect run beats manual bisect

Manual bisect has a specific failure mode: you get bored, you start guessing, and you mark a commit good or bad based on a hunch rather than an actual test run. Even when you're disciplined, the process is slow. Each step requires a context switch — checkout, build, run test, interpret result, mark commit, repeat. For a range of 200 commits, that's eight iterations of the binary search, each one costing you a few minutes of attention. Multiply that by every regression you hunt in a quarter and it adds up to real time.

git bisect run collapses all of that into one command. You write a script that encapsulates your definition of "broken," then hand it to Git. Git checks out the midpoint of your range, runs the script, reads the exit code, and uses that to narrow the search. The script runs in a clean state each time, so you don't have to worry about stale build artifacts or leftover state from the previous commit. The binary search proceeds automatically until Git finds the first commit where your script's exit code flips from "good" to "bad."

This is especially valuable for regressions that aren't a single failing test. Maybe the bug is a performance degradation that only shows up when a benchmark crosses a threshold. Maybe it's a build failure that only happens with certain dependency versions. Maybe it's a flaky test that fails intermittently — you can wrap the test runner in a script that retries a few times before declaring the commit bad. The script is the abstraction; git bisect run just needs a reliable boolean answer per commit.

Writing a script that git bisect run understands

The contract is simple: your script exits 0 if the commit is good, and any exit code from 1 to 127 (excluding 125) if the commit is bad. Exit code 125 is special — it means "I can't test this commit," and Git will skip it and move on. More on that in a moment.

The script itself can be anything executable: a shell script, a Node.js one-liner, a Makefile target wrapped in a script. The key is that it must be deterministic and reasonably fast, because Git will run it once per bisect step. If your full test suite takes twenty minutes, git bisect run will take twenty minutes times the number of steps. That's why the best bisect scripts target a narrow slice of the problem.

Here's what a minimal script looks like:

#!/usr/bin/env bash
set -euo pipefail
 
# Build first, so we're testing the actual code
npm run build
 
# Run the specific test that reproduces the bug
npx jest --testNamePattern='regression: checkout flow'

The set -euo pipefail is important. If the build fails, the script exits non-zero immediately, which Git interprets as "bad." That's usually wrong for a build failure — you want 125 so Git skips the commit rather than blaming it for a bug it might not even have. We'll fix that in the real example below.

Another common pattern is a compile check. If you're hunting a TypeScript regression, a script that runs tsc --noEmit and exits 125 on failure, then runs a specific test, is a solid approach. The script should be idempotent: running it twice on the same commit should give the same result. Avoid anything that writes to shared state, like a database or a cache directory, unless you're confident it's clean between runs.

Setup: marking the range and starting the run

The setup is the same as a manual bisect, just condensed. You tell Git the bad commit (where the bug exists) and a good commit (where it doesn't). Then you pass your script to git bisect run.

# Start the bisect session
git bisect start
 
# Mark the known bad and good commits
git bisect bad main
git bisect good v1.2.0
 
# Run the automated bisect with your script
git bisect run ./scripts/check-regression.sh

Git will check out the midpoint, run your script, read the exit code, and automatically mark the commit good or bad. It prints progress after each step, showing how many revisions remain. When it's done, it prints the first bad commit — the one that introduced the regression.

One thing to watch: the script path. If you pass a relative path like ./check-regression.sh, Git runs it from the repository root, but only if the script is at that path in the current checkout. Since bisect checks out different commits, a script that's not committed to the repo will disappear when Git switches to an older commit. The safest approach is to commit the script to the repository, or place it outside the working tree entirely, like /usr/local/bin/check-regression.sh. Relative paths fail silently in the worst way: the script isn't found, Git sees a non-zero exit, and it marks a commit bad that might be perfectly fine.

Exit code 125: skipping unbuildable or untestable commits

Exit code 125 is the escape hatch. It tells Git, "I can't make a determination on this commit, skip it." Git will mark the commit as untestable and move to another candidate in the range. This is essential for real-world repositories where not every commit compiles.

Consider a commit that introduced a syntax error in a file that your script depends on. Without 125, your script would exit non-zero, Git would mark that commit as "bad," and the binary search would proceed as if that commit introduced the regression. But the commit might be ancient, and the actual regression might be much later — the syntax error just prevented your script from running. With 125, Git skips that commit and keeps searching, and the final result points at the commit where the test actually started failing.

Here's the pattern I use:

#!/usr/bin/env bash
set -uo pipefail
 
# If the build fails, we can't test this commit
if ! npm run build; then
  exit 125
fi
 
# Run the specific test
npx jest --testNamePattern='regression: checkout flow'

Notice I dropped set -e. The if ! construct handles the build failure explicitly. If jest fails, the script exits with its non-zero code, and Git marks the commit bad. If the test passes, the script exits 0, and Git marks it good. The 125 path is only for the case where the build itself is broken.

Without 125, you'll get false positives. A commit that breaks the build will be blamed for the regression even if the actual bug was introduced later. The binary search will still terminate, but it'll point at the wrong commit, and you'll waste time investigating a commit that was never runnable.

Real example: finding a regression in a TypeScript project

Let's put this together with a concrete scenario. You have a TypeScript project, and a test that validates the checkout flow started failing last week. You know it passed at the v2.3.0 tag. The bug is somewhere in the last fifty commits.

Here's a script that handles both build failures and the test itself:

#!/usr/bin/env bash
set -uo pipefail
 
# Step 1: Type-check the project. If this fails, the commit is untestable.
if ! npx tsc --noEmit; then
  echo "TypeScript build failed, skipping commit"
  exit 125
fi
 
# Step 2: Run the specific failing test. Exit code determines good/bad.
npx jest --testNamePattern='checkout flow: applies discount code'

Run it:

git bisect start
git bisect bad main
git bisect good v2.3.0
git bisect run ./scripts/bisect-checkout.sh

Git will churn through the range. On commits where tsc fails, it prints a message and skips. On commits where the test passes, it marks good. On commits where the test fails, it marks bad. After about six steps (for fifty commits), Git prints something like:

c4a7f2e3 is the first bad commit

You check out that commit, look at the diff, and find the offending change. The whole process took a few minutes of wall-clock time, mostly spent waiting for tsc and jest to run.

One refinement: if your test is flaky, wrap it in a retry loop. A single flaky failure can send the bisect down the wrong path, and you'll end up investigating a commit that was never actually broken. A simple retry that runs the test up to three times before declaring it bad adds a few seconds per step but saves you from a wrong answer.

Failure modes and gotchas

The most common failure is a script that isn't deterministic. If your test depends on network access, a shared database, or any external state, it might pass on one commit and fail on another for reasons unrelated to the code. The bisect will happily converge on a commit that looks guilty but isn't. The fix is to make the script as self-contained as possible: mock external services, use a local test database, or at least verify the environment is clean before each run.

The relative path problem I mentioned earlier is worth repeating. If your script is at ./scripts/check.sh and you run git bisect run ./scripts/check.sh, Git will look for that path in each checked-out commit. If the script isn't committed, older commits won't have it, and the script will fail with "No such file or directory" — which Git reads as a bad commit. Commit the script, or use an absolute path.

Another gotcha: forgetting to reset the bisect state. When you're done, git bisect reset returns your working tree to the branch you started from. If you don't reset, you're left on a detached HEAD at the first bad commit, which is confusing when you try to switch branches later. Make it a habit to run git bisect reset immediately after the run finishes, even if the result looks wrong.

Finally, be aware that bisect stores its state in .git/BISECT_* files. If you have a CI system or a script that runs git clean -fdx, it might wipe those files mid-run. It's rare, but it happens. If your bisect suddenly forgets where it was, check whether something cleaned the repo.

Alternatives and complementary tools

git bisect visualize opens a graphical view of the remaining range, which is handy when you want to see how many commits are left or spot a suspicious merge. git bisect log replays a session — useful if you need to pause and resume later, or share the bisect state with a colleague.

For teams, consider running bisects in CI. A GitHub Actions workflow that triggers on a failing test, checks out the range, and runs git bisect run can automate regression hunting entirely. It's a heavier setup, but it means the machine does the work while developers stay focused on fixes. If you're designing such a workflow, the principles around fail-fast and clear output apply directly — see Designing GitHub Actions That Fail Fast and Explain Why for the details.

During a bisect, you'll often hit merge conflicts when a commit touches files that were later refactored. git rerere (reuse recorded resolution) remembers how you resolved a conflict and applies it automatically on subsequent encounters. It's enabled with git config --global rerere.enabled true, and it can save you from resolving the same conflict a dozen times during a long bisect.

For the script itself, a well-tested shell function can be the core of your bisect logic. I keep a bisect_build_and_test function in my dotfiles that handles the 125 logic and the retry loop — it's saved me from rewriting the same script for every project. If you're looking for inspiration, Shell functions that earn their dotfiles spot has some patterns worth borrowing.

The official git-bisect documentation covers the full command surface, including the run subcommand and exit code semantics. If you're debugging a flaky test during a bisect, the Jest documentation on --testNamePattern is the reference for targeting a single test. And for the 125 exit code specifically, the Git source code is where the behavior is defined, though the man page is usually enough.

Key takeaways

  • git bisect run replaces the manual good/bad loop with a single command, provided your script returns 0 for good, non-zero for bad, and 125 for untestable.
  • Always handle build failures with exit 125 so Git skips unbuildable commits instead of blaming them for the regression.
  • Keep the script deterministic and fast — you'll run it once per bisect step, so a slow or flaky script defeats the purpose.
  • Commit the script to the repo or use an absolute path; a relative path to an uncommitted script will fail silently on older commits.
  • Run git bisect reset when you're done to return to your branch, and consider git bisect log if you need to pause and resume a long hunt.

Frequently asked questions

What exit codes does git bisect run expect?
Exit 0 means good. Exit 1-127 means bad. Exit 125 means the commit cannot be tested (skip it). Any other exit code aborts the bisect.
Can I use a script that requires user input?
No — git bisect run expects a non-interactive script. If your test needs input, redesign it to accept flags or environment variables.
How do I bisect a merge commit?
By default bisect skips merge commits. Use --first-parent to follow only the first parent, or handle them manually with git bisect skip.
#git#bisect#automation#debugging
Share

Keep reading