# Branch Prediction in .NET

You can run the same C# loop over two arrays containing exactly the same values and get different timings. The number of elements hasn't changed. The calculation hasn't changed. Neither version allocates anything while scanning. All you've changed is the order of the data. One possible explanation is branch prediction. Your CPU tries to predict which instructions it will need next, including which way a conditional branch will go. When the prediction is right, it can keep useful work moving through the processor. When it's wrong, some of that work has to be discarded and execution redirected.

In the [cache-lines article](https://dotnetdigest.com/cache-lines-the-invisible-boundary-your-net-code-keeps-crossing), I looked at how the position of data affects the cost of accessing it. There's another part of the story, the pattern of decisions your code makes while processing that data. A simple `if` inside a busy loop gives us a useful place to start. But before we blame it for a performance problem, we need to establish whether that `if` even survives as a conditional branch in the generated machine code.

## A decision repeated a million times

Imagine a batch processor scanning submission statuses to count the records requiring review:

```csharp
public static int CountRequiringReview(int[] statuses)
{
    var count = 0;
    for (var i = 0; i < statuses.Length; i++)
    {
        if (statuses[i] == 1)
        {
            count++;
        }
    }
    return count;
}
```

Assume that zero means a record can proceed automatically and one means it requires review. Half the records have each status. Now consider three arrangements. In the first, every automatic record appears before every review record. In the second, the statuses alternate. In the third, the same statuses have been shuffled. Each arrangement produces the same count. Each reads the array sequentially. Each executes the same C# method. Yet if the status check becomes a conditional branch, the CPU encounters a different sequence of branch outcomes in each case.

Long runs of the same outcome are usually easy to predict. Alternating outcomes have a regular pattern that modern predictors can often learn. A shuffled sequence with roughly equal numbers of both outcomes is generally more difficult. Thats useful, a condition being true half the time tells you very little about how predictable it is. A perfectly alternating sequence and a shuffled sequence can have the same proportion of true values while behaving differently at the processor level.

## Why the CPU predicts anything

Modern processors overlap work. They can fetch and decode upcoming instructions while earlier instructions are still being executed. They can also execute instructions speculatively, before it has been established that those instructions belong to the path the program will actually take. A conditional branch introduces uncertainty. The next useful instruction might be on the path that handles a review record, or it might be on the path that skips it. Waiting for every decision to resolve would restrict how far the processor can work ahead.

Branch prediction helps it choose a path early. If the prediction is confirmed, the speculative work can contribute to the result. If the prediction is wrong, work from the incorrect path is discarded and the processor continues along the correct path. The program still produces the correct result, but recovering from the wrong prediction costs time. Microsoft's discussion of branching in .NET describes this relationship between pipelining, prediction and branch elimination.

![](https://cdn.hashnode.com/uploads/covers/67c36038c69a4b7143c5fc49/1f34e671-9896-4a05-a8f3-60f45f008077.png align="center")

This is a simplified view. Real processors have several prediction mechanisms, and their behaviour varies between architectures and processor generations. You don't need to reproduce those mechanisms to investigate a .NET application. You need to understand that predictability can influence throughput, then collect evidence from your own code. There also isn't one fixed penalty you can apply to every misprediction. The processor, surrounding instructions, dependencies and memory behaviour all affect the observed cost. Multiplying the number of `if` statements by an assumed cycle count won't give you a useful performance estimate.

## The C# condition might disappear

The counting example is deliberately small. It's also a good example of why source code alone can't establish a branch prediction problem. The JIT may turn a conditional increment into instructions that calculate a zero or one and add it to the count. On x64, you might see an instruction such as `sete`, which produces a value from comparison flags without a conditional jump. Other architectures have their own ways to express conditional operations. If that happens, the status check no longer asks the branch predictor to choose between incrementing and skipping. The loop still has control flow, but the data dependent decision we're interested in has changed shape.

Writing this instead doesn't settle the issue:

```csharp
count += statuses[i] == 1 ? 1 : 0;
```

The ternary expression might produce the same machine code as the original `if`. It might produce different code. Runtime version, architecture and the surrounding method all influence the result. Microsoft's .NET performance articles show examples where JIT improvements replace branching with conditional operations. Those transformations are a reason to inspect the generated code rather than assume a particular spelling of C# is faster.

For this experiment, we'll include both a simple counting loop and a second loop that conditionally calls a method. The latter makes retaining a branch more likely, although we still need to check the disassembly.

## A benchmark that changes the decision pattern

The experiment below uses one array per benchmark case. All three patterns contain exactly the same number of zeros and ones. The only change is their order. Data generation happens in `GlobalSetup`, outside the timed methods. That avoids measuring random number generation, shuffling or array construction alongside the operation we're trying to understand.

Create a .NET 10 console project and add BenchmarkDotNet:

```bash
dotnet new console -n BranchPredictionDemo -f net10.0
cd BranchPredictionDemo
dotnet add package BenchmarkDotNet
```

Replace `Program.cs` with the following:

```csharp
using System.Runtime.CompilerServices;
using BenchmarkDotNet.Attributes;
using BenchmarkDotNet.Running;

BenchmarkSwitcher
    .FromAssembly(typeof(BranchPredictionBenchmarks).Assembly)
    .Run(args);

public enum StatusPattern
{
    Grouped,
    Alternating,
    Shuffled
}

[MemoryDiagnoser]
[DisassemblyDiagnoser(maxDepth: 2)]
public class BranchPredictionBenchmarks
{
    private int[] _statuses = null!;

    [Params(4_096, 1_048_576)]
    public int Count { get; set; }

    [Params(
        StatusPattern.Grouped,
        StatusPattern.Alternating,
        StatusPattern.Shuffled)]
    public StatusPattern Pattern { get; set; }

    [GlobalSetup]
    public void Setup()
    {
        _statuses = new int[Count];
        for (var i = 0; i < _statuses.Length; i++)
        {
            _statuses[i] = i < Count / 2 ? 0 : 1;
        }

        if (Pattern == StatusPattern.Alternating)
        {
            for (var i = 0; i < _statuses.Length; i++)
            {
                _statuses[i] = i & 1;
            }
        }
        else if (Pattern == StatusPattern.Shuffled)
        {
            var random = new Random(42);

            for (var i = _statuses.Length - 1; i > 0; i--)
            {
                var other = random.Next(i + 1);
                (_statuses[i], _statuses[other]) =
                    (_statuses[other], _statuses[i]);
            }
        }

        var expected = Count / 2;

        if (CountWithIf() != expected ||
            CountWithConditionalExpression() != expected ||
            CountWithAddition() != expected ||
            ReviewSelected() != expected ||
            ReviewEveryRecord() != expected)
        {
            throw new InvalidOperationException(
                "The benchmark methods returned different results.");
        }
    }

    [Benchmark]
    public int CountWithIf()
    {
        var count = 0;
        var statuses = _statuses;

        for (var i = 0; i < statuses.Length; i++)
        {
            if (statuses[i] == 1)
            {
                count++;
            }
        }
        return count;
    }

    [Benchmark]
    public int CountWithConditionalExpression()
    {
        var count = 0;
        var statuses = _statuses;

        for (var i = 0; i < statuses.Length; i++)
        {
            count += statuses[i] == 1 ? 1 : 0;
        }
        return count;
    }

    [Benchmark]
    public int CountWithAddition()
    {
        var count = 0;
        var statuses = _statuses;

        for (var i = 0; i < statuses.Length; i++)
        {
            count += statuses[i];
        }
        return count;
    }

    [Benchmark]
    public int ReviewSelected()
    {
        var count = 0;
        var statuses = _statuses;

        for (var i = 0; i < statuses.Length; i++)
        {
            if (statuses[i] == 1)
            {
                count += Review(statuses[i]);
            }
        }
        return count;
    }

    [Benchmark]
    public int ReviewEveryRecord()
    {
        var count = 0;
        var statuses = _statuses;

        for (var i = 0; i < statuses.Length; i++)
        {
            count += Review(statuses[i]);
        }
        return count;
    }

    [MethodImpl(MethodImplOptions.NoInlining)]
    private static int Review(int status) => status;
}
```

Run it in Release mode:

```bash
dotnet run -c Release -- --filter "*"
```

The `Review` method is artificial on purpose. Preventing inlining preserves a call boundary and helps give the conditional call experiment a branch to investigate. It doesn't represent the cost of a real review operation, and `NoInlining` is being used as an experimental control rather than an optimisation recommendation. The setup checks that every method returns the expected result. Both array sizes are even, so every pattern contains exactly half zeros and half ones. Shuffling preserves those counts, and the fixed seed makes that shuffled input reproducible. This article doesn't attach measured timings to this example. Run it on the runtime and hardware you care about, the following sections explain how to interpret the results without assuming a particular winner.

## What each comparison tells you

Start by comparing `CountWithIf` across the three patterns at the same array size. If the timings are similar, check whether the JIT removed the conditional jump associated with the status check. A branch free inner calculation would explain why changing the outcome pattern has little effect on that decision. Next, compare `ReviewSelected` across the patterns. Every case makes exactly the same number of calls to `Review`. If the shuffled case is slower and the status check remains a conditional jump, branch prediction is a plausible explanation. Hardware counters can help establish whether mispredictions actually increased.

`CountWithConditionalExpression` tests whether replacing the `if` with a ternary expression changes anything useful. If the assembly is identical, similar timings are unsurprising. A shorter expression hasn't necessarily changed the work performed by the CPU. `CountWithAddition` removes the status decision by using our data representation directly. Adding zeros and ones gives the number of review records. That equivalence depends on the input contract. If two later becomes another status, summing statuses no longer counts records whose status equals one. Then, `ReviewEveryRecord` explores the cost of performing work unconditionally. It calls `Review` for every element, including zeros, so it makes twice as many calls as `ReviewSelected`. The returned total is still the same because this deliberately simple method returns its input.

An unpredictable branch can cost enough that doing some extra cheap work is competitive. A predictable branch can make skipping that work more attractive. This experiment lets you examine the trade off, but doesn't establish that unconditional processing is generally faster. In a real processor, review could involve validation, allocation, logging or external calls. Executing those operations for records that don't require them could be expensive or change the application's behaviour. The numerical equivalence in this experiment doesn't give you permission to remove conditions around side effects.

## Read the assembly before explaining the result

On a supported platform, BenchmarkDotNet's disassembly diagnoser produces reports under `BenchmarkDotNet.Artifacts/results`. Find the inner loop and identify the instructions associated with the status comparison. For x64 code, conditional jumps include instructions such as `je` and `jne`. A comparison followed by a jump around the call to `Review` would support the interpretation that this is a data dependent branch. Instructions such as `sete` and `cmov` represent different ways of using a condition. Seeing those instead of a jump can change your explanation of the timing results. On ARM64, instruction names and code shapes differ, so interpret the report for that architecture.

Also distinguish the status check from the loop's own control flow. A loop usually still needs to decide whether to continue. Bounds checks and other generated checks can introduce additional branches too. Finding any conditional jump in the method doesn't establish that the source level `if` is responsible. If disassembly isn't available in your environment, you can still run the timing experiment. Treat the explanation as a hypothesis until another supported tool or environment gives you the missing evidence.

## Timing shows the difference; counters help explain it

Elapsed time establishes whether one arrangement performed differently. It doesn't, by itself, establish why. On supported Windows environments, BenchmarkDotNet can collect branch instructions and branch mispredictions. Its documentation lists restrictions including administrator access and limitations in virtualised environments.

For that additional run, add the Windows diagnostics package:

```bash
dotnet add package BenchmarkDotNet.Diagnostics.Windows
```

Then import the diagnoser namespace and add the hardware counter attribute to the benchmark class:

```csharp
using BenchmarkDotNet.Diagnosers;

[HardwareCounters(
    HardwareCounter.BranchInstructions,
    HardwareCounter.BranchMispredictions)]
```

Keep the existing class attributes as well. Follow the package's requirements for your environment and inspect any diagnostic warnings. A missing counter result doesn't mean the code had no mispredictions. On Linux, tools such as `perf` can collect hardware events when the host and permissions allow it. For a focused harness, an example command is:

```bash
perf stat -e branches,branch-misses -- dotnet FocusedHarness.dll
```

That measures the process, including startup, rather than automatically isolating one method. The harness should warm up the target operation and execute one selected pattern repeatedly so useful work dominates the measurement. Don't treat counters from a whole run containing every BenchmarkDotNet case as the counters for one particular loop. Even a focused counter report needs interpretation. Calls, returns and loop branches contribute to the totals. A misprediction percentage is a property of the measured execution, rather than a direct percentage of C# `if` statements that went wrong.

## The benchmark has its own workload

The small array occupies 16 KiB of element data. The larger array occupies 4 MiB. Those sizes let you examine the same operation with different memory footprints, although the cache levels they fit into depend on your CPU and what else is using those caches. Sequential access keeps the experiment close to the cache friendly loops from the earlier article. It reduces the complications introduced by following object references, but it doesn't remove memory effects altogether. A sufficiently large scan can become constrained by data movement, reducing the relative visibility of branch costs.

The shuffled array is also reused. Each benchmark invocation sees the same sequence, which lets the processor encounter that sequence repeatedly. Modern predictors may learn some repeated patterns, especially with small inputs. A fixed shuffled sequence is reproducible, it isn't a fresh source of unpredictability on every invocation. If production inputs vary between batches, a follow up experiment can rotate through several pre generated arrays. Keep input generation outside the timed operation and account for the larger working set. Otherwise you can accidentally introduce a memory footprint difference while trying to study branch behaviour.

The 50/50 distribution is another deliberate choice. Production data may contain only a small proportion of review records. Try representative proportions as a separate experiment, preserving the same values across the different arrangements. A mostly false condition can be relatively predictable even when the occasional true outcome appears in irregular positions.

## Grouping has a price

After seeing a grouped input perform well, it's tempting to sort every batch before processing it. That adds work which the main benchmark deliberately excludes. Sorting, partitioning or building separate collections consumes CPU time and may allocate memory. Processing can also have ordering requirements. Changing record order can affect downstream behaviour even when the final count stays the same. For a batch scanned once, the cost of rearranging it can outweigh the saving in the scan. For a batch processed repeatedly, that cost may be paid once and recovered across several passes. If ingestion already separates record categories, you may get predictable processing without adding another rearrangement step.

![](https://cdn.hashnode.com/uploads/covers/67c36038c69a4b7143c5fc49/c103c085-9597-4086-86d6-6dec6dbd347f.png align="center")

For an optimisation decision, benchmark the operation the application actually performs. If that operation includes preparing groups, include preparation in the measurement. A faster inner loop is useful only when it improves the surrounding work enough to justify the change.

## Short circuiting changes the work too

Branch behaviour also appears in ordinary validation expressions:

```csharp
if (record.IsActive && PassesDetailedValidation(record))
{
    Process(record);
}
```

The `&&` operator skips the second condition when the first is false. Placing a cheap condition first can avoid expensive validation for records that fail immediately. Reordering conditions can therefore improve performance for reasons beyond branch prediction. It changes how often later work happens. Any comparison needs to account for both the pattern of outcomes and the amount of work being skipped. Replacing `&&` with `&` evaluates both operands. That can remove some short circuit control flow, but it also forces the validation call to happen for every record. It may introduce exceptions or side effects that the original expression avoided.

Keep the short circuit form when it expresses the required behaviour. For a measured hot path, investigate condition order only when changing that order preserves correctness. The cost and selectivity of each check often give you a more useful starting point than counting branches.

## The runtime learns as well

CPU branch prediction and dynamic profile guided optimisation operate at different levels. The CPU predicts control flow while executing machine code. Dynamic PGO gathers information about execution and lets the JIT use that information when producing optimised code. It can influence decisions such as code layout, helping keep frequently executed paths together. .NET 10 improves how profile information feeds into code layout. Microsoft's runtime article describes how hot and cold blocks, loop placement and branch directions influence that layout.

This helps explain why warm up and representative workloads are useful. A loop that has barely executed may be running different generated code from one that has become hot. A benchmark trained on one input distribution may also present a different profile from production. PGO doesn't make an irregular sequence of outcomes predictable by definition. It gives the runtime information it can use to improve generated code. The processor still has to execute that code against the values it receives. For an initial investigation, leave normal runtime optimisations enabled. Changing PGO settings can be a useful controlled follow up, but the first result should describe the configuration your application actually uses.

## Where this earns attention

Branch prediction is worth investigating when you have a small amount of work repeated very frequently, parsing bytes, scanning statuses, classifying telemetry, filtering numeric data or applying simple checks across large batches. It's less likely to explain the bulk of an operation dominated by database requests, network latency, large allocations or expensive calculations. The branches still exist, but their contribution may be too small for changing them to produce a useful improvement.

In business applications, the practical opportunity often sits inside one busy processing stage. You can keep the normal domain model around it and give that stage a compact representation suited to the work. As with cache locality, the scope of the optimisation can remain small. The earlier cache lines benchmark used alternating active flags. That arrangement was useful for comparing object and struct layouts with a consistent pattern. It wasn't intended to establish how either representation would behave under every possible decision pattern. If your production workload has irregular outcomes, vary those outcomes as another dimension of the investigation.

## Follow the evidence through the loop

Two loops can perform the same logical calculation while presenting different work to the processor. Data layout influences how values arrive. Generated instructions determine how decisions are expressed. Input order influences the sequence of outcomes the processor encounters. For a hot loop, compare representative inputs, inspect the generated code and use counters when they're available. If grouping helps, include its cost. If unconditional work helps, check that performing it preserves behaviour. If the JIT has already removed the branch, let that evidence guide the next experiment. A simple `if` can be cheap when its outcome is predictable, expensive when the processor repeatedly chooses the wrong path, or absent from the machine code altogether. Establishing which case you have is a much stronger basis for a change than rewriting the expression and hoping it becomes faster.

[.NET 10 performance improvements](https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-10/#code-layout)

[BenchmarkDotNet diagnosers](https://benchmarkdotnet.org/articles/configs/diagnosers.html)

[NET 8 performance improvements](https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-8/#branching)
