Hardening Linux Kernel eBPF: Blue Team Defense Guide

Hardening Linux Kernel eBPF: Blue Team Defense Guide

Overview & Defensive Context#

In our morning research breakdown, we deconstructed the root-cause mechanics of Linux Kernel eBPF Verifier Logic Flaws (such as CVE-2023-2163, CVE-2021-3490, and CVE-2020-8835). We analyzed how abstract interpretation and range deduction bugs in kernel/bpf/verifier.c create a fatal state divergence between the verifier's mathematical simulation and the physical CPU's Just-In-Time (JIT) execution.

By exploiting discrepancies in sub-register truncation (ALU32 vs. ALU64) and conditional branch evaluations (scalar32_min_max_and), an unprivileged attacker can trick the verifier into believing that a register holds a constant value of zero. At runtime, the physical CPU evaluates that register to an arbitrary non-zero integer. When added to a legitimate BPF map pointer (PTR_TO_MAP_VALUE), this bounds confusion primitive yields an out-of-bounds kernel pointer. Because the verifier mathematically approved the pointer based on its flawed state, the program gains arbitrary read and write capabilities across kernel space (Ring 0), enabling the attacker to leak the kernel base address (bypassing KASLR) and overwrite process credentials (task_struct->cred) to achieve instant root privilege escalation.

For Linux platform architects, kernel security engineers, and enterprise blue teams, defending against verifier vulnerabilities exposes critical architectural realities:

  • Hardware Security Mitigations Are Ineffective: Verifier logic bugs execute natively within authorized in-kernel execution environments. Hardware-enforced protections—such as Supervisor Mode Execution Prevention (SMEP), Supervisor Mode Access Prevention (SMAP), and Kernel Page Table Isolation (KPTI)—are entirely bypassed because the exploit manipulates internal kernel heap objects and memory translation tables directly without executing userspace shellcode or corrupting control-flow returns.
  • The Complexity Paradox of in-Kernel Static Analysis: The eBPF verifier is an abstract interpretation engine comprising tens of thousands of lines of complex C code running directly in the kernel core. As eBPF expands to support bounded loops, 32-bit sub-registers, atomic operations, and dynamic pointers, the mathematical state space expands exponentially. Subtle range deduction and branch pruning oversights become mathematically inevitable.
  • Permissive Default Distribution Settings: Historically, many enterprise Linux distributions left unprivileged eBPF enabled (kernel.unprivileged_bpf_disabled = 0), allowing any local account or unprivileged container process to submit raw bytecode to the kernel verifier.

[!WARNING] If unprivileged eBPF is enabled on a system, any unprivileged local user or containerized workload can submit crafted bytecode to probe for verifier state divergence bugs, converting an unprivileged shell into full kernel root.

This guide delivers an enterprise-grade blue team defense blueprint: locking down unprivileged eBPF via immutable sysctl parameters, enforcing container seccomp profiles, implementing mandatory access control via BPF LSM, deploying production-grade Sigma detection rules, and executing a complete incident response triage playbook.


Architecture Hardening#

Securing Linux infrastructure against eBPF verifier vulnerabilities requires establishing defense-in-depth boundaries that neutralize the attack surface across four tiers: Immutable Sysctl Lockdown, Container Seccomp Filtering, BPF LSM Access Control Policies, and JIT Memory Blinding.

Defense Architecture Flowchart#

flowchart TD
    classDef attacker fill:#1e293b,stroke:#ef4444,stroke-width:2px,color:#f8fafc
    classDef sysctl fill:#0f172a,stroke:#3b82f6,stroke-width:2px,color:#93c5fd
    classDef seccomp fill:#064e3b,stroke:#10b981,stroke-width:2px,color:#6ee7b7
    classDef lsm fill:#4c1d95,stroke:#8b5cf6,stroke-width:2px,color:#ddd6fe
    classDef drop fill:#450a0a,stroke:#dc2626,stroke-width:2px,color:#fca5a5
    classDef success fill:#065f46,stroke:#34d399,stroke-width:2px,color:#a7f3d0

    LocalProcess[Local Process / Container Workload]:::attacker --> SysctlCheck{Ring 1: Sysctl unprivileged_bpf_disabled == 2}:::sysctl
    
    SysctlCheck -- Unprivileged User & Sysctl Enabled --> Drop1[Action: EPERM Syscall Blocked Immuted]:::drop
    SysctlCheck -- Process Has CAP_BPF / CAP_SYS_ADMIN --> SeccompGate{Ring 2: Container Seccomp Profile}:::seccomp
    
    SeccompGate -- Seccomp Blocks syscall bpf --> Drop2[Action: EPERM / SIGSYS Container Intercept]:::drop
    SeccompGate -- Authorized Host Process --> LSMCheck{Ring 3: BPF LSM Mandatory Access Control}:::lsm
    
    LSMCheck -- Process Comm Not in Whitelisted Daemon List --> Drop3[Action: BPF_PROG_LOAD Denied by LSM]:::drop
    LSMCheck -- Verified Observability Daemon e.g. Cilium --> Verifier[Ring 4: eBPF Verifier Abstract Interpretation]:::lsm
    
    Verifier --> JITHarden{Ring 5: JIT Hardening net.core.bpf_jit_harden == 2}:::sysctl
    JITHarden -- Constant Blinding Active --> NativeCode[Native JIT Code Emitted with Obfuscated Constants]:::success

Layer 1: Immutable Sysctl Lockdown (unprivileged_bpf_disabled = 2)#

The single most effective defense against eBPF verifier exploitation is disabling unprivileged access system-wide. The Linux kernel provides the kernel.unprivileged_bpf_disabled sysctl parameter with three distinct operational states:

  • 0: Unprivileged eBPF is enabled for all users (legacy default).
  • 1: Unprivileged eBPF is disabled, but can be re-enabled at runtime by root.
  • 2: Unprivileged eBPF is permanently disabled. Once set to 2, the setting becomes immutable and cannot be reverted back to 0 or 1 without a full system reboot, even by the root superuser.

Enforce immutable unprivileged eBPF disabling along with JIT constant blinding:

BASH
# /etc/sysctl.d/60-ebpf-hardening.conf

## Permanently disable unprivileged eBPF (cannot be undone without reboot)
kernel.unprivileged_bpf_disabled = 2

## Enable constant blinding for all programs to prevent JIT spraying
net.core.bpf_jit_harden = 2

## Ensure BPF JIT compiler is enabled to avoid interpreter vulnerabilities
net.core.bpf_jit_enable = 1

## Bound memory allocatable by the BPF JIT compiler (in bytes)
net.core.bpf_jit_limit = 268435456

Apply the configuration immediately:

BASH
sudo sysctl -p /etc/sysctl.d/60-ebpf-hardening.conf

Layer 2: Container Seccomp Confinement & Dropping Capabilities#

In cloud-native and Kubernetes environments, containerized workloads must be barred from invoking the bpf() system call:

  1. Drop Capabilities: Ensure container runtimes strip CAP_SYS_ADMIN, CAP_BPF, and CAP_PERFMON from all pod security contexts:
YAML
# Kubernetes Pod Security Context Hardening
apiVersion: v1
kind: Pod
metadata:
  name: hardened-app
spec:
  containers:
  - name: web-service
    image: company/web:latest
    securityContext:
      allowPrivilegeEscalation: false
      capabilities:
        drop:
        - ALL
  1. Enforce Docker / Containerd Seccomp Profile: Ensure default seccomp profiles explicitly block the bpf system call:
JSON
{
  "defaultAction": "SCMP_ACT_ALLOW",
  "architectures": [
    "SCMP_ARCH_X86_64",
    "SCMP_ARCH_AARCH64"
  ],
  "syscalls": [
    {
      "names": [
        "bpf"
      ],
      "action": "SCMP_ACT_ERRNO",
      "args": [],
      "comment": "Block bpf syscall to eliminate kernel verifier attack surface"
    }
  ]
}

Layer 3: BPF LSM Mandatory Access Control#

For enterprise environments running legitimate eBPF monitoring tools (e.g., Cilium, Falco, Datadog), administrators can use BPF LSM to restrict program loading strictly to approved binary paths:

C
// BPF LSM Policy: Restrict BPF_PROG_LOAD to authorized binaries (bpf_guard.bpf.c)
#include <vmlinux.h>
#include <bpf/bpf_helpers.h>
#include <bpf/bpf_tracing.h>

char LICENSE[] SEC("license") = "GPL";

SEC("lsm/bpf")
int BPF_PROG(restrict_bpf, int cmd, union bpf_attr *attr, unsigned int size) {
    // Only inspect program load operations
    if (cmd != BPF_PROG_LOAD) {
        return 0;
    }

    char comm[16];
    bpf_get_current_comm(&comm, sizeof(comm));

    // Whitelist legitimate monitoring daemons
    if (comm[0] == 'c' && comm[1] == 'i' && comm[2] == 'l' && comm[3] == 'i') {
        return 0; // Allow Cilium
    }
    if (comm[0] == 'f' && comm[1] == 'a' && comm[2] == 'l' && comm[3] == 'c') {
        return 0; // Allow Falco
    }

    // Deny all other processes from loading eBPF bytecode
    return -1; // EPERM
}

Production Detection Queries#

Blue teams must deploy detection covering bpf() system call invocations, anomalous process lineage, and kernel audit logs.

Production Sigma Rule: Unauthorized BPF System Call Execution#

The following Sigma rule detects unauthorized processes invoking the bpf system call:

YAML
title: Suspicious Invocations of BPF System Call by Non-Standard Binaries
id: a8b4c2e1-4821-4f90-bc32-918237dc1042
status: production
description: |
  Detects instances where the bpf system call (syscall 321 on x86_64) is invoked by processes
  other than known system monitoring tools, container networking daemons, or authorized security agents.
references:
  - https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
  - https://ebpf.io/

tags:
  - attack.privilege_escalation
  - attack.t1068
  - attack.defense_evasion
  - attack.t1562.001
logsource:
  product: linux
  service: auditd
detection:
  selection:
    type: 'SYSCALL'
    arch: 'c000003e' # x86_64
    syscall:
      - '321' # bpf syscall number
  filter_known_daemons:
    exe:
      - '/usr/bin/bpftool'
      - '/usr/sbin/cilium-agent'
      - '/usr/bin/falco'
      - '/opt/datadog-agent/embedded/bin/system-probe'
      - '/usr/bin/systemd'
  condition: selection and not filter_known_daemons
fields:
  - exe
  - comm
  - uid
  - ppid
  - success
falsepositives:
  - Custom internal performance profiling utilities or developer testing in non-production environments.
level: high

Host-Level Auditd Rules for BPF Syscall Monitoring#

Deploy auditd rules to generate immediate telemetry whenever the bpf syscall is executed:

BASH
# Monitor 64-bit bpf syscall invocations
-a always,exit -F arch=b64 -S bpf -F key=bpf_syscall_audit

## Monitor 32-bit bpf syscall invocations
-a always,exit -F arch=b32 -S bpf -F key=bpf_syscall_audit

## Monitor modifications to kernel sysctl configuration files
-w /etc/sysctl.conf -p wa -k sysctl_tamper
-w /etc/sysctl.d/ -p wa -k sysctl_tamper

## Lock auditd configuration
-e 2

eBPF Runtime Kernel Probe: Detecting Suspicious bpf() Syscall Invocations#

Below is an eBPF tracepoint probe (bpf_syscall_tracer.bpf.c) that hooks sys_enter_bpf, logging the command argument and calling UID to catch stealthy unprivileged attempts:

C
#include <vmlinux.h>
#include <bpf/bpf_helpers.h>
#include <bpf/bpf_tracing.h>

struct event_t {
    __u32 pid;
    __u32 uid;
    __u32 cmd;
    char comm[16];
};

struct {
    __uint(type, BPF_MAP_TYPE_RINGBUF);
    __uint(max_entries, 1 << 24);
} events SEC(".maps");

SEC("tracepoint/syscalls/sys_enter_bpf")
int trace_bpf_enter(struct trace_event_raw_sys_enter *ctx) {
    struct event_t *event = bpf_ringbuf_reserve(&events, sizeof(struct event_t), 0);
    if (!event) {
        return 0;
    }

    event->pid = bpf_get_current_pid_tgid() >> 32;
    event->uid = bpf_get_current_uid_gid() & 0xFFFFFFFF;
    event->cmd = (__u32)ctx->args[0]; // BPF command: BPF_MAP_CREATE, BPF_PROG_LOAD, etc.
    bpf_get_current_comm(&event->comm, sizeof(event->comm));

    // Submit alert if caller is non-root or cmd is BPF_PROG_LOAD (5)
    if (event->uid != 0 || event->cmd == 5) {
        bpf_ringbuf_submit(event, 0);
    } else {
        bpf_ringbuf_discard(event, 0);
    }
    return 0;
}

char LICENSE[] SEC("license") = "GPL";

Enterprise Mitigation Matrix#

When evaluating defense strategies against eBPF verifier exploitation, infrastructure architects must balance observability requirements with attack surface reduction:

Remediation Strategy Implementation Effort Blast Radius & Operational Risk Detection & Prevention Efficacy Performance Overhead Architectural Longevity
Workaround: Sysctl Lockdown (unprivileged_bpf_disabled = 2) 5 - 15 Minutes
(Immediate sysctl deployment via Ansible or Puppet)
Low
Standard enterprise applications do not require unprivileged eBPF.
Absolute for Unprivileged Exploitation
Blocks unprivileged attackers from reaching verifier.c.
Zero
No performance impact.
Permanent Best Practice
Mandatory baseline configuration across all production servers.
Hotfix: Vendor Kernel Patch & Upstream Upgrade 2 - 4 Hours
(Kernel update requiring node reboot)
Moderate
Requires rolling node reboots in Kubernetes clusters.
Absolute for Patched CVE
Corrects specific range deduction and ALU32 logic bugs.
Zero
Baseline operational metrics maintained.
Mandatory Milestone
Eliminates the specific known vulnerability code path.
Architecture Fix: BPF LSM & Container Seccomp Profiles 1 - 2 Weeks
(Testing seccomp and LSM whitelist policies against workloads)
Moderate
Must ensure legitimate monitoring daemons are properly whitelisted.
Comprehensive (100%)
Enforces least-privilege mandatory access control on eBPF execution.
Negligible
(LSM entry check < 0.1% CPU).
Permanent State
Zero-trust kernel architecture resilient against future 0-days.

[!IMPORTANT] Setting kernel.unprivileged_bpf_disabled = 2 is an immediate, zero-downtime operational mitigation that completely neutralizes unprivileged verifier exploitation while allowing authorized system observability daemons to function normally.


Incident Response & Verification Playbook#

If a Linux server or container host is suspected of being targeted by an eBPF verifier exploit, execute the following four-phase incident response playbook:

flowchart TD
    classDef check fill:#1e293b,stroke:#f59e0b,stroke-width:2px,color:#fef3c7
    classDef alert fill:#450a0a,stroke:#dc2626,stroke-width:2px,color:#fca5a5
    classDef clean fill:#064e3b,stroke:#10b981,stroke-width:2px,color:#6ee7b7
    classDef triage fill:#0f172a,stroke:#3b82f6,stroke-width:2px,color:#93c5fd

    Start[Step 1: Check Sysctl unprivileged_bpf_disabled Status]:::triage --> CheckSysctl{Is unprivileged_bpf_disabled == 2?}:::check
    CheckSysctl -- No --> AlertExposure[Warning: Unprivileged Kernel Exposure Active]:::alert
    CheckSysctl -- Yes --> Step2[Step 2: Inspect Active eBPF Programs via bpftool]:::triage
    
    AlertExposure --> Step2
    Step2 --> CheckProgs{Anomalous or Unrecognized BPF Programs Loaded?}:::check
    CheckProgs -- Yes --> DeclareIncident[Declare Severity 1 Incident: Kernel Compromise]:::alert
    CheckProgs -- No --> Step3[Step 3: Audit Kernel Syslog for Verifier Warnings]:::triage
    
    DeclareIncident --> IsolateNode[Isolate Node & Dump Volatile Kernel Memory]:::alert
    IsolateNode --> UnpinProgs[Detach and Unload Malicious BPF Progs]:::alert
    
    Step3 --> CheckDmesg{Verifier State Warnings or JIT Errors in dmesg?}:::check
    CheckDmesg -- Yes --> DeclareIncident
    CheckDmesg -- No --> Step4[Step 4: Enforce Sysctl Lockdown & Seccomp Rules]:::clean

Phase 1: Live eBPF Program & Map Audit#

  1. Enumerate All Loaded eBPF Programs: Use bpftool to inspect all active BPF programs in the kernel:
BASH
# List all active eBPF programs with full metadata
sudo bpftool prog list

## Inspect loaded program types and verify loaded IDs against authorized daemons
sudo bpftool prog show --json | jq '.[] | {id: .id, type: .type, name: .name, loaded_at: .loaded_at}'
  1. Audit Active BPF Maps: Inspect BPF maps to identify unexpected array maps or pinned map handles in /sys/fs/bpf:
BASH
# List all BPF maps
sudo bpftool map list

## Inspect pinned objects in the BPF filesystem
ls -la /sys/fs/bpf/

Phase 2: Kernel Log & Memory Anomaly Triage#

  1. Inspect Kernel Ring Buffer for Verifier and JIT Errors: Exploitation attempts often trigger verifier rejections or JIT compilation warnings during weaponization:
BASH
# Search dmesg for BPF verifier warnings, page allocation failures, or JIT alerts
sudo dmesg -T | grep -iE "bpf|verifier|out of bounds|unrecognized|jit"
  1. Verify Process Credential Integrity: Inspect running processes to identify suspicious binaries running with UID 0 that originated from non-root parent processes:
BASH
# Search for processes running as root with unprivileged parent lineages
ps -eo pid,ppid,user,euser,comm,args | grep -E "^ *[0-9]+ +[0-9]+ +root"

Phase 3: Containment & Remediation#

  1. Unload Rogue eBPF Programs: If an unauthorized BPF program is identified, detach and unload it using bpftool:
BASH
# Remove pinned links in /sys/fs/bpf
sudo rm -f /sys/fs/bpf/

## Terminate the controlling user process to trigger automatic garbage collection of unpinned BPF programs
sudo kill -9 
  1. Enforce Immediate Immutable Sysctl Lockdown: Lock down the host immediately without rebooting:
BASH
# Set unprivileged eBPF to permanently disabled
sudo sysctl -w kernel.unprivileged_bpf_disabled=2
sudo sysctl -w net.core.bpf_jit_harden=2

Phase 4: Post-Remediation Hardening Verification#

Before returning the system to production service:

  • Verify cat /proc/sys/kernel/unprivileged_bpf_disabled outputs 2.
  • Verify cat /proc/sys/net/core/bpf_jit_harden outputs 2.
  • Test that unprivileged users receive -EPERM when invoking bpf():
BASH
# Test as unprivileged user
python3 -c 'import ctypes; libc=ctypes.CDLL(None); print("Syscall return:", libc.syscall(321, 0, 0, 0))'
## Expected output: Syscall return: -1 (with errno 1: Operation not permitted)
  • Confirm container seccomp profiles drop the bpf syscall across all production workloads.
  • Ensure auditd rules for bpf_syscall_audit are active and logging to the SIEM.

Authoritative References#

  1. Linux Kernel Documentation: BPF Sysctl Settings and Verifier Security Architecture — Kernel.org
  2. Red Hat Product Security: Hardening eBPF to Mitigate Local Kernel Privilege Escalation Vulnerabilities — Red Hat Customer Portal
  3. CISA & NSA Cybersecurity Guidance: Mitigating Linux Kernel Vulnerabilities via Kernel Lockdown and Sysctl Controls — CISA Best Practices
  4. Google Project Zero Research: Jit-picking: BPF verifier bug to kernel execution by Jann Horn — Project Zero Blog

Comments