eBPF in Production: Kernel Superpowers Without Kernel Modules
eBPF started as a packet filter and grew into a general-purpose, sandboxed virtual machine inside the Linux kernel. In production, it lets teams observe system calls, enforce security policy, route network traffic, and profile latency without loading custom kernel modules or restarting workloads. That combination of deep visibility and low overhead is why eBPF now sits under many cloud-native platforms, including Cilium, Falco, Tetragon, Pixie, and Parca. But eBPF is not magic. It has a strict verifier, kernel-version dependencies, privilege requirements, and real operational limits. Treating it as a production platform rather than a collection of clever scripts is the difference between useful telemetry and a cluster-wide outage.
What eBPF Actually Is
eBPF is a Linux kernel subsystem that runs bytecode in a restricted environment. A program is compiled from a supported language, usually C or Rust, into eBPF instructions. The kernel verifier proves that the program is safe: it cannot loop forever, dereference arbitrary memory, or crash the kernel. Once accepted, the program is attached to a hook such as a tracepoint, kprobe, cgroup, network interface, or Linux Security Module hook.
- Programs are event-driven. They run when the hook fires and must finish quickly.
- Maps are kernel-resident data structures used for state, counters, histograms, and communication with user space.
- Helper functions are stable kernel APIs for tasks such as reading process context, getting timestamps, manipulating packets, and writing to ring buffers.
- The verifier is the safety gate. It rejects unsafe memory access, unbounded loops, and invalid helper calls.
- The JIT compiler translates verified bytecode into native instructions for near-native performance.
A typical production lifecycle is: compile with Clang and libbpf, embed or load the object file, let the verifier check it, attach it to one or more hooks, read events from a ring buffer or maps, and detach cleanly during upgrades. With BTF and CO-RE, the same object can often run across multiple kernel versions without recompiling for every distro.
Why Production Teams Care
Traditional observability often depends on application agents, sidecars, or static instrumentation. eBPF can observe kernel events directly, which means it can see activity across all processes, containers, and network namespaces on a node. It is especially valuable when you cannot modify application code, when you need host-level context, or when you want to reduce the cost and complexity of per-pod agents.
- Lower overhead for high-frequency events: filtering and aggregation can happen in the kernel before data reaches user space.
- Broad coverage: one node-level agent can observe many workloads without language-specific SDKs.
- Runtime enforcement: security tools can block or audit actions at the kernel boundary, not just after the fact.
- Network efficiency: XDP and TC programs can process packets before they traverse the full networking stack.
- Dynamic deployment: many probes can be attached and detached without rebooting or rebuilding the kernel.
The tradeoff is that eBPF operates at a privileged layer. A bug, a misconfigured map, or an overly broad hook can affect every workload on the node. Production adoption requires guardrails, canaries, and clear ownership.
Observability Without Bespoke Instrumentation
eBPF observability usually combines three data types: traces, metrics, and profiles. Tracepoints and kprobes capture syscalls, scheduler events, file I/O, and network connections. Maps aggregate counts and latency histograms. Stack traces and perf events feed CPU profiles. Instead of asking every service to emit the same telemetry, you can derive a baseline from the kernel and then correlate it with application-level traces.
For example, a quick bpftrace one-liner can show which processes are opening files:
bpftrace -e 'tracepoint:syscalls:sys_enter_openat { printf("%s %s", comm, str(args->filename)); }'
In production, you would not run ad hoc one-liners everywhere. You would package eBPF programs into an agent with bounded maps, sampling, filtering, and a stable event schema. Tools such as Grafana Beyla, Coroot, and Pixie use eBPF for auto-instrumentation, while Cilium Hubble focuses on network flow visibility. The key is to start with specific questions: Which pods are making unexpected DNS queries? Which syscalls add latency during a deployment? Which nodes show TCP retransmits under load?
Security: Runtime Enforcement, Not Just Alerts
Security is one of the fastest-growing eBPF use cases because it shifts detection and enforcement closer to the kernel. Tools like Falco, Tetragon, and Tracee use eBPF to observe process execution, file access, network connections, privilege changes, and namespace escapes. Some can not only alert but also kill a process or deny an action using LSM hooks or signal-based enforcement.
- Process execution monitoring: detect unexpected shells, package managers, or crypto miners in containers.
- File integrity and access: watch sensitive paths such as /etc/shadow, /proc, and Kubernetes service account tokens.
- Network policy: enforce allowed egress and ingress at the socket or cgroup level.
- Privilege escalation: trace setuid, capabilities, and namespace transitions.
- Host intrusion detection: correlate kernel events with container and Kubernetes metadata.
The advantage over kernel modules is safety and maintainability: eBPF programs are verified before they run. The disadvantage is that the agent still needs significant privileges to attach to many hooks. A compromised privileged eBPF agent can be a serious risk, so run it with least privilege, isolate its credentials, and audit its deployment like any other security-critical component.
Networking and the Sidecar-Free Data Plane
eBPF is the foundation of modern Kubernetes networking in projects like Cilium. Instead of relying on iptables rules that grow with service count, eBPF programs can implement load balancing, network policy, NAT, and observability in the kernel. XDP programs run early in the receive path and can drop, redirect, or modify packets before the kernel allocates a socket buffer. TC programs run later and can shape, mark, or forward traffic. Socket and cgroup hooks can apply policy per pod or per connection.
- kube-proxy replacement: eBPF maps provide O(1) service lookup and reduce iptables overhead.
- Distributed load balancing: direct server return and Maglev-style hashing can improve throughput and tail latency.
- Network policy: identity-based policy enforcement without sidecars.
- Flow visibility: Hubble and similar tools export network flows with Kubernetes context.
- Bandwidth management: rate limiting and priority scheduling at the cgroup or interface level.
Sidecar-free service mesh is an attractive promise, but it is not automatic. You still need to understand packet paths, MTU, conntrack, encryption, and failure modes. eBPF networking can be extremely fast, but debugging a dropped packet in an XDP program is different from reading proxy logs.
Toolchain and Architecture
The modern eBPF toolchain has stabilized around libbpf, BTF, and CO-RE. Clang compiles C into eBPF bytecode. libbpf loads programs, creates maps, and handles attachments. BTF provides type information from the kernel, and CO-RE uses that information to relocate field offsets at load time. This means you can compile once and run across many kernel versions, as long as the required fields and hooks exist.
- bpftool inspects programs, maps, links, and BTF.
- BCC offers Python and Lua frontends, useful for rapid prototyping.
- bpftrace provides a high-level tracing language for one-liners and scripts.
- cilium/ebpf is a pure Go library for loading and managing eBPF.
- Aya is a Rust framework for eBPF with a focus on developer experience.
- libbpf-rs provides Rust bindings over libbpf.
For production, prefer CO-RE over hardcoded struct offsets. Prefer stable tracepoints and BTF-enabled fentry/fexit hooks over raw kprobes when possible. Keep a kernel-version test matrix in CI, because a program that loads on Ubuntu 22.04 may fail on a hardened distribution or an older vendor kernel.
Production Deployment Patterns
Most eBPF agents run as a DaemonSet on Kubernetes or as a systemd service on virtual machines. They typically need host PID and network namespaces, access to the BPF filesystem, and capabilities such as CAP_BPF, CAP_PERFMON, CAP_NET_ADMIN, and sometimes CAP_SYS_ADMIN. The exact set depends on program types and kernel version. Treat these privileges as a security boundary, not a convenience.
- Start small: attach to a few namespaces or cgroups before enabling node-wide hooks.
- Canary by node pool: roll out to a small group and compare CPU, memory, and event drop rates.
- Bound every map: set max entries and use LRU or per-CPU maps where appropriate.
- Use ring buffers carefully: they are efficient but still finite. Monitor drops and backpressure.
- Pin maps when state must survive restarts: use a dedicated bpffs path and clean up stale pins.
- Respect cgroup v2: many modern hooks expect unified hierarchy and may behave differently on cgroup v1.
- Plan for SELinux and AppArmor: mandatory access control can block loading or attaching even with root.
In multi-tenant clusters, avoid giving tenants the ability to load arbitrary eBPF. Unprivileged eBPF is restricted for good reason. If you expose eBPF-powered features to users, wrap them in a controlled API and validate inputs.
Performance and Overhead
eBPF is fast, but it is not free. Every attached program adds work to its hook. A kprobe on a hot function or an XDP program on every packet can become the bottleneck if it does too much. The goal is to filter early, aggregate in kernel, and send only useful data to user space.
- Prefer sampling over full capture for high-volume events such as packet headers or syscalls.
- Use per-CPU maps to reduce contention for counters and histograms.
- Use ring buffers for event delivery when ordering and efficiency matter, but watch for backpressure.
- Avoid bpf_printk in production. It is useful for debugging, not for high-rate logging.
- Benchmark with realistic traffic. A microbenchmark on an idle node rarely predicts performance under load.
- Consider driver-mode XDP for high packet rates, but verify NIC support and fallback behavior.
- Watch tail latency. A program that adds microseconds per packet can still hurt when multiplied by millions of packets.
Verifier complexity is also a performance and reliability factor. Large programs with many branches may fail to load or take longer to verify. Break complex logic into multiple programs and maps when possible.
Reliability of the eBPF Layer
Once eBPF is in the critical path, it needs the same operational discipline as any other infrastructure component. Monitor load failures, verifier logs, map utilization, ring buffer drops, and CPU usage per program. Use bpftool to inspect live state. Export metrics for your own agent.
- Program load failures: track by kernel version, node image, and program name.
- Map pressure: alert before maps hit max entries.
- Event drops: measure ring buffer and perf event losses.
- Agent health: expose readiness and liveness checks that verify attachment, not just process uptime.
- Upgrade safety: support detach and rollback without rebooting nodes.
- CI matrix: test across supported kernels, architectures, and security profiles.
Design for graceful degradation. If an eBPF program cannot load, the node should not become unmanageable. Fall back to a reduced feature set, emit a clear error, and avoid crash loops.
Common Pitfalls
- Assuming kernel portability: CO-RE helps, but hooks, helper availability, and BTF completeness still vary.
- Attaching too many programs to hot hooks: small overheads compound.
- Forgetting map limits: unbounded state can exhaust kernel memory.
- Leaking file descriptors or pinned maps: agents that restart frequently can leave stale resources.
- Using unstable tracepoints: internal tracepoints can change between kernel versions.
- Ignoring namespace context: container IDs, cgroups, and network namespaces must be resolved correctly.
- Treating eBPF as unprivileged: most useful production programs require elevated capabilities.
- Skipping canaries: a bad program can affect every pod on a node.
Minimal Example: Trace execve
The following eBPF program prints the command name when a process executes a new program. It uses a tracepoint, which is more stable than a raw kprobe. In production, replace bpf_printk with a ring buffer event and include container and pod metadata.
#include <linux/bpf.h>
#include <bpf/bpf_helpers.h>
SEC("tracepoint/syscalls/sys_enter_execve")
int handle_execve(void *ctx) {
char comm[16];
bpf_get_current_comm(&comm, sizeof(comm));
bpf_printk("exec: %s", comm);
return 0;
}
char LICENSE[] SEC("license") = "GPL";
Compile it with Clang targeting bpf, then load it with bpftool, libbpf, or a Go/Rust loader. The verifier will reject it if the tracepoint is unavailable or if the program violates safety rules. That immediate feedback is one of eBPF greatest strengths: unsafe code does not reach the kernel.
Operational Checklist
- Define the exact question the eBPF program answers.
- Choose stable hooks and CO-RE where possible.
- Set map sizes, sampling rates, and ring buffer limits.
- Run with least privilege and document required capabilities.
- Canary on a small node pool before broad rollout.
- Monitor load failures, drops, CPU, and memory.
- Test across supported kernels, architectures, and security profiles.
- Provide a fallback mode that does not crash the workload.
- Treat the agent as security-critical code.
Conclusion
eBPF gives production teams a powerful way to see and shape system behavior from inside the kernel. It can replace brittle instrumentation, simplify Kubernetes networking, and enforce security policy at runtime. The winning approach is not to load every clever program you can find. It is to treat eBPF as a platform: stable toolchain, bounded resources, least privilege, canary rollouts, and honest observability of the observability layer itself. Do that, and eBPF becomes a durable part of your infrastructure rather than a fragile trick.

