Kernel kallsyms lookups sped up 7x for BPF tracing loads
Jim Cromie's three-patch series cuts bulk symbol attach from hundreds of milliseconds to tens, after production stalls in fleet observability tools.
A three-patch series from Jim Cromie accelerates Linux kallsyms name lookups by about 7x, cutting the cost of resolving tens of thousands of kernel symbols that BPF tracers and security agents attach at boot or service start.
Modern observability and security daemons such as CrowdStrike Falcon, Cilium, Datadog, and Falco resolve large batches of kernel functions by name. CrowdStrike hit the bottleneck in production: attaching only 50 kprobe.session programs stalled for 858 ms with a quarter of CPU time spent inside kallsyms. A BPF selftest that attaches across 64,000 symbols spent roughly 390 ms in raw lookup spin.
Cromie's changes drop average lookup latency from about 6,102 ns to 866 ns on a kernel with roughly 184,000 symbols, shrinking that 64k-symbol attach from ~390 ms to ~55 ms. The work attacks two inner-loop costs left after the 2022 move from linear scan to binary search. Lookups no longer fully decompress every candidate name into a stack buffer before comparing; a new on-the-fly compare walks compressed tokens and bails at the first mismatched character. Marker density in the packed name table rises from one every 256 symbols to one every 16, cutting average sequential hops per probe from 127.5 to 7.5 at a cost of about 42 KiB of read-only data. A small follow-on inlines and unrolls 24-bit sequence reconstruction on the hot path.
Andrew Morton welcomed the series on the BPF list, calling the result awesome. The trade-off favors the growing class of tools that perform bulk name resolution, where previous marker spacing tuned for rare oops backtraces had become the dominant cost.