freenode
Kernel & Low-Level

NUMA balancing filters were stranding hot pages on CXL memory

Meta reports major latency and bandwidth gains after decoupling tier promotion from socket-placement heuristics in the kernel scheduler.

Filters meant to stop wasteful NUMA placement scans were also blocking memory-tier promotion, leaving hot pages stuck on slower CXL memory in production. Gregory Price of Meta posted a backportable fix series on the Linux kernel list after measuring the damage on a host with 768 GB of DRAM and 256 GB of CXL running two roughly 430 GB database workloads under combined NUMA balancing.

Before the changes, the machine saturated CXL at 40-45 GB/s while DRAM stayed around 150-200 GB/s, and request latency sat above 5 ms with long tails. Afterward, DRAM sustained 250 GB/s or more, CXL traffic fell to about 5-10 GB/s, and latency dropped into the 800 us to 2 ms range.

The root problem is that the global balancing mode only says which mechanisms are on. It cannot express whether a given VMA walk is for east-west socket placement or north-south tier promotion. Heuristics that skip inactive VMAs, read-only file mappings, and certain shared folios make sense for placement, but they permanently strand hot data on the slow tier when tiering is enabled. In one case a hot 20 GB hash table sat entirely on CXL; in another, most of a shared main binary never left the slow tier.

The series separates scan intent so promotion-only walks can proceed past those placement filters, while ordinary placement behavior is preserved. Shared and executable folios may still promote from a low tier to the top tier without reopening east-west bounce paths. PID-inactive and read-only file VMAs receive promotion scans when tiering is on, and placement eligibility is tracked apart from scan continuation so large multi-threaded VMAs are not starved.

David Hildenbrand has acked several of the core fixes. The functional changes are aimed at stable backports; two follow-on cleanups are independent.