freenode
Kernel & Low-Level

Linux shmem race corrupts page cache during hole punch

Fault-around can re-map large folios mid-punch, leaving stale mappings and exact 512-page RSS imbalances on production hosts.

A race in the Linux kernel's shared-memory (shmem/tmpfs) paths can leave pages mapped after they are deleted from the page cache and throw process RSS counters out of balance. Production hosts that punch holes in memfds for reclaim while other threads fault the same MAP_SHARED mapping have hit it on 6.12 and 6.18 with transparent huge pages enabled for shmem.

Ayush Ranjan reported the failures after seeing "still mapped when deleted" page-cache warnings (the kernel taints but does not oops) and, more often, paired RSS counter bugs of exactly one PMD-order folio: MM_FILEPAGES off by -512 and MM_SHMEMPAGES off by +512 when the mm is torn down. An earlier report of the same corruption needed roughly a hundred ballooning VMs to trip and stalled without a fix. Ranjan's single-memfd reproducer hits it in minutes on a large machine when khugepaged is tuned to re-collapse punched ranges aggressively; fork is not required.

Shmem already makes ordinary faults wait out an in-progress hole punch. Its fault-around path, which pre-maps neighboring pages, does not join that protocol and does not take the invalidate lock regular filesystems use against truncation. During a partial punch of a large folio the folio can split, and the pieces can stay briefly visible in the page cache. Fault-around can install PTEs on those pages before the punch finishes removing them, producing the stale mappings and the mismatched counters.

Andrew Morton posted a fix that serializes shmem fault-around against the hole-punch unmap and truncate sequence with the mapping invalidate lock. Because fault-around runs under RCU, the fault side takes a shared trylock and simply skips ahead to a normal fault if a punch holds the lock exclusively. Baolin Wang reproduced the race on current development trees and confirmed the patch stops it. Maintainers are still tightening the exact sequence that leaves a folio mapped at deletion, and Morton has added targeted diagnostics while that is clarified. Hosts that combine concurrent hole punch on shared memfds with aggressive shmem huge-page collapse remain exposed until a fix reaches stable kernels.