freenode
Kernel & Low-Level

KVM hardens teardown against host memory corruption

Nested VMX could write guest state into the wrong process address space while a VM dies, and the fix also closes a PowerPC shadow-page leak.

Sean Christopherson has posted a second round of KVM patches that stop the hypervisor from corrupting an unrelated process's memory when a virtual machine is destroyed.

The failure mode is simple in effect even if the trigger is narrow. When a VM is dying, the task tearing it down often no longer runs in the original userspace address space that owns the guest. If KVM still tries to write "guest" memory at that moment, ordinary userspace copy helpers will write into whatever mapping happens to sit at the same address in the current process. Jim Mattson reported the concrete case: on nested VMX, freeing a vCPU that is still in L2 forces a synthetic nested VM-Exit, which flushes a cached shadow VMCS back through those copy paths and can therefore scribble on a completely different task.

Christopherson's series rejects guest-memory uaccess whenever the current mm does not match the VM's, and it tears memslots down early, replacing them with empty dummies so later arch teardown cannot usefully resolve real guest mappings. The same work also repairs PowerPC, which had been ignoring the common shadow-flush hook and therefore left guest page tables standing after the owning process exited. Remaining nested-VMX paths that still abuse the exit flow are short-circuited so the new checks can fire as warnings instead of silent corruption, and the uaccess guards are marked for stable kernels.

The practical result is that a KVM bug during VM destruction no longer turns into memory corruption of an innocent host process, and any leftover offenders become noisy rather than silent.