In the Linux kernel, the following vulnerability has been resolved:
mm/slab: take n->list_lock in __slab_try_return_freelist() to avoid race
Commit ba7425312607 ("mm, slab: add an optimistic
__slab_try_return_freelist()") incorrectly assumed that nobody has freed
an object to the slab as long as slab->freelist is NULL and cmpxchg
succeeds.
However, as reported by Hyunwoo Kim [1], other CPUs might have freed
an object to the slab, insert the slab to the partial list, then
allocated an object from the slab, and be in the middle of removing
the slab from the list under n->list_lock.
Since __refill_objects_node() puts the slab back on pc.slabs
outside n->list_lock, it might insert the slab into that list while
the slab is concurrently being removed from n->partial.
This led to a list corruption [1]:
list_add corruption. next->prev should be prev
(ffff888100000248), but was dead000000000122.
(next=ffffea000416e410).
kernel BUG at lib/list_debug.c:29!
Oops: invalid opcode: 0000 [#1] SMP NOPTI
CPU: 1 UID: 65534 PID: 144 Comm: poc Not tainted
7.2.0-16172-gcf72cbb39da8-dirty #1 PREEMPT(lazy)
RIP: 0010:__list_add_valid_or_report+0x80/0xd0
...
Call Trace:
alloc_from_new_slab+0x183/0x300
slaballoc+0x31c/0x890
kmalloc_noprof+0x3d4/0x800
lsm_blob_alloc+0x2d/0x50
security_msg_msg_alloc+0x26/0x90
load_msg+0x1aa/0x210
do_msgsnd+0x91/0x800
do_syscall_64+0x109/0x5d0
entry_SYSCALL_64_after_hwframe+0x77/0x7f
...
Kernel panic - not syncing: Fatal exception
This is a classic ABA problem where cmpxchg succeeds but the state has
changed since __refill_objects_node() took the freelist from the slab.
As Vlastimil Babka mentioned [2], it should be rare to return more than
one slab (due to the racy read of slab->counters in
get_partial_node_bulk()). Therefore, instead of introducing additional
complexity, acquire and release n->list_lock twice in the worst case.
Return the slab directly to the partial list and hold n->list_lock
across the cmpxchg and add_partial(). This is similar to the initial
version of commit ba7425312607 [3]. This is enough to avoid the race as
the list manipulation is serialized by n->list_lock. While at it,
bring back unlikely() hint now that the condition is unlikely.
CVSS Vector: CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H
CVSS Score: 7.8
AV:L - The race is in __refill_objects_node()/__slab_try_return_freelist() in mm/slub.c, which runs when the slab allocator refills its per-CPU object cache for kmalloc. It is reached from local syscalls such as msgsnd() (do_msgsnd -> load_msg -> __kmalloc -> ___slab_alloc in the reporter's trace). No remote data is involved.
AC:L - The attacker's own threads on several CPUs do the allocating and freeing (e.g. msgsnd/msgrcv) that returns the slab to n->partial and pulls it off again while the refill re-links it outside n->list_lock. The attacker drives both sides of the race, and the reporter's reproducer hit it.
PR:L - No capability is needed. Any unprivileged process that does kmalloc-backed syscalls reaches the refill path, and the reported crash came from a process running as UID 65534 (nobody) calling msgsnd.
UI:N - The attacker triggers the race entirely with their own allocation and free syscalls. No other user has to do anything.
S:U - The corruption is in the kernel's own slab lists, and the impact stays within the kernel's security authority as ordinary local privilege escalation or crash.
C:H - The ABA race leaves slab->slab_list linked on both the local pc.slabs list and n->partial. The same slab can then be handed out twice, giving overlapping kmalloc objects (a use-after-free-class primitive) that can be groomed to read kernel memory.
I:H - A slab handed out twice lets an attacker put a controlled object (e.g. msg_msg) on top of a live kernel object in a general kmalloc cache. That enables arbitrary write and control-flow hijack, and the corrupted list pointers are themselves a write primitive.
A:H - The reproducer hits a list_add corruption BUG (next->prev = dead000000000122) in alloc_from_new_slab and the kernel panics with "Fatal exception". An unprivileged user can repeat this at will.
| Attack Vector |
Local |
Scope |
Unchanged |
| Attack Complexity |
Low |
Confidentiality Impact |
High |
| Privileges Required |
Low |
Integrity Impact |
High |
| User Interaction |
None |
Availability Impact |
High |
AV:L - The race is in __refill_objects_node()/__slab_try_return_freelist() in mm/slub.c, which runs when the slab allocator refills its per-CPU object cache for kmalloc. It is reached from local syscalls such as msgsnd() (do_msgsnd -> load_msg -> __kmalloc -> ___slab_alloc in the reporter's trace). No remote data is involved.
AC:L - The attacker's own threads on several CPUs do the allocating and freeing (e.g. msgsnd/msgrcv) that returns the slab to n->partial and pulls it off again while the refill re-links it outside n->list_lock. The attacker drives both sides of the race, and the reporter's reproducer hit it.
PR:L - No capability is needed. Any unprivileged process that does kmalloc-backed syscalls reaches the refill path, and the reported crash came from a process running as UID 65534 (nobody) calling msgsnd.
UI:N - The attacker triggers the race entirely with their own allocation and free syscalls. No other user has to do anything.
S:U - The corruption is in the kernel's own slab lists, and the impact stays within the kernel's security authority as ordinary local privilege escalation or crash.
C:H - The ABA race leaves slab->slab_list linked on both the local pc.slabs list and n->partial. The same slab can then be handed out twice, giving overlapping kmalloc objects (a use-after-free-class primitive) that can be groomed to read kernel memory.
I:H - A slab handed out twice lets an attacker put a controlled object (e.g. msg_msg) on top of a live kernel object in a general kmalloc cache. That enables arbitrary write and control-flow hijack, and the corrupted list pointers are themselves a write primitive.
A:H - The reproducer hits a list_add corruption BUG (next->prev = dead000000000122) in alloc_from_new_slab and the kernel panics with "Fatal exception". An unprivileged user can repeat this at will.
CVSS 3.1