In the Linux kernel, the following vulnerability has been resolved:
net/rds: acquire RDS_IN_XMIT in rds_tcp_reset_callbacks()
rds_tcp_reset_callbacks() quiesces the transmit path by setting the
path state to RDS_CONN_RESETTING and then waiting for RDS_IN_XMIT to
be sampled clear before swapping the underlying socket and calling
rds_send_path_reset().
Sampling the bit clear is not the same as owning it: rds_send_xmit()
can re-acquire RDS_IN_XMIT right after the wait_event() returns. Its
state recheck after taking the lock is a store-buffering pattern (the
resetter writes the state and reads the bit, the sender writes the
bit and reads the state) and acquire_in_xmit() is only an acquire
operation, so on weakly ordered architectures both sides can miss
each other's write and the transmit path then runs concurrently with
rds_send_path_reset() rewriting cp_xmit_* state - which is exactly
what the comment above rds_send_path_reset() tells its callers to
prevent.
Take the lock instead, hold it across the socket swap and
rds_send_path_reset(), and release it with a wake-up at the end. The
lock-ordering constraint documented above the wait still holds: the
lock is acquired before lock_sock(), so a sender inside tcp_sendmsg()
can never be waited on while we hold the socket lock.
Two details of the old code go away with the same change:
-
t_sock is now read only after the lock is acquired. The old code
cached it before waiting; the teardown in rds_conn_shutdown()
releases that socket and clears t_sock, so a pointer cached before
the wait can be stale by the time the accept path resumes. Reading
it under RDS_IN_XMIT is what makes the exclusion complete once the
teardown owns the same lock, which the next patch arranges; until
then the teardown still only samples the bit, and the two paths
remain as exposed to each other as they are today.
-
The old !osock early path called rds_send_path_reset() with no
serialization at all. It now runs under the lock like the normal
path. The conditional RDS_CONN_RESETTING transition of the
previous patch happens before the socket check either way: a path
found without a socket is either still connecting (its reconnect
worker blocked on t_conn_path_lock) and legitimately goes
RESETTING -> UP on the new socket, or it has been torn down
meanwhile and is dropped.
The in-function comment describing the old wait-based quiesce is
rewritten to describe the lock-based one, and the stale block comment
above the function (which still described a return value and an
incomplete list of t_sock writers) is refreshed to name all four
writers - the connect, accept, teardown and swap paths - and what
serializes each of them.
CVSS Vector: CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H
CVSS Score: 8.1
AV:N - rds_tcp_reset_callbacks() runs when a remote peer's TCP SYN to the RDS/TCP listener is accepted by rds_tcp_accept_one() while the local path already has t_sock set (the duelling-SYN case). The peer's routable TCP connection is what sets off the unsynchronized socket swap and rds_send_path_reset().
AC:H - The attacker must win a narrow store-buffering race between the resetter's test_bit(RDS_IN_XMIT) and rds_send_xmit()'s acquire_in_xmit()/state recheck, which only fails on weakly ordered CPUs. It also needs a duelling SYN that arrives while the local path is still CONNECTING, which the attacker cannot fully control.
PR:N - The RDS/TCP listener accepts incoming connections with no authentication. The accept path through rds_tcp_accept_one() into rds_tcp_reset_callbacks() needs only a reachable IP address and TCP port.
UI:N - No local user action is needed. The reset path runs from the listen socket's data_ready callback and accept worker, and rds_send_xmit() is driven by the send worker or by pongs to peer pings.
S:U - The corruption stays inside kernel memory under a single security authority. Crossing no VM or hardware isolation boundary.
C:H - rds_send_path_reset() puts cp_xmit_rm and zeroes the cp_xmit_ offsets while a concurrent rds_send_xmit() still dereferences that rds_message. This is a use-after-free, and a stale osock cached before the wait can be released twice. Reclaiming the freed object could leak kernel memory.
I:H - The use-after-free of the rds_message and of the stale socket, plus the torn cp_xmit_ transmit state, corrupt kernel heap objects that could be reclaimed and used for controlled writes.
A:H - A sender using a freed cp_xmit_rm, or a double sock_release() of the stale osock, will oops or panic the kernel.
| Attack Vector |
Network |
Scope |
Unchanged |
| Attack Complexity |
High |
Confidentiality Impact |
High |
| Privileges Required |
None |
Integrity Impact |
High |
| User Interaction |
None |
Availability Impact |
High |
AV:N - rds_tcp_reset_callbacks() runs when a remote peer's TCP SYN to the RDS/TCP listener is accepted by rds_tcp_accept_one() while the local path already has t_sock set (the duelling-SYN case). The peer's routable TCP connection is what sets off the unsynchronized socket swap and rds_send_path_reset().
AC:H - The attacker must win a narrow store-buffering race between the resetter's test_bit(RDS_IN_XMIT) and rds_send_xmit()'s acquire_in_xmit()/state recheck, which only fails on weakly ordered CPUs. It also needs a duelling SYN that arrives while the local path is still CONNECTING, which the attacker cannot fully control.
PR:N - The RDS/TCP listener accepts incoming connections with no authentication. The accept path through rds_tcp_accept_one() into rds_tcp_reset_callbacks() needs only a reachable IP address and TCP port.
UI:N - No local user action is needed. The reset path runs from the listen socket's data_ready callback and accept worker, and rds_send_xmit() is driven by the send worker or by pongs to peer pings.
S:U - The corruption stays inside kernel memory under a single security authority. Crossing no VM or hardware isolation boundary.
C:H - rds_send_path_reset() puts cp_xmit_rm and zeroes the cp_xmit_ offsets while a concurrent rds_send_xmit() still dereferences that rds_message. This is a use-after-free, and a stale osock cached before the wait can be released twice. Reclaiming the freed object could leak kernel memory.
I:H - The use-after-free of the rds_message and of the stale socket, plus the torn cp_xmit_ transmit state, corrupt kernel heap objects that could be reclaimed and used for controlled writes.
A:H - A sender using a freed cp_xmit_rm, or a double sock_release() of the stale osock, will oops or panic the kernel.
CVSS 3.1