In the Linux kernel, the following vulnerability has been resolved:
xprtrdma: Decouple req recycling from RPC completion
rl_kref formerly served two distinct lifetimes through a single
refcount: it gated when a Reply could wake its RPC task, and it
gated when an rpcrdma_req could return to its free pool. The
marshal path took the Send-side reference only when SGEs needed
DMA-unmap (sc_unmap_count > 0), which made a Send carrying only
pre-registered buffers an exception: the Reply handler dropped
rl_kref from 1 to 0 and freed the req while the HCA might still
be DMA-reading from its send buffer.
Give rl_kref a narrower job. The RPC layer takes one reference
when slot allocation hands a req out. rpcrdma_prepare_send_sges()
takes a Send-side reference unconditionally after WR preparation
succeeds. xprt_rdma_free_slot() and xprt_rdma_bc_free_rqst() drop
the RPC-layer reference; rpcrdma_sendctx_unmap() drops the
Send-side reference. The req returns to its free pool only after
both owners have signed off.
The existing kref_init(&req->rl_kref) call in
rpcrdma_prepare_send_sges() is removed. Initialization moves to
the slot-allocation paths (xprt_rdma_alloc_slot and
rpcrdma_bc_rqst_get), and the release callback re-arms rl_kref
before the req returns to a free pool. A re-init in the marshal
path would discard the RPC-layer reference that already exists
on entry.
Three invariants follow:
-
Any rpcrdma_req held by an rpc_rqst has rl_kref >= 1.
xprt_rdma_alloc_slot(), rpcrdma_bc_rqst_get(), and the
backlog-wake branch in xprt_rdma_alloc_slot() each kref_init
rl_kref before publishing the req. Without this invariant,
an RPC task that aborts between slot allocation and marshal
(gss_refresh failure or signal during call_connect, for
example) would drive xprt_release() ->
xprt_rdma_free_slot() -> kref_put against a refcount of
zero, saturating refcount_t and stranding the slot.
-
The Send-side reference is taken only after WR prep
succeeds. A mapping failure in rpcrdma_prepare_send_sges()
runs rpcrdma_sendctx_cancel(), which DMA-unmaps the sendctx
and clears sc_req without touching rl_kref. The sendctx
ring walks in rpcrdma_sendctx_put_locked() and
rpcrdma_sendctxs_destroy() skip entries with sc_req == NULL,
so a burst of -EIO marshal failures cannot hold reqs off
rb_send_bufs.
-
The release callback re-arms rl_kref so the next consumer
enters with the invariant satisfied.
Replies now complete the RPC directly. rpcrdma_reply_handler()
calls rpcrdma_complete_rqst() in place of kref_put on the
non-LocalInv branch. The LocalInv branch already completes the
RPC from frwr_unmap_async() and is unaffected.
Because Send-side references can now outlive RPC completion,
connection teardown drains sendctx entries whose unsignaled
Sends never had a later signaled completion to walk the ring.
rpcrdma_sendctxs_destroy() walks the active range and runs
rpcrdma_sendctx_unmap() on each entry with a non-NULL sc_req
before the request buffers are reset, and is moved ahead of
rpcrdma_reqs_reset() in rpcrdma_xprt_disconnect() so the reqs
are still in their pre-reset state when the Send-side refs are
released.
The drain creates a teardown-ordering hazard on the backchannel
path. With the new lifetime, releasing a bc_prealloc req from
rpcrdma_req_release() re-adds it to bc_pa_list. The disconnect
in xprt_rdma_destroy() runs after xprt_destroy_backchannel() has
already emptied bc_pa_list, so the drained reqs would otherwise
leak. xprt_rdma_destroy() now runs xprt_rdma_bc_destroy(xprt, 0)
a second time after the disconnect to reclaim them.
CVSS Vector: CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H
CVSS Score: 9.8
AV:N - The bug is in the RPC-over-RDMA client (xprtrdma); a remote NFS/RDMA peer delivers Receive completions that invoke rpcrdma_reply_handler() over InfiniBand/RoCE/iWARP, so exploitation is via network protocol traffic from a malicious or compromised server.
AC:L - For inline Sends using only pre-registered buffers (sc_unmap_count==0), the reply path completes the RPC and returns the req to the free pool before Send completion; a remote peer can reliably trigger this by replying immediately to normal small NFS RPCs.
PR:N - Exploitation requires no privileges on the victim host; a remote malicious or compromised NFS/RDMA server can send crafted/fast replies over the established RDMA connection to drive the vulnerable completion path.
UI:N - No end-user interaction is needed at exploit time; once a host uses NFS-over-RDMA, triggering the bug is automatic during ordinary client I/O against the attacker-controlled server.
S:U - The flaw causes kernel heap UAF/DMA corruption within the NFS client's kernel context and does not by itself cross hypervisor, VM, or sandbox security boundaries.
C:H - Recycling rpcrdma_req while the HCA may still DMA-read its send buffers is a use-after-free that can expose or leak kernel memory when the slot is reallocated and buffers are reused under active DMA.
I:H - Premature return of rpcrdma_req to rb_send_bufs while Send WRs are in flight permits heap corruption and typical UAF exploitation paths for arbitrary kernel memory writes or control-flow hijack.
A:H - Concurrent DMA against freed/reused request buffers can provoke kernel oops/panic or wedged RPC/RDMA transport state, giving high availability impact on NFS-over-RDMA clients.
| Attack Vector |
Network |
Scope |
Unchanged |
| Attack Complexity |
Low |
Confidentiality Impact |
High |
| Privileges Required |
None |
Integrity Impact |
High |
| User Interaction |
None |
Availability Impact |
High |
AV:N - The bug is in the RPC-over-RDMA client (xprtrdma); a remote NFS/RDMA peer delivers Receive completions that invoke rpcrdma_reply_handler() over InfiniBand/RoCE/iWARP, so exploitation is via network protocol traffic from a malicious or compromised server.
AC:L - For inline Sends using only pre-registered buffers (sc_unmap_count==0), the reply path completes the RPC and returns the req to the free pool before Send completion; a remote peer can reliably trigger this by replying immediately to normal small NFS RPCs.
PR:N - Exploitation requires no privileges on the victim host; a remote malicious or compromised NFS/RDMA server can send crafted/fast replies over the established RDMA connection to drive the vulnerable completion path.
UI:N - No end-user interaction is needed at exploit time; once a host uses NFS-over-RDMA, triggering the bug is automatic during ordinary client I/O against the attacker-controlled server.
S:U - The flaw causes kernel heap UAF/DMA corruption within the NFS client's kernel context and does not by itself cross hypervisor, VM, or sandbox security boundaries.
C:H - Recycling rpcrdma_req while the HCA may still DMA-read its send buffers is a use-after-free that can expose or leak kernel memory when the slot is reallocated and buffers are reused under active DMA.
I:H - Premature return of rpcrdma_req to rb_send_bufs while Send WRs are in flight permits heap corruption and typical UAF exploitation paths for arbitrary kernel memory writes or control-flow hijack.
A:H - Concurrent DMA against freed/reused request buffers can provoke kernel oops/panic or wedged RPC/RDMA transport state, giving high availability impact on NFS-over-RDMA clients.
CVSS 3.1