In the Linux kernel, the following vulnerability has been resolved:
KVM: x86: hyper-v: Clamp stimer deadline to avoid livelock
Fix an issue where userspace or the guest can program an Hyper-V
synthetic timer to have a deadline in the past via integer overflow,
preventing the CPU from making progress and triggering an RCU stall.
Hyper-V's SynIC exposes 4 per-vCPU synthetic timers to the
guest, which are emulated by KVM. Each is programmed through the
HV_X64_MSR_STIMERi_CONFIG and HV_X64_MSR_STIMERi_COUNT MSRs. Depending
on CONFIG, COUNT represents either the absolute expiration time or the
period of a periodic timer, both expressed in 100ns ticks. These timers
may be set both by the guest (WRMSR) and the host (KVM_SET_MSRS).
When the timer is enabled, stimer_start() translates COUNT to an
absolute monotonic deadline and arms an hrtimer. If COUNT is set to a
value close to U64_MAX, the deadline calculation can overflow.
<pre>
ktime_add_ns(ktime_now, 100 * (stimer->exp_time - time_now))
</pre>
This can result in a CPU livelock. stimer_start() arms the timer
via hrtimer_start() with a deadline in the past, which causes it to
immediately fire. The stimer callback then raises KVM_RQ_HV_STIMER, with
the intention of causing KVM to deliver a synthetic interrupt on the
next vCPU guest enter.
Then, once userspace issues KVM_RUN, vcpu_enter_guest() consumes the
request, calling kvm_hv_process_stimers(). This would normally disable
the timer via stimer_expiration() once the deadline is in the past.
However, the deadline comparison is done between the KVM reference
counter and stime->exp_time, which is a big value close to U64_MAX, so
this never happens for a few thousand years.
kvm_hv_process_timers() then re-arms the timer via stimer_start(), since
it was not disabled, which again fires immediately. Before entering
the guest, kvm_vcpu_exit_request() checks kvm_request_pending(),
which returns true due to the newly raised KVM_REQ_HV_STIMER. Then
vcpu_enter_guest() aborts the guest entry, returning early into
vcpu_run(), which loops back again into vcpu_enter_guest(), restarting
the cycle.
Since there are no manual yields in this loop, a task with SCHED_FIFO
may starve RCU grace-period kthreads, which exposes the stalls found
by syzcaller:
<pre>
rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
rcu: (detected by 1, t=10502 jiffies, g=14269, q=1142 ncpus=2)
rcu: All QSes seen, last rcu_preempt kthread activity 10500 (4294965239-4294954739), jiffies_till_next_fqs=1, root ->qsmask 0x0
rcu: rcu_preempt kthread starved for 10500 jiffies! g14269 f0x2 RCU_GP_WAIT_FQS(5) ->state=0x0 ->cpu=0
rcu: Unless rcu_preempt kthread gets sufficient CPU time, OOM is now expected behavior.
( ... )
Call Trace:
<IRQ>
__run_hrtimer kernel/time/hrtimer.c:1773 [inline]
__hrtimer_run_queues+0x408/0xc30 kernel/time/hrtimer.c:1841
hrtimer_interrupt+0x45b/0xaa0 kernel/time/hrtimer.c:1903
local_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1045 [inline]
__sysvec_apic_timer_interrupt+0x102/0x3e0 arch/x86/kernel/apic/apic.c:1062
instr_sysvec_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1056 [inline]
sysvec_apic_timer_interrupt+0xa1/0xc0 arch/x86/kernel/apic/apic.c:1056
</IRQ>
<TASK>
asm_sysvec_apic_timer_interrupt+0x1a/0x20 arch/x86/include/asm/idtentry.h:697
RIP: 0010:__raw_spin_unlock_irqrestore include/linux/spinlock_api_smp.h:152 [inline]
RIP: 0010:_raw_spin_unlock_irqrestore+0xa8/0x110 kernel/locking/spinlock.c:194
Code: 74 05 e8 0b f4 5f f6 48 c7 44 24 20 00 00 00 00 9c 8f 44 24 20 f6 44 24 21 02 75 4f f7 c3 00 02 00 00 74 01 fb bf 01 00 00 00 <e8> 23 6b 27 f6 65 8b 05 7c 60 5a 07 85 c0 74 40 48 c7 04 24 0e 36
RSP: 0018:ffffc900040a7320 EFLAGS: 00000206
RAX: 5de15cb931505900 RBX: 0000000000000a06 RCX: 5de15cb931505900
RDX: 0000000000000007 RSI: ffffffff8daa9dc3 RDI: 0000000000000001
RBP: ffffc900040a73b0 R08: ffffffff8fc3d0
</pre>
---truncated---
CVSS Vector: CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:H
CVSS Score: 7.1
AV:L - Guest WRMSR to HV_X64_MSR_STIMERi_CONFIG/COUNT is trapped into KVM (kvm_emulate_wrmsr -> kvm_hv_set_msr_common -> stimer_set_config/count); host userspace can also program the same MSRs via KVM_SET_MSRS, then KVM_RUN. Both paths are local KVM VM-exit/ioctl, not network, adjacent-radio, or physical.
AC:L - The attacker fully controls COUNT/CONFIG. A one-shot COUNT near U64_MAX deterministically overflows 100*(exp_time-time_now), so hrtimer_start() uses a past deadline and kvm_hv_process_stimers() re-arms forever. CONFIG_KVM_HYPERV defaults to Y and hv-synic/hv-stimer are standard on Windows KVM guests; no race or rare config is required.
PR:N - The highest-impact case is a cloud x86 KVM tenant whose VM already has Hyper-V CPUID and KVM_CAP_HYPERV_SYNIC (normal for Windows/hv-stimer guests). They WRMSR the stimer MSRs from guest CPL0 with no host root, init-namespace capability, or /dev/kvm access.
UI:N - The guest programs the synthetic timer during ordinary KVM_RUN; the host vCPU thread livelocks on the same vcpu_enter_guest path with no additional victim action such as mounting a filesystem or opening a file.
S:C - The retry loop runs in the host KVM vCPU thread (vcpu_run/vcpu_enter_guest) and can stall host CPUs, starve RCU, and OOM the machine, taking down the hypervisor and co-resident VMs and crossing the KVM guest-to-host security boundary.
C:N - This is an integer overflow in the stimer deadline calculation that only mis-arms an hrtimer; there is no out-of-bounds read, use-after-free, or other host memory disclosure primitive.
I:N - Host kernel memory is not written or corrupted; the overflowed deadline only produces a KVM_REQ_HV_STIMER/hrtimer retry loop with no arbitrary write or control-flow hijack.
A:H - stimer_start() re-arms an immediately expiring hrtimer, kvm_vcpu_exit_request() aborts VM-entry, and vcpu_run() loops without a dedicated yield, livelocking the host vCPU thread. A SCHED_FIFO VMM can starve RCU (syzkaller RCU stall and expected OOM), denying service to the host and co-located VMs.
| Attack Vector |
Local |
Scope |
Changed |
| Attack Complexity |
Low |
Confidentiality Impact |
None |
| Privileges Required |
None |
Integrity Impact |
None |
| User Interaction |
None |
Availability Impact |
High |
AV:L - Guest WRMSR to HV_X64_MSR_STIMERi_CONFIG/COUNT is trapped into KVM (kvm_emulate_wrmsr -> kvm_hv_set_msr_common -> stimer_set_config/count); host userspace can also program the same MSRs via KVM_SET_MSRS, then KVM_RUN. Both paths are local KVM VM-exit/ioctl, not network, adjacent-radio, or physical.
AC:L - The attacker fully controls COUNT/CONFIG. A one-shot COUNT near U64_MAX deterministically overflows 100*(exp_time-time_now), so hrtimer_start() uses a past deadline and kvm_hv_process_stimers() re-arms forever. CONFIG_KVM_HYPERV defaults to Y and hv-synic/hv-stimer are standard on Windows KVM guests; no race or rare config is required.
PR:N - The highest-impact case is a cloud x86 KVM tenant whose VM already has Hyper-V CPUID and KVM_CAP_HYPERV_SYNIC (normal for Windows/hv-stimer guests). They WRMSR the stimer MSRs from guest CPL0 with no host root, init-namespace capability, or /dev/kvm access.
UI:N - The guest programs the synthetic timer during ordinary KVM_RUN; the host vCPU thread livelocks on the same vcpu_enter_guest path with no additional victim action such as mounting a filesystem or opening a file.
S:C - The retry loop runs in the host KVM vCPU thread (vcpu_run/vcpu_enter_guest) and can stall host CPUs, starve RCU, and OOM the machine, taking down the hypervisor and co-resident VMs and crossing the KVM guest-to-host security boundary.
C:N - This is an integer overflow in the stimer deadline calculation that only mis-arms an hrtimer; there is no out-of-bounds read, use-after-free, or other host memory disclosure primitive.
I:N - Host kernel memory is not written or corrupted; the overflowed deadline only produces a KVM_REQ_HV_STIMER/hrtimer retry loop with no arbitrary write or control-flow hijack.
A:H - stimer_start() re-arms an immediately expiring hrtimer, kvm_vcpu_exit_request() aborts VM-entry, and vcpu_run() loops without a dedicated yield, livelocking the host vCPU thread. A SCHED_FIFO VMM can starve RCU (syzkaller RCU stall and expected OOM), denying service to the host and co-located VMs.
CVSS 3.1