summaryrefslogtreecommitdiffstats
path: root/include
AgeCommit message (Collapse)Author
11 dayshrtimer: Use the mask to clear TIF_HRTIMER_REARM from the exit workKarl Mehltretter
hrtimer_rearm_deferred_user_irq() removes the rearm bit from the local copy of the TIF work with: *tif_work &= ~TIF_HRTIMER_REARM; TIF_HRTIMER_REARM is the bit number (12), not the mask. That clears TIF_NOTIFY_SIGNAL (bit 2) and TIF_MEMDIE (bit 3) from the copy and leaves bit 12 set. So the function never returns true, and the copy handed to exit_to_user_mode_loop() lacks TIF_NOTIFY_SIGNAL. If that was the only work for the loop, the task returns to user space with the task work still pending. hrtimer_interrupt() sets TIF_HRTIMER_REARM every time, so the next tick repeats that. Task work queued with TWA_SIGNAL from a hrtimer callback stays pending until the task does a syscall, takes an interrupt without deferred rearm or has to reschedule. A task which polls the io_uring completion ring in user space sees a timeout completion after 30ms (median) instead of 70us. 2 CPU QEMU guest, HZ=1000. Use the mask. Fixes: 15dd3a948855 ("hrtimer: Push reprogramming timers into the interrupt return path") Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Assisted-by: LLM Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260919144027.70250-1-kmehltretter@gmail.com
13 daysMerge tag 'sched-urgent-2026-09-27' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull scheduler fixes from Ingo Molnar: - Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang) - Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen) - Skip kernel threads for cache aware scheduling to rubustify the code (Chen Yu) - Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug (Davi Chaves Azevedo) - Account PSI IRQ time to the execution context, not the scheduling context, to fix proxy scheduling accounting bug (Zhan Xusheng) * tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: sched/core: Account PSI IRQ time to the execution context, not the scheduling context sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug sched/cache: Skip kernel threads for cache aware scheduling to rubustify the code sched/cache: Introduce task_struct->sched_cache_grp to fix UAF sched/cache: Decouple sched_cache_group from mm to fix UAF sched/cache: Honor migrate_llc_task semantics in active load balance, to fix LLC mis-scheduling bug sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC mis-scheduling bug
13 daysMerge tag 'perf-urgent-2026-09-27' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull perf events fixes from Ingo Molnar: - Fixes for KVM guest PEBS virtualization (Sean Christopherson) - Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi) - Fix Intel Panther Cove event scheduling constraints (Dapeng Mi) - Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi) - Rename two confusingly named PMU attributes (Dapeng Mi) - Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim) - Fix NULL pointer dereference crash in __perf_pmu_sched_task() (Puranjay Mohan) - Fix CPU-wide event scheduling (Puranjay Mohan) - Fix x86 LBR branch entry generation (Puranjay Mohan) * tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: perf/core: Fill branch entries with a single assignment perf/core: Run sched_task() for PMUs with only CPU-wide events perf/core: Fix NULL pmu_ctx passed to pmu->sched_task() perf/core: Fix a refcount leak in attach_perf_ctx_data() perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3 perf/x86/intel: Delete dead NVL PEBS data-source initcall perf/x86/intel: Fix Panther Cove PEBS data-source snoop states perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints perf/x86/intel: Update arw_latency_data() mem-op direction handling perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs() perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
14 daysMerge tag 'ata-7.3-rc5' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux Pull ata fixes from Niklas Cassel: - Extend the quirk "no LPM on ATI" quirk, that is currently only applied for Samsung drives, to include AMD controllers as well. The AMD AHCI controllers are newer versions of the ATI AHCI controllers, and these controllers still have LPM issues with Samsung drives - LPM works with drives from other vendors (me) - Fix errors in the libata.force parameter documentation (me) - Verify the sense data descriptor lengths for ATA PASS-THROUGH command, so that a malicious device cannot write past the buffer length (Matthias) - Mention the libata for-next branch in MAINTAINERS such that the git ls-remote command done by get_maintainer.pl --self-test=scm can verify it (Matthias) * tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux: MAINTAINERS: name the libata/linux for-next branch ata: libata-scsi: bound the ATA passthru sense descriptor writes ata: libata: Correct libata.force parameter documentation ata: libata-core: Extend Samsung LPM quirk to AMD controllers
14 daysMerge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvmLinus Torvalds
Pull kvm fixes from Paolo Bonzini: "Arm: - Invalidate the ITS translation cache when the guest changes the base address of the ITS tables (Fuad Tabba) - Skip saving ITS devices with device IDs that are out-of-bounds rather than failing the entire ITS save ioctl (Fuad Tabba) - Close race between VM teardown and invalidations of nested MMUs when handling MMU operations that are allowed to block (Lorenzo Stoakes) - Various fixes for the handling of the host's untrusted SVE configuration in pKVM (Fuad Tabba) - Make sure that empty SMCCC ranges based at 0 are rejected by the kvm_smccc_set_filter() (Karl Mehltretter) - Revoke the host mapping for pKVM's private stack pages, along with a new sanity check that all mappings in the hyp's private VA range have been correctly marked as hyp-owned (Fuad Tabba) - Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that concurrent vCPU initialization cannot relocate in-use MMUs. Defer the freeing of shadow stage-2 MMUs to the point that no other users (e.g. MMU notifier) could reference them (Marc Zyngier) - Drop useless WARN when rejecting an unsupported ioctl for pKVM (Fuad Tabba) - Fix the steal_time selftest to install correctly-sized mappings for non-4K hosts (Sebastian Ott) - Correct mapping of fine-grained trap for GCSPOPX instruction (Mark Brown) - Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit guests (Karl Mehltretter) RISC-V: - Synchronize hrtimer during VCPU teardown - Fix the conversion between vsip and hvip values - Serialize IMSIC attributes with vCPU migration - Release unused page after MMU invalidation - Propagate interrupted G-stage faults to KVM user-space as EINTR - Fix nested acceleration hfence entry update order - Fix sdata leak and stale snapshot_addr in snapshot_set_shmem - Preserve firmware counter value across PMU counter stop/start - Report PMU snapshot write failure to the guest - Fix perf-backed counter accounting across PMU stop and read - Correctly propagate error of a hart status SBI call s390: - Ensure that accesses through kvm_arch_set_irq_inatomic mark as dirty the pages that contain indicator and summary bits - Fix compile warning for kvm_s390_update_cmma_dirty() - Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to userspace - Move s390_kvm_mmu_commit_memory_region() into s390_kvm_mmu_prepare_memory_region() so that it can fail instead of WARN - Add missing srcu in kvm_s390_set_irq_state() - Fix potential races in storage functions - Fix race in _destroy_pages_crste() - Fix issues in the handling of KVM interrupt and page resources, when a queue that is assigned to a mediated device (mdev) is removed from the host's AP configuration - Fix loop condition in uv_find_secrets - Prevent potential out-of-bounds read x86: - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs as valid on AMD - Fix a regression in the hardware disable selftest where it checked the wrong macro when detecting glibc support (breaks at least musl) - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially important for KVM_BUG_ON() flows, which often guard more dangerous bugs - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug where KVM would let userspace run a broken setup with stale vmcs12 pages - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if getting nested pages failed - Treat reserved entries in the memory attributes xarray as "no attributes", to fix false positives when checking for mixed attributes - Fix memcg accounting for the memory attributes xarray (the xarray library subtly requires the xarray to be configured for accounting upfront; the gfp flags taken at runtime are used only rarely) - Don't pre-reserve xarray entries when storing empty attributes, as storing NULL must not require memory allocation (KVM and other subsystems heavily rely on this behavior) - Fix a memory leak and a cache maintenance issue related to doing intra-host migration on an SEV guest" * tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits) KVM: SEV: Do cache maintenance on the source VM during intra-host migration KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg KVM: Don't treat reserved xarray entries as having memory attributes KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails KVM: arm64: Fix AArch32 DBGBXVR<n> handling KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP KVM: selftests: fix steal_time for arm64 with host page size > 4K KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction KVM: arm64: nv: Fix life cycle of the nested_mmus array KVM: arm64: Check every private mapping is hyp-owned at pKVM init KVM: arm64: Move the private VA allocation cursor to __io_map_next KVM: arm64: Match hyp text by physical address in fix_host_ownership() KVM: arm64: Transfer the hyp stack pages out of the host stage-2 KVM: arm64: selftests: Test empty SMCCC filter range at base 0 KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0 KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2 ...
2026-09-26Merge tag 'kvm-x86-fixes-7.3-rc5' of https://github.com/kvm-x86/linux into HEADPaolo Bonzini
KVM fixes for 7.3-rcN - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs as valid on AMD. - Fix a regression in the hardware disable selftest where it checked the wrong macro when detecting glibc support (breaks at least musl). - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially important for KVM_BUG_ON() flows, which often guard more dangerous bugs. - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug where KVM would let userspace run a broken setup with stale vmcs12 pages. - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if getting nested pages failed. - Treat reserved entries in the memory attributes xarray as "no attributes", to fix false positives when checking for mixed attributes. - Fix memcg accounting for the memory attributes xarray (the xarray library subtly requires the xarray to be configured for accounting upfront; the gfp flags taken at runtime are used only rarely). - Don't pre-reserve xarray entries when storing empty attributes, as storing NULL must not require memory allocation (KVM and other subsystems heavily rely on this behavior).
2026-09-26KVM: SEV: Do cache maintenance on the source VM during intra-host migrationgraftedSean Christopherson
Manually perform cache maintenance on the source VM during intra-host migration to ensure no stale data is left in CPU caches after the VM is destroyed. Because the source VM is "converted" to a non-SEV VM, KVM's memory reclaim flows won't trigger cache maintenance, e.g. when all guest memory is reclaimed in response to detaching from the mmu_notifier. Note, relying on the destination VM to do cache maintenance isn't an option as KVM doesn't require identical guest memory configurations, i.e. the source VM may have access to memory that the destination VM does not. Enforcing equivalent memory configurations is infeasible, as it would require a *deep* comparison of memslots, e.g. to verify that not only are the memslot identical, but what the memslots point at is also identical. Fixes: b56639318bb2 ("KVM: SEV: Add support for SEV intra host migration") Cc: stable@vger.kernel.org Reported-by: Stefan Teodorescu <fane@google.com> Signed-off-by: Sean Christopherson <seanjc@google.com> Message-ID: <20260923163721.1584779-3-seanjc@google.com> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-09-25Merge tag 'ipe-pr-20260925' of ↵graftedLinus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/wufan/ipe Pull IPE fixes from Fan Wu: "Two fixes for use-after-free issues found by recent LLM-assisted code analysis. - move successful policy load auditing under the new policy directory's inode lock, preventing a concurrent policy deletion from freeing the policy while it is still being audited - protect the dm-verity root hash with RCU, preventing policy evaluation from racing with root hash replacement during preresume" * tag 'ipe-pr-20260925' of git://git.kernel.org/pub/scm/linux/kernel/git/wufan/ipe: ipe: protect the dm-verity root hash with RCU ipe: fix use-after-free when auditing a newly loaded policy
2026-09-23perf/core: Fill branch entries with a single assignmentPuranjay Mohan
perf_clear_branch_entry_bitfields() clears the bitfields of struct perf_branch_entry one by one and leaves from/to alone, since callers overwrite those straight away. The list has to be kept in sync with the struct by hand and has already fallen behind: new_type and priv were added to perf_branch_entry and never added here. Only BRBE writes those two, and neither for every record. brbe_set_perf_entry_type() leaves new_type alone for a branch type it does not recognise, and priv is not set for source-only records. arm_pmuv3.c allocates the per-CPU branch stack with kmalloc(), so such a record reaches userspace with whatever the slot held: uninitialised kmalloc() data on the first pass over the buffer, the previous record's values after that. Nothing under arch/x86/events/ writes either field, so only arm64 is affected. Assign the whole entry at each site instead. Everything not named is then zero, and there is no list to keep in sync. The bitfields add up to exactly 64 bits, so the struct has no padding to leave undefined. perf_clear_branch_entry_bitfields() has no callers left, so remove it. perf_entry_from_brbe_regset() assigns an empty literal instead, since it fills from/to conditionally. PERF_BR_SPEC_NA is 0, so dropping the explicit spec assignment changes nothing. Fixes: b190bc4ac9e6 ("perf: Extend branch type classification") Fixes: 5402d25aa571 ("perf: Capture branch privilege information") Suggested-by: Peter Zijlstra <peterz@infradead.org> Signed-off-by: Puranjay Mohan <puranjay@kernel.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Tested-by: Yifan Wu <wuyifan50@huawei.com> Link: https://patch.msgid.link/20260810133540.1947118-4-puranjay@kernel.org
2026-09-22KVM: Don't pre-reserve xarray entries when storing empty/NULL attributesgraftedZeng Chi
Skip the xarray reservation loop when clearing all memory attributes, as storing NULL only erases the entry and never needs to allocate, so no reservation (and no cleanup of a failed one) is required in that case. Suggested-by: Sean Christopherson <seanjc@google.com> Cc: David Ballesteros <davimaba.v@proton.me> Signed-off-by: Zeng Chi <zengchi@kylinos.cn> Link: https://patch.msgid.link/20260921102442.1232375-1-zeng_chi911@163.com [sean: split to separate patch] Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-09-22perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rspgraftedDapeng Mi
DMR introduces Offmodule Response events in place of the legacy Offcore Response events, but it still exposes the inherited offcore_rsp PMU attribute for programming the corresponding MSR data. Rename the DMR PMU attribute to offmodule_rsp so the sysfs interface matches the underlying event name and avoids user & tooling confusion. Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Link: https://patch.msgid.link/20260917015234.981153-12-dapeng1.mi@linux.intel.com
2026-09-22sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity ↵Davi Chaves Azevedo
underestimation bug The scheduler scales LLC capacity by the fraction of cache-sharing CPUs covered by a domain: llc_bytes = cache_size * span_weight / shared_weight During CPU teardown, sched_cpu_deactivate() rebuilds scheduler domains before cacheinfo_cpu_pre_down() removes the CPU from shared_cpu_map. The new domains therefore use the old sharing weight. The later call to sched_update_llc_bytes() looks up the departing CPU's sd_llc, which has already been detached, and returns without correcting the surviving CPUs. On a Ryzen 5 7535U with twelve logical CPUs sharing a 16 MiB LLC, offlining one SMT sibling left the remaining CPUs with: llc_bytes = floor(16777216 * 11 / 12) = 15379114 bytes The correct capacity is still 16777216 bytes. On systems with active cache-aware scheduling, an underestimated capacity can cause exceed_llc_capacity() to reject aggregation for a process whose footprint would fit. Unchanged cpuset partitions sharing the physical cache can also retain stale capacity when a CPU comes online in another partition. Pass the cache-sharing mask already retained by cacheinfo to the scheduler update. Refresh every surviving CPU using its own LLC domain so that each partition receives the correct share. This also preserves the correction needed as cache-sharing maps grow during boot. Keep the existing CPU-hotplug and scheduler-domain synchronization. The update remains on the hotplug path; no steady-state scheduling operation or persistent allocation is added. Fixes: 7030513a0877 ("sched/cache: Calculate the LLC size and store it in sched_domain") Signed-off-by: Davi Chaves Azevedo <davichazbh@gmail.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Chen Yu <yu.c.chen@intel.com> Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com> Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com> Tested-by: Chen Yu <yu.c.chen@intel.com> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com> Cc: <stable@kernel.org> # v7.2.x Link: https://patch.msgid.link/6751d93e15889e624796c74db0bfe66603d60b1b.1790035273.git.tim.c.chen@linux.intel.com
2026-09-22sched/cache: Introduce task_struct->sched_cache_grp to fix UAFTim Chen
Add a sched_cache_grp pointer to task_struct so that scheduler code can access the cache group directly via the task, without going through mm->sched_cache_grp. This decouples the scheduler's hot-path accesses from the mm_struct. Each task holds its own refcount on the sched_cache_group, separate from the reference held by its mm_struct. The reference is acquired in copy_mm() (fork) and exec_mmap() (exec), and released in exit_mm(). This fixes the use-after-free when account_mm_sched() reaches the group through a task whose mm is being switched, as reported by Hyunwoo: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Convert all scheduler code in fair.c and exit.c to use p->sched_cache_grp instead of p->mm->sched_cache_grp. Keep the fork/exec/exit reference management out of the generic mm paths: add sched_cache_fork(), sched_cache_fork_cleanup(), sched_cache_exec_mmap() and sched_cache_exit_mm() in kernel/sched/cache_sched.c (with empty stubs for !CONFIG_SCHED_CACHE), so fs/exec.c, kernel/fork.c and kernel/exit.c each call one helper instead of open-coding the refcounting under #ifdef. Also add sched_cache_group_get() and task_cache_group_get(). Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing") Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/ Reported-by: Hyunwoo Kim <imv4bel@gmail.com> Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev> Co-developed-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@kernel.org> #7.2.x Link: https://patch.msgid.link/ae7081dc54736bf115215f9867abb2711a7403fb.1790035273.git.tim.c.chen@linux.intel.com
2026-09-22sched/cache: Decouple sched_cache_group from mm to fix UAFTim Chen
Currently the sched cache grouping is by mm and the scheduling statistics sched_cache_stat lives in the mm structure. This ties the life cycle of scheduling stats with mm. In account_mm_sched(), the scheduling stats are accessed by task->mm->sc_stat. However, a task may be switching mm on one CPU when another CPU is running account_mm_sched(), and possibly accessing the old mm that was freed. This problem was found when running tests with KASAN by Hyunwoo: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Instead of serializing the mm access by introducing extra acquisition of rq lock in the mm free path, extract sched_cache_stat from mm_struct, rename it as sched_cache_group and manage its life cycle apart from mm_struct with its own ref counting. This allows us in the next patch access sched_cache_group directly from task, and add a refcount on sched_cache_group when a task links to it. This prevents the use after free issue when accessing stale and released old mm and its sched cache stat a task switches to a new mm while account_mm_sched() is done elsewhere. The other benefit of this restructure is in the future, the grouping of tasks to a LLC would have the flexibility to be associated with a user defined grouping, or cgroup, cookie group, numa_group or others instead of just with a single mm address space. Rename sched_cache_stat to sched_cache_group and turn it into a refcounted object allocated from mm_struct. The mm_struct now holds a pointer (sched_cache_grp) to this object instead of embedding it. Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing") Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/ Reported-by: Hyunwoo Kim <imv4bel@gmail.com> Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev> Co-developed-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@kernel.org> #7.2.x Link: https://patch.msgid.link/91fd1e3266707c865bc9abecfb3e17bc676712df.1790035273.git.tim.c.chen@linux.intel.com
2026-09-22sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC ↵graftedTim Chen
mis-scheduling bug alb_break_llc() decides whether to break LLC preference during active load balance. It does so by testing that every runnable fair task on the source rq prefers its LLC: env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_runnable But the two counters cover different sets. nr_pref_llc_running is updated in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued, so it follows queued tasks. h_nr_runnable is updated in set_delayed()/ clear_delayed() and drops delay-dequeued tasks. So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks, alb_break_llc() returns false, and active balance is free to pull a task off its preferred LLC. Active balance only moves runnable tasks, and this is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB skips the per-task test in can_migrate_task(). The runnable set is the one we want. Fix it on the counter side. A task should be counted in nr_pref_llc_running exactly while it is both queued on its preferred LLC (pref_llc_queued) and runnable (!sched_delayed). Define that membership once in task_pref_llc_runnable(), and adjust the counter only through pref_llc_running_inc()/pref_llc_running_dec() from the four sites that change either input: account_llc_enqueue(), account_llc_dequeue(), set_delayed() and clear_delayed(). Gating every update on the same predicate keeps the delay, wake and dequeue paths from double-counting or underflowing; see the comments at those sites for the ordering. nr_llc_running and sd->llc_counts are not touched and stay on queued semantics. Fixes: 714059f79ff0 ("sched/cache: Handle moving single tasks to/from their preferred LLC") Closes: https://lore.kernel.org/lkml/20260827135000.735138-1-zhanxusheng@xiaomi.com/ Reported-by: Zhan Xusheng <zhanxusheng@xiaomi.com> Suggested-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Kayra Cizmeci <kayracizmeci@gmail.com> Cc: <stable@kernel.org> # v7.2.x Link: https://patch.msgid.link/06af61afedac32e6477f57feb4d658f6c411c3af.1790035273.git.tim.c.chen@linux.intel.com
2026-09-21ata: libata-core: Extend Samsung LPM quirk to AMD controllersNiklas Cassel
A Samsung SSD 870 QVO 8TB connected to an AMD 600 Series chipset SATA controller is reported to time out on STANDBY IMMEDIATE during system suspend with med_power_with_dipm enabled. The command completes when using max_performance instead. The existing Samsung LPM quirk only matches ATI controllers, leaving AMD controllers unaffected. Rename it to ATA_QUIRK_NO_LPM_ON_ATI_AND_AMD and extend the vendor check to AMD for the same Samsung SSD model patterns. Keep LPM behavior unchanged for other controller vendors, including Intel. Leave ATA_QUIRK_NO_NCQ_ON_ATI restricted to ATI, since the reported AMD issue concerns LPM rather than NCQ. Link: https://bugzilla.kernel.org/show_bug.cgi?id=221986 Reviewed-by: Damien Le Moal <dlemoal@kernel.org> Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org>> --- Link: https://lore.kernel.org/r/20260918124030.1962773-5-cassel@kernel.org Signed-off-by: Niklas Cassel <cassel@kernel.org>
2026-09-16ata: libata-scsi: fix ata_dsm_trim_pages() kernel-docgraftedNiklas Cassel
make htmldocs fails with: Documentation/driver-api/libata:604: ./drivers/ata/libata-scsi.c:2301: ERROR: Unexpected indentation. [docutils] The Return: section of ata_dsm_trim_pages() ends its first sentence with a colon and continues with an indented bullet list. reStructuredText requires a blank line before an indented block, so docutils chokes on the list. Simply adding the missing blank line does not work either: Return: is a kernel-doc "special section", which is terminated by the first blank line, so the bullet list would end up in the Description section, detached from the sentence introducing it. Spell the two bounds out as prose instead, so that the Return: section stays self-contained. Documentation-only change. Fixes: e64e6b5dc867 ("ata: libata-scsi: scale DSM TRIM payload by MAX PAGES PER DSM COMMAND") Reported-by: Thomas Huth <thuth@redhat.com> Closes: https://lore.kernel.org/linux-ide/b0a0b8e8-cc4f-4c7b-8bc5-fee0712d405a@redhat.com/ Reviewed-by: Damien Le Moal <dlemoal@kernel.org> Reviewed-by: Thomas Huth <thuth@redhat.com> Link: https://lore.kernel.org/r/20260916130031.29990-2-cassel@kernel.org Signed-off-by: Niklas Cassel <cassel@kernel.org>