| Age | Commit message (Collapse) | Author |
|
xiic_smbus_block_read_setup() recalculates i2c->rx_msg->len based on the
length byte returned by the device, but historically clobbered the PEC
byte expectation the SMBus core had baked into msg->len. That dropped
the PEC byte from the caller's buffer on the normal and chunked
receive-fifo branches.
Compute pec_len up-front as (i2c->rx_msg->len - 1) -- the trailing bytes
the caller has already accounted for beyond the length byte, 1 when the
SMBus core enabled PEC, and possibly more for an I2C_M_RECV_LEN request
coming from i2c-dev -- and add it to the new length in every branch:
- chunked: the trailing bytes do not fit in the Rx FIFO, so drain in
chunks. The guard becomes (rxmsg_len + pec_len > IIC_RX_FIFO_DEPTH)
rather than rxmsg_len alone, both because pec_len bytes also have to
fit and because it is what bounds rfd_set in the else branch below
to the 4 bits of XIIC_RFD_REG_OFFSET.
- padded (1 + rxmsg_len + pec_len < SMBUS_BLOCK_READ_MIN_LEN): the
hardware needs at least 3 bytes on the bus to exit the read cleanly
(the second byte is already being clocked in by the time the ISR
reads the length byte and is too late to NACK), so we still pad
rx_msg->len up to SMBUS_BLOCK_READ_MIN_LEN. The dummy trailing byte
that gets drained must then be trimmed off before handing the
message back to the SMBus core; otherwise i2c_smbus_check_pec()
reads buf[len-1] (= dummy) instead of the real PEC byte at buf[1]
and rejects every clean zero-length block read with -EBADMSG.
Record the true valid byte count in a new field
i2c->smbus_actual_len and trim rx_msg->len down to it in
xiic_smbus_trim_len(), called from both completion sites that clear
rx_msg: xiic_process()'s RX_FULL branch and xiic_recv_atomic(),
which drains the FIFO with interrupts off.
smbus_actual_len is per-receive state, so xiic_start_recv() clears
it before every receive. Only the padded branch ever sets it, and a
block read aborted by arbitration loss or a TX error never reaches
the completion site, so without that clear a stale value would trim
the length of an unrelated later read.
The condition is expressed in total bytes rather than the old
"(rxmsg_len == 1) || (rxmsg_len == 0)" so that a request carrying
more than one trailing byte does not get padded: padding records a
length the drain never reaches, which would hand the caller a byte
that was never received.
- normal: all trailing bytes fit in one FIFO fill. rfd_set gains
pec_len for the same reason the length does. Because the padded
branch above has already taken every case with fewer than
SMBUS_BLOCK_READ_MIN_LEN total bytes, rxmsg_len + pec_len is at
least 2 here and the subtraction cannot underflow the u8.
Fixes: e4c1ff772e1a ("i2c: xiic: Add smbus_block_read functionality")
Signed-off-by: Abdurrahman Hussain <abdurrahman@nexthop.ai>
Cc: <stable@vger.kernel.org> # v6.3+
Acked-by: Michal Simek <michal.simek@amd.com>
Signed-off-by: Andi Shyti <andi.shyti@kernel.org>
Link: https://patch.msgid.link/20260924-i2c-xiic-v7-1-df7e752332ef@nexthop.ai
|
|
mc_probe() acquires a reference to the remote processor with
rproc_get_by_phandle(), but mc_remove() does not release the reference.
rproc_shutdown() only balances the power reference acquired by rproc_boot();
it does not drop the device reference acquired by rproc_get_by_phandle(). As
a result, successful driver removal leaves the remoteproc reference
unbalanced.
Call rproc_put() during removal to release the reference acquired in
mc_probe().
[ bp: Massage commit message. ]
Fixes: d5fe2fec6c40d ("EDAC: Add a driver for the AMD Versal NET DDR controller")
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Reviewed-by: Radhey Shyam Pandey <radhey.shyam.pandey@amd.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260913053532.1324671-1-lgs201920130244@gmail.com
|
|
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/driver-core/driver-core
Pull driver core fix from Danilo Krummrich:
- Suppress spurious "debugfs is not initialized yet" boot warnings when
the caller passes an error parent to debugfs file creation; callers
propagating an earlier failure should not trigger the warning
* tag 'driver-core-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/driver-core/driver-core:
debugfs: don't warn about uninitialized debugfs for an error parent
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/andi.shyti/linux
Pull i2c fix from Andi Shyti:
- qcom-geni: select the correct source clock table entry
* tag 'i2c-fixes-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/andi.shyti/linux:
i2c: qcom-geni: Fix hardcoded clock index in SE_GENI_CLK_SEL
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tj/wq
Pull workqueue fix from Tejun Heo:
- Fix a NULL dereference in the flush dependency check when a worker
flushes outside a work item, such as from the OOM path during worker
creation
* tag 'wq-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/wq:
workqueue: Fix NULL current_pwq deref in flush dependency check
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup
Pull cgroup fix from Tejun Heo:
- A cpuset partition could claim CPUs an ancestor partition already
held exclusively. Restore the rejection an earlier change had turned
into a warning.
* tag 'cgroup-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup:
cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext
Pull sched_ext fix from Tejun Heo:
- The CPU topology helper for BPF schedulers took no buffer size, so
its structure couldn't grow without breaking schedulers built against
the older layout. Add a size argument.
* tag 'sched_ext-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext:
sched_ext: Add a size argument to scx_bpf_cid_topo() so struct scx_cid_topo can grow
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 fixes from Ingo Molnar:
- Fix preemption bugs in the SVSM vTPM guest implementation
(Melody Wang)
- Fix MCE-triggered hardware debug register corruption on
task migration (Masami Hiramatsu)
* tag 'x86-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
x86/mce: Fix hardware debug register corruption on task migration
x86/sev: Make vTPM SVSM calls preemption-safe
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull scheduler fixes from Ingo Molnar:
- Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang)
- Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen)
- Skip kernel threads for cache aware scheduling to rubustify the code
(Chen Yu)
- Refresh LLC capacity across CPU hotplug, to fix capacity
underestimation bug (Davi Chaves Azevedo)
- Account PSI IRQ time to the execution context, not the scheduling
context, to fix proxy scheduling accounting bug (Zhan Xusheng)
* tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
sched/core: Account PSI IRQ time to the execution context, not the scheduling context
sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug
sched/cache: Skip kernel threads for cache aware scheduling to rubustify the code
sched/cache: Introduce task_struct->sched_cache_grp to fix UAF
sched/cache: Decouple sched_cache_group from mm to fix UAF
sched/cache: Honor migrate_llc_task semantics in active load balance, to fix LLC mis-scheduling bug
sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC mis-scheduling bug
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fixes for KVM guest PEBS virtualization (Sean Christopherson)
- Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi)
- Fix Intel Panther Cove event scheduling constraints (Dapeng Mi)
- Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi)
- Rename two confusingly named PMU attributes (Dapeng Mi)
- Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim)
- Fix NULL pointer dereference crash in __perf_pmu_sched_task()
(Puranjay Mohan)
- Fix CPU-wide event scheduling (Puranjay Mohan)
- Fix x86 LBR branch entry generation (Puranjay Mohan)
* tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf/core: Fill branch entries with a single assignment
perf/core: Run sched_task() for PMUs with only CPU-wide events
perf/core: Fix NULL pmu_ctx passed to pmu->sched_task()
perf/core: Fix a refcount leak in attach_perf_ctx_data()
perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL
perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3
perf/x86/intel: Delete dead NVL PEBS data-source initcall
perf/x86/intel: Fix Panther Cove PEBS data-source snoop states
perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints
perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints
perf/x86/intel: Update arw_latency_data() mem-op direction handling
perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs()
perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused
perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED
perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
|
|
check_flush_dependency() uses current_wq_worker() to determine whether
the caller is a workqueue worker and then dereferences worker->current_pwq
to test whether the current workqueue is WQ_MEM_RECLAIM.
current_wq_worker() only means that %current has PF_WQ_WORKER set. A
kworker can reach check_flush_dependency() while it is not executing a
work item. One such path is worker_thread() acting as the pool manager,
where create_worker() does GFP_KERNEL allocation and the allocation path
invokes the OOM notifier. In that state worker->current_pwq is NULL
because current_pwq is set only by process_one_work() and cleared again
after the work function returns.
[ 416.760634][ T375] Call trace:
[ 416.760638][ T375] check_flush_dependency+0x80/0x120 (P)
[ 416.760648][ T375] __flush_work+0x98/0x224
[ 416.760657][ T375] flush_work+0x30/0x44
[ 416.760665][ T375] ...
[ 416.760710][ T375] blocking_notifier_call_chain+0x58/0xa0
[ 416.760719][ T375] out_of_memory+0xb4/0x458
[ 416.760730][ T375] __alloc_pages_may_oom+0x11c/0x1a8
[ 416.760739][ T375] __alloc_pages_slowpath+0x314/0x46c
[ 416.760746][ T375] __alloc_frozen_pages_noprof+0x110/0x1a4
[ 416.760753][ T375] new_slab+0x12c/0x484
[ 416.760759][ T375] ___slab_alloc+0x7a8/0xc7c
[ 416.760765][ T375] __slab_alloc+0x74/0xd8
[ 416.760772][ T375] __kmalloc_cache_node_noprof+0x2ac/0x304
[ 416.760779][ T375] alloc_worker+0x28/0x60
[ 416.760785][ T375] create_worker+0x4c/0x20c
[ 416.760790][ T375] worker_thread+0xe8/0x2b8
[ 416.760796][ T375] kthread+0x1a8/0x200
[ 416.760805][ T375] ret_from_fork+0x10/0x20
Guard the WQ_MEM_RECLAIM-worker warning with worker->current_pwq. If the
kworker is not currently executing a work item, there is no current
workqueue to diagnose with that warning. The PF_MEMALLOC warning is left
unchanged so explicit reclaim context flushing a !WQ_MEM_RECLAIM target
is still reported.
Fixes: fca839c00a12 ("workqueue: warn if memory reclaim tries to flush !WQ_MEM_RECLAIM workqueue")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Pavankumar Kondeti <pavan.kondeti@oss.qualcomm.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
can grow
scx_bpf_cid_topo() copies struct scx_cid_topo into a buffer the BPF program
sized from its own vmlinux.h while the verifier sizes the write from the
running kernel's BTF. The struct may grow and each growth then breaks every
scheduler built against the older layout, rejected at load or written past
its buffer. This is the usual hole for a struct handed to BPF, closed
elsewhere with a size argument, and it was missed here.
Take the buffer size, copy the smaller of it and the kernel's struct and set
the rest to -1. Accesses to the copy are CO-RE relocated, so the struct can
grow by appending fields, which its comment now states. The kfunc changes in
place: the cid interface is still being finalized and no released scheduler
uses the current form.
Fixes: e9b55af47edf ("sched_ext: Add topological CPU IDs (cids)")
Cc: stable@vger.kernel.org # v7.2+
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
|
|
Under altr_portb_setup() and socfpga_init_sdmmc_ecc(),
of_find_compatible_node() was being used to look up the sdmmc-ecc
node. This node wasn't being dropped using of_node_put().
altr_portb_setup() did not drop its reference under its success path
or on any error path.
socfpga_init_sdmmc_ecc() did an early return thereby skipping the
common exit label and thus leaking the reference.
Add the missing of_node_put() calls in altr_portb_setup(), and route
socfpga_init_sdmmc_ecc()'s success path through the common exit label.
Fixes: 911049845d70 ("EDAC, altera: Add Arria10 SD-MMC EDAC support")
Fixes: 788586efd116 ("EDAC/altera: Initialize peripheral FIFOs in probe()")
Closes: https://sashiko.dev/#/patchset/20260708091135.94114-1-rounakdas2025%40gmail.com
Signed-off-by: Rounak Das <rounakdas2025@gmail.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Acked-by: Dinh Nguyen <dinguyen@kernel.org>
Cc: stable@vger.kernel.org # 6.18+
Link: https://patch.msgid.link/20260926120846.35716-1-rounakdas2025@gmail.com
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux
Pull ata fixes from Niklas Cassel:
- Extend the quirk "no LPM on ATI" quirk, that is currently only
applied for Samsung drives, to include AMD controllers as well.
The AMD AHCI controllers are newer versions of the ATI AHCI
controllers, and these controllers still have LPM issues with
Samsung drives - LPM works with drives from other vendors (me)
- Fix errors in the libata.force parameter documentation (me)
- Verify the sense data descriptor lengths for ATA PASS-THROUGH
command, so that a malicious device cannot write past the buffer
length (Matthias)
- Mention the libata for-next branch in MAINTAINERS such that the
git ls-remote command done by get_maintainer.pl --self-test=scm
can verify it (Matthias)
* tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux:
MAINTAINERS: name the libata/linux for-next branch
ata: libata-scsi: bound the ATA passthru sense descriptor writes
ata: libata: Correct libata.force parameter documentation
ata: libata-core: Extend Samsung LPM quirk to AMD controllers
|
|
subpartitions_cpus conflict
When a remote partition is created underneath an existing local partition
via a non-partition (PRS_MEMBER) intermediate cgroup, update_prstate() sees
parent->partition_root_state == PRS_MEMBER and calls
remote_partition_enable().
Commit 86888c7bd117 ("cgroup/cpuset: Add warnings to catch inconsistency
in exclusive CPUs") replaced the cpumask_intersects(tmp->new_cpus,
subpartitions_cpus) error check in remote_partition_enable() with
WARN_ON_ONCE(). As a result, remote_partition_enable() emits a warning
and proceeds to enable the remote partition on CPUs that are already
owned by the ancestor local partition in subpartitions_cpus.
This can be reproduced on Linux 7.3.0-rc3 with:
mkdir -p /tmp/cg1
mount -t cgroup2 none /tmp/cg1
echo "+cpuset" > /tmp/cg1/cgroup.subtree_control
mkdir /tmp/cg1/A
echo 1 > /tmp/cg1/A/cpuset.cpus
echo 1 > /tmp/cg1/A/cpuset.cpus.exclusive
echo root > /tmp/cg1/A/cpuset.cpus.partition
echo "+cpuset" > /tmp/cg1/A/cgroup.subtree_control
mkdir /tmp/cg1/A/B
echo 1 > /tmp/cg1/A/B/cpuset.cpus
echo 1 > /tmp/cg1/A/B/cpuset.cpus.exclusive
echo "+cpuset" > /tmp/cg1/A/B/cgroup.subtree_control
mkdir /tmp/cg1/A/B/D
echo 1 > /tmp/cg1/A/B/D/cpuset.cpus
echo 1 > /tmp/cg1/A/B/D/cpuset.cpus.exclusive
echo root > /tmp/cg1/A/B/D/cpuset.cpus.partition
which triggers:
WARNING: kernel/cgroup/cpuset.c:1594 at remote_partition_enable+0x1c1/0x300
and leaves both /tmp/cg1/A and /tmp/cg1/A/B/D as active root partitions
claiming exclusive CPU 1.
Fix this by returning PERR_NOCPUS when tmp->new_cpus intersects
subpartitions_cpus in remote_partition_enable(), matching the error code
used by remote_cpus_update() for the same subpartitions_cpus conflict, and
add a regression test case to
tools/testing/selftests/cgroup/test_cpuset_prs.sh.
Tested in QEMU on Linux 7.3.0-rc3 using the reproducer above and
tools/testing/selftests/cgroup/test_cpuset_prs.sh.
Fixes: 86888c7bd117 ("cgroup/cpuset: Add warnings to catch inconsistency in exclusive CPUs")
Suggested-by: Guopeng Zhang <guopeng.zhang@linux.dev>
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Reviewed-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci
Pull PCI fixes from Bjorn Helgaas:
- Make BAR resize work even for devices where no upstream bridge is
visible to the OS, which fixes an amdgpu regression on SolidRun
HoneyComb, which doesn't expose Root Ports to the OS (Liz Fong-Jones)
- Omit bus properties in dynamic OF nodes when a bridge has no
subordinate bus, which fixes early boot hangs caused by NULL pointer
dereferences with CONFIG_PCI_DYNAMIC_OF_NODES enabled (Angel J)
- Disable enhanced atomics on AMD NBIO 7.7 and 7.11 to avoid silent
data corruption on 64-bit DMAs (Mario Limonciello)
* tag 'pci-v7.3-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci:
x86/PCI: Disable enhanced atomics on AMD NBIO 7.7 and 7.11
PCI: of_property: Omit bus properties without a subordinate bus
PCI: Fix BAR resize for devices on a root bus
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace
Pull probe fixes from Masami Hiramatsu:
- kprobes: Fix permanent hang when flushing the kprobe optimizer
Fix a deadlock when disabling kprobe optimization via sysctl or
debugfs where flushers hung waiting for optimizer_completion.
Replaced the completion with an optimizer_passes counter and
wait_var_event_mutex() under kprobe_mutex so concurrent flushers can
wait and wake up safely.
- fprobe: Terminate the fgraph_data list when the reservation is not
filled
Fix an issue where unused shadow stack data left uninitialized by
fprobe_fgraph_entry() was misparsed as stale fprobe headers on
return. Explicitly write a zero word to terminate the list and update
read_fprobe_header() to handle the zeroed slot properly.
- ftracetest: Fix unique symbol check in kprobe_non_uniq_symbol.tc
Fix false test failures in kprobe_non_uniq_symbol.tc on architectures
like s390 where a symbol exists once in core kernel but also in
modules. Anchor the /proc/kallsyms search regex to the end of the
line so that module symbols are not incorrectly counted.
* tag 'probes-fixes-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace:
kprobes: Fix permanent hang when flushing the kprobe optimizer
fprobe: Terminate the fgraph_data list when the reservation is not filled
selftests/ftrace: Fix unique symbol check in kprobe_non_uniq_symbol.tc
|
|
Pull kvm fixes from Paolo Bonzini:
"Arm:
- Invalidate the ITS translation cache when the guest changes the
base address of the ITS tables (Fuad Tabba)
- Skip saving ITS devices with device IDs that are out-of-bounds
rather than failing the entire ITS save ioctl (Fuad Tabba)
- Close race between VM teardown and invalidations of nested MMUs
when handling MMU operations that are allowed to block (Lorenzo
Stoakes)
- Various fixes for the handling of the host's untrusted SVE
configuration in pKVM (Fuad Tabba)
- Make sure that empty SMCCC ranges based at 0 are rejected by the
kvm_smccc_set_filter() (Karl Mehltretter)
- Revoke the host mapping for pKVM's private stack pages, along with
a new sanity check that all mappings in the hyp's private VA range
have been correctly marked as hyp-owned (Fuad Tabba)
- Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
concurrent vCPU initialization cannot relocate in-use MMUs. Defer
the freeing of shadow stage-2 MMUs to the point that no other users
(e.g. MMU notifier) could reference them (Marc Zyngier)
- Drop useless WARN when rejecting an unsupported ioctl for pKVM
(Fuad Tabba)
- Fix the steal_time selftest to install correctly-sized mappings for
non-4K hosts (Sebastian Ott)
- Correct mapping of fine-grained trap for GCSPOPX instruction (Mark
Brown)
- Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
guests (Karl Mehltretter)
RISC-V:
- Synchronize hrtimer during VCPU teardown
- Fix the conversion between vsip and hvip values
- Serialize IMSIC attributes with vCPU migration
- Release unused page after MMU invalidation
- Propagate interrupted G-stage faults to KVM user-space as EINTR
- Fix nested acceleration hfence entry update order
- Fix sdata leak and stale snapshot_addr in snapshot_set_shmem
- Preserve firmware counter value across PMU counter stop/start
- Report PMU snapshot write failure to the guest
- Fix perf-backed counter accounting across PMU stop and read
- Correctly propagate error of a hart status SBI call
s390:
- Ensure that accesses through kvm_arch_set_irq_inatomic mark as
dirty the pages that contain indicator and summary bits
- Fix compile warning for kvm_s390_update_cmma_dirty()
- Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
userspace
- Move s390_kvm_mmu_commit_memory_region() into
s390_kvm_mmu_prepare_memory_region() so that it can fail instead of
WARN
- Add missing srcu in kvm_s390_set_irq_state()
- Fix potential races in storage functions
- Fix race in _destroy_pages_crste()
- Fix issues in the handling of KVM interrupt and page resources,
when a queue that is assigned to a mediated device (mdev) is
removed from the host's AP configuration
- Fix loop condition in uv_find_secrets
- Prevent potential out-of-bounds read
x86:
- Fix a brown paper bag bug where KVM would incorrectly treat Intel
PMU MSRs as valid on AMD
- Fix a regression in the hardware disable selftest where it checked
the wrong macro when detecting glibc support (breaks at least musl)
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is
especially important for KVM_BUG_ON() flows, which often guard more
dangerous bugs
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a
bug where KVM would let userspace run a broken setup with stale
vmcs12 pages
- Fix a class of bugs where KVM would fail to fill kvm_run exit
fields if getting nested pages failed
- Treat reserved entries in the memory attributes xarray as "no
attributes", to fix false positives when checking for mixed
attributes
- Fix memcg accounting for the memory attributes xarray (the xarray
library subtly requires the xarray to be configured for accounting
upfront; the gfp flags taken at runtime are used only rarely)
- Don't pre-reserve xarray entries when storing empty attributes, as
storing NULL must not require memory allocation (KVM and other
subsystems heavily rely on this behavior)
- Fix a memory leak and a cache maintenance issue related to doing
intra-host migration on an SEV guest"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits)
KVM: SEV: Do cache maintenance on the source VM during intra-host migration
KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes
KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg
KVM: Don't treat reserved xarray entries as having memory attributes
KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths
KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails
KVM: arm64: Fix AArch32 DBGBXVR<n> handling
KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP
KVM: selftests: fix steal_time for arm64 with host page size > 4K
KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode
KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction
KVM: arm64: nv: Fix life cycle of the nested_mmus array
KVM: arm64: Check every private mapping is hyp-owned at pKVM init
KVM: arm64: Move the private VA allocation cursor to __io_map_next
KVM: arm64: Match hyp text by physical address in fix_host_ownership()
KVM: arm64: Transfer the hyp stack pages out of the host stage-2
KVM: arm64: selftests: Test empty SMCCC filter range at base 0
KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0
KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2
...
|
|
KVM fixes for 7.3-rcN
- Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
as valid on AMD.
- Fix a regression in the hardware disable selftest where it checked the wrong
macro when detecting glibc support (breaks at least musl).
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
important for KVM_BUG_ON() flows, which often guard more dangerous bugs.
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
where KVM would let userspace run a broken setup with stale vmcs12 pages.
- Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
getting nested pages failed.
- Treat reserved entries in the memory attributes xarray as "no attributes",
to fix false positives when checking for mixed attributes.
- Fix memcg accounting for the memory attributes xarray (the xarray library
subtly requires the xarray to be configured for accounting upfront; the gfp
flags taken at runtime are used only rarely).
- Don't pre-reserve xarray entries when storing empty attributes, as storing
NULL must not require memory allocation (KVM and other subsystems heavily
rely on this behavior).
|
|
Manually perform cache maintenance on the source VM during intra-host
migration to ensure no stale data is left in CPU caches after the VM is
destroyed. Because the source VM is "converted" to a non-SEV VM, KVM's
memory reclaim flows won't trigger cache maintenance, e.g. when all guest
memory is reclaimed in response to detaching from the mmu_notifier.
Note, relying on the destination VM to do cache maintenance isn't an option
as KVM doesn't require identical guest memory configurations, i.e. the
source VM may have access to memory that the destination VM does not.
Enforcing equivalent memory configurations is infeasible, as it would
require a *deep* comparison of memslots, e.g. to verify that not only are
the memslot identical, but what the memslots point at is also identical.
Fixes: b56639318bb2 ("KVM: SEV: Add support for SEV intra host migration")
Cc: stable@vger.kernel.org
Reported-by: Stefan Teodorescu <fane@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260923163721.1584779-3-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
In both altr_edac_a10_device_add() and altr_portb_setup(), the error path
freed the dci structure before releasing the devres group. Since the managed
single and double bit IRQ handlers use altdev(dci->pvt_info) as their data, an
IRQ firing between freeing dci and unregistering the IRQs could dereference
the freed memory.
Release the devres group first so the managed IRQs are unregistered
before the dci structure is freed.
Fixes: 911049845d70 ("EDAC, altera: Add Arria10 SD-MMC EDAC support")
Fixes: 588cb03ea208 ("EDAC, altera: Add Arria10 L2 Cache ECC handling")
Closes: https://sashiko.dev/#/patchset/20260719211238.589402-1-rosenp%40gmail.com
Assisted-by: LLM
Signed-off-by: Dinh Nguyen <dinguyen@kernel.org>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Cc: stable@vger.kernel.org ## 6.18+
Link: https://patch.msgid.link/20260911120627.2634225-5-dinguyen@kernel.org
|
|
Sashiko reports:
"If devres_open_group() fails, the function returns -ENOMEM without freeing the
dci structure allocated earlier with edac_device_alloc_ctl_info()."
Free the dci structure if devres_open_group() fails.
Fixes: c3eea1942a16 ("EDAC, altera: Add Altera L2 cache and OCRAM support")
Signed-off-by: Dinh Nguyen <dinguyen@kernel.org>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Cc: stable@vger.kernel.org # 6.18+
Link: https://patch.msgid.link/20260911120627.2634225-4-dinguyen@kernel.org
|
|
Sashiko reports:
"Does suppressing sysfs unbinding fully prevent the execution of freed __init
memory? If altr_sysmgr_regmap_lookup_by_phandle() returns -EPROBE_DEFER, the
probe is deferred until after __init memory is freed."
The a10 EDAC .setup callbacks (sdmmc, ethernet, nand, dma, usb, qspi) and
their helpers (altr_init_a10_ecc_device_type, altr_init_a10_ecc_block) were
marked __init. These run from the probe path, which may execute after init
memory is freed -- e.g. a probe deferred via -EPROBE_DEFER that only succeeds
once a late/module dependency appears, or a manual unbind/rebind. Calling
__init code then dereferences freed memory. Remove __init so these functions
remain valid at runtime.
Fixes: 788586efd116 ("EDAC/altera: Initialize peripheral FIFOs in probe()")
Assisted-by: LLM
Signed-off-by: Dinh Nguyen <dinguyen@kernel.org>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Cc: stable@vger.kernel.org # 6.18+
Link: https://patch.msgid.link/20260911120627.2634225-3-dinguyen@kernel.org
|
|
The driver must remain bound; unbinding and re-binding it would erase active
system memory.
Remove the .remove functions because they will not ever get used.
Fixes: 588cb03ea208 ("EDAC, altera: Add Arria10 L2 Cache ECC handling")
Signed-off-by: Dinh Nguyen <dinguyen@kernel.org>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Cc: stable@vger.kernel.org # 6.18+
Link: https://patch.msgid.link/20260911120627.2634225-2-dinguyen@kernel.org
|
|
Pull drm fixes from Dave Airlie:
"While most of this is AI inspired fixes for error handling paths,
leaks and use after frees, there are some normal things.
nouveau has probably the biggest changes with some fixes to stabilise
runtime suspend/resume on 570 firmware which regressed after we moved
from 535, there are some fixes to stackframe issues seen with amdgpu,
and otherwise the usual bunch of i915/xe/amdgpu fixes, and some
virtio-gpu fixes.
Hopefully it will start to quiten down a bit from here.
client:
- fix restore of partially initialized client
i915:
- Fix incorrect RCU teardown order leading to endless loop
- Fix DP MST TU and FEC handling for disconnected streams
- Fix selective fetch disable, again
- Fix export namespace for kunit helpers
- Workaround eDP flicker on a specific laptop model
xe:
- CRI throttle reasons report
- TLB invalidation at wedge
- SVM eviction and VM close
- Display corruption on LNL on Xen PV
- W/a fix and addition
amdgpu:
- Display ref count fix
- Userq fixes
- VCN 4, 5 reset fixes
- Fixes for various error paths
- Stack frame size fixes for various combinations of compilers and
configs
amdkfd:
- Possible UAF fix
nouveau:
- runtime suspend/resume fixes for newer firmware
- rcu free the scheduler
- fix VRAM pinning
- fix double free
- fix reference leaks
- fix runtime PM leak
- fix cursor list usage problems
- fix HDMI config rejection without SCDC
virtio:
- fix a bunch of object/memory leaks in failure paths
- add pixel blend mode property to cursor plane
- revert prime buffers import
- sync shmem backing on guest transfers
imagination:
- propogate map failures properly
- fix page count in map interface
- clamp freelist reconstruction requests
ivpu:
- use separate flag for job timeout
bridge:
- samsung-dsim: fix TE GPIO lifetime for host attach"
* tag 'drm-fixes-2026-09-26' of https://gitlab.freedesktop.org/drm/kernel: (60 commits)
drm/amd/display: Bump frame warning limit for all builds of dml
drm/imagination: clamp freelist reconstruction requests
drm/imagination: Fix page count for page table for map() interface
drm/imagination: Propagate map failures correctly from pvr_mmu_map_sgl()
drm/amd/display: Bump frame warning limit for clang builds of dml
drm/amd/display: Relax DML frame limit with UBSAN
drm/amdgpu: Fix runtime PM leak in amdgpu_debugfs_test_ib_show()
drm/amdgpu: Fix last_update fence leak in amdgpu_vm_init()
drm/amdgpu: Fix acpi device leak in amdgpu_acpi_enumerate_xcc()
drm/amdgpu: Fix vmid_wait fence leak in amdgpu_ring_init()
drm/amdkfd: fix use-after-free and multi-container gap in kfd_dev_mapping
drm/amdgpu/vcn4.0.3: fix video_timeout unit mismatch in jpeg reset wait
drm/amdgpu/vcn5.0.1: fix video_timeout unit mismatch in jpeg reset wait
drm/amdgpu/userq: fix double jiffies conversion in hang detect timeout
drm/amdgpu: move userq fence wait out of signalling section
drm/amd/display: Fix dc stream excess put in dm_update_crtc_state()
drm/xe: Add wa_14025941587 to xe2, xe3 and xe3p platforms
drm/xe: harden adjust_idledly() against divide-by-zero and overflow
drm/xe: Limit sg segment size to PAGE_SIZE on Xen PV
drm/i915: fix incorrect RCU teardown order
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/wufan/ipe
Pull IPE fixes from Fan Wu:
"Two fixes for use-after-free issues found by recent LLM-assisted code
analysis.
- move successful policy load auditing under the new policy
directory's inode lock, preventing a concurrent policy deletion
from freeing the policy while it is still being audited
- protect the dm-verity root hash with RCU, preventing policy
evaluation from racing with root hash replacement during preresume"
* tag 'ipe-pr-20260925' of git://git.kernel.org/pub/scm/linux/kernel/git/wufan/ipe:
ipe: protect the dm-verity root hash with RCU
ipe: fix use-after-free when auditing a newly loaded policy
|
|
https://gitlab.freedesktop.org/drm/misc/kernel into drm-fixes
A number of fixes:
- bridge:
- samsung-dsim: fix GPIO lifetime
- client: Null pointer dereference fix
- imagination: error handling fix, page handling fix
- nouveau: fix reference leaks, double-frees, out-of-bounds accesses,
use-after-frees, don't reject config without SCDC, a number of
workarounds
- virtio: fix memory leak, reference leaks, null pointer dereference,
add pixel blend mode, cache coherency fix
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Maxime Ripard <self@mripard.dev>
Link: https://patch.msgid.link/arU22zzqUGDEco1y@houat
|
|
Writing 0 to /proc/sys/debug/kprobes-optimization while a kprobe is
jump-optimized never returns. The writer sleeps in D state forever with
kprobe_sysctl_mutex held, so any later read or write of that sysctl
hangs as well. For example, with vfs_read+9 as an optimizable address
in this build:
# cd /sys/kernel/tracing
# echo 'p:myprobe vfs_read+9' >> kprobe_events
# echo 1 > events/kprobes/myprobe/enable
# # wait until /sys/kernel/debug/kprobes/list shows [OPTIMIZED]
# echo 0 > /proc/sys/debug/kprobes-optimization
INFO: task sh:246 blocked for more than 10 seconds.
Call Trace:
<TASK>
__schedule+0x1176/0x4f70
schedule+0xdc/0x2c0
schedule_timeout+0x17b/0x260
wait_for_completion+0x173/0x3c0
wait_for_kprobe_optimizer_locked+0xbc/0x130
proc_kprobes_optimization_handler+0x156/0x1b0
proc_sys_call_handler+0x324/0x490
vfs_write+0x52d/0xfe0
ksys_write+0xff/0x200
do_syscall_64+0x106/0x630
entry_SYSCALL_64_after_hwframe+0x77/0x7f
</TASK>
...
INFO: task cat:265 is blocked on a mutex likely owned by task sh:246.
wait_for_kprobe_optimizer_locked() reinitializes optimizer_completion,
asks the optimizer thread to flush and sleeps in wait_for_completion().
The thread drains the (un)optimizing lists, but calls complete() only
if completion_done() is true, i.e. if the completion is already done,
which never happens while someone waits. disarm_all_kprobes() and
kprobe_trace_self_tests_init() wait the same way.
Calling complete() unconditionally would not be enough: the waiter
drops kprobe_mutex while it sleeps, and nothing else serializes the
sysctl handler against the debugfs "enabled" file. A second flusher
that still finds the lists non-empty, e.g. because a disabled probe is
queued for unoptimizing, reinitializes the completion under the first:
sysctl write debugfs "enabled" write
unoptimize_all_kprobes()
wait_for_kprobe_optimizer_locked()
init_completion(c)
mutex_unlock(&kprobe_mutex)
wait_for_completion(c)
disarm_all_kprobes()
wait_for_kprobe_optimizer_locked()
init_completion(c)
// c->wait is reset, the first
// waiter is off the queue
mutex_unlock(&kprobe_mutex)
wait_for_completion(c)
kprobe_optimizer()
complete(c)
// wakes the debugfs writer only
where c is &optimizer_completion. Lining up the two writes during an
optimizer pass loses the sysctl writer this way.
Replace the completion with a counter of optimizer passes, bumped at the
end of each pass and signalled with wake_up_var_locked(), both under
kprobe_mutex. A flusher samples the count and waits with
wait_var_event_mutex(), which drops kprobe_mutex only while sleeping, so
a new count means a whole pass ran in the meantime. Nothing is
reinitialized, so several flushers can sleep in the wait at once.
Link: https://lore.kernel.org/all/20260924092142.199198-1-parri.andrea@gmail.com/
Fixes: 73c12f209462 ("kprobes: Use dedicated kthread for kprobe optimizer")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Andrea Parri <parri.andrea@gmail.com>
Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
|
|
The T: entry for LIBATA SUBSYSTEM (Serial and Parallel ATA drivers)
names libata/linux without a branch. The repository's HEAD pointer
points to branch master, which has no active development. Active
development is on the for-next branch. Name the branch so the entry
identifies where development happens.
Documentation/process/submitting-patches.rst sends contributors to the
T: entry to find the tree to prepare patches against, so a branch-less
entry whose HEAD is already in mainline points them to the wrong branch.
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
Link: https://lore.kernel.org/r/20260925052329.2683619-1-matthias.goergens@gmail.com
Signed-off-by: Niklas Cassel <cassel@kernel.org>
|
|
qcom_geni_i2c_conf() writes a hardcoded 0 to SE_GENI_CLK_SEL, which
selects an index from the hardware clock performance table. This always
picks the first table entry regardless of the actual source clock
configuration. On platforms where the matching entry is not at index 0,
the wrong source clock divider is active and the I2C bus runs at an
incorrect frequency.
Use geni_se_clk_freq_match() in geni_i2c_clk_map_idx() to find the
performance table index for the source clock (32 MHz or 19.2 MHz). Store
the resolved index in a new clk_idx field in geni_i2c_dev and write it
to SE_GENI_CLK_SEL instead of the hardcoded 0.
Fixes: 37692de5d523 ("i2c: i2c-qcom-geni: Add bus driver for the Qualcomm GENI I2C controller")
Signed-off-by: Viken Dadhaniya <viken.dadhaniya@oss.qualcomm.com>
Cc: <stable@vger.kernel.org> # v4.19+
Reviewed-by: Mukesh Kumar Savaliya <mukesh.savaliya@oss.qualcomm.com>
Signed-off-by: Andi Shyti <andi.shyti@kernel.org>
Link: https://patch.msgid.link/20260921-i2c-fix-se-clk-conf-v2-1-8b5537ceff2d@oss.qualcomm.com
|
|
A fork rejected by the pids controller increments the counter reported by
pids.events. When local event accounting is selected, however, pids_event()
returns after notifying only events_local_file, leaving pids.events pollers
asleep.
On legacy hierarchies, pids.events.local does not exist. With
pids_localevents, pids.events reports the same local counter. In both
cases, pids.events changes without generating a notification.
This can be reproduced with a pids_localevents mount:
mkdir /tmp/test
mount -t cgroup2 -o pids_localevents none /tmp/test
mkdir /tmp/test/t
echo 1 > /tmp/test/t/pids.max
cat /tmp/test/t/pids.events # max 0
timeout 3 inotifywait -e modify /tmp/test/t/pids.events &
sh -c 'echo $$ > /tmp/test/t/cgroup.procs; (true &)' 2>/dev/null
wait
cat /tmp/test/t/pids.events # max 1
Without this patch, inotifywait times out without reporting an event.
Notify pids.events before returning from the local event path.
Fixes: 3f26a885a068 ("cgroup/pids: Add pids.events.local")
Cc: stable@vger.kernel.org # v6.11+
Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
Commit 42bc6935339b ("thunderbolt: stream: Support IOCB_NOWAIT in
non-blocking I/O as well") added support for IOCB_NOWAIT but forgot to
actually announce it as part of the file->f_mode. Add this now so users
such as io_uring can actually take advantage of IOCB_NOWAIT.
Fixes: 42bc6935339b ("thunderbolt: stream: Support IOCB_NOWAIT in non-blocking I/O as well")
Signed-off-by: Mika Westerberg <mika.westerberg@linux.intel.com>
|
|
arch_freq_get_on_cpu() computes the product of the frequency scale and
the reference frequency as a u64, but assigns it to an unsigned int
before shifting it back down:
freq = scale * arch_scale_freq_ref(cpu);
freq >>= SCHED_CAPACITY_SHIFT;
The product is truncated to 32 bits before the shift, so the result
wraps once arch_scale_freq_ref() exceeds 2^32 / SCHED_CAPACITY_SCALE,
i.e. 4194304 kHz.
On a Snapdragon X2 Elite (Glymur) laptop, whose boost OPP is 4723200
kHz, cpuinfo_avg_freq reports 524283 kHz instead of ~4723200 kHz while
the CPU demonstrably runs at the boost frequency: a fixed workload
completes in 1.72 s at the 4723200 kHz OPP versus 2.01 s at 4032000
kHz, matching the 1.171 frequency ratio.
Shift the u64 product and narrow only at the return.
Fixes: 16d1e27475f6 ("arm64: Provide an AMU-based version of arch_freq_get_on_cpu")
Reviewed-by: Dietmar Eggemann <dietmar.eggemann@arm.com>
Signed-off-by: Oleg Keri <okerixx@gmail.com>
Signed-off-by: Will Deacon <will@kernel.org>
|
|
modify_prot_start_ptes() performs the break-before-make TLB invalidation
required by erratum 2645198 with __flush_tlb_range(), whose third
argument is the end address of the range. It passes nr * PAGE_SIZE
instead of addr + nr * PAGE_SIZE, so __do_flush_tlb_range() computes the
page count as (nr * PAGE_SIZE - addr) >> PAGE_SHIFT.
For addr > nr * PAGE_SIZE that subtraction underflows, the page count
exceeds the batching limit and the flush degenerates to flush_tlb_mm(),
so a single-page mprotect broadcasts an ASID-wide invalidation and a
full-range mmu notifier call.
For addr <= nr * PAGE_SIZE only [addr, nr * PAGE_SIZE) is invalidated,
and when the cleared batch starts below nr * PAGE_SIZE the tail is left
in the TLB. The workaround then no longer covers the whole batch, and
for addr == nr * PAGE_SIZE the flush is empty.
On affected Cortex-A715 CPUs, this can corrupt ESR_ELx and FAR_ELx on the
next instruction abort caused by a permission fault.
Pass addr + nr * PAGE_SIZE as the end address.
Fixes: 7efa1cd5f89b5 ("arm64: add batched versions of ptep_modify_prot_start/commit")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Andrea Parri <parri.andrea@gmail.com>
Reviewed-by: Dev Jain <dev.jain@arm.com>
Signed-off-by: Will Deacon <will@kernel.org>
|
|
scheduling context
psi_account_irqtime() has two callers which share rq->psi_irq_time, and
they disagree about the context: __schedule() passes the outgoing rq->curr,
sched_tick() passes rq->donor. Under proxy execution the donor is blocked
on a mutex while rq->curr burns the CPU.
The tick charges PSI_IRQ_FULL to the donor's cgroup and advances the
timestamp, so the call from __schedule() then finds delta <= 0 and charges
nothing. The delta is not counted twice, it lands on the wrong cgroup.
Pass rq->curr, which is what the call read before commit af0c8b2bf67b
("sched: Split scheduler and execution contexts") renamed 'curr' to
'donor' across sched_tick(). Without CONFIG_SCHED_PROXY_EXEC the two rq
members are a union, so this only changes anything where that option is set,
and it depends on EXPERT.
Fixes: af0c8b2bf67b ("sched: Split scheduler and execution contexts")
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Link: https://patch.msgid.link/20260918132915.1236312-1-zhanxusheng@xiaomi.com
|
|
When an ATA PASS-THROUGH command to an ATAPI device fails, the sense
buffer holds the device's REQUEST SENSE reply, and
ata_scsi_set_passthru_sense_fields() trusts its additional length
byte, sb[7], when adding the ATA Status Return descriptor. A faulty
or malicious device can use that to make the kernel read and write
past the 96-byte buffer in three ways:
- scsi_sense_desc_find() is passed sb[7] + 8 as the buffer length, so
its clamp against sb[7] does nothing and the walk runs off the end.
- A type-9 descriptor found near the end is filled in unchecked.
- A new descriptor at sb[8 + len] needs len + 22 bytes, not len + 14,
so len 75..82 writes up to 8 bytes past the end.
Reproduced with KASAN under qemu, with the emulated ATAPI REQUEST SENSE
reply patched:
BUG: KASAN: slab-out-of-bounds in scsi_sense_desc_find+0x1a5/0x210
BUG: KASAN: slab-out-of-bounds in ata_scsi_qc_complete+0x1a15/0x1a50
Both are gone with this patch, and a valid descriptor is still filled
in.
Fixes: 97981926224a ("ata: libata-scsi: Do not overwrite valid sense data when CK_COND=1")
Cc: stable@vger.kernel.org
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
Link: https://lore.kernel.org/r/20260923175203.1576825-1-matthias.goergens@gmail.com
Signed-off-by: Niklas Cassel <cassel@kernel.org>
|
|
Multiple users report data corruption during 64-bit DMA transfers on
systems with AMD NBIO 7.7 and 7.11 controllers.
This occurs when BIOS enables AMD "enhanced atomic operations" on PCIe Root
Ports. When enhanced atomics are enabled, any 64-bit DMA access may be
corrupted.
Disable enhanced atomics using SMN for NBIO 7.7 and 7.11 based models.
Reported-by: Mikael Etienne <mikael1022bzh@gmail.com>
Closes: https://lore.kernel.org/178789300872.392066.15963676631650361573@gmail.com/
Reported-by: Arthur Husband <artmoty@gmail.com>
Closes: https://lore.kernel.org/20260406222335.379935-1-artmoty@gmail.com/
Reported-by: Alvin Lim <alvinwylim@gmail.com>
Closes: https://lore.kernel.org/20260621100844.1224301-1-alvinwylim@gmail.com/
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
[bhelgaas: commit log, s/IOVA/DMA/ in comment]
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Cc: stable@vger.kernel.org
Cc: David Laight <david.laight.linux@gmail.com>
Cc: John Smith <imjohnsmith4000@gmail.com>
Cc: Lennert Buytenhek <kernel@wantstofly.org>
Cc: Niklas Cassel <cassel@kernel.org>
Cc: Roland Waltersson <roland.waltersson@netinsight.net>
Link: https://patch.msgid.link/20260908190600.226485-2-mario.limonciello@amd.com
|
|
In exc_machine_check_user(), local_db_save() and local_db_restore() are
invoked in the outer entry stubs (DEFINE_IDTENTRY_MCE_USER,
DEFINE_FREDENTRY_MCE, and DEFINE_IDTENTRY_RAW), surrounding
exc_machine_check_user().
However, exc_machine_check_user() calls irqentry_exit_to_user_mode(), which
handles pending thread work and may schedule() if TIF_NEED_RESCHED is set. If
the task migrates to another CPU during schedule(), local_db_restore() runs on
the new CPU with the dr7 state saved from the old CPU. This corrupts the new
CPU's DR7 hardware debug register and leaves the old CPU's DR7 disabled. In
short, local_db_save() and local_db_restore() pair must be run on the same
CPU.
To fix this, move local_db_save() and local_db_restore() inside
exc_machine_check_user() and exc_machine_check_kernel(). In
exc_machine_check_user(), DR7 is saved and restored strictly around
do_machine_check() to avoid schedule() during migration. In
exc_machine_check_kernel(), local_db_save() is called at the entry point to
prevent early memory accesses from triggering nested #DB exceptions, and
restored on all exits.
Fixes: cd840e424f27 ("x86/entry, mce: Disallow #DB during #MC")
Assisted-by: LLM
Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Cc: <stable@kernel.org>
Link: https://patch.msgid.link/179005109564.388919.3937970081044095776.stgit@devnote2
|
|
When a descendant scheduler enters bypass mode, its tasks are parked in
the bypass DSQs of the nearest non-bypassing ancestor, which is then
responsible for running them. On behalf of such a non-bypassing host,
scx_dispatch_sched() consumes those bypass DSQs from two places: the
attempt made every SCX_BYPASS_HOST_NTH dispatches, and the
end-of-dispatch fallback that keeps the CPU from going idle while
bypassed descendants still have tasks queued.
The former increments SCX_EV_SUB_BYPASS_DISPATCH but the latter does
not, even though both perform the same scx_consume_dispatch_q() on the
same bypass DSQ. The descendant bypass dispatches done by the fallback
are therefore missing from the counter exposed via sysfs,
scx_dump_state() and the scx_bpf_events() kfunc, which under-reports the
actual number of such dispatches.
Add the missing __scx_add_event() so the fallback counts them too. When
@sch itself is bypassing, scx_dispatch_sched() takes the earlier
self-bypass branch and returns before reaching these host paths; that
mode is accounted for by SCX_EV_BYPASS_DISPATCH at enqueue time and is
intentionally left unchanged.
Fixes: 025b1bd41965 ("sched_ext: Implement hierarchical bypass mode")
Signed-off-by: Liang Luo <luoliang@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
__init_el2_fgt2() writes one mask to both HDFGRTR2_EL2 and HDFGWTR2_EL2.
PMZR_EL0 is write-only, so its trap bit, nPMZR_EL0, exists only in
HDFGWTR2_EL2 and is therefore never set: a PMZR_EL0 write from the host
traps to EL2, where the nVHE hypervisor has no handler and BUG()s. The
kernel never writes PMZR_EL0, but kernel.perf_user_access=1 has the PMU
driver set PMUSERENR_EL0.UEN for a task with a user-read event, so a
write from EL0 reaches the trap and takes the host down without a panic
message.
Accumulate the HDFGWTR2_EL2 bits separately, as __init_el2_fgt() already
does for HDFGWTR_EL2, and set nPMZR_EL0 with the other FEAT_PMUv3p9
bits.
Fixes: 858c7bfcb35e1 ("arm64/boot: Enable EL2 requirements for FEAT_PMUv3p9")
Cc: stable@vger.kernel.org
Signed-off-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Anshuman Khandual <anshuman.khandual@arm.com>
Reviewed-by: Oliver Upton <oupton@kernel.org>
Signed-off-by: Will Deacon <will@kernel.org>
|
|
perf_clear_branch_entry_bitfields() clears the bitfields of struct
perf_branch_entry one by one and leaves from/to alone, since callers
overwrite those straight away. The list has to be kept in sync with the
struct by hand and has already fallen behind: new_type and priv were
added to perf_branch_entry and never added here.
Only BRBE writes those two, and neither for every record.
brbe_set_perf_entry_type() leaves new_type alone for a branch type it
does not recognise, and priv is not set for source-only records.
arm_pmuv3.c allocates the per-CPU branch stack with kmalloc(), so such a
record reaches userspace with whatever the slot held: uninitialised
kmalloc() data on the first pass over the buffer, the previous record's
values after that. Nothing under arch/x86/events/ writes either field,
so only arm64 is affected.
Assign the whole entry at each site instead. Everything not named is
then zero, and there is no list to keep in sync. The bitfields add up to
exactly 64 bits, so the struct has no padding to leave undefined.
perf_clear_branch_entry_bitfields() has no callers left, so remove it.
perf_entry_from_brbe_regset() assigns an empty literal instead, since it
fills from/to conditionally. PERF_BR_SPEC_NA is 0, so dropping the
explicit spec assignment changes nothing.
Fixes: b190bc4ac9e6 ("perf: Extend branch type classification")
Fixes: 5402d25aa571 ("perf: Capture branch privilege information")
Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Tested-by: Yifan Wu <wuyifan50@huawei.com>
Link: https://patch.msgid.link/20260810133540.1947118-4-puranjay@kernel.org
|
|
perf_pmu_sched_task() returns early when cpuctx->task_ctx is set and
leaves the work to perf_ctx_sched_task_cb(), which only walks
ctx->pmu_ctx_list. A PMU whose events are all CPU-wide is not on that
list, so nothing calls its sched_task(). With
perf record -b -e cycles -a -- ls
armv8pmu_sched_task() is skipped on every switch to a task that has a
perf context but no event on that PMU, and BRBE records leak across the
task boundary. intel_pmu_lbr_add() calls perf_sched_cb_inc()
unconditionally too, so LBR records leak the same way on x86.
Drop the early return and skip only the CPCs that
perf_ctx_sched_task_cb() handles. That one needs a gate of its own to
make the split exact: it tests cpc->sched_cb_usage, which
perf_sched_cb_inc() sets per CPU for every branch stack user, so a task
with an event for that PMU pinned to another CPU would be handled twice.
On x86 the second __intel_pmu_lbr_restore() finds lbr_stack_state ==
LBR_NONE and calls intel_pmu_lbr_reset(), throwing away the callstack
the first one restored.
cpc->task_epc is set only while a task context is scheduled in, and
there is one epc per PMU on ctx->pmu_ctx_list, so the two gates are
inverses.
For the CPCs perf_pmu_sched_task() picks up, the callback now runs
outside the perf_ctx_disable() and perf_ctx_enable() pair in
perf_event_context_sched_in(). __perf_pmu_sched_task() disables the PMU
around the call itself.
Fixes: bd2756811766 ("perf: Rewrite core context handling")
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Tested-by: Yifan Wu <wuyifan50@huawei.com>
Link: https://patch.msgid.link/20260810133540.1947118-3-puranjay@kernel.org
Cc: stable@vger.kernel.org
|
|
perf_pmu_sched_task() returns early when cpuctx->task_ctx is set, and
cpc->task_epc is only non-NULL while a task context is scheduled in on
this CPU. __perf_pmu_sched_task() therefore always passes NULL:
Unable to handle kernel NULL pointer dereference at virtual address 00
pc : armv8pmu_sched_task+0x14/0x50
Call trace:
armv8pmu_sched_task+0x14/0x50 (P)
perf_pmu_sched_task+0xac/0x108
__perf_event_task_sched_out+0x6c/0xe0
Pass &cpc->epc instead, the CPU-wide context for this PMU, which the
function already dereferences a few lines up to find pmu.
armv8pmu_sched_task() is the only in-tree implementation that
dereferences the argument, and it only reads ->pmu, so the oops needs
BRBE, added in v6.17.
Fixes: bd2756811766 ("perf: Rewrite core context handling")
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Tested-by: Yifan Wu <wuyifan50@huawei.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260810133540.1947118-2-puranjay@kernel.org
|
|
The attach_perf_ctx_data() can race on global and !global cases. The
global case is protected by global_ctx_data_rwsem and shares a single
reference count using perf_ctx_data.global field.
But when it races with !global case, it may miss to set the global field
and result in a reference count leak.
CPU1 CPU2
----------------------------------------------------------------
attach_task_ctx_data(.global=1) attach_task_ctx_data(.global=0)
cd1 = alloc_perf_ctx_data(); cd2 = alloc_perf_ctx_data();
// { .global = 0, .refcount = 1 };
try_cmpxchg(); // success,
// task->perf_ctx_data = cd2
try_cmpxhg(); // fail; old = cd2
refcount_inc_not_zero(&old->refcount); // success
// old.refcount = 2
free_perf_ctx_data(cd1);
Then later detach_global_ctx_data() will see the data but it's not
marked as global, so it won't call detach_task_ctx_data().
Fixes: 506e64e710ff ("perf: attach/detach PMU specific data")
Assisted-by: Sashiko.dev:Gemini-3.1-pro
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://patch.msgid.link/20260920231639.11910-1-namhyung@kernel.org
|
|
Skip the xarray reservation loop when clearing all memory attributes, as
storing NULL only erases the entry and never needs to allocate, so no
reservation (and no cleanup of a failed one) is required in that case.
Suggested-by: Sean Christopherson <seanjc@google.com>
Cc: David Ballesteros <davimaba.v@proton.me>
Signed-off-by: Zeng Chi <zengchi@kylinos.cn>
Link: https://patch.msgid.link/20260921102442.1232375-1-zeng_chi911@163.com
[sean: split to separate patch]
Signed-off-by: Sean Christopherson <seanjc@google.com>
|
|
NVL introduces Offmodule Response events in place of the legacy
Offcore Response events, but it still exposes the inherited
offcore_rsp PMU attribute for programming the corresponding MSR data.
Rename the NVL PMU attribute to offmodule_rsp so the sysfs interface
matches the underlying event name and avoids user & tooling confusion.
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Link: https://patch.msgid.link/20260917015234.981153-13-dapeng1.mi@linux.intel.com
|
|
DMR introduces Offmodule Response events in place of the legacy
Offcore Response events, but it still exposes the inherited
offcore_rsp PMU attribute for programming the corresponding MSR data.
Rename the DMR PMU attribute to offmodule_rsp so the sysfs interface
matches the underlying event name and avoids user & tooling confusion.
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Link: https://patch.msgid.link/20260917015234.981153-12-dapeng1.mi@linux.intel.com
|
|
underestimation bug
The scheduler scales LLC capacity by the fraction of cache-sharing CPUs
covered by a domain:
llc_bytes = cache_size * span_weight / shared_weight
During CPU teardown, sched_cpu_deactivate() rebuilds scheduler domains
before cacheinfo_cpu_pre_down() removes the CPU from shared_cpu_map. The
new domains therefore use the old sharing weight. The later call to
sched_update_llc_bytes() looks up the departing CPU's sd_llc, which has
already been detached, and returns without correcting the surviving CPUs.
On a Ryzen 5 7535U with twelve logical CPUs sharing a 16 MiB LLC,
offlining one SMT sibling left the remaining CPUs with:
llc_bytes = floor(16777216 * 11 / 12) = 15379114 bytes
The correct capacity is still 16777216 bytes. On systems with active
cache-aware scheduling, an underestimated capacity can cause
exceed_llc_capacity() to reject aggregation for a process whose footprint
would fit. Unchanged cpuset partitions sharing the physical cache can
also retain stale capacity when a CPU comes online in another partition.
Pass the cache-sharing mask already retained by cacheinfo to the
scheduler update. Refresh every surviving CPU using its own LLC domain
so that each partition receives the correct share. This also preserves
the correction needed as cache-sharing maps grow during boot.
Keep the existing CPU-hotplug and scheduler-domain synchronization. The
update remains on the hotplug path; no steady-state scheduling operation
or persistent allocation is added.
Fixes: 7030513a0877 ("sched/cache: Calculate the LLC size and store it in sched_domain")
Signed-off-by: Davi Chaves Azevedo <davichazbh@gmail.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Reviewed-by: Chen Yu <yu.c.chen@intel.com>
Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com>
Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com>
Tested-by: Chen Yu <yu.c.chen@intel.com>
Tested-by: K Prateek Nayak <kprateek.nayak@amd.com>
Cc: <stable@kernel.org> # v7.2.x
Link: https://patch.msgid.link/6751d93e15889e624796c74db0bfe66603d60b1b.1790035273.git.tim.c.chen@linux.intel.com
|
|
code
Kernel thread should not be covered by cache aware scheduling as
it borrows the statistics from the user space thread. Filter the
kernel thread in account_mm_sched().
In theory a kernel thread does not have any valid
cache group, so !grp should gate the kernel thread.
Add the PF_KTHREAD check explicitly here for safety
reasons, to guard against future modifications and
to pair with task_tick_cache().
Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Cc: <stable@kernel.org> # 7.2.x
Link: https://patch.msgid.link/058f0c6ea7b991c177a17de347fa3157f25489a7.1790035273.git.tim.c.chen@linux.intel.com
|