| Age | Commit message (Collapse) | Author |
|
Add a sched_cache_grp pointer to task_struct so that scheduler code
can access the cache group directly via the task, without going
through mm->sched_cache_grp. This decouples the scheduler's hot-path
accesses from the mm_struct.
Each task holds its own refcount on the sched_cache_group, separate
from the reference held by its mm_struct. The reference is acquired
in copy_mm() (fork) and exec_mmap() (exec), and released in exit_mm().
This fixes the use-after-free when account_mm_sched() reaches the group
through a task whose mm is being switched, as reported by Hyunwoo:
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Convert all scheduler code in fair.c and exit.c to use
p->sched_cache_grp instead of p->mm->sched_cache_grp.
Keep the fork/exec/exit reference management out of the generic mm
paths: add sched_cache_fork(), sched_cache_fork_cleanup(),
sched_cache_exec_mmap() and sched_cache_exit_mm() in
kernel/sched/cache_sched.c (with empty stubs for !CONFIG_SCHED_CACHE),
so fs/exec.c, kernel/fork.c and kernel/exit.c each call one helper
instead of open-coding the refcounting under #ifdef. Also add
sched_cache_group_get() and task_cache_group_get().
Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/
Reported-by: Hyunwoo Kim <imv4bel@gmail.com>
Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev>
Co-developed-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Cc: <stable@kernel.org> #7.2.x
Link: https://patch.msgid.link/ae7081dc54736bf115215f9867abb2711a7403fb.1790035273.git.tim.c.chen@linux.intel.com
|
|
Currently the sched cache grouping is by mm and the scheduling statistics
sched_cache_stat lives in the mm structure. This ties the life cycle
of scheduling stats with mm.
In account_mm_sched(), the scheduling stats are accessed by
task->mm->sc_stat. However, a task may be switching mm on one CPU when
another CPU is running account_mm_sched(), and possibly accessing the
old mm that was freed. This problem was found when running tests with
KASAN by Hyunwoo:
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Instead of serializing the mm access by introducing extra acquisition of
rq lock in the mm free path, extract sched_cache_stat from mm_struct,
rename it as sched_cache_group and manage its life cycle apart from
mm_struct with its own ref counting. This allows us in the next patch
access sched_cache_group directly from task, and add a refcount
on sched_cache_group when a task links to it. This prevents the use
after free issue when accessing stale and released old mm and its
sched cache stat a task switches to a new mm while account_mm_sched()
is done elsewhere.
The other benefit of this restructure is in the future, the grouping of
tasks to a LLC would have the flexibility to be associated with a user
defined grouping, or cgroup, cookie group, numa_group or others instead
of just with a single mm address space.
Rename sched_cache_stat to sched_cache_group and turn it into a refcounted
object allocated from mm_struct. The mm_struct now holds a pointer
(sched_cache_grp) to this object instead of embedding it.
Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/
Reported-by: Hyunwoo Kim <imv4bel@gmail.com>
Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev>
Co-developed-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Cc: <stable@kernel.org> #7.2.x
Link: https://patch.msgid.link/91fd1e3266707c865bc9abecfb3e17bc676712df.1790035273.git.tim.c.chen@linux.intel.com
|
|
LLC mis-scheduling bug
Cache aware scheduling introduced the migrate_llc_task migration type to direct
tasks toward their preferred LLC, but its semantics can be lost when passive
load balance falls back to active load balance (ALB). This may allow ALB to
select a candidate whose preferred LLC does not match the destination, moving
it away from its preferred LLC.
Example scenario:
src_rq has two runnable tasks, p1 and p2. p1 prefers dst_rq (dst_llc), while p2
prefers src_rq (src_llc). In this case, migrate_llc_task is set because src_rq
has at least one task, p1, that wants to migrate to dst_rq. In ALB,
can_migrate_task() finds p2 and returns true for it, thus moving p2 out of its
preferred LLC.
Solution:
The CPU stopper in ALB constructs a fresh lb_env that does not inherit
migration_type from the passive load-balance pass. Two approaches are
possible:
(a) Add a new member to struct rq so ALB can inherit migrate_llc_task
from the passive LB that triggered it.
(b) Define a new flag LBF_ACTIVE_LB_LLC and select the stopper callback
at kick time to preserve the migration semantics across the
asynchronous boundary.
We choose (b) because it avoids passing migration_type through the
stopper, which would affect the meaning of migration_type for
delayed-dequeue tasks.
Fixes: e4c9a4cb244a ("sched/cache: Add migrate_llc_task migration type for cache-aware balancing")
Suggested-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Lu Wang <wanglu.priv@gmail.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com>
Reviewed-by: Chen Yu <yu.c.chen@intel.com>
Cc: <stable@kernel.org> # v7.2.x
Link: https://patch.msgid.link/cb39f64a17fc2b76097264aaec74a2d6dfff4315.1790035273.git.tim.c.chen@linux.intel.com
|
|
mis-scheduling bug
alb_break_llc() decides whether to break LLC preference during active
load balance. It does so by testing that every runnable fair task on the
source rq prefers its LLC:
env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_runnable
But the two counters cover different sets. nr_pref_llc_running is updated
in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued,
so it follows queued tasks. h_nr_runnable is updated in set_delayed()/
clear_delayed() and drops delay-dequeued tasks.
So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted
in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks,
alb_break_llc() returns false, and active balance is free to pull a task
off its preferred LLC. Active balance only moves runnable tasks, and this
is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB
skips the per-task test in can_migrate_task(). The runnable set is the one
we want.
Fix it on the counter side. A task should be counted in
nr_pref_llc_running exactly while it is both queued on its preferred LLC
(pref_llc_queued) and runnable (!sched_delayed). Define that membership
once in task_pref_llc_runnable(), and adjust the counter only through
pref_llc_running_inc()/pref_llc_running_dec() from the four sites that
change either input: account_llc_enqueue(), account_llc_dequeue(),
set_delayed() and clear_delayed(). Gating every update on the same
predicate keeps the delay, wake and dequeue paths from double-counting
or underflowing; see the comments at those sites for the ordering.
nr_llc_running and sd->llc_counts are not touched and stay on queued
semantics.
Fixes: 714059f79ff0 ("sched/cache: Handle moving single tasks to/from their preferred LLC")
Closes: https://lore.kernel.org/lkml/20260827135000.735138-1-zhanxusheng@xiaomi.com/
Reported-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Suggested-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Reviewed-by: Kayra Cizmeci <kayracizmeci@gmail.com>
Cc: <stable@kernel.org> # v7.2.x
Link: https://patch.msgid.link/06af61afedac32e6477f57feb4d658f6c411c3af.1790035273.git.tim.c.chen@linux.intel.com
|
|
Two functions in the SVSM vTPM guest implementation do not disable
preemption when fetching the SVSM Calling Area Address (CAA).
The SVSM CAA is a per-CPU structure. When a thread is preempted and migrated
to a different CPU after fetching the per-CPU CAA, the SVSM call will execute
on the new CPU with the original CPU's CAA. Which is wrong.
Move the CAA fetching operation inside svsm_perform_call_protocol() which
disables interrupts around the SVSM call and thus runs preemption-safe.
Fixes: 770de678bc28 ("x86/sev: Add SVSM vTPM probe/send_command functions")
Signed-off-by: Melody Wang <huibo.wang@amd.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Reviewed-by: Stefano Garzarella <sgarzare@redhat.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/a5bc0d4a2c462a0089109e145c21626b244b2ff0.1789345277.git.huibo.wang@amd.com
|
|
Align the documented libata.force options with their implementation.
The force table accepts PIO modes 0 through 6, not mode 7, and ncqati
controls NCQ generally rather than only queued TRIM.
The max_sec_1024 and max_sec_lba48 options only set transfer size
limits. Remove the misleading claim that they can also clear them.
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Link: https://lore.kernel.org/r/20260918124030.1962773-6-cassel@kernel.org
Signed-off-by: Niklas Cassel <cassel@kernel.org>
|
|
A Samsung SSD 870 QVO 8TB connected to an AMD 600 Series chipset SATA
controller is reported to time out on STANDBY IMMEDIATE during system
suspend with med_power_with_dipm enabled. The command completes when
using max_performance instead.
The existing Samsung LPM quirk only matches ATI controllers, leaving
AMD controllers unaffected. Rename it to
ATA_QUIRK_NO_LPM_ON_ATI_AND_AMD and extend the vendor check to AMD for
the same Samsung SSD model patterns. Keep LPM behavior unchanged for
other controller vendors, including Intel.
Leave ATA_QUIRK_NO_NCQ_ON_ATI restricted to ATI, since the reported AMD
issue concerns LPM rather than NCQ.
Link: https://bugzilla.kernel.org/show_bug.cgi?id=221986
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org>> ---
Link: https://lore.kernel.org/r/20260918124030.1962773-5-cassel@kernel.org
Signed-off-by: Niklas Cassel <cassel@kernel.org>
|
|
On the Fujitsu LIFEBOOK U7410 the internal keyboard and touchpad die
shortly after boot if i8042 multiplexing controller is probed and
switched to muxed mode. These multiplexing setup commands disturb the EC
emulated keyboard and atkbd reports "Spurious ACK" and "Unknown key
pressed" warnings, after which both keyboard and touchpad stop
delivering events. Both touchpad and keyboard work in BIOS and GRUB.
Disabling multiplexer using i8042.nomux=1 makes both devices work on
warm/cold boot.
Add a DMI quirk to force nomux mode for this model.
Reported-by: Heiko Brey <cico0815@gmx.de>
Link: https://lore.kernel.org/all/89f36f3f5e4c2d07b9ff72bc0b355875c336e498.camel@gmx.de/
Signed-off-by: Lovekesh Solanki <lovekeshsolanki00@gmail.com>
Link: https://patch.msgid.link/20260920113541.1484808-1-lovekeshsolanki00@gmail.com
Cc: stable@vger.kernel.org
Signed-off-by: Dmitry Torokhov <dmitry.torokhov@gmail.com>
|
|
Lenovo ThinkPad T490 has a Synaptics Touchpad that supports SMBus/RMI4
mode, but is not listed in smbus_pnp_ids table.
Add LEN205b to smbus_pnp_ids[] passlist.
Reported-by: John Rosencutter <rosencutter@gmail.com>
Closes: https://lore.kernel.org/all/CAOGuaRh025eKcBf9P8UrvtXNp68vQ0HVF2EbNMp-kpdaC1HNpA@mail.gmail.com/
Signed-off-by: Lovekesh Solanki <lovekeshsolanki00@gmail.com>
Link: https://patch.msgid.link/20260920200730.1836756-1-lovekeshsolanki00@gmail.com
Cc: stable@vger.kernel.org
Signed-off-by: Dmitry Torokhov <dmitry.torokhov@gmail.com>
|
|
|
|
Since the MHI HELLO exchange was relocated, it is sent only at device
registration. During a suspend-resume cycle, the firmware in WiFi
cards such as WCN7850 indefinitely waits for another HELLO,
triggering:
ath12k_wifi7_pci 0004:01:00.0: timeout while waiting for restart complete
ath12k_wifi7_pci 0004:01:00.0: failed to resume core: -110
Fix this by triggering the handshake from resume_early in the MHI
transport.
Validated on Qualcomm X1E-801800 on Lenovo Slim 7x across 10
suspend-resume cycles.
Fixes: 544d85de4dc2 ("net: qrtr: Send HELLO message on endpoint register")
Signed-off-by: Daniel J Blueman <daniel@quora.org>
Reviewed-by: Manivannan Sadhasivam <manivannan.sadhasivam@oss.qualcomm.com>
Reported-by: Jeff Johnson <jeff.johnson@oss.qualcomm.com>
Link: https://lore.kernel.org/all/6257c447-788d-4362-851e-0d552bcf7c56@oss.qualcomm.com/
Tested-by: Jeff Johnson <jeff.johnson@oss.qualcomm.com>
Reported-by: Vlastimil Babka (SUSE) <vbabka@suse.com>
Link: https://lore.kernel.org/all/ab1491bb-cca5-4145-ac7d-31c966abf7b4@suse.com/
Tested-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Reported-by: Takashi Iwai <tiwai@suse.de>
Link: https://lore.kernel.org/all/87a4plsg4w.wl-tiwai@suse.de/
Tested-by: Takashi Iwai <tiwai@suse.de>
Signed-off-by: Thorsten Leemhuis <linux@leemhuis.info>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/vkoul/dmaengine
Pull dmaengine fixes from Vinod Koul:
- A couple of fixes in core around dma_chan_put() for kref underflow,
use-after-free and waiting for rcu readers for dma devices
- mmp sg length and wrong extended DRCMR base for SpacemiT K3
- hardware buffer descriptor chain fix for xilinx dma
- sun6i fixes for status behaviour and dma position registers
- runtime pm reference leak fix for sprd driver
* tag 'dmaengine-fix-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/vkoul/dmaengine:
dmaengine: mmp_pdma: fix wrong sg length in mmp_pdma_prep_slave_sg()
dmaengine: xilinx_dma: Fix hardware buffer descriptor chain after cyclic DMA
dmaengine: pxa: fix double counting of the hw descriptors
dmaengine: sun6i: fix undefined behaviour in sun6i_dma_tx_status
dmaengine: sun6i: fix non-atomic read of DMA position registers
dmaengine: xilinx_dma: Fix hardware buffer descriptor reuse order
dmaengine: wait for RCU readers before releasing dma_device
dmaengine: fix use-after-free in dma_chan_put() and dma_release_channel()
dmaengine: Fix device kref underflow in dma_chan_put()
dmaengine: add dma_device_get() helper
dmaengine: sprd: Fix runtime PM reference leak in probe
dmaengine: ti: k3-udma-glue: fix NULL dereference in k3_udma_glue_release_rx_chn()
dmaengine: mmp_pdma: fix wrong extended DRCMR base for SpacemiT K3
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/phy/linux-phy
Pull phy fixes from Vinod Koul:
- avoid atomic context delay in renesas driver
- TMDS and PLL rate calculation fixes for mediatek driver
* tag 'phy-fixes-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/phy/linux-phy:
phy: mediatek: phy-mtk-hdmi-mt8195: Fix TMDS clk bit ratio setting
phy: mediatek: phy-mtk-hdmi-mt8195: Fix PLL calc divisor overflow
phy: renesas: rcar-gen3-usb2: Avoid long delay in atomic context
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/vkoul/soundwire
Pull soundwire fixes from Vinod Koul:
- Cadence: ensure work completion before clock stop
- Disable ghost Realtek on Asus GX651AX
* tag 'soundwire-7.3-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vkoul/soundwire:
soundwire: cadence_master: wait and cancel cdns->work before clock stop
soundwire: dmi-quirks: Disable ghost Realtek on Asus GX651AX
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 fixes from Ingo Molnar:
- Reject the loading of a potentially problematic microcode version
on Intel Granite Rapids systems (Chang S. Bae)
- On FRED, reconstruct the proper #GP context for rejected INT
instructions, to fix a signal ABI regression (Matthew Schwartz)
- Add a test for this signal ABI regression the x86
self-test suite (Matthew Schwartz)
- Don't emit the new and not yet properly supported EGPR instructions
(%r16-%r31) on CONFIG_X86_NATIVE_CPU=y builds (Chang S. Bae)
* tag 'x86-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
x86/build/64: Prevent native builds from generating EGPR use
selftests/x86: Check signal state for rejected software interrupts
x86/fred: Reconstruct the #GP context for rejected INT instructions
x86/microcode/intel: Reject problematic loading on Granite Rapids systems
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull timer race fixes from Ingo Molnar:
- Fix timer signal <-> exec() race, to prevent UAF (Thomas Gleixner)
- Clean up POSIX CPU timers right after de_thread(), to prevent UAF
(Hyunwoo Kim)
- Fix POSIX CPU timers race between expiry and timer_settime(),
to prevent UAF (Thomas Gleixner)
* tag 'timers-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
posix-cpu-timers: Prevent freeing a timer which is queued on the expiry list
exec: Cleanup POSIX timers right after de_thread()
signal: Prevent exec() race
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull scheduler fix from Ingo Molnar:
- Avoid false positive migration warning for proxy donors
(Andrea Righi)
* tag 'sched-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
sched/core: Avoid false migration warning for proxy donors
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fix crash when probing CS CALL instructions (Jinke Han)
- Fix NULL pointer crash during module unload (Vinay Belgaumkar)
* tag 'perf-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf: Fix null pointer access in is_include_guest_event()
x86/kprobes: Fix crash when probing CS CALL instructions
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull objtool fix from Ingo Molnar:
- Fix objtool build error on systems where libopcodes is
present, but development headers (binutils-dev) are not
(Ulises Mendez Martinez)
* tag 'objtool-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
objtool: Validate disassembler headers in libopcodes probe
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull futex fix from Ingo Molnar:
- Also allocate a default private futex hash on vfork() as well, to
avoid races with (private) futex waiters (Peter Zijlstra)
* tag 'locking-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
futex: Also allocate private hash on vfork()
|
|
Kijo analyzed another race in the POSIX CPU timer code:
Commit bf635681c906 converted cpu_timer::firing from a tristate value to a
boolean. This lost the distinction between "not owned by the firing list"
and "still owned, but delivery was canceled". The resulting race is:
expiry handler timer_settime() timer_delete()
-------------- --------------- --------------
collect timer onto
private firing list
firing = true
observes firing = true
firing = false
return TIMER_RETRY
wait for handler
observes firing = false
finish deletion
unhash and free timer
resume list traversal
read freed elist.next
-> UAF
The firing bit is clearly the wrong indicator since that commit.
Check whether the timer is queued on the expiry list or not instead. If it
is queued clear the firing bit to prevent signal delivery as before and
return TIMER_RETRY so the caller unlocks the timer which allows the expiry
code to make progress and remove it from the list.
Fixes: bf635681c906 ("posix-cpu-timers: Cleanup the firing logic")
Reported-by: Kijo Park <red993688@gmail.com>
Debugged-by: Kijo Park <red993688@gmail.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Tested-by: Kijo Park <red993688@gmail.com>
Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
Cc: stable@vger.kernel.org
|
|
cid-form ops.enable() now hands the task's cmask to the scheduler and
set_cmask() repeats it right after. Add a cid-form selftest that checks both
against p->cpus_ptr, that they match each other, that the initial
set_cmask() lands before set_weight() and before the task first becomes
runnable, and that set_cmask() never precedes enable(), across class-switch
enables, fork-path enables and live affinity changes.
v2: Mismatch details returned through a caller-local struct instead of
globals, alloc_words validated in the header check, loop bounded by nr_cids
directly (Andrea Righi).
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
The cid-form API has an obvious hole. A task's cid mask is only visible
through ops.set_cmask(), which fires on affinity changes and class switches
but not when a task enters a scheduler through fork, sub-sched enable or
re-home, and there is no p->cpus_ptr equivalent to fall back on. Schedulers
work around it by seeding the mask in ops.init_task() from p->cpus_ptr cid
by cid, which is subtly wrong: on sub-sched enable and re-home, an affinity
change between init_task() and enable() is delivered to the sched the task
is still on, and nothing corrects the new sched's copy afterwards.
Fix it by adding struct scx_enable_args to cid-form ops.enable() carrying
the task's cmask, built in the per-cpu scratch under the rq lock as the task
enters the scheduler, and calling set_cmask() with the same mask right after
enable(), ahead of set_weight(). A scheduler can then track affinity in
set_cmask() alone, and scx_qmap drops its init_task() seed. set_cmask() no
longer fires for a cid-form task before it is enabled, and the class-switch
republish in switching_to_scx() is limited to the cpu form.
This changes the cid-form ops.enable() signature, which is fine as the
cid-form API is still considered unreleased. An args struct rather than a
bare cmask argument leaves room for more initial state without another
signature change, and the cmask travels as a plain arena address because BTF
can't type arena struct members yet.
v2: The cmask travels as a u64 arena address, cmask_arena_addr, instead of a
kernel-typed pointer, with the typing limitation and the planned typed alias
documented (Sashiko review).
v3: The initial set_cmask() is delivered before set_weight() so that the
mask is in place when weight-dependent state is derived (Andrea Righi).
Selftest added.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
|
|
Omar reports that CONFIG_X86_NATIVE_CPU=y allows builds to opportunistically
emit instructions using %r16-%r31 (EGPRs) when the build host supports APX
since the commit:
ea1dcca1de12 ("x86/kbuild/64: Add the CONFIG_X86_NATIVE_CPU option to locally optimize the kernel with '-march=native'")
But the kernel is not yet prepared to use new registers internally. For
example, there is no context-switch support for general in-kernel use.
Explicitly disable EGPR use when building with -march=native.
For C, since GCC 14 and Clang 18, both compilers support suppressing EGPR
use with -mno-apx-features=egpr, whose availability can be detected via
cc-option.
For Rust, pass features=-apxf through the generated JSON to avoid
unstable-feature warnings, see
https://github.com/rust-lang/rust/issues/139284
Note Rust only accepts the option to disable APX instructions entirely or not.
Support for this gating also depends on the Rust/LLVM combination. Rust
1.88 introduced the `apxf` feature option, but versions prior to 1.93 may
emit an `apxf` attribute to the backend that only LLVM 23 or later can
interpret. Restrict native Rust builds accordingly.
Fixes: ea1dcca1de12 ("x86/kbuild/64: Add the CONFIG_X86_NATIVE_CPU option to locally optimize the kernel with '-march=native'")
Reported-by: Omar Avelar <omar.avelar@intel.com>
Signed-off-by: Chang S. Bae <chang.seok.bae@intel.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Reviewed-by: Nathan Chancellor <nathan@kernel.org>
Acked-by: Miguel Ojeda <ojeda@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260916230003.1144622-1-chang.seok.bae@intel.com
|
|
A bridge (a device with a Type 1 header) may not have a secondary bus
allocated (pdev->subordinate), e.g., if there are no available bus numbers
or the bridge secondary/subordinate bus numbers are not writable.
The dynamic OF helpers of_pci_prop_bus_range() and of_pci_prop_intr_map()
dereference pdev->subordinate without checking it. When
CONFIG_PCI_DYNAMIC_OF_NODES is enabled, this can cause a NULL pointer
dereference and early boot hang.
Generate 'bus-range' and 'interrupt-map' properties only when a subordinate
bus exists. Keep the node and its remaining properties for bridges without
one.
The problem was latent since 407d1a51921e ("PCI: Create device tree node
for bridge"), but wasn't reachable until 1f340724419e ("PCI: of: Create
device tree PCI host bridge node"), which appeared in v6.15. Before
1f340724419e, of_pci_make_dev_node() returned early because the parent OF
node was missing.
Fixes: 407d1a51921e ("PCI: Create device tree node for bridge")
Signed-off-by: Angel J <iamanaws@httpd.dev>
[bhelgaas: move pdev->subordinate test to callees, commit log]
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Cc: stable@vger.kernel.org # v6.6+
Link: https://patch.msgid.link/20260918195540.GA1187209@bhelgaas
|
|
pci_do_resource_release_and_resize() releases device BARs that share a
bridge window with the BAR being resized, but when the device sits directly
on a root bus (pdev->bus->self == NULL) it then skips resource assignment
entirely and returns success, leaving the BARs it just released unassigned
(IORESOURCE_UNSET).
Skipping pbus_reassign_bridge_resources() is correct in that case -- there
is no bridge window to adjust -- but the device BARs still have to be
reassigned. Before the BAR release was consolidated into the PCI core, this
case worked for amdgpu because the driver released the BARs itself and then
called pci_assign_unassigned_bus_resources() unconditionally after the
resize, which assigns unassigned device BARs also on a root bus. Commit
db92e3fef53e ("drm/amdgpu: Remove driver side BAR release before resize")
removed that call, so nothing assigns the released BARs anymore.
This breaks amdgpu completely on the SolidRun HoneyComb LX2K (NXP LX2160A,
arm64, ACPI), where ACPI doesn't expose the Root Port so the GPU endpoint
appears directly on a "root bus" of its segment:
amdgpu 0004:01:00.0: BAR 0 [mem 0xa400000000-0xa40fffffff 64bit pref]: releasing
amdgpu 0004:01:00.0: BAR 2 [mem 0xa410000000-0xa4101fffff 64bit pref]: releasing
amdgpu 0004:01:00.0: sw_init of IP block <gmc_v8_0> failed -19
amdgpu 0004:01:00.0: amdgpu_device_ip_init failed
amdgpu 0004:01:00.0: Fatal error during GPU init
No error is logged because the resize path reports success; amdgpu then
finds BAR 0 IORESOURCE_UNSET and bails out with -ENODEV.
When there is no upstream bridge, call pci_bus_assign_resources() on the
root bus to place the BARs released above, using the same alignment-sorted
algorithm as normal enumeration instead of a manual per-BAR loop. This also
walks the rest of the hierarchy under the root bus, as
pci_assign_unassigned_bus_resources() used to for amdgpu before commit
db92e3fef53e ("drm/amdgpu: Remove driver side BAR release before resize")
removed that call -- the core-side fix that commit asked for ("such a
problem should be fixed inside pci_resize_resource() instead").
pci_bus_assign_resources() returns void, so failure is detected by checking
whether the released BARs are still assigned afterward; if not, roll back
as in the bridged case. This is stricter than the bridged path -- it fails
on any unplaced resource, not just required ones -- since a root bus
typically has one shared window, and failing loudly seemed better than
leaving something silently unassigned.
The root bus path also had a locking bug that any fix here necessarily
touches: the old "goto out" jumped to up_read(&pci_bus_sem) without a
matching down_read() (as does the "goto restore" taken when
pci_dev_res_add_to_list() fails in the release loop). Take pci_bus_sem
before the BAR release loop so every path through the function holds it
exactly once.
Fixes: 337b1b566db0 ("PCI: Fix restoring BARs on BAR resize rollback path")
Link: https://bugs.launchpad.net/ubuntu/+source/linux-hwe-7.0/+bug/2159596
Suggested-by: Ilpo Järvinen <ilpo.jarvinen@linux.intel.com>
Assisted-by: Claude:claude-fable-5 checkpatch
Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Liz Fong-Jones <lizf@honeycomb.io>
[bhelgaas: commit log]
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Reviewed-by: Ilpo Järvinen <ilpo.jarvinen@linux.intel.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260918035633.566823-1-lizf@honeycomb.io
|
|
RISC-V SpacemiT clock fixes for v7.3
- Add CPU PLL rate tables
|
|
Add a test of the signal ABI for INT instructions in both 32-bit and
64-bit processes. Check the signal number, trap number, error code,
si_code, si_addr, instruction pointer and RF/TF state against legacy
IDT behavior. Include a 15-byte prefixed INT to check that IP uses the
hardware instruction length. Exercise both INT3 encodings, INT4, UD2
and HLT to cover the unchanged trap and fault paths.
Run each instruction with TF clear and set. Resume at a known NOP after
handling the signal and check that single-stepping traps after the NOP.
Also drive INT 0x2d under ptrace, which resumes through the fault frame
rather than sigreturn and so exposes a stale FRED software event flag.
Start from an INT3 stop, whose FRED frame has no software event flag,
instead of the syscall frame of raise(SIGSTOP). Single-step into the INT
and check that the fault reports its address. Then suppress SIGSEGV and
resume at the NOP, once with PTRACE_SINGLESTEP and once with PTRACE_CONT
and TF set. Section 6.2.3 of the Intel FRED specification [1] specifies
the immediate single-step trap caused by returning with both that flag
and TF set. Check that each trap occurs after the NOP, rather than at
its address.
Report whether the CPU supports FRED, since a pass looks the same on
either entry path. INT 0x80 with IA32 emulation disabled and a 64-bit
tracer of a 32-bit tracee are not covered.
Both variants pass all 29 checks on a non-FRED AMD host and on Panther
Lake with FRED enabled and the fix applied. With the same binaries on
unpatched Panther Lake, 16 signal-context checks fail and the first
ptrace check reports the IP after the INT. The two dependent ptrace
resume checks are not reached.
[1] Intel Flexible Return and Event Delivery (FRED) Specification,
revision 9.0 (346446-009US), section 6.2.3.
Signed-off-by: Matthew Schwartz <matthew.schwartz@linux.dev>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://cdrdv2.intel.com/v1/dl/getContent/678938 # [1]
Link: https://patch.msgid.link/20260917230907.2080792-3-matthew.schwartz@linux.dev
|
|
FRED event delivery does not use the IDT, so the gate DPL check that
rejects a user INT n falls to software (Intel FRED specification [1],
section 8.3). fred_intx() rejects the same vectors as IDT delivery, but
reports a zero error code and the IP after the INT. This breaks the
signal ABI. Wine uses the error code to recognize INT 0x2d, so the
changed context turns a handled breakpoint into an access violation in
Elden Ring.
Rewind IP using the instruction length in the augmented SS and
synthesize the IDT selector error code, (vector << 3) | 2. Set RF in the
saved flags, as the CPU does for a #GP fault. Section 5.2.1 defines the
saved vector, instruction length and RF state. The supplied length
handles prefixes without reading user memory. Limit the changes to
already-rejected software interrupts, preserving the accepted INT3, INT4
and enabled INT80 paths and hardware exceptions. With IA32 emulation
disabled, INT 0x80 now reports the same #GP as the DPL 0 gate IDT
installs there. The rewound IP also stops fixup_iopl_exception() from
inspecting the byte after the INT.
Also clear the software event flag. Section 6.2.3 specifies that ERETU
with this flag and TF set traps before executing any user instruction. A
tracer that suppresses SIGSEGV and resumes with TF set expects the next
instruction to run first, as after IRET. The sigreturn path clears the
same flag for this reason in prevent_single_step_upon_eretu().
[1] Intel Flexible Return and Event Delivery (FRED) Specification,
revision 9.0 (346446-009US), sections 5.2.1, 6.2.3 and 8.3.
Fixes: 14619d912b65 ("x86/fred: FRED entry/exit and dispatch code")
Closes: https://gitlab.freedesktop.org/mesa/mesa/-/work_items/15745
Closes: https://gitlab.freedesktop.org/mesa/mesa/-/work_items/16132
Reported-by: Paul Gofman <pgofman@codeweavers.com>
Signed-off-by: Matthew Schwartz <matthew.schwartz@linux.dev>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Reviewed-by: H. Peter Anvin <hpa@zytor.com>
Link: https://cdrdv2.intel.com/v1/dl/getContent/678938 # [1]
Link: https://patch.msgid.link/20260917230907.2080792-2-matthew.schwartz@linux.dev
|
|
Proxy execution can move a blocked donor's scheduling context to the
lock owner's CPU even when the donor is migration-disabled. The donor
does not execute there, and its original execution CPU remains recorded
in wake_cpu.
set_task_cpu() warns unconditionally for migration-disabled tasks, so a
subsequent proxy migration or the wakeup path returning the donor home
triggers a false positive: moving a blocked scheduling context does not
violate the migration-disabled execution context.
For example, creating a mutex owner on CPU1 and a migration-disabled
waiter on CPU0 can trigger the following warning:
proxy_migrate_repro: donor blocking on CPU0 with migration disabled
proxy_migrate_repro: donor moved from CPU0 to CPU1
WARNING: kernel/sched/core.c:3389 at set_task_cpu+0x1d3/0x280
...
Call Trace:
try_to_wake_up+0x43f/0x780
__mutex_unlock_slowpath+0x330/0x540
owner_fn+0x9f/0xc0 [proxy_migrate_repro]
...
proxy_migrate_repro: donor woke on CPU0, task_cpu=0
proxy_migrate_repro: completed
Exclude blocked proxy donors from the warning. The proxy wakeup path
restores an executable placement before clearing the blocked state.
Fixes: b049b81bdff6 ("sched: Handle blocked-waiter migration (and return migration)")
Signed-off-by: Andrea Righi <arighi@nvidia.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Acked-by: John Stultz <jstultz@google.com>
Link: https://patch.msgid.link/20260915184101.2621252-1-arighi@nvidia.com
|
|
A typical module unload occurring event when there is an active perf
connection leads to freeing of the pmu pointer. The call log is something
like:
..
__pmu_detach_event
pmu_detach_event
pmu_detach_events
perf_pmu_unregister
..
__pmu_detach_event() sets event->pmu to null. When the perf connection
finally is closed, the following stack trace is observed:
Oops: general protection fault, kernel NULL pointer dereference
...
RIP: 0010:_free_event+0x3e/0x370
...
Call Trace:
...
perf_event_release_kernel+0x260/0x2d0
perf_release+0x12/0x20
A call to mediated_pmu_unaccount_event() inside _free_event() is the root
cause of this crash. Adding a check inside is_include_guest_event() ensures
we don't accidentally access a null pmu ptr. In addition to this, we will
now call mediated_pmu_unaccount_event() before clearing the pmu ptr so that
nr_include_guest_events counts are maintained correctly.
Fixes: eff95e170275 ("perf: Add APIs to create/release mediated guest vPMUs")
Assisted-by: Claude:Claude-Sonnet-5
Signed-off-by: Vinay Belgaumkar <vinay.belgaumkar@intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
Link: https://patch.msgid.link/20260904181625.1394082-1-vinay.belgaumkar@intel.com
|
|
fprobe_fgraph_entry() reserves shadow stack space for every fprobe with
an exit handler, but only fills it for those whose entry handler returns
0. fgraph_reserve_data() does not clear the area, so fprobe_return()
parses the unused tail as headers left over from an earlier call, and an
exit handler can run twice or despite its entry handler asking to skip
it.
Write a zero word after the last entry to terminate the walk. A zeroed
slot does not decode to a NULL fprobe on the arches that encode the
header into one unsigned long, since arch_decode_fprobe_header_fp() ORs
in FPROBE_HEADER_MSB_PATTERN, so make read_fprobe_header() return NULL
for a zeroed slot.
Link: https://lore.kernel.org/all/20260917212407.384468-1-devnexen@gmail.com/
Fixes: e0a384434ae1 ("tracing: fprobe: do not zero out unused fgraph_data")
Cc: stable@vger.kernel.org
Suggested-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Signed-off-by: David Carlier <devnexen@gmail.com>
Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
|
|
Microcode updates can usually jump revisions. However, there is an erratum on
Granite Rapids systems. If they "jump over" revision 0x1000405, they result in
an #MC. Avoid it.
Signed-off-by: Chang S. Bae <chang.seok.bae@intel.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Reviewed-by: Dave Hansen <dave.hansen@linux.intel.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260916225939.1144524-1-chang.seok.bae@intel.com
|
|
Add a scheduler whose ops.dequeue() iterates the source user DSQ with
bpf_iter_scx_dsq. The iteration takes the DSQ's raw spinlock; on a
kernel that runs ops.dequeue() while the consume path still holds that
lock, the first task consumed self-deadlocks the CPU with IRQs
disabled. The watchdog cannot recover from that state, so on an
unfixed kernel this test wedges the system instead of failing cleanly.
On a fixed kernel the scheduler runs clean and the test passes.
Signed-off-by: fangqiurong <fangqiurong@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
ops.dequeue() is invoked with the source user DSQ's lock still held on
the consume and move paths (scx_consume_dispatch_q(),
move_task_between_dsqs()). A BPF scheduler which locks the source user
DSQ from ops.dequeue() - e.g. by iterating it with bpf_iter_scx_dsq -
self-deadlocks.
ops.dequeue() can only call the "any" kfuncs and none of them can lock a
builtin DSQ, so the global and bypass paths can't deadlock; however,
all DSQ locks share one lockdep class, so iterating any user DSQ from
ops.dequeue() on those paths trips the recursion check.
Move the invocation after the DSQ unlock on all three paths.
SCX_TASK_IN_CUSTODY is cleared under the lock serializing the transfer
so that the callback is invoked exactly once.
Fixes: ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics")
Cc: stable@vger.kernel.org # v7.1+
Acked-by: Andrea Righi <arighi@nvidia.com>
Signed-off-by: fangqiurong <fangqiurong@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
schedule_deferred_locked() skips scheduling a deferred action while
SCX_RQ_IN_WAKEUP is set and relies on the task_woken_scx() call that follows
a wakeup enqueue to run it. enqueue_task_scx() sets the flag from the merged
enqueue flags, which include the flags stashed for a remote activation.
move_remote_task_to_local_dsq() thus sets SCX_RQ_IN_WAKEUP on the
destination rq when the moved task was woken up, although no
task_woken_scx() follows that activation.
An IMMED insert into a busy destination requests a local reenqueue during
that enqueue. The request gets linked but not scheduled and stays pending
until an unrelated wakeup or preemption on that CPU runs the deferred
actions. The IMMED task sits behind the running task in the meantime. If
nothing runs them before the scheduler is disabled, the request outlives the
scheduler and points into its freed per-cpu area, which the next scheduler
dereferences from run_deferred().
Test the core enqueue flags for the wakeup bit. Only the core's wakeup path
is followed by task_woken_scx().
Fixes: 57ccf5ccdc56 ("sched_ext: Fix enqueue_task_scx() truncation of upper enqueue flags")
Cc: stable@vger.kernel.org # v7.1+
Reported-by: Andrea Righi <arighi@nvidia.com>
Link: https://lore.kernel.org/all/20260916145807.3250167-1-arighi@nvidia.com/
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
|
|
When using eBPF to probe CS CALL instructions within a function,
a crash can be triggered.
The eBPF tool probes offset 257 of the __hrtimer_run_queues()
function:
<__hrtimer_run_queues+249>: nopl 0x0(%rax,%rax,1)
<__hrtimer_run_queues+254>: mov %r14,%rdi
<__hrtimer_run_queues+257>: cs call <__x86_indirect_thunk_r12>
<__hrtimer_run_queues+263>: mov %eax,%r12d
<__hrtimer_run_queues+266>: xchg %ax,%ax
<__hrtimer_run_queues+268>: mov %r13,%rdi
Which triggers this crash:
BUG: unable to handle page fault for address: 00000000000f41c9
#PF: supervisor write access in kernel mode
#PF: error_code(0x0002) - not-present page
PGD 0 P4D 0
Oops: 0002 [#1] SMP NOPTI
CPU: 1 PID: 0 Comm: swapper/1 Kdump: loaded Tainted: P
RIP: 0010:__hrtimer_run_queues+0x106/0x230
Note that __hrtimer_run_queues+0x106 is __hrtimer_run_queues+262, which is
at the 6th byte of the above CS CALL instruction. Since the CS CALL
instruction occupies 6 bytes, the exception occurred in the middle of that
call instruction.
The root cause is that when using eBPF tools to probe in the middle of a
function, a kprobe with INT3 is used as the underlying implementation.
During single-step emulation of the original CALL instruction,
int3_emulate_call() assumes that the probed CALL instruction is 5 bytes
long. However, the actual CS-prefixed CALL instruction occupies 6 bytes,
so it constructs an incorrect exception return address. When the CPU
returns from the kprobe handler, the next instruction to be executed is at
the address of the last byte of that CS CALL instruction. Coincidentally,
starting from that address, the CPU fetches and decodes a completely
different instruction, which ultimately triggers a kernel crash.
Fix the issue by using the actual instruction length obtained from
the instruction decoder when constructing the exception return
address, rather than relying on the hardcoded CALL_INSN_SIZE macro.
[ mingo: Refined the changelog ]
Fixes: 6256e668b7af ("x86/kprobes: Use int3 instead of debug trap for single-step")
Suggested-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Signed-off-by: Jinke Han <jinkehan@didiglobal.com>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Reviewed-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Acked-by: Yafang Shao <laoar.shao@gmail.com>
Acked-by: Borislav Petkov <bp@alien8.de>
Cc: Peter Zijlstra <peterz@infradead.org>
Link: https://patch.msgid.link/20260908073742.GA10517@didi-ThinkCentre-M920t-N000
|
|
O_TMPFILE, like O_CREAT, needs the third argument. Without it glibc
refuses the call at compile time as soon as fortification is on:
In function 'open',
inlined from 'get_temp_fd' at test_memcontrol.c:33:9:
/usr/include/bits/fcntl2.h:52:11: error: call to '__open_missing_mode'
declared with attribute error: open with O_CREAT or O_TMPFILE in
second argument needs 3 arguments
The fortify checks take effect only once the compiler optimises, and
cgroup/Makefile builds with "-Wall -pthread" alone, so this goes
unnoticed in a plain build. Building the tests with the flags
distributions commonly use, -O2 -D_FORTIFY_SOURCE=3, loses
test_memcontrol entirely.
Fixes: 84092dbcf901 ("selftests: cgroup: add memory controller self-tests")
Signed-off-by: Eva Kurchatova <eva.kurchatova@virtuozzo.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
|
|
rust_binderfs is missing the equivalent of commit f37b55ded8ed ("binder:
add transaction_report feature entry"), which adds "transaction_report"
to the binderfs feature list. This helps userspace determine if the
BINDER_CMD_REPORT from the generic netlink API is supported.
Cc: stable <stable@kernel.org>
Fixes: f14e0c8183bc ("rust_binder: report netlink transactions")
Signed-off-by: Carlos Llamas <cmllamas@google.com>
Link: https://patch.msgid.link/20260916120833.593407-1-cmllamas@google.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
When a thread exits via BINDER_THREAD_EXIT, its pending work items are
cancelled. If a thread exits while holding a pending node refcount
increment (e.g. pushed as deferred work to that thread), the refcount
increment was previously dropped because Node::cancel() and
NodeWrapper::cancel() were no-ops.
Dropping the refcount update leaves the node's delivery state and count
state desynchronized, and userspace will not receive the notification,
which can cause the node to never be freed from the process's nodes tree
when all external references are dropped.
Fix this by implementing DeliverToRead::cancel() for Node and NodeWrapper
to move the pending refcount update to the process's work queue on thread
exit so another thread can deliver it to userspace.
Cc: stable <stable@kernel.org>
Fixes: eafedbc7c050 ("rust_binder: add Rust Binder driver")
Signed-off-by: Alice Ryhl <aliceryhl@google.com>
Link: https://patch.msgid.link/20260903-binder-thread-exit-node-v1-1-be09ff14f6a4@google.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
If there are deferred work items on the thread todo list, then they are
not cleaned up in the Thread::release() method. Thus, update the code to
clean up the work items even if they are deferred.
This can happen if the thread dies while it has an active outgoing
transaction.
Cc: stable <stable@kernel.org>
Fixes: eafedbc7c050 ("rust_binder: add Rust Binder driver")
Signed-off-by: Alice Ryhl <aliceryhl@google.com>
Link: https://patch.msgid.link/20260903-binder-exit-get-work-v1-1-2d6129a238df@google.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
binderfs_binder_device_create() publishes the new dentry with
d_make_persistent() and then calls simple_done_creating(), which drops
the parent directory lock and the creator's dentry reference. It then
calls binder_add_device() to register the device in the global
binder_devices list.
After simple_done_creating() releases the parent directory lock, a
concurrent unlinkat() can remove the new device entry. Dropping the
creator's dentry reference can then trigger binderfs_evict_inode(),
freeing the device. binder_add_device() later accesses the freed
object, causing UAF write.
Found by a modified Syzkaller:
BUG: KASAN: slab-use-after-free in hlist_add_head include/linux/list.h:1073 [inline]
BUG: KASAN: slab-use-after-free in binder_add_device+0xa9/0xc0 drivers/android/binder.c:7068
Write of size 8 at addr ffff88805b340c00 by task syz.1.532/11389
CPU: 0 UID: 0 PID: 11389 Comm: syz.1.532 Not tainted 7.2.0 #4 PREEMPT(full)
Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
Call Trace:
<TASK>
__dump_stack lib/dump_stack.c:94 [inline]
dump_stack_lvl+0x10e/0x1f0 lib/dump_stack.c:120
print_address_description mm/kasan/report.c:378 [inline]
print_report+0xf7/0x600 mm/kasan/report.c:482
kasan_report+0xe4/0x120 mm/kasan/report.c:595
hlist_add_head include/linux/list.h:1073 [inline]
binder_add_device+0xa9/0xc0 drivers/android/binder.c:7068
binderfs_binder_device_create.isra.0+0x724/0x990 drivers/android/binderfs.c:196
binder_ctl_ioctl+0x186/0x1b0 drivers/android/binderfs.c:241
vfs_ioctl fs/ioctl.c:51 [inline]
__do_sys_ioctl fs/ioctl.c:597 [inline]
__se_sys_ioctl fs/ioctl.c:583 [inline]
__x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7ff9027a833d
Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 b0 ff ff ff f7 d8 64 89 01 48
RSP: 002b:00007ff903674018 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
RAX: ffffffffffffffda RBX: 00007ff902a35fa0 RCX: 00007ff9027a833d
RDX: 0000200000000500 RSI: 00000000c1086201 RDI: 0000000000000004
RBP: 00007ff902850733 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
R13: 00007ff902a36038 R14: 00007ff902a35fa0 R15: 00007ffd7166bfa0
</TASK>
Allocated by task 11389:
kasan_save_stack+0x33/0x60 mm/kasan/common.c:57
kasan_save_track+0x14/0x30 mm/kasan/common.c:78
poison_kmalloc_redzone mm/kasan/common.c:398 [inline]
__kasan_kmalloc+0xaa/0xb0 mm/kasan/common.c:415
kasan_kmalloc include/linux/kasan.h:263 [inline]
__kmalloc_cache_noprof+0x2e4/0x6f0 mm/slub.c:5489
_kmalloc_noprof include/linux/slab.h:988 [inline]
_kzalloc_noprof include/linux/slab.h:1309 [inline]
binderfs_binder_device_create.isra.0+0x17a/0x990 drivers/android/binderfs.c:148
binder_ctl_ioctl+0x186/0x1b0 drivers/android/binderfs.c:241
vfs_ioctl fs/ioctl.c:51 [inline]
__do_sys_ioctl fs/ioctl.c:597 [inline]
__se_sys_ioctl fs/ioctl.c:583 [inline]
__x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
Freed by task 11389:
kasan_save_stack+0x33/0x60 mm/kasan/common.c:57
kasan_save_track+0x14/0x30 mm/kasan/common.c:78
kasan_save_free_info+0x3b/0x60 mm/kasan/generic.c:584
poison_slab_object mm/kasan/common.c:253 [inline]
__kasan_slab_free+0x5f/0x80 mm/kasan/common.c:285
kasan_slab_free include/linux/kasan.h:235 [inline]
slab_free_hook mm/slub.c:2677 [inline]
slab_free mm/slub.c:6377 [inline]
kfree+0x2fc/0x6e0 mm/slub.c:6692
binderfs_evict_inode+0x1e8/0x260 drivers/android/binderfs.c:268
evict+0x3c2/0xad0 fs/inode.c:825
iput_final fs/inode.c:2019 [inline]
iput fs/inode.c:2068 [inline]
iput+0x79a/0xd30 fs/inode.c:2031
dentry_unlink_inode+0x27f/0x460 fs/dcache.c:479
dentry_kill+0x25d/0xc20 fs/dcache.c:826
finish_dput fs/dcache.c:1001 [inline]
dput.part.0+0xce/0x230 fs/dcache.c:1042
dput+0x1f/0x30 fs/dcache.c:1037
end_dirop+0x7d/0xa0 fs/namei.c:2956
binderfs_binder_device_create.isra.0+0x71c/0x990 drivers/android/binderfs.c:194
binder_ctl_ioctl+0x186/0x1b0 drivers/android/binderfs.c:241
vfs_ioctl fs/ioctl.c:51 [inline]
__do_sys_ioctl fs/ioctl.c:597 [inline]
__se_sys_ioctl fs/ioctl.c:583 [inline]
__x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
The buggy address belongs to the object at ffff88805b340c00
which belongs to the cache kmalloc-512 of size 512
The buggy address is located 0 bytes inside of
freed 512-byte region [ffff88805b340c00, ffff88805b340e00)
Fix by calling binder_add_device() before d_make_persistent(),
while the parent directory lock is still held and the dentry
cannot be discarded.
Cc: stable <stable@kernel.org>
Fixes: b89aa544821d ("convert binderfs")
Signed-off-by: Peiyang He <peiyang_he@smail.nju.edu.cn>
Assisted-by: Codex:gpt-5.5
Acked-by: Carlos Llamas <cmllamas@google.com>
Link: https://patch.msgid.link/F6EF5FB778E87C98+20260913085645.1639558-1-peiyang_he@smail.nju.edu.cn
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
When a TF_UPDATE_TXN transaction supersedes a pending async transaction,
binder_release_entire_buffer() is called with is_failure=false. Since the
superseded transaction was never delivered, binder_apply_fd_fixups() was
never called and no fds were installed in the target process.
With is_failure=false, the BINDER_TYPE_FDA cleanup handler interprets
stale buffer contents as installed fd numbers and passes them to
binder_deferred_fd_close(), closing unrelated file descriptors.
Pass is_failure=true since the transaction was never delivered to the
target, matching the semantics of all other undelivered-transaction
cleanup paths.
Fixes: 9864bb480133 ("Binder: add TF_UPDATE_TXN to replace outdated txn")
Cc: stable <stable@kernel.org>
Signed-off-by: Tomer Pomeranc <tomerpo@gmail.com>
Acked-by: Carlos Llamas <cmllamas@google.com>
Link: https://patch.msgid.link/20260812195316.259136-3-tomerpo@gmail.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
When a TF_UPDATE_TXN transaction supersedes a pending async transaction
in a frozen process, the outdated transaction is freed with kfree()
directly. This skips binder_free_txn_fixups(), leaking all
binder_txn_fd_fixup entries and their fget()'d struct file references.
The leaked file refcounts never reach zero, so the struct file objects
are permanently pinned in memory. They survive process exit and
accumulate across invocations until file-max exhaustion.
Every other transaction cleanup path (binder_free_transaction(),
binder_transaction() error paths, binder_release_work()) correctly
calls binder_free_txn_fixups(). Add the missing call before kfree()
in the t_outdated cleanup block.
Fixes: 9864bb480133 ("Binder: add TF_UPDATE_TXN to replace outdated txn")
Cc: stable <stable@kernel.org>
Signed-off-by: Tomer Pomeranc <tomerpo@gmail.com>
Acked-by: Carlos Llamas <cmllamas@google.com>
Reviewed-by: Alice Ryhl <aliceryhl@google.com>
Link: https://patch.msgid.link/20260812195316.259136-2-tomerpo@gmail.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
A per-thread CPU timer holds a reference to the PID of the thread it is
attached to and, while it is armed, its node is queued in that thread's
posix_cputimers. The task is looked up by that PID.
When a non-leader thread exec()s, de_thread() changes which task owns
that PID. pid_task(timer->it.cpu.pid, PIDTYPE_PID) then returns NULL,
but the node is still queued on tsk, which is alive. timer_lock_sighand()
takes a failed lookup to mean that the node is already dequeued, so it
has nothing to undo.
begin_new_exec() calls posix_cpu_timers_exit(me) right after
exec_task_namespaces() and that removes the leftover node, so the state
normally stays invisible. But bprm->point_of_no_return is set before
de_thread(), so if unshare_files(), set_mm_exe_file(), exec_mmap() or
exec_task_namespaces() fails, the task dies before it gets there.
exit_itimers() then frees the k_itimer while its node is still queued,
and reaping tsk later erases that freed node from the rbtree.
In short:
the non-leader thread B the parent
timer_create(CLOCK_THREAD_CPUTIME_ID)
timer_settime()
arm_timer() // the node is queued on B
execve()
de_thread(B)
exchange_tids(B, leader) // B's PID now belongs to the leader
release_task(leader)
__exit_signal(leader)
posix_cpu_timers_exit(leader) // cleans leader's queue, not B's
__unhash_process(leader) // that PID has no task anymore
exec_mmap()
mmap_read_lock_killable(old_mm)
kill(B, SIGKILL)
// -EINTR
get_signal()
do_exit()
exit_itimers()
posix_timer_delete()
posix_cpu_timer_del()
posix_timer_unhash_and_free() // freed while still queued
wait4()
release_task(B)
posix_cpu_timers_exit(B)
cleanup_timerqueue()
timerqueue_del() // use-after-free
Move the POSIX timer cleanup right after de_thread() before any of the
later failure conditions brings the task into do_exit().
[ tglx: Move the cleanup right after de_thread() ]
Fixes: 55e8c8eb2c7b ("posix-cpu-timers: Store a reference to a pid not a task")
Signed-off-by: Hyunwoo Kim <imv4bel@gmail.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Tested-by: Kijo Park <red993688@gmail.com>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/ao7Q8miiuLAPVnWv@v4bel
Link: https://patch.msgid.link/20260911090541.627712075@kernel.org
|
|
Hyunwoo debugged the following KASAN UAF splat:
BUG: KASAN: slab-use-after-free in __send_signal_locked+0xb27/0xba0
Write of size 8 at addr ffff888007ed80c8 by task poc/79
...
Call Trace:
__send_signal_locked+0xb27/0xba0
do_send_sig_info+0xa7/0x160
do_send_specific+0x76/0xa0
__x64_sys_tgkill+0x193/0x270
...
Allocated by task 80:
do_timer_create+0x1a4/0x1030
__x64_sys_timer_create+0x145/0x190
...
Freed by task 12:
kmem_cache_free_bulk+0x1f8/0x4a0
kvfree_rcu_bulk+0x14f/0x1c0
kfree_rcu_work+0x128/0x1a0
...
Last potentially related work creation:
kvfree_call_rcu+0x39/0x390
__flush_itimer_signals+0x211/0x320
flush_itimer_signals+0x47/0x90
begin_new_exec+0xa6b/0x28c0
It turned out that this happens with a non-leader exec() as Hyunwoo
explained:
de_thread() calls exchange_tids() before release_task(leader), so the
struct pid held by a SIGEV_THREAD_ID timer created against the leader's tid
now points to the thread which called execve(). pid_task() returns that
thread and lock_task_sighand() on it succeeds.
If the timer signal is blocked, its sigqueue stays queued on the leader's
task::pending. The next expiry of that timer can then run while
release_task() flushes the queue.
posixtimer_send_sigqueue() checks whether the sigqueue is already queued
with a plain list_empty(), which only reads list_head::next.
list_del_init() is not atomic and INIT_LIST_HEAD() stores list_head::next
before list_head::prev, so the check can pass in between. list_add_tail()
queues the entry on the task::pending of the live thread, and the
list_head::prev store from the flush then overwrites the list_head::prev
link that list_add_tail() has just set.
__flush_itimer_signals() does not undo that either. With list_head::prev
pointing at the entry itself, its list_del_init() only stores the same
values again, so the entry is not removed from the list. It is still there
after the last reference is dropped and the timer is freed by RCU, and the
list_add_tail() of a later tgkill() follows that list_head::prev into the
freed timer.
This problem surfaced with the recent commit which moved the sigqueue flush
out of the sighand lock held region.
Hyonwoo proposed to fix this by using list_del_init_careful(), but that
just papers over the problem. After some disucssions and various attempts
to solve it, Eric pointed out that there is no reason to flush
task::pending late in release_task() and it should be done in
exit_signals() already.
As nothing can collect and deliver signals which are queued in a dying
task's pending queue, there is no reason to delay it further.
But it has to be ensured that no signals can be queued into it after that
point. exit_signals() sets PF_EXITING in task::flags, which can be used as
an indicator for this.
Cure it by:
- Preventing signal queueing for task private signals (PIDTYPE_PID) when
the task has PF_EXITING set in __send_signal_locked() and in
posixtimer_send_sigqueue().
- Protecting the unlocked setting of PF_EXITING in exit_signals() for the
task group empty and the group exit case with sighand lock
- Flushing task::pending signals right there.
Optimize that by moving the whole pending list to an on-stack list head
under sighand lock and free the signals without the lock held.
There has been quite some discussion about the lockless flush and the
non-leader exec case on weakly ordered systems. The problem is that a third
party which tries to send a posix timer signal relies on the PID lookup to
find the target task and that lookup might result in the new leader when
the signal was originaly directed to the old leader. In case that the
signal was queued on the old leader then the lockless flush raised a
concern over the following situation:
old_leader new_leader third party
A: flush_list() // list_del_init() stores to sigqueue
LOCK (tasklist)
old_leader->exit_state = EXIT_ZOMBIE;
B: UNLOCK (tasklist)
C: LOCK (tasklist)
if (old_leader->exit_state)
transfer_tids()
D: store PID
posix_timer_send_sigqueue()
// Observes #D so t = new_leader
E: t = get_target()
F: LOCK (sighand)
G: if (list_empty(sigqueue))
list_add(sigqueue)
The concern was that the third party might observe #D but not observe #A
and therefore would proceed to #G while the list_del() stores (#A) in
flush_list() are not visible yet, which could result in list corruption.
That would be possible if looking at it solely from a RELEASE+ACQUIRE
ordering point of view, but B-C is a UNLOCK+LOCK hand-over, which is not
the same as RELEASE+ACQUIRE:
RELEASE+ACQUIRE: RCpc, only the CPUs involved agree on the ordering
UNLOCK+LOCK: RCtso, the hand-over is store-ordering
As B-C is UNLOCK+LOCK, which is RCtso and that does impose store order,
A stores must happen before the D store.
Combine with E-F, which has a data dependency from the LOAD to the LOCK and
thereby constraints later LOADs, those sigqueue loads in G that come after
F must in fact observe the A stores.
Fixes: fb3bbcfe344e ("exit: change the release_task() paths to call flush_sigqueue() lockless")
Reported-by: Hyunwoo Kim <imv4bel@gmail.com>
Debugged-by: Hyunwoo Kim <imv4bel@gmail.com>
Suggested-by: "Eric W. Biederman" <ebiederm@xmission.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Tested-by: Kijo Park <red993688@gmail.com>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260911090541.572536604@kernel.org
Closes: https://patch.msgid.link/aok1rdkBgZsynHZB@v4bel
|
|
make htmldocs fails with:
Documentation/driver-api/libata:604: ./drivers/ata/libata-scsi.c:2301:
ERROR: Unexpected indentation. [docutils]
The Return: section of ata_dsm_trim_pages() ends its first sentence with
a colon and continues with an indented bullet list. reStructuredText
requires a blank line before an indented block, so docutils chokes on
the list.
Simply adding the missing blank line does not work either: Return: is a
kernel-doc "special section", which is terminated by the first blank
line, so the bullet list would end up in the Description section,
detached from the sentence introducing it.
Spell the two bounds out as prose instead, so that the Return: section
stays self-contained. Documentation-only change.
Fixes: e64e6b5dc867 ("ata: libata-scsi: scale DSM TRIM payload by MAX PAGES PER DSM COMMAND")
Reported-by: Thomas Huth <thuth@redhat.com>
Closes: https://lore.kernel.org/linux-ide/b0a0b8e8-cc4f-4c7b-8bc5-fee0712d405a@redhat.com/
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Reviewed-by: Thomas Huth <thuth@redhat.com>
Link: https://lore.kernel.org/r/20260916130031.29990-2-cassel@kernel.org
Signed-off-by: Niklas Cassel <cassel@kernel.org>
|
|
This reverts commit 5a8b2cc36e796d595c9d97eab23c2be22805cc1c.
Since the introduction of this QoS change, X1 machines based on the
Purwa SoC (such as the Lenovo Mini X) take a hard reset at boot time.
Other machines based on Hamoa (such as the X1E001DE Snapdragon Devkit)
end-up with a similar hard reset while under load, most likely due to
the PCIe ports being starved of traffic.
Revert the whole thing until someone figures out what magic parameters
allow QoS to work in a sensible manner, as a usable machine is somehow
preferable to one that crashes efficiently.
Link: https://lore.kernel.org/all/apgeBNSiP01YfNcs@google.com
Link: https://lore.kernel.org/all/86ld9d4h4n.wl-maz@kernel.org
Signed-off-by: Marc Zyngier <maz@kernel.org>
Cc: Georgi Djakov <djakov@kernel.org>
Cc: Raviteja Laggyshetty <raviteja.laggyshetty@oss.qualcomm.com>
Cc: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com>
Cc: Abel Vesa <abelvesa@kernel.org>
Cc: Bjorn Andersson <andersson@kernel.org>
Cc: Mostafa Saleh <smostafa@google.com>
Cc: Thorsten Leemhuis <regressions@leemhuis.info>
Reviewed-by: Abel Vesa <abel.vesa@oss.qualcomm.com>
Reviewed-by: Mostafa Saleh <smostafa@google.com>
Tested-by: Mostafa Saleh <smostafa@google.com>
Closes: https://krzk.eu/#/builders/102/builds/70/steps/23/logs/warnings__3_
Link: https://patch.msgid.link/20260915123703.3228329-1-maz@kernel.org
Signed-off-by: Georgi Djakov <djakov@kernel.org>
|
|
for AMD USB4 routers"
This commit has caused a deadlock at shutdown. A proper fix with
another approach will be coming later. Revert commit
f1de1fc5f632cdeae1f5c2984572ab710d4dfcaa for now.
Reported-by: juan.martinez@amd.com
Closes: https://lore.kernel.org/linux-usb/20260825214237.4179813-1-juan.martinez@amd.com/
Cc: Sanath S <Sanath.S@amd.com>
Cc: Basavaraj Natikar <Basavaraj.Natikar@amd.com>
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Mika Westerberg <mika.westerberg@linux.intel.com>
|
|
ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics") moved the final
ops_state store in scx_dispatch_enqueue() after the DSQ unlock so that the
custody update and ops.dequeue() precede it. A task can thus be found on a
DSQ while still SCX_OPSS_DISPATCHING.
The dequeue and core-sched pick paths wait for the state to clear in
ops_dequeue() but the reenqueue paths don't. A reenqueue in that window runs
ops.enqueue() and sets SCX_OPSS_QUEUED before the dispatch has completed.
The dispatcher's final store then overwrites it with SCX_OPSS_NONE and
finish_dispatch() drops every later dispatch of the task.
Wait for SCX_OPSS_DISPATCHING to clear before dequeueing a task for
reenqueue, the same way ops_dequeue() does.
Fixes: ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics")
Cc: stable@vger.kernel.org # v7.1+
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
|