summaryrefslogtreecommitdiffstats
AgeCommit message (Collapse)Author
2026-09-22sched/cache: Introduce task_struct->sched_cache_grp to fix UAFTim Chen
Add a sched_cache_grp pointer to task_struct so that scheduler code can access the cache group directly via the task, without going through mm->sched_cache_grp. This decouples the scheduler's hot-path accesses from the mm_struct. Each task holds its own refcount on the sched_cache_group, separate from the reference held by its mm_struct. The reference is acquired in copy_mm() (fork) and exec_mmap() (exec), and released in exit_mm(). This fixes the use-after-free when account_mm_sched() reaches the group through a task whose mm is being switched, as reported by Hyunwoo: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Convert all scheduler code in fair.c and exit.c to use p->sched_cache_grp instead of p->mm->sched_cache_grp. Keep the fork/exec/exit reference management out of the generic mm paths: add sched_cache_fork(), sched_cache_fork_cleanup(), sched_cache_exec_mmap() and sched_cache_exit_mm() in kernel/sched/cache_sched.c (with empty stubs for !CONFIG_SCHED_CACHE), so fs/exec.c, kernel/fork.c and kernel/exit.c each call one helper instead of open-coding the refcounting under #ifdef. Also add sched_cache_group_get() and task_cache_group_get(). Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing") Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/ Reported-by: Hyunwoo Kim <imv4bel@gmail.com> Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev> Co-developed-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@kernel.org> #7.2.x Link: https://patch.msgid.link/ae7081dc54736bf115215f9867abb2711a7403fb.1790035273.git.tim.c.chen@linux.intel.com
2026-09-22sched/cache: Decouple sched_cache_group from mm to fix UAFTim Chen
Currently the sched cache grouping is by mm and the scheduling statistics sched_cache_stat lives in the mm structure. This ties the life cycle of scheduling stats with mm. In account_mm_sched(), the scheduling stats are accessed by task->mm->sc_stat. However, a task may be switching mm on one CPU when another CPU is running account_mm_sched(), and possibly accessing the old mm that was freed. This problem was found when running tests with KASAN by Hyunwoo: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Instead of serializing the mm access by introducing extra acquisition of rq lock in the mm free path, extract sched_cache_stat from mm_struct, rename it as sched_cache_group and manage its life cycle apart from mm_struct with its own ref counting. This allows us in the next patch access sched_cache_group directly from task, and add a refcount on sched_cache_group when a task links to it. This prevents the use after free issue when accessing stale and released old mm and its sched cache stat a task switches to a new mm while account_mm_sched() is done elsewhere. The other benefit of this restructure is in the future, the grouping of tasks to a LLC would have the flexibility to be associated with a user defined grouping, or cgroup, cookie group, numa_group or others instead of just with a single mm address space. Rename sched_cache_stat to sched_cache_group and turn it into a refcounted object allocated from mm_struct. The mm_struct now holds a pointer (sched_cache_grp) to this object instead of embedding it. Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing") Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/ Reported-by: Hyunwoo Kim <imv4bel@gmail.com> Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev> Co-developed-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Cc: <stable@kernel.org> #7.2.x Link: https://patch.msgid.link/91fd1e3266707c865bc9abecfb3e17bc676712df.1790035273.git.tim.c.chen@linux.intel.com
2026-09-22sched/cache: Honor migrate_llc_task semantics in active load balance, to fix ↵Lu Wang
LLC mis-scheduling bug Cache aware scheduling introduced the migrate_llc_task migration type to direct tasks toward their preferred LLC, but its semantics can be lost when passive load balance falls back to active load balance (ALB). This may allow ALB to select a candidate whose preferred LLC does not match the destination, moving it away from its preferred LLC. Example scenario: src_rq has two runnable tasks, p1 and p2. p1 prefers dst_rq (dst_llc), while p2 prefers src_rq (src_llc). In this case, migrate_llc_task is set because src_rq has at least one task, p1, that wants to migrate to dst_rq. In ALB, can_migrate_task() finds p2 and returns true for it, thus moving p2 out of its preferred LLC. Solution: The CPU stopper in ALB constructs a fresh lb_env that does not inherit migration_type from the passive load-balance pass. Two approaches are possible: (a) Add a new member to struct rq so ALB can inherit migrate_llc_task from the passive LB that triggered it. (b) Define a new flag LBF_ACTIVE_LB_LLC and select the stopper callback at kick time to preserve the migration semantics across the asynchronous boundary. We choose (b) because it avoids passing migration_type through the stopper, which would affect the meaning of migration_type for delayed-dequeue tasks. Fixes: e4c9a4cb244a ("sched/cache: Add migrate_llc_task migration type for cache-aware balancing") Suggested-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Lu Wang <wanglu.priv@gmail.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com> Reviewed-by: Chen Yu <yu.c.chen@intel.com> Cc: <stable@kernel.org> # v7.2.x Link: https://patch.msgid.link/cb39f64a17fc2b76097264aaec74a2d6dfff4315.1790035273.git.tim.c.chen@linux.intel.com
2026-09-22sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC ↵graftedTim Chen
mis-scheduling bug alb_break_llc() decides whether to break LLC preference during active load balance. It does so by testing that every runnable fair task on the source rq prefers its LLC: env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_runnable But the two counters cover different sets. nr_pref_llc_running is updated in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued, so it follows queued tasks. h_nr_runnable is updated in set_delayed()/ clear_delayed() and drops delay-dequeued tasks. So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks, alb_break_llc() returns false, and active balance is free to pull a task off its preferred LLC. Active balance only moves runnable tasks, and this is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB skips the per-task test in can_migrate_task(). The runnable set is the one we want. Fix it on the counter side. A task should be counted in nr_pref_llc_running exactly while it is both queued on its preferred LLC (pref_llc_queued) and runnable (!sched_delayed). Define that membership once in task_pref_llc_runnable(), and adjust the counter only through pref_llc_running_inc()/pref_llc_running_dec() from the four sites that change either input: account_llc_enqueue(), account_llc_dequeue(), set_delayed() and clear_delayed(). Gating every update on the same predicate keeps the delay, wake and dequeue paths from double-counting or underflowing; see the comments at those sites for the ordering. nr_llc_running and sd->llc_counts are not touched and stay on queued semantics. Fixes: 714059f79ff0 ("sched/cache: Handle moving single tasks to/from their preferred LLC") Closes: https://lore.kernel.org/lkml/20260827135000.735138-1-zhanxusheng@xiaomi.com/ Reported-by: Zhan Xusheng <zhanxusheng@xiaomi.com> Suggested-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Kayra Cizmeci <kayracizmeci@gmail.com> Cc: <stable@kernel.org> # v7.2.x Link: https://patch.msgid.link/06af61afedac32e6477f57feb4d658f6c411c3af.1790035273.git.tim.c.chen@linux.intel.com
2026-09-21x86/sev: Make vTPM SVSM calls preemption-safeMelody Wang
Two functions in the SVSM vTPM guest implementation do not disable preemption when fetching the SVSM Calling Area Address (CAA). The SVSM CAA is a per-CPU structure. When a thread is preempted and migrated to a different CPU after fetching the per-CPU CAA, the SVSM call will execute on the new CPU with the original CPU's CAA. Which is wrong. Move the CAA fetching operation inside svsm_perform_call_protocol() which disables interrupts around the SVSM call and thus runs preemption-safe. Fixes: 770de678bc28 ("x86/sev: Add SVSM vTPM probe/send_command functions") Signed-off-by: Melody Wang <huibo.wang@amd.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Stefano Garzarella <sgarzare@redhat.com> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/a5bc0d4a2c462a0089109e145c21626b244b2ff0.1789345277.git.huibo.wang@amd.com
2026-09-21ata: libata: Correct libata.force parameter documentationNiklas Cassel
Align the documented libata.force options with their implementation. The force table accepts PIO modes 0 through 6, not mode 7, and ncqati controls NCQ generally rather than only queued TRIM. The max_sec_1024 and max_sec_lba48 options only set transfer size limits. Remove the misleading claim that they can also clear them. Reviewed-by: Damien Le Moal <dlemoal@kernel.org> Link: https://lore.kernel.org/r/20260918124030.1962773-6-cassel@kernel.org Signed-off-by: Niklas Cassel <cassel@kernel.org>
2026-09-21ata: libata-core: Extend Samsung LPM quirk to AMD controllersNiklas Cassel
A Samsung SSD 870 QVO 8TB connected to an AMD 600 Series chipset SATA controller is reported to time out on STANDBY IMMEDIATE during system suspend with med_power_with_dipm enabled. The command completes when using max_performance instead. The existing Samsung LPM quirk only matches ATI controllers, leaving AMD controllers unaffected. Rename it to ATA_QUIRK_NO_LPM_ON_ATI_AND_AMD and extend the vendor check to AMD for the same Samsung SSD model patterns. Keep LPM behavior unchanged for other controller vendors, including Intel. Leave ATA_QUIRK_NO_NCQ_ON_ATI restricted to ATI, since the reported AMD issue concerns LPM rather than NCQ. Link: https://bugzilla.kernel.org/show_bug.cgi?id=221986 Reviewed-by: Damien Le Moal <dlemoal@kernel.org> Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org>> --- Link: https://lore.kernel.org/r/20260918124030.1962773-5-cassel@kernel.org Signed-off-by: Niklas Cassel <cassel@kernel.org>
2026-09-20Input: i8042 - add nomux quirk for Fujitsu LIFEBOOK U7410Lovekesh Solanki
On the Fujitsu LIFEBOOK U7410 the internal keyboard and touchpad die shortly after boot if i8042 multiplexing controller is probed and switched to muxed mode. These multiplexing setup commands disturb the EC emulated keyboard and atkbd reports "Spurious ACK" and "Unknown key pressed" warnings, after which both keyboard and touchpad stop delivering events. Both touchpad and keyboard work in BIOS and GRUB. Disabling multiplexer using i8042.nomux=1 makes both devices work on warm/cold boot. Add a DMI quirk to force nomux mode for this model. Reported-by: Heiko Brey <cico0815@gmx.de> Link: https://lore.kernel.org/all/89f36f3f5e4c2d07b9ff72bc0b355875c336e498.camel@gmx.de/ Signed-off-by: Lovekesh Solanki <lovekeshsolanki00@gmail.com> Link: https://patch.msgid.link/20260920113541.1484808-1-lovekeshsolanki00@gmail.com Cc: stable@vger.kernel.org Signed-off-by: Dmitry Torokhov <dmitry.torokhov@gmail.com>
2026-09-20Input: synaptics - add LEN205b entry to smbus_pnp_ids for ThinkPad T490Lovekesh Solanki
Lenovo ThinkPad T490 has a Synaptics Touchpad that supports SMBus/RMI4 mode, but is not listed in smbus_pnp_ids table. Add LEN205b to smbus_pnp_ids[] passlist. Reported-by: John Rosencutter <rosencutter@gmail.com> Closes: https://lore.kernel.org/all/CAOGuaRh025eKcBf9P8UrvtXNp68vQ0HVF2EbNMp-kpdaC1HNpA@mail.gmail.com/ Signed-off-by: Lovekesh Solanki <lovekeshsolanki00@gmail.com> Link: https://patch.msgid.link/20260920200730.1836756-1-lovekeshsolanki00@gmail.com Cc: stable@vger.kernel.org Signed-off-by: Dmitry Torokhov <dmitry.torokhov@gmail.com>
2026-09-20Linux 7.3-rc4v7.3-rc4Linus Torvalds
2026-09-20net: qrtr: resend HELLO on MHI resumeDaniel J Blueman
Since the MHI HELLO exchange was relocated, it is sent only at device registration. During a suspend-resume cycle, the firmware in WiFi cards such as WCN7850 indefinitely waits for another HELLO, triggering: ath12k_wifi7_pci 0004:01:00.0: timeout while waiting for restart complete ath12k_wifi7_pci 0004:01:00.0: failed to resume core: -110 Fix this by triggering the handshake from resume_early in the MHI transport. Validated on Qualcomm X1E-801800 on Lenovo Slim 7x across 10 suspend-resume cycles. Fixes: 544d85de4dc2 ("net: qrtr: Send HELLO message on endpoint register") Signed-off-by: Daniel J Blueman <daniel@quora.org> Reviewed-by: Manivannan Sadhasivam <manivannan.sadhasivam@oss.qualcomm.com> Reported-by: Jeff Johnson <jeff.johnson@oss.qualcomm.com> Link: https://lore.kernel.org/all/6257c447-788d-4362-851e-0d552bcf7c56@oss.qualcomm.com/ Tested-by: Jeff Johnson <jeff.johnson@oss.qualcomm.com> Reported-by: Vlastimil Babka (SUSE) <vbabka@suse.com> Link: https://lore.kernel.org/all/ab1491bb-cca5-4145-ac7d-31c966abf7b4@suse.com/ Tested-by: Vlastimil Babka (SUSE) <vbabka@kernel.org> Reported-by: Takashi Iwai <tiwai@suse.de> Link: https://lore.kernel.org/all/87a4plsg4w.wl-tiwai@suse.de/ Tested-by: Takashi Iwai <tiwai@suse.de> Signed-off-by: Thorsten Leemhuis <linux@leemhuis.info> Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2026-09-20Merge tag 'dmaengine-fix-7.3' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/vkoul/dmaengine Pull dmaengine fixes from Vinod Koul: - A couple of fixes in core around dma_chan_put() for kref underflow, use-after-free and waiting for rcu readers for dma devices - mmp sg length and wrong extended DRCMR base for SpacemiT K3 - hardware buffer descriptor chain fix for xilinx dma - sun6i fixes for status behaviour and dma position registers - runtime pm reference leak fix for sprd driver * tag 'dmaengine-fix-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/vkoul/dmaengine: dmaengine: mmp_pdma: fix wrong sg length in mmp_pdma_prep_slave_sg() dmaengine: xilinx_dma: Fix hardware buffer descriptor chain after cyclic DMA dmaengine: pxa: fix double counting of the hw descriptors dmaengine: sun6i: fix undefined behaviour in sun6i_dma_tx_status dmaengine: sun6i: fix non-atomic read of DMA position registers dmaengine: xilinx_dma: Fix hardware buffer descriptor reuse order dmaengine: wait for RCU readers before releasing dma_device dmaengine: fix use-after-free in dma_chan_put() and dma_release_channel() dmaengine: Fix device kref underflow in dma_chan_put() dmaengine: add dma_device_get() helper dmaengine: sprd: Fix runtime PM reference leak in probe dmaengine: ti: k3-udma-glue: fix NULL dereference in k3_udma_glue_release_rx_chn() dmaengine: mmp_pdma: fix wrong extended DRCMR base for SpacemiT K3
2026-09-20Merge tag 'phy-fixes-7.3' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/phy/linux-phy Pull phy fixes from Vinod Koul: - avoid atomic context delay in renesas driver - TMDS and PLL rate calculation fixes for mediatek driver * tag 'phy-fixes-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/phy/linux-phy: phy: mediatek: phy-mtk-hdmi-mt8195: Fix TMDS clk bit ratio setting phy: mediatek: phy-mtk-hdmi-mt8195: Fix PLL calc divisor overflow phy: renesas: rcar-gen3-usb2: Avoid long delay in atomic context
2026-09-20Merge tag 'soundwire-7.3-fixes' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/vkoul/soundwire Pull soundwire fixes from Vinod Koul: - Cadence: ensure work completion before clock stop - Disable ghost Realtek on Asus GX651AX * tag 'soundwire-7.3-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vkoul/soundwire: soundwire: cadence_master: wait and cancel cdns->work before clock stop soundwire: dmi-quirks: Disable ghost Realtek on Asus GX651AX
2026-09-20Merge tag 'x86-urgent-2026-09-20' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull x86 fixes from Ingo Molnar: - Reject the loading of a potentially problematic microcode version on Intel Granite Rapids systems (Chang S. Bae) - On FRED, reconstruct the proper #GP context for rejected INT instructions, to fix a signal ABI regression (Matthew Schwartz) - Add a test for this signal ABI regression the x86 self-test suite (Matthew Schwartz) - Don't emit the new and not yet properly supported EGPR instructions (%r16-%r31) on CONFIG_X86_NATIVE_CPU=y builds (Chang S. Bae) * tag 'x86-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: x86/build/64: Prevent native builds from generating EGPR use selftests/x86: Check signal state for rejected software interrupts x86/fred: Reconstruct the #GP context for rejected INT instructions x86/microcode/intel: Reject problematic loading on Granite Rapids systems
2026-09-20Merge tag 'timers-urgent-2026-09-20' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull timer race fixes from Ingo Molnar: - Fix timer signal <-> exec() race, to prevent UAF (Thomas Gleixner) - Clean up POSIX CPU timers right after de_thread(), to prevent UAF (Hyunwoo Kim) - Fix POSIX CPU timers race between expiry and timer_settime(), to prevent UAF (Thomas Gleixner) * tag 'timers-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: posix-cpu-timers: Prevent freeing a timer which is queued on the expiry list exec: Cleanup POSIX timers right after de_thread() signal: Prevent exec() race
2026-09-20Merge tag 'sched-urgent-2026-09-20' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull scheduler fix from Ingo Molnar: - Avoid false positive migration warning for proxy donors (Andrea Righi) * tag 'sched-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: sched/core: Avoid false migration warning for proxy donors
2026-09-20Merge tag 'perf-urgent-2026-09-20' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull perf events fixes from Ingo Molnar: - Fix crash when probing CS CALL instructions (Jinke Han) - Fix NULL pointer crash during module unload (Vinay Belgaumkar) * tag 'perf-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: perf: Fix null pointer access in is_include_guest_event() x86/kprobes: Fix crash when probing CS CALL instructions
2026-09-20Merge tag 'objtool-urgent-2026-09-20' of ↵Linus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull objtool fix from Ingo Molnar: - Fix objtool build error on systems where libopcodes is present, but development headers (binutils-dev) are not (Ulises Mendez Martinez) * tag 'objtool-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: objtool: Validate disassembler headers in libopcodes probe
2026-09-20Merge tag 'locking-urgent-2026-09-20' of ↵graftedLinus Torvalds
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull futex fix from Ingo Molnar: - Also allocate a default private futex hash on vfork() as well, to avoid races with (private) futex waiters (Peter Zijlstra) * tag 'locking-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: futex: Also allocate private hash on vfork()
2026-09-19posix-cpu-timers: Prevent freeing a timer which is queued on the expiry listThomas Gleixner
Kijo analyzed another race in the POSIX CPU timer code: Commit bf635681c906 converted cpu_timer::firing from a tristate value to a boolean. This lost the distinction between "not owned by the firing list" and "still owned, but delivery was canceled". The resulting race is: expiry handler timer_settime() timer_delete() -------------- --------------- -------------- collect timer onto private firing list firing = true observes firing = true firing = false return TIMER_RETRY wait for handler observes firing = false finish deletion unhash and free timer resume list traversal read freed elist.next -> UAF The firing bit is clearly the wrong indicator since that commit. Check whether the timer is queued on the expiry list or not instead. If it is queued clear the firing bit to prevent signal delivery as before and return TIMER_RETRY so the caller unlocks the timer which allows the expiry code to make progress and remove it from the list. Fixes: bf635681c906 ("posix-cpu-timers: Cleanup the firing logic") Reported-by: Kijo Park <red993688@gmail.com> Debugged-by: Kijo Park <red993688@gmail.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Tested-by: Kijo Park <red993688@gmail.com> Reviewed-by: Frederic Weisbecker <frederic@kernel.org> Cc: stable@vger.kernel.org
2026-09-19selftests/sched_ext: Check the cmask cid-form ops.enable() receivesTejun Heo
cid-form ops.enable() now hands the task's cmask to the scheduler and set_cmask() repeats it right after. Add a cid-form selftest that checks both against p->cpus_ptr, that they match each other, that the initial set_cmask() lands before set_weight() and before the task first becomes runnable, and that set_cmask() never precedes enable(), across class-switch enables, fork-path enables and live affinity changes. v2: Mismatch details returned through a caller-local struct instead of globals, alloc_words validated in the header check, loop bounded by nr_cids directly (Andrea Righi). Signed-off-by: Tejun Heo <tj@kernel.org>
2026-09-19sched_ext: Pass the initial cmask to cid-form ops.enable()Tejun Heo
The cid-form API has an obvious hole. A task's cid mask is only visible through ops.set_cmask(), which fires on affinity changes and class switches but not when a task enters a scheduler through fork, sub-sched enable or re-home, and there is no p->cpus_ptr equivalent to fall back on. Schedulers work around it by seeding the mask in ops.init_task() from p->cpus_ptr cid by cid, which is subtly wrong: on sub-sched enable and re-home, an affinity change between init_task() and enable() is delivered to the sched the task is still on, and nothing corrects the new sched's copy afterwards. Fix it by adding struct scx_enable_args to cid-form ops.enable() carrying the task's cmask, built in the per-cpu scratch under the rq lock as the task enters the scheduler, and calling set_cmask() with the same mask right after enable(), ahead of set_weight(). A scheduler can then track affinity in set_cmask() alone, and scx_qmap drops its init_task() seed. set_cmask() no longer fires for a cid-form task before it is enabled, and the class-switch republish in switching_to_scx() is limited to the cpu form. This changes the cid-form ops.enable() signature, which is fine as the cid-form API is still considered unreleased. An args struct rather than a bare cmask argument leaves room for more initial state without another signature change, and the cmask travels as a plain arena address because BTF can't type arena struct members yet. v2: The cmask travels as a u64 arena address, cmask_arena_addr, instead of a kernel-typed pointer, with the typing limitation and the planned typed alias documented (Sashiko review). v3: The initial set_cmask() is delivered before set_weight() so that the mask is in place when weight-dependent state is derived (Andrea Righi). Selftest added. Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-09-18x86/build/64: Prevent native builds from generating EGPR useChang S. Bae
Omar reports that CONFIG_X86_NATIVE_CPU=y allows builds to opportunistically emit instructions using %r16-%r31 (EGPRs) when the build host supports APX since the commit: ea1dcca1de12 ("x86/kbuild/64: Add the CONFIG_X86_NATIVE_CPU option to locally optimize the kernel with '-march=native'") But the kernel is not yet prepared to use new registers internally. For example, there is no context-switch support for general in-kernel use. Explicitly disable EGPR use when building with -march=native. For C, since GCC 14 and Clang 18, both compilers support suppressing EGPR use with -mno-apx-features=egpr, whose availability can be detected via cc-option. For Rust, pass features=-apxf through the generated JSON to avoid unstable-feature warnings, see https://github.com/rust-lang/rust/issues/139284 Note Rust only accepts the option to disable APX instructions entirely or not. Support for this gating also depends on the Rust/LLVM combination. Rust 1.88 introduced the `apxf` feature option, but versions prior to 1.93 may emit an `apxf` attribute to the backend that only LLVM 23 or later can interpret. Restrict native Rust builds accordingly. Fixes: ea1dcca1de12 ("x86/kbuild/64: Add the CONFIG_X86_NATIVE_CPU option to locally optimize the kernel with '-march=native'") Reported-by: Omar Avelar <omar.avelar@intel.com> Signed-off-by: Chang S. Bae <chang.seok.bae@intel.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Nathan Chancellor <nathan@kernel.org> Acked-by: Miguel Ojeda <ojeda@kernel.org> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260916230003.1144622-1-chang.seok.bae@intel.com
2026-09-18PCI: of_property: Omit bus properties without a subordinate busAngel J
A bridge (a device with a Type 1 header) may not have a secondary bus allocated (pdev->subordinate), e.g., if there are no available bus numbers or the bridge secondary/subordinate bus numbers are not writable. The dynamic OF helpers of_pci_prop_bus_range() and of_pci_prop_intr_map() dereference pdev->subordinate without checking it. When CONFIG_PCI_DYNAMIC_OF_NODES is enabled, this can cause a NULL pointer dereference and early boot hang. Generate 'bus-range' and 'interrupt-map' properties only when a subordinate bus exists. Keep the node and its remaining properties for bridges without one. The problem was latent since 407d1a51921e ("PCI: Create device tree node for bridge"), but wasn't reachable until 1f340724419e ("PCI: of: Create device tree PCI host bridge node"), which appeared in v6.15. Before 1f340724419e, of_pci_make_dev_node() returned early because the parent OF node was missing. Fixes: 407d1a51921e ("PCI: Create device tree node for bridge") Signed-off-by: Angel J <iamanaws@httpd.dev> [bhelgaas: move pdev->subordinate test to callees, commit log] Signed-off-by: Bjorn Helgaas <bhelgaas@google.com> Cc: stable@vger.kernel.org # v6.6+ Link: https://patch.msgid.link/20260918195540.GA1187209@bhelgaas
2026-09-18PCI: Fix BAR resize for devices on a root busLiz Fong-Jones
pci_do_resource_release_and_resize() releases device BARs that share a bridge window with the BAR being resized, but when the device sits directly on a root bus (pdev->bus->self == NULL) it then skips resource assignment entirely and returns success, leaving the BARs it just released unassigned (IORESOURCE_UNSET). Skipping pbus_reassign_bridge_resources() is correct in that case -- there is no bridge window to adjust -- but the device BARs still have to be reassigned. Before the BAR release was consolidated into the PCI core, this case worked for amdgpu because the driver released the BARs itself and then called pci_assign_unassigned_bus_resources() unconditionally after the resize, which assigns unassigned device BARs also on a root bus. Commit db92e3fef53e ("drm/amdgpu: Remove driver side BAR release before resize") removed that call, so nothing assigns the released BARs anymore. This breaks amdgpu completely on the SolidRun HoneyComb LX2K (NXP LX2160A, arm64, ACPI), where ACPI doesn't expose the Root Port so the GPU endpoint appears directly on a "root bus" of its segment: amdgpu 0004:01:00.0: BAR 0 [mem 0xa400000000-0xa40fffffff 64bit pref]: releasing amdgpu 0004:01:00.0: BAR 2 [mem 0xa410000000-0xa4101fffff 64bit pref]: releasing amdgpu 0004:01:00.0: sw_init of IP block <gmc_v8_0> failed -19 amdgpu 0004:01:00.0: amdgpu_device_ip_init failed amdgpu 0004:01:00.0: Fatal error during GPU init No error is logged because the resize path reports success; amdgpu then finds BAR 0 IORESOURCE_UNSET and bails out with -ENODEV. When there is no upstream bridge, call pci_bus_assign_resources() on the root bus to place the BARs released above, using the same alignment-sorted algorithm as normal enumeration instead of a manual per-BAR loop. This also walks the rest of the hierarchy under the root bus, as pci_assign_unassigned_bus_resources() used to for amdgpu before commit db92e3fef53e ("drm/amdgpu: Remove driver side BAR release before resize") removed that call -- the core-side fix that commit asked for ("such a problem should be fixed inside pci_resize_resource() instead"). pci_bus_assign_resources() returns void, so failure is detected by checking whether the released BARs are still assigned afterward; if not, roll back as in the bridged case. This is stricter than the bridged path -- it fails on any unplaced resource, not just required ones -- since a root bus typically has one shared window, and failing loudly seemed better than leaving something silently unassigned. The root bus path also had a locking bug that any fix here necessarily touches: the old "goto out" jumped to up_read(&pci_bus_sem) without a matching down_read() (as does the "goto restore" taken when pci_dev_res_add_to_list() fails in the release loop). Take pci_bus_sem before the BAR release loop so every path through the function holds it exactly once. Fixes: 337b1b566db0 ("PCI: Fix restoring BARs on BAR resize rollback path") Link: https://bugs.launchpad.net/ubuntu/+source/linux-hwe-7.0/+bug/2159596 Suggested-by: Ilpo Järvinen <ilpo.jarvinen@linux.intel.com> Assisted-by: Claude:claude-fable-5 checkpatch Assisted-by: Claude:claude-sonnet-5 Signed-off-by: Liz Fong-Jones <lizf@honeycomb.io> [bhelgaas: commit log] Signed-off-by: Bjorn Helgaas <bhelgaas@google.com> Reviewed-by: Ilpo Järvinen <ilpo.jarvinen@linux.intel.com> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260918035633.566823-1-lizf@honeycomb.io
2026-09-18Merge tag 'tags/spacemit-clk-fixes-for-7.3-1' into clk-fixesBrian Masney
RISC-V SpacemiT clock fixes for v7.3 - Add CPU PLL rate tables
2026-09-18selftests/x86: Check signal state for rejected software interruptsMatthew Schwartz
Add a test of the signal ABI for INT instructions in both 32-bit and 64-bit processes. Check the signal number, trap number, error code, si_code, si_addr, instruction pointer and RF/TF state against legacy IDT behavior. Include a 15-byte prefixed INT to check that IP uses the hardware instruction length. Exercise both INT3 encodings, INT4, UD2 and HLT to cover the unchanged trap and fault paths. Run each instruction with TF clear and set. Resume at a known NOP after handling the signal and check that single-stepping traps after the NOP. Also drive INT 0x2d under ptrace, which resumes through the fault frame rather than sigreturn and so exposes a stale FRED software event flag. Start from an INT3 stop, whose FRED frame has no software event flag, instead of the syscall frame of raise(SIGSTOP). Single-step into the INT and check that the fault reports its address. Then suppress SIGSEGV and resume at the NOP, once with PTRACE_SINGLESTEP and once with PTRACE_CONT and TF set. Section 6.2.3 of the Intel FRED specification [1] specifies the immediate single-step trap caused by returning with both that flag and TF set. Check that each trap occurs after the NOP, rather than at its address. Report whether the CPU supports FRED, since a pass looks the same on either entry path. INT 0x80 with IA32 emulation disabled and a 64-bit tracer of a 32-bit tracee are not covered. Both variants pass all 29 checks on a non-FRED AMD host and on Panther Lake with FRED enabled and the fix applied. With the same binaries on unpatched Panther Lake, 16 signal-context checks fail and the first ptrace check reports the IP after the INT. The two dependent ptrace resume checks are not reached. [1] Intel Flexible Return and Event Delivery (FRED) Specification, revision 9.0 (346446-009US), section 6.2.3. Signed-off-by: Matthew Schwartz <matthew.schwartz@linux.dev> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://cdrdv2.intel.com/v1/dl/getContent/678938 # [1] Link: https://patch.msgid.link/20260917230907.2080792-3-matthew.schwartz@linux.dev
2026-09-18x86/fred: Reconstruct the #GP context for rejected INT instructionsMatthew Schwartz
FRED event delivery does not use the IDT, so the gate DPL check that rejects a user INT n falls to software (Intel FRED specification [1], section 8.3). fred_intx() rejects the same vectors as IDT delivery, but reports a zero error code and the IP after the INT. This breaks the signal ABI. Wine uses the error code to recognize INT 0x2d, so the changed context turns a handled breakpoint into an access violation in Elden Ring. Rewind IP using the instruction length in the augmented SS and synthesize the IDT selector error code, (vector << 3) | 2. Set RF in the saved flags, as the CPU does for a #GP fault. Section 5.2.1 defines the saved vector, instruction length and RF state. The supplied length handles prefixes without reading user memory. Limit the changes to already-rejected software interrupts, preserving the accepted INT3, INT4 and enabled INT80 paths and hardware exceptions. With IA32 emulation disabled, INT 0x80 now reports the same #GP as the DPL 0 gate IDT installs there. The rewound IP also stops fixup_iopl_exception() from inspecting the byte after the INT. Also clear the software event flag. Section 6.2.3 specifies that ERETU with this flag and TF set traps before executing any user instruction. A tracer that suppresses SIGSEGV and resumes with TF set expects the next instruction to run first, as after IRET. The sigreturn path clears the same flag for this reason in prevent_single_step_upon_eretu(). [1] Intel Flexible Return and Event Delivery (FRED) Specification, revision 9.0 (346446-009US), sections 5.2.1, 6.2.3 and 8.3. Fixes: 14619d912b65 ("x86/fred: FRED entry/exit and dispatch code") Closes: https://gitlab.freedesktop.org/mesa/mesa/-/work_items/15745 Closes: https://gitlab.freedesktop.org/mesa/mesa/-/work_items/16132 Reported-by: Paul Gofman <pgofman@codeweavers.com> Signed-off-by: Matthew Schwartz <matthew.schwartz@linux.dev> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Reviewed-by: H. Peter Anvin <hpa@zytor.com> Link: https://cdrdv2.intel.com/v1/dl/getContent/678938 # [1] Link: https://patch.msgid.link/20260917230907.2080792-2-matthew.schwartz@linux.dev
2026-09-18sched/core: Avoid false migration warning for proxy donorsAndrea Righi
Proxy execution can move a blocked donor's scheduling context to the lock owner's CPU even when the donor is migration-disabled. The donor does not execute there, and its original execution CPU remains recorded in wake_cpu. set_task_cpu() warns unconditionally for migration-disabled tasks, so a subsequent proxy migration or the wakeup path returning the donor home triggers a false positive: moving a blocked scheduling context does not violate the migration-disabled execution context. For example, creating a mutex owner on CPU1 and a migration-disabled waiter on CPU0 can trigger the following warning: proxy_migrate_repro: donor blocking on CPU0 with migration disabled proxy_migrate_repro: donor moved from CPU0 to CPU1 WARNING: kernel/sched/core.c:3389 at set_task_cpu+0x1d3/0x280 ... Call Trace: try_to_wake_up+0x43f/0x780 __mutex_unlock_slowpath+0x330/0x540 owner_fn+0x9f/0xc0 [proxy_migrate_repro] ... proxy_migrate_repro: donor woke on CPU0, task_cpu=0 proxy_migrate_repro: completed Exclude blocked proxy donors from the warning. The proxy wakeup path restores an executable placement before clearing the blocked state. Fixes: b049b81bdff6 ("sched: Handle blocked-waiter migration (and return migration)") Signed-off-by: Andrea Righi <arighi@nvidia.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Acked-by: John Stultz <jstultz@google.com> Link: https://patch.msgid.link/20260915184101.2621252-1-arighi@nvidia.com
2026-09-18perf: Fix null pointer access in is_include_guest_event()Vinay Belgaumkar
A typical module unload occurring event when there is an active perf connection leads to freeing of the pmu pointer. The call log is something like: .. __pmu_detach_event pmu_detach_event pmu_detach_events perf_pmu_unregister .. __pmu_detach_event() sets event->pmu to null. When the perf connection finally is closed, the following stack trace is observed: Oops: general protection fault, kernel NULL pointer dereference ... RIP: 0010:_free_event+0x3e/0x370 ... Call Trace: ... perf_event_release_kernel+0x260/0x2d0 perf_release+0x12/0x20 A call to mediated_pmu_unaccount_event() inside _free_event() is the root cause of this crash. Adding a check inside is_include_guest_event() ensures we don't accidentally access a null pmu ptr. In addition to this, we will now call mediated_pmu_unaccount_event() before clearing the pmu ptr so that nr_include_guest_events counts are maintained correctly. Fixes: eff95e170275 ("perf: Add APIs to create/release mediated guest vPMUs") Assisted-by: Claude:Claude-Sonnet-5 Signed-off-by: Vinay Belgaumkar <vinay.belgaumkar@intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com> Link: https://patch.msgid.link/20260904181625.1394082-1-vinay.belgaumkar@intel.com
2026-09-18fprobe: Terminate the fgraph_data list when the reservation is not filledDavid Carlier
fprobe_fgraph_entry() reserves shadow stack space for every fprobe with an exit handler, but only fills it for those whose entry handler returns 0. fgraph_reserve_data() does not clear the area, so fprobe_return() parses the unused tail as headers left over from an earlier call, and an exit handler can run twice or despite its entry handler asking to skip it. Write a zero word after the last entry to terminate the walk. A zeroed slot does not decode to a NULL fprobe on the arches that encode the header into one unsigned long, since arch_decode_fprobe_header_fp() ORs in FPROBE_HEADER_MSB_PATTERN, so make read_fprobe_header() return NULL for a zeroed slot. Link: https://lore.kernel.org/all/20260917212407.384468-1-devnexen@gmail.com/ Fixes: e0a384434ae1 ("tracing: fprobe: do not zero out unused fgraph_data") Cc: stable@vger.kernel.org Suggested-by: Masami Hiramatsu (Google) <mhiramat@kernel.org> Signed-off-by: David Carlier <devnexen@gmail.com> Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
2026-09-17x86/microcode/intel: Reject problematic loading on Granite Rapids systemsChang S. Bae
Microcode updates can usually jump revisions. However, there is an erratum on Granite Rapids systems. If they "jump over" revision 0x1000405, they result in an #MC. Avoid it. Signed-off-by: Chang S. Bae <chang.seok.bae@intel.com> Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de> Reviewed-by: Dave Hansen <dave.hansen@linux.intel.com> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260916225939.1144524-1-chang.seok.bae@intel.com
2026-09-17selftests/sched_ext: Test that ops.dequeue() can iterate the consumed DSQfangqiurong
Add a scheduler whose ops.dequeue() iterates the source user DSQ with bpf_iter_scx_dsq. The iteration takes the DSQ's raw spinlock; on a kernel that runs ops.dequeue() while the consume path still holds that lock, the first task consumed self-deadlocks the CPU with IRQs disabled. The watchdog cannot recover from that state, so on an unfixed kernel this test wedges the system instead of failing cleanly. On a fixed kernel the scheduler runs clean and the test passes. Signed-off-by: fangqiurong <fangqiurong@kylinos.cn> Signed-off-by: Tejun Heo <tj@kernel.org>
2026-09-17sched_ext: Don't run ops.dequeue() with a DSQ lock heldfangqiurong
ops.dequeue() is invoked with the source user DSQ's lock still held on the consume and move paths (scx_consume_dispatch_q(), move_task_between_dsqs()). A BPF scheduler which locks the source user DSQ from ops.dequeue() - e.g. by iterating it with bpf_iter_scx_dsq - self-deadlocks. ops.dequeue() can only call the "any" kfuncs and none of them can lock a builtin DSQ, so the global and bypass paths can't deadlock; however, all DSQ locks share one lockdep class, so iterating any user DSQ from ops.dequeue() on those paths trips the recursion check. Move the invocation after the DSQ unlock on all three paths. SCX_TASK_IN_CUSTODY is cleared under the lock serializing the transfer so that the callback is invoked exactly once. Fixes: ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics") Cc: stable@vger.kernel.org # v7.1+ Acked-by: Andrea Righi <arighi@nvidia.com> Signed-off-by: fangqiurong <fangqiurong@kylinos.cn> Signed-off-by: Tejun Heo <tj@kernel.org>
2026-09-17sched_ext: Derive SCX_RQ_IN_WAKEUP from the core enqueue flagsTejun Heo
schedule_deferred_locked() skips scheduling a deferred action while SCX_RQ_IN_WAKEUP is set and relies on the task_woken_scx() call that follows a wakeup enqueue to run it. enqueue_task_scx() sets the flag from the merged enqueue flags, which include the flags stashed for a remote activation. move_remote_task_to_local_dsq() thus sets SCX_RQ_IN_WAKEUP on the destination rq when the moved task was woken up, although no task_woken_scx() follows that activation. An IMMED insert into a busy destination requests a local reenqueue during that enqueue. The request gets linked but not scheduled and stays pending until an unrelated wakeup or preemption on that CPU runs the deferred actions. The IMMED task sits behind the running task in the meantime. If nothing runs them before the scheduler is disabled, the request outlives the scheduler and points into its freed per-cpu area, which the next scheduler dereferences from run_deferred(). Test the core enqueue flags for the wakeup bit. Only the core's wakeup path is followed by task_woken_scx(). Fixes: 57ccf5ccdc56 ("sched_ext: Fix enqueue_task_scx() truncation of upper enqueue flags") Cc: stable@vger.kernel.org # v7.1+ Reported-by: Andrea Righi <arighi@nvidia.com> Link: https://lore.kernel.org/all/20260916145807.3250167-1-arighi@nvidia.com/ Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-09-17x86/kprobes: Fix crash when probing CS CALL instructionsgraftedJinke Han
When using eBPF to probe CS CALL instructions within a function, a crash can be triggered. The eBPF tool probes offset 257 of the __hrtimer_run_queues() function: <__hrtimer_run_queues+249>: nopl 0x0(%rax,%rax,1) <__hrtimer_run_queues+254>: mov %r14,%rdi <__hrtimer_run_queues+257>: cs call <__x86_indirect_thunk_r12> <__hrtimer_run_queues+263>: mov %eax,%r12d <__hrtimer_run_queues+266>: xchg %ax,%ax <__hrtimer_run_queues+268>: mov %r13,%rdi Which triggers this crash: BUG: unable to handle page fault for address: 00000000000f41c9 #PF: supervisor write access in kernel mode #PF: error_code(0x0002) - not-present page PGD 0 P4D 0 Oops: 0002 [#1] SMP NOPTI CPU: 1 PID: 0 Comm: swapper/1 Kdump: loaded Tainted: P RIP: 0010:__hrtimer_run_queues+0x106/0x230 Note that __hrtimer_run_queues+0x106 is __hrtimer_run_queues+262, which is at the 6th byte of the above CS CALL instruction. Since the CS CALL instruction occupies 6 bytes, the exception occurred in the middle of that call instruction. The root cause is that when using eBPF tools to probe in the middle of a function, a kprobe with INT3 is used as the underlying implementation. During single-step emulation of the original CALL instruction, int3_emulate_call() assumes that the probed CALL instruction is 5 bytes long. However, the actual CS-prefixed CALL instruction occupies 6 bytes, so it constructs an incorrect exception return address. When the CPU returns from the kprobe handler, the next instruction to be executed is at the address of the last byte of that CS CALL instruction. Coincidentally, starting from that address, the CPU fetches and decodes a completely different instruction, which ultimately triggers a kernel crash. Fix the issue by using the actual instruction length obtained from the instruction decoder when constructing the exception return address, rather than relying on the hardcoded CALL_INSN_SIZE macro. [ mingo: Refined the changelog ] Fixes: 6256e668b7af ("x86/kprobes: Use int3 instead of debug trap for single-step") Suggested-by: Masami Hiramatsu (Google) <mhiramat@kernel.org> Signed-off-by: Jinke Han <jinkehan@didiglobal.com> Signed-off-by: Ingo Molnar <mingo@kernel.org> Reviewed-by: Masami Hiramatsu (Google) <mhiramat@kernel.org> Acked-by: Yafang Shao <laoar.shao@gmail.com> Acked-by: Borislav Petkov <bp@alien8.de> Cc: Peter Zijlstra <peterz@infradead.org> Link: https://patch.msgid.link/20260908073742.GA10517@didi-ThinkCentre-M920t-N000
2026-09-16selftests: cgroup: give the O_TMPFILE open in get_temp_fd() a modeEva Kurchatova
O_TMPFILE, like O_CREAT, needs the third argument. Without it glibc refuses the call at compile time as soon as fortification is on: In function 'open', inlined from 'get_temp_fd' at test_memcontrol.c:33:9: /usr/include/bits/fcntl2.h:52:11: error: call to '__open_missing_mode' declared with attribute error: open with O_CREAT or O_TMPFILE in second argument needs 3 arguments The fortify checks take effect only once the compiler optimises, and cgroup/Makefile builds with "-Wall -pthread" alone, so this goes unnoticed in a plain build. Building the tests with the flags distributions commonly use, -O2 -D_FORTIFY_SOURCE=3, loses test_memcontrol entirely. Fixes: 84092dbcf901 ("selftests: cgroup: add memory controller self-tests") Signed-off-by: Eva Kurchatova <eva.kurchatova@virtuozzo.com> Signed-off-by: Tejun Heo <tj@kernel.org>
2026-09-16rust_binderfs: add transaction_report feature entryCarlos Llamas
rust_binderfs is missing the equivalent of commit f37b55ded8ed ("binder: add transaction_report feature entry"), which adds "transaction_report" to the binderfs feature list. This helps userspace determine if the BINDER_CMD_REPORT from the generic netlink API is supported. Cc: stable <stable@kernel.org> Fixes: f14e0c8183bc ("rust_binder: report netlink transactions") Signed-off-by: Carlos Llamas <cmllamas@google.com> Link: https://patch.msgid.link/20260916120833.593407-1-cmllamas@google.com Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2026-09-16rust_binder: reschedule node refcount update on thread exitAlice Ryhl
When a thread exits via BINDER_THREAD_EXIT, its pending work items are cancelled. If a thread exits while holding a pending node refcount increment (e.g. pushed as deferred work to that thread), the refcount increment was previously dropped because Node::cancel() and NodeWrapper::cancel() were no-ops. Dropping the refcount update leaves the node's delivery state and count state desynchronized, and userspace will not receive the notification, which can cause the node to never be freed from the process's nodes tree when all external references are dropped. Fix this by implementing DeliverToRead::cancel() for Node and NodeWrapper to move the pending refcount update to the process's work queue on thread exit so another thread can deliver it to userspace. Cc: stable <stable@kernel.org> Fixes: eafedbc7c050 ("rust_binder: add Rust Binder driver") Signed-off-by: Alice Ryhl <aliceryhl@google.com> Link: https://patch.msgid.link/20260903-binder-thread-exit-node-v1-1-be09ff14f6a4@google.com Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2026-09-16rust_binder: cancel deferred work items in thread exitAlice Ryhl
If there are deferred work items on the thread todo list, then they are not cleaned up in the Thread::release() method. Thus, update the code to clean up the work items even if they are deferred. This can happen if the thread dies while it has an active outgoing transaction. Cc: stable <stable@kernel.org> Fixes: eafedbc7c050 ("rust_binder: add Rust Binder driver") Signed-off-by: Alice Ryhl <aliceryhl@google.com> Link: https://patch.msgid.link/20260903-binder-exit-get-work-v1-1-2d6129a238df@google.com Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2026-09-16binderfs: fix UAF write in binder_add_devicePeiyang He
binderfs_binder_device_create() publishes the new dentry with d_make_persistent() and then calls simple_done_creating(), which drops the parent directory lock and the creator's dentry reference. It then calls binder_add_device() to register the device in the global binder_devices list. After simple_done_creating() releases the parent directory lock, a concurrent unlinkat() can remove the new device entry. Dropping the creator's dentry reference can then trigger binderfs_evict_inode(), freeing the device. binder_add_device() later accesses the freed object, causing UAF write. Found by a modified Syzkaller: BUG: KASAN: slab-use-after-free in hlist_add_head include/linux/list.h:1073 [inline] BUG: KASAN: slab-use-after-free in binder_add_device+0xa9/0xc0 drivers/android/binder.c:7068 Write of size 8 at addr ffff88805b340c00 by task syz.1.532/11389 CPU: 0 UID: 0 PID: 11389 Comm: syz.1.532 Not tainted 7.2.0 #4 PREEMPT(full) Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014 Call Trace: <TASK> __dump_stack lib/dump_stack.c:94 [inline] dump_stack_lvl+0x10e/0x1f0 lib/dump_stack.c:120 print_address_description mm/kasan/report.c:378 [inline] print_report+0xf7/0x600 mm/kasan/report.c:482 kasan_report+0xe4/0x120 mm/kasan/report.c:595 hlist_add_head include/linux/list.h:1073 [inline] binder_add_device+0xa9/0xc0 drivers/android/binder.c:7068 binderfs_binder_device_create.isra.0+0x724/0x990 drivers/android/binderfs.c:196 binder_ctl_ioctl+0x186/0x1b0 drivers/android/binderfs.c:241 vfs_ioctl fs/ioctl.c:51 [inline] __do_sys_ioctl fs/ioctl.c:597 [inline] __se_sys_ioctl fs/ioctl.c:583 [inline] __x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f RIP: 0033:0x7ff9027a833d Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 b0 ff ff ff f7 d8 64 89 01 48 RSP: 002b:00007ff903674018 EFLAGS: 00000246 ORIG_RAX: 0000000000000010 RAX: ffffffffffffffda RBX: 00007ff902a35fa0 RCX: 00007ff9027a833d RDX: 0000200000000500 RSI: 00000000c1086201 RDI: 0000000000000004 RBP: 00007ff902850733 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000 R13: 00007ff902a36038 R14: 00007ff902a35fa0 R15: 00007ffd7166bfa0 </TASK> Allocated by task 11389: kasan_save_stack+0x33/0x60 mm/kasan/common.c:57 kasan_save_track+0x14/0x30 mm/kasan/common.c:78 poison_kmalloc_redzone mm/kasan/common.c:398 [inline] __kasan_kmalloc+0xaa/0xb0 mm/kasan/common.c:415 kasan_kmalloc include/linux/kasan.h:263 [inline] __kmalloc_cache_noprof+0x2e4/0x6f0 mm/slub.c:5489 _kmalloc_noprof include/linux/slab.h:988 [inline] _kzalloc_noprof include/linux/slab.h:1309 [inline] binderfs_binder_device_create.isra.0+0x17a/0x990 drivers/android/binderfs.c:148 binder_ctl_ioctl+0x186/0x1b0 drivers/android/binderfs.c:241 vfs_ioctl fs/ioctl.c:51 [inline] __do_sys_ioctl fs/ioctl.c:597 [inline] __se_sys_ioctl fs/ioctl.c:583 [inline] __x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f Freed by task 11389: kasan_save_stack+0x33/0x60 mm/kasan/common.c:57 kasan_save_track+0x14/0x30 mm/kasan/common.c:78 kasan_save_free_info+0x3b/0x60 mm/kasan/generic.c:584 poison_slab_object mm/kasan/common.c:253 [inline] __kasan_slab_free+0x5f/0x80 mm/kasan/common.c:285 kasan_slab_free include/linux/kasan.h:235 [inline] slab_free_hook mm/slub.c:2677 [inline] slab_free mm/slub.c:6377 [inline] kfree+0x2fc/0x6e0 mm/slub.c:6692 binderfs_evict_inode+0x1e8/0x260 drivers/android/binderfs.c:268 evict+0x3c2/0xad0 fs/inode.c:825 iput_final fs/inode.c:2019 [inline] iput fs/inode.c:2068 [inline] iput+0x79a/0xd30 fs/inode.c:2031 dentry_unlink_inode+0x27f/0x460 fs/dcache.c:479 dentry_kill+0x25d/0xc20 fs/dcache.c:826 finish_dput fs/dcache.c:1001 [inline] dput.part.0+0xce/0x230 fs/dcache.c:1042 dput+0x1f/0x30 fs/dcache.c:1037 end_dirop+0x7d/0xa0 fs/namei.c:2956 binderfs_binder_device_create.isra.0+0x71c/0x990 drivers/android/binderfs.c:194 binder_ctl_ioctl+0x186/0x1b0 drivers/android/binderfs.c:241 vfs_ioctl fs/ioctl.c:51 [inline] __do_sys_ioctl fs/ioctl.c:597 [inline] __se_sys_ioctl fs/ioctl.c:583 [inline] __x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f The buggy address belongs to the object at ffff88805b340c00 which belongs to the cache kmalloc-512 of size 512 The buggy address is located 0 bytes inside of freed 512-byte region [ffff88805b340c00, ffff88805b340e00) Fix by calling binder_add_device() before d_make_persistent(), while the parent directory lock is still held and the dentry cannot be discarded. Cc: stable <stable@kernel.org> Fixes: b89aa544821d ("convert binderfs") Signed-off-by: Peiyang He <peiyang_he@smail.nju.edu.cn> Assisted-by: Codex:gpt-5.5 Acked-by: Carlos Llamas <cmllamas@google.com> Link: https://patch.msgid.link/F6EF5FB778E87C98+20260913085645.1639558-1-peiyang_he@smail.nju.edu.cn Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2026-09-16binder: fix is_failure flag for superseded transaction cleanupTomer Pomeranc
When a TF_UPDATE_TXN transaction supersedes a pending async transaction, binder_release_entire_buffer() is called with is_failure=false. Since the superseded transaction was never delivered, binder_apply_fd_fixups() was never called and no fds were installed in the target process. With is_failure=false, the BINDER_TYPE_FDA cleanup handler interprets stale buffer contents as installed fd numbers and passes them to binder_deferred_fd_close(), closing unrelated file descriptors. Pass is_failure=true since the transaction was never delivered to the target, matching the semantics of all other undelivered-transaction cleanup paths. Fixes: 9864bb480133 ("Binder: add TF_UPDATE_TXN to replace outdated txn") Cc: stable <stable@kernel.org> Signed-off-by: Tomer Pomeranc <tomerpo@gmail.com> Acked-by: Carlos Llamas <cmllamas@google.com> Link: https://patch.msgid.link/20260812195316.259136-3-tomerpo@gmail.com Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2026-09-16binder: fix leaked fd fixups on TF_UPDATE_TXN supersedeTomer Pomeranc
When a TF_UPDATE_TXN transaction supersedes a pending async transaction in a frozen process, the outdated transaction is freed with kfree() directly. This skips binder_free_txn_fixups(), leaking all binder_txn_fd_fixup entries and their fget()'d struct file references. The leaked file refcounts never reach zero, so the struct file objects are permanently pinned in memory. They survive process exit and accumulate across invocations until file-max exhaustion. Every other transaction cleanup path (binder_free_transaction(), binder_transaction() error paths, binder_release_work()) correctly calls binder_free_txn_fixups(). Add the missing call before kfree() in the t_outdated cleanup block. Fixes: 9864bb480133 ("Binder: add TF_UPDATE_TXN to replace outdated txn") Cc: stable <stable@kernel.org> Signed-off-by: Tomer Pomeranc <tomerpo@gmail.com> Acked-by: Carlos Llamas <cmllamas@google.com> Reviewed-by: Alice Ryhl <aliceryhl@google.com> Link: https://patch.msgid.link/20260812195316.259136-2-tomerpo@gmail.com Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2026-09-16exec: Cleanup POSIX timers right after de_thread()Hyunwoo Kim
A per-thread CPU timer holds a reference to the PID of the thread it is attached to and, while it is armed, its node is queued in that thread's posix_cputimers. The task is looked up by that PID. When a non-leader thread exec()s, de_thread() changes which task owns that PID. pid_task(timer->it.cpu.pid, PIDTYPE_PID) then returns NULL, but the node is still queued on tsk, which is alive. timer_lock_sighand() takes a failed lookup to mean that the node is already dequeued, so it has nothing to undo. begin_new_exec() calls posix_cpu_timers_exit(me) right after exec_task_namespaces() and that removes the leftover node, so the state normally stays invisible. But bprm->point_of_no_return is set before de_thread(), so if unshare_files(), set_mm_exe_file(), exec_mmap() or exec_task_namespaces() fails, the task dies before it gets there. exit_itimers() then frees the k_itimer while its node is still queued, and reaping tsk later erases that freed node from the rbtree. In short: the non-leader thread B the parent timer_create(CLOCK_THREAD_CPUTIME_ID) timer_settime() arm_timer() // the node is queued on B execve() de_thread(B) exchange_tids(B, leader) // B's PID now belongs to the leader release_task(leader) __exit_signal(leader) posix_cpu_timers_exit(leader) // cleans leader's queue, not B's __unhash_process(leader) // that PID has no task anymore exec_mmap() mmap_read_lock_killable(old_mm) kill(B, SIGKILL) // -EINTR get_signal() do_exit() exit_itimers() posix_timer_delete() posix_cpu_timer_del() posix_timer_unhash_and_free() // freed while still queued wait4() release_task(B) posix_cpu_timers_exit(B) cleanup_timerqueue() timerqueue_del() // use-after-free Move the POSIX timer cleanup right after de_thread() before any of the later failure conditions brings the task into do_exit(). [ tglx: Move the cleanup right after de_thread() ] Fixes: 55e8c8eb2c7b ("posix-cpu-timers: Store a reference to a pid not a task") Signed-off-by: Hyunwoo Kim <imv4bel@gmail.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Tested-by: Kijo Park <red993688@gmail.com> Reviewed-by: Oleg Nesterov <oleg@redhat.com> Reviewed-by: Frederic Weisbecker <frederic@kernel.org> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/ao7Q8miiuLAPVnWv@v4bel Link: https://patch.msgid.link/20260911090541.627712075@kernel.org
2026-09-16signal: Prevent exec() raceThomas Gleixner
Hyunwoo debugged the following KASAN UAF splat: BUG: KASAN: slab-use-after-free in __send_signal_locked+0xb27/0xba0 Write of size 8 at addr ffff888007ed80c8 by task poc/79 ... Call Trace: __send_signal_locked+0xb27/0xba0 do_send_sig_info+0xa7/0x160 do_send_specific+0x76/0xa0 __x64_sys_tgkill+0x193/0x270 ... Allocated by task 80: do_timer_create+0x1a4/0x1030 __x64_sys_timer_create+0x145/0x190 ... Freed by task 12: kmem_cache_free_bulk+0x1f8/0x4a0 kvfree_rcu_bulk+0x14f/0x1c0 kfree_rcu_work+0x128/0x1a0 ... Last potentially related work creation: kvfree_call_rcu+0x39/0x390 __flush_itimer_signals+0x211/0x320 flush_itimer_signals+0x47/0x90 begin_new_exec+0xa6b/0x28c0 It turned out that this happens with a non-leader exec() as Hyunwoo explained: de_thread() calls exchange_tids() before release_task(leader), so the struct pid held by a SIGEV_THREAD_ID timer created against the leader's tid now points to the thread which called execve(). pid_task() returns that thread and lock_task_sighand() on it succeeds. If the timer signal is blocked, its sigqueue stays queued on the leader's task::pending. The next expiry of that timer can then run while release_task() flushes the queue. posixtimer_send_sigqueue() checks whether the sigqueue is already queued with a plain list_empty(), which only reads list_head::next. list_del_init() is not atomic and INIT_LIST_HEAD() stores list_head::next before list_head::prev, so the check can pass in between. list_add_tail() queues the entry on the task::pending of the live thread, and the list_head::prev store from the flush then overwrites the list_head::prev link that list_add_tail() has just set. __flush_itimer_signals() does not undo that either. With list_head::prev pointing at the entry itself, its list_del_init() only stores the same values again, so the entry is not removed from the list. It is still there after the last reference is dropped and the timer is freed by RCU, and the list_add_tail() of a later tgkill() follows that list_head::prev into the freed timer. This problem surfaced with the recent commit which moved the sigqueue flush out of the sighand lock held region. Hyonwoo proposed to fix this by using list_del_init_careful(), but that just papers over the problem. After some disucssions and various attempts to solve it, Eric pointed out that there is no reason to flush task::pending late in release_task() and it should be done in exit_signals() already. As nothing can collect and deliver signals which are queued in a dying task's pending queue, there is no reason to delay it further. But it has to be ensured that no signals can be queued into it after that point. exit_signals() sets PF_EXITING in task::flags, which can be used as an indicator for this. Cure it by: - Preventing signal queueing for task private signals (PIDTYPE_PID) when the task has PF_EXITING set in __send_signal_locked() and in posixtimer_send_sigqueue(). - Protecting the unlocked setting of PF_EXITING in exit_signals() for the task group empty and the group exit case with sighand lock - Flushing task::pending signals right there. Optimize that by moving the whole pending list to an on-stack list head under sighand lock and free the signals without the lock held. There has been quite some discussion about the lockless flush and the non-leader exec case on weakly ordered systems. The problem is that a third party which tries to send a posix timer signal relies on the PID lookup to find the target task and that lookup might result in the new leader when the signal was originaly directed to the old leader. In case that the signal was queued on the old leader then the lockless flush raised a concern over the following situation: old_leader new_leader third party A: flush_list() // list_del_init() stores to sigqueue LOCK (tasklist) old_leader->exit_state = EXIT_ZOMBIE; B: UNLOCK (tasklist) C: LOCK (tasklist) if (old_leader->exit_state) transfer_tids() D: store PID posix_timer_send_sigqueue() // Observes #D so t = new_leader E: t = get_target() F: LOCK (sighand) G: if (list_empty(sigqueue)) list_add(sigqueue) The concern was that the third party might observe #D but not observe #A and therefore would proceed to #G while the list_del() stores (#A) in flush_list() are not visible yet, which could result in list corruption. That would be possible if looking at it solely from a RELEASE+ACQUIRE ordering point of view, but B-C is a UNLOCK+LOCK hand-over, which is not the same as RELEASE+ACQUIRE: RELEASE+ACQUIRE: RCpc, only the CPUs involved agree on the ordering UNLOCK+LOCK: RCtso, the hand-over is store-ordering As B-C is UNLOCK+LOCK, which is RCtso and that does impose store order, A stores must happen before the D store. Combine with E-F, which has a data dependency from the LOAD to the LOCK and thereby constraints later LOADs, those sigqueue loads in G that come after F must in fact observe the A stores. Fixes: fb3bbcfe344e ("exit: change the release_task() paths to call flush_sigqueue() lockless") Reported-by: Hyunwoo Kim <imv4bel@gmail.com> Debugged-by: Hyunwoo Kim <imv4bel@gmail.com> Suggested-by: "Eric W. Biederman" <ebiederm@xmission.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Tested-by: Kijo Park <red993688@gmail.com> Reviewed-by: Oleg Nesterov <oleg@redhat.com> Reviewed-by: Frederic Weisbecker <frederic@kernel.org> Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260911090541.572536604@kernel.org Closes: https://patch.msgid.link/aok1rdkBgZsynHZB@v4bel
2026-09-16ata: libata-scsi: fix ata_dsm_trim_pages() kernel-docgraftedNiklas Cassel
make htmldocs fails with: Documentation/driver-api/libata:604: ./drivers/ata/libata-scsi.c:2301: ERROR: Unexpected indentation. [docutils] The Return: section of ata_dsm_trim_pages() ends its first sentence with a colon and continues with an indented bullet list. reStructuredText requires a blank line before an indented block, so docutils chokes on the list. Simply adding the missing blank line does not work either: Return: is a kernel-doc "special section", which is terminated by the first blank line, so the bullet list would end up in the Description section, detached from the sentence introducing it. Spell the two bounds out as prose instead, so that the Return: section stays self-contained. Documentation-only change. Fixes: e64e6b5dc867 ("ata: libata-scsi: scale DSM TRIM payload by MAX PAGES PER DSM COMMAND") Reported-by: Thomas Huth <thuth@redhat.com> Closes: https://lore.kernel.org/linux-ide/b0a0b8e8-cc4f-4c7b-8bc5-fee0712d405a@redhat.com/ Reviewed-by: Damien Le Moal <dlemoal@kernel.org> Reviewed-by: Thomas Huth <thuth@redhat.com> Link: https://lore.kernel.org/r/20260916130031.29990-2-cassel@kernel.org Signed-off-by: Niklas Cassel <cassel@kernel.org>
2026-09-16Revert "interconnect: qcom: x1e80100: enable QoS configuration"Marc Zyngier
This reverts commit 5a8b2cc36e796d595c9d97eab23c2be22805cc1c. Since the introduction of this QoS change, X1 machines based on the Purwa SoC (such as the Lenovo Mini X) take a hard reset at boot time. Other machines based on Hamoa (such as the X1E001DE Snapdragon Devkit) end-up with a similar hard reset while under load, most likely due to the PCIe ports being starved of traffic. Revert the whole thing until someone figures out what magic parameters allow QoS to work in a sensible manner, as a usable machine is somehow preferable to one that crashes efficiently. Link: https://lore.kernel.org/all/apgeBNSiP01YfNcs@google.com Link: https://lore.kernel.org/all/86ld9d4h4n.wl-maz@kernel.org Signed-off-by: Marc Zyngier <maz@kernel.org> Cc: Georgi Djakov <djakov@kernel.org> Cc: Raviteja Laggyshetty <raviteja.laggyshetty@oss.qualcomm.com> Cc: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com> Cc: Abel Vesa <abelvesa@kernel.org> Cc: Bjorn Andersson <andersson@kernel.org> Cc: Mostafa Saleh <smostafa@google.com> Cc: Thorsten Leemhuis <regressions@leemhuis.info> Reviewed-by: Abel Vesa <abel.vesa@oss.qualcomm.com> Reviewed-by: Mostafa Saleh <smostafa@google.com> Tested-by: Mostafa Saleh <smostafa@google.com> Closes: https://krzk.eu/#/builders/102/builds/70/steps/23/logs/warnings__3_ Link: https://patch.msgid.link/20260915123703.3228329-1-maz@kernel.org Signed-off-by: Georgi Djakov <djakov@kernel.org>
2026-09-16Revert "thunderbolt: Add quirk to reset host interface on DMA path teardown ↵Mario Limonciello
for AMD USB4 routers" This commit has caused a deadlock at shutdown. A proper fix with another approach will be coming later. Revert commit f1de1fc5f632cdeae1f5c2984572ab710d4dfcaa for now. Reported-by: juan.martinez@amd.com Closes: https://lore.kernel.org/linux-usb/20260825214237.4179813-1-juan.martinez@amd.com/ Cc: Sanath S <Sanath.S@amd.com> Cc: Basavaraj Natikar <Basavaraj.Natikar@amd.com> Signed-off-by: Mario Limonciello <mario.limonciello@amd.com> Signed-off-by: Mika Westerberg <mika.westerberg@linux.intel.com>
2026-09-15sched_ext: Wait for SCX_OPSS_DISPATCHING before reenqueueing a taskTejun Heo
ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics") moved the final ops_state store in scx_dispatch_enqueue() after the DSQ unlock so that the custody update and ops.dequeue() precede it. A task can thus be found on a DSQ while still SCX_OPSS_DISPATCHING. The dequeue and core-sched pick paths wait for the state to clear in ops_dequeue() but the reenqueue paths don't. A reenqueue in that window runs ops.enqueue() and sets SCX_OPSS_QUEUED before the dispatch has completed. The dispatcher's final store then overwrites it with SCX_OPSS_NONE and finish_dispatch() drops every later dispatch of the task. Wait for SCX_OPSS_DISPATCHING to clear before dequeueing a task for reenqueue, the same way ops_dequeue() does. Fixes: ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics") Cc: stable@vger.kernel.org # v7.1+ Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>