| Age | Commit message (Collapse) | Author |
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull timer fix from Ingo Molnar:
- Fix task work flags management regression in the hrtimer
rearming code that can leave task work items unprocessed
(Karl Mehltretter)
* tag 'timers-urgent-2026-10-04' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
hrtimer: Use the mask to clear TIF_HRTIMER_REARM from the exit work
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fix race between perf_event_exit_task() and perf_pending_task()
(Luo Gengkun)
- Fix perf header output management regressions (Ian Rogers)
- Require kernel access for text poke events (Zhengchuan Liang)
* tag 'perf-urgent-2026-10-04' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf: Require kernel access for text poke events
perf: Replace perf_event_header__init_id with full header init
perf: Fix race between perf_event_exit_task() and perf_pending_task()
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/tty
Pull tty/serial fixes from Greg KH:
"Here are some small tty/serial driver fixes for 7.3-rc6. Nothing major
here, just lots of small fixes for reported issues, some of them very
long-standing:
- tty hangup fixes that have been there since the BKL days and kept
tripping people up over time.
- vt selection bugfix
- other vt bugfixes (memory leaks and screen update fixes)
- n_gsm bugfix
- qcom-geni serial driver bugfix
- 8250 serial driver bugfixes
- other tiny serial driver fixes
All of these have been in linux-next, the last few only a few days but
testing here seems solid (this pull request was generated on that
tree)"
* tag 'tty-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/tty: (37 commits)
tty: add missing driver flag kernel-doc colon
vt: selection: Fix unsigned underflow and slab-out-of-bounds read in paste_selection()
vt: skip screen update for DEC alignment test on backgroup consoles
vc_screen: reload vc pointer before if (ret) in vcs_write() to avoid UAF
serial: sc16is7xx: reduce TX refill rate with half-FIFO trigger
serial: sc16is7xx: refill TX FIFO below trigger using fresh TXLVL
serial: tegra: don't clear the Tx FIFO on an Rx-only reset
serial: sc16is7xx: fix TX gap caused by kfifo circular buffer wrap-around
tty: fix saved termios reset race
tty: serial: mpc52xx_uart: move static declarations up.
tty: serial: max3100: shut down timer before freeing port
tty: add break_wait kernel-doc
serial: qcom-geni: keep registered console runtime active
serial: qcom-geni: Fix unbalanced runtime PM resume for no_console_suspend
serial: qcom-geni: avoid unused-function warning
tty: serial: qcom_geni_serial: Keep console RX functional after deep idle
soc: qcom: geni-se: Correct QUP Core ICC vote constants
serial: 8250_bcm7271: fix use-after-free in brcmuart_remove()
serial: vt8500: Fix clock reference leak in vt8500_serial_probe()
kgdboc: Fix tty driver reference leak in configure_kgdboc()
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/usb
Pull USB/Thunderbolt fixes from Greg KH:
"Here is a big set of USB and Thunderbolt driver fixes for 7.3-rc6.
They were delayed on my side due to conference travel, not the fault
of the submitters at all. Included in here are:
- lots of small thunderbolt fixes for reported issues due to more
testing and devices and a few reverts as well based on that work
- more usb-serial device ids added
- usb-serial and cdc-acm driver hangup and other fixes
- dwc3 driver fixes for reported problems
- lots of usb gadget driver fixes as people again fuzz these drivers
and send in fixes, which is nice to finally see
- typec driver fixes for reported problems
- octeon-hcd driver fixes for reported problems
- more usb-storage quirks added
- other small USB driver bugs resolved for reported problems
All of these have been in linux-next without any reported issues"
* tag 'usb-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/usb: (63 commits)
usb: dwc3: gadget: fix IRQ storm on invalid event buffer count
Revert "usb: dwc3: gadget: fix IRQ storm on invalid event buffer count"
USB: gadget: dummy-hcd: Fix wait for outstanding request completions
usb: typec: port-mapper: Only match USB4 port if host interface is available
usb: cdns3: Fix NULL pointer dereference in cdns3_pci_probe
usb: dwc3: gadget: fix IRQ storm on invalid event buffer count
usb: typec: ucsi: Get the connector fwnode based on reg value
usb: gadget: f_uac1_legacy: validate bRequest index in generic_{set,get}_cmd
usb: core: clear both ep_in and ep_out for non-ep0 control endpoints
USB: cdc-acm: skip URB restart in port_shutdown if disconnected
usb: gadget: f_fs: Fix NULL pointer dereference in FUNCTIONFS_ENDPOINT_DESC
usb: gadget: aspeed-vhub: cancel wake work on device removal
thunderbolt: Disable CL states for the Anker Prime TB5 dock
thunderbolt: stream: Announce support for FMODE_NOWAIT
usb: typec: ucsi: displayport: Current CAM OOB index fixup
usb: ohci-st: disable controller wakeup on removal
usb: ohci-spear: disable controller wakeup on removal
usb: ohci-s3c2410: disable controller wakeup on removal
usb: ohci-da8xx: disable controller wakeup on cleanup
usb: cdns3: fix use-after-free in cdns3_gadget_exit()
...
|
|
Pull smb client fixes from Paulo Alcantara:
"Fix a series of data corruption and I/O error bugs found by running
generic/363 (fsx) in a loop against Windows Server 2022 and Samba.
- Stop data dirtied past EOF through an mmap from reappearing as file
content once the file is extended by a write, truncate, zero range,
copy range or clone range
- Flush dirty data and drain in-flight I/O before operations that
assume the pagecache and the server agree on the file: querying
allocated ranges, the O_TRUNC open, interior zero range, and
server-side copy/clone
- Stop a genuine size-extending zero range or preallocate from being
refused with -EOPNOTSUPP when the inode is not read caching, by
querying the server's authoritative EOF instead of trusting a stale
cached i_size
- Zero the untransferred tail of a short read, both in the netfs
read-gaps path (where stale folio content could otherwise be
written back to the server) and in the DIO/unbuffered read
collector, and tell a real EOF apart from a stale cached
remote_i_size after a lease downgrade
- Require stable pages on signed connections so a buffered write
can't modify a folio whose signature has already been computed and
is in flight, which the server rejected with STATUS_ACCESS_DENIED
and the client surfaced as -EIO
- Split several cifsFileInfo flags out of a shared bitfield byte so
concurrent updates taken under different locks no longer clobber
each other through a byte-level RMW"
* tag 'cifs-fixes-7.3-rc6' of https://git.manguebit.org/linux:
smb: client: split cifsFileInfo bitfields to avoid shared-byte RMW races
smb: client: require stable pages for signed connections
smb: client: distinguish real EOF from a stale remote_i_size on read
netfs: zero the tail of a short DIO/unbuffered read
smb: client: only require read lease for size-extending preallocate
netfs: zero gaps in read-gaps folio to avoid writing back stale data
smb: client: only require read lease for size-extending zero range
smb: client: drain and invalidate before server-side copy/clone
smb: client: flush dirty data before zeroing a range
smb: client: drain outstanding I/O before truncating on O_TRUNC open
smb: client: flush and commit data before querying allocated ranges
smb: client: discard post-EOF pagecache when extending a file via clone range
smb: client: discard post-EOF pagecache when extending a file via copy range
smb: client: discard post-EOF pagecache when extending a file via zero range
smb: client: clear post-EOF pagecache when extending a file via truncate
netfs: clear post-EOF pagecache when extending a file via write
|
|
Pull bpf fixes from Alexei Starovoitov:
- Fix overflow of backward jump offset in constant blinding
(Alexei Starovoitov)
- Fix packet range of packet pointers sharing an id when var_off
tightens umax of one pointer and not the other (Alexei Starovoitov)
- Fix objects stuck in free_by_rcu_ttrace list of bpf memalloc
(Alexei Starovoitov)
- Fix use-after-free of progs detached from busy trampolines: wait for
an RCU tasks grace period before freeing trampoline progs, and patch
detached progs out of trampoline images that are still in use
(Florent Revest)
- Hold map BTF for the memory allocator destructor record to fix UAF in
deferred bpf_mem_alloc destruction (Kumar Kartikeya Dwivedi)
- Fix missing migration protection in resizable hashtab
lookup_and_delete batch operation (Ömer Mete Kaya)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf:
bpf: Fix missing migration protection in __rhtab_map_lookup_and_delete_batch()
selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace
bpf: Fix objects stuck in free_by_rcu_ttrace
bpf: Factor out __do_call_rcu_ttrace()
selftests/bpf: Test packet range of pointers sharing an id
bpf: Fix packet range of pointers sharing an id
selftests/bpf: Detach a trampoline prog while a task sleeps before it
bpf: Skip detached progs in trampoline images that are still in use
bpf: Wait for an RCU tasks grace period before freeing trampoline progs
bpf: Hold map BTF for the memory allocator destructor record
bpf: Fix overflow of jump offset in constant blinding
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux
Pull io_uring fixes from Jens Axboe:
- Fix a task_work add use-after-free with SQPOLL.
The sqpoll thread could pop and complete the last request while
io_req_normal_work_add() was still looking at them after the mpscq
push.
Use the same approach as DEFER_TASKRUN to protect from that, holding
an RCU read lock across the add, and have exit wait for an RCU grace
period for SQPOLL rings as well.
- CQE32 ring fixes: correct the free entry check for 32b CQEs, zero the
big_cqe for aux CQEs, and only post the dummy skip CQE on CQE_MIXED
rings
- Mark the source filter table as COW when cloning bpf filters, so
registering another filter on the source doesn't modify the shared
table in place
- Initialize the task context before running the BPF loop
- Requeue zcrx multishot receives stopped by a local resource
- End a TX_TIMESTAMP multishot cmd when the CQ is full (lollipopkit)
* tag 'io_uring-7.3-20261002' of git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux:
io_uring: fix task_work add use-after-free with SQPOLL
io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full
io_uring/zcrx: requeue multishot receives stopped by a local resource
io_uring: initialize task context before running the BPF loop
io_uring: zero big_cqe for aux CQEs on CQE32 rings
io_uring: fix free entry check for 32b CQEs on CQE32 rings
io_uring: only post the dummy skip CQE on CQE_MIXED rings
io_uring/bpf_filter: mark source as COW when cloning filters
|
|
A recent change adding a new TTY driver flag left out a colon required
for well-formed kernel-doc.
Fixes: 8df07fe93573 ("tty: fix saved termios reset race")
Reported-by: Randy Dunlap <rdunlap@infradead.org>
Link: https://lore.kernel.org/ec3a09d6-4ba4-41b8-abb1-b59e265f4532@infradead.org
Signed-off-by: Johan Hovold <johan@kernel.org>
Link: https://patch.msgid.link/20261002072736.2063004-1-johan@kernel.org
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Add tests where two packet pointers share an id and tightening one
pointer's umax from its var_off would put it less than their constant
distance from the other's umax: with an index & 0x38 capped at 50, the
base pointer keeps umax 50, so the pointer 8 bytes further on must keep
umax 58, even though its known bits allow at most 56.
These refused a valid program or accepted an out-of-bounds access before
the fix:
- check the advanced copy, load through the base: valid, was refused;
- check the base, load the byte at base + 1 through a copy advanced by
8: was accepted;
- check base + 4, load 4 bytes at base + 2 through base + 8: reads two
bytes past the checked range, was accepted;
- the same as the second with data_meta pointers checked against data:
was accepted.
These pass with and without the fix and cover nearby paths:
- subtract an unknown scalar from a checked pointer and load below it
(the range is kept across a new id);
- reach a load through two paths whose checks cover 8 and 7 bytes after
the loaded pointer; the second path must not be pruned by the first;
- spill a copy of a pointer, check the pointer, fill the copy and load
one byte past the checked range: the load is refused, and the copy
has the range of the check.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20261001145255.855630-2-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
perf_iterate_sb() invokes its callback for each matching perf_event on
the CPU and task context, passing a shared caller-allocated event
structure.
perf_event_header__init_id() mutated header->size in place by adding
event->id_header_size, requiring sideband output callbacks to save and
restore header fields across iterations. Three sideband callbacks failed
to save and restore header.size around perf_event_header__init_id():
- perf_event_ksymbol_output()
- perf_event_bpf_output()
- perf_event_text_poke_output()
When multiple events with attr.ksymbol, attr.bpf_event, or
attr.text_poke and sample_id_all are active on the same CPU, each
subsequent event receives a record whose header.size is inflated by all
preceding events' id_header_size values while only a single id_sample is
written, leaving uninitialized ring-buffer bytes at the end of the
record and causing userspace perf to fail with -EFAULT ("Bad address")
when parsing the sample_id trailer.
Similarly, perf_event_mmap_output() set PERF_RECORD_MISC_MMAP_BUILD_ID
in mmap_event->event_id.header.misc when event->attr.build_id was
enabled, but only saved and restored header.size and header.type. If an
event with attr.build_id was followed by an event with attr.mmap2 and
!attr.build_id, the second event received PERF_RECORD_MISC_MMAP_BUILD_ID
in header.misc while its payload contained maj/min/ino/ino_generation
instead of a build ID.
Rather than splitting header initialization between callers and output
callbacks and saving/restoring mutated header fields, replace
perf_event_header__init_id() with perf_event_header__init(), which
initializes header->type, header->misc, and header->size alongside the
sample_id fields on each invocation.
Fixes: 76193a94522f ("perf, bpf: Introduce PERF_RECORD_KSYMBOL")
Fixes: 6ee52e2a3fe4 ("perf, bpf: Introduce PERF_RECORD_BPF_EVENT")
Fixes: e17d43b93e54 ("perf: Add perf text poke event")
Fixes: 88a16a130933 ("perf: Add build id data in mmap2 event")
Assisted-by: Antigravity:gemini-3.1-pro
Signed-off-by: Ian Rogers <irogers@google.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://patch.msgid.link/20260929222332.973435-1-irogers@google.com
Cc: stable@vger.kernel.org
|
|
Resetting saved termios state on device registration is needed where a
minor number can be reused for an entirely different device and where
the old settings may prevent the port from even being opened (e.g. when
CLOCAL is not set).
Not all TTY drivers guarantee that the minor number is no longer in use
when registering devices however, something which can lead to a
use-after-free when closing a TTY (and saving its termios) races with
re-registration.
Add a new TTY_DRIVER_RESET_SAVED_TERMIOS flag to request that any saved
termios state is reset on registration and only set it for drivers that
make sure that the minor number is no longer in use.
Fixes: 93857edd9829 ("tty: reset termios state on device registration")
Reported-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Link: https://lore.kernel.org/20260926184154.3017929-1-nicoyip.dev@gmail.com
Cc: stable@kernel.org # 4.12
Signed-off-by: Johan Hovold <johan@kernel.org>
Link: https://patch.msgid.link/20260930131845.1809256-1-johan@kernel.org
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Avoid a compilation error so that they're not used before being
declared.
Fixes: 4d105880666a ("tty: serial: mpc52xx_uart: add bounds check for psc_num array index")
Reported-by: kernel test robot <lkp@intel.com>
Closes: https://lore.kernel.org/oe-kbuild-all/202609260356.tJWbd1WU-lkp@intel.com/
Signed-off-by: Rosen Penev <rosenp@gmail.com>
Link: https://patch.msgid.link/20260927203501.14105-1-rosenp@gmail.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
generic_set_cmd() and generic_get_cmd() use the low nibble of
ctrl->bRequest as an index into con->data[]:
u8 cmd = (ctrl->bRequest & 0x0F); /* 0 .. 15 */
...
con->data[cmd] = value; /* OOB when cmd >= 5 */
struct usb_audio_control (include/linux/usb/audio.h) declares data as
a 5-element array, so indices 5 through 15 write (or read) up to 44
bytes past the end of the array on the heap.
A malicious USB host can craft a class-specific SET_CUR / GET_CUR
request with an arbitrary bRequest value, triggering the out-of-bounds
access from an IRQ completion handler with no further preconditions.
Add an ARRAY_SIZE() guard to both functions.
Fixes: c47d7b09891a ("USB: audio: add USB audio class definitions")
Cc: stable <stable@kernel.org>
Reviewed-by: Weibin Liu <liuwb@xiaopeng.com>
Signed-off-by: Liu Chao <liuc63@xiaopeng.com>
Reviewed-by: Ivy Lopez <skunkolee@gmail.com>
Link: https://patch.msgid.link/20260921072551.3708191-1-liuc63@xiaopeng.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
The buffered read collector zero-fills the tail of a short read that
stops below the inode's i_size (netfs_clear_unread()), so a read that
races an extending write still returns the expected number of bytes.
The non-buffered collector path does no such thing: it just records
how much was transferred.
Add the same zero-fill for the non-buffered case, gated on
NETFS_SREQ_CLEAR_TAIL: a subreq's source sets that flag when a short
result from it is known to be safe to treat as a hole, as opposed to
NETFS_SREQ_HIT_EOF, which means the read genuinely ran off the end of
the file and should be reported short as-is. Only CLEAR_TAIL should
zero-fill here; a real EOF must stay a real short read.
No source currently sets CLEAR_TAIL on an unbuffered/DIO subrequest,
so this is inert on its own -- a following change teaches cifs to set
it in the one case that needs it.
Fixes: e2d46f2ec332 ("netfs: Change the read result collector to only use one work item")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
|
|
hrtimer_rearm_deferred_user_irq() removes the rearm bit from the local
copy of the TIF work with:
*tif_work &= ~TIF_HRTIMER_REARM;
TIF_HRTIMER_REARM is the bit number (12), not the mask. That clears
TIF_NOTIFY_SIGNAL (bit 2) and TIF_MEMDIE (bit 3) from the copy and
leaves bit 12 set.
So the function never returns true, and the copy handed to
exit_to_user_mode_loop() lacks TIF_NOTIFY_SIGNAL. If that was the only
work for the loop, the task returns to user space with the task work
still pending. hrtimer_interrupt() sets TIF_HRTIMER_REARM every time,
so the next tick repeats that. Task work queued with TWA_SIGNAL from a
hrtimer callback stays pending until the task does a syscall, takes an
interrupt without deferred rearm or has to reschedule.
A task which polls the io_uring completion ring in user space sees a
timeout completion after 30ms (median) instead of 70us. 2 CPU QEMU
guest, HZ=1000.
Use the mask.
Fixes: 15dd3a948855 ("hrtimer: Push reprogramming timers into the interrupt return path")
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Assisted-by: LLM
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260919144027.70250-1-kmehltretter@gmail.com
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull scheduler fixes from Ingo Molnar:
- Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang)
- Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen)
- Skip kernel threads for cache aware scheduling to rubustify the code
(Chen Yu)
- Refresh LLC capacity across CPU hotplug, to fix capacity
underestimation bug (Davi Chaves Azevedo)
- Account PSI IRQ time to the execution context, not the scheduling
context, to fix proxy scheduling accounting bug (Zhan Xusheng)
* tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
sched/core: Account PSI IRQ time to the execution context, not the scheduling context
sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug
sched/cache: Skip kernel threads for cache aware scheduling to rubustify the code
sched/cache: Introduce task_struct->sched_cache_grp to fix UAF
sched/cache: Decouple sched_cache_group from mm to fix UAF
sched/cache: Honor migrate_llc_task semantics in active load balance, to fix LLC mis-scheduling bug
sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC mis-scheduling bug
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fixes for KVM guest PEBS virtualization (Sean Christopherson)
- Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi)
- Fix Intel Panther Cove event scheduling constraints (Dapeng Mi)
- Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi)
- Rename two confusingly named PMU attributes (Dapeng Mi)
- Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim)
- Fix NULL pointer dereference crash in __perf_pmu_sched_task()
(Puranjay Mohan)
- Fix CPU-wide event scheduling (Puranjay Mohan)
- Fix x86 LBR branch entry generation (Puranjay Mohan)
* tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf/core: Fill branch entries with a single assignment
perf/core: Run sched_task() for PMUs with only CPU-wide events
perf/core: Fix NULL pmu_ctx passed to pmu->sched_task()
perf/core: Fix a refcount leak in attach_perf_ctx_data()
perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL
perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3
perf/x86/intel: Delete dead NVL PEBS data-source initcall
perf/x86/intel: Fix Panther Cove PEBS data-source snoop states
perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints
perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints
perf/x86/intel: Update arw_latency_data() mem-op direction handling
perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs()
perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused
perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED
perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux
Pull ata fixes from Niklas Cassel:
- Extend the quirk "no LPM on ATI" quirk, that is currently only
applied for Samsung drives, to include AMD controllers as well.
The AMD AHCI controllers are newer versions of the ATI AHCI
controllers, and these controllers still have LPM issues with
Samsung drives - LPM works with drives from other vendors (me)
- Fix errors in the libata.force parameter documentation (me)
- Verify the sense data descriptor lengths for ATA PASS-THROUGH
command, so that a malicious device cannot write past the buffer
length (Matthias)
- Mention the libata for-next branch in MAINTAINERS such that the
git ls-remote command done by get_maintainer.pl --self-test=scm
can verify it (Matthias)
* tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux:
MAINTAINERS: name the libata/linux for-next branch
ata: libata-scsi: bound the ATA passthru sense descriptor writes
ata: libata: Correct libata.force parameter documentation
ata: libata-core: Extend Samsung LPM quirk to AMD controllers
|
|
Pull kvm fixes from Paolo Bonzini:
"Arm:
- Invalidate the ITS translation cache when the guest changes the
base address of the ITS tables (Fuad Tabba)
- Skip saving ITS devices with device IDs that are out-of-bounds
rather than failing the entire ITS save ioctl (Fuad Tabba)
- Close race between VM teardown and invalidations of nested MMUs
when handling MMU operations that are allowed to block (Lorenzo
Stoakes)
- Various fixes for the handling of the host's untrusted SVE
configuration in pKVM (Fuad Tabba)
- Make sure that empty SMCCC ranges based at 0 are rejected by the
kvm_smccc_set_filter() (Karl Mehltretter)
- Revoke the host mapping for pKVM's private stack pages, along with
a new sanity check that all mappings in the hyp's private VA range
have been correctly marked as hyp-owned (Fuad Tabba)
- Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
concurrent vCPU initialization cannot relocate in-use MMUs. Defer
the freeing of shadow stage-2 MMUs to the point that no other users
(e.g. MMU notifier) could reference them (Marc Zyngier)
- Drop useless WARN when rejecting an unsupported ioctl for pKVM
(Fuad Tabba)
- Fix the steal_time selftest to install correctly-sized mappings for
non-4K hosts (Sebastian Ott)
- Correct mapping of fine-grained trap for GCSPOPX instruction (Mark
Brown)
- Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
guests (Karl Mehltretter)
RISC-V:
- Synchronize hrtimer during VCPU teardown
- Fix the conversion between vsip and hvip values
- Serialize IMSIC attributes with vCPU migration
- Release unused page after MMU invalidation
- Propagate interrupted G-stage faults to KVM user-space as EINTR
- Fix nested acceleration hfence entry update order
- Fix sdata leak and stale snapshot_addr in snapshot_set_shmem
- Preserve firmware counter value across PMU counter stop/start
- Report PMU snapshot write failure to the guest
- Fix perf-backed counter accounting across PMU stop and read
- Correctly propagate error of a hart status SBI call
s390:
- Ensure that accesses through kvm_arch_set_irq_inatomic mark as
dirty the pages that contain indicator and summary bits
- Fix compile warning for kvm_s390_update_cmma_dirty()
- Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
userspace
- Move s390_kvm_mmu_commit_memory_region() into
s390_kvm_mmu_prepare_memory_region() so that it can fail instead of
WARN
- Add missing srcu in kvm_s390_set_irq_state()
- Fix potential races in storage functions
- Fix race in _destroy_pages_crste()
- Fix issues in the handling of KVM interrupt and page resources,
when a queue that is assigned to a mediated device (mdev) is
removed from the host's AP configuration
- Fix loop condition in uv_find_secrets
- Prevent potential out-of-bounds read
x86:
- Fix a brown paper bag bug where KVM would incorrectly treat Intel
PMU MSRs as valid on AMD
- Fix a regression in the hardware disable selftest where it checked
the wrong macro when detecting glibc support (breaks at least musl)
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is
especially important for KVM_BUG_ON() flows, which often guard more
dangerous bugs
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a
bug where KVM would let userspace run a broken setup with stale
vmcs12 pages
- Fix a class of bugs where KVM would fail to fill kvm_run exit
fields if getting nested pages failed
- Treat reserved entries in the memory attributes xarray as "no
attributes", to fix false positives when checking for mixed
attributes
- Fix memcg accounting for the memory attributes xarray (the xarray
library subtly requires the xarray to be configured for accounting
upfront; the gfp flags taken at runtime are used only rarely)
- Don't pre-reserve xarray entries when storing empty attributes, as
storing NULL must not require memory allocation (KVM and other
subsystems heavily rely on this behavior)
- Fix a memory leak and a cache maintenance issue related to doing
intra-host migration on an SEV guest"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits)
KVM: SEV: Do cache maintenance on the source VM during intra-host migration
KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes
KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg
KVM: Don't treat reserved xarray entries as having memory attributes
KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths
KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails
KVM: arm64: Fix AArch32 DBGBXVR<n> handling
KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP
KVM: selftests: fix steal_time for arm64 with host page size > 4K
KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode
KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction
KVM: arm64: nv: Fix life cycle of the nested_mmus array
KVM: arm64: Check every private mapping is hyp-owned at pKVM init
KVM: arm64: Move the private VA allocation cursor to __io_map_next
KVM: arm64: Match hyp text by physical address in fix_host_ownership()
KVM: arm64: Transfer the hyp stack pages out of the host stage-2
KVM: arm64: selftests: Test empty SMCCC filter range at base 0
KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0
KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2
...
|
|
KVM fixes for 7.3-rcN
- Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
as valid on AMD.
- Fix a regression in the hardware disable selftest where it checked the wrong
macro when detecting glibc support (breaks at least musl).
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
important for KVM_BUG_ON() flows, which often guard more dangerous bugs.
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
where KVM would let userspace run a broken setup with stale vmcs12 pages.
- Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
getting nested pages failed.
- Treat reserved entries in the memory attributes xarray as "no attributes",
to fix false positives when checking for mixed attributes.
- Fix memcg accounting for the memory attributes xarray (the xarray library
subtly requires the xarray to be configured for accounting upfront; the gfp
flags taken at runtime are used only rarely).
- Don't pre-reserve xarray entries when storing empty attributes, as storing
NULL must not require memory allocation (KVM and other subsystems heavily
rely on this behavior).
|
|
Manually perform cache maintenance on the source VM during intra-host
migration to ensure no stale data is left in CPU caches after the VM is
destroyed. Because the source VM is "converted" to a non-SEV VM, KVM's
memory reclaim flows won't trigger cache maintenance, e.g. when all guest
memory is reclaimed in response to detaching from the mmu_notifier.
Note, relying on the destination VM to do cache maintenance isn't an option
as KVM doesn't require identical guest memory configurations, i.e. the
source VM may have access to memory that the destination VM does not.
Enforcing equivalent memory configurations is infeasible, as it would
require a *deep* comparison of memslots, e.g. to verify that not only are
the memslot identical, but what the memslots point at is also identical.
Fixes: b56639318bb2 ("KVM: SEV: Add support for SEV intra host migration")
Cc: stable@vger.kernel.org
Reported-by: Stefan Teodorescu <fane@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260923163721.1584779-3-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/wufan/ipe
Pull IPE fixes from Fan Wu:
"Two fixes for use-after-free issues found by recent LLM-assisted code
analysis.
- move successful policy load auditing under the new policy
directory's inode lock, preventing a concurrent policy deletion
from freeing the policy while it is still being audited
- protect the dm-verity root hash with RCU, preventing policy
evaluation from racing with root hash replacement during preresume"
* tag 'ipe-pr-20260925' of git://git.kernel.org/pub/scm/linux/kernel/git/wufan/ipe:
ipe: protect the dm-verity root hash with RCU
ipe: fix use-after-free when auditing a newly loaded policy
|
|
perf_clear_branch_entry_bitfields() clears the bitfields of struct
perf_branch_entry one by one and leaves from/to alone, since callers
overwrite those straight away. The list has to be kept in sync with the
struct by hand and has already fallen behind: new_type and priv were
added to perf_branch_entry and never added here.
Only BRBE writes those two, and neither for every record.
brbe_set_perf_entry_type() leaves new_type alone for a branch type it
does not recognise, and priv is not set for source-only records.
arm_pmuv3.c allocates the per-CPU branch stack with kmalloc(), so such a
record reaches userspace with whatever the slot held: uninitialised
kmalloc() data on the first pass over the buffer, the previous record's
values after that. Nothing under arch/x86/events/ writes either field,
so only arm64 is affected.
Assign the whole entry at each site instead. Everything not named is
then zero, and there is no list to keep in sync. The bitfields add up to
exactly 64 bits, so the struct has no padding to leave undefined.
perf_clear_branch_entry_bitfields() has no callers left, so remove it.
perf_entry_from_brbe_regset() assigns an empty literal instead, since it
fills from/to conditionally. PERF_BR_SPEC_NA is 0, so dropping the
explicit spec assignment changes nothing.
Fixes: b190bc4ac9e6 ("perf: Extend branch type classification")
Fixes: 5402d25aa571 ("perf: Capture branch privilege information")
Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Tested-by: Yifan Wu <wuyifan50@huawei.com>
Link: https://patch.msgid.link/20260810133540.1947118-4-puranjay@kernel.org
|
|
Skip the xarray reservation loop when clearing all memory attributes, as
storing NULL only erases the entry and never needs to allocate, so no
reservation (and no cleanup of a failed one) is required in that case.
Suggested-by: Sean Christopherson <seanjc@google.com>
Cc: David Ballesteros <davimaba.v@proton.me>
Signed-off-by: Zeng Chi <zengchi@kylinos.cn>
Link: https://patch.msgid.link/20260921102442.1232375-1-zeng_chi911@163.com
[sean: split to separate patch]
Signed-off-by: Sean Christopherson <seanjc@google.com>
|
|
DMR introduces Offmodule Response events in place of the legacy
Offcore Response events, but it still exposes the inherited
offcore_rsp PMU attribute for programming the corresponding MSR data.
Rename the DMR PMU attribute to offmodule_rsp so the sysfs interface
matches the underlying event name and avoids user & tooling confusion.
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Link: https://patch.msgid.link/20260917015234.981153-12-dapeng1.mi@linux.intel.com
|
|
underestimation bug
The scheduler scales LLC capacity by the fraction of cache-sharing CPUs
covered by a domain:
llc_bytes = cache_size * span_weight / shared_weight
During CPU teardown, sched_cpu_deactivate() rebuilds scheduler domains
before cacheinfo_cpu_pre_down() removes the CPU from shared_cpu_map. The
new domains therefore use the old sharing weight. The later call to
sched_update_llc_bytes() looks up the departing CPU's sd_llc, which has
already been detached, and returns without correcting the surviving CPUs.
On a Ryzen 5 7535U with twelve logical CPUs sharing a 16 MiB LLC,
offlining one SMT sibling left the remaining CPUs with:
llc_bytes = floor(16777216 * 11 / 12) = 15379114 bytes
The correct capacity is still 16777216 bytes. On systems with active
cache-aware scheduling, an underestimated capacity can cause
exceed_llc_capacity() to reject aggregation for a process whose footprint
would fit. Unchanged cpuset partitions sharing the physical cache can
also retain stale capacity when a CPU comes online in another partition.
Pass the cache-sharing mask already retained by cacheinfo to the
scheduler update. Refresh every surviving CPU using its own LLC domain
so that each partition receives the correct share. This also preserves
the correction needed as cache-sharing maps grow during boot.
Keep the existing CPU-hotplug and scheduler-domain synchronization. The
update remains on the hotplug path; no steady-state scheduling operation
or persistent allocation is added.
Fixes: 7030513a0877 ("sched/cache: Calculate the LLC size and store it in sched_domain")
Signed-off-by: Davi Chaves Azevedo <davichazbh@gmail.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Reviewed-by: Chen Yu <yu.c.chen@intel.com>
Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com>
Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com>
Tested-by: Chen Yu <yu.c.chen@intel.com>
Tested-by: K Prateek Nayak <kprateek.nayak@amd.com>
Cc: <stable@kernel.org> # v7.2.x
Link: https://patch.msgid.link/6751d93e15889e624796c74db0bfe66603d60b1b.1790035273.git.tim.c.chen@linux.intel.com
|
|
Add a sched_cache_grp pointer to task_struct so that scheduler code
can access the cache group directly via the task, without going
through mm->sched_cache_grp. This decouples the scheduler's hot-path
accesses from the mm_struct.
Each task holds its own refcount on the sched_cache_group, separate
from the reference held by its mm_struct. The reference is acquired
in copy_mm() (fork) and exec_mmap() (exec), and released in exit_mm().
This fixes the use-after-free when account_mm_sched() reaches the group
through a task whose mm is being switched, as reported by Hyunwoo:
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Convert all scheduler code in fair.c and exit.c to use
p->sched_cache_grp instead of p->mm->sched_cache_grp.
Keep the fork/exec/exit reference management out of the generic mm
paths: add sched_cache_fork(), sched_cache_fork_cleanup(),
sched_cache_exec_mmap() and sched_cache_exit_mm() in
kernel/sched/cache_sched.c (with empty stubs for !CONFIG_SCHED_CACHE),
so fs/exec.c, kernel/fork.c and kernel/exit.c each call one helper
instead of open-coding the refcounting under #ifdef. Also add
sched_cache_group_get() and task_cache_group_get().
Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/
Reported-by: Hyunwoo Kim <imv4bel@gmail.com>
Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev>
Co-developed-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Cc: <stable@kernel.org> #7.2.x
Link: https://patch.msgid.link/ae7081dc54736bf115215f9867abb2711a7403fb.1790035273.git.tim.c.chen@linux.intel.com
|
|
Currently the sched cache grouping is by mm and the scheduling statistics
sched_cache_stat lives in the mm structure. This ties the life cycle
of scheduling stats with mm.
In account_mm_sched(), the scheduling stats are accessed by
task->mm->sc_stat. However, a task may be switching mm on one CPU when
another CPU is running account_mm_sched(), and possibly accessing the
old mm that was freed. This problem was found when running tests with
KASAN by Hyunwoo:
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Instead of serializing the mm access by introducing extra acquisition of
rq lock in the mm free path, extract sched_cache_stat from mm_struct,
rename it as sched_cache_group and manage its life cycle apart from
mm_struct with its own ref counting. This allows us in the next patch
access sched_cache_group directly from task, and add a refcount
on sched_cache_group when a task links to it. This prevents the use
after free issue when accessing stale and released old mm and its
sched cache stat a task switches to a new mm while account_mm_sched()
is done elsewhere.
The other benefit of this restructure is in the future, the grouping of
tasks to a LLC would have the flexibility to be associated with a user
defined grouping, or cgroup, cookie group, numa_group or others instead
of just with a single mm address space.
Rename sched_cache_stat to sched_cache_group and turn it into a refcounted
object allocated from mm_struct. The mm_struct now holds a pointer
(sched_cache_grp) to this object instead of embedding it.
Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/
Reported-by: Hyunwoo Kim <imv4bel@gmail.com>
Reported-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev>
Co-developed-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Cc: <stable@kernel.org> #7.2.x
Link: https://patch.msgid.link/91fd1e3266707c865bc9abecfb3e17bc676712df.1790035273.git.tim.c.chen@linux.intel.com
|
|
mis-scheduling bug
alb_break_llc() decides whether to break LLC preference during active
load balance. It does so by testing that every runnable fair task on the
source rq prefers its LLC:
env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_runnable
But the two counters cover different sets. nr_pref_llc_running is updated
in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued,
so it follows queued tasks. h_nr_runnable is updated in set_delayed()/
clear_delayed() and drops delay-dequeued tasks.
So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted
in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks,
alb_break_llc() returns false, and active balance is free to pull a task
off its preferred LLC. Active balance only moves runnable tasks, and this
is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB
skips the per-task test in can_migrate_task(). The runnable set is the one
we want.
Fix it on the counter side. A task should be counted in
nr_pref_llc_running exactly while it is both queued on its preferred LLC
(pref_llc_queued) and runnable (!sched_delayed). Define that membership
once in task_pref_llc_runnable(), and adjust the counter only through
pref_llc_running_inc()/pref_llc_running_dec() from the four sites that
change either input: account_llc_enqueue(), account_llc_dequeue(),
set_delayed() and clear_delayed(). Gating every update on the same
predicate keeps the delay, wake and dequeue paths from double-counting
or underflowing; see the comments at those sites for the ordering.
nr_llc_running and sd->llc_counts are not touched and stay on queued
semantics.
Fixes: 714059f79ff0 ("sched/cache: Handle moving single tasks to/from their preferred LLC")
Closes: https://lore.kernel.org/lkml/20260827135000.735138-1-zhanxusheng@xiaomi.com/
Reported-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Suggested-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Reviewed-by: Kayra Cizmeci <kayracizmeci@gmail.com>
Cc: <stable@kernel.org> # v7.2.x
Link: https://patch.msgid.link/06af61afedac32e6477f57feb4d658f6c411c3af.1790035273.git.tim.c.chen@linux.intel.com
|
|
A Samsung SSD 870 QVO 8TB connected to an AMD 600 Series chipset SATA
controller is reported to time out on STANDBY IMMEDIATE during system
suspend with med_power_with_dipm enabled. The command completes when
using max_performance instead.
The existing Samsung LPM quirk only matches ATI controllers, leaving
AMD controllers unaffected. Rename it to
ATA_QUIRK_NO_LPM_ON_ATI_AND_AMD and extend the vendor check to AMD for
the same Samsung SSD model patterns. Keep LPM behavior unchanged for
other controller vendors, including Intel.
Leave ATA_QUIRK_NO_NCQ_ON_ATI restricted to ATI, since the reported AMD
issue concerns LPM rather than NCQ.
Link: https://bugzilla.kernel.org/show_bug.cgi?id=221986
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org>> ---
Link: https://lore.kernel.org/r/20260918124030.1962773-5-cassel@kernel.org
Signed-off-by: Niklas Cassel <cassel@kernel.org>
|
|
make htmldocs fails with:
Documentation/driver-api/libata:604: ./drivers/ata/libata-scsi.c:2301:
ERROR: Unexpected indentation. [docutils]
The Return: section of ata_dsm_trim_pages() ends its first sentence with
a colon and continues with an indented bullet list. reStructuredText
requires a blank line before an indented block, so docutils chokes on
the list.
Simply adding the missing blank line does not work either: Return: is a
kernel-doc "special section", which is terminated by the first blank
line, so the bullet list would end up in the Description section,
detached from the sentence introducing it.
Spell the two bounds out as prose instead, so that the Return: section
stays self-contained. Documentation-only change.
Fixes: e64e6b5dc867 ("ata: libata-scsi: scale DSM TRIM payload by MAX PAGES PER DSM COMMAND")
Reported-by: Thomas Huth <thuth@redhat.com>
Closes: https://lore.kernel.org/linux-ide/b0a0b8e8-cc4f-4c7b-8bc5-fee0712d405a@redhat.com/
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Reviewed-by: Thomas Huth <thuth@redhat.com>
Link: https://lore.kernel.org/r/20260916130031.29990-2-cassel@kernel.org
Signed-off-by: Niklas Cassel <cassel@kernel.org>
|