| Age | Commit message (Collapse) | Author |
|
Pull bpf fixes from Alexei Starovoitov:
- Fix overflow of backward jump offset in constant blinding
(Alexei Starovoitov)
- Fix packet range of packet pointers sharing an id when var_off
tightens umax of one pointer and not the other (Alexei Starovoitov)
- Fix objects stuck in free_by_rcu_ttrace list of bpf memalloc
(Alexei Starovoitov)
- Fix use-after-free of progs detached from busy trampolines: wait for
an RCU tasks grace period before freeing trampoline progs, and patch
detached progs out of trampoline images that are still in use
(Florent Revest)
- Hold map BTF for the memory allocator destructor record to fix UAF in
deferred bpf_mem_alloc destruction (Kumar Kartikeya Dwivedi)
- Fix missing migration protection in resizable hashtab
lookup_and_delete batch operation (Ömer Mete Kaya)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf:
bpf: Fix missing migration protection in __rhtab_map_lookup_and_delete_batch()
selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace
bpf: Fix objects stuck in free_by_rcu_ttrace
bpf: Factor out __do_call_rcu_ttrace()
selftests/bpf: Test packet range of pointers sharing an id
bpf: Fix packet range of pointers sharing an id
selftests/bpf: Detach a trampoline prog while a task sleeps before it
bpf: Skip detached progs in trampoline images that are still in use
bpf: Wait for an RCU tasks grace period before freeing trampoline progs
bpf: Hold map BTF for the memory allocator destructor record
bpf: Fix overflow of jump offset in constant blinding
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux
Pull block fixes from Jens Axboe:
- NVMe fixes via Keith:
- Fix an out-of-bounds write in nvmet_auth_challenge(), where
sizeof() on a void pointer undercounted the challenge header and
let a short AUTH_RECEIVE buffer pass the check
- nvme-multipath fixes for an ANA log bounds check underflow, the
command effects log lifetime for multipath heads, and only
setting BLK_FEAT_ZONED after the zone info is known.
- nvmet fixes for ns->enabled teardown ordering, rejecting I/O
after the percpu ns reference is killed, device path preservation
on allocation failure, and too-short SGL segments in pci-epf
- nvme-tcp: revert the per-socket dynamic lockdep keys, and delay
the socket reclassification
- A DMA pool alignment quirk for the Micron 4100AT
- Controller state/reset race fixes, and -Wformat-security
workarounds
- blk-mq: set RQF_USE_SCHED when the operation is known, and allow
cached requests to be used for flush operations
- Reject polled dio with user integrity metadata
- Save the IRQ state in blkg_tryget_closest()
- Set the zone write granularity in virtio_blk
- ublk selftest fixes
* tag 'block-7.3-20261002' of git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux: (23 commits)
virtio_blk: set the zone write granularity
nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known
nvme: fix command effects log lifetime for multipath heads
nvmet: don't allow I/O admission after percpu ns reference is killed
nvmet: defer setting ns->enabled to false in nvmet_ns_disable()
nvmet: copy the hostid into the ctrl before creating PR pc_refs
nvmet-auth: fix out-of-bounds write in nvmet_auth_challenge()
nvmet: pci-epf: reject too-short SGL segments
nvme-multipath: fix underflow in ANA log bounds checks
nvme: work around all -Wformat-security warnings
nvme: work around -Wformat-security warning
nvme: do not reset controllers in NVME_CTRL_NEW state
nvme-tcp: delay nvme_tcp_reclassify_socket()
Revert "nvme-tcp: lockdep: use dynamic lockdep keys per socket instance"
drbd: remove unused drbd_nl_mcgrps[] array
blk-mq: allow cached requests to be used for flush operations
blk-mq: set RQF_USE_SCHED when the operation is known
block: reject polled dio with user integrity metadata
selftests: ublk: fix unused_result error
blk-cgroup: save IRQ state in blkg_tryget_closest()
...
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux
Pull io_uring fixes from Jens Axboe:
- Fix a task_work add use-after-free with SQPOLL.
The sqpoll thread could pop and complete the last request while
io_req_normal_work_add() was still looking at them after the mpscq
push.
Use the same approach as DEFER_TASKRUN to protect from that, holding
an RCU read lock across the add, and have exit wait for an RCU grace
period for SQPOLL rings as well.
- CQE32 ring fixes: correct the free entry check for 32b CQEs, zero the
big_cqe for aux CQEs, and only post the dummy skip CQE on CQE_MIXED
rings
- Mark the source filter table as COW when cloning bpf filters, so
registering another filter on the source doesn't modify the shared
table in place
- Initialize the task context before running the BPF loop
- Requeue zcrx multishot receives stopped by a local resource
- End a TX_TIMESTAMP multishot cmd when the CQ is full (lollipopkit)
* tag 'io_uring-7.3-20261002' of git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux:
io_uring: fix task_work add use-after-free with SQPOLL
io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full
io_uring/zcrx: requeue multishot receives stopped by a local resource
io_uring: initialize task context before running the BPF loop
io_uring: zero big_cqe for aux CQEs on CQE32 rings
io_uring: fix free entry check for 32b CQEs on CQE32 rings
io_uring: only post the dummy skip CQE on CQE_MIXED rings
io_uring/bpf_filter: mark source as COW when cloning filters
|
|
Delete all elements of BPF_F_NO_PREALLOC hash map in one batch. The first
free_bulk() starts RCU tasks trace GP and the rest of the elements are
freed while it's in flight. Wait for call_rcu_ttrace_in_progress to clear
in bpf_mem_cache of every cpu and check that free_by_rcu_ttrace and
waiting_for_gp_ttrace lists are empty.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20260930095920.601738-4-alexei.starovoitov@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
|
|
virtblk_read_zoned_limits() reads the write granularity that the device
reports in virtio_blk_zoned_characteristics and assigns it to the
physical block size and to io_min, but never to the limit that is named
after it. queue_limits.zone_write_granularity is left at zero, so
blk_validate_zoned_limits() raises it to the logical block size:
if (lim->zone_write_granularity < lim->logical_block_size)
lim->zone_write_granularity = lim->logical_block_size;
A device that reports a granularity coarser than its logical block size,
which is what the field exists to express, therefore has it silently
reduced. A 512e host managed disk passed through to a guest reports a
logical block size of 512 and a write granularity of 4096, and the guest
ends up with a zone write granularity of 512.
bio_split_alignment() returns lim->zone_write_granularity if it is non-zero
and bio_split_io_at() may split a bio with as per bio_split_alignment().
This can real to the write getting rejected by the host drive, as the write
is not aligned to the physical block size.
zonefs also takes its block size from bdev_zone_write_granularity(), so it
would incorrectly use 512 on a disk that requires 4096.
sd_zbc_read_zones() sets the limit from the physical block size for the
same reason. NVMe ZNS and null_blk leave it unset, but the fallback
gives the right answer for them, as their write granularity is the
logical block size. virtio carries a separate value that may exceed it.
Set the zone write granularity from the value that the device reports.
Fixes: 95bfec41bd3d ("virtio-blk: add support for zoned block devices")
Signed-off-by: Niklas Cassel <cassel@kernel.org>
Reviewed-by: Stefan Hajnoczi <stefanha@redhat.com>
Link: https://patch.msgid.link/20260918140641.2031075-2-cassel@kernel.org
Signed-off-by: Jens Axboe <axboe@kernel.dk>
|