| Age | Commit message (Collapse) | Author |
|
for AMD USB4 routers"
This commit has caused a deadlock at shutdown. A proper fix with
another approach will be coming later. Revert commit
f1de1fc5f632cdeae1f5c2984572ab710d4dfcaa for now.
Reported-by: juan.martinez@amd.com
Closes: https://lore.kernel.org/linux-usb/20260825214237.4179813-1-juan.martinez@amd.com/
Cc: Sanath S <Sanath.S@amd.com>
Cc: Basavaraj Natikar <Basavaraj.Natikar@amd.com>
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Mika Westerberg <mika.westerberg@linux.intel.com>
|
|
When connected to another host and then unplugging cable lockdep
triggers following:
======================================================
WARNING: possible circular locking dependency detected
7.1.0-rc2+ #1775 Tainted: G U
------------------------------------------------------
kworker/u16:6/312 is trying to acquire lock:
ffff8881179c70a8 ((work_completion)(&ring->work)){+.+.}-{0:0}, at: __flush_work+0x3cf/0xd10
but task is already holding lock:
ffff8881a8b810b0 (&net->connection_lock){+.+.}-{4:4}, at: tbnet_tear_down+0x110/0x720 [thunderbolt_net]
which lock already depends on the new lock.
the existing dependency chain (in reverse order) is:
-> #1 (&net->connection_lock){+.+.}-{4:4}:
__mutex_lock+0x19a/0x2490
mutex_lock_nested+0x1b/0x30
tbnet_handle_packet+0x74c/0xd70 [thunderbolt_net]
tb_xdomain_handle_request+0x37c/0x4b0 [thunderbolt]
tb_domain_event_cb+0xc9/0x140 [thunderbolt]
tb_ctl_handle_event+0xd6/0x2c0 [thunderbolt]
tb_ctl_rx_callback+0x22c/0xa10 [thunderbolt]
ring_work+0x715/0xcb0 [thunderbolt]
process_one_work+0x902/0x1790
worker_thread+0x5cd/0xfe0
kthread+0x339/0x420
ret_from_fork+0x79a/0x9d0
ret_from_fork_asm+0x1a/0x30
-> #0 ((work_completion)(&ring->work)){+.+.}-{0:0}:
__lock_acquire+0x1592/0x2640
lock_acquire+0x1a3/0x300
__flush_work+0x3e9/0xd10
flush_work+0x21/0x30
tb_ring_stop+0x240/0x840 [thunderbolt]
tbnet_tear_down+0x2ff/0x720 [thunderbolt_net]
tbnet_stop+0x47/0x1a0 [thunderbolt_net]
__dev_close_many+0x19e/0x4e0
netif_close_many+0x1e8/0x640
unregister_netdevice_many_notify+0x6d3/0x22d0
unregister_netdevice_queue+0x2b9/0x3a0
unregister_netdev+0x1c/0x70
tbnet_remove+0x52/0xb0 [thunderbolt_net]
tb_service_remove+0x8a/0xe0 [thunderbolt]
device_remove+0xc5/0x190
device_release_driver_internal+0x3db/0x590
device_release_driver+0x12/0x20
bus_remove_device+0x2c1/0x580
device_del+0x3d9/0x9f0
device_unregister+0x17/0xc0
unregister_service+0x46/0x60 [thunderbolt]
device_for_each_child_reverse+0xfa/0x180
tb_xdomain_unregister+0x57/0xe0 [thunderbolt]
unregister_unplugged_xdomain+0x101/0x1a0 [thunderbolt]
bus_for_each_dev+0x111/0x1a0
tb_domain_unregister_unplugged_xdomains+0x98/0xe0 [thunderbolt]
tb_handle_hotplug+0xc3/0x2bb0 [thunderbolt]
process_one_work+0x902/0x1790
worker_thread+0x5cd/0xfe0
kthread+0x339/0x420
ret_from_fork+0x79a/0x9d0
ret_from_fork_asm+0x1a/0x30
other info that might help us debug this:
Possible unsafe locking scenario:
CPU0 CPU1
---- ----
lock(&net->connection_lock);
lock((work_completion)(&ring->work));
lock(&net->connection_lock);
lock((work_completion)(&ring->work));
This in fact is false positive because they involve unrelated rings (and
unrelated work structures). In the first one it is ring 0 which is used
for control traffic and in the second it is dealing with another ring
used for the high-speed traffic.
Fix this by using separate lock class for each ring worker.
Signed-off-by: Mika Westerberg <mika.westerberg@linux.intel.com>
|
|
There is a slight race between tb_remove_work() and tb_domain_remove()
which leads to dereferencing a NULL tb->root_switch pointer inside
tb_free_unplugged_xdomains():
Thread A Thread B
tb_remove_work()
tb_domain_remove()
mutex_lock(&tb->lock)
tb_stop()
/* doesn't cancel a running callback */
cancel_delayed_work(&tcm->remove_work)
...
tb_switch_remove(tb->root_switch)
tb->root_switch = NULL
mutex_unlock(&tb->lock)
mutex_lock(&tb->lock)
...
/* without checking ->root_switch */
tb_free_unplugged_xdomains(tb->root_switch)
mutex_unlock(&tb->lock)
Commit a8937f35cf39 ("thunderbolt: Remove XDomain from the bus without
holding tb->lock") doesn't seem right to move tb_free_unplugged_xdomains()
out of the &tb->lock section and the check for tb->root_switch, in
particular. It states:
For this reason separate removing the XDomain from the topology data
structures (where we need the lock) from unregistering the device from
the bus (where remove callbacks of the drivers are being called).
tb_free_unplugged_xdomains() belongs to the former group of functions
requiring the lock. And it also calls tb_xdomain_remove() which should
only be called with &tb->lock held.
Found by Linux Verification Center (linuxtesting.org) with Svace static
analysis tool.
Fixes: a8937f35cf39 ("thunderbolt: Remove XDomain from the bus without holding tb->lock")
Cc: stable@vger.kernel.org
Signed-off-by: Fedor Pchelkin <pchelkin@ispras.ru>
Signed-off-by: Mika Westerberg <mika.westerberg@linux.intel.com>
|