commit 469fa055664bfd8717b087a80d59c03055789c17 Author: Stanley Jhu Date: Mon Sep 28 11:58:16 2026 +0800 scsi: ufs: core: Decouple CQ sweep from request iterator in MCQ In MCQ mode, ufshcd_mcq_compl_pending_transfer() uses blk_mq_tagset_busy_iter() to iterate over busy requests during error recovery and host reset. However, both iterator callbacks perform whole-queue operations redundantly for each visited request: - force_compl == true: ufshcd_mcq_force_compl_one() calls ufshcd_mcq_compl_all_cqes_lock() on every busy request, sweeping the entire completion ring (hwq->max_entries slots) once per active request under spin_lock_irqsave even though the first sweep already cleared all completion entries. - force_compl == false: ufshcd_mcq_compl_one() acquires cq_lock and polls CQTPy over MMIO via ufshcd_mcq_poll_cqe_lock() for every busy request without doing any per-request work. Sweep or poll each hardware queue (hba->uhq[i]) once at the start of ufshcd_mcq_compl_pending_transfer(). When force_compl is true, run blk_mq_tagset_busy_iter() afterward to complete residual in-flight requests with DID_REQUEUE, and remove the now-unused ufshcd_mcq_compl_one() callback. Fixes: ab248643d3d6 ("scsi: ufs: core: Add error handling for MCQ mode") Cc: stable@vger.kernel.org Reviewed-by: Bart Van Assche Signed-off-by: Stanley Jhu Reviewed-by: Peter Wang Link: https://patch.msgid.link/20260928035816.1294326-3-stanleyjhu@google.com Signed-off-by: Martin K. Petersen (Oracle) commit 375f3a5691dd46b5d05b90a8968872ec8b2248c5 Author: Stanley Jhu Date: Mon Sep 28 11:58:15 2026 +0800 scsi: ufs: core: Avoid unsafe MMIO reads in ufshcd_mcq_compl_all_cqes_lock() During MCQ host reset, ufshcd_host_reset_and_restore() stops the host controller via ufshcd_hba_stop() (HCE = 0) before calling ufshcd_complete_requests(hba, true) -> ufshcd_mcq_compl_pending_transfer(hba, true) -> ufshcd_mcq_force_compl_one() -> ufshcd_mcq_compl_all_cqes_lock(). Because ufshcd_mcq_force_compl_one() is its sole caller, ufshcd_mcq_compl_all_cqes_lock() always runs with HCE = 0. Despite the comment above ufshcd_mcq_compl_all_cqes_lock() stating that reading CQTPy may not be safe with the controller disabled, the function still calls ufshcd_mcq_update_cq_tail_slot() at the end of its sweep: 1. Unsafe CQTPy MMIO read: Calling ufshcd_mcq_update_cq_tail_slot() at the end of the sweep reads CQTPy over MMIO while HCE = 0, directly contradicting the function's documented contract (commit 1373df88d535 ("scsi: ufs: core: Add a comment block above ufshcd_mcq_compl_all_cqes_lock()")) that reading CQTPy may not be safe with the controller disabled. 2. Spurious error logs on empty slots: Sweeping all max_entries slots visits empty entries where command_desc_base_addr is 0, causing ufshcd_mcq_process_cqe() to log unguarded dev_err(hba->dev, "Abnormal CQ entry!\n") messages. Fix both issues in ufshcd_mcq_compl_all_cqes_lock(): - Remove the ufshcd_mcq_update_cq_tail_slot() call and the redundant hwq->cq_head_slot = hwq->cq_tail_slot assignment without replacement. The two indices are already equal after the sweep: they are equal when the sweep starts, since ufshcd_mcq_poll_cqe_lock() consumes entries until cq_head_slot reaches cq_tail_slot, and the sweep advances cq_head_slot by exactly one full ring. Both indices are also reinitialized before the queue is reused. - Extract ufshcd_mcq_compl_cqe() and invoke it only on non-empty slots during full-ring sweeps, keeping "Abnormal CQ entry!" logging strictly for unexpected empty entries in ufshcd_mcq_poll_cqe_lock(). Fixes: ab248643d3d6 ("scsi: ufs: core: Add error handling for MCQ mode") Cc: stable@vger.kernel.org Reviewed-by: Peter Wang Signed-off-by: Stanley Jhu Reviewed-by: Bart Van Assche Link: https://patch.msgid.link/20260928035816.1294326-2-stanleyjhu@google.com Signed-off-by: Martin K. Petersen (Oracle) commit 27c0f718d3b3c99e6d315308e95d987d0382a8a8 Author: Ahmed Abdelhaleem Ahmed Date: Tue Sep 22 15:46:46 2026 +0000 scsi: ch: Do not keep references to data transfer element devices ch_readconfig() looks up the scsi_device of every data transfer element whose SCSI id the changer reports in READ ELEMENT STATUS, and stores it in ch->dt[]. scsi_device_lookup() takes a reference, and nothing ever drops it: ch_destroy() frees the array with kfree(). ch->dt[] is read nowhere else - it only supplies the vendor, model and revision printed in the same loop. Once such a drive is removed, its scsi_device can never be released. It stays on the host's device list at its address, so a new device there is refused by anything that walks the list - target_core_pscsi reports "scsi_device_get() failed for H:C:T:L" - and the low-level driver's module can no longer be unloaded. Only a reboot recovers. It shows with any changer that reports its drives' ids; with the mhvtl virtual library (IBM 3573-TL personality), each create and remove of a library with four drives leaves four references behind, counted by the module's use count in lsmod. With ch not bound the count is unchanged, and with this patch applied it is unchanged too. Drop the reference as soon as the name has been printed, and remove the now unused dt[] array. That also removes the array leaked when ch_probe() fails after ch_readconfig(). The same leak was reported with an RFC patch in 2022, which was not merged. Fixes: daa6eda65a53 ("[SCSI] add scsi changer driver") Cc: stable@vger.kernel.org Link: https://lore.kernel.org/linux-scsi/20220719075442.6215-1-yanghao_ht@163.com/ Signed-off-by: Ahmed Abdelhaleem Ahmed Reviewed-by: Laurence Oberman Fixes: daa6eda65a53 ("[SCSI] add scsi changer driver") Link: https://patch.msgid.link/20260929-ch-dt-leak-v2-1-ddda8d635dca@gmail.com Signed-off-by: Martin K. Petersen (Oracle) commit 78e04810542da416d3bead707b60cec773f1524d Author: Pengpeng Hou Date: Sun Sep 20 11:49:53 2026 +0800 scsi: ufs: mediatek: Handle mPHY power-on failures The MediaTek helper enables VA09 before powering on the mPHY, but ignores the PHY result. A failure can leave VA09 enabled and publish mphy_powered_on as true, so a later power-on attempt is skipped. Check phy_power_on(), attempt to unwind VA09 on failure, and leave the state flag unchanged. Propagate the helper result from host initialization as well as its existing resume caller. Keep the original PHY error if the supply rollback also fails, and report that rollback separately. The issue was found by our static-analysis tool. Fixes: 561e3a8726b2 ("scsi: ufs-mediatek: Fix unbalanced clock on/off") Fixes: cf137b3ea49a ("scsi: ufs-mediatek: Support VA09 regulator operations") Assisted-by: gpt 5 Signed-off-by: Pengpeng Hou Reviewed-by: Peter Wang Reviewed-by: Stanley Jhu Link: https://patch.msgid.link/20260920034953.19535-1-hppiscas@163.com Signed-off-by: Martin K. Petersen (Oracle) commit 85de8b3a5eaa23d153aff4c7c83dddbc949064c9 Author: Runyu Xiao Date: Thu Sep 10 17:24:05 2026 +0800 scsi: qedi: Initialize callback state before registration qedi_get_protocol_tlv_data() can be called asynchronously by QED after qedi_ops->register_ops(). It takes stats_lock and reads ll2_mtu, but __qedi_probe() currently registers the callback before initializing the mutex and assigning the default MTU on the normal probe path. Initialize the callback-visible state before registering qedi_cb_ops. Keep the recovery path from reinitializing state because it reuses the existing qedi context. Fixes: 3cc5746e5ad7 ("scsi: qedi: Initialize the stats mutex lock") Cc: stable@vger.kernel.org Assisted-by: LLM Codex Signed-off-by: Runyu Xiao Link: https://patch.msgid.link/20260910092405.1300129-1-runyu.xiao@seu.edu.cn Signed-off-by: Martin K. Petersen (Oracle) commit 02a3178319639f650428164d4d003de4dc03e326 Author: Naomi Chu Date: Wed Sep 9 17:10:45 2026 +0800 scsi: ufs: core: Hold a clock reference across the probe Clock gating becomes possible as soon as ufshcd_init_clk_gating() has run, and from that point on the probe keeps accessing host registers without ever taking a clock reference. This has been safe only because of the state check in __ufshcd_release(): gate_work is not queued unless hba->ufshcd_state is UFSHCD_STATE_OPERATIONAL, and the promotion to that state used to happen after the last register access of the probe, at the end of ufshcd_probe_hba(). That is fragile: it only works while the promotion happens after the register accesses. Commit a390e6677f41 ("scsi: ufs: core: Expand the ufshcd_device_init(hba, true) call") changed that ordering by moving the promotion into ufshcd_init(), which schedules ufshcd_async_scan() afterwards. ufshcd_probe_hba() therefore now runs with the state already promoted, and it accesses host registers without holding a clock reference: - On hosts with UFSHCD_QUIRK_REINIT_AFTER_MAX_GEAR_SWITCH it calls ufshcd_hba_stop() and ufshcd_hba_enable() before ufshcd_device_init() sets the state back to UFSHCD_STATE_RESET, and both read REG_CONTROLLER_ENABLE, so gated clocks stall there instead of just losing a write. - It ends with an ufshcd_configure_auto_hibern8() write, which its other callers do take a clock reference for. Take a clock reference as soon as clock gating has been initialised and keep it until the probe is over. It then does not matter who drops a clock reference while the probe is running, and the register accesses of the probe no longer depend on hba->ufshcd_state. The reference is dropped by ufshcd_async_scan() once the scan has finished, or by the new out_release label if the probe fails after it was taken. Fixes: a390e6677f41 ("scsi: ufs: core: Expand the ufshcd_device_init(hba, true) call") Signed-off-by: Naomi Chu Reviewed-by: Peter Wang Link: https://patch.msgid.link/20260909091045.1134956-1-naomi.chu@mediatek.com Signed-off-by: Martin K. Petersen (Oracle)