[SRU][R][PATCH 1/1] scsi: megaraid_sas: Limit NVMe request size to the PRP chain frame

Alex Shi alex.shi at canonical.com
Wed Sep 9 02:26:44 UTC 2026


Any comments for this patchset?

Thanks


On 2026/9/8 09:24, Alex Shi wrote:
> The 'CC' email are all unnecessary, please just reply our ubuntu list.
>
> Thanks!
>
> On 2026/9/7 16:53, Alex Shi wrote:
>> From: Thomas Lamprecht <t.lamprecht at proxmox.com>
>>
>> BugLink: https://bugs.launchpad.net/bugs/2148534
>>
>> megasas_make_prp_nvme() builds a command's PRP list in cmd->sg_frame, a
>> DMA pool buffer of instance->max_chain_frame_sz bytes, spending one
>> entry per NVMe page of the transfer plus one per page of the buffer for
>> the chain pointer. The loop runs until the transfer is described and
>> never checks the buffer bound.
>>
>> max_hw_sectors comes straight from the MDTS the firmware reports for the
>> drive. On drives with a large MDTS the only thing keeping the list
>> inside the buffer was the block layer default of 1280 KiB, which needs
>> 320 entries, which fit into a 4 KiB frame as that holds 512. But since
>> commit 9b8b84879d4a ("block: Increase BLK_DEF_MAX_SECTORS_CAP") that
>> default is 4 MiB, and such a transfer needs 1025 entries, so the list
>> runs a full page past the end of the frame:
>>
>>    sd 1:0:1:0: [sdb] tag#630 page boundary ptr_sgl: 0x00000000ba62d13f
>>    BUG: unable to handle page fault for address: ff663bcb81e7c000
>>    #PF: supervisor write access in kernel mode
>>    #PF: error_code(0x0002) - not-present page
>>    RIP: 0010:megasas_build_and_issue_cmd_fusion+0xeaa/0x1870 
>> [megaraid_sas]
>>
>> If the page after the frame happens to be mapped, the overrun does not
>> fault but silently corrupts the neighbouring pool entry, which is
>> another in-flight command's PRP list.
>>
>> Cap max_hw_sectors at what the chain frame can describe, less one page
>> for transfers that do not start on a page boundary and so need one entry
>> more. This is the megaraid_sas counterpart of commit 04631f55afc5
>> ("scsi: mpt3sas: Limit NVMe request size to 2 MiB"), but derives the
>> limit from max_chain_frame_sz rather than hardcoding it.
>>
>> Cc: stable at vger.kernel.org
>> Fixes: 9b8b84879d4a ("block: Increase BLK_DEF_MAX_SECTORS_CAP")
>> Reported-by: Lukasz Magiera <me at magik.net>
>> Closes: 
>> https://lore.kernel.org/all/GPhsSM0vkgyIrs0DIZ62qeUZX7X4RxwQXVKiuvMx-lHQVSPDxpztUyQOGS0xikqvJ-Z94hMV-dW_5KN_0CX2hsfV7kTf_t0MTf6vdAAaSEc=@magik.net/
>> Reported-by: Mira Limbeck <m.limbeck at proxmox.com>
>> Closes: 
>> https://lore.kernel.org/all/d171cc76-bf25-48ce-b482-d344669dfc24@proxmox.com/
>> Suggested-by: Martin K. Petersen <martin.petersen at oracle.com>
>> Link: https://lore.kernel.org/all/yq17bmzd5jr.fsf@ca-mkp.ca.oracle.com/
>> Signed-off-by: Thomas Lamprecht <t.lamprecht at proxmox.com>
>> Closes: 
>> https://lore.kernel.org/linux-scsi/20260827182106.535D61F000E9@smtp.kernel.org
>> Link: 
>> https://patch.msgid.link/20260827175743.734593-1-t.lamprecht@proxmox.com
>> Signed-off-by: Martin K. Petersen (Oracle) <mkp at kernel.org>
>> (cherry picked from commit af8c27375733fb6a6df9fa484cda77cc3dd0cb80)
>> Signed-off-by: Alex Shi <alex.shi at canonical.com>
>> ---
>>   drivers/scsi/megaraid/megaraid_sas_base.c | 13 ++++++++++++-
>>   1 file changed, 12 insertions(+), 1 deletion(-)
>>
>> diff --git a/drivers/scsi/megaraid/megaraid_sas_base.c 
>> b/drivers/scsi/megaraid/megaraid_sas_base.c
>> index ac71ea4898b2..436d4e0a15c4 100644
>> --- a/drivers/scsi/megaraid/megaraid_sas_base.c
>> +++ b/drivers/scsi/megaraid/megaraid_sas_base.c
>> @@ -1973,12 +1973,23 @@ megasas_set_nvme_device_properties(struct 
>> scsi_device *sdev,
>>   {
>>       struct megasas_instance *instance;
>>       u32 mr_nvme_pg_size;
>> +    u64 max_prp_io;
>>         instance = (struct megasas_instance *)sdev->host->hostdata;
>>       mr_nvme_pg_size = max_t(u32, instance->nvme_page_size,
>>                   MR_DEFAULT_NVME_PAGE_SIZE);
>>   -    lim->max_hw_sectors = max_io_size / 512;
>> +    /*
>> +     * megasas_make_prp_nvme() builds the PRP list in cmd->sg_frame 
>> without
>> +     * bounding it against that buffer, and spends one entry per 
>> page of
>> +     * it on the chain pointer. Cap the transfer at what the buffer 
>> holds,
>> +     * less one page for lists that start off a page boundary.
>> +     */
>> +    max_prp_io = (u64)((instance->max_chain_frame_sz / sizeof(u64)) -
>> +               (instance->max_chain_frame_sz / mr_nvme_pg_size) - 1) *
>> +             mr_nvme_pg_size;
>> +
>> +    lim->max_hw_sectors = min_t(u64, max_io_size, max_prp_io) >> 
>> SECTOR_SHIFT;
>>       lim->virt_boundary_mask = mr_nvme_pg_size - 1;
>>   }
>




More information about the kernel-team mailing list