Skip to content

FROMLIST: nvme: Add adaptive PCIe link rate switching - #2008

Open
Hariharan Sreedhar (hariharan-sree) wants to merge 6 commits into
qualcomm-linux:tech/bus/pci/allfrom
hariharan-sree:nvme_dynamic_BW_scale
Open

Hariharan Sreedhar (hariharan-sree) wants to merge 6 commits into
qualcomm-linux:tech/bus/pci/allfrom
hariharan-sree:nvme_dynamic_BW_scale

Conversation

@hariharan-sree

Copy link
Copy Markdown

Add support for dynamically adjusting the PCIe link rate of NVMe devices based on I/O activity. The link is downgraded to a
configurable minimum rate during periods of low activity to save power, and upgraded to the maximum supported rate as soon as the workload exceeds a computed threshold, so that peak performance is maintained under load.

The feature is currently only supported on Hygon platforms.
Link : https://patch.msgid.link/9e99e007bf05a608b1b82da740c2b33935a27175.1790222172.git.liaoxuan@hygon.cn

Liao Xuan added 3 commits October 9, 2026 17:13
Add support for dynamically adjusting the PCIe link rate of NVMe
devices based on I/O activity. The link is downgraded to a
configurable minimum rate during periods of low activity to save
power, and upgraded to the maximum supported rate as soon as the
workload exceeds a computed threshold, so that peak performance is
maintained under load.

The amount of transferred data is accumulated on the I/O submission
path and aggregated by a per-controller timer. The threshold is
derived from the bandwidth of the minimum and maximum rates such that
the extra time required to transfer the data at the minimum rate does
not exceed the estimated link retraining time.

The feature is currently only supported on Hygon platforms.

Four NVMe devices were tested on the link with this patch. The power
benefit was measured by sampling the voltage and current of the power
monitor every 1 second, each run lasted 3 minutes and the test was
repeated 5 times.

Test run    Power Benefit(w)
1           5.3236
2           5.2285
3           5.4176
4           5.4538
5           5.3444
Average     5.3536

The adaptive link rate switching saves about 5.3w of system power
on average.

The feature was tested with Gen1 as the minimum link rate. The workload
runs 10 seconds of random read followed by 0.5 seconds of idle, repeated
100 times.

                              Avg IOPS (k)  Avg BW (GiB/s)  Avg clat (us)
With speed switching          2671.2        10.19           378.1
Without speed switching       2681.5        10.23           377.0

The differences are within measurement noise (< 0.5%), which shows
that the adaptive link rate switching does not degrade the baseline
performance and keeps the I/O stable.

In summary, the adaptive link rate switching keeps the I/O performance
essentially unchanged (the differences in IOPS, bandwidth and latency
are all within 0.5% measurement noise) while saving about 5w of
system power, providing an effective way to reduce the platform power
consumption under light I/O without sacrificing throughput or latency.

Signed-off-by: Liao Xuan <liaoxuan@open-hieco.net>
Signed-off-by: Zheng Tan <tanzheng@kylinos.cn>
Link: https://patch.msgid.link/9e99e007bf05a608b1b82da740c2b33935a27175.1790222172.git.liaoxuan@hygon.cn
Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
The rate switching decision currently toggles the link rate in every
monitoring window based on the I/O of that window alone, which can
cause excessive rate switching when the workload fluctuates around the
threshold.

Add counters that require the threshold to be exceeded for a
consecutive number of windows before the rate is upgraded, and to stay
below the threshold for a consecutive number of windows before it is
downgraded. The upgrade happens immediately (threshold 0 by default)
so that latency is not sacrificed, while the downgrade requires 10
consecutive windows.

Monitoring is stopped after 10 seconds of inactivity and is re-armed by
the next I/O.

Signed-off-by: Liao Xuan <liaoxuan@open-hieco.net>
Link: https://patch.msgid.link/de23e8b9699b5a7a77e7ee215e8e4c9f530ce5da.1790222172.git.liaoxuan@hygon.cn
Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
Add a speed attribute group under /sys/class/nvme/nvmeX/ that
exposes the adaptive link rate switching parameters for runtime tuning:

  enable           - toggle the feature on/off (0/1)
  monitor_interval - I/O monitoring window in milliseconds (>= 100)
  min_speed        - minimum PCIe generation to downgrade to (1-5)
  up_threshold     - windows above the threshold needed to upgrade
  down_threshold   - windows below the threshold needed to downgrade

The attributes are read/write and are removed when the controller is
torn down.

Signed-off-by: Liao Xuan <liaoxuan@open-hieco.net>
Link: https://patch.msgid.link/be224cd8385a5c28a20db65d635000a7be4da91a.1790222172.git.liaoxuan@hygon.cn
Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
@qcomlnxci
qcomlnxci requested review from a team, krishnachaitanya-linux and Matthew Leung (meleung) and removed request for a team October 9, 2026 11:57
… around link retraining

PCIe host bridge controllers may need their operating point raised before
retraining to a higher link speed so that hardware resources (e.g., RPMh
votes on Qualcomm platforms) are available at the requested data rate.
After retraining, the operating point must be updated to reflect the
actual negotiated speed.

Add pcie_set_opp() to look up an OPP on the host bridge parent device
using a key of (per-lane frequency in kHz, LNKCTL2 Target Link Speed
level).  Keying by generation rather than total bandwidth lets OPP tables
remain width-independent.

In pcie_set_target_speed(), call pcie_set_opp() before retraining only
when upscaling (speed_req > cur_bus_speed), since only raising the
operating point requires pre-staging hardware.  After retraining, call
pcie_set_opp() unconditionally with the actual cur_bus_speed to settle
the votes.  Both calls are skipped for downstream ports of PCIe switches,
as those are outside the host controller's scope.

Some controllers also require ASPM to be disabled around link retraining.
Add a disable_aspm_for_retrain flag to pci_host_bridge; when set,
pcie_set_target_speed() saves the child device's ASPM state, disables all
ASPM link states before retraining, and restores them afterward.

Link: https://patch.msgid.link/20260819-bwscale-v5-1-6dea79786b37@oss.qualcomm.com
Signed-off-by: Krishna Chaitanya Chundru <krishna.chundru@oss.qualcomm.com>
Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
Set the disable_aspm_for_retrain flag in qcom_pcie_host_init() to ensure
ASPM is disabled during link retraining operations. This prevents potential
issues with link state transitions during bandwidth scaling on Qualcomm
PCIe controllers.

Signed-off-by: Krishna Chaitanya Chundru <krishnac@codeaurora.org>
Link: https://patch.msgid.link/20260819-bwscale-v5-6-6dea79786b37@oss.qualcomm.com
Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
The Qualcomm PCIe controller hardwires the Link Bandwidth Notification
Capability (LBNC) bit to 0, even though the Root Port supports retraining
the link to the speeds advertised in LNKCAP2.

The PCIe port bandwidth controller relies on this capability to determine
whether link bandwidth control is supported. With LBNC cleared, the
bandwidth controller is not registered and link speed cannot be managed
through the associated thermal cooling device.

Override the read-only LNKCAP register through the DBI RO write interface
and set PCI_EXP_LNKCAP_LBNC during controller initialization.

This allows the PCIe bandwidth controller to bind to the Root Port and
enables link speed scaling through the thermal framework.

Signed-off-by: Krishna Chaitanya Chundru <krishna.chundru@oss.qualcomm.com>
Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants