[v2,2/3] clk: sunxi-ng: sun55i-a523: GPU clock divider is fractional, not linear

Message ID 20260803180755.288793-3-juanmanuellopezcarrillo@gmail.com (mailing list archive)
State New
Headers
Series clk: sunxi-ng: fix the A523/T527 GPU clock model, enable GPU DVFS |

Commit Message

Juan Manuel López Carrillo Aug. 3, 2026, 6:07 p.m. UTC
The GPU mod clock (0x670) was modelled as a linear M+1 divider, but the
M factor of this register is a cycle-masking divider: GPU_CLK = Clock
Source * ((16-M)/16) (T527 user manual v0.92, section 2.7.6.58).

With the linear model every OPP that needed M > 0 silently ran the GPU
faster than requested.  Measured on an Orange Pi 4A (T527) with the
Mali cycle counter against the programmed register:

	OPP request	programmed	real rate
	150 MHz		600M, M=3	487.5 MHz
	200 MHz		800M, M=3	650   MHz
	300 MHz		600M, M=1	562.5 MHz
	400 MHz		800M, M=1	750   MHz
	600 MHz		600M, M=0	600   MHz

i.e. the "400 MHz" OPP ran the GPU at 750 MHz, 25% above the vendor
ceiling of 600 MHz, at the low-OPP voltage.  Thermal throttling to
"400 MHz" actually overclocked the GPU.

Switch the clock to the maskdiv type.  With least-masking preference
the vendor OPP set now resolves to 600/400/300/200 MHz taken undivided
from their periph outputs and 150 MHz = pll-periph0-200M * 12/16, all
verified exact on hardware with the same cycle-counter method.

Drop pll-periph0-800M from the selectable parents (the mux table skips
hardware index 1): the vendor BSP removed it from its parent list with
the comment "If GPU use pll-peri0-800m, gpu will occur job fault", and
with the masking semantics every vendor OPP matches exactly from the
800M parent first, so it would otherwise always be chosen.

Also drop CLK_SET_RATE_PARENT: every OPP is reachable from the fixed
pll-periph0 outputs, and pll-gpu must never be reprogrammed through this
mux.  Once the GPU moves off pll-gpu the PLL is no longer prepared, so
it loses the rate protection of CLK_SET_RATE_GATE; a propagated rate
request would then reprogram the PLL while its gate is off (the lock
bit never asserts, 70 ms poll timeout per transition) and switch the
running GPU onto it before it locks.

Fixes: 6702d17f54a8 ("clk: sunxi-ng: a523: add video mod clocks")
Signed-off-by: Juan Manuel López Carrillo <juanmanuellopezcarrillo@gmail.com>
---
 drivers/clk/sunxi-ng/ccu-sun55i-a523.c | 32 +++++++++++++++++++++-----
 1 file changed, 26 insertions(+), 6 deletions(-)
  

Comments

sashiko-bot@kernel.org Aug. 3, 2026, 6:25 p.m. UTC | #1
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:

New issues:
- [High] Integer truncation in `ccu_maskdiv_find_m()` causes the hardware to drift lower than the framework's requested rate when the parent rate is not a perfect multiple of the cycle width (16). When `determine_rate` computes a new rate using `new_rate = (parent_rate * (16 - m)) / 16`, integer division truncates the remainder. When `set_rate` subsequently attempts to recover `m` by computing `kept = (new_rate * 16) / parent_rate`, the loss of the remainder strictly reduces `kept` by 1. This causes `set_rate` to program `m_found = m + 1` into the hardware, masking an additional cycle and driving the clock at a lower rate than what was reported and committed by `determine_rate`.

Pre-existing issues:
- [High] The newly introduced sparse `gpu_mux_table` triggers a pre-existing flaw in `ccu_mux_helper_get_parent()`. If the vendor bootloader leaves the GPU mux at hardware index 1 (`pll-periph0-800M`), `ccu_mux_helper_get_parent()` fails to find `1` in the sparse table (`{ 0, 2, 3, 4, 5 }`) and falls back to returning the raw hardware index `1`. The clock framework interprets this `1` as a software array index and incorrectly maps it to `gpu_parents[1]` (`pll_periph0_600M_clk`). If the kernel subsequently requests a 600MHz GPU rate, the framework falsely assumes the mux is already on the correct parent, skipping the mux switch and only setting the divider. This causes the hardware to remain at the undivided 800MHz parent, triggering the exact GPU job faults this patch intends to fix.
- [High] `pll_gpu_clk` is defined with the `CLK_SET_RATE_GATE` flag, which requires the clock to be unprepared (gated) before its rate can be changed. However, as noted in the commit message, reprogramming this PLL while its gate is off causes a 70ms lock timeout because the lock bit never asserts when the clock is disabled. This means any rate change on `pll_gpu` while it is gated will stall the kernel for 70ms. The patch works around this for the `gpu` downstream consumer by dropping `CLK_SET_RATE_PARENT`, but `pll_gpu_clk` itself remains incorrectly modeled. It should use `CLK_SET_RATE_UNGATE` to inform the framework that the clock must be powered on to change its rate.
--

--- Patch [2]: [PATCH v2 2/3] clk: sunxi-ng: sun55i-a523: GPU clock divider is fractional, not linear ---
commit 5635b9edbbd86c7ae0bef62dc457d0e280dc656a
Author: Juan Manuel López Carrillo <juanmanuellopezcarrillo@gmail.com>

clk: sunxi-ng: sun55i-a523: GPU clock divider is fractional, not linear

This patch changes the GPU clock modeling on sun55i-a523 from a linear divider
to a fractional cycle-masking divider. It updates the parent list, removes
the 800M parent due to hardware faults, and drops CLK_SET_RATE_PARENT to
prevent reprogramming the GPU PLL while it is gated.

> Also drop CLK_SET_RATE_PARENT: every OPP is reachable from the fixed
> pll-periph0 outputs, and pll-gpu must never be reprogrammed through this
> mux.  Once the GPU moves off pll-gpu the PLL is no longer prepared, so
> it loses the rate protection of CLK_SET_RATE_GATE; a propagated rate
> request would then reprogram the PLL while its gate is off (the lock
> bit never asserts, 70 ms poll timeout per transition) and switch the
> running GPU onto it before it locks.

[Severity: High]
This is a pre-existing issue, but does pll_gpu_clk use the wrong rate flag?

The commit message notes that the lock bit never asserts when the gate is
off, causing a 70ms timeout when reprogrammed. However, pll_gpu_clk is
declared in this file with CLK_SET_RATE_GATE:

drivers/clk/sunxi-ng/ccu-sun55i-a523.c:
    ...
                               CLK_SET_RATE_GATE),
    ...

Using CLK_SET_RATE_GATE enforces that the clock must be gated (unprepared)
before its rate can be changed. Since the hardware requires the PLL to be
ungated to lock, this contradicts the hardware's requirements.

Should pll_gpu_clk be updated to use CLK_SET_RATE_UNGATE instead, to inform
the framework that the clock must be powered on to change its rate?

> diff --git a/drivers/clk/sunxi-ng/ccu-sun55i-a523.c b/drivers/clk/sunxi-ng/ccu-sun55i-a523.c
> index 20dad06b37ca3..979e53e635228 100644
> --- a/drivers/clk/sunxi-ng/ccu-sun55i-a523.c
> +++ b/drivers/clk/sunxi-ng/ccu-sun55i-a523.c

[ ... ]

> @@ -442,18 +443,37 @@ static SUNXI_CCU_GATE_HWS(bus_g2d_clk, "bus-g2d", ahb_hws, 0x63c, BIT(0), 0);
>  
>  static const struct clk_hw *gpu_parents[] = {
>  	&pll_gpu_clk.common.hw,
> -	&pll_periph0_800M_clk.common.hw,
>  	&pll_periph0_600M_clk.hw,
>  	&pll_periph0_400M_clk.hw,
>  	&pll_periph0_300M_clk.hw,
>  	&pll_periph0_200M_clk.hw,
>  };
>  
> -static SUNXI_CCU_M_HW_WITH_MUX_GATE(gpu_clk, "gpu", gpu_parents, 0x670,
> -				    0, 4,	/* M */
> -				    24, 3,	/* mux */
> -				    BIT(31),	/* gate */
> -				    CLK_SET_RATE_PARENT);
> +/*
> + * Mux index 1 (pll-periph0-800M) is skipped: the vendor BSP removed it
> + * from the parent list ("If GPU use pll-peri0-800m, gpu will occur job
> + * fault"), and with the masking divider every OPP would match exactly
> + * from it first.
> + */
> +static const u8 gpu_mux_table[] = { 0, 2, 3, 4, 5 };

[Severity: High]
This is a pre-existing issue, but does skipping index 1 here expose a flaw
in ccu_mux_helper_get_parent()?

If the bootloader leaves the GPU mux at hardware index 1, the helper fails
to find it in this sparse table and falls back to returning the raw index:

drivers/clk/sunxi-ng/ccu_mux.c:ccu_mux_helper_get_parent() {
    ...
        for (i = 0; i < num_parents; i++)
            if (cm->table[i] == parent)
                return i;
    }

    return parent;
}

The clock framework then interprets this returned 1 as a software array
index and incorrectly maps it to gpu_parents[1] (which is now
pll_periph0_600M_clk).

If the kernel subsequently requests a 600MHz GPU rate, the framework
might falsely assume the mux is already on the correct parent and skip the
switch. This would leave the hardware at the undivided 800MHz parent,
triggering the exact GPU job faults this patch intends to fix.

> +
> +/*
> + * The M factor is a cycle-masking (fractional) divider, not a linear
> + * one: rate = source * (16 - M) / 16 (T527 manual, GPU_CLK_REG).
> + *
> + * No CLK_SET_RATE_PARENT: every GPU OPP is reachable from the fixed
> + * pll-periph0 outputs, and pll-gpu must never be reprogrammed through this mux.
> + * Once the GPU moves off pll-gpu the PLL is no longer prepared, so it loses
> + * the rate protection of CLK_SET_RATE_GATE; a propagated rate request would
> + * then reprogram the PLL while its gate is off (the lock bit never asserts,
> + * 70 ms timeout) and switch the running GPU onto it before it locks.
> + */
> +static SUNXI_CCU_MASKDIV_HW_WITH_MUX_TABLE_GATE(gpu_clk, "gpu", gpu_parents,
> +						gpu_mux_table, 0x670,
> +						0, 4,	/* M */
> +						24, 3,	/* mux */
> +						BIT(31),	/* gate */
> +						0);

[Severity: High]
Does using the maskdiv implementation here cause the hardware to drift
lower than the requested rate when the parent rate is not evenly divisible
by 16?

In ccu_maskdiv_find_m(), integer truncation occurs when recovering the
multiplier:

drivers/clk/sunxi-ng/ccu_maskdiv.c:ccu_maskdiv_find_m() {
    ...
    kept = div64_ul((u64)rate * n, parent_rate);
    ...
}

When determine_rate computes new_rate = (parent_rate * (16 - m)) / 16,
the remainder is truncated. When set_rate subsequently attempts to recover
m using the truncated rate:

drivers/clk/sunxi-ng/ccu_maskdiv.c:ccu_maskdiv_set_rate() {
    ...
    m = ccu_maskdiv_find_m(parent_rate, rate, cmd->width);
    ...
}

The loss of the remainder strictly reduces kept by 1. This causes set_rate
to program m + 1 into the hardware, masking an additional cycle and driving
the clock at a lower rate than what was reported by determine_rate.

Could DIV_ROUND_UP_ULL be used to prevent this truncation?
  

Patch

diff --git a/drivers/clk/sunxi-ng/ccu-sun55i-a523.c b/drivers/clk/sunxi-ng/ccu-sun55i-a523.c
index 20dad06b37ca..979e53e63522 100644
--- a/drivers/clk/sunxi-ng/ccu-sun55i-a523.c
+++ b/drivers/clk/sunxi-ng/ccu-sun55i-a523.c
@@ -21,6 +21,7 @@ 
 
 #include "ccu_div.h"
 #include "ccu_gate.h"
+#include "ccu_maskdiv.h"
 #include "ccu_mp.h"
 #include "ccu_mult.h"
 #include "ccu_nk.h"
@@ -442,18 +443,37 @@  static SUNXI_CCU_GATE_HWS(bus_g2d_clk, "bus-g2d", ahb_hws, 0x63c, BIT(0), 0);
 
 static const struct clk_hw *gpu_parents[] = {
 	&pll_gpu_clk.common.hw,
-	&pll_periph0_800M_clk.common.hw,
 	&pll_periph0_600M_clk.hw,
 	&pll_periph0_400M_clk.hw,
 	&pll_periph0_300M_clk.hw,
 	&pll_periph0_200M_clk.hw,
 };
 
-static SUNXI_CCU_M_HW_WITH_MUX_GATE(gpu_clk, "gpu", gpu_parents, 0x670,
-				    0, 4,	/* M */
-				    24, 3,	/* mux */
-				    BIT(31),	/* gate */
-				    CLK_SET_RATE_PARENT);
+/*
+ * Mux index 1 (pll-periph0-800M) is skipped: the vendor BSP removed it
+ * from the parent list ("If GPU use pll-peri0-800m, gpu will occur job
+ * fault"), and with the masking divider every OPP would match exactly
+ * from it first.
+ */
+static const u8 gpu_mux_table[] = { 0, 2, 3, 4, 5 };
+
+/*
+ * The M factor is a cycle-masking (fractional) divider, not a linear
+ * one: rate = source * (16 - M) / 16 (T527 manual, GPU_CLK_REG).
+ *
+ * No CLK_SET_RATE_PARENT: every GPU OPP is reachable from the fixed
+ * pll-periph0 outputs, and pll-gpu must never be reprogrammed through this mux.
+ * Once the GPU moves off pll-gpu the PLL is no longer prepared, so it loses
+ * the rate protection of CLK_SET_RATE_GATE; a propagated rate request would
+ * then reprogram the PLL while its gate is off (the lock bit never asserts,
+ * 70 ms timeout) and switch the running GPU onto it before it locks.
+ */
+static SUNXI_CCU_MASKDIV_HW_WITH_MUX_TABLE_GATE(gpu_clk, "gpu", gpu_parents,
+						gpu_mux_table, 0x670,
+						0, 4,	/* M */
+						24, 3,	/* mux */
+						BIT(31),	/* gate */
+						0);
 
 static SUNXI_CCU_GATE_HWS(bus_gpu_clk, "bus-gpu", ahb_hws, 0x67c, BIT(0), 0);