On 9/24/26 11:10, [email protected] wrote:
Thank you for your contribution! Sashiko AI review found 1 potential issue(s)
to consider:
- [Medium] ftrace: time_max TOCTOU data race clobbers maximum latency spikes
--
commit 9fb82451733a5ec95aac95499945b5c01e39d82b
Author: Yun Zhou <[email protected]>
ftrace: add max field to function profiler stats
This commit adds a max field to the function profiler statistics to record the
maximum single-call duration. This allows users to better investigate
occasional latency spikes when reviewing function execution times.
diff --git a/kernel/trace/ftrace.c b/kernel/trace/ftrace.c
index 53d5db60bfa56..dfd25e3c87f71 100644
--- a/kernel/trace/ftrace.c
+++ b/kernel/trace/ftrace.c
[ ... ]
@@ -920,6 +923,8 @@ static void profile_graph_return(struct ftrace_graph_ret
*trace,
if (rec) {
rec->time += calltime;
rec->time_squared += calltime * calltime;
+ if (calltime > rec->time_max)
+ rec->time_max = calltime;
[Severity: Medium]
Could this non-atomic check and update open a time-of-check to time-of-use
data race against concurrent nested interrupt context tracing on the same CPU?
Since profile_graph_return() is executed with local interrupts enabled, an
interrupt or NMI could fire exactly between the condition check and the
assignment. If the interrupt handler traces the identical function and
records a massive latency spike, it would update rec->time_max. When the
interrupt returns, the normal context resumes and executes the assignment,
overwriting the true maximum latency spike with its smaller value.
Because ftrace explicitly permits tracing recursion from different contexts
like hardirq or NMI, does this require an atomic compare-and-exchange loop
or similar protection to avoid dropping the genuine latency spikes this
patch intends to capture?
I think this is intentional and consistent with the profiler's design.
profile_graph_return() runs under preempt_notrace() and the existing
time, time_squared and counter updates are all likewise non-atomic best-
effort accounting; the recursion guard only protects record allocation,
not the stat updates. This is a hot path - every traced function return
passes through it, and keeping the observation overhead minimal is more
important here than never dropping a single sample.