linux-stable.git/tools/perf/bench, branch v6.15

perf bench sched pipe: fix enforced blocking reads in worker_thread

2025-03-24T06:20:37+00:00

The function worker_thread() is programmed in a way that roughly
doubles the number of expectable context switches, because it enforces
blocking reads:

 Performance counter stats for 'perf bench sched pipe':

         2,000,004      context-switches

      11.859548321 seconds time elapsed

       0.674871000 seconds user
       8.076890000 seconds sys

The result of this behavior is that the blocking reads by far dominate
the performance analysis of 'perf bench sched pipe':

Samples: 78K of event 'cycles:P', Event count (approx.): 27964965844
Overhead  Command     Shared Object         Symbol
  25.28%  sched-pipe  [kernel.kallsyms]     [k] read_hpet
   8.11%  sched-pipe  [kernel.kallsyms]     [k] retbleed_untrain_ret
   2.82%  sched-pipe  [kernel.kallsyms]     [k] pipe_write

From the code, it is unclear if that behavior is wanted but the log
says that at least Ingo Molnar aims to mimic lmbench's lat_ctx, that
doesn't handle the pipe ends that way
(https://sourceforge.net/p/lmbench/code/HEAD/tree/trunk/lmbench2/src/lat_ctx.c)

Fix worker_thread() by always first feeding the write ends of the pipes
and then trying to read.

This roughly halves the context switches and runtime of pure
'perf bench sched pipe':

 Performance counter stats for 'perf bench sched pipe':

         1,005,770      context-switches

       6.033448041 seconds time elapsed

       0.423142000 seconds user
       4.519829000 seconds sys

And the blocking reads do no longer dominate the analysis at the above
extreme:

Samples: 40K of event 'cycles:P', Event count (approx.): 14309364879
Overhead  Command     Shared Object         Symbol
  12.20%  sched-pipe  [kernel.kallsyms]     [k] read_hpet
   9.23%  sched-pipe  [kernel.kallsyms]     [k] retbleed_untrain_ret
   3.68%  sched-pipe  [kernel.kallsyms]     [k] pipe_write

Signed-off-by: Dirk Gouders 
Acked-by: Ingo Molnar 
Link: https://lore.kernel.org/r/20250323140316.19027-2-dirk@gouders.net
Signed-off-by: Namhyung Kim

perf bench: Fix perf bench syscall loop count

2025-03-05T17:19:23+00:00

Command 'perf bench syscall fork -l 100000' offers option -l to run for
a specified number of iterations. However this option is not always
observed. The number is silently limited to 10000 iterations as can be
seen:

Output before:
 # perf bench syscall fork -l 100000
 # Running 'syscall/fork' benchmark:
 # Executed 10,000 fork() calls
     Total time: 23.388 [sec]

    2338.809800 usecs/op
            427 ops/sec
 #

When explicitly specified with option -l or --loops, also observe
higher number of iterations:

Output after:
 # perf bench syscall fork -l 100000
 # Running 'syscall/fork' benchmark:
 # Executed 100,000 fork() calls
     Total time: 716.982 [sec]

    7169.829510 usecs/op
            139 ops/sec
 #

This patch fixes the issue for basic execve fork and getpgid.

Fixes: ece7f7c0507c ("perf bench syscall: Add fork syscall benchmark")
Signed-off-by: Thomas Richter 
Acked-by: Sumanth Korikkar 
Tested-by: Athira Rajeev 
Cc: Tiezhu Yang 
Link: https://lore.kernel.org/r/20250304092349.2618082-1-tmricht@linux.ibm.com
Signed-off-by: Namhyung Kim

perf bench: Fix undefined behavior in cmpworker()

2025-01-18T18:14:36+00:00

The comparison function cmpworker() violates the C standard's
requirements for qsort() comparison functions, which mandate symmetry
and transitivity:

Symmetry: If x < y, then y > x.
Transitivity: If x < y and y < z, then x < z.

In its current implementation, cmpworker() incorrectly returns 0 when
w1->tid < w2->tid, which breaks both symmetry and transitivity. This
violation causes undefined behavior, potentially leading to issues such
as memory corruption in glibc [1].

Fix the issue by returning -1 when w1->tid < w2->tid, ensuring
compliance with the C standard and preventing undefined behavior.

Link: https://www.qualys.com/2024/01/30/qsort.txt [1]
Fixes: 121dd9ea0116 ("perf bench: Add epoll parallel epoll_wait benchmark")
Cc: stable@vger.kernel.org
Signed-off-by: Kuan-Wei Chiu 
Reviewed-by: James Clark 
Link: https://lore.kernel.org/r/20250116110842.4087530-1-visitorckw@gmail.com
Signed-off-by: Namhyung Kim

perf bench: Remove reference to cmd_inject

2024-12-18T19:24:33+00:00

Avoid `perf bench internals inject-build-id` referencing the
cmd_inject sub-command that requires perf-bench to backward reference
internals of builtins. Replace the reference to cmd_inject with a call
to main. To avoid python.c needing to link with something providing
main, drop the libperf-bench library from the python shared object.

Signed-off-by: Ian Rogers 
Tested-by: Arnaldo Carvalho de Melo 
Cc: Adrian Hunter 
Cc: Alexander Shishkin 
Cc: Andi Kleen 
Cc: Athira Rajeev 
Cc: Colin Ian King 
Cc: Dapeng Mi 
Cc: Howard Chu 
Cc: Ilya Leoshkevich 
Cc: Ingo Molnar 
Cc: James Clark 
Cc: Jiri Olsa 
Cc: Josh Poimboeuf 
Cc: Kan Liang 
Cc: Mark Rutland 
Cc: Michael Petlan 
Cc: Namhyung Kim 
Cc: Peter Zijlstra 
Cc: Thomas Richter 
Cc: Veronika Molnarova 
Cc: Weilin Wang 
Link: https://lore.kernel.org/r/20241119011644.971342-17-irogers@google.com
Signed-off-by: Arnaldo Carvalho de Melo

perf header: Move is_cpu_online to numa bench

2024-11-16T19:36:47+00:00

The helper function is only used in the NUMA benchmark as typically
online CPUs are determined through perf_cpu_map__new_online_cpus().

Reduce the scope of the function for now.

Reviewed-by: James Clark 
Signed-off-by: Ian Rogers 
Tested-by: Xu Yang 
Cc: Adrian Hunter 
Cc: Albert Ou 
Cc: Alexander Shishkin 
Cc: Alexandre Ghiti 
Cc: Athira Rajeev 
Cc: Ben Zong-You Xie 
Cc: Benjamin Gray 
Cc: Bibo Mao 
Cc: Clément Le Goffic 
Cc: Dima Kogan 
Cc: Dr. David Alan Gilbert 
Cc: Huacai Chen 
Cc: Ingo Molnar 
Cc: Jiri Olsa 
Cc: John Garry 
Cc: Kan Liang 
Cc: Leo Yan 
Cc: Mark Rutland 
Cc: Masami Hiramatsu 
Cc: Mike Leach 
Cc: Namhyung Kim 
Cc: Palmer Dabbelt 
Cc: Paul Walmsley 
Cc: Peter Zijlstra 
Cc: Ravi Bangoria 
Cc: Sandipan Das 
Cc: Will Deacon 
Cc: Yicong Yang 
Cc: linux-arm-kernel@lists.infradead.org
Cc: linux-riscv@lists.infradead.org
Link: https://lore.kernel.org/r/20241107162035.52206-3-irogers@google.com
Signed-off-by: Arnaldo Carvalho de Melo

perf tools: sched-pipe bench: add (-n) nonblocking benchmark

2024-10-22T04:23:01+00:00

The -n mode will benchmark pipes in a non-blocking mode using
epoll_wait.

This specific mode was added to demonstrate the broken sync nature
of epoll: https://lore.kernel.org/lkml/20240426-zupfen-jahrzehnt-5be786bcdf04@brauner

Signed-off-by: Brian Geffon 
Reviewed-by: Ian Rogers 
Cc: Steven Rostedt 
Link: https://lore.kernel.org/r/20241016190009.866615-1-bgeffon@google.com
Signed-off-by: Namhyung Kim

perf tool: Constify tool pointers

2024-08-12T21:05:14+00:00

The tool pointer (to a struct largely of function pointers) is passed
around but is unchanged except at initialization. Change parameter and
variable types to be const to lower the possibilities of what could
happen with a tool.

Reviewed-by: Adrian Hunter 
Signed-off-by: Ian Rogers 
Tested-by: Adrian Hunter 
Tested-by: Leo Yan 
Cc: Alexander Shishkin 
Cc: Anshuman Khandual 
Cc: Athira Rajeev 
Cc: Huacai Chen 
Cc: Ilkka Koskinen 
Cc: Ingo Molnar 
Cc: James Clark 
Cc: Jiri Olsa 
Cc: John Garry 
Cc: Jonathan Cameron 
Cc: Kan Liang 
Cc: Leo Yan 
Cc: Mark Rutland 
Cc: Mike Leach 
Cc: Namhyung Kim 
Cc: Nick Desaulniers 
Cc: Nick Terrell 
Cc: Oliver Upton 
Cc: Peter Zijlstra 
Cc: Song Liu 
Cc: Sun Haiyong 
Cc: Suzuki Poulouse 
Cc: Will Deacon 
Cc: Yanteng Si 
Cc: Yicong Yang 
Cc: linux-arm-kernel@lists.infradead.org
Link: https://lore.kernel.org/r/20240812204720.631678-4-irogers@google.com
Signed-off-by: Arnaldo Carvalho de Melo

perf bench: Make bench its own library

2024-06-26T18:07:28+00:00

Make the benchmark code into a library so it may be linked against
things like the python module to avoid compiling code twice.

Signed-off-by: Ian Rogers 
Reviewed-by: James Clark 
Cc: Suzuki K Poulose 
Cc: Kees Cook 
Cc: Palmer Dabbelt 
Cc: Albert Ou 
Cc: Nick Terrell 
Cc: Gary Guo 
Cc: Alex Gaynor 
Cc: Boqun Feng 
Cc: Wedson Almeida Filho 
Cc: Ze Gao 
Cc: Alice Ryhl 
Cc: Andrei Vagin 
Cc: Yicong Yang 
Cc: Jonathan Cameron 
Cc: Guo Ren 
Cc: Miguel Ojeda 
Cc: Will Deacon 
Cc: Mike Leach 
Cc: Leo Yan 
Cc: Oliver Upton 
Cc: John Garry 
Cc: Benno Lossin 
Cc: Björn Roy Baron 
Cc: Andreas Hindborg 
Cc: Paul Walmsley 
Signed-off-by: Namhyung Kim 
Link: https://lore.kernel.org/r/20240625214117.953777-6-irogers@google.com

tools/perf: Fix timing issue with parallel threads in perf bench wake-up-parallel

2024-06-14T04:27:49+00:00

perf bench futex fails as below and hangs intermittently when
attempted to run on on a powerpc system:

./perf bench futex wake-parallel
 Running 'futex/wake-parallel' benchmark:
 Run summary [PID 88588]: blocking on 640 threads (at [private] futex 0x10464b8c), 640 threads waking up 1 at a time.

[Run 1]: Avg per-thread latency (waking 1/640 threads) in 0.1309 ms (+-53.27%)
[Run 2]: Avg per-thread latency (waking 1/640 threads) in 0.0120 ms (+-31.16%)
[Run 3]: Avg per-thread latency (waking 1/640 threads) in 0.1474 ms (+-92.47%)
[Run 4]: Avg per-thread latency (waking 1/640 threads) in 0.2883 ms (+-67.75%)
[Run 5]: Avg per-thread latency (waking 1/640 threads) in 0.4108 ms (+-39.60%)
[Run 6]: Avg per-thread latency (waking 1/640 threads) in 0.7843 ms (+-78.98%)
perf: couldn't wakeup all tasks (0/1)
perf: couldn't wakeup all tasks (0/1)
perf: couldn't wakeup all tasks (0/1)
perf: couldn't wakeup all tasks (0/1)
perf: couldn't wakeup all tasks (0/1)
perf: couldn't wakeup all tasks (0/1)

In the system, where perf bench wake-up-parallel is has system
configuration of 640 cpus. After debugging, this turned out to be
a timing issue. The benchmark creates threads equal to number of
cpus and issues a futex_wait. Then it does a usleep for .1 second
before initiating futex_wake. In system configuration with more
threads, the usleep time is not enough. Patch changes the usleep
from 100000 to 200000

With the patch, ran multiple iterations and there were no issues
further seen

Reported-by: Disha Goel 
Signed-off-by: Athira Rajeev 
Reviewed-by: Ian Rogers 
Tested-by: Disha Goel 
Cc: akanksha@linux.ibm.com
Cc: kjain@linux.ibm.com
Cc: maddy@linux.ibm.com
Cc: linuxppc-dev@lists.ozlabs.org
Signed-off-by: Namhyung Kim 
Link: https://lore.kernel.org/r/20240607044354.82225-3-atrajeev@linux.vnet.ibm.com

tools/perf: Fix perf bench epoll to enable the run when some CPU's are offline

2024-06-14T04:27:26+00:00

Perf bench epoll fails as below when attempted to run on
on a powerpc system:

   ./perf bench epoll wait
   Running 'epoll/wait' benchmark:
   Run summary [PID 627653]: 79 threads monitoring on 64 file-descriptors for 8 secs.

   perf: pthread_create: No such file or directory

In the setup where this perf bench was ran, difference was that
partition had 640 CPU's, but not all CPUs were online. 80 CPUs
were online. While creating threads and using epoll_wait , code
sets the affinity using cpumask. The cpumask size used is 80
which is picked from "nrcpus = perf_cpu_map__nr(cpu)". Here the
benchmark reports fail while setting affinity for cpu number which
is greater than 80 or higher, because it attempts to set a bit
position which is not allocated on the cpumask. Fix this by changing
the size of cpumask to number of possible cpus and not the number
of online cpus.

Signed-off-by: Athira Rajeev 
Reviewed-by: Ian Rogers 
Tested-by: Disha Goel 
Cc: akanksha@linux.ibm.com
Cc: kjain@linux.ibm.com
Cc: maddy@linux.ibm.com
Cc: linuxppc-dev@lists.ozlabs.org
Signed-off-by: Namhyung Kim 
Link: https://lore.kernel.org/r/20240607044354.82225-2-atrajeev@linux.vnet.ibm.com