Profile
Back to NewsBack
GitHub Trending 3 min
Reader Mode
perf-labs/perf: perf-labs/perf

perf-labs/perf: perf-labs/perf

11 hours ago

perf-labs/perf

profile → benchmark → analyze → optimize ↻

Features

  • Benchmark assembly snippets, functions and regions of any binary with hardware performance counters and CPU state 'control'.
  • Pass a program's own arguments, so a main is measured the way a shell calls it.
  • Profile live binaries without re-compiling.
  • Analyze machine code together with measured data.
  • Compare data via central limit theorem and null-hypothesis tests.
  • View/Plot in the terminal, via interactive mode or jupyter notebooks.

Documentation

| Name | Description | | ---- | ---- | | CLI reference | Command line options, configuration | | Python API | Architecture, synopsis | | Code Markers | Zero-instruction code (C, C++, Rust, Zig) markers | | Skills | The Human–Agent performance engineering loop | | Studies | Top-down Microarchitecture Analysis Method studies: retiring, bad-speculation, frontend-bound, backend-bound | | How it works | Symbolic exploration, data synthesis, JIT harness, CPU 'control' (cache/branch/...), code patching, performance hardware counters | | References | Specifcations, publications, manuals |

Requirements

  • x86-64 Linux, kernel 6.x+
  • Python 3.11+ (see pyproject.toml)
  • linux-perf and user-space rdpmc access (see Setup)

Install

pip install git+https://github.com/perf-labs/perf.git

See Development.

Setup

echo 2 | sudo tee /sys/devices/{cpu_core,cpu_atom}/rdpmc

This enables user-space RDPMC hardware counters without syscalls on the hot path. This is not required for duration_time (Time-Stamp Counter).

Quick start

# benchmark
perf benchmark 'imul eax, 0' -m latency -e cycles
perf benchmark snippet.s:label
perf benchmark a.out:func -e 'topdown-*'
perf benchmark a.out:0x401000..0x401020
perf benchmark a.out:func -m throughput -e cycles,instructions
perf benchmark a.out:hot_begin..hot_end # see code markers
perf benchmark /usr/bin/tree:main -m latency -- .

explore

perf benchmark a.out:func --data.rdi=15 perf benchmark a.out:func --config.branch=predictable perf benchmark a.out:func --config.dcache=cold -e cache-misses,cycles perf benchmark 'mov rax, [rdi]' --data.rdi=0x42000000000 --data[0x42000000000]=123

analyze

perf analyze a.out:func --filter 'latency > 4' perf analyze a.out:func -e 'index,assembly,data*' perf analyze a.out:func -e assembly | llvm-mca perf analyze a.out:func --filter '15 in data.rdi' perf analyze a.out:func -- perf.data profile.json

profile

perf profile -f begin..end -e cycles -o profile.json -- ./a.out

view/plot

perf benchmark a.out:func | perf view -s p99 perf benchmark a.out:func | perf plot -t ecdf

compare # Central Limit Theorem -> null-hypothesis test

perf benchmark clang.out:func_v1 --output data/ perf benchmark gcc.out:func_v2 --output data/ perf compare --baseline func_v1 -- data/ perf plot -e instructions/cycles -- data/

See CLI reference.

import perf

df = perf.benchmark(["a.out", "func"]) df.duration_time.describe()

df = perf.benchmark(code="mov eax, 42", mode=["latency"], event=["cycles"]) df.cycles.plot()

df = perf.benchmark(code=["/usr/bin/tree", "main"], mode=["latency"], argv=["."]) perf.to_json("tree.json", df)

See Python API.

Cite

@software{jusiak2026perf,
  author  = {Kris Jusiak},
  title   = {perf},
  year    = {2026},
  url     = {https://github.com/perf-labs/perf}
}
Chat with me