perf record Turns a Mystery CPU Spike Into a Function Name
The alert said CPU saturation on one node. Our APM showed the usual functions, nothing dominant. The flame graph from the agent looked like a flat city skyline with no towers. Yet the machine was pinned and requests were slowing.
The gap was that our agent sampled at a polite rate and only instrumented application frames. The thing eating the CPU was spending its time where the agent barely looked. Thirty seconds of perf record found it.
perf is the Linux kernel's profiler. It samples the hardware performance counters, so it sees everything the CPU is doing, in your code, in libraries and in the kernel, with near zero overhead. You do not need to restart anything or add instrumentation.
This was Ubuntu 24.04 with perf from the linux-tools package, on a Node 22.14 service and a Go 1.23 service. Both are covered.
The thirty second recipe
Record for a fixed window while the spike is happening:
sudo perf record -F 99 -g -p $(pidof node) -- sleep 30
| Flag | Meaning |
|---|---|
-F 99 |
99 samples per second. Odd number avoids lockstep with timers |
-g |
Capture call stacks, not just leaves |
-p PID |
Limit to the suspect process, keeps the output small |
-- sleep 30 |
Stop after thirty seconds |
Then look at it:
sudo perf report --stdio | head -40
That prints the hottest stacks in order. The top entry is the function the CPU was in when sampled most often, with its full call chain because of -g.
For a shareable picture, turn it into a flame graph. The classic tooling is Brendan Gregg's FlameGraph scripts:
sudo perf script | ./stackcollapse-perf.pl | ./flamegraph.pl > flame.svg
The wide plateaus in the SVG are your CPU. You are looking for a tower you did not expect.
What it finds that your APM will not
Kernel and syscall time. If your service is spending its CPU in page faults, in TLS handshakes, in memcpy for a huge buffer, or in a spinlock, an application level profiler shows a thin or missing frame. perf shows the kernel stack underneath your call. I found a service pinned by clear_page because it was allocating and freeing huge buffers per request. The app profiler showed nothing unusual.
Native code in your runtime. Regex engines, JSON parsers, crypto. These run as native code inside the runtime. perf attributes them correctly and shows you which regex or which parse is hot.
Compiler output surprises. Sometimes the hot function is one you wrote but that the compiler turned into something unexpected, or an inlined loop that dominates. The stack tells you the source frame and the leaf tells you the reality.
Off CPU versus on CPU. The default perf records on CPU time, which is what you want for a saturation alert. If instead the mystery is latency without CPU, you want off CPU tracing, which is the domain of bpftrace, covered in ten bpftrace one liners.
Reading the output like a diagnosis
The report gives you a percentage. Treat it as the answer to "what fraction of CPU was spent here". Two shapes matter.
A single tall tower means one function dominates. You have found it. Go read that code. This is the satisfying case and it is more common than you would think.
A flat skyline means the CPU is spread across many functions, which usually means the workload genuinely grew rather than a single defect appeared. That points you at capacity or traffic, not at a code change. Knowing which of the two you have is itself the diagnosis.
The container caveat
perf reads hardware counters, and in a container those counters are per host CPU, not per cgroup. So two caveats.
First, sampling a container process with -p works fine and is what you want. But a system wide perf record -a in a container sees the whole host's CPUs, which is noise and also a security boundary you should not be crossing.
Second, the percentage in perf report is a fraction of the samples for your process, which is what you care about. But do not try to reconcile it with cgroup CPU accounting directly. They measure different things.
You also need the right permissions. In many locked down environments perf_event_paranoid is set high, and you will need either root or a capability. Check:
cat /proc/sys/kernel/perf_event_paranoid
A value of 1 or lower generally lets you profile your own processes without root.
The languages
For Node and Go, perf sees native stacks, but symbolication of JITted or Go frames needs a little help.
Go symbols are usually present in the binary, so perf report on a Go process is already readable. Add --call-graph dwarf if frame pointers are omitted.
For Node, the JIT means your JavaScript functions appear as anonymous native code unless you expose symbols. The modern fix is node --perf-basic-prof plus perf's JIT support, or, more simply, use the built in node --prof and --prof-process for the application layer and keep perf for the native and kernel layer. The two together cover the whole stack.
The rule
When CPU is the symptom, start with the kernel's own profiler before trusting any agent. It has the least bias, the widest view and no sampling budget to protect. Thirty seconds of recording during the spike, then read the top of the report.
Your APM is for trends and for the application layer. perf is for the moment you need to know, precisely, what the CPU was doing. They answer different questions, and the alert you are chasing is usually perf's question.
If the top of the report is a function you recognise as cheap, the next step is to look at the allocation and buffer patterns around it, which is the heap and snapshot work in Node memory leaks.