auditd syscall rules on a KVM host: 12x CPU overhead while guests idle
At 10,000 feet
Setting: a Linux KVM hypervisor running the audit daemon (auditd) for
security logging. Each guest vCPU is an ordinary host thread calling into the
kernel through the KVM interface.
Problem: audit rules logging process execution (execve) made one VM’s host
process burn ~1200% CPU (~12 cores) and ~100W more at the wall while the
guest inside sat idle at 0.8%.
What this note establishes: auditd syscall/exit rules bill per system
call, and a vCPU issues millions of ioctls per second, so the accounting
explodes. eBPF on the execve tracepoint gives the same visibility for
~0% CPU. It fires per exec (hundreds per hour), not per syscall.
Takeaway: never put auditd syscall/exit execve rules on a KVM hypervisor.
Use an eBPF execve tracer, keeping auditd only for cheap file-integrity -w
watches.
Notes. Measured on a Linux KVM hypervisor host with an otherwise-idle guest. The mechanism (per-syscall accounting vs. per-exec tracing) is standard Linux auditing behavior. No warranty, corrections welcome.
Why per-syscall auditing explodes on a hypervisor
An auditd exit rule (-a exit,always -S execve style) does bookkeeping on
every matching syscall. Each guest vCPU traps into the host kernel via
ioctl(KVM_RUN) whenever the guest needs mediation: millions of ioctls per
second per vCPU. The cost lands on the host, attributed to the VM’s process,
invisible from inside the guest, hence the phantom time. Flushing the rules
dropped host CPU to 9% instantly (measured), isolating them as the cause.
An eBPF program on the execve tracepoint (e.g. bpftrace on
syscalls:sys_enter_execve) bills per-exec, not per-syscall: a structural
difference, not a tuning win. auditd -w file watches hook filesystem events,
not the syscall path, and stay cheap.
Honesty: measured vs. mechanism
- Measured: ~1200% CPU, ~100W, the 0.8% idle guest, the drop to 9%, the ~0% eBPF cost.
- Mechanism: vCPU
ioctlvolume feeding per-syscall accounting is standard documented behavior. The contribution is the magnitude on a hypervisor, and that the fix is a different tool, not a quieter rule.
Licensed under the site footer’s CC BY 4.0. If this saved you a debugging session, the optional thanks link in the footer is appreciated, no obligation.
Validation (2026-08-31)
On a production multi-socket Linux/KVM host (live guests, baseline -w watches
loaded): a static C binary looping ioctl(fd, FIONREAD) on /dev/null, timed
via getrusage, without and then with one narrow rule (-a always,exit -F
arch=b64 -S ioctl -F exe=<test-binary>) auditing only the test process. Ruleset
byte-identical afterward. Audit lost counter 0.
| condition | per-ioctl wall |
per-ioctl process CPU |
|---|---|---|
| FIM watches only (baseline) | 631–642 ns | 631–641 ns |
+ narrow always,exit -S ioctl rule |
26,217–27,624 ns | 7,415–8,391 ns |
- ~12x process-CPU, ~42x wall per audited syscall. The gap is record emission
and backlog throttling billed to
kauditd/auditd, so system-wide cost exceeds the 12x the process sees. - ~7 µs added per syscall × ~1M+
ioctls/sec per busy vCPU = multiple cores of pure audit overhead per VM, the field incident’s regime. - Not re-run: the ~1200% / ~100W figure (needs the busy-guest state and wall-power metering). It stands as the original field measurement.
Verdict: reproduced. Per-syscall cost and scaling confirmed, from which the hypervisor exposure follows.