Repository navigation
perf(ebpf): wake the proc_events reader per event instead of polling every 10 ms - #366
Merged
Merged
Conversation
…every 10 ms Perf readers are created with WakeupEvents 100: the kernel wakes a reader once a CPU's buffer holds 100 events, and below that the events wait for the reader's next deadline. Process events come a few at a time, so to act on an exec before the new program connects, the proc_events deadline was cut to 10 ms. Polling 100 times a second cost ~5 millicores on an idle node (agent CPU profile: the perf reader's epoll loop went from 0.57 s to 1.0 s per 90 s). proc_events is now created with WakeupEvents 1, so an exec wakes the reader at once, and its deadline is back to the default 100 ms.
There was a problem hiding this comment.
Code Review
This pull request introduces a wakeupEvents configuration for perfMap in ebpftracer/tracer.go to wake up the reader on every event for proc_events instead of polling, aiming to reduce idle CPU usage. However, feedback points out that because runEventsReader still defaults a zero readTimeout to 100 * time.Millisecond and sets a deadline on every iteration, the idle wakeups are not fully eliminated. It is recommended to update runEventsReader to only set a deadline if readTimeout > 0 and explicitly configure a readTimeout for other maps.
proc_events is created with WakeupEvents 1, but runEventsReader still gave it the default 100 ms deadline, so an idle node woke its reader 10 times a second for nothing. A reader the kernel wakes for every event now has no deadline; the others keep 100 ms unless set otherwise. Close still interrupts a blocked Read.
RamanKharchee
approved these changes
Oct 6, 2026
blue4209211
approved these changes
Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Perf readers are created with
WakeupEvents: 100: the kernel wakes a reader once a CPU's buffer holds 100 events, and below that the events wait for the reader's next deadline. Process events come a few at a time. So #353 cut theproc_eventsdeadline from 100 ms to 10 ms, to act on an exec before the new program connects. Polling 100 times a second cost about 5 millicores on an idle node.proc_eventsis now created withWakeupEvents: 1, so an exec wakes the reader at once. Because the kernel wakes it for every event, it has no read deadline at all;Closestill interrupts a blocked read. The other readers keep their deadlines: 100 ms by default, 10 ms fortcp_connect_events.This is the agent-CPU follow-up from the #353 review.
Engineering detail
Diagnosis: an idle CPU profile of the agent before and after fix(tls): count TLS capture losses and fix the gaps they exposed #353 shows the extra time entirely in the perf reader's epoll loop (
runEventsReader→ReadInto→EpollWait): 0.57 s → 1.0 s per 90 s.Scope: only
proc_eventschanges.tcp_connect_eventskeeps its 10 ms deadline, because connects can be frequent enough that a wakeup per event would cost more than the polling.CPU, from the agent's own cgroup: idle went from 22 to 13–15 millicores, against 18–19 before fix(tls): count TLS capture losses and fix the gaps they exposed #353. With ~3.3 short-lived Go CLI execs a second it went from 35 to 28–33.
Regression suite: 394 captured from 390 short-lived TLS clients (some make two requests; main: 388–390), same-path 131, wrapper+exec 21 (main: 17–18), no races.
Local e2e: built agent images from this branch and from main and ran each in a local Docker VM (kernel 6.10). I measured the agent's own cgroup CPU over a 180s idle window, then over 300s while a stripped Go CLI linking crypto/tls ran ~3.3 times a second. Idle: main 22, this branch 13–15 millicores. Under the CLI: main 35, this branch 28–33. Runs that overlapped my own builds in the same VM are excluded. The short-lived-process scenario suite is at parity or better, with no data races.