You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Port upstream coroot-node-agent v1.29.0 → v1.36.2 #369
Everything up to upstream v1.28.3 is already here.Merge upstream coroot/coroot-node-agent changes #201 brought it in as a squash, so git doesn't know: git merge-base still points at 2025-06-03. That's why a plain git merge upstream/main conflicts everywhere.
49 upstream commits since v1.28.3. Each was cherry-picked on its own onto main and the conflict hunks counted, ignoring the generated ebpftracer/ebpf.go. The counts are in the tables below.
Upstream's Windows refactor (coroot/coroot-node-agent@34de3a0) moved files. It split flags/, node/ and gpu/ into *_linux.go / *_windows.go. Later upstream commits that touch flags/flags_linux.go or windows/* need their paths remapped by hand.
The upstream log commits need logparser ≥ v1.4.2. Our logparser fork is 6 upstream commits behind.
How to work through it
One PR per batch, in order.
git cherry-pick -x so every commit message keeps the upstream SHA.
After any .c change, regenerate with cd ebpftracer && make build.
Each PR needs a local end-to-end run, and a rollout to a test node checked for up == 0 and duplicate-series gather errors before the next batch starts.
Decided: 50 FQDNs. --min-container-age ships at 0 (off), changed from upstream's 30s after review: with 30s, crash-looping pods at the maximum backoff and jobs under 30s never reported.
Done when: the series count per node is measured before and after, and containers living longer than 30s keep all their series.
Decided: the TLS attach delay ships at 0. Upstream's 30s would drop the first 30s of every process's TLS traffic, and all of it for processes that live less than 30s. Measure attach CPU with #352 in place before changing that. Upstream uses a single flag for both delays, so keeping Python/Node at 30s with TLS at 0 needs a separate flag; settle that in the PR.
Done when:
The uprobe count on a test node stays flat over 24h.
TLS-attach CPU is no worse than before.
Plaintext captured for long-lived processes doesn't drop.
It adds container_*_inbound_* metrics labelled status (plus method for RabbitMQ/NATS). It has to be re-checked against our variable-length ring records (#356) and the HTTP/2 resync and stream-cap changes (#364, #365).
Done when:
A request from an uninstrumented client produces container_http_inbound_requests_total.
With both sides instrumented, server-side MySQL, Redis and Postgres counts match client-side counts within 1%.
Ring-buffer drops don't regress.
The programs still load on 5.8.
B7: logs (logparser fork synced first)
Sync the 6 missing upstream logparser commits into our logparser fork
Done when: process start times are unchanged and the taskstats calls are gone from the CPU profile.
After the port
Once every commit above is ported or skipped, open a PR with git merge -s ours <upstream tag>. It changes no files; it gives git the right merge base, so the next sync starts from v1.36 instead of 2025-06-03.
Not porting (for now)
Commits
Reason
34de3a0, 481feb0, and the windows/* parts of later commits
JVM async-profiler and Go heap profiling: revisit with the profiling work. The async-profiler is also injected into JVMs
f6813ff
docker → moby client: do it when the docker client deprecation requires it
de67715, ae3e2e6
group AWS/GCP endpoints by FQDN: overlaps the FQDN resolution in #313–#315; compare naming first
f97a3f8
ping timestamping: the pinger is off by default here
80d4530
RHEL 7 / kernel 3.10: our L7 path needs BPF ring buffers (5.8+, #363), so this needs a perf-buffer fallback design. First test whether RHEL 8's backported ring buffer loads our programs
2de271a
removes the Vagrant test setup, which we don't have
Tracking issue for porting coroot/coroot-node-agent v1.29.0 → v1.36.2 into this fork.
Where we are
git merge-basestill points at 2025-06-03. That's why a plaingit merge upstream/mainconflicts everywhere.mainand the conflict hunks counted, ignoring the generatedebpftracer/ebpf.go. The counts are in the tables below.ssl_pending,*_last_read_fd). Here they have to be ported onto fix(ebpf): capture OpenSSL traffic regardless of BIO type and OpenSSL version #346'sssl_write_pending/ssl_read_pending/ssl_fds.flags/,node/andgpu/into*_linux.go/*_windows.go. Later upstream commits that touchflags/flags_linux.goorwindows/*need their paths remapped by hand.How to work through it
git cherry-pick -xso every commit message keeps the upstream SHA..cchange, regenerate withcd ebpftracer && make build.up == 0and duplicate-series gather errors before the next batch starts.Batches
B1: push path and service lifecycle (Go only)
--ca-filefor TLS verification (2) (fix: port upstream push-path and container-tracking fixes (B1 of #369) #370)podruntime.slice(clean) (fix: port upstream push-path and container-tracking fixes (B1 of #369) #370)coroot/coroot-node-agent@1e261c3:already covered by fix: log each agent message once, and skip untracked sockets before the map lookup #358--log-level(2)Done when:
CA_FILE.WAL_DIRis discarded and sending resumes.systemctl restart <unit>keeps the unit's series continuous.B2: cardinality guards
container_dns_requests_total; the overflow goes to~other(4) (fix: cap DNS domain labels per container, add opt-in --min-container-age (B2 of #369) #371)--min-container-age(3) (fix: cap DNS domain labels per container, add opt-in --min-container-age (B2 of #369) #371)Decided: 50 FQDNs.
--min-container-ageships at 0 (off), changed from upstream's 30s after review: with 30s, crash-looping pods at the maximum backoff and jobs under 30s never reported.Done when: the series count per node is measured before and after, and containers living longer than 30s keep all their series.
B3: L7 parser correctness (eBPF)
Done when:
B4: uprobe lifecycle (port; overlaps #349, #352, #354)
Process.Close(9)Decided: the TLS attach delay ships at 0. Upstream's 30s would drop the first 30s of every process's TLS traffic, and all of it for processes that live less than 30s. Measure attach CPU with #352 in place before changing that. Upstream uses a single flag for both delays, so keeping Python/Node at 30s with TLS at 0 needs a separate flag; settle that in the PR.
Done when:
B5: TLS runtimes (port onto the #346 hooks)
The Java agent is injected into running JVMs. It stays opt-in, off by default.
Done when:
HttpClientcalling HTTPS produce HTTP metrics with the right status, and spans with paths.B6: inbound L7 (port; depends on B3–B5)
It adds
container_*_inbound_*metrics labelledstatus(plusmethodfor RabbitMQ/NATS). It has to be re-checked against our variable-length ring records (#356) and the HTTP/2 resync and stream-cap changes (#364, #365).Done when:
container_http_inbound_requests_total.B7: logs (logparser fork synced first)
open()fd is already gone (4)windows/parts)windows/parts)trace_id/span_idin JSON logs (2)Done when:
B8: process start and boot time
/proc/<pid>/statinstead of taskstats (1)Done when: process start times are unchanged and the taskstats calls are gone from the CPU profile.
After the port
git merge -s ours <upstream tag>. It changes no files; it gives git the right merge base, so the next sync starts from v1.36 instead of 2025-06-03.Not porting (for now)
windows/*parts of later commits