Skip to content

Port upstream coroot-node-agent v1.29.0 → v1.36.2 #369

Description

@mayankpande88

Tracking issue for porting coroot/coroot-node-agent v1.29.0 → v1.36.2 into this fork.

Where we are

How to work through it

  • One PR per batch, in order.
  • git cherry-pick -x so every commit message keeps the upstream SHA.
  • After any .c change, regenerate with cd ebpftracer && make build.
  • Each PR needs a local end-to-end run, and a rollout to a test node checked for up == 0 and duplicate-series gather errors before the next batch starts.
  • Compare cumulative counters, not short windows.

Batches

B1: push path and service lifecycle (Go only)

Done when:

  • Remote-write to a store behind TLS signed by a private CA works with CA_FILE.
  • An empty or corrupt file in WAL_DIR is discarded and sending resumes.
  • systemctl restart <unit> keeps the unit's series continuous.

B2: cardinality guards

Decided: 50 FQDNs. --min-container-age ships at 0 (off), changed from upstream's 30s after review: with 30s, crash-looping pods at the maximum backoff and jobs under 30s never reported.

Done when: the series count per node is measured before and after, and containers living longer than 30s keep all their series.

B3: L7 parser correctness (eBPF)

Done when:

  • The programs load on a 5.8 kernel and on a current kernel.
  • Verified instruction counts stay within limits.
  • A ClickHouse client and a RabbitMQ client are classified correctly.
  • L7 drop and unknown counters don't regress over 24h.

B4: uprobe lifecycle (port; overlaps #349, #352, #354)

Decided: the TLS attach delay ships at 0. Upstream's 30s would drop the first 30s of every process's TLS traffic, and all of it for processes that live less than 30s. Measure attach CPU with #352 in place before changing that. Upstream uses a single flag for both delays, so keeping Python/Node at 30s with TLS at 0 needs a separate flag; settle that in the PR.

Done when:

  • The uprobe count on a test node stays flat over 24h.
  • TLS-attach CPU is no worse than before.
  • Plaintext captured for long-lived processes doesn't drop.

B5: TLS runtimes (port onto the #346 hooks)

The Java agent is injected into running JVMs. It stays opt-in, off by default.

Done when:

  • A rustls client and a Java HttpClient calling HTTPS produce HTTP metrics with the right status, and spans with paths.
  • With the Java agent enabled, the JVM stays stable under load.

B6: inbound L7 (port; depends on B3–B5)

It adds container_*_inbound_* metrics labelled status (plus method for RabbitMQ/NATS). It has to be re-checked against our variable-length ring records (#356) and the HTTP/2 resync and stream-cap changes (#364, #365).

Done when:

  • A request from an uninstrumented client produces container_http_inbound_requests_total.
  • With both sides instrumented, server-side MySQL, Redis and Postgres counts match client-side counts within 1%.
  • Ring-buffer drops don't regress.
  • The programs still load on 5.8.

B7: logs (logparser fork synced first)

Done when:

  • A log storm into one container doesn't push the agent's CPU past the cap.
  • JSON log lines become OTLP records with attributes and trace context.

B8: process start and boot time

Done when: process start times are unchanged and the taskstats calls are gone from the CPU profile.

After the port

  • Once every commit above is ported or skipped, open a PR with git merge -s ours <upstream tag>. It changes no files; it gives git the right merge base, so the next sync starts from v1.36 instead of 2025-06-03.

Not porting (for now)

Commits Reason
34de3a0, 481feb0, and the windows/* parts of later commits Windows support is out of scope
e5818fc, e73d36c, b9fc7fd, 675758e, 835ae85, 0c934cd, 8992a5b JVM async-profiler and Go heap profiling: revisit with the profiling work. The async-profiler is also injected into JVMs
f6813ff docker → moby client: do it when the docker client deprecation requires it
de67715, ae3e2e6 group AWS/GCP endpoints by FQDN: overlaps the FQDN resolution in #313–#315; compare naming first
f97a3f8 ping timestamping: the pinger is off by default here
80d4530 RHEL 7 / kernel 3.10: our L7 path needs BPF ring buffers (5.8+, #363), so this needs a perf-buffer fallback design. First test whether RHEL 8's backported ring buffer loads our programs
2de271a removes the Vagrant test setup, which we don't have
7efe671, 91bb874 already covered (#346, Dockerfile)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions