Skip to content

feat(tls): attribute TLS losses per container and find sockets behind wrapped Go connections - #354

Merged
mayankpande88 merged 3 commits into
fix/tls-capture-gapsfrom
feat/tls-drop-attribution
Oct 5, 2026
Merged

mayankpande88 merged 3 commits into
fix/tls-capture-gapsfrom
feat/tls-drop-attribution

Conversation

@mayankpande88

Copy link
Copy Markdown
Contributor

Stacked on #353. It targets that branch until #353 merges.

Problem

node_agent_tls_plaintext_dropped_total (added in #353) says how much TLS plaintext the kernel could not attribute to a socket, but not which process it came from. A loss could not be traced to a workload without a debug build.

Once it could be, almost all of the loss on a multi-node cluster came from two programs that run crypto/tls over their own net.Conn wrappers:

  • A reverse proxy. Its TLS connections are three wrappers deep, each embedding the next at offset 0.
  • A metrics scraper. Its wrapper puts an int32 counter before the embedded net.Conn, at offset 8.

The Go TLS probe only unwrapped one level, at offset 0, so all of their TLS plaintext was dropped.

Change

Attribution (feat(tls): attribute TLS plaintext losses...)

  • The TLS library uprobes also count each loss per process and reason, in an LRU map. The syscall programs are untouched.
  • The agent reads and clears those counts every stats tick. It also drains a process's entries when the process exits, because a short-lived process is gone before the next tick. That is why the process's executable is recorded when its probes are first checked.
  • container_tls_plaintext_dropped_total{reason} counts losses per container.
  • The first loss per binary and reason is logged once, as a warning naming the executable and what the reason usually means:
    TLS plaintext of /usr/local/bin/<proxy> (pid N, container …) is not being captured: N go_fd_unknown events. crypto/tls runs over a connection whose socket the probe cannot find: …
    

Wrapped connections (fix(ebpf): find the socket behind wrapped Go connections)

  • The probe descends up to four levels, taking the embedded net.Conn at offset 0 or 8.
  • When a level counts as the socket:
    • Exact: its type is the binary's *net.TCPConn itab.
    • Otherwise (stripped binaries, or PIE binaries whose symbol table holds the itab unrelocated), the fd it yields is accepted only if it names an IPv4 or IPv6 socket of the process.
  • Same shape: gRPC's syscallConn has the same layout and goes through the same walk.
  • No socket: TLS over in-memory connections (net.Pipe, gRPC bufconn) has none, so it is still counted as go_fd_unknown.

Verification

  • gofmt, go vet ./..., golangci-lint and go test (CI package set, plus containers locally) pass on linux/arm64. ebpf.go is regenerated with make build.
Local end-to-end run

Local e2e: built the agent from this branch and ran it privileged in a local Linux VM (kernel 6.10, standalone mode), then ran the previous build the same way.

  • Clients: HTTPS clients whose crypto/tls runs over the two wrapper shapes, built with and without -s -w, plus a plain *net.TCPConn control and TLS over net.Pipe.
  • Each client: three runs of ten requests, each on a new connection.
client previous build: captured (dropped) this branch: captured (dropped)
plain *net.TCPConn 27 (0) 27 (0)
one wrapper, net.Conn at offset 8 0 (54) 27 (0)
three wrappers, offset 0 0 (54) 27 (0)
one wrapper, stripped 0 (53) 27 (0)
three wrappers, stripped 0 (53) 26 (0)
TLS over net.Pipe 0 (40) 0 (40)

Cluster run. Both builds ran as a DaemonSet on a multi-node cluster; rates are per minute over windows taken after each rollout settled.

attribution only this branch
go_fd_unknown (cluster) 37,970 11
…from the reverse proxy / the scraper 34,541 / 3,436 0 / 0
HTTP/2-over-TLS L7 events, server direction 48.6k 170k
HTTP/2-over-TLS L7 events, client direction 21.6k 47.3k
HTTP/1-over-TLS L7 events 325 640

All pods ran with 0 restarts and no scrape or eBPF load errors. Agent CPU on the busiest node is 88m; the fleet averages 41m and 231 MiB.

node_agent_tls_plaintext_dropped_total says how much TLS plaintext the
kernel could not attribute to a socket, but not which process it came
from, so a loss could not be traced to a workload without a debug build.

- The TLS library uprobes also count losses per process and reason in an
  LRU map. The syscall programs are left alone.
- The agent reads and clears those counts every stats tick, and drains a
  process's entries when it exits, since a short-lived process is gone
  before the next tick.
- container_tls_plaintext_dropped_total{reason} counts them per container.
- The first loss per binary and reason is logged once as a warning, naming
  the executable and what the reason usually means.
The Go TLS probe found a connection's socket only if crypto/tls ran
directly over *net.TCPConn, or over a wrapper embedding the connection
at offset 0, one level deep. Proxies and scrapers wrap deeper. One
reverse proxy's TLS connections are three wrappers deep, each embedding
the next at offset 0. A scraper's wrapper puts an int32 counter before
the embedded net.Conn, at offset 8. All of their TLS plaintext was
dropped.

The probe now descends up to four levels, taking the embedded net.Conn
at offset 0 or 8. A level is accepted when its type is the binary's
*net.TCPConn itab. Otherwise, as in stripped binaries, or PIE binaries
whose symbol table holds the itab unrelocated, the fd it yields is
accepted only if it names an IPv4 or IPv6 socket of the process. gRPC's
syscallConn has the same shape and goes through the same walk. TLS over
in-memory connections still has no socket and is still counted as
go_fd_unknown.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces tracking and reporting for TLS plaintext dropped events, attributing them to specific processes and containers, and exposing a new container_tls_plaintext_dropped_total metric. It also refactors Go TLS connection unwrapping in eBPF to handle nested wrappers and verify socket ownership. Feedback on these changes highlights a potential bug where untracked processes could globally suppress warnings across different containers due to shared zero-value cache keys, and suggests refactoring duplicate anonymous dropKey structs in the tracer into a single package-level struct for better type safety.

Comment thread containers/tls_drops.go Outdated
Comment thread ebpftracer/tracer.go
Comment thread ebpftracer/tracer.go
… per node

An unknown process has no executable identity, so its log key was the
zero value and the first such warning on a node silenced the rest. Also
name the per-pid drop map key type.
@mayankpande88
mayankpande88 merged commit 200ebbc into fix/tls-capture-gaps Oct 5, 2026
2 checks passed
@mayankpande88
mayankpande88 deleted the feat/tls-drop-attribution branch October 5, 2026 08:15
blue4209211 pushed a commit that referenced this pull request Oct 6, 2026
* fix(tls): count TLS capture losses and fix the gaps they exposed

TLS capture could fail at attach time, in the kernel, or in the agent with
nothing to show for it at default verbosity. This adds counters for each
stage and fixes the losses they revealed.

Visibility:
- node_agent_tls_attach_total{lib,result} for every attach attempt, and
  one log line per binary and outcome instead of one per process.
- node_agent_tls_plaintext_dropped_total{reason}: TLS plaintext a library
  hook saw but could not attribute to a socket.
- node_agent_l7_ringbuf_drops_total: L7 events lost to a full ring buffer.
- node_agent_l7_events_dropped_total{reason,protocol,tls}: events dropped
  in the agent before protocol parsing.

Fixes:
- Register a process on its first socket event, so its first connection
  is tracked and its TLS probes attach before its first request.
- Give a connection built from an event's socket tuple the event's
  timestamp, and let a newer event replace the closed predecessor on a
  reused fd. Both cases used to be dropped as stale.
- Key the cache of binaries that cannot be probed by file identity, not
  by path, which is only meaningful inside one container.
- Re-attach after a process execs a different binary.
- Only a write that starts with a TLS record may claim pending OpenSSL
  plaintext, and that write's ciphertext now counts toward bytes sent.
- Guard Process state shared between the instrumentation goroutine, the
  event loop and the scrape path.

* fix(containers): read Python and Node.js stats under the container lock

Collect read c.pythonStats and c.nodejsStats without c.lock while the
registry's event loop updates them under it.

* fix(containers): filter and NAT-resolve connections built from socket tuples

A connection created from an L7 event's socket tuple, because the event
arrived before the connection's open event, skipped the filters and the
NAT resolution onConnectionOpen applies. With these connections now kept
(previous commit), that tracked traffic to ignored destinations and
labelled requests with a service's virtual IP as their actual destination.

- Both paths now share connectionKey: the same port, loopback, ignored
  workload and connection filters, and the same destination key.
- The socket-tuple path takes the actual destination from the kernel's
  conntrack-derived actual_destinations map, as open events do.
- An exec is re-checked on every new socket, not only when a throttle
  shared with the periodic sweep allows it.
- The go_fd_unknown help text says it includes TLS over in-memory
  connections, which have no socket.

* feat(tls): attribute TLS losses per container and find sockets behind wrapped Go connections (#354)

* feat(tls): attribute TLS plaintext losses to the container and binary

node_agent_tls_plaintext_dropped_total says how much TLS plaintext the
kernel could not attribute to a socket, but not which process it came
from, so a loss could not be traced to a workload without a debug build.

- The TLS library uprobes also count losses per process and reason in an
  LRU map. The syscall programs are left alone.
- The agent reads and clears those counts every stats tick, and drains a
  process's entries when it exits, since a short-lived process is gone
  before the next tick.
- container_tls_plaintext_dropped_total{reason} counts them per container.
- The first loss per binary and reason is logged once as a warning, naming
  the executable and what the reason usually means.

* fix(ebpf): find the socket behind wrapped Go connections

The Go TLS probe found a connection's socket only if crypto/tls ran
directly over *net.TCPConn, or over a wrapper embedding the connection
at offset 0, one level deep. Proxies and scrapers wrap deeper. One
reverse proxy's TLS connections are three wrappers deep, each embedding
the next at offset 0. A scraper's wrapper puts an int32 counter before
the embedded net.Conn, at offset 8. All of their TLS plaintext was
dropped.

The probe now descends up to four levels, taking the embedded net.Conn
at offset 0 or 8. A level is accepted when its type is the binary's
*net.TCPConn itab. Otherwise, as in stripped binaries, or PIE binaries
whose symbol table holds the itab unrelocated, the fd it yields is
accepted only if it names an IPv4 or IPv6 socket of the process. gRPC's
syscallConn has the same shape and goes through the same walk. TLS over
in-memory connections still has no socket and is still counted as
go_fd_unknown.

* fix(tls): log drops of unknown processes once per container, not once per node

An unknown process has no executable identity, so its log key was the
zero value and the first such warning on a node silenced the rest. Also
name the per-pid drop map key type.

* fix(containers): drop an fd's HTTP/2 parser when a new connection takes it

HTTP/2 parsers are keyed by pid+fd alone. When a recycled fd got a new
connection, either from its open event or from an L7 event that arrived
first, the new connection inherited the previous one's parser and its
HPACK dynamic table. Its headers then failed to decode, or decoded wrong.

The parser is now dropped when the fd's connection is replaced. A parser
already tagged with the new connection's timestamp is kept: it was created
for this connection by an L7 event that beat the open event.

* fix(tls): confirm the socket the Go TLS connection walk finds

A level of the walk that was not a *net.TCPConn was still read as one. For a
wrapper struct the first word is an itab pointer, so the "Sysfd" read was the
itab's type hash, and when a value read that way named another live socket of
the process, the socket check accepted it: the session's plaintext went to
that connection, and it was marked TLS, every time, for that binary. On
kernels without BTF the socket check was skipped, so all four levels were
trusted unchecked.

A level found without the itab must now point to a net.netFD whose family is
AF_INET or AF_INET6 and whose sotype is SOCK_STREAM. The check reads only the
application's memory, so it holds without kernel BTF. Where the kernel's
struct offsets are known, the fd must also be a SOCK_STREAM socket of that
family. The netFD offsets come from DWARF and default to 56 and 64, which
Go 1.17 through 1.26 share. The walk no longer reads the gRPC syscallConn
itab, so its discovery is gone.

New metrics make the kernel-dependent part visible:
- node_agent_go_tls_fd_resolved_total{method,depth}: method is itab, socket
  (netFD check plus kernel check) or shape (netFD check only, no BTF)
- node_agent_ebpf_info{program_variant,btf,socket_offsets}
- node_agent_ebpf_program_instructions and
  node_agent_ebpf_program_verified_instructions, per program

TestProgramsLoad loads the variant the running kernel gets and logs each
program's size and verifier cost. TestGoTLSFdWalk runs TLS over a plain, a
wrapped and a decoy connection and checks which socket each one is
attributed to. Both need root and VM=1.

* fix(containers): drop a timestamp-less HTTP/2 parser when a tracked connection takes its fd

A parser without a connection timestamp was created by events from a socket
the kernel was not tracking, such as a connection older than the agent. A
connection opening on its fd is tracked, and the kernel stamps every event of
a tracked connection with its timestamp, so the parser cannot be the new
connection's. It was kept, and the new connection inherited its HPACK table.

* feat(tls): attach Go TLS probes at exec, and log where L7 events are dropped (#355)

* feat(tls): attach Go TLS probes at exec, and log where L7 events are dropped

A process's Go TLS probes were attached on its first connection, after
its connect event had been read. That was late for short-lived programs:
many exited first, and the rest had made their first requests unprobed.
A program that a wrapper exec'd kept probes for the wrapper's binary until
a throttled re-check noticed the change.

- A sched_process_exec tracepoint reports execs. The agent attaches Go
  TLS probes for the new image at once, and drops those of the old one.
  OpenSSL is still attached on the first connection, because the loader
  maps libssl after the exec.
- Process events are read every 10 ms, like connect events, instead of
  every 100 ms.
- The periodic executable re-check stays as a fallback for lost exec
  events, at a 10 s interval.
- Every L7 event dropped before parsing is logged once per container,
  reason and protocol, with the pid, fd and destination, so the counter's
  reasons can be traced to a workload.

* fix(containers): report L7 events on non-IP sockets as no_ip_socket

Events on sockets without an IP tuple, such as gRPC over a Unix socket,
were counted as unknown_connection, which reads as a capture loss. The
agent does not track those sockets, so they now get their own reason.

* chore(ebpf): regenerate ebpf.go after rebasing onto the review fixes
mayankpande88 added a commit that referenced this pull request Oct 6, 2026
…he map lookup (#358)

* fix(tls): count TLS capture losses and fix the gaps they exposed

TLS capture could fail at attach time, in the kernel, or in the agent with
nothing to show for it at default verbosity. This adds counters for each
stage and fixes the losses they revealed.

Visibility:
- node_agent_tls_attach_total{lib,result} for every attach attempt, and
  one log line per binary and outcome instead of one per process.
- node_agent_tls_plaintext_dropped_total{reason}: TLS plaintext a library
  hook saw but could not attribute to a socket.
- node_agent_l7_ringbuf_drops_total: L7 events lost to a full ring buffer.
- node_agent_l7_events_dropped_total{reason,protocol,tls}: events dropped
  in the agent before protocol parsing.

Fixes:
- Register a process on its first socket event, so its first connection
  is tracked and its TLS probes attach before its first request.
- Give a connection built from an event's socket tuple the event's
  timestamp, and let a newer event replace the closed predecessor on a
  reused fd. Both cases used to be dropped as stale.
- Key the cache of binaries that cannot be probed by file identity, not
  by path, which is only meaningful inside one container.
- Re-attach after a process execs a different binary.
- Only a write that starts with a TLS record may claim pending OpenSSL
  plaintext, and that write's ciphertext now counts toward bytes sent.
- Guard Process state shared between the instrumentation goroutine, the
  event loop and the scrape path.

* fix(containers): read Python and Node.js stats under the container lock

Collect read c.pythonStats and c.nodejsStats without c.lock while the
registry's event loop updates them under it.

* fix(containers): filter and NAT-resolve connections built from socket tuples

A connection created from an L7 event's socket tuple, because the event
arrived before the connection's open event, skipped the filters and the
NAT resolution onConnectionOpen applies. With these connections now kept
(previous commit), that tracked traffic to ignored destinations and
labelled requests with a service's virtual IP as their actual destination.

- Both paths now share connectionKey: the same port, loopback, ignored
  workload and connection filters, and the same destination key.
- The socket-tuple path takes the actual destination from the kernel's
  conntrack-derived actual_destinations map, as open events do.
- An exec is re-checked on every new socket, not only when a throttle
  shared with the periodic sweep allows it.
- The go_fd_unknown help text says it includes TLS over in-memory
  connections, which have no socket.

* feat(tls): attribute TLS losses per container and find sockets behind wrapped Go connections (#354)

* feat(tls): attribute TLS plaintext losses to the container and binary

node_agent_tls_plaintext_dropped_total says how much TLS plaintext the
kernel could not attribute to a socket, but not which process it came
from, so a loss could not be traced to a workload without a debug build.

- The TLS library uprobes also count losses per process and reason in an
  LRU map. The syscall programs are left alone.
- The agent reads and clears those counts every stats tick, and drains a
  process's entries when it exits, since a short-lived process is gone
  before the next tick.
- container_tls_plaintext_dropped_total{reason} counts them per container.
- The first loss per binary and reason is logged once as a warning, naming
  the executable and what the reason usually means.

* fix(ebpf): find the socket behind wrapped Go connections

The Go TLS probe found a connection's socket only if crypto/tls ran
directly over *net.TCPConn, or over a wrapper embedding the connection
at offset 0, one level deep. Proxies and scrapers wrap deeper. One
reverse proxy's TLS connections are three wrappers deep, each embedding
the next at offset 0. A scraper's wrapper puts an int32 counter before
the embedded net.Conn, at offset 8. All of their TLS plaintext was
dropped.

The probe now descends up to four levels, taking the embedded net.Conn
at offset 0 or 8. A level is accepted when its type is the binary's
*net.TCPConn itab. Otherwise, as in stripped binaries, or PIE binaries
whose symbol table holds the itab unrelocated, the fd it yields is
accepted only if it names an IPv4 or IPv6 socket of the process. gRPC's
syscallConn has the same shape and goes through the same walk. TLS over
in-memory connections still has no socket and is still counted as
go_fd_unknown.

* fix(tls): log drops of unknown processes once per container, not once per node

An unknown process has no executable identity, so its log key was the
zero value and the first such warning on a node silenced the rest. Also
name the per-pid drop map key type.

* fix(containers): drop an fd's HTTP/2 parser when a new connection takes it

HTTP/2 parsers are keyed by pid+fd alone. When a recycled fd got a new
connection, either from its open event or from an L7 event that arrived
first, the new connection inherited the previous one's parser and its
HPACK dynamic table. Its headers then failed to decode, or decoded wrong.

The parser is now dropped when the fd's connection is replaced. A parser
already tagged with the new connection's timestamp is kept: it was created
for this connection by an L7 event that beat the open event.

* fix(tls): confirm the socket the Go TLS connection walk finds

A level of the walk that was not a *net.TCPConn was still read as one. For a
wrapper struct the first word is an itab pointer, so the "Sysfd" read was the
itab's type hash, and when a value read that way named another live socket of
the process, the socket check accepted it: the session's plaintext went to
that connection, and it was marked TLS, every time, for that binary. On
kernels without BTF the socket check was skipped, so all four levels were
trusted unchecked.

A level found without the itab must now point to a net.netFD whose family is
AF_INET or AF_INET6 and whose sotype is SOCK_STREAM. The check reads only the
application's memory, so it holds without kernel BTF. Where the kernel's
struct offsets are known, the fd must also be a SOCK_STREAM socket of that
family. The netFD offsets come from DWARF and default to 56 and 64, which
Go 1.17 through 1.26 share. The walk no longer reads the gRPC syscallConn
itab, so its discovery is gone.

New metrics make the kernel-dependent part visible:
- node_agent_go_tls_fd_resolved_total{method,depth}: method is itab, socket
  (netFD check plus kernel check) or shape (netFD check only, no BTF)
- node_agent_ebpf_info{program_variant,btf,socket_offsets}
- node_agent_ebpf_program_instructions and
  node_agent_ebpf_program_verified_instructions, per program

TestProgramsLoad loads the variant the running kernel gets and logs each
program's size and verifier cost. TestGoTLSFdWalk runs TLS over a plain, a
wrapped and a decoy connection and checks which socket each one is
attributed to. Both need root and VM=1.

* fix(containers): drop a timestamp-less HTTP/2 parser when a tracked connection takes its fd

A parser without a connection timestamp was created by events from a socket
the kernel was not tracking, such as a connection older than the agent. A
connection opening on its fd is tracked, and the kernel stamps every event of
a tracked connection with its timestamp, so the parser cannot be the new
connection's. It was kept, and the new connection inherited its HPACK table.

* feat(tls): attach Go TLS probes at exec, and log where L7 events are dropped (#355)

* feat(tls): attach Go TLS probes at exec, and log where L7 events are dropped

A process's Go TLS probes were attached on its first connection, after
its connect event had been read. That was late for short-lived programs:
many exited first, and the rest had made their first requests unprobed.
A program that a wrapper exec'd kept probes for the wrapper's binary until
a throttled re-check noticed the change.

- A sched_process_exec tracepoint reports execs. The agent attaches Go
  TLS probes for the new image at once, and drops those of the old one.
  OpenSSL is still attached on the first connection, because the loader
  maps libssl after the exec.
- Process events are read every 10 ms, like connect events, instead of
  every 100 ms.
- The periodic executable re-check stays as a fallback for lost exec
  events, at a 10 s interval.
- Every L7 event dropped before parsing is logged once per container,
  reason and protocol, with the pid, fd and destination, so the counter's
  reasons can be traced to a workload.

* fix(containers): report L7 events on non-IP sockets as no_ip_socket

Events on sockets without an IP tuple, such as gRPC over a Unix socket,
were counted as unknown_connection, which reads as a capture loss. The
agent does not track those sockets, so they now get their own reason.

* chore(ebpf): regenerate ebpf.go after rebasing onto the review fixes

* fix: log each agent message once, and skip untracked sockets before the map lookup

klog.SetOutput hands every severity the same writer, and klog writes a
message to its own severity's writer and to every lower one's, so each
warning reached the log twice and each error three times. one_output
writes it once.

createConnectionFromSocketInfo parsed the socket tuple and looked up the
kernel's NAT map before applying the connection filters. A TLS server's
accepted sockets have a client's ephemeral port as their destination,
are never tracked, and came through this path on every read and write
the server made. The port filter is now checked first.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant