Skip to content

Latest commit

 

History

239 Commits

Folders and files

Repository files navigation

lattice-node-agent

Outbound node daemon for Lattice.

The agent has no inbound listener. It authenticates with a per-node token, reports metrics and slow-changing HostFacts inventory telemetry, polls for queued tasks, executes bounded tasks only when explicitly enabled, and posts results back to the server. It can also report proxy-core traffic counters from a bounded local JSON snapshot file or a loopback-only HTTP JSON source for the server-owned usage rollup. Interactive browser terminal sessions are supported only when explicitly enabled; they are outbound, agent-side PTY sessions, not an inbound SSH service. Node tokens are sent in the Authorization: Bearer header, not in JSON bodies. For rollback-protected firewall apply tasks, the binary also supports --selfcheck-controlplane, a one-shot unauthenticated /api/health reachability check used after nft commit; this mode does not require or send the node token.

HostFacts are best-effort advisory facts (OS, arch, CPU cores/model, memory/swap, platform, kernel, hostname, boot time, virtualization hint). They are collected with stdlib and local platform files such as /proc and /etc/os-release; missing fields are left empty and never block the agent.

NetGuard reality reporting is on by default (it was opt-in from 0.3.5 to 0.3.9-alpha.3, which left the whole fleet silent); set LATTICE_REPORT_GUARD_REALITY=0 or pass -report-guard-reality=false to opt a node out. Each agent interval the agent collects one read-only snapshot and posts it to /api/agent/guard-reality. Collection invokes local ss, ip, and nft commands with bounded output and a shared 10-second collection-and-report deadline. ss and ip are required; a node without nft reports its listeners and interfaces with the ruleset facts left empty, so the console shows what it can prove instead of a stale snapshot. Core task, monitor, and log-source polling runs before this optional report, so a degraded collector cannot delay that work in the current cycle. Reported facts are low-trust input for display, drift, and suggestions only. They never mutate nftables or author policy.

The same snapshot carries sshd facts for SSH Guard under reality.sshd: the effective sshd -T configuration (password_authentication, pubkey_authentication, permit_root_login as sshd prints it, max_auth_tries, every port line under ports, every listenaddress line under listen_addresses) plus observed_at. The agent reads it only when it runs as root and sshd resolves, searching /bin, /sbin, /usr/bin, /usr/sbin, /usr/local/bin and /usr/local/sbin in that order, to a root-owned file nobody else can write whose every parent directory is a real, root-owned directory nobody else can write. That is the same ancestry test the sing-box liveness probe applies to /proc/<pid>/exe: a binary is trusted by the integrity of its path, not by which directory it sits in. The probe runs under a 3-second deadline. Anything else leaves sshd out of the report and puts the reason in reality.sshd_note; the agent never fills in a default it did not read, and the rest of the snapshot still posts.

Run

go run ./cmd/lattice-agent \
  -server http://127.0.0.1:8088 \
  -node-id demo-node \
  -token '<enrollment-token>' \
  -allow-exec=false

-allow-exec=false is the safe default. Use -allow-exec=true only on nodes where remote script execution is acceptable. Task stdout and stderr are capped per task; if either stream exceeds the server-provided output limit, the agent returns the capped bytes and marks the task result with an explicit truncation error instead of reporting silent success. Each metrics heartbeat also reports a task_sandbox runtime profile so operators can distinguish disabled execution, root-refused execution, and the Linux rlimit/process-group hardened path from the dashboard. This is visibility, not a policy bypass: -allow-exec, -allow-root-exec, and -no-exec remain the authoritative execution gates. On Linux, task interpreters also run with no_new_privs set, so a script cannot gain additional privilege through setuid or file-capability executables.

On Linux hosts with a delegated cgroup v2 service cgroup, operators can add per-task resident-memory, process/thread, and CPU caps:

LATTICE_TASK_CGROUP_ROOT=auto \
LATTICE_TASK_CGROUP_MEMORY_MAX=536870912 \
LATTICE_TASK_CGROUP_PIDS_MAX=64 \
LATTICE_TASK_CGROUP_CPU_MAX="100000 100000"

auto resolves to a lattice-tasks child under the agent's own cgroup. An absolute writable cgroup root can be used instead. Cgroup caps are off by default; once configured, task launch fails closed if the agent cannot prepare or join the task cgroup, so a node does not silently run without the requested cap.

Operators can also pin per-task working directories under a dedicated root:

LATTICE_TASK_WORK_ROOT=/opt/lattice/state/tasks

Each task still gets a fresh private subdirectory that is removed after the task. The agent sets HOME, TMPDIR, and XDG_RUNTIME_DIR to that directory so task-local temporary files do not spill into a shared temp path. The configured root must be absolute and not group/world writable; otherwise the task fails before the script runs. This is workdir containment, not a full mount namespace: scripts can still read or write other host paths allowed to the agent user.

NetGuard apply leases explicitly marked durable_result by the server are journaled before execution and upload. This scope matches NetGuard's atomic server-side result/approval/binding transition; heterogeneous legacy tasks keep their historical one-shot delivery semantics. The outbox uses a private, server-and-node-specific subdirectory under LATTICE_LOG_STATE_DIR; manual runs without that setting use the current user's cache directory. Override the base directory with LATTICE_TASK_OUTBOX_DIR or -task-outbox-dir. If a marked lease cannot be written, it does not run. After a restart, completed marked results are retried first and an interrupted marked task is reported as an unknown outcome rather than executed a second time.

For least-privilege Linux systemd installs, set LATTICE_AGENT_RUN_USER before running scripts/install.sh:

LATTICE_AGENT_RUN_USER=lattice-agent \
LATTICE_AGENT_RUN_GROUP=lattice-agent \
sh scripts/install.sh

The installer creates the user/group when needed, writes User=/Group= into the systemd unit, keeps the token env file root-only, and grants the service user ownership of the agent state directory. This profile is best for monitoring, inventory, terminal, and non-privileged tasks. Host mutation tasks such as nft/WireGuard apply or agent self-update still require a root-capable service profile or a separately delegated privileged helper.

The node token is sent in the Authorization: Bearer header on every request. The loopback http://127.0.0.1:8088 URL above is safe because the token never leaves the host. For a remote server the agent refuses to start on a cleartext http:// URL (it would leak the token). Use https:// instead. The -allow-insecure-http flag exists only as a deliberate escape hatch and is off by default.

Debug diagnostics:

go run ./cmd/lattice-agent \
  -server https://lattice.example.com \
  -node-id demo-node \
  -token '<enrollment-token>' \
  -debug

LATTICE_AGENT_DEBUG=1 enables the same mode for systemd environments. Debug logs include poll progress, request paths, payload key names, metrics summaries, monitor counts, task IDs, and task exit status. They do not print the node token, task script body, proxy usage secret, or client secret values.

The server can also enable debug mode per node for lattice-agent 0.2.1+ through /api/nodes/debug or the dashboard node detail panel. Server-controlled debug writes to the node's normal service logs and, by default, also uploads the same debug lines to the server Logs store under agent-debug://<node_id>. Set collect=false on the server policy to keep debug enabled on the node while preventing central collection.

Current topology is hub-and-spoke: every agent points directly at the primary lattice-server. role and tags are metadata for filtering/planning; there is no production group-leader or relay-agent mode yet.

Interactive terminal sessions:

go run ./cmd/lattice-agent \
  -server https://lattice.example.com \
  -node-id demo-node \
  -token '<enrollment-token>' \
  -allow-terminal=true

LATTICE_AGENT_ALLOW_TERMINAL=1 enables the same mode for systemd environments. Terminal mode is off by default and is separate from reviewed batch tasks: it opens a short-lived PTY on the node, polls the server for input, and posts output back to the dashboard. The agent still has no inbound listener, does not accept SSH connections, and does not store SSH credentials. Dashboard access requires the operator scope terminal:open.

If the agent process runs as root, terminal mode is refused unless -allow-root-exec=true is also set. Prefer running the agent as a dedicated least-privilege service user when browser terminal access is needed.

Firewall apply selfcheck:

lattice-agent --selfcheck-controlplane -server https://203.0.113.99

The selfcheck exits 0 only when GET /api/health returns HTTP 200. It reuses the same transport safety guard as normal startup, so remote cleartext http:// is refused unless deliberately allowed.

Firewall apply domain-set update:

lattice-agent --update-nft-domain-set \
  -host lattice.example.com \
  -family inet \
  -table lattice_policy \
  -set lattice_control4 \
  -set6 lattice_control6

This mode resolves the hostname with Go's resolver, splits answers into IPv4 and IPv6 sets, sorts/deduplicates each family, then updates the existing nft named sets using direct nft argv calls. -set may be used alone for the legacy IPv4-only path; -set + -set6 updates both control-plane sets and requires at least one A or AAAA answer. It does not require or send the node token. It is intended for server-rendered, rollback-protected apply scripts; empty resolution or invalid nft identifiers exit non-zero so the task can roll back.

Proxy usage reporting bridge (file source):

lattice-agent \
  -server https://lattice.example.com \
  -node-id gmami-jp1 \
  -token '<node-token>' \
  -proxy-usage-file /run/lattice/proxy-usage.json

The file is read once per agent loop and posted to /api/agent/proxy-usage. The agent overrides any node_id in the file with its configured node id, defaults at when omitted, rejects empty user ids, rejects negative counters, and refuses files over 1 MiB. The server performs monotonic diffing, per-profile user eligibility filtering, quota status updates, and audit.

Minimal file shape:

{
  "core_uptime_sec": 12345,
  "user_bytes": {
    "alice": 1048576,
    "bob": 2097152
  }
}

This is an interim stable contract for sidecar collectors and future direct sing-box/xray collectors; it is not a general log or metrics ingestion channel.

Proxy usage reporting bridge (loopback HTTP JSON source):

lattice-agent \
  -server https://lattice.example.com \
  -node-id gmami-jp1 \
  -token '<node-token>' \
  -proxy-usage-url http://127.0.0.1:19090/stats \
  -proxy-usage-secret-file /etc/lattice/proxy-usage.secret

The proxy usage source is a single choice, not a pair: -proxy-usage-file, -proxy-usage-url, -proxy-usage-xray-api, and -singbox-stats-api are all mutually exclusive, and the agent picks the first one set in that order (cmd/lattice-agent/main.go). The URL must use http:// or https:// and a loopback host (127.0.0.0/8, ::1, or localhost); remote hosts and URL userinfo are refused before any request is sent. The optional local bearer secret is sent as Authorization: Bearer to that local source only. For persistent services, prefer -proxy-usage-secret-file / LATTICE_PROXY_USAGE_SECRET_FILE over -proxy-usage-secret so the secret is not exposed in process arguments or shell history. Responses are capped at 1 MiB and fetched with -proxy-usage-timeout (default 3s).

The HTTP source accepts the same Lattice snapshot shape as the file source, an envelope {"snapshot": ...}, or V2Ray-style stats output:

{
  "stat": [
    {"name": "user>>>alice>>>traffic>>>uplink", "value": 1048576},
    {"name": "user>>>alice>>>traffic>>>downlink", "value": 2097152}
  ]
}

The agent sums uplink/downlink per user and posts the normalized ProxyUsageSnapshot to /api/agent/proxy-usage. This keeps the server's monotonic diffing, eligibility filtering, quota state, and audit as the authoritative layer. Direct sing-box/xray gRPC adapters can reuse this parser later without changing the server ingest contract.

On-box sing-box discovery

-singbox-discover / LATTICE_SINGBOX_DISCOVER=1 reports existing sing-box inbounds to /api/agent/singbox-inventory without enabling generic task execution. This is the read-only adoption path for machines that already run VPN configs outside Lattice.

Discovery order:

  1. Try the 233boy management interface, sb --json list, plus best-effort sb --json provision for version metadata.
  2. If that interface is missing or returns non-JSON, fall back to parsing the running standard sing-box config set. The agent discovers sing-box run arguments from /proc/*/cmdline, honors -c/--config and -C/--config-directory, then falls back to /etc/sing-box/config.json and /etc/sing-box/conf/*.json.

The runtime-config fallback emits only safe inbound metadata: tag/name, protocol, network, public address, port, SNI, and listen host. It does not emit private keys or invent credential-bearing share URLs from raw config files. Nodes that run with -allow-exec=false should use this discovery path instead of dashboard manual probe tasks.

With discovery on, the agent also runs sb --json caps at startup and then at most every 10 minutes (5 s timeout), and adds what the script can do to the capabilities in its hello and heartbeat as sb:<cap>, for example sb:user-del-by-name. The server uses these to pick a code path per node, such as removing a deleted user from an adopted line by name. Only names from a fixed allowlist pass (the seven lr00rl/sing-box v1.24.3-alpha.8 lists), so a script cannot add arbitrary strings. A script without the verb (alpha.7 and older), a non-zero exit, a timeout, or output that is not the expected JSON means no sb: capabilities; the agent writes one debug line and the heartbeat is unchanged.

Dashboard manual probe is different from continuous discovery: it queues a bounded task and asks the on-box sb --json list/provision interface first, then falls back to parsing the running sing-box config set. Older management scripts may print human error text before fallback JSON (for example when --addr is unsupported); the server parser extracts the first JSON object from each probe section and records a bounded error summary plus the task id when the probe still fails. Full stdout/stderr stay in Task History so operators can debug incompatible local sing-box layouts without losing the last good continuous-discovery inventory.

Installer-persisted launch profile

The dashboard's enroll and reconfigure commands set lattice-agent startup behavior through environment variables. scripts/install.sh persists most of these into /opt/lattice/lattice-agent.env (or the platform equivalent) so the service keeps the same behavior after restart.

Three of the keys listed below are not persisted, as of 2026-08-19: LATTICE_SINGBOX_META, LATTICE_SINGBOX_STATS_API, and LATTICE_PROXY_USAGE_SECRET_FILE appear nowhere in scripts/install.sh, in neither the env heredoc nor the reconfigure preserve list. Set one of them and the agent honors it for the current process only; the next install or service restart drops it silently. Until the installer is fixed, write those three into the env file by hand after install, and re-apply them after any reconfigure.

The full list:

The installer downloads release artifacts through HTTPS-only curl/wget options and only when it can also download SHA256SUMS and verify the selected binary with sha256sum, shasum, or FreeBSD sha256. Missing checksum tooling or a missing checksum manifest aborts the install before the binary is written.

  • LATTICE_AGENT_ALLOW_EXEC=1 enables bounded task execution.
  • LATTICE_AGENT_ALLOW_ROOT_EXEC=1 permits task execution while the agent runs as root.
  • LATTICE_TASK_OUTBOX_DIR overrides the durable NetGuard result-journal base. The installer creates a private task-outbox leaf beneath that base and preserves it across reconfiguration; otherwise journals share LATTICE_LOG_STATE_DIR.
  • LATTICE_NO_EXEC=1 is the hard kill switch and overrides execution/terminal enablement.
  • LATTICE_AGENT_RUN_USER / LATTICE_AGENT_RUN_GROUP configure an optional non-root systemd service identity during install or reconfigure.
  • LATTICE_AGENT_ALLOW_TERMINAL=1 enables audited browser terminal sessions.
  • LATTICE_TERMINAL_TRANSPORT=poll|stream selects the terminal transport.
  • LATTICE_SSH_ALERTS=1 reports accepted sshd logins.
  • LATTICE_SINGBOX_DISCOVER=1 and LATTICE_SINGBOX_BIN=sb enable sing-box discovery. LATTICE_SINGBOX_META overrides the design-15 sidecar path (default /etc/sing-box/lattice-metadata.json), which annotates each line with its control-plane line_uuid and declared downstream chain edge.
  • LATTICE_PROXY_USAGE_FILE, LATTICE_PROXY_USAGE_URL, and LATTICE_PROXY_USAGE_XRAY_API configure proxy usage reporting sources.
  • LATTICE_SINGBOX_STATS_API=127.0.0.1:8080 enables the sing-box stats collector (read-only loopback gRPC against the core's experimental API; vendored proto, ADR-004). When it is unset and no other usage source is configured, an agent with LATTICE_SINGBOX_DISCOVER=1 reads experimental.v2ray_api.listen from the sing-box config it already discovers and uses that address if it is loopback; the agent logs the address once when it appears. The collector reports user_bytes plus the direction-split inbound_traffic, user_traffic, and outbound_traffic maps, each capped at 4096 entries with the overflow counted in ignored_counters.

If a task-backed dashboard action reports agent task execution disabled, rerun the node detail page's generated reconfigure command with allow_exec=true instead of enrolling the same machine as a second node.

Agent 0.2.7+ reports a non-secret agent_runtime object with every metrics heartbeat. The dashboard uses it as runtime proof for exec/root/terminal transport/ssh-alert/sing-box-discovery state, while agent_launch remains only the saved desired installer profile. A node can therefore show three distinct states: runtime now, saved desired, and unsaved local draft.

Terminal transport modes:

  • poll is the legacy HTTP store-and-forward path. It is slower but works with older agents and avoids a long-lived browser-to-agent stream.
  • stream attaches the browser WebSocket to an agent-dialed WebSocket bridge. It is lower latency and better for interactive shells. The agent keeps the PTY alive across short WebSocket drops, redials the server, and replays recent output from a bounded ring using the browser's rendered byte offset. Browser paste is handled through a single xterm paste path so native paste events and Cmd/Ctrl+Shift+V do not duplicate input; bracketed paste is still honored. Explicit dashboard close sends a stream close control frame so the node-side PTY is torn down immediately instead of waiting for detach cleanup.

Default Linux install layout:

  • Binary: /opt/lattice/lattice-agent
  • Environment file: /opt/lattice/lattice-agent.env
  • State directory: /opt/lattice/state
  • systemd unit: lattice-agent.service

Legacy beta nodes may still use /opt/lattice/node-agent/lattice-agent, /opt/lattice/node-agent/agent.env, and lattice-node-agent.service. The installer adopts that existing service and env file when rerun, preserving the node token and upgrading the binary that systemd actually starts. This avoids creating a duplicate agent under the canonical path. If a reconfigure command targets a different node id than the token stored in the existing env file, the installer refuses to run rather than cross-wiring one node's token to another node id.

Keepalive and supervision

The heartbeat (POST /api/agent/metrics) runs on its own goroutine every interval with a 10 s deadline, apart from the work loop that fetches config and posts usage, inventory, task polls, monitors, log sources, trace, debug lines and guard reality. A slow control plane slows those reports, not the liveness signal. Each beat carries loop_health: process start, the last cycle's start, end and duration, the step in progress and since when, per step the last success, the last error (one bounded line) and the consecutive error count, how long a task batch has been in flight, the monitor result queue depth and drop count, whether the watchdog is armed, and, while durable linechain recovery refuses to proceed, its reason and start. While recovery is blocked the beat does not advertise durable-task-result-v1. When recovery blocks startup, the heartbeat starts before hello so the node reads online with the reason instead of offline.

On systemd the installer writes /etc/systemd/system/lattice-agent.service.d/10-lattice-keepalive.conf when the installed binary lists sd-notify-v1 in -compat-json:

[Service]
NotifyAccess=main
RuntimeDirectory=lattice-agent
RuntimeDirectoryMode=0700

With that notify socket the agent sends READY=1 once its local state is open. Once durable linechain recovery at startup has finished, it arms a 120 s watchdog for itself with WATCHDOG_USEC= (a unit that sets WatchdogSec decides the timeout instead) and sends WATCHDOG=1 at half the timeout only while the work loop and the heartbeat have both moved within five minutes, or three intervals when the interval is longer. Startup recovery restarts sing-box once per interrupted journal and has no deadline of its own, so it runs before the watchdog rather than under it; the heartbeat is judged from the moment it starts. Every request has a timeout, so an unreachable control plane never stops the keepalive; a step that never returns does, and systemd restarts the agent under Restart=always. The unit itself stays Type=simple without WatchdogSec, so an older binary installed under the same drop-in runs as before.

After its first successful hello the agent writes its version to $RUNTIME_DIRECTORY/healthy (/run/lattice-agent/healthy), the health-marker-v1 contract. A server-managed update whose plan names the update guard arms a transient timer that restores the previous binary and restarts the service when the new agent has not written its version there within 300 s; the server writes the same drop-in for a binary that advertises sd-notify-v1.

On openrc the installer runs the agent under supervise-daemon (respawn_delay=10, unlimited respawns) where OpenRC provides it, so a crashed agent comes back as it does under systemd and launchd.

Monitor results are queued and sent each interval in batches of up to 200 on POST /api/agent/monitor-results. A server without that route answers 404 and the agent posts one result per request on /api/agent/monitor-result for 30 minutes before it tries the batch route again. The queue holds up to 2000 results through a control plane outage; past that the oldest are dropped. The queue lives in memory: a clean stop (an update, an update guard's restore, an operator restart) sends it once more within 5 s, while a crash or a watchdog kill loses it. Every result the server never stored (overflow, a batch or a single result the server refused, or one a batch answer lists as dropped) is logged and counted in loop_health.monitor_results_dropped.

Control-plane witness

Lattice cannot report its own outage, so one node watches it from outside: lattice-agent -witness /etc/lattice-witness/witness.json runs as its own process under lattice-witness.service. Both the unit and the config are written by an approved control-plane plan (controlplane-witness), never by hand, and the binary the unit runs is a copy of the agent taken when the plan applied (/usr/local/lib/lattice-witness/lattice-agent), so an agent update or an update guard's restore never changes the safety net.

It is a separate process because the agent exits when its first hello fails: after a reboot or a restart during an outage, an in-process witness would be gone exactly when it is needed. The witness holds no node token, carries no fleet data, and talks to three things only: the control plane's public readiness URL (<public URL>/readyz, reached the way any client reaches it, never the agent's private path), one to three reference URLs, and the bark-server on the node's own loopback interface. The Bark device key is read from a root-only file the config names (mode 0600, owned by the witness's user, the shape of a Bark key), at start and at every push; the config holds only its path. The config must belong to the witness's user too and must not be writable by group or others. Both files are checked on the opened file, not the path, so the file checked is the file read.

Every interval (30 s by default) it asks the readiness URL; only HTTP 200 counts. When that fails it asks the references, and any HTTP answer from one of them proves the node's own network. Then:

  • The control plane keeps failing with the network up for the hold window (3 minutes and at least two checks by default): one push, at the configured Bark level (critical by default), naming the node, the host, how long the failures have lasted, the last classified reason, and when the first failure was by the node's own clock. It is titled "not ready" when the last check got an HTTP answer other than 200 (a 503 from /readyz when its store or audit check fails, or a status from a proxy in front of the server), and "unreachable" when nothing answered.
  • The control plane and every reference fail together: the node's own network is down, the check is not counted as a failure, does not end the run, and nothing is pushed. The hold keeps running on the clock from the run's first failure; the first failed check after the network returns only confirms, and the next one may push.
  • The control plane answers again without a break for the recovery window (1 minute and at least two checks): one recovery push at level active. A control plane that flaps during recovery sends nothing more; one that recovers before the alert was delivered sends nothing at all.
  • A push the bark-server did not accept is owed again at every check until it is delivered; it is logged once, not at every check.

The state file (/var/lib/lattice-witness/status.json, 0644, no secret) is written atomically after every check. A witness restarted in the middle of an outage reads it back, so it neither pushes twice nor forgets the recovery; a witness that was itself stopped for longer than the hold window (or three intervals) starts counting again. While it runs, the hold and the recovery window are measured on the monotonic clock, so a step of the node's wall clock (NTP correcting drift, a wrong RTC fixed after boot) neither pages early nor holds a page back; a wall clock that moved ahead by more than that gap while the monotonic clock stood still (the machine was suspended) starts the count again too. Across a restart only the saved wall time is left, so a saved last check later than the node's clock now starts the count again as well. A check or a push cut short because the witness is stopping is not recorded: the push is owed again after the restart. The main agent attaches that file to its heartbeat as witness, re-encoded from its typed form, so the console can show the last check and the last push; it reads the file and nothing else and never starts, stops or configures the witness. LATTICE_WITNESS_STATUS_FILE moves where it looks. The relay adds relayed_at, this node's clock when the agent read the file. A witness whose service stopped or wedged leaves its last status on disk and the agent keeps relaying it, so the server compares relayed_at with the last check, on the same clock, against the interval and shows such a witness as stopped rather than watching.

The witness has no mute and no quiet hours. Once it is applied, any stop of the control plane longer than the hold (a server switch, the stop for a backup) pages at the configured level. Before a planned stop, either expect that page or first apply a witness plan with a hold longer than the stop; only a remove plan silences the witness.

lattice-agent -witness <config> -witness-check validates the config and the key file and prints a summary with the config's SHA-256, never the key. The apply script runs it before it enables the unit. Binaries with witness mode list control-plane-witness-v1 in hello capabilities and in -compat-json features; the server plans a witness only for a node that advertises it.

Execution Limits

  • Interpreter allowlist: sh, bash, python3, node.
  • Default timeout: 30 seconds.
  • Maximum timeout: 10 minutes.
  • Output cap: up to 256 KiB.
  • Server-side task creation enforces the same interpreter, timeout, output, and script-size limits before a task can be leased.
  • Server-side result ingestion also rejects stdout, stderr, or error text that exceeds the task's output cap.
  • Temporary working directory and minimal environment. HOME, TMPDIR, and XDG_RUNTIME_DIR are bound to the task's private workdir; set LATTICE_TASK_WORK_ROOT to place those per-task directories under an operator-controlled root.
  • Linux tasks inherit umask 077 from the rlimit shim so files created by task scripts default to owner-only access.
  • Leased tasks carry a server-issued lease_id; the agent returns it with the result and exposes it to the task as LATTICE_TASK_LEASE_ID for traceability.
  • A lease must be durably journaled before execution. Completed or unknown-outcome results remain in the outbox until the server acknowledges them, and are flushed before the agent fetches any new task.
  • Leased task payloads contain only execution fields; control-plane actor/token metadata is not sent to agents.
  • A server-managed agent update runs as an ordinary task. Its script replaces the binary and schedules the service restart through a transient systemd timer that fires 3 s after the script exits, so the agent being replaced posts the task result first. A stop that lands while that upload is still in flight waits for the upload (taskShutdownGrace) rather than dropping it, and the control plane also confirms the update from the version the new agent reports in its first hello, so a lost result does not leave the approval unresolved.

Development

go test ./...
go build ./cmd/lattice-agent

Releases

Push a semver tag to publish Linux and Darwin binaries. During alpha/beta development, use prerelease tags such as v0.2.10-alpha.1; stable-looking vX.Y.Z tags are reserved for deliberate stable releases. Prerelease releases are marked GitHub prerelease and explicitly not Latest, so server target_version=latest will not auto-select them.

NEXT_AGENT=v0.2.9
git tag "$NEXT_AGENT"
git push origin "$NEXT_AGENT"

The release workflow builds:

lattice-agent-linux-amd64
lattice-agent-linux-arm64
lattice-agent-darwin-amd64
lattice-agent-darwin-arm64
SHA256SUMS

The tag version is injected into the binary with -X main.version=..., so:

lattice-agent -version

must match the server update policy target version. Compatibility metadata is separate so update scripts can keep strict version checks:

lattice-agent -compat-json

Release builds inject the minimum supported server/dashboard channel into that JSON. Use the matching artifact URL and SHA-256 digest from SHA256SUMS when configuring server-controlled agent updates.

About

Outbound node agent for Lattice fleet monitoring and bounded automation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages