Skip to content

Reduce monitoring ingestion costs - #2640

Merged
ejsmith merged 1 commit into
mainfrom
feature/reduce-monitoring-ingestion-costs
Oct 4, 2026
Merged

ejsmith merged 1 commit into
mainfrom
feature/reduce-monitoring-ingestion-costs

Conversation

@ejsmith

@ejsmith ejsmith commented Oct 4, 2026

Copy link
Copy Markdown
Member

Reduce Azure ingestion costs while retaining signals used to investigate incidents. These manifests record the monitoring changes already applied and checked on October 3–4.

  • Restrict Azure's annotated-pod Prometheus collection to 31 certificate, controller, ingress endpoint, queue, and leader-election metrics. Keep the one-minute interval. The repository previously disabled this scrape; enabling it with an allowlist aligns it with the live configuration, where scraping was already enabled without filtering.
  • Default both ECK Fleet policies to warning logging, retaining warnings/errors and agent log/metric monitoring. Application logging is unchanged.

The tradeoff is less historical controller runtime/latency detail and no informational Elastic Agent logs. No application API or public contract changes.

Verification:

  • All YAML documents parse; semantic checks confirm the 31 unique metric names, interval, logging levels, and retained monitoring. All other parsed configuration matches main; git diff --check passes.
  • Prior live checks retained all 31 metric names and native Azure diagnostics. Six complete hours showed about 87% less custom-metric ingestion than the seven-day baseline.
  • All six Elastic Agents were healthy at warning level; 57 Elastic metric datasets remained fresh. Informational agent logs stopped in the observed window, with warnings still emitted and no application or pod restarts.

Projected savings are about $200 per 30 days ($168 metrics plus $33 logs), based on measured billable usage at $2.30/GB. These are projections, not finalized invoice savings.

Deployment note: preconfigured Fleet defaults apply when policies are created. Existing live policies were updated separately through Fleet; one agent also required a logging-level action to work around its existing restart-related setting issue.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-04T18:42:33.503513Z f299fc8 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

Code Coverage

Package Line Rate Branch Rate Complexity Health
Exceptionless.AppHost 23% 23% 128 ❌
Exceptionless.Core 77% 68% 10825 ✔
Exceptionless.Insulation 51% 43% 370 ➖
Exceptionless.Web 86% 70% 9173 ✔
Summary 80% (27805 / 34874) 68% (13736 / 20064) 20496 ✔

@ejsmith
ejsmith merged commit 067ed96 into main Oct 4, 2026
21 checks passed
@ejsmith
ejsmith deleted the feature/reduce-monitoring-ingestion-costs branch October 4, 2026 19:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant