Skip to content

Draft: CloudHub application sizing and OutOfMemoryError troubleshooting (W-24093783) - #906

Open
kevintroller wants to merge 1 commit into
latestfrom
W-24093783-ch-sizing-memory-oom-kt
Open

Draft: CloudHub application sizing and OutOfMemoryError troubleshooting (W-24093783)#906
kevintroller wants to merge 1 commit into
latestfrom
W-24093783-ch-sizing-memory-oom-kt

Conversation

@kevintroller

Copy link
Copy Markdown
Contributor

Summary

  • First draft for W-24093783: add an Application Sizing and Memory Management section on cloudhub-architecture.adoc (the ticket’s cloudhub/sizing-guide URL 404s; that page does not exist).
  • Draft uses only already published behavior: worker/heap table, 0.1/0.2 vs 1+ vCores for consistent performance, Anypoint Monitoring JVM/Overview charts, automatic restarts, impaired-worker Unrecoverable Runtime Error after five failed starts, streaming/batch OOM causes, and CloudHub 2.0 replica xrefs.
  • Do not merge yet. Engineering confirmation is still open in #cloudhub-doc-review. After ENG replies, we can add or drop claims (production-minimum 1 vCore, exit 137/SIGKILL, saw-tooth heap, OOM → Unrecoverable sequence, RTF → CH2 size mapping).

Request vs. draft

Ticket request Status
vCore memory limits (~500 MB / 1 GB heap) Already on this page in the CloudHub Workers table — not duplicated; new section points to it
vCores, memory, worker count Covered; vertical vs horizontal scale with xref to Worker Scale-out
Fractional vCores unsuitable for production / min 1 vCore Not added — kept existing “1 or more vCores provide performance consistency” pending ENG
Diagnose OOM (exit code 137, SIGKILL) Not added — unsourced; asked ENG
Saw-tooth heap pattern Not added — unsourced; asked ENG
Anypoint Monitoring heap / GC / message load Added UI steps from published Monitoring docs
Streaming / batch for large payloads Cross-linked Mule runtime topics; not copied
RTF → CloudHub 2.0 sizing Not added — no official mapping; xref CH2 replica table and CH→CH2 migration only
Exceeding memory → OS kill → Unrecoverable runtime error Not added as a direct chain — xref impaired-worker (five failed starts) pending ENG

Style

Active voice, contractions (can't, don't), no “the following”. Matched existing CloudHub Architecture heading levels and reused existing xrefs instead of new pages.

Sources

  • GUS W-24093783 — case-reduction request (Hive report); target URL 404
  • GUS W-24093648 — parent “Create work items from report”
  • Published docs: CloudHub Workers table, Impaired Worker Monitoring, Application Monitoring and Automatic Restarts, Anypoint Monitoring built-in dashboards, Mule FATAL_JVM_ERROR / streaming / batch
  • Slack #cloudhub-doc-review — ENG confirmation questions (4 Sep 2026)

Test plan

  • Local Antora build: new === / ==== headings render under CloudHub Workers
  • xrefs resolve: cloudhub-fabric, worker-monitoring, cloudhub-impaired-worker, mule-runtime::mule-error-concept, monitoring::app-dashboards, cloudhub-2::ch2-architecture, ch-ch2-migration-configuration
  • Heap-dump partial renders
  • After ENG replies: update section for confirmed OOM symptoms only, then SME review

Give customers a path from worker sizing to OutOfMemoryError diagnosis using published heap limits, Monitoring charts, and existing restart/impaired-worker docs. Draft only; engineering confirmation still needed for production-minimum vCore, OOM kill signals, and RTF-to-CH2 mapping.
@kevintroller
kevintroller requested a review from a team as a code owner September 4, 2026 14:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant