Skip to content

openstack_test: serial coda for Machine leak, suite flakes, and topology - #44

Merged
openshift-merge-bot[bot] merged 7 commits into
mainfrom
fix/openstack-test-machines-leak-cluster-serial
Sep 28, 2026
Merged

openshift-merge-bot[bot] merged 7 commits into
mainfrom
fix/openstack-test-machines-leak-cluster-serial

Conversation

@tusharjadhav3302

@tusharjadhav3302 tusharjadhav3302 commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Under randomized OTE batch, several tests flake or race:

  • Machine leak trio (bogus Machine / port cleanup cluster)
  • suite-load flakes (egressIP FIP failover AAP wait, prometheus PVC resize)
  • CSI topology pair (AZ check vs enable_topology=false mutator)

Solution

Split ordered matchers out of list_of_tests_to_run.txt, run the main suite in batch, then run a serial coda. No-op when matchers are absent (e.g. lb_tests allowlists).

Coda order (8)

  1. bz_2073398 MachineSet scale-in port leak
    2–4. Machine trio (phase Running, MachineSet replica, ProviderSpec)
  2. egressIP FIP failover
  3. prometheus PVC resize
  4. topology AZ (identical compute/volume AZs)
  5. enable_topology=false

Matcher uniqueness

When any coda test is present in the filtered list, every matcher must resolve to exactly one test; missing or non-unique matches fail hard. Zero hits keep the allowlist no-op (empty serial file, batch unchanged).

Intentional behavior

Serial coda runs after a failed batch (+ must-gather on the rescue path). The coda task sits outside the batch block/rescue in run_openstack_test.yml, so serial isolation and junit merge still run when the batch is UNSTABLE.

Notes

Absorbs former PR #45 (egressIP / prometheus / topology coda matchers).

Keep the main openstack_test OTE suite in batch for wall time, but split
bz_2073398 and the Machine trio into an ordered serial coda so DeferCleanup
from openstack-test#321 can finish before Machine Running / MachineSet
replica / ProviderSpec. No-ops when those tests are absent (lb allowlists).

Co-authored-by: Cursor <cursoragent@cursor.com>
tusharjadhav3302 and others added 2 commits September 26, 2026 21:33
Shell-splitting multi-word matcher argv pulled accidental tests into the
serial coda (leak_cluster=7). Write one matcher per line and invoke the
splitter with a matchers file path instead.

Co-authored-by: Cursor <cursoragent@cursor.com>
Pre-count expected tests from the serial list file instead of incrementing
in the while loop. The in-loop counter stayed at 1 after four successful
runs (passed=4 expected=1), falsely marking the coda UNSTABLE.

Co-authored-by: Cursor <cursoragent@cursor.com>
@tusharjadhav3302

Copy link
Copy Markdown
Contributor Author

@ekuris-redhat ekuris-redhat left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

/lgtm

@IlanZuckerman IlanZuckerman left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please confirm whether this is intentional that coda run after batch failure + must-gather in addition to both of mine comments.

Comment thread collection/stages/roles/openstack_test/files/split_machines_leak_cluster.py Outdated
Comment thread collection/stages/roles/openstack_test/files/split_machines_leak_cluster.py Outdated
Co-authored-by: Cursor <cursoragent@cursor.com>
@tusharjadhav3302

Copy link
Copy Markdown
Contributor Author

@IlanZuckerman Confirming: yes, serial coda after batch failure + must-gather is intentional. The coda task sits outside the batch block/rescue in run_openstack_test.yml, so serial isolation and junit merge still run when the batch is UNSTABLE; must-gather stays on the rescue path. Style nits addressed in the latest commit.

tusharjadhav3302 and others added 3 commits September 28, 2026 10:31
Extend Option B (PR #44) so egressIP FIP failover and prometheus PVC
resize leave the randomized batch and run one-by-one after the Machine
leak cluster. Both passed alone on warm serval71 but timed out under
suite load (10s AAP wait / 3m pod-ready wait).

Related: #44
Co-authored-by: Cursor <cursoragent@cursor.com>
Pull the topology AZ check and the enable_topology=false mutator out of
the randomized batch into the Option B serial coda. Order AZ check first
so it still sees enable_topology=true, then the mutator. Avoids the race
seen on warm serval71 TP (Expected false to equal true at topology.go:89).

Related: #45
Co-authored-by: Cursor <cursoragent@cursor.com>
Fail hard when any coda matcher is present but missing or ambiguous,
and drop leftover Option B wording in favor of serial coda.

Co-authored-by: Cursor <cursoragent@cursor.com>
@tusharjadhav3302 tusharjadhav3302 changed the title osp_verification: serial coda for bogus Machine leak cluster openstack_test: serial coda for Machine leak, suite flakes, and topology Sep 28, 2026
@tusharjadhav3302

Copy link
Copy Markdown
Contributor Author

@ekuris-redhat Matcher uniqueness from your #45 review is addressed here in d900268: when any serial-coda matcher hits, every matcher must resolve to exactly one test; missing or non-unique matches fail hard. Zero hits keep the allowlist no-op.

@IlanZuckerman

Copy link
Copy Markdown

/lgtm

@ekuris-redhat

Copy link
Copy Markdown
Contributor

/approve

@openshift-ci

openshift-ci Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: ekuris-redhat

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@openshift-merge-bot
openshift-merge-bot Bot merged commit db20cfe into main Sep 28, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Development

Successfully merging this pull request may close these issues.

3 participants