openstack_test: serial coda for Machine leak, suite flakes, and topology - #44
Conversation
Keep the main openstack_test OTE suite in batch for wall time, but split bz_2073398 and the Machine trio into an ordered serial coda so DeferCleanup from openstack-test#321 can finish before Machine Running / MachineSet replica / ProviderSpec. No-ops when those tests are absent (lb allowlists). Co-authored-by: Cursor <cursoragent@cursor.com>
Shell-splitting multi-word matcher argv pulled accidental tests into the serial coda (leak_cluster=7). Write one matcher per line and invoke the splitter with a matchers file path instead. Co-authored-by: Cursor <cursoragent@cursor.com>
Pre-count expected tests from the serial list file instead of incrementing in the while loop. The in-loop counter stayed at 1 after four successful runs (passed=4 expected=1), falsely marking the coda UNSTABLE. Co-authored-by: Cursor <cursoragent@cursor.com>
|
ci-framework-testproject PR : https://gitlab.cee.redhat.com/ci-framework/ci-framework-testproject/-/merge_requests/2734 |
IlanZuckerman
left a comment
There was a problem hiding this comment.
Please confirm whether this is intentional that coda run after batch failure + must-gather in addition to both of mine comments.
Co-authored-by: Cursor <cursoragent@cursor.com>
|
@IlanZuckerman Confirming: yes, serial coda after batch failure + must-gather is intentional. The coda task sits outside the batch block/rescue in |
Extend Option B (PR #44) so egressIP FIP failover and prometheus PVC resize leave the randomized batch and run one-by-one after the Machine leak cluster. Both passed alone on warm serval71 but timed out under suite load (10s AAP wait / 3m pod-ready wait). Related: #44 Co-authored-by: Cursor <cursoragent@cursor.com>
Pull the topology AZ check and the enable_topology=false mutator out of the randomized batch into the Option B serial coda. Order AZ check first so it still sees enable_topology=true, then the mutator. Avoids the race seen on warm serval71 TP (Expected false to equal true at topology.go:89). Related: #45 Co-authored-by: Cursor <cursoragent@cursor.com>
Fail hard when any coda matcher is present but missing or ambiguous, and drop leftover Option B wording in favor of serial coda. Co-authored-by: Cursor <cursoragent@cursor.com>
|
@ekuris-redhat Matcher uniqueness from your #45 review is addressed here in |
|
/lgtm |
|
/approve |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: ekuris-redhat The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
Problem
Under randomized OTE batch, several tests flake or race:
enable_topology=falsemutator)Solution
Split ordered matchers out of
list_of_tests_to_run.txt, run the main suite in batch, then run a serial coda. No-op when matchers are absent (e.g.lb_testsallowlists).Coda order (8)
bz_2073398MachineSet scale-in port leak2–4. Machine trio (
phase Running, MachineSet replica, ProviderSpec)enable_topology=falseMatcher uniqueness
When any coda test is present in the filtered list, every matcher must resolve to exactly one test; missing or non-unique matches fail hard. Zero hits keep the allowlist no-op (empty serial file, batch unchanged).
Intentional behavior
Serial coda runs after a failed batch (+ must-gather on the rescue path). The coda task sits outside the batch block/rescue in
run_openstack_test.yml, so serial isolation and junit merge still run when the batch is UNSTABLE.Notes
Absorbs former PR #45 (egressIP / prometheus / topology coda matchers).