diff --git a/CHANGELOG.md b/CHANGELOG.md index c52a2bd..7e21c3d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -25,6 +25,7 @@ Sections dated before 2026-09-19 predate the cycle and stay as they are. - fix(project-bin/exec.sh): **the delta gate keys an error on code + message + location, counted, instead of the message text alone.** On a real 22-error baseline (Mendix 11.12.1), 20 errors shared two "Could not find widget" messages and CE0117 always reads "Error(s) in expression.", so a script adding another CE0117 in a new microflow, or another copy of a baseline error, was judged "identical" and kept. Now each Error is a hash of errorCode, message and every location's module / document / element, and the subset test is a multiset test. Names, not GUIDs, so a CREATE OR REPLACE that leaves a pre-existing error in place still reads as pre-existing. Field-run on a scratch copy of a greenfield PoC model — MendixMau - learn(workflow-structure-rules, learned-workflow-patterns, module-folder-convention, learned-mdl-preflight, bug-logs): **five workflow rules re-probed on mxcli v0.24.0 / Mendix 11.13.0.** A forward `JUMP TO` builds clean (the CE6681 row taught direction as a platform rule; it was the jump-named-after-its-target defect, fixed upstream in v0.21.0); an interrupting boundary timer ending in `jump to` or `end workflow` builds clean (BUG-109 stamped no-longer-reproduces); BUG-76 splits — the bare-enum `DECISION` is now refused by `MDL-WF03` and still corrupts when forced, a Boolean `true`/`false` decision builds at 0 errors; `create or modify workflow … folder` places the workflow. New: one interrupting boundary per activity (`MDL-WF15`/CE6697) and no `end workflow` on a non-interrupting path (`MDL-WF08`/CE1844); `SET TASK OUTCOME` needs a signed-in named user ("Only named users can complete user task"), with the run-as-session Java action pattern. Evidence: a six-probe, four-control construct run (check, exec, native `mx check`, describe read-back) and a live proof test — card-disbursement requirements-driven build - fix(project-tests/e2e/otel.js): **`capture()` pages back to t0 instead of reading only the newest 400 traces.** Jaeger answers the newest `limit` traces and says nothing about the rest; with timers running, 400 traces were two minutes, so a capture with t0 15 minutes back reported 0 errors over 24 real ERROR spans. It now pages by `end` while a page is full and still after t0, dedupes by trace, caps at `OTEL_MAX_PAGES` (25) and marks the result `.truncated` when the cap stops it short. Field run against the live Jaeger: t0 −15 min, 462 spans / 0 errors → 15,439 spans / 24 errors in 11 pages; a step's own capture stays one page — card-disbursement requirements-driven build +- fix(exec.sh): **when `exec.sh` undoes a script on a model that already had errors, it now names only the errors that script added.** Before, the restore printed every error on the model (22 on the field model, the one new one unmarked), and then said the errors were "PRE-EXISTING … This script is NOT the cause. Nothing will exec until that is cleared" and logged "blocked". Both were false: the script did add the error, and the delta gate keeps any later script that adds nothing new. Now the headline reads "this script added 1 new error(s) [CE0109] (22 in total, 21 were already there)", the new errors are listed in full, the old ones are summarised one line per code, and the build log records "1 new error(s): CE0109 — script rolled back; 22 pre-existing remain". New and old are split with the same code + message + location key the gate uses, so the two cannot disagree. With no measured baseline (`SKIP_BASELINE=1`) the output is unchanged. — MendixMau - fix(project-bin/constants-audit.sh): **a `__SET_ME__`-style sentinel default reports `SENTINEL`, not `MODEL-SECRET`.** The audit only knew SET vs EMPTY, so the placeholder `learned-constants-and-secrets.md` Step 3 prescribes for a must-override secret got the same verdict as a real password in git and needed a waiver. The test reads the value's shape (`__[A-Za-z0-9_]+__`) and still prints nothing; the skill's verdict table gains the row. Field run on a copy of the model: 4 findings → 3, the hub password constant SENTINEL — card-disbursement requirements-driven build - fix(review-report.js, coverage-check-all.sh): **a module with one coverage ledger per BRD no longer carries a permanent coverage FAULT over a clean count.** `coverage-check-all.sh` keeps `leaves:` on each per-BRD line, and `review-report.js` sums the counters across those lines (null, never zero, when a BRD line is FAULT / NO LEDGER, lacks a counter, or the lines miss the stated BRD count); an old installed copy without `leaves:` is named in the fault. Field run on two real ledgers: FAULT → pass, leaves 337 / 610, 947 summed over both — card-disbursement requirements-driven build - fix(project-bin/check-page-shell.sh): **the page-body scan skips quoted strings and block comments.** An unbalanced brace inside a string literal (a JSON payload preview `Content: '{{ … }'`) left the depth up, the scan ran on into the next pages and a page with one H1 was reported as declaring 3; a `--` inside a string cut the line and did the same. Field run: 18 pages across 87 scripts, 1 violation (the false positive) → 0 — card-disbursement requirements-driven build diff --git a/project-bin/exec.sh b/project-bin/exec.sh index 3206ae8..ca44f21 100755 --- a/project-bin/exec.sh +++ b/project-bin/exec.sh @@ -563,22 +563,67 @@ err_codes() { # Names, not elementId GUIDs: a CREATE OR REPLACE re-mints the GUIDs of a microflow whose # pre-existing error it leaves in place, and that must still read as pre-existing. A rename # does read as new, and restores — the safe direction. Hex tokens: no "|" or glob characters, -# whatever the message says. +# whatever the message says. ERRKEY_PY is the one definition of that key: err_set and +# err_report both prepend it, so "new" in the verdict and "NEW" in the report cannot drift. +ERRKEY_PY='import hashlib +def errkey(x): + locs = sorted("\x1e".join(str(l.get(k) or "") for k in ("module", "document", "element")) + for l in (x.get("locations") or [])) + raw = "\x1f".join([x.get("errorCode") or "", x.get("message") or ""] + locs) + return hashlib.sha1(raw.encode("utf-8")).hexdigest()[:16]' err_set() { [ -n "$PY" ] || { echo ""; return; } - "$PY" - "$(native_path "$1")" <<'ERRSETPY' 2>/dev/null || echo "" -import hashlib, json, sys + { printf '%s\n' "$ERRKEY_PY"; cat <<'ERRSETPY'; } | "$PY" - "$(native_path "$1")" 2>/dev/null || echo "" +import json, sys d = json.load(open(sys.argv[1])) -keys = [] +print('|'.join(sorted(errkey(x) for x in d.get('problems', []) if x.get('severity') == 'Error'))) +ERRSETPY +} + +# err_report count|print — splits the post-exec errors against BASELINE_SET, +# counted the same way is_subset_of counts. Only when the baseline was measured +# (BASE_KNOWN=1); otherwise every error is listed as before, with no new/old split. +# count → " ", or "?" when the baseline is unknown. +# print → the NEW errors in full (code, message, where, expression detail), then the +# pre-existing ones as one line per code + message with a count. +# Why (field run, 2026-09-29): a restore on a 22-error model printed "22 error(s) found" +# and all 22 with locations; the one the script added was not marked, and the count was +# the same as before the script ran. +err_report() { + [ -n "$PY" ] || { [ "$1" = count ] && echo "?"; return; } + { printf '%s\n' "$ERRKEY_PY"; cat <<'ERRREPPY'; } | "$PY" - "$1" "$(native_path "$2")" "$BASELINE_SET" "${BASE_KNOWN:-0}" 2>/dev/null || { [ "$1" = count ] && echo "?"; } +import json, sys +from collections import Counter +mode, known = sys.argv[1], sys.argv[4] == '1' +d = json.load(open(sys.argv[2])) +left = Counter(k for k in sys.argv[3].split('|') if k) +new, old = [], [] for x in d.get('problems', []): if x.get('severity') != 'Error': continue - locs = sorted('\x1e'.join(str(l.get(k) or '') for k in ('module', 'document', 'element')) - for l in (x.get('locations') or [])) - raw = '\x1f'.join([x.get('errorCode') or '', x.get('message') or ''] + locs) - keys.append(hashlib.sha1(raw.encode('utf-8')).hexdigest()[:16]) -print('|'.join(sorted(keys))) -ERRSETPY + k = errkey(x) + if known and left[k] > 0: + left[k] -= 1 + old.append(x) + else: + new.append(x) +if mode == 'count': + print('%d %s' % (len(new), ','.join(sorted(set(x.get('errorCode') or '?' for x in new)))) if known else '?') + sys.exit(0) +if known: + print(' NEW, from this script (%d):' % len(new)) +for e in new: + print(' ', e.get('errorCode', '?'), e.get('message', '')) + for loc in e.get('locations', []) or []: + print(' at', loc.get('module', '-'), '/', loc.get('document', '-'), '/', loc.get('element', '-')) + md = e.get('metadata') or {} + if md.get('expressionErrors'): + print(' expression:', md['expressionErrors']) +if old: + print(' Already there before this script, not why it was undone (%d):' % len(old)) + for (code, msg), n in sorted(Counter((e.get('errorCode', '?'), e.get('message', '')) for e in old).items()): + print(' ', code, msg, ('x%d' % n) if n > 1 else '') +ERRREPPY } # is_subset_of — both are err_set's "|"-joined sorted key lists. @@ -638,6 +683,7 @@ fi # "this script broke it" from "it was already broken". Costs one extra mxbuild; # skip with SKIP_BASELINE=1 when you know the tree is clean. BASELINE_SET="" +BASE_KNOWN=0 if [ "${SKIP_BASELINE:-0}" != "1" ] && [ -x "$MXBUILD" ] && [ -x "$JAVA_EXE" ]; then echo "→ Pre-flight: checking whether the model already has errors..." _BF=$(mktemp /tmp/mxbuild-baseline.XXXXXX) @@ -664,6 +710,7 @@ if [ "${SKIP_BASELINE:-0}" != "1" ] && [ -x "$MXBUILD" ] && [ -x "$JAVA_EXE" ]; [ "$_BC" = "0" ] && [ "$MXTK_MXBUILD_EXIT" -ne 0 ] && _BC="?" fi rm -f "$_BF" + case "$_BC" in ''|*[!0-9]*) ;; *) BASE_KNOWN=1 ;; esac if [ "$_BC" = "?" ]; then echo " ⚠ Baseline NOT measured: mxbuild exited $MXTK_MXBUILD_EXIT without checking the model." [ -n "${MXTK_MXBUILD_WHY:-}" ] && echo " mxbuild says: $MXTK_MXBUILD_WHY" @@ -818,8 +865,22 @@ if [ -x "$MXBUILD" ] && [ -x "$JAVA_EXE" ]; then rm -f "$ERRORS_FILE" else GATE_STATE="fail" - echo " ✗ mxbuild: $CE_COUNT error(s) found — restoring snapshot to avoid loading a corrupt MPR." CE_CODES=$(err_codes "$ERRORS_FILE") + read -r NEW_COUNT NEW_CODES </dev/null || true -import json, sys -d = json.load(open(sys.argv[1])) -for e in d.get('problems', []): - if e.get('severity') != 'Error': - continue - print(' ', e.get('errorCode', '?'), e.get('message', '')) - for loc in e.get('locations', []) or []: - print(' at', loc.get('module', '-'), '/', loc.get('document', '-'), '/', loc.get('element', '-')) - md = e.get('metadata') or {} - if md.get('expressionErrors'): - print(' expression:', md['expressionErrors']) -PYEOF + # With a measured baseline, the NEW errors come first and the rest are summarised. + err_report print "$ERRORS_FILE" || true cp "$ERRORS_FILE" "$LAST_ERRS" echo " → Full error detail: .mpr-snapshots/last-mxbuild-errors.json" @@ -870,7 +920,16 @@ PYEOF rm -f "$BASE_ERRS" echo "" - if [ "$BASE_COUNT" != "0" ] && [ "$BASE_COUNT" != "?" ]; then + if [ "$NEW_COUNT" != "?" ] && [ "$BASE_COUNT" != "0" ] && [ "$BASE_COUNT" != "?" ]; then + # The baseline was measured and this script added errors on top of it. The old + # branch below said "NOT the cause. Nothing will exec until that is cleared" here, + # which was false twice over: the script added the error, and the delta gate keeps + # any later script that adds nothing new (field run, 2026-09-29). + echo " → Undone: this script added $NEW_COUNT new error(s) [$NEW_CODES]. Fix those and re-run." + echo " The model is back to its $BASE_COUNT earlier error(s) [$BASE_CODES]. Those did not cause this; a script" + echo " that adds nothing new is kept even while they remain." + log_build "❌ gate failed" "$NEW_COUNT new error(s): $NEW_CODES — script rolled back; $BASE_COUNT pre-existing remain" + elif [ "$BASE_COUNT" != "0" ] && [ "$BASE_COUNT" != "?" ]; then echo " ⚠ PRE-EXISTING: the restored model already fails with $BASE_COUNT error(s) [$BASE_CODES]." echo " This script is NOT the cause. Nothing will exec until that is cleared." case "$BASE_CODES" in