Skip to content

Testing

github-actions[bot] edited this page Oct 2, 2026 · 17 revisions

Testing

wolfTrust separates host behavior, architecture-accurate emulation, and physical-hardware evidence. A result from one environment must not be reported as a result from another.

Validation layers

Environment What it validates What it does not validate
Native host State machines, manifests, IPC ownership, copied transfers, services, storage, crypto integration, recovery decisions, and negative inputs Cortex-M exception return, CMSE, SAU, MPU, GTZC, or physical flash behavior
M33MU Cortex-M33 instruction flow, TrustZone transitions, CMSE gateway calls, MPU faults, guest scheduling, authenticated boot, and target service interactions STM32H563 peripherals, real option bytes, WRP, ST-Link, or silicon timing
STM32H563 The actual NUCLEO-H563ZI boot chain, Secure/Non-secure attribution, faults, flash, UART, and selected end-to-end behavior Other devices, other provisioning states, peer-flash confidentiality, or adversarial peripheral and Non-secure NVIC ownership

Host tests

Run the complete native set:

make test

The suite list is generated by:

make -s -C tests/host print-suites

Current suites cover domain and manifest validation, lifecycle, guest verification, rollback decisions, IPC and FF-M behavior, SPM policy, gateway vectors, Secure Partition layout and recovery, crypto-engine relay and key isolation, vault and storage services, attestation and COSE integration, firmware update, runtime remeasurement, linked Secure layout, VNET, public PSA headers, boot-handoff record consumption, and negative paths. The attestation IAK suite runs wolfHSM NVM with both the default 8-byte and STM32H5 16-byte flash programming units.

Additional host checks:

make test-compilers
make test-sanitize
make test-valgrind

Valgrind requires the tool on the host. make test-compilers reruns all suites with the selected CC (default cc); it does not select multiple compilers. Run it once per compiler, for example make test-compilers CC=clang. Sanitizer support depends on the local toolchain.

LTO validation

Every Secure link generates wolftrust.map and runs the linked-image layout check. Compare optimized and diagnostic builds with separate output trees:

make BUILD_DIR=build-lto size-report
make BUILD_DIR=build-no-lto WT_LTO=0 size-report

The default M33MU and hardware commands exercise the LTO image. A release-size result is valid only when the applicable target scenarios pass without a fault marker. The cross-compilation workflow rebuilds one output tree from explicit WT_LTO=0 to the default WT_LTO=1, verifies the build stamp and object formats change, checks that isolation-critical objects remain non-LTO, and requires the optimized ELF to be smaller. It also injects a forbidden heap symbol twice to prove a rejected ELF, flat binary, map, and CMSE import library are deleted instead of being reused by the next Make invocation.

The per-PR core/port split workflow builds the CONFIG_VNET=y Secure image. The M33MU smoke tier builds and runs the WT_CONFORMANCE=1 confboot layout on every PR; the VNET layout runs with the full matrix (ci:h5 label, push to main, nightly), so both optional isolation-band configurations stay under the linked-image check in CI.

PSA FF conformance

make test-conformance

This target fetches the pinned Arm PSA architecture tests. When M33MU is available it runs the target FF-M IPC suite. Without M33MU it runs only the host-side client and policy subset and prints explicit warnings. Do not treat that fallback as target conformance.

The dedicated guest configurations also exercise the PSA Crypto, Storage, and Initial Attestation validation applications. Those are target scenarios, not part of a plain make test result.

M33MU scenarios

The baseline target command is:

make test-target

The engine defaults to native. Set WT_ENGINE to exercise the same target path with either backend:

WT_ENGINE=native make test-target
WT_ENGINE=hsm make test-target

It runs the port's smoke tier (STM32H563: positive, gtzcneg, crossdomain, bothpsa, confboot, devcrypto; WT_TIER=full runs every scenario). Detection accepts m33mu on PATH or a path in M33MU. If the emulator is unavailable, the target reports a skip rather than a pass.

Additional focused runs use:

WT_ENGINE=native tests/target/run_m33mu_scenario.sh positive
WT_ENGINE=hsm tests/target/run_m33mu_scenario.sh positive

The runner's usage output is the authoritative scenario list. It includes authenticated-boot failure, rollback, runtime remeasurement, Secure Partition recovery, key and vault isolation, storage recovery, attestation negatives, firmware update, manifest rejection, GTZC behavior, and VNET paths.

Five scenarios cover processor-state isolation, and the emulator proves less than their names suggest:

  • fpneg proves containment only. A floating-point instruction in the SERVICE_HSM partition takes the NOCP UsageFault and does not escalate. M33MU ends the run when it raises NOCP, so partition restart and guest survival are not shown here; the STM32H563 fpneg run checks them.
  • sealneg and sealhaltneg prove only the SPM's software check. In sealneg a partition overwrites its own stack-top seal, that partition alone faults at its next resume, and the guests keep running. In sealhaltneg a partition's seal is overwritten before its first dispatch (the same check runs on every dispatch) and the platform halts before that partition runs. No architectural unstack fault is exercised, on the emulator or by the STM32H563 sealneg run; M33MU models neither the seal nor the function return integrity check.
  • sealbootneg damages one main-stack seal word in the reset path; the boot halts on the production panic before any partition or guest runs.
  • sealpivotneg issues a partition's blocking wait with the stack pointer parked on its stack top, so the exception frame lands on the seal words; that partition alone faults and restarts, and the guests keep running.

bandneg1 through bandneg6 prove the vault, attestation, and crypto partitions cannot reach each other's data band:

Scenario Prober Band touched
bandneg1 crypto vault
bandneg2 crypto attestation
bandneg3 attestation vault
bandneg4 attestation crypto
bandneg5 vault attestation
bandneg6 vault crypto

The prober reads the band, is restarted, writes the band, and is restarted again. Both accesses must fault on the prober's own stack, and the full positive lifecycle must still complete. The crypto and attestation probers also confirm the keystore services they do not own are refused.

restartneg1, restartneg2, and restartneg3 prove a restarted crypto, attestation, or vault partition starts from its band's link-time image. The partition changes initialized and zero-initialized state in its own band and faults; its restarted instance faults again if either value survived.

manifestneg removes a required feature, manifestneg2 declares isolation level 2, and manifestneg3 composes a partition table that reaches another partition's band. Each must halt the boot before anything is scheduled.

Each numbered probe variant is its own matrix row; CI packs each family into one job.

Three more cover the SPM's own fault handling:

  • mspovfneg pushes on the Secure main stack in the reset path until MSPLIM_S raises STKOF. M33MU escalates the entry-time STKOF to HardFault and ends the run there without executing the handler, so the emulator proves only the limit; the STM32H563 run reads the SPM fault latch and checks that no guest ran.
  • xnneg makes the privileged SVC gate call a thunk copied into SPM .bss while a partition thread domain is installed; the execute-never cover faults the fetch. M33MU pends that synchronous fault instead of escalating it past the active SVC, so the emulator proves only the denied fetch; the STM32H563 run checks the SPM fault latch and the halt.
  • svcneg has the ITS partition issue the scheduler's internal guest-return SVC; the partition alone is panicked and restarted, and the lifecycle completes.

busfaultneg (the SERVICE_HSM partition reads an MPU-permitted window past the end of physical SRAM) and nsbusfaultneg (guest0 turns off its own MPU and reads an unmapped Non-secure peripheral hole; the monitor restarts it to its limit while guest1 runs) run only on the STM32H563: M33MU turns an unmapped data access into a MemManage fault and never vectors a data BusFault, so these scenarios have no emulator row until the pinned emulator models it.

VNET has convenience targets:

make test-vnet
make test-vnet-target

test-vnet is host-only. test-vnet-target launches two authenticated wolfIP guests under M33MU.

The wolfIP guests poll once per guest millisecond and idle between ticks so empty copied IPC calls leave time for the guest clock to advance. The vnet and vnetneg M33MU scenarios allow 180 wall-clock seconds for the initial ARP delay in guest time and the copied IPC round trip on shared runners. Both require the mediated ping reply and clean breakpoint exit; vnetneg also requires both isolation faults and recovery.

Engine matrix

The full CI scenario list adds engine: [native, hsm] as a matrix dimension. Every scenario runs with both engines except the wolfHSM-only ones listed in HSM_ONLY in tests/target/lib/scenario_matrix.py, which run under hsm only. The generator is the source of truth for the exact list: python3 tests/target/lib/scenario_matrix.py --port stm32h563 --tier full --json.

hsmattackneg drives the raw wolfHSM protocol from a compromised-guest probe. It checks that a forged wolfHSM client ID cannot select the attestation key and that a wolfHSM NVM-group packet cannot reach the rollback store. The native engine does not link the wolfHSM client wire, server, or message handlers, so that exact attack surface does not exist there. Native key and namespace behavior remains covered by the common positive, cross-domain, keystore, storage, attestation, and Crypto-validation rows.

Validation of the engine split completed under both engines with:

  • the applicable M33MU scenario matrix;
  • the Arm FF-M IPC suite at 85 passed, 4 heap-dependent tests skipped, and 0 failed, test for test as recorded in tests/target/ffm_ipc_results.txt;
  • the current dev_apis Crypto schedule at 64 passed, 13 skipped, and 0 failed (77 scheduled tests; c047 is configuration-skipped in addition to the upstream schedule); and
  • the STM32H563 positive, restart, cross-domain, and conformance hardware suite.

The engine dimension changes crypto dispatch, not what M33MU proves. Emulator results still do not establish STM32 attribution or physical flash behavior.

MIMXRT700 emulator scenarios

M33MU also models the i.MX RT700 (--cpu imxrt700), so the MIMXRT700 port has an emulator gate that mirrors its hardware runner:

make test-target TARGET=mimxrt700
tests/target/run_rt700_m33mu.sh positive
tests/target/run_rt700_m33mu.sh ahbscneg

positive boots the whole chain: wolfBoot verifies the wolfTrust image, and both bare-metal guests launch, reach the SPM through the SG veneers, and finish. ahbscneg adds the CPU isolation negative: guest 0 stores into guest 1's RAM window, which the per-dispatch SAU window keeps Secure, so the SAU refuses the store; the monitor contains the fault, relaunches guest 0 through its restart budget, quarantines it, and guest 1 keeps running. The runner asserts every step from the emulator log (the guest console lines and M33MU's protection-unit trace), never from a debugger.

The runner also carries the shared Secure-verdict negatives from tests/target/lib/scenario.sh, the per-scenario table both the STM32H563 and MIMXRT700 runners draw their Secure-image probe flags from. Each one ends on the verdict breakpoint its probe emits, asserted port-independently: rollbackneg (a downgraded boot is refused fail-closed and no guest enters a domain), remeasureneg (guest 0 measures clean at launch verification, its image is then tampered in flash, and the on-demand re-measure before dispatch must catch it and quarantine the domain), manifestneg (a corrupted manifest halts boot before anything is scheduled), and spbudgetneg (restart-budget exhaustion escalates to the fail-closed platform recovery). crossdomain and keystoreneg prove the Secure Partition MPU domains: an unprivileged storage-SP read of SPM-private RAM, or of the shared keystore band, MemManage-faults at the port's band address (shown by a second traced boot), the SP's wake never serves the guests' storage connect, and the guests' own lifecycle rides it out. spfaultneg and panicneg prove Secure Partition recovery: the crypto relay runs an undefined instruction on its first entry, or the storage SP closes an error-status handle, which the SPM must panic it for. Either way the SP UsageFaults exactly once, the SPM restarts it in place, the restarted SP serves the guests that follow, and both guests finish with no escalation. fpneg proves containment only on M33MU: the relay's FP instruction takes the NOCP UsageFault and M33MU ends the run there. The STM32H563 hardware run also checks partition restart and guest survival. restart makes guest 0 read Secure RAM on every launch: the SAU refuses it, the monitor relaunches guest 0 through its restart budget and quarantines it, and guest 1 runs on. authneg flips one byte of guest 0's image after its digest was pinned, so launch verification refuses guest 0 while guest 1 boots and runs normally.

The remaining scenarios run the portable PSA test guest (tests/firmware/psa-guest/) in both Non-secure windows. It is the STM32H563 Zephyr guest's client lifecycle with no operating system underneath: wolfPSA over the SPM-mediated client for the PSA Crypto API, the OS-neutral FF-M, storage, and attestation clients, and the wolfCOSE verifier for the token. A port supplies a linker window and a console (boards/<target>/), nothing else. bothpsa runs the whole lifecycle from both guests in one boot: the mediated SHA-256 KAT, ITS and sealed PS set/get, volatile P-256 key-ops with a cross-key refusal, RNG, AES-CTR, and an attestation token verified against the IAK public key and wolfBoot's measurement of the signed Secure image. bothiso proves the SPM rejects a forged handle, an oversized vector, a vector inside the peer guest's window, and an unknown SID from both guests. attestneg adds the attestation negatives (invalid requests refused with the statuses Arm's tests expect; tampered and misattributed tokens fail the guest verify). fwustage stages a candidate into the wolfBoot update partition through SERVICE_FWU, arms it, and proves reject/clean restore READY. hsmattackneg (hsm engine only, it drives the raw wolfHSM client wire) proves a forged client id cannot reach the IAK and an NVM-group packet never reaches the server.

The conformance scenarios are the same drop-in proof the STM32H563 gives: confboot hosts Arm's unmodified psa-arch-tests FF-M IPC suite in the PSA guest against the conformance Secure image (WT_CONFORMANCE=1, the manifest that adds Arm's server, driver, and client partitions) and expects the same 85 passed, 4 heap tests skipped, 0 failed; devstorage, devcrypto, devattest, and devattestqcbor run the dev_apis storage, crypto, and initial-attestation suites; vaultrecover proves a foreign vault pool self-heals on a development device (the crypto suite passes after the reformat), and vaultrecoversec forces the SECURED lifecycle and proves the refusal is graceful on the emulator: no fault, and the guest still starts (the vault-not-wiped and attestation-degraded counters are hardware-runner evidence, read over the debug port as on the STM32H563). The port supplies port/mimxrt700/manifest-conformance.json and port/mimxrt700/conformance/ (the PAL bindings: console, watchdog, the flash-backed NVMEM boot flag, and the isolation MMIO windows); the val framework, its PAL, and the test lists are the upstream sources the Secure build fetches and generates.

The runner builds its own pinned emulator and wolfBoot first stage. The emulator is upstream M33MU at M33MU_REF, unpatched. Its Secure AHBSC SRAM rules apply to CPU0 as documented, which is stricter than the EVK measured. wolfBoot is WOLFBOOT_REF built from config/examples/imx-rt700-tz.config plus tests/target/wolfboot-imxrt700-lifecycle.patch (the RT700 HAL's PSA lifecycle hook, which the attestation service needs; it goes once wolfBoot carries it), linked at the NOR base, the same offset the wolfBoot emulator tests use, so building it needs the MCUXpresso SDK or DFP like any RT700 wolfBoot build; set RT700_WOLFBOOT_DIR to reuse an existing one (the runner stops if that tree has no wolfboot.bin, rather than replace it). WT_ENGINE selects the crypto engine as for any build. A console boot ends on the emulator's wall-clock budget because the guests idle once done, which M33MU reports as exit status 127; a Secure-verdict boot stops on its breakpoint with status 0, as does a traced boot at its first delivered fault; any other status fails the run.

These are emulator results. They prove the SAU attribution and the monitor's containment on a faithful core model; the silicon ahbscneg run on the EVK is recorded separately.

STM32H563 hardware

The published hardware run requires:

  • a Linux host with Bash and GNU userland;
  • Docker for the documented container build;
  • Python 3 on the host for the flash-phase measurement check;
  • STM32_Programmer_CLI, found at its default install path or through STM32_CLI;
  • pyOCD with STM32H563 support;
  • a host-visible arm-none-eabi-nm, or an override in ARM_NM;
  • a serial VCP, default /dev/ttyACM0 or H5_SERIAL; and
  • a connected ST-Link.

The detector checks the board/programmer path, not every required host tool. It verifies the ST-Link USB ID only when lsusb is available. Without lsusb, a missing probe can be reported later as a flash failure rather than an initial skip.

Run the default hardware set with:

WT_ENGINE=native \
WT_H5_DOCKER_IMAGE=ghcr.io/wolfssl/wolfboot-ci-m33mu:v1.15 \
make test-hardware
WT_ENGINE=hsm \
WT_H5_DOCKER_IMAGE=ghcr.io/wolfssl/wolfboot-ci-m33mu:v1.15 \
make test-hardware

The target skips if hardware detection fails. It builds and flashes the positive, restart, cross-domain, and conformance scenarios by default. The suite wrapper forwards WT_ENGINE into its build container. Select a narrower set with WT_H5_SCENARIOS:

WT_H5_SCENARIOS="positive bootupdate" \
WT_H5_DOCKER_IMAGE=ghcr.io/wolfssl/wolfboot-ci-m33mu:v1.15 \
make test-hardware

The optional gtzcneg scenario must be selected explicitly. It verifies the GTZC peer-RAM curtain after a Non-secure MPU bypass; it does not establish peer flash confidentiality or adversarial peripheral and Non-secure NVIC ownership.

periphneg has a privileged Non-secure guest read and clear the SPM's RNG through its Non-secure alias and run Non-secure GPDMA copies out of the Secure image and Secure SRAM. Nothing may be read or changed, and Secure entropy must still work before the normal lifecycle completes. periphspneg (M33MU only) has the storage partition read the SPM's RNG registers; no partition domain maps a peripheral the port does not assign, so the read must MemManage-fault.

For the hardened guest-flash configuration, explicitly forward the build flag into the container and repeat it for the host flash run:

docker run --rm \
    -e WT_GUEST_FLASH_WRP=1 \
    -v "$PWD":/workspace \
    -w /workspace \
    ghcr.io/wolfssl/wolfboot-ci-m33mu:v1.15 \
    bash tests/target/run_h5_hardware.sh build positive
WT_GUEST_FLASH_WRP=1 \
    tests/target/run_h5_hardware.sh flash positive

With that variable set, the flash runner clears WRP so it can program the guests, then re-applies WRP before boot. The Secure image independently reads the live WRP register and refuses an incompletely protected guest. The supported provisioning_ctrl.sh set-wrp workflow deliberately applies WRP only while the board is Open; that is a conservative project workflow, not the complete silicon rule. The hardware runner does not check product state before clearing or reapplying WRP. Current ST guidance makes WRP nonmodifiable in TZ-Closed, Closed, and Locked, and RM0481 separately defines the FLASH_WRPSGNxR.UNLOCK condition. The current suite wrapper does not forward this variable into its Docker build, so the shorter make test-hardware form does not build a WRP-enforcing Secure image.

Use the disposable container for Secure and guest builds. The current direct-host build path adds safe.directory '*' to the user's global Git configuration and is not suitable as a published workflow.

Hardware test commands write flash and can reset the board. Review STM32H5 Guide before running them.

CI coverage

The workflows under .github/workflows/ separately run:

  • host unit tests;
  • compiler variants, sanitizers, and Valgrind;
  • Cortex-M33 cross-compilation of both crypto engines;
  • dependency integration;
  • the core/port split guard and the docs guard (no internal-ledger or home-directory references in the published docs);
  • fuzz targets;
  • selected and nightly M33MU scenarios; and
  • the MIMXRT700 chain under M33MU's RT700 model (the RT700 scenario matrix).

Trigger routing

The host, compiler, sanitizer, Valgrind, cross-compilation, integration, and core/port split checks run on every pull request, including drafts. The fuzz target also runs on pull requests as a 60-second libFuzzer smoke pass; the nightly schedule and manual dispatch run the 600-second soak instead.

The M33MU workflow (wolfBoot plus both guest lifecycles, then a select job that picks each port's tier and one matrix job that runs exactly those scenario groups) is tiered. Every pull request runs each port's smoke tier on both crypto engines (STM32H563: positive, gtzcneg, crossdomain, bothpsa, confboot, devcrypto; MIMXRT700: positive, ahbscneg, crossdomain, bothpsa, confboot, devcrypto). The full matrix runs on a push to main, on the nightly schedule, on manual dispatch (with a port input), and on a pull request that carries the ci:h5, ci:rt700, or ci:all label. The scenario groups per port and tier are in tests/target/lib/scenario_matrix.py; make test-target TARGET=<port> runs the same smoke tier locally and WT_TIER=full every scenario.

Running M33MU off a pull request

To run a port's full matrix against a branch without a PR, dispatch the workflow:

gh workflow run m33mu.yml --ref <branch> -f port=mimxrt700

To run a single scenario locally, use tests/target/run_m33mu_scenario.sh <key> (for example positive or crossdomain).

The GitHub-hosted workflows do not establish a physical-board result. Hardware output must come from the STM32H563 runner attached to a board.

Interpreting failures

  • A host pass plus target failure usually points to architecture glue, image assembly, linker placement, or hardware policy rather than a neutral state machine.
  • A missing target prints SKIP. It is not a successful run.
  • Preserve the first failing marker and the generated log before rebuilding; target runners place detailed output under logs/ or the configured scenario log.
  • Confirm that Secure, guest, manifest, flash, and emulator addresses were built from the same configuration.

Clone this wiki locally