active_inference_power is a reproducible, investigator-facing simulation
suite for statistical power in adaptive studies that embed an active-inference
agent. Before simulation, the investigator declares the agent-side generative
model, evaluator-side generative process, testing setting, policy, and
replication plan. The agent then acts only on its visible history; the evaluator
retains simulated hidden truth solely to score traces after the fact. Power is
therefore a conditional operating characteristic of the declared design—not a
universal property of an agent or of its task environment.
The fixed-horizon method and calibration lanes compare declared observation streams, evidence objects, and correction procedures; they do not by themselves instantiate an action-controlled agent--task-environment loop. The action-loop lane adds an embedded agent that acts under uncertainty within a fixed synthetic process, while the investigator estimates operating characteristics across replications. In neither lane is power a capability score for the agent.
The action-loop process is a synthetic task environment. “Simulated niche” is only informal shorthand for that declared process; it does not denote an empirical ecological, social, or deployment population, and the scenario grid is not a distribution over such niches. See the evaluation perspective for the role boundary and the exact action-loop estimand.
The suite includes a tested family of correction procedures, replication-level
FDP, true-positive-rate, and any-false-rejection-event scoring aggregated as
FDR, power, and FWER, analytic z/t power, the Genovese--Wasserman BH fixed
point, seeded dependence-aware simulations, real inferactively-pymdp==1.0.3
inference, and action-in-the-loop environments with adaptive sensing, target
selection, costs, latent context transitions, and stopping.
The project’s central rule is simple: a posterior belief threshold is not a family-level error guarantee. Under the declared null, the observed stream is mapped to a calibrated Binomial p-value before BH is applied in the fixed-horizon reference; the action-loop and sequential tracks use separately declared evidence contracts and validity labels.
In the action loop, a target-specific e-process crossing can support only the declared per-target optional-stopping interpretation. It is not an FWER or FDR guarantee for the target family; the all-action likelihood-ratio product is a diagnostic display, and finite observed FDR/FWER values are estimands rather than control claims.
Prepare the locked development environment once. The dev extra supplies the
test, coverage, lint, and type-check tools used below; the optional demo
extra is needed only for the Streamlit application.
env -u VIRTUAL_ENV uv sync --locked --extra devStart with a fresh, disposable fast-profile bundle rather than overwriting a retained result or campaign receipt:
quickstart_parent="$(mktemp -d "${TMPDIR:-/tmp}/active_inference_power_fast.XXXXXX")"
quickstart_root="$quickstart_parent/results"
./.venv/bin/python scripts/generate_results.py --profile fast --workers 1 \
--checkpoint-every 500 --output-root "$quickstart_root"
./.venv/bin/python scripts/audit_figure_accessibility.py --output-root "$quickstart_root"
./.venv/bin/python scripts/smoke_demo.py --payload "$quickstart_root/demo/payload.json"This smoke bundle exercises the real project contracts but is not a release or publication candidate. For a full isolated candidate and the PDF/HTML review sequence, use the quickstart and rendering pipeline.
For the stronger non-release fast-bundle check used by continuous integration, keep the same fresh root and add a passing quality receipt and strict hydration before running the read-only integrity audit:
./.venv/bin/python scripts/generate_quality_report.py \
--output-root "$quickstart_root"
./.venv/bin/python scripts/z_generate_manuscript_variables.py \
--output-root "$quickstart_root"
./.venv/bin/python scripts/audit_manuscript_presentation.py \
--variables-file "$quickstart_root/data/manuscript_variables.json" \
--resolved-manuscript-dir "$quickstart_root/manuscript"
./.venv/bin/python scripts/audit_fast_bundle.py --output-root "$quickstart_root"This check verifies the fast profile, current-source witnesses, result-derived files, figure registry, quality receipt, and manifest bytes. It neither renders nor marks a bundle complete, so it remains distinct from a release candidate.
Validate source changes with:
./.venv/bin/python -m compileall -q src scripts tests examples
./.venv/bin/ruff check src scripts tests examples
./.venv/bin/python -m pytest tests/ --cov=active_inference_power --cov-branch --cov-fail-under=95
env -u VIRTUAL_ENV uv run mypy src scripts tests examples --ignore-missing-importsAt the selected output root, analysis writes results.json, certificates.json,
the configured publication PNGs plus figures/figure_registry.json,
review/dashboard.html, and manifest.json.
The registry carries the caption/alternative-text pair plus reader question,
visible comparison, takeaway, visual encoding, and validity metadata; a
mixed-status multi-panel figure also records panel-specific validity so an
exception cannot be hidden by a whole-figure label.
It also writes formalism_registry.json, which connects formal objects
to public symbols, tests, assumptions, evidence classes, and non-claims.
It also writes run_log.json with stage durations, seeds, and
replication accounting. The result schema includes a global-null p-value
calibration audit and conservative Monte-Carlo precision budgets.
Release-derived files under output/ are generated and should be regenerated
rather than hand-edited. Do not discard the named preserved campaign,
precision, and campaign-v2 receipts: they are hash-bound historical evidence
with their own retention rules, not disposable scratch output.
The manuscript title page uses the source artwork at
manuscript/assets/active_inference_power_cover.png. Its provenance and role
are documented beside the asset; it is publication design, not scientific
evidence.
Start with the quickstart, then use the documentation hub for the package contract, architecture, the investigator/agent/task-environment distinction, action-loop model/process/setting contract, research roadmap, and failure boundaries. The release review guide distinguishes read-only verification from rendered-bundle finalization; the FAQ explains the scientific scope and retained campaign failure.
This is an orientation map rather than an exhaustive module inventory. The
public exports are defined in
src/active_inference_power/__init__.py;
the formalism-to-code map and
package contract provide the complete audited
symbol, test, and boundary cross-reference.
src/active_inference_power/corrections.py: Bonferroni, Šidák, Holm, Hochberg, BH, BY, Storey, a conservative fixed-lambda null-count-bound diagnostic, adaptive BH, and fixed-weight weighted BH.src/active_inference_power/error_rates.py: confusion counts, FDP, TPR, FWER, and aggregation.src/active_inference_power/power.py: analytic z/t power, sample-size inversion, and BH asymptotics.src/active_inference_power/simulation.py: two-groups Gaussian model and paired Monte Carlo.src/active_inference_power/active_inference.py: real pinnedpymdpsequential testing plus calibrated evidence p-values.src/active_inference_power/certificates.py: seeded numeric certificates with explicit bounds.src/active_inference_power/dependence.py: independent, negative, block, and factor regimes.src/active_inference_power/sequential.py: fixed-horizon p-values and time-uniform e-processes.src/active_inference_power/online_fdr.py: LORD++, SAFFRON, and e-LOND online schedules.src/active_inference_power/computational_complexity.py: deterministic source-work proxies, declared action-trace accounting, and validation for an isolated host-specific timing profile.src/active_inference_power/selection.py: visible-filtration selection traces and independent split-stream confirmation.src/active_inference_power/learned_policy.py: deterministic train/evaluate policy artifacts with frozen hashes.src/active_inference_power/agent_sampling.py: realpymdppolicy/stopping traces.src/active_inference_power/scenario.py: composable generative-model, generative-process, and testing-setting contracts with stable hashes; policies are evaluated separately against those contracts.src/active_inference_power/action_loop.py: action/observation/belief/ evidence/cost/stopping traces and conditional power summaries.src/active_inference_power/suite.py: scenario surfaces and adaptive target-selection accounting.src/active_inference_power/action_statistics.py: replication-level true/false rejection accounting, mean-TPR power and separately named any-true-detection rates, cost/stopping intervals, power-per-cost summaries, and paired policy contrasts.examples/: deterministic offline reference, environment-control, and adaptive-target examples.src/active_inference_power/demo.py: optional Streamlit review application.src/active_inference_power/benchmarks.py: checksum-verified offline benchmark adapters.src/active_inference_power/figures/: deterministic publication figures, scenario-aware captions, and a lightweight dashboard.src/active_inference_power/formalism.py: executable formalism-to-code traceability registry and generated manuscript rows.
The public API is re-exported from src/active_inference_power/__init__.py, so import active_inference_power works
without test-specific path manipulation. See
docs/PACKAGE_CONTRACT.md for invariants.
The publication PDF/HTML renderer is a separately provisioned controlled release environment, not a public package dependency. This repository reproduces results, figures, and the hydrated manuscript; the immutable release PDF and accessible HTML are distributed as release assets. See the controlled-renderer boundary and the maintainer rendering pipeline.
Version 1.0.0 is released at
ActiveInferenceInstitute/active_inference_power
and its exact GitHub release is
v1.0.0.
Use the concept DOI 10.5281/zenodo.21695160
for a citation that follows the release family, and the immutable version DOI
10.5281/zenodo.21695161 for this
specific v1 archive. The matching Zenodo record is
21695161. The source-bound release
chain and its deliberately bounded scientific claims are documented in
docs/publication_metadata.md.
Project source is licensed under the MIT License. The compact
offline benchmark fixtures retain their own source provenance and license
contexts; they are not relicensed by the project MIT grant. See
THIRD_PARTY_NOTICES.md and the checksum-pinned
data/benchmarks/manifest.yaml before
redistributing or extending those fixtures.
The Monte-Carlo certificates are evidence under stated assumptions, not formal
proofs. Negative and block dependence are stress regimes; BY is the
arbitrary-dependence comparator, not a simulation proof. Fixed-horizon
posterior thresholds remain distinct from e-process optional-stopping claims.
The fixed-horizon binary model remains the numerical reference baseline. The
action-loop suite is deliberately conditional: environment-level control,
adaptive target selection, and process shifts are stress/operating-characteristic
tracks unless the declared conditional evidence assumptions are satisfied. The
power-cost frontier is a conditional decision-design comparison, not a policy
optimality or utility theorem. The manuscript’s
docs/scholarship_map.md records how active-inference
process theory, discrete-state inference, multiple-testing theory, and the
sequential martingale/e-value literature map onto implemented methods and
explicit non-claims.
The release's computational-work ledger is a version-specific source-work
accounting aid, not a processor, memory, or timing benchmark. It keeps the
dimensionless cost_unit and
cost_scale_id agree. An optional docs/formalism_to_code.md, and the release claim
review is recorded in docs/claim_audit.md.
The staged research agenda is documented in
docs/research_roadmap.md. Stage 1—predictable
adaptive target selection has an independent split-stream confirmation path.
The formal conditional split-confirmation result uses a separately declared,
hash-bound static channel outside the operating grid; an action-loop example's
split result remains a context-specific conditional diagnostic. Chronological
selection remains diagnostic; richer process dynamics, learned policies, and
external benchmarks remain separately labeled conditional or stress tracks.
The original canonical campaign profile produced the preserved
output/campaign/ receipt for a broader effect/dependence/reliability surface.
That all-terminal campaign completed its simulations but failed the
same-vector Storey FDR certificate, so manifest.complete remains false.
It is retained as a method-sensitivity finding; never reuse that output root
or promote it into the publication manuscript.
The separately frozen v2 configuration at
data/research_lanes/conditional_power_campaign_v2.yaml
has a retained canonical receipt at output/campaign-v2/. Its campaign-specific
policy keeps BH FDR, Bonferroni FWER, positive dependence, and independent
split-stream Storey confirmation terminal, while retaining the same-vector
Storey power-gain record as a visible diagnostic. The v2 receipt is complete
for that frozen campaign contract only: it neither changes the canonical
release's all-terminal policy nor promotes a release or publication claim. See
docs/storey_certificate_decision.md
for the exact decision and receipt boundary. It is historical, hash-bound
campaign evidence rather than a source-current release candidate: a
source-contract refresh of this exact frozen lane first archives its prior
canonical receipt, and a genuinely new campaign requires a new reviewed
configuration and a previously absent output namespace.
The package also exposes storey_qvalue_conservative. It uses a fixed lambda
and an exact one-sided binomial upper confidence bound for the null count,
preventing the zero-null plug-in degeneracy seen in tiny families. Because the
bound is estimated from the same p-values used for rejection, this is an
assumption-gated diagnostic. An independent split avoids reusing the selection
stream but supplies only the declared conditional calibration evidence, not a
general finite-sample FDR theorem.