Skip to content

Replay/adjoint linearises about a state the forward run never visited: snapshots restore fields but not parameter values #773

Description

@lmoresi

What

model.save_state() / load_state() restore fields but not the values of uw.expression parameters. TranscriptAdjoint replays a run by restoring each step's snapshot and re-solving, so any parameter that changed later in the run is replayed at its final value, not the value that step actually used.

The adjoint then assembles dR/du about a state the forward run never visited, and returns a gradient for a different problem. No error, no warning.

Demonstrated

st.solve()                       # eta = 1.0
snap = model.save_state()
eta.sym = sympy.Float(99.0)      # a later-step parameter change
model.load_state(snap)
eta.sym                          # 99.0   <-- not 1.0
at save_state          : eta = 1.00000000000000
after change           : eta = 99.0000000000000
after load_state(snap) : eta = 99.0000000000000

Why this is severe

Any run that changes a parameter mid-run is affected, and those are exactly the runs people take adjoints of:

  • a continuation or homotopy ladder (the parameter IS the ladder);
  • a ramped boundary condition or body force;
  • time-dependent material properties;
  • an inversion that steps a control between iterations.

The failure is silent and plausible: the replayed solve converges, the transcript records it as normal (see #772 — the transcript never records parameter values either, so nothing in the record contradicts it), and the gradient is simply wrong. A Taylor test done at the END of the run would pass, because there the final value IS the current value; the error only appears for steps before the last change.

Relationship to #772

#772 is the reporting half: the transcript fingerprints the symbolic form and never the values, so an invisible parameter change leaves no trace. This is the correctness half: replay cannot restore what was never recorded. Recording the change (the preferred fix in #772) is also what makes replay correct, since the replay can then re-apply each step's parameter values before re-solving.

They should probably be fixed together: log expr.sym assignments as transcript events, and have load_state / the replay path re-apply the values in force at that step.

Acceptance

A test that: solves at p = a, snapshots, changes to p = b, replays the snapshot, and asserts the replayed solve used a — and a TranscriptAdjoint gradient over a run with a mid-run parameter change that passes a Taylor test at an EARLY step, not only the last one.

Activity

  1. lmoresi commented on Oct 7, 2026

    @lmoresi
    MemberAuthor

    Not reproducible on development (7c9cbbfd) — the adjoint API is not there:

    uw.adjoint present:        False
    Stokes adjoint API:        []
    files with 'adjoint' in the path: 0
    

    Only scattered references remain (one line in petsc_generic_snes_solvers.pyx, five in mcp/__init__.py, eight in transcript_query.py), not the implementation. So this defect lives in unmerged work — feature/discrete-adjoint (#744, currently conflicting and red), feature/adjoint-rotated-bc (#751, based on #744), or scratch/adjoint-merge.

    That is why it has sat untouched: there is nothing on the main line to fix. It is not fixed-in-PR — the opposite. It has to be fixed on the branch before that branch lands, or it lands with the defect.

    Checked in the untouched-issue triage. Blocked behind #744, which needs its conflicts resolved and a CI re-run (its red check is test_1063_constrained_traction, whose timeout #814 addressed when it merged on 2026-10-06).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions