Skip to content

test(replay): replay a Predbat log forwards from a debug yaml - exact and simulated, with a plan timeline - #5363

Draft
chalfontchubby wants to merge 32 commits into
feat/log-replay-inputsfrom
tools/debug-replay-forward
Draft

chalfontchubby wants to merge 32 commits into
feat/log-replay-inputsfrom
tools/debug-replay-forward

Conversation

@chalfontchubby

@chalfontchubby chalfontchubby commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Opened by Claude on Rik's behalf. Draft: a working tool we expect to iterate on for a while. Stacked on #5427 (the Replay input: logging). Description updated 7 Oct 2026 to match the branch.

Status: a work in progress across this and later PRs

The replay is a tool we will keep improving over several PRs. It will not reproduce every situation exactly yet, and this PR does not try to close every gap: each case it misses tells us what to log or derive next. Exact replay needs a log with the Replay input: lines from #5427; older logs replay approximately. Known gaps are listed below and will be taken one at a time.

Summary

A debug yaml is one moment in time; the question in a bug report is usually "why did the plan do that at 09:25?" when the yaml is from 05:30. This replays a Predbat log forwards from a debug yaml, one live run at a time, re-planning wherever the live run re-planned, and compares what the replay planned with what the log says live planned.

cd coverage
./run_all --debug_file <predbat_debug.yaml> --replay_log <predbat.log> \
    [--replay_until HH:MM] [--replay_simulate] [--replay_chart out.png] [--override name=value]

Two modes

  • Exact (default): the SoC, clock and inputs come from the log on every run, so errors cannot compound. This is the fidelity check: with full inputs in the log it should match live line for line.
  • Simulated (--replay_simulate): the battery is stepped by Predbat's own prediction model using the PV and load that actually happened. This is for trying code or setting changes against real days, and its SoC error against the log measures the battery model itself. --override name=value changes a setting for a what-if run.

What it reads from the log

Clock, SoC and day counters from the normal log lines, and, where the log has them, Replay input: lines with the exact values the plan used: rates, PV forecast, load forecast, plan starting state, the car state, the inverter's programmed state and the load divergence. Those lines are added by #5427, which this PR is stacked on. Derived state the live loop rebuilds every cycle (the p90 guard, rate windows, a car charging outside its plan) is rebuilt with the live code rather than logged.

It carries across midnight, follows invalid-plan and sensor-triggered re-plans, and marks rows after a logged version change.

Older logs

Logs from before the Replay input: lines replay too, approximately, back to at least v8.8.13 (Dec 2024). The replay prints a note saying so. It reads the older SoC, day-total and re-plan line formats, and the all-inverter SoC total where a log has it. Where the exact lines are missing it takes the car's planned slots and flags from Car N charging plan is (since v5.1) and the Cars line (since v7.0). It takes the Intelligent dispatch list from Octopus slots changed (since v8.27.27), with the import rates rebuilt from it by the live code. The forecasts stay as the yaml had them.

Overnight 21:30 -> 09:15, 6 Oct, every Replay input: line stripped Before Now
Adopted plans identical / same first export start 5 / 14 of 66 20 / 66 of 66
Simulated SoC RMS 11.2% 2.0%

Real issue logs from v8.8.13, v8.18.7, v8.25.5, v8.32.14, v8.46.4 and v9.0.3 all parse, and a 20-hour v9.0.3 log (#2716) replays.

Output

  • Per re-plan: the export windows live adopted next to the replay's, and the candidate plans, with a summary count.
  • --replay_chart: actual SoC (and simulated SoC) against each plan's export target, plus a plan timeline per side in the web plan's terms (Chrg, HoldChrg, FrzChrg, Exp, HoldExp, FrzExp, car) and a lane for the car's planned charging.
  • The replay's own log carries Replay of run at ... markers so it can be read next to the live log.

Fidelity on a live system (Sigenergy, IOG, 6 Oct 2026)

Replay Mode Adopted identical Candidates identical Notes
21:30 -> 09:15 overnight, car dispatches, midnight exact 66 / 66 96 / 96 live and replay plan timelines identical
same simulated SoC 1.99% RMS
10:49 -> 12:50, after the inverter-state logging was deployed exact 11 / 11 13 / 13 every pass's base/best metrics and costs identical
same simulated SoC 0.35% RMS
08:30 -> 11:30, across two restarts exact 13 / 14 23 / 23 miss is the first run after a restart (below)

Known gaps

  • Restarts: a restart resets the plan in force; the replay carries the pre-restart plan, so the first keep-or-switch decision after a restart can differ. Fix planned: detect the restart in the log and reset the plan state.
  • apps.yaml changes applied by a restart are not followed; start a new replay from a yaml written after it.
  • Dynamic load baseline when the load status is "high" is not replayed yet.
  • Multi-pass runs: when an input changes mid-run (a car plugged in, a restart), live re-plans within the run; the replay folds the passes into one.
  • Older logs (e.g. the 3-inverter system in Predbat set to manual Demand but planned (and executed) to force export #5423, v9.3.6): the SoC is now totalled across inverters, but the forecasts and the plan in force come from the yaml, so plans after the first few re-plans are approximate.
  • Simulated mode gained more charge in 02:30-03:00 on 6 Oct than its own 5.5 kW charge rate allows in 30 minutes. The cause is not yet found (charging past the window edges around car slots is the leading guess), and until it is, the simulation can't be used to fit battery losses.

Not in this PR

Changes

  • apps/predbat/tests/replay_forward.py (new): log parser, replay loop, comparison, simulated mode, chart.
  • apps/predbat/tests/test_replay_forward.py (new): unit tests for parsing, every logged input, history shifting, midnight, the plan timeline and the car modelling.
  • apps/predbat/dummy_inverter.py (new): a simulated inverter and battery component.
  • apps/predbat/tests/test_single_debug.py: helpers shared with run_single_debug, whose behaviour is unchanged (debug_cases goldens pass), plus --override.
  • apps/predbat/unit_test.py: the replay flags and test registration.
  • docs/developing.md: how to use the replay.

Test plan

  • ./run_all --test replay_forward --test debug_cases
  • ./run_all --quick
  • coverage/run_pre_commit
  • Replayed the live logs above in both modes.

🤖 Generated with Claude Code

@chalfontchubby chalfontchubby changed the title test(replay): replay a Predbat log forwards from a debug yaml test(replay): replay a Predbat log forwards from a debug yaml - exact and simulated, with a plan timeline Oct 6, 2026
chalfontchubby and others added 29 commits October 7, 2026 11:42
A debug yaml is one moment; bug reports usually also carry the following hours of log. The new
--replay_log mode restores the yaml, then steps through the log run by run, setting the clock, SoC,
load/PV/import/export history and the inverter's export window from the log, re-planning on the runs
where the log re-planned, and printing the logged export windows next to the replayed ones.

run_single_debug's state restore and load/PV model rebuild are extracted into helpers so both modes share them.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…plan

The candidate charge/export windows start at the current slot, and fetch rebuilds them every cycle. The
replay never did, so every re-plan optimised the yaml's original windows and the plan froze. Extract the
--redo rate rescan into rescan_rate_windows() and call it before each replayed re-plan.

On a 3 Oct Sigenergy log replayed from the 07:00 yaml the first export window's start now matches the log
on 12 of 15 re-plans (9 before).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…xport targets

--replay_chart PNG draws the log's actual SoC with the export target of the live plan and the replayed plan
on the same % scale, plus a lane per plan showing whether it says export (or freeze) at each run. Uses the
Agg backend, so it never opens a window.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two modes now:
- exact (default): every re-plan starts from the SoC the log recorded, to check the replay reproduces the
  live plans;
- simulated (--replay_simulate): the battery is stepped forward under the replayed plan with the actual PV
  and load from the log, using Predbat's own battery and inverter model (rate curves, losses, reserve,
  inverter and export limits). This is the mode for trying a code change, since the SoC then follows the
  changed plan.

Plan windows are now parsed with their date, as minutes from the run's midnight. Previously tomorrow's
06:00 window read as today's, so the chart showed exports that were never planned for today.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
calculate_plan recomputes the load divergence from the load history, and the replay can only rebuild that
history at the log's 5-minute resolution, so the divergence came out smoother than the live system's. The
replay now overrides get_load_divergence with the value logged for each run (restored afterwards). The cloud
factor is no longer injected: calculate_plan derives it from the forecasts, which the replay already has.

Issue 4865 from 05:30: 22 of 51 re-plans now match the logged first export window start (20 before). Rik's
3 Oct log from 04:00: 8 of 34 plans identical (5 before).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y SoC RMS error

A dummy_inverter apps.yaml block enables a component that publishes inverter sensors and controls and
wires Predbat to them, so Predbat plans and controls it like a real inverter - for demos, and as an
independent inverter model for log replay. Behind the entities is a minute-by-minute hybrid model:
rate limits and losses, PV and battery sharing the inverter's AC limit (PV first), export capped at
the export limit, clipped PV counted, charge/export windows obeyed (zero rate = freeze), Eco otherwise.
A DUMMY row in INVERTER_DEF carries its capabilities.

The replay now also prints the RMS difference between simulated and logged SoC: 3.18% on the #4865
day, 0.50% on Rik's 3 Oct log.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…uilds each run

Fidelity fixes found by stepping through the first divergences of the #4865 replay:
- compare the adopted plan with the log's adopted plan (the next run's first "Best export window"), and
  the candidate plan with the log's candidate ("Export windows filtered") - before, the log's candidate
  was compared with the replay's adopted plan, which differ whenever a re-plan is rejected;
- rebuild the weighted-bucket historical load forecast before each re-plan, as fetch does;
- age the car charging and iBoost histories with the load history, since the load filter subtracts
  them - unshifted, overnight car charging was misaligned and the load forecast drifted upwards;
- take the cost so far today and the day counters from the log, as every metric starts from them;
- set the inverter's export limits alongside its export window, and publish while re-planning so the
  replay logs the same "predict base/best" lines as the live system for per-run comparison.

Rik's 3 Oct log from 04:00: 20 of 34 adopted plans now identical to the log (8 before).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A faithful replay needs the code that wrote the log. The log records the running version every run, so
the replay now stops at the first run whose version differs and says so - the #4865 log, for one, spans
an upgrade from v9.3.1 to v9.3.3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rows from the first run on a new version are flagged, the summary counts them separately, and the chart
marks the change. A replay past it is still useful but not expected to match as closely.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Predbat versions with feat/log-replay-inputs write "Replay input:" lines for the PV forecast (on each
change) and the load forecast (each cycle). The replay now applies them: a logged PV forecast replaces
the yaml's from its half hour onwards, and a logged load forecast replaces the rebuilt one before the
re-plan. Logs without the lines replay as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… replays

Applied after the debug yaml is restored, repeatable, with values read as numbers or booleans where they
look like them; an unknown setting name is rejected. First use: replaying the #4865 morning in simulated
mode with pv_metric90_weight at 0.15, 0.25 and 0.5.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e debug replays too

What-if overrides are useful for a single debug yaml as much as for a forward replay, so the option is now
--override name=value and run_single_debug applies it after restoring the yaml. The shared helper
apply_overrides lives beside restore_debug_state.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…put lines, the 8-hourly config refresh re-plan

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…w on every replayed re-plan

Fetch recomputes them each run over the horizon from now, so a passed peak drops out. The replay kept the
yaml's values, which after midnight would still include yesterday's rates.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the one logged after executing it

A run logs the inverter's SoC before planning and again after executing the plan; the parser kept the
last, so every replayed re-plan started a few minutes' battery movement away from the live one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
At midnight the replay moves everything held as minutes from midnight (rates, PV and load forecasts, keeps,
plan and inverter windows, manual times) back a day and advances midnight_utc; plan_last_updated_minutes
is left alone so calculate_plan forces the start-of-day re-plan as the live system does. Day counters that
restart at midnight count from zero, the midnight run's tariff comparison plans are ignored, and an export
window over midnight seen in the small hours started yesterday. --replay_until HH:MM means the first such
time after the yaml.

Replaying the 23:00 yaml of 3 Oct through to 09:50 on 4 Oct: 28 of 65 adopted plans identical to the log,
49 with the same first export start; simulated SoC within 2.6% RMS of the log over 131 runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When the 8-hourly config refresh is pending the run's adopted plan is marked invalid, and the next run
(at the same minutes_now) re-plans and adopts without comparing to the old plan. The replay now does the
same, and no longer reads that run's working window list as the plan in force. Without it the replay kept
the old plan under metric_min_improvement_plan and stayed on it for hours.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…es; mark each replayed run in the replay's log

fetch_sensor_data resets pv_forecast_minute90_signatures every cycle, so the plan's stale-p90 guard only
compares within one cycle. The replay never cleared them, so a logged PV update that moved p50 but left p90's
values alone read as a p90 left behind: the plan swapped p90 for p50, dropped to the proportional cloud
model and ran 5-7p off the live metric from then on. On the 3-4 Oct log this takes adopted plans identical
to the log from 29 to 55 of 71.

The replay's own predbat.log now carries a line per replayed run, so it can be read beside the live log.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Predbat now logs "Replay input: state" every cycle (SoC, soc_max, in-day adjustment, cost so far and the day
counters, each as a repr that reads back as the same float) and "Replay input: rates changed" whenever the
import/export rates or their base curves change. The replay now:
- takes the plan's starting values from the state line over the rounded human-readable lines, and uses its
  exact counters for the history shifts and the simulated battery's PV and load;
- rebuilds the four rate series from the logged change points, so rates fetched after the yaml (the next
  day's Agile prices, a new futurerate prediction) reach the re-plans. Rates are logged only on change, so a
  line from a run the parser drops carries to the next run, moved back a day if that run is past midnight.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"Will recompute the plan as it is invalid" is logged for every recompute, whatever triggered it, so the replay
treated each sensor-triggered re-plan as invalid and adopted the new plan without the old-vs-new comparison
the live run made. Only calculate_plan's "Recompute, previous plan is invalid..." marks a really invalid plan;
the first line now only stops a recomputing run's fresh window list being read as the plan in force.

Replay of the 11:30 yaml to 16:50 on 4 Oct: adopted plans 26 of 28 identical (22 before), candidates 31 of 31.
Also corrects the journal's 8-hourly refresh bullet, which still said the replay could not follow that re-plan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Newer logs carry, for every step the plan builds, the two cumulative load forecast values step_data_history()
reads (the step's start and the minute after), exactly. The replay sets those and leaves other minutes alone;
older logs' per-slot line is still read when the exact one is absent.

Checked by injecting the line, built from the live 14:50 debug yaml, into the 4 Oct log: the replay's
load_minutes_step (and 10/90) then match live's at 14:50 exactly, as do predict_metric_best and predict_soc_best.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Newer logs give p50/p10/p90 per minute from now to the end of the plan, exactly, as runs of [kWh, minutes];
the replay sets those minutes (removing ones logged None) and leaves the rest. The older half-hourly line is
still read when the exact one is absent.

Checked by injecting the exact load and PV lines, built from the live 14:50 debug yaml, into the 4 Oct log:
the replay's 14:50 re-plan then matches live's line for line, the previous plan's re-score included, and its
load/PV step series, predict_metric_best and predict_soc_best equal live's exactly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…only on re-plans

The replay applied the logged load forecast only on runs that re-planned, so between re-plans the instance
held the previous re-plan's forecast while live held a fresh one. No re-plan changes, but a state dump at a
non-re-plan run (the hourly debug yamls fall on :15, re-plans on :x0) now compares like for like.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"Replay input: cars changed" carries the run-time car fields the plan reads, logged only on change; the
replay sets them each run it applies, and carries a line from a dropped run on to the next kept run, as it
does the rates.

Checked on the 5 Oct log by injecting the car state from the live 18:15 debug yaml at 17:25, when the car
was plugged in: adopted plans 47 of 48 identical (43 of 46 without it), and the 18:15 re-plan, where the
injected state is the real one, matches live line for line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the exact load divergence

Where a log carries "Replay input: inverter changed", the replay sets the inverter's programmed state from it and
keeps it between changes, instead of rebuilding the export window from the previous run's lines. The rates in
force come from the state line, and the exact divergence wins over the rounded percentage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mperature from the state line

The compact line counts each value in whole tenths of a Wh and divides by 10000 only at the end, which gives
exactly the dp4 float the live forecast held: the 268 forecasts in an overnight log (195,640 values) all rebuild
bit for bit. Logs with the full cumulative kWh line still replay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rt and car - for live and replay

The chart's bottom strip showed only export and freeze. It now gives each plan's state at every run in the web
plan's terms and colours (Chrg, HoldChrg, FrzChrg, Exp, HoldExp, FrzExp, car), from both the export and the
charge windows in force, plus a lane for the car's planned charging. The live charge windows are the "Best
charge window" line at the start of the next run, the replay's its own adopted list. plan_state_now() replaces
export_mode_now().

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… dynamic_load() does live

Live dynamic_load() rebuilds car_charging_now_slots every cycle, so the plan holds the battery for a car that
reports charging with no planned slot covering it. The replay never did, so at 00:05 on 6 Oct it planned a
battery charge 00:00-00:30 that the live plan did not have. Export windows still matched, which hid it, but a
simulated replay charged the battery there and drifted: 8.8% SoC RMS overnight, 1.99% with this. The replay now
derives the slots with the live dynamic_load_car_charging_now() from the logged car state.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…plan timeline chart

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…logs it when divergence is off

The logging now writes the divergence exactly as the plan uses it, which is None when metric_load_divergence_enable
is off. The replay applied the rounded percentage line instead, giving the plan a divergence live never used.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
chalfontchubby added a commit that referenced this pull request Oct 7, 2026
… a log can be replayed exactly

A debug yaml captures one moment; a forward replay of the log after it (draft PR #5363) needs the values each
later plan started from, which the human-readable lines round or omit. Each line is prefixed "Replay input:",
reads back exactly (repr floats, or dicts that ast.literal_eval reads), and is logged on change where the value
rarely changes. None of it changes what Predbat does: every logger catches its own errors, and a value that
cannot be logged warns once per change. Tariff comparison runs (save=False / publish=False) log none of it, as
their inputs are not the live plan's.

- PV forecast, on change: p50/p10/p90 per minute from now to the end of the forecast, as runs of [kWh, minutes].
- Load forecast, every cycle: the two values per 5-minute step the plan reads (step start and the minute after).
  Rounded to 0.1 Wh, so logged compactly as Wh per step from a cumulative base, about half the size of the full
  cumulative kWh line it falls back to for a forecast that is not in whole tenths of a Wh.
- Load ML predictions, on change, exactly as read from the sensor.
- Rates, on change: import/export and their base rates from now, as change points.
- Plan starting state every cycle: SoC, SoC max, in-day adjustment, cost so far, the day counters, the charge
  and discharge rates in force and the battery temperature.
- Car state and the inverter's programmed state (windows, charging/exporting and targets, reserve, rate limits),
  each on change; and the load divergence exactly as the plan uses it (None when divergence is off).

The on-change lines share one helper (log_replay_changed). On a live system the replay lines are about 5% of
the log.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chalfontchubby
chalfontchubby force-pushed the tools/debug-replay-forward branch from 92a60d7 to 07a1416 Compare October 7, 2026 12:44
@chalfontchubby
chalfontchubby changed the base branch from main to feat/log-replay-inputs October 7, 2026 12:44
chalfontchubby and others added 2 commits October 7, 2026 14:15
…oximately, back to v8.8

A log from before the replay-input logging could not be replayed: the run's SoC, day totals and re-plan marker
had all changed format over the versions, so most older logs gave no runs at all. The replay now reads:

- the SoC from "Inverter 0 SOC: 1.9kW ..." / "SoC: 6.85kW ..." as well as today's line, and the all-inverter
  total from "Found N inverters totals" where a log has it (so a multi-inverter system starts from the total);
- the older spaced day totals ("load 5.04 kWh import 14.07 kWh ... pv 0.0 kWh");
- a re-plan from "Filtered charge windows", for versions that do not log the filtered export windows;
- where a run has no exact cars line, the car's planned slots from "Car N charging plan is" (since v5.1) and its
  planned and charging-now flags (since v7.0);
- where the log has no rates lines, the Intelligent dispatch list from "Octopus slots changed" (since v8.27.27),
  with the import rates rebuilt from it by the live rate_add_io_slots.

A log without any replay-input lines gets a note that the plans are approximate, and one with no readable runs a
clear error. On the 5-6 Oct overnight log with every replay-input line stripped, adopted plans identical rose from
5 to 20 of 66 (same first export start 14 to 66) and the simulated SoC error fell from 11.2% to 2.0%; real issue logs
from v8.8.13 to v9.0.3 now parse, and a 20-hour v9.0.3 log replays. The full-input replay is unchanged (66 of 66).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… line

A multi-inverter log logs each inverter's own SoC line, and newer versions dropped the "Found N inverters totals"
line, so the replay planned from inverter 0's SoC alone: on the 3-inverter system in #5423 that was 6.46 kWh of a
real 18.54 kWh at 18:45, and no re-plan matched. The run's SoC is now the sum of each inverter's first SoC line of
the run (matching live's own total, 26.86 vs 26.865 kWh at 17:30), and each inverter is set to its own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant