The training-label fix in #2 corrects a reproduced error: nested extension lists could lose the first accepted identifier. Singleton/edgeless graph handling was corrected at the same time. Existing checkpoints and the files under results/ predate that fix; the current CPU smoke tests verify execution and labels but do not establish the validity of historical benchmark metrics.
Regenerate the experiment artifacts using the corrected loader, retaining the old artifacts as explicitly historical until replacement is reproducible.
Acceptance criteria:
- Identify the original train/validation/test graph lists, solution semantics and source corpus version.
- Record the dataset checksums, code commit, random seeds, dependencies and full commands for each run.
- Retrain the refined model and the claimed baselines on those documented splits.
- Export comparable metrics, per-run logs and checkpoint metadata, and link each result to its producing run.
- Update result summaries only after the new artifacts have been independently reproduced.
The original split manifests are not present in this checkout, so a synthetic smoke example cannot substitute for this validation. No new benchmark claim is made by #2.
The training-label fix in #2 corrects a reproduced error: nested extension lists could lose the first accepted identifier. Singleton/edgeless graph handling was corrected at the same time. Existing checkpoints and the files under
results/predate that fix; the current CPU smoke tests verify execution and labels but do not establish the validity of historical benchmark metrics.Regenerate the experiment artifacts using the corrected loader, retaining the old artifacts as explicitly historical until replacement is reproducible.
Acceptance criteria:
The original split manifests are not present in this checkout, so a synthetic smoke example cannot substitute for this validation. No new benchmark claim is made by #2.