Skip to content

Regenerate historical experiment artifacts after the training-label correction #3

Description

@lmlearning

The training-label fix in #2 corrects a reproduced error: nested extension lists could lose the first accepted identifier. Singleton/edgeless graph handling was corrected at the same time. Existing checkpoints and the files under results/ predate that fix; the current CPU smoke tests verify execution and labels but do not establish the validity of historical benchmark metrics.

Regenerate the experiment artifacts using the corrected loader, retaining the old artifacts as explicitly historical until replacement is reproducible.

Acceptance criteria:

  • Identify the original train/validation/test graph lists, solution semantics and source corpus version.
  • Record the dataset checksums, code commit, random seeds, dependencies and full commands for each run.
  • Retrain the refined model and the claimed baselines on those documented splits.
  • Export comparable metrics, per-run logs and checkpoint metadata, and link each result to its producing run.
  • Update result summaries only after the new artifacts have been independently reproduced.

The original split manifests are not present in this checkout, so a synthetic smoke example cannot substitute for this validation. No new benchmark claim is made by #2.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions