Summary
On the latest main (3409b82), the Stage 2 action-aware GRPO pipeline cannot run end to end. The rollout scripts, reward and GRPO loss are present. However, three pieces of code they depend on appear to have been removed by the sync commit 42bcb0c. All three still exist in the initial commit 3028419 under posttrain/InfinityStar-main/.
Environment: clean checkout of 3409b82, Python 3.10, torch 2.5.1, transformers from requirements.txt, 1 GPU.
1. GRPO training arguments are not defined in Worldmodel/runtime/infinity/utils/arg_util.py
I ran the official launcher:
bash action_aware_grpo/scripts/run_stageb_partialfreeze.sh PARTIAL_FREEZE_MODE=smoke RUN_ID=repro
It parsed the arguments without any GRPO fields; the printed args dump contains no grpo_* keys. I then parsed the exact argv that GRPO_stageB_train.sh passes to train.py, using the repo's own parser (Args(explicit_bool=True).parse_args(argv, known_only=True), the same call as arg_util.py:318):
extra_args (unrecognized flags): 25
['--trainer_type', '--shuffle_batches', '--hybrid_step_on_role', '--grpo_hybrid_rl_coef', '--max_train_iters', '--extra_epochs_after_resume', '--grpo_adv_clip', '--grpo_ratio_eps', '--grpo_kl_beta', '--grpo_lambda_act', '--grpo_lambda_task', '--grpo_lambda_ce', '--grpo_alpha_decay', '--grpo_require_old_logprob', '--grpo_new_logprob_mode', '--grpo_require_nonnegative_adv', '--grpo_log_weight_stats', '--grpo_replay_save_on_cpu', '--grpo_replay_save_on_cpu_pin', '--grpo_replay_checkpointing', '--grpo_pg_only', '--grpo_weight_mode', '--grpo_aux_sft_coef', '--freeze_chunk_prefix', '--partial_freeze_print_summary']
getattr(args,'trainer_type','sft') -> sft
accessing args.freeze_chunk_prefix (train.py:183) ...
AttributeError: 'Args' object has no attribute 'freeze_chunk_prefix'
Consequences:
Worldmodel/runtime/train.py:183 crashes with this AttributeError.
- Even if that line is bypassed,
trainer_type falls back to sft, so Stage B would silently train as plain SFT instead of GRPO.
2. gen_one_example has no return_trace parameter
action_aware_grpo/grpo_server.py:1000 and :1028 call gen_one_example(..., return_trace=True). gen_one_example is imported from tools.run_infinity (grpo_server.py:130-146), and its definition at Worldmodel/runtime/tools/run_infinity.py:94 has no such parameter.
Calling the real function with the same keyword arguments as grpo_server.py:
TypeError: gen_one_example() got an unexpected keyword argument 'return_trace'
So every Stage A rollout fails on its first segment. In 3028419, posttrain/InfinityStar-main/tools/run_infinity.py:124 had return_trace.
3. Trace output and replay log-prob removed from Infinity
Infinity has ar_infer_infinity_elegant_replay_logprob: False
ar_infer_infinity_elegant accepts return_trace: False
Worldmodel/runtime/infinity/trainer/sft_trainer.py:987 calls ar_infer_infinity_elegant_replay_logprob, which is used in trace_replay mode.
- Both this method and the
return_trace / per-clip log-prob output of ar_infer_infinity_elegant exist in 3028419:posttrain/InfinityStar-main/infinity/models/infinity.py (around lines 562–782).
Also: merge-conflict markers in the backbone training launcher
$ bash -n train/scripts/train_from_base.sh
train/scripts/train_from_base.sh: line 50: syntax error near unexpected token `<<<'
train/scripts/train_from_base.sh: line 50: `<<<<<<< Updated upstream'
Lines 50–53 still contain <<<<<<< Updated upstream / ======= / >>>>>>> Stashed changes.
Request
Could you restore the GRPO argument definitions and the trace / replay-log-prob code in Worldmodel/runtime/, or point to the commit that contains the working Stage 2 code? Thanks!
Summary
On the latest
main(3409b82), the Stage 2 action-aware GRPO pipeline cannot run end to end. The rollout scripts, reward and GRPO loss are present. However, three pieces of code they depend on appear to have been removed by the sync commit42bcb0c. All three still exist in the initial commit3028419underposttrain/InfinityStar-main/.Environment: clean checkout of
3409b82, Python 3.10, torch 2.5.1, transformers fromrequirements.txt, 1 GPU.1. GRPO training arguments are not defined in
Worldmodel/runtime/infinity/utils/arg_util.pyI ran the official launcher:
bash action_aware_grpo/scripts/run_stageb_partialfreeze.sh PARTIAL_FREEZE_MODE=smoke RUN_ID=reproIt parsed the arguments without any GRPO fields; the printed args dump contains no
grpo_*keys. I then parsed the exact argv thatGRPO_stageB_train.shpasses totrain.py, using the repo's own parser (Args(explicit_bool=True).parse_args(argv, known_only=True), the same call asarg_util.py:318):Consequences:
Worldmodel/runtime/train.py:183crashes with thisAttributeError.trainer_typefalls back tosft, so Stage B would silently train as plain SFT instead of GRPO.2.
gen_one_examplehas noreturn_traceparameteraction_aware_grpo/grpo_server.py:1000and:1028callgen_one_example(..., return_trace=True).gen_one_exampleis imported fromtools.run_infinity(grpo_server.py:130-146), and its definition atWorldmodel/runtime/tools/run_infinity.py:94has no such parameter.Calling the real function with the same keyword arguments as
grpo_server.py:So every Stage A rollout fails on its first segment. In
3028419,posttrain/InfinityStar-main/tools/run_infinity.py:124hadreturn_trace.3. Trace output and replay log-prob removed from
InfinityWorldmodel/runtime/infinity/trainer/sft_trainer.py:987callsar_infer_infinity_elegant_replay_logprob, which is used intrace_replaymode.return_trace/ per-clip log-prob output ofar_infer_infinity_elegantexist in3028419:posttrain/InfinityStar-main/infinity/models/infinity.py(around lines 562–782).Also: merge-conflict markers in the backbone training launcher
Lines 50–53 still contain
<<<<<<< Updated upstream/=======/>>>>>>> Stashed changes.Request
Could you restore the GRPO argument definitions and the trace / replay-log-prob code in
Worldmodel/runtime/, or point to the commit that contains the working Stage 2 code? Thanks!