Skip to content

name fully-async checkpoints after the last trained step - #2446

Open
coreyhu wants to merge 4 commits into
NovaSky-AI:mainfrom
coreyhu:fully-async-checkpoint-step
Open

coreyhu wants to merge 4 commits into
NovaSky-AI:mainfrom
coreyhu:fully-async-checkpoint-step

Conversation

@coreyhu

@coreyhu coreyhu commented Oct 8, 2026 •

Copy link
Copy Markdown

Summary

The fully-async trainer advanced global_step before its final saves, and before the save it makes when an epoch runs out of groups early (sample_full_batch), so those checkpoints and HF exports were named one step past the last trained step. They are now saved under the last trained step, as in RayPPOTrainer, and a step that already has a checkpoint is not saved again. Around that:

  • If prompts are filtered after the last trained step's checkpoint and the epoch then ends early, only that checkpoint's fully_async_state.pt is rewritten, so resume skips them. Model weights are not saved again, and a checkpoint loaded on resume is never modified.
  • Nothing is saved before the first trained step (no global_step_0).
  • A resumed run with nothing left to train (step budget spent, or every epoch done) returns before syncing weights, generating or saving, like preserve checkpoint publication and dataloader resume state #2362's max_training_steps early return.

Why

On main, two trained steps produce global_step_1, global_step_2 and an extra global_step_3, and max_training_steps=3 ends at global_step_4, so a resumed run starts one step late.

Testing

  • pytest tests/train/test_fully_async_trainer.py: 23 passed; 9 of the 10 new cases fail on main (the ninth checks global_step after a failed save).
  • CPU suite (tests/train/ tests/backends/skyrl_train/, -m "not vllm"): 2232 passed, 41 skipped (2222 on main); the same 2 Megatron import failures occur on main. pre-commit clean.

This overlaps #2362 (the early return and the fully-async save_checkpoints), so whichever lands second needs a small merge in fully_async_trainer.py.

AI assistance: Claude Code ported this from our patch set, wrote the tests, and drafted this description.


Note

Medium Risk
Changes checkpoint/resume semantics and when fully-async state is rewritten; mistakes could skew resume step counts or skip/repeat training, though behavior is heavily covered by new unit tests.

Overview
Aligns fully-async checkpoint and HF export naming with RayPPOTrainer: saves use the last trained step instead of the loop’s “next” step, skips duplicate saves for a step already checkpointed, and avoids global_step_0 when nothing has trained yet.

The train() loop tracks last_ckpt_step / last_hf_step (and last_ckpt_dir for in-run checkpoints). On early epoch end under sample_full_batch, it decrements global_step for the save, then restores it; if prompts were filtered after the last checkpoint, it may rewrite only fully_async_state.pt on a checkpoint this run wrote—never the weights or a checkpoint loaded on resume. Resume can exit immediately when _nothing_left_to_train() (new is_epoch_complete() on the async dataloader) says the step budget or epochs are done.

Checkpoint helpers are split into _fully_async_state() and _write_fully_async_state(). Tests add stubbed train() cases for naming, early epoch end, filtered-UID state updates, and no-op resume.

Reviewed by Cursor Bugbot for commit 8f700eb. Bugbot is set up for automated code reviews on this repo. Configure here.

The fully-async loop advances global_step before its final saves and before the save on an early epoch end, so those checkpoints and HF models were named one step past the last trained step. Save them under the last trained step, as the synchronous trainer does, and skip a save when that step already has one.

Signed-off-by: Corey Hu <corey.hu@scale.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request prevents duplicate checkpoint and model saves by tracking the last saved steps and adjusting 'self.global_step' to ensure saves are named after the last trained step. It also adds corresponding unit tests. The reviewer suggested wrapping the save operations in a try-finally block during early epoch termination to ensure 'self.global_step' is always restored even if an exception occurs.

Comment thread skyrl/train/fully_async_trainer.py Outdated
@greptile-apps

greptile-apps Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 3/5

[Medium risk] Adjusts checkpoint naming logic in async training loop.

Fix the lost resume state before merging; resumed runs can redraw discarded prompts or restart completed epochs.

Findings

  1. P1 Discarded prompts return on resume ▶
  2. P1 Empty training runs restart ▶
  3. P2 Completed resumes rewrite model exports ▶
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart TD
    A[Train step N] --> B[Save checkpoint N]
    B --> C[Collect more groups]
    C --> D[Discard remaining groups and end epoch]
    D --> E[Skip save because step N already exists]
    E --> F[Resume checkpoint N]
    F --> G[Draw discarded prompts again]
Loading

Reviews (1) · Last reviewed commit: "Name fully-async checkpoints after the l..." · Reviewed by Greptile

Comment thread skyrl/train/fully_async_trainer.py Outdated
Comment thread skyrl/train/fully_async_trainer.py
Comment thread skyrl/train/fully_async_trainer.py Outdated

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

Reviewed by Cursor Bugbot for commit 4ba4dda. Configure here.

Comment thread skyrl/train/fully_async_trainer.py Outdated
The early epoch-end path lowers global_step to name the save after the last trained step. Put the restore in a finally block so a failed save leaves global_step naming the next step, as any other failure in the loop does.

Signed-off-by: Corey Hu <corey.hu@scale.com>
Three cases left a resumed run with the wrong state after checkpoints
were named after the last trained step:

- Prompts filtered after the last trained step's checkpoint, when the
  epoch then ends early, were not recorded. Rewrite only that
  checkpoint's fully_async_state.pt so they are, without saving the
  model again.
- An epoch that ended before any step trained saved global_step_0,
  which resume ignores. Nothing is saved before the first trained step.
- Resuming a run with nothing left to train (step budget spent or every
  epoch done, including an epoch that ran out of prompts early) re-saved
  its last HF export. It now returns before syncing weights, generating
  or saving, like the max_training_steps early return in NovaSky-AI#2362.

Signed-off-by: Corey Hu <corey.hu@scale.com>
A checkpoint loaded on resume can belong to another run or be a shared base (resume_mode=from_path), so it is never modified. If an epoch ends early after resuming and before any new save, the prompts filtered since are redrawn on the next resume instead.

Signed-off-by: Corey Hu <corey.hu@scale.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant