[Feature] Add opt-in KV-cache continuation for DualVLN evaluation - #368
Open
maxfanxd wants to merge 2 commits into
Open
[Feature] Add opt-in KV-cache continuation for DualVLN evaluation#368maxfanxd wants to merge 2 commits into
maxfanxd wants to merge 2 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
This adds
--kv-cache-continuationto the existing Habitat DualVLN evaluator. It reuses the generation KV cache for latent readout, avoiding repeated image encoding and prefix computation.The flag is off by default. Without it, evaluation uses the original
generate_latentspath. Existing configs, dependencies and the root README are unchanged.Usage
Validation
23 CPU regression tests pass, covering default behavior, numerical agreement, cache boundaries and cleanup.
A small Habitat R2R test with the complete model on RTX A6000 (BF16):
This is 17.30x faster latent readout, not 17.30x faster navigation. All 433 simulator actions matched in the two runs; mean latent query cosine was 0.999700.
Impact on success rate: Both modes recorded SR/SPL of 0 on these two episodes, which hit the configured 40-step cap without success. The matching action sequences provide a small integration check, not evidence of unchanged full-dataset success rate. A full SR/SPL comparison remains to be done.
BF16 outputs are not bitwise equal and could affect other trajectories, so the optimization remains opt-in. Robot deployment has not been tested.
Tests and measurement details