Repository navigation
Archiver 2GB byte-array ceiling fix - #145
Merged
Merged
Conversation
…builder so that
`-Dhoodie.*=...` in spark.driver.extraJavaOptions reaches the archiver
config. No effect when no such properties are set.
✅ Snyk checks have passed. No issues have been found so far.
💻 Catch issues earlier using the plugins for VS Code, JetBrains IDEs, Visual Studio, and Eclipse. |
schung507
approved these changes
Sep 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
HoodieTimelineArchiver OOMs at HoodieAvroDataBlock.serializeRecords while wrapping clean instants into an archive log block. Root cause: HoodieCleanerPlan.filePathsToBeDeletedPerPartition and HoodieCleanMetadata.partitionMetadata each carry per-partition maps that, at 6.7M partitions, produce serialized instants whose individual Avro block exceeds Integer.MAX_VALUE bytes. Lowering the archival batch size doesn't help — a single clean instant is already >2GB.
Changes
Two opt-in features + one CLI plumbing fix, all default-off so behaviour is unchanged unless explicitly enabled.
hoodie.archive.trim.clean.action.metadata(HoodieArchivalConfig) — new boolean, default false. When true, strips per-partition file lists from HoodieCleanerPlan. Retention boundary, policy, timestamps, summary counts, and extraMetadata are preserved.MetadataConversionUtils.createMetaWrapper— new 3-arg overload accepting trimCleanActionMetadata. Old 2-arg signature kept as a delegating shim. HoodieTimelineArchiver.convertToAvroRecord reads the config and passes it through. In-place mutation on the freshly-deserialized objects avoids allocating a rebuilt copy of the very maps we're shrinking.ArchiveExecutorUtils.archive— added .withProperties(System.getProperties()) to the HoodieWriteConfig builder. Previously the ARCHIVE CLI's config was built from only the five positional args and silently ignored any -Dhoodie.*=… passed via spark.driver.extraJavaOptions. Now fork-side configs (this trim flag, existing options like hoodie.commits.archival.batch.size) can be tuned per-invocation without a code change.Verified
Manual spark-submit of the ARCHIVE command against events-v1 with -Dhoodie.archive.trim.clean.action.metadata=true completes in ~57 min through the previously-failing Feb 2023 clean instants; nine archive-file rollovers, ~1GB each, well under the ceiling. Prior runs OOM'd at the same fourth clean instant every time.
Compatibility