A small Nextflow component that reads an Illumina samplesheet (bcl2fastq or
Local Run Manager format) and emits a generic nf-core samplesheet
(sample,fastq_1,fastq_2) that can be fed straight into an nf-core module.
Pure Groovy — no Python, no extra runtime dependencies. Nextflow already
ships Groovy, and the parser is in lib/SamplesheetParser.groovy.
Given:
- an Illumina
SampleSheet.csvin any of three auto-detected formats:- bcl2fastq (V1, IEM) — has a
[Data]section (NovaSeq 6000 and earlier Illumina platforms) - BCLConvert (V2) — has a
[BCLConvert_Data]or[Cloud_Data]section (NovaSeq X series and newer Illumina platforms). V2-specific config sections like[BCLConvert_Settings]andOverrideCyclesare recognised for auto-detection but not used by the reshape use case — the per-sample data rows have the same structure as V1. - Local Run Manager (LRM) — headerless variant, first line is the header
- bcl2fastq (V1, IEM) — has a
- a directory of fastq files produced by
bcl2fastq(or any tool that names files as{SampleID}_S{N}_L00{Lane}_R{1,2}_001.fastq.gz),
it produces:
- a CSV with header
sample,fastq_1,fastq_2where each row is one sample. Multiple R1 (or R2) files for a sample — e.g. across lanes — are joined with a comma in the same cell, per nf-core convention. Single-end samples get an emptyfastq_2cell.
.
├── lib/ # Nextflow auto-loads all .groovy here
│ ├── SamplesheetReshape.groovy # public API facade
│ ├── SamplesheetParser.groovy # bcl2fastq / LRM parsing
│ ├── SamplesheetReshaper.groovy # file matching + nf-core CSV output
│ ├── SamplesheetValidator.groovy # bcl2fastq structural checks
│ └── CsvSupport.groovy # shared utilities (asFile)
├── subworkflows/
│ └── nf-core/
│ └── reshape_samplesheet/ # the nf-core sub-workflow (publish layout)
│ ├── main.nf
│ ├── meta.yml
│ └── tests/main.nf.test
├── modules/
│ └── nf-core/
│ ├── samplesheet_validate/ # pre-flight check (≥2 modules rule)
│ │ ├── main.nf
│ │ └── meta.yml
│ └── samplesheet_reshape/ # the actual reshape + CSV write
│ ├── main.nf
│ └── meta.yml
├── tests/
│ ├── data/ # sample inputs
│ ├── fastqs/ # generated by bin/make_test_fastqs.sh
│ ├── test_runner.groovy # shared TestRunner (loaded via `evaluate`)
│ ├── test_parser.groovy # parseSamplesheet + malformed input
│ ├── test_reshaper.groovy # reshape + writeReshaped
│ ├── test_validator.groovy # validate() + validateBcl2fastq()
│ ├── test_options.groovy # recursive + strandedness opts
│ └── test_matching.groovy # default + pattern matching
├── bin/
│ ├── make_test_fastqs.sh
│ └── test.sh # test runner (full or smoke)
├── main.nf # demo entry point
├── nextflow.config
└── mise.toml # dev tool pins
--samplesheet, --fastq_dir, --outdir, and --recursive are
required; a bare nextflow run main.nf fails loudly with a clear
message. --strandedness is optional — omit it to keep the default
3-column output, or set it to add a fourth column.
The pipeline's main.nf calls validate() before doing anything
expensive, so a malformed samplesheet or a missing fastq dir fails
fast (before any compute is submitted) with a single error message
listing all the problems at once.
# 3-column output (generic nf-core format)
nextflow run main.nf \
--samplesheet tests/data/illumina_bcl2fastq.csv \
--fastq_dir tests/fastqs \
--outdir results \
--recursive false
# 4-column output with strandedness, useful for nf-core/rnaseq
nextflow run main.nf \
--samplesheet /path/to/SampleSheet.csv \
--fastq_dir /path/to/bcl2fastq/output/ \
--outdir /path/to/where/you/want/the/output \
--recursive false \
--strandedness unstranded # or forward, reverse, autoThe sub-workflow emits two views of the same data: a csv channel
(the canonical artifact, good for inspection and external tools)
and a samples channel (a pre-parsed channel of [meta, fastq_1, fastq_2] tuples — the "three lines to a real analysis module"
hand-off). Pick whichever fits your consumer.
include { RESHAPE_SAMPLESHEET } from './subworkflows/nf-core/reshape_samplesheet/main.nf'
workflow {
RESHAPE_SAMPLESHEET(
file(params.illumina_samplesheet),
file(params.fastq_dir),
file(params.outdir),
[recursive: false]
)
// Option A: use the CSV as a file artifact (e.g. for an
// external tool that consumes a samplesheet, or for human
// inspection).
RESHAPE_SAMPLESHEET.out.csv.view { csv -> println "Reshaped CSV: ${csv}" }
// Option B: use the pre-parsed `samples` emit to feed a
// downstream nf-core module — three lines, no inline parser.
RESHAPE_SAMPLESHEET.out.samples
.map { meta, fq1, fq2 -> [meta, fq1 + fq2] }
.set { ch_fastqc_input }
// FASTQC(ch_fastqc_input) // or any (meta, reads) consumer
}examples/poc/ is a self-contained proof-of-concept that takes
the next step: it reads the CSV RESHAPE_SAMPLESHEET emits and
feeds it straight into nf-core/fastqc
to produce per-sample QC reports — including the multi-lane
sample_A, which produces 4 reports (2 R1 lanes + 2 R2 lanes).
The shape that makes this useful is the new samples emit
on RESHAPE_SAMPLESHEET: a channel of [meta, fastq_1, fastq_2]
tuples, built by parsing the CSV through the lib's
ReshapedCsvParser. The POC's consumer code is three lines:
RESHAPE_SAMPLESHEET.out.samples
.map { meta, fq1, fq2 -> [meta, fq1 + fq2] }
.set { ch_fastqc_input }
FASTQC(ch_fastqc_input)No inline parser, no RFC-4180 handling, no file matching — the
sub-workflow already did that and exposes ready-to-use tuples.
The POC has its own nf-test, parser unit tests, and smoke
test, and includes the one-time
setup.sh that clones nf-core/modules at a pinned SHA.
# One-time setup (stages lib/ + clones nf-core/modules@<pinned SHA>)
examples/poc/bin/setup.sh
# Run the POC end-to-end (uses Docker to run fastqc)
nextflow run examples/poc/main.nf \
--samplesheet tests/data/illumina_bcl2fastq.csv \
--fastq_dir tests/fastqs \
--outdir results-poc \
--recursive false
# Or, run the test stack
bin/test.sh --poc-smoke # end-to-end: real fastqc against the test fixtures
bin/test.sh --poc-test # full stack: nf-test (stub) + the real fastqc smokeThe sub-workflow takes four arguments:
| input | type | meaning |
|---|---|---|
illumina_samplesheet |
path | the Illumina CSV |
fastq_dir |
path | directory containing the fastq files |
output_dir |
path | where to write the reshaped CSV |
opts |
Map | options. Recognised keys: recursive (Boolean), strandedness (String). See below. |
| key | type | default | effect |
|---|---|---|---|
recursive |
Boolean | false |
walk fastq_dir recursively to find fastqs in per-sample subdirectories |
strandedness |
String | null |
when non-null, a fourth strandedness column is added with this value for every sample. The lib coerces the following to null: Boolean true, empty string, the String "true", the String "false" — these are all easy CLI footguns (e.g. --strandedness '' is collapsed by Nextflow to the String "true") |
pattern |
String | null |
a regex template for matching fastq filenames to Sample_IDs. The template may reference the Sample_ID as ${sampleId} (regex-escaped). See the "How samples are matched" section. |
validateStructure |
Boolean | false |
when true, run bcl2fastq structural checks on the samplesheet (Sample_ID uniqueness, non-empty I7/I5 index values, I7/I5 index format, I7/I5 length consistency across samples, I7+I5 index uniqueness) in addition to the per-sample fastq checks. Useful for a combined pre-flight that validates everything in one pass. See the "Pre-bcl2fastq validation" section. |
hammingDistanceAsError |
Boolean | false |
only meaningful when validateStructure: true. Controls whether Hamming-distance index violations (two indices in the same column differing by fewer than 2 bases) are reported as errors (thrown) or warnings (printed to System.err, the pipeline continues). A Hamming distance of 1 is a soft risk that bcl2fastq may or may not handle depending on the configured mismatch tolerance, and some NovaSeq X series UMI-style demultiplexing workflows deliberately use close indices, so the warning default is deliberate. Set to true for production runs where every close-index pair is worth investigating. |
Unknown keys in opts are silently ignored. Pass null (or [:]) for all defaults.
…and emits a single value on the csv channel. (Nextflow auto-wraps
the emit value in a single-element channel, so downstream code should
treat it as a channel — e.g. .view, .first(), .subscribe — not as
a raw value.)
The library accepts any path-like type — java.io.File, String, or
java.nio.file.Path (the parent of Nextflow's wrapper path type) — so
the caller doesn't need to do any explicit conversion.
// Non-recursive, no strandedness column (default)
def csv = SamplesheetReshape.reshape(samplesheetFile, fastqDirFile)
// Recursive
def csv = SamplesheetReshape.reshape(samplesheetFile, fastqDirFile, [recursive: true])
// With strandedness column
def csv = SamplesheetReshape.reshape(samplesheetFile, fastqDirFile, [strandedness: 'unstranded'])
// Both options
def csv = SamplesheetReshape.reshape(samplesheetFile, fastqDirFile,
[recursive: true, strandedness: 'forward'])The Map opts form is forward-compatible: unknown keys are silently
ignored, so future options can be added without breaking callers.
include { RESHAPE_SAMPLESHEET } from './subworkflows/nf-core/reshape_samplesheet/main.nf'
workflow {
// Build the opts Map in a named variable — Nextflow's DSL parser
// can choke on Map literals in certain call positions.
def opts = new LinkedHashMap()
opts.recursive = false
// opts.strandedness = 'unstranded' // optional
RESHAPE_SAMPLESHEET(
file(params.illumina_samplesheet),
file(params.fastq_dir),
file(params.outdir),
opts
)
}Default (strict): each row's Sample_ID must be a prefix of the
fastq filename, and the character immediately after it must be a
separator (. or _) or the end of the string. So sample_A matches
sample_A_S1_L001_R1_001.fastq.gz and sample_A.fastq.gz but NOT
X_sample_A_… or sample_A1_…. This avoids the substring-matching
footgun where a short Sample_ID like A would otherwise false-match
ABC_S1_L001_R1_001.fastq.gz.
Custom regex (opts.pattern): you can override the default with a
regex template. The template may reference the Sample_ID as
${sampleId} (any number of times), and the substitution is
regex-escaped so a Sample_ID like v1.0 matches literally:
// bcl2fastq convention
def csv = SamplesheetReshape.reshape(samplesheet, fastqDir,
[pattern: '^${sampleId}_S\\d+_L\\d+_R[12]_\\d+\\.fastq\\.gz$'])
// any file starting with the Sample_ID (looser than the default)
def csv = SamplesheetReshape.reshape(samplesheet, fastqDir,
[pattern: '^${sampleId}'])
// embed the sample ID twice
def csv = SamplesheetReshape.reshape(samplesheet, fastqDir,
[pattern: '${sampleId}_v1_${sampleId}_R[12]\\.fastq\\.gz$'])${sampleId} is matched against the full filename (==~ semantics),
so the pattern is implicitly anchored.
R1 vs R2 is detected per filename with _R{1,2}[._] (also accepts
_1.fastq / _2.fastq shorthand). Files that don't match any sample
are silently ignored.
Missing samples / fastq_dir issues (strict, see "Input handling" below):
- if no fastq file matches a
Sample_ID, reshape throwsIllegalArgumentExceptionwith the complete list of unmatched sample IDs in one error message — never silently drop a sample - if
fastq_diris missing or is not a directory, reshape throwsIllegalArgumentExceptionimmediately - if
fastq_direxists but contains no.fastq/.fq(.gz)files, reshape throwsIllegalArgumentException
In all three cases the file path is included in the error message so you can find the offending input in a batch run.
The library is strict by default — it throws
IllegalArgumentException for malformed input rather than producing
silently-wrong output. The file path is always included in the error
message so you can find the offending file in a batch.
| Input | Behaviour |
|---|---|
| File missing | IllegalArgumentException |
| File empty / whitespace-only | IllegalArgumentException (empty file is almost certainly a mistake) |
bcl2fastq with [Header]/[Reads]/[Manifests] but no [Data] |
IllegalArgumentException |
| Empty header row | IllegalArgumentException |
| Empty header cell | IllegalArgumentException with the 1-based column number |
| Duplicate header column names | IllegalArgumentException with both positions |
Header missing Sample_ID column |
IllegalArgumentException listing the columns found |
| Malformed row (wrong cell count) | IllegalArgumentException with the 1-based line number — all bad rows are collected and reported in one error so the user sees the complete list, not just the first |
Sample with no Sample_ID value (empty cell) |
IllegalArgumentException with the 1-based line number |
Duplicate Sample_ID (bcl2fastq check) |
IllegalArgumentException listing all duplicates |
| Empty I7 or I5 value (bcl2fastq check) | IllegalArgumentException naming the sample with the empty index — bcl2fastq would have nothing to demultiplex by |
| I7/I5 index with invalid characters (bcl2fastq check) | IllegalArgumentException showing the bad index sequence; lowercase is normalised to uppercase |
| I7 or I5 length inconsistent across samples (bcl2fastq check) | IllegalArgumentException naming the offending sample — bcl2fastq requires consistent lengths in a run |
| Duplicate I7+I5 combination (bcl2fastq check) | IllegalArgumentException listing all colliding samples — demultiplexing-critical |
| I7 or I5 indices with Hamming distance < 2 (bcl2fastq check) | by default, WARNING: printed to System.err; with validateStructure: true, hammingDistanceAsError: true in the opts, IllegalArgumentException listing all collisions |
fastq_dir missing or not a directory |
IllegalArgumentException with the path |
fastq_dir contains no .fastq/.fq(.gz) files |
IllegalArgumentException with the path |
| Sample in the samplesheet has no matching fastq files | IllegalArgumentException listing all such samples in one error so the user sees the complete list, not just the first |
| Multi-line quoted CSV field | supported (the field's embedded newline is preserved) |
| UTF-8 BOM at start of file | stripped |
| CRLF / LF / CR line endings | all handled |
Non-ASCII content (e.g. 样品 as a sample name) |
read as UTF-8 explicitly (no platform-default surprises) |
Recommended: call validate(samplesheet, fastqDir) at the entry
point of any pipeline that uses reshape. The strict behaviour throws
on the first bad sample, so without a pre-flight check, partial
output (e.g. "20 of 21 samples processed") only surfaces after the
run has billed compute.
// 1-arg: just parse the samplesheet (catches malformed file/header/rows)
SamplesheetReshape.validate(samplesheet)
// 2-arg: full pre-flight (samplesheet + fastq matching)
SamplesheetReshape.validate(samplesheet, fastqDir)
// 3-arg: with options (same opts as reshape)
SamplesheetReshape.validate(samplesheet, fastqDir, [recursive: true])The 2- and 3-arg versions run exactly the same checks as reshape —
including the per-sample fastq matching — but discard the output.
If they return without throwing, the actual reshape (or
writeReshaped) call that follows is guaranteed to succeed.
Example error from a real call:
Cannot reshape samplesheet /data/Sheet.csv: 2 sample(s) have no matching
fastq files in /data/fastqs:
sample_X
sample_Y
The user sees both missing samples in one error and can fix them before the pipeline runs.
If you want to validate a samplesheet before running bcl2fastq —
when you don't have any FASTQ files yet — call
validateBcl2fastq(samplesheet). It runs the parse-level checks
plus the bcl2fastq-specific structural checks:
- Sample_ID uniqueness — bcl2fastq hard-fails on duplicates
- I7 (
index) sequence format — onlyA/C/G/T/Nallowed; lowercase is normalised to uppercase soatcgacgtis accepted - I5 (
index2) sequence format — same rules - I7+I5 combination uniqueness — two samples on the same lane with the same I7 AND the same I5 cannot be demultiplexed; the error lists every colliding sample so you can fix them in one pass
- I7 and I5 Hamming distance ≥ 2 — every pair of indices in
the same column (I7 and I5 separately) must differ by at least
2 bases. Catches demultiplexing risks where one sequencing
error could cross-assign reads to the wrong sample. Default
behaviour is a warning (printed to
System.err, the pipeline continues) — see thehammingDistanceAsErroropt below.
// Just the samplesheet, no FASTQ files needed
SamplesheetReshape.validateBcl2fastq(samplesheet)
// Promote Hamming violations to errors (default: warn)
SamplesheetReshape.validateBcl2fastq(samplesheet, [hammingDistanceAsError: true])Run this check before a 24-hour sequencing run, so problems
like an ATXG index sequence or two samples sharing I7+I5 don't
produce data you'll have to discard.
For an end-of-pipeline pre-flight that runs both the bcl2fastq checks and the fastq matching in one pass, use:
SamplesheetReshape.validate(samplesheet, fastqDir, [validateStructure: true])Example error from a real call:
Illumina samplesheet has bcl2fastq validation issues (file: /data/Sheet.csv):
Sample 'sample_X': I7 (index) contains invalid characters: 'ATXGACGT' (only A, C, G, T, N allowed)
Duplicate Sample_ID: 'sample_Y'
Duplicate index combination (I7='ATCGACGT', I5='GCTAGCTA') used by samples: sample_Z, sample_W
All three problems are reported in one error, so the user fixes them in one pass instead of discovering them one at a time.
bin/test.sh # nf-test suite if available, else smoke test
bin/test.sh --nf-test # nf-test suite only (error if nf-test missing)
bin/test.sh --smoke # Nextflow-based smoke test only| mode | what it does | needs on PATH |
|---|---|---|
| (default) | runs unit + nf-test (or unit + smoke if nf-test is missing) | groovy and (nf-test or nextflow) |
--unit |
runs the 79-test bespoke Groovy unit suite (tests/test_*.groovy) — fast, in-process, no Nextflow |
groovy |
--nf-test |
runs unit + nf-test against subworkflows/nf-core/reshape_samplesheet/tests/main.nf.test (5 nf-test cases) |
groovy + nf-test + nextflow |
--smoke |
runs unit + nextflow run main.nf end-to-end (asserts the emitted CSV exists with the right header) |
groovy + nextflow |
The unit suite is pure Groovy and runs in ~10 seconds without installing anything heavy. Install standalone Groovy with one of:
sdk install groovy # SDKMAN
brew install groovy # Homebrew
mise use --yes groovy@latest # mise (also pins java 21)The nf-test suite needs nf-test on PATH in addition to
Nextflow itself. Install with:
curl -fsSL https://get.nf-test.com | bash # current stable
pipx install nf-test # via pipxbin/test.sh will tell you which install you need if anything is
missing.
- unit = fast, cheap, in-process. Catches every parser / validate / matching / coercion edge case at the lib level. Doesn't exercise the Nextflow wire (no module wrapping, no path staging, no emit channel).
- nf-test = slow, full integration. Catches module-wrapping regressions, the bash/Groovy subprocess invocation, Nextflow's path staging, and the emit-channel shape.
They're not redundant — each layer catches a different class of bug. Together they give complete coverage.
# via SDKMAN
curl -s https://get.sdkman.io | bash
sdk install java 21.0.2-tem
sdk install groovy
# via Homebrew
brew install groovy
# via mise (also pins java 21, which Nextflow 26 requires)
mise use --yes java@21 groovy@latestThe mise.toml in this repo does the last option automatically.
Restored alongside nf-test in v0.2.1 — they cover complementary
concerns. Five files, 79 cases total, run via bin/test.sh --unit:
-
test_parser.groovy— 22 cases. Samplesheet parsing edge cases across both formats: missing file, empty file, whitespace-only, bcl2fastq without[Data], duplicate header columns, missingSample_IDcolumn, empty header cell, CRLF line endings, UTF-8 BOM, non-ASCII characters, multi-line quoted fields, lowercase column names. Plus row-level errors (wrong cell count, emptySample_ID) and the "all malformed rows collected into one exception" contract. -
test_reshaper.groovy— 15 cases. Reshape andwriteReshapedbehaviour: paired-end, multi-lane aggregation into a comma-joined cell, single-end emits emptyfastq_2, unrelated files ignored, LRM and bcl2fastq produce the same rows for the same fastq_dir, strict failure modes (lists ALL missing samples; non-existent and empty fastq_dir; trailing newline; opts=null is the default). -
test_validator.groovy— 20 cases. EveryvalidateBcl2fastqcheck: duplicateSample_IDs, invalid I7/I5 characters, mixed-case normalisation, duplicate I7+I5, empty I7/I5 values, inconsistent I7/I5 lengths across samples, dual-indexed distinguishability,opts.validateStructure: trueintegration through reshape. -
test_options.groovy— 14 cases. Everyoptsfootgun: the four--strandedness ''coercions (Boolean true, empty string, String"true", String"false"),recursive=true/false, opts=null == empty, named-arg sugar equivalence. -
test_matching.groovy— 8 cases. The fastq-filename matching algorithm: strict default matching (catches the substring trap), separator handling,opts.patternregex meta-char escaping, multiple${sampleId}references.
The nf-test suite in subworkflows/nf-core/reshape_samplesheet/tests/main.nf.test
exercises the full sub-workflow emit path (validate → reshape →
emit) against the bundled fixtures. Uses structural assertions (not
snapshots) because the CSV contains absolute fastq paths that include
the workdir hash and would produce a different MD5 every run:
- bcl2fastq — happy path — default opts — asserts
workflow.success, the emit channel is non-empty, the emitted filename ends withillumina_bcl2fastq.nfcore.csv, and versions are emitted. - bcl2fastq — with strandedness column — same plus asserts the
CSV header starts with
sample,fastq_1,fastq_2,strandednessand every row ends with,reverse(the strandedness value). - LRM — happy path — default opts — same plus asserts the CSV row count is 4 (matches bcl2fastq).
- missing fastq match — should fail — uses
illumina_lrm_with_orphan.csv; assertsworkflow.failedand the error report containssample_orphan. - stub — runs with
options "-stub"so the processes produce empty placeholder outputs without actually shelling out to Groovy. Required by the nf-core convention.
Each test cleans up its own scratch directory
(tests/.nfcore_test_output/) via the cleanup block.
| Layer | Catches | Speed |
|---|---|---|
unit (bin/test.sh --unit) |
semantic regressions in lib/ — every parser, validator, matching, opts-coercion case |
~10 s |
nf-test (bin/test.sh --nf-test) |
wire regressions — module wrapping, bash/Groovy subprocess, Nextflow path staging, emit channel shape | ~30 s |
smoke (bin/test.sh --smoke) |
full end-to-end via main.nf |
~60 s |
Each layer catches a different class of bug. CI runs all three on every PR.
- Strandedness: set
--strandedness <value>to add a fourthstrandednesscolumn; omit the flag to keep the 3-column output. Same value is applied to every sample row — per-sample strandedness is not supported out of the box (you'd need to post-process or extendreshapeImpl). - Output filename collision:
writeReshapedwrites${samplesheet_basename}.nfcore.csvinto the output dir. Two samplesheets with the same basename (e.g. both namedSampleSheet.csvfrom different runs) will overwrite each other. If you need to keep both, rename one of the input files or write to a differentoutdir. - Matching strategy: the default (prefix + separator) is strict
enough to avoid the common footguns. If you have a non-standard
filename layout, use
opts.patternto override — see the "How samples are matched" section for examples. - New options: add a key to the
optsMap inreshapeImpland default it conservatively. Unknown keys are silently ignored at call sites, so adding a new option is backwards-compatible.
Because Nextflow already runs on the JVM and ships Groovy in its fat JAR,
the parser and the sub-workflow can both be pure Groovy. The only
"external" runtime dependency is therefore a JVM (which Nextflow
already requires). No Python, no Node, no separate CSV library — just
a small handwritten RFC-4180 parser in lib/SamplesheetParser.groovy.