Generate large synthetic CSV files — bounded by row count or byte size — or transform an existing CSV by adding, overwriting and reshaping columns.
Everything is a transducer. The size limit, the data generators and the CSV
formatting compose into a single transformation chain that is run lazily against
an output stream, so producing a 2 GB file costs the same memory as producing a
2 KB one. The Makefile runs it under -Xmx256m for exactly that reason.
$ data-plant csv 5rows "id uuid, kind (oneOf alpha beta gamma), n pos-int"
id,kind,n
abe0fa96-2af9-48af-8df1-78ef78832758,gamma,1091407559
4a86ff49-b123-44bc-a28e-0b854f595d26,beta,555077742
ed98c0a4-0fc0-4ca5-95fa-84d3a3173dd5,gamma,1379519777
4b4f2f47-90b6-4905-a5ee-b5c2590dac7c,beta,264819990Output is written to stdout — redirect it to get a file. Note that
5rowsproduced 4 data rows: the header counts as a row. That and other sharp edges are catalogued in doc/usage.md; read it before trusting the output size.
$ docker run --rm ghcr.io/kacurez/data-plant csv 1krows "id uuid, n int" > out.csv$ lein run -- csv 1krows "id uuid, n int" > out.csv$ lein uberjar # or: make jar
$ java -jar target/data-plant.jar csv 1krows "id uuid, n int" > out.csvThe Makefile wraps this with a deliberately small heap, demonstrating that
output size is decoupled from memory use:
$ make run # java -XX:+UseG1GC -Xms1m -Xmx256m -jar target/data-plant.jar
$ make test-jar # builds, then generates a 10-row sampleproject.clj carries a lein-native-image configuration, but it does not
build on a current GraalVM — see
doc/usage.md. Use the jar or the
Docker image.
data-plant csv <source> <definition> [options]
csv is the only implemented subcommand.
| Option | Default | Meaning |
|---|---|---|
-h, --help |
Print usage | |
-d, --delimiter DELIMITER |
, |
Column separator |
-e, --enclosure ENCLOSURE |
" |
Quote character |
-f, --file |
off | Treat <source> as an input CSV path to transform |
-x, --xform-spec |
off | Treat <definition> as a raw Clojure transducer |
-g, --gzip-output |
off | Compress the output with gzip |
Help and error messages are printed to stderr. Top-level --help exits 0;
csv --help exits 1.
Either a size budget, or — with -f — a path to an existing CSV.
The size grammar is <number>[K|M|G]<rows|b|bytes>, case-insensitive:
| Example | Meaning |
|---|---|
500rows |
500 lines, including the header |
50Krows |
50 000 lines |
9Mrows |
9 000 000 lines |
145b / 145bytes |
~145 bytes (see the byte-accounting caveat) |
1KB |
~1 000 bytes |
50MB |
~50 000 000 bytes |
2GB |
~2 000 000 000 bytes |
K/M/G are decimal — 1 KB is 1 000 bytes, not 1 024. There must be no
space between the parts; 2 rows, 1234 and 2rb are all rejected.
A sequence of column-name value pairs. Surrounding braces are optional and
commas count as whitespace, so these are equivalent:
"id uuid n int"
"id uuid, n int"
"{id uuid, n int}"
A value is either a generator symbol, a constant, or a oneOf choice.
| Generator | Produces |
|---|---|
int |
A 32-bit integer — may be negative |
pos-int |
A non-negative integer |
neg-int |
A negative integer |
float |
A floating-point number |
boolean |
true / false |
string |
A random-length string of random printable characters |
date |
A date |
uuid |
A random UUID |
Constants are numbers, quoted strings, booleans and characters — n 42,
env "prod", ok true. Any symbol that is not in the table above is silently
treated as a literal string, so a typo like intt yields the text intt in
every row rather than an error.
(oneOf a b c) picks one option per row, uniformly. Options may themselves be
generators or nested oneOf forms:
"kind (oneOf alpha beta gamma), mixed (oneOf int string 0)"
Fixed values, custom delimiter and quote character:
$ data-plant csv 4rows "a-a 1 b 2" -d'|' -e'-'
-a--a-|b
1|2
1|2
1|2Add a column to an existing file (.gz input is decompressed automatically):
$ data-plant csv in.csv "batch 42" -f
id,name,batch
1429069737,"j$J|9MBC|UQ.91$R=,EB^[TNW""(5[Roj)b:HDd1(Zv:%e0",42
...Overwrite an existing column — naming a column that already exists replaces it, which makes redaction a one-liner:
$ data-plant csv in.csv.gz "name REDACTED" -f
id,name
1429069737,REDACTED
...Supply a raw transducer instead of a definition map:
$ data-plant csv 4rows '(map #(assoc % "id" 7 "tag" "x"))' -x
id,tag
7,x
7,x
7,x⚠ -x evaluates its argument as Clojure code. Only pass strings you wrote
yourself.
More recipes: doc/usage.md.
Each CLI string is parsed into a transducer; csv/core.clj composes them and
output.clj runs the composed chain against a stream.
| Layer | Namespace | Role |
|---|---|---|
| Entry | cli.clj |
-main → validate-args → try-parse-subcommand |
| Command | csv/cli_command.clj |
Options, usage, prepare-run-command |
| Parsers | parsers/{size,data,xform}_definition.clj |
Each turns a CLI string into a transducer via parse-to-xform |
| Core | csv/core.clj |
generate-random-csv->stream, transform-csv-file->stream |
| Transducers | csv/transducers.clj |
maps ↔ CSV lines, header handling, quoting, byte limiter |
| Generators | generators.clj |
Random value primitives |
| Output | output.clj |
transduce-coll->stream and friends — the sink |
Generation feeds an infinite (repeat {}) through the chain; the size limit is
the transducer that terminates it. Transformation swaps the infinite source for
a lazily-read CSV file.
$ lein test # 25 tests, 209 assertions
$ make jar # uberjar → target/data-plant.jar
$ make cleanlein test also runs inside the Docker build, so a failing test fails the image
build in CI.
Images are published to ghcr.io/kacurez/data-plant for linux/amd64 and
linux/arm64 by .github/workflows/docker.yml:
| Trigger | Tags |
|---|---|
Push to master |
sha-<short> |
Push of a v* tag |
latest, the matching semver (e.g. 0.2.0 and 0.2), sha-<short> |
| Pull request | none — builds and smoke-tests without pushing |
latest therefore follows releases, not master.
Copyright © 2018 kacurez
Distributed under the Eclipse Public License either version 1.0 or (at your option) any later version.