Skip to content

About

data generation lib

Resources

Stars

0 stars

Watchers

1 watching

Forks

Repository files navigation

data-plant

docker

Generate large synthetic CSV files — bounded by row count or byte size — or transform an existing CSV by adding, overwriting and reshaping columns.

Everything is a transducer. The size limit, the data generators and the CSV formatting compose into a single transformation chain that is run lazily against an output stream, so producing a 2 GB file costs the same memory as producing a 2 KB one. The Makefile runs it under -Xmx256m for exactly that reason.

$ data-plant csv 5rows "id uuid, kind (oneOf alpha beta gamma), n pos-int"
id,kind,n
abe0fa96-2af9-48af-8df1-78ef78832758,gamma,1091407559
4a86ff49-b123-44bc-a28e-0b854f595d26,beta,555077742
ed98c0a4-0fc0-4ca5-95fa-84d3a3173dd5,gamma,1379519777
4b4f2f47-90b6-4905-a5ee-b5c2590dac7c,beta,264819990

Output is written to stdout — redirect it to get a file. Note that 5rows produced 4 data rows: the header counts as a row. That and other sharp edges are catalogued in doc/usage.md; read it before trusting the output size.

Running it

Docker (no toolchain needed)

$ docker run --rm ghcr.io/kacurez/data-plant csv 1krows "id uuid, n int" > out.csv

From source

$ lein run -- csv 1krows "id uuid, n int" > out.csv

As a jar

$ lein uberjar                       # or: make jar
$ java -jar target/data-plant.jar csv 1krows "id uuid, n int" > out.csv

The Makefile wraps this with a deliberately small heap, demonstrating that output size is decoupled from memory use:

$ make run          # java -XX:+UseG1GC -Xms1m -Xmx256m -jar target/data-plant.jar
$ make test-jar     # builds, then generates a 10-row sample

Native binary

project.clj carries a lein-native-image configuration, but it does not build on a current GraalVM — see doc/usage.md. Use the jar or the Docker image.

Usage

data-plant csv <source> <definition> [options]

csv is the only implemented subcommand.

Option Default Meaning
-h, --help Print usage
-d, --delimiter DELIMITER , Column separator
-e, --enclosure ENCLOSURE " Quote character
-f, --file off Treat <source> as an input CSV path to transform
-x, --xform-spec off Treat <definition> as a raw Clojure transducer
-g, --gzip-output off Compress the output with gzip

Help and error messages are printed to stderr. Top-level --help exits 0; csv --help exits 1.

<source>

Either a size budget, or — with -f — a path to an existing CSV.

The size grammar is <number>[K|M|G]<rows|b|bytes>, case-insensitive:

Example Meaning
500rows 500 lines, including the header
50Krows 50 000 lines
9Mrows 9 000 000 lines
145b / 145bytes ~145 bytes (see the byte-accounting caveat)
1KB ~1 000 bytes
50MB ~50 000 000 bytes
2GB ~2 000 000 000 bytes

K/M/G are decimal — 1 KB is 1 000 bytes, not 1 024. There must be no space between the parts; 2 rows, 1234 and 2rb are all rejected.

<definition>

A sequence of column-name value pairs. Surrounding braces are optional and commas count as whitespace, so these are equivalent:

"id uuid n int"
"id uuid, n int"
"{id uuid, n int}"

A value is either a generator symbol, a constant, or a oneOf choice.

Generator Produces
int A 32-bit integer — may be negative
pos-int A non-negative integer
neg-int A negative integer
float A floating-point number
boolean true / false
string A random-length string of random printable characters
date A date
uuid A random UUID

Constants are numbers, quoted strings, booleans and characters — n 42, env "prod", ok true. Any symbol that is not in the table above is silently treated as a literal string, so a typo like intt yields the text intt in every row rather than an error.

(oneOf a b c) picks one option per row, uniformly. Options may themselves be generators or nested oneOf forms:

"kind (oneOf alpha beta gamma), mixed (oneOf int string 0)"

Examples

Fixed values, custom delimiter and quote character:

$ data-plant csv 4rows "a-a 1 b 2" -d'|' -e'-'
-a--a-|b
1|2
1|2
1|2

Add a column to an existing file (.gz input is decompressed automatically):

$ data-plant csv in.csv "batch 42" -f
id,name,batch
1429069737,"j$J|9MBC|UQ.91$R=,EB^[TNW""(5[Roj)b:HDd1(Zv:%e0",42
...

Overwrite an existing column — naming a column that already exists replaces it, which makes redaction a one-liner:

$ data-plant csv in.csv.gz "name REDACTED" -f
id,name
1429069737,REDACTED
...

Supply a raw transducer instead of a definition map:

$ data-plant csv 4rows '(map #(assoc % "id" 7 "tag" "x"))' -x
id,tag
7,x
7,x
7,x

⚠ -x evaluates its argument as Clojure code. Only pass strings you wrote yourself.

More recipes: doc/usage.md.

How it works

Each CLI string is parsed into a transducer; csv/core.clj composes them and output.clj runs the composed chain against a stream.

Layer Namespace Role
Entry cli.clj -main → validate-args → try-parse-subcommand
Command csv/cli_command.clj Options, usage, prepare-run-command
Parsers parsers/{size,data,xform}_definition.clj Each turns a CLI string into a transducer via parse-to-xform
Core csv/core.clj generate-random-csv->stream, transform-csv-file->stream
Transducers csv/transducers.clj maps ↔ CSV lines, header handling, quoting, byte limiter
Generators generators.clj Random value primitives
Output output.clj transduce-coll->stream and friends — the sink

Generation feeds an infinite (repeat {}) through the chain; the size limit is the transducer that terminates it. Transformation swaps the infinite source for a lazily-read CSV file.

Development

$ lein test          # 25 tests, 209 assertions
$ make jar           # uberjar → target/data-plant.jar
$ make clean

lein test also runs inside the Docker build, so a failing test fails the image build in CI.

Images are published to ghcr.io/kacurez/data-plant for linux/amd64 and linux/arm64 by .github/workflows/docker.yml:

Trigger Tags
Push to master sha-<short>
Push of a v* tag latest, the matching semver (e.g. 0.2.0 and 0.2), sha-<short>
Pull request none — builds and smoke-tests without pushing

latest therefore follows releases, not master.

License

Copyright © 2018 kacurez

Distributed under the Eclipse Public License either version 1.0 or (at your option) any later version.

About

data generation lib

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages