A modern, from-scratch C++20 implementation of NASA's CDF (Common Data Format) with full Python bindings.
📖 Documentation: quickstart, guides for reading and for producing ISTP-compliant files, cookbook, and C++ guide. The examples run live in your browser.
Why another CDF library? NASA's official C implementation has no multi-thread support (global shared state), an aging C89 interface, and a license incompatible with most Linux distribution policies. CDFpp solves all three: it is thread-safe, idiomatic C++20, and MIT-licensed.
- C++20 library, mostly headers — see the C++ guide for the compiler flags
- Complete read/write support — CDF versions 2.2 through 3.x, row and column major, compressed files and variables (GZip, RLE)
- Python bindings (
pycdfpp) via pybind11 — zero-copy NumPy integration, GIL-free I/O - SIMD-accelerated time conversions — AVX512/AVX2/SSE2 runtime dispatch for TT2000, EPOCH, EPOCH16
- Lazy loading — variable data is read on first access, not at file open
- Fast — 5.6× to several hundred times faster than spacepy and cdflib at reading, 3.7× to 14× at writing compressed files (see below); SIMD time conversions at up to 8 billion epochs/s (TT2000, AVX-512)
- Runs everywhere — Linux, Windows, macOS (x86_64 + ARM64), and WebAssembly (Pyodide / emscripten-forge)
- In-browser app — CDFpp Explorer inspects, plots, and ISTP-validates CDF files entirely client-side, no install
| Linux x86_64 | Linux aarch64 | Windows x86_64 | macOS x86_64 | macOS ARM64 | WASM (Pyodide) | |
|---|---|---|---|---|---|---|
| Wheels | ||||||
| Tests |
Also available on emscripten-forge for use in JupyterLite and other Emscripten-based environments.
CDFpp compiles to WebAssembly, so the whole library runs client-side — no install, no server, and your files never leave your machine. CDFpp Explorer is a small web app built on it:
- Inspect — browse variables and attributes grouped by ISTP
VAR_TYPE, with live search and a value preview - Plot — line plots and spectrograms, ISTP-aware (
DEPEND_0time axis,DISPLAY_TYPE,SCALETYP, fill/valid masking), with CSV / JSON export - Validate — hand the file to AstraLint for ISTP conformance checking
The WebAssembly wrapper lives in wacdfpp/ and ships TypeScript declarations (wacdfpp/cdfpp.d.ts) for embedding CDF read/write — with zero-copy typed arrays — in your own JavaScript app.
pip install pycdfppSee Adding CDFpp to your project in the C++ guide.
python -m build .
# wheel is in dist/import pycdfpp
cdf = pycdfpp.load("my_data.cdf")
# Variables expose numpy arrays (zero-copy when possible)
data = cdf["variable_name"].values
# Global attributes
print(cdf.attributes["Project"][0])
# Variable attributes
print(cdf["variable_name"].attributes["UNITS"].value)
# Iterate over all variables
for name, var in cdf.items():
print(f"{name}: shape={var.shape}, type={var.type}")import pycdfpp
# Load from any bytes-like object (useful with HTTP responses, S3, etc.)
with open("my_data.cdf", "rb") as f:
cdf = pycdfpp.load(f.read())CDFpp handles all three CDF time types (EPOCH, EPOCH16, TT2000) and converts them to numpy datetime64[ns] or Python datetime:
import pycdfpp
import numpy as np
cdf = pycdfpp.load("my_data.cdf")
# Convert any CDF time variable to numpy datetime64 (fast, ~2ns/element)
times = pycdfpp.to_datetime64(cdf["Epoch"])
# Or to Python datetime objects
times_dt = pycdfpp.to_datetime(cdf["Epoch"])
# Format as strings (e.g. for PDS4 compliance)
time_strings = pycdfpp.to_time_string(cdf["Epoch"], "%Y-%m-%dT%H:%M:%SZ")
# array([b'2020-02-01T00:00:00.000000000Z', ...])
# Convert Python/numpy times to CDF types
tt2000_values = pycdfpp.to_tt2000(np.array(['2020-01-01', '2020-06-15'], dtype='datetime64[ns]'))
epoch_values = pycdfpp.to_epoch(np.array(['2020-01-01', '2020-06-15'], dtype='datetime64[ns]'))import pycdfpp
import numpy as np
from datetime import datetime
cdf = pycdfpp.CDF()
# Add global attributes (each entry is a list of values)
cdf.add_attribute("Project", ["MyMission"])
cdf.add_attribute("PI_name", ["Jane Doe"])
# Add a time variable with TT2000 encoding
times = np.arange('2020-01-01', '2020-01-02', dtype='datetime64[h]').astype('datetime64[ns]')
cdf.add_variable("Epoch", values=times, data_type=pycdfpp.DataType.CDF_TIME_TT2000)
# Add a data variable with variable attributes
cdf.add_variable("B_GSM",
values=np.random.randn(24, 3).astype(np.float32),
attributes={
"FIELDNAM": "Magnetic Field",
"UNITS": "nT",
"DEPEND_0": "Epoch",
})
# Save to disk
pycdfpp.save(cdf, "output.cdf")
# Or save to memory (returns bytes)
data = pycdfpp.save(cdf)import pycdfpp
import numpy as np
cdf = pycdfpp.CDF()
cdf.add_variable("data", values=np.zeros(10000, dtype=np.float64))
# Whole-file GZip compression
cdf.compression = pycdfpp.CompressionType.gzip_compression
pycdfpp.save(cdf, "compressed.cdf")
# Or per-variable compression
cdf2 = pycdfpp.CDF()
cdf2.add_variable("data",
values=np.zeros(10000, dtype=np.float64),
compression=pycdfpp.CompressionType.gzip_compression)import pycdfpp
cdf = pycdfpp.load("large_file.cdf")
# Keep only specific variables and attributes (returns a new CDF)
filtered = cdf.filter(variables=["Epoch", "B_GSM"], attributes=["Project"])
# Filter with a regex pattern
filtered = cdf.filter(variables="B_.*", attributes=".*")
# Filter with a callable
filtered = cdf.filter(variables=lambda var: var.name.startswith("B_"))import pycdfpp
src = pycdfpp.load("source.cdf")
dst = pycdfpp.CDF()
# Clone a variable (deep copy, including its attributes)
dst.add_variable(src["Epoch"])
dst.add_variable(src["B_GSM"])Variables implement the Python buffer protocol, so they work directly with NumPy and any library that accepts array-like objects:
import pycdfpp
import numpy as np
cdf = pycdfpp.load("my_data.cdf")
# Direct numpy array construction (zero-copy for numeric types)
arr = np.array(cdf["B_GSM"])#include "cdfpp/cdf-io/cdf-io.hpp"
#include <iostream>
int main()
{
// cdf::io::load returns std::optional<CDF>
if (auto cdf = cdf::io::load("my_data.cdf"))
{
for (const auto& [name, variable] : cdf->variables)
std::cout << name << " shape: " << variable.shape() << "\n";
for (const auto& [name, attribute] : cdf->attributes)
std::cout << name << "\n";
}
}#include "cdfpp/cdf-io/cdf-io.hpp"
int main()
{
cdf::CDF my_cdf;
// Save to file (returns bool)
cdf::io::save(my_cdf, "output.cdf");
// Or save to memory (returns a vector<char>)
auto data = cdf::io::save(my_cdf);
}#include "cdfpp/cdf-io/cdf-io.hpp"
#include <vector>
void process(const std::vector<char>& buffer)
{
if (auto cdf = cdf::io::load(buffer.data(), buffer.size()))
{
// ...
}
}Everyday tasks on real CDAWeb files, with each library used the way its documentation shows, including its fastest time conversion. All three return the same values.
| Task | Data | pycdfpp | spacepy.pycdf | cdflib |
|---|---|---|---|---|
| Open a file, list variables, read all attributes | MMS FPI electron distribution, 178 MB | 0.6 ms | 264 ms (407×) | 9.9 ms (15×) |
Read B and its time axis as datetime64 |
MMS FGM survey, 1.2 M points, gzip, TT2000 | 7.8 ms | 3.77 s (483×) | 306 ms (39×) |
Read B and its time axis as datetime64 |
Wind MFI, 0.9 M points, CDF_EPOCH | 3.6 ms | 40.2 ms (11×) | 13.3 s (3718×) |
| Read every variable of a file | MMS FPI electron distribution, 178 MB, gzip | 110 ms | 960 ms (8.7×) | 620 ms (5.6×) |
| Read every variable of a folder | 23 CDAWeb files, 11 missions, 528 MB | 391 ms | 2.37 s (6.1×) | 2.81 s (7.2×) |
| Same folder, 8 threads | 23 CDAWeb files, 11 missions, 528 MB | 161 ms | not thread-safe | 2.41 s (15×) |
| Write B and its time axis, gzip | MMS FGM survey, 1.2 M points, 29 MB | 46.0 ms | 458 ms (10×) | 172 ms (3.7×) |
| Write B and its time axis, uncompressed | MMS FGM survey, 1.2 M points, 29 MB | 14.1 ms | 12.3 ms (0.9×) | 17.4 ms (1.2×) |
| Write a particle distribution file, gzip | MMS FPI electron distribution, 210 MB | 275 ms | 3.93 s (14×) | 1.87 s (6.8×) |
(N×) = N times longer than pycdfpp. Median of 5 runs, files in the page cache. AMD Ryzen 7 5800X (AVX2, no AVX-512), Python 3.13, pycdfpp 0.15.0, spacepy 0.7.0 (NASA CDF 3.9.0), cdflib 1.3.14.
Why is pycdfpp faster?
- Opening a file only parses the headers, in C++. NASA's library, used by spacepy, hashes the whole file to check its MD5 checksum on every open. pycdfpp skips that check.
- Time conversion runs in C++ with SIMD, straight to
datetime64[ns]. For TT2000, spacepy creates one Pythondatetimeper value. cdflib converts CDF_EPOCH in a Python loop. - Gzip blocks are decompressed on all cores at once (an MMS FPI distribution variable has 640 of them), with libdeflate, itself 1.5–1.7× faster than zlib on these files. Big buffers use 2 MB huge pages.
- Threads work: pycdfpp releases the GIL while reading and decompressing. cdflib is mostly Python, so it holds the GIL. NASA's library keeps global state.
- Writing compresses 256 KB blocks on all cores, with libdeflate. The other two compress with zlib, on one thread. Without compression there is little to gain. With
copy=False, pycdfpp writes straight from the arrays, like spacepy. The rest of the gap is the time axis: pycdfpp converts it fromdatetime64(1.3 ms), spacepy is given TT2000 integers.
Reading scales to about 2× with threads, and to 3.5 GB/s on big files; writing reaches 710 MB/s with gzip, where spacepy stays at 42 MB/s. See the scaling results.
Details and caveats in the performance page. Reproduce with benchmarks/python_libs/compare.py.
Release builds (-O3). Source code in benchmarks/.
Converting CDF time types to nanoseconds since 1970 (epochs/s, higher is better, one thread). CDFpp picks the best instruction set the CPU has at run time (AVX-512, AVX2 or SSE2).
AMD Ryzen 7 7840U/HS laptop CPU (Zen 4, 5.1 GHz boost, 16 MB L3), AVX-512, measured with pycdfpp 0.13, before CDF_EPOCH conversions became exact (see below):
| Conversion | 64 | 1K | 64K | 1M | 64M |
|---|---|---|---|---|---|
| TT2000 scalar | 7.8e+08 | 8.7e+08 | 8.8e+08 | 8.6e+08 | 5.9e+08 |
| TT2000 SIMD | 2.5e+09 | 8.1e+09 | 5.3e+09 | 3.4e+09 | 1.5e+09 |
| EPOCH scalar | 2.2e+09 | 2.3e+09 | 2.2e+09 | 2.1e+09 | 1.1e+09 |
| EPOCH SIMD | 9.6e+09 | 1.4e+10 | 6.7e+09 | 3.8e+09 | 1.5e+09 |
AMD Ryzen 7 5800X desktop CPU (Zen 3, 32 MB L3), AVX2 (median of 3 runs):
| Conversion | 64 | 1K | 64K | 1M | 64M |
|---|---|---|---|---|---|
| TT2000 scalar | 8.4e+08 | 9.7e+08 | 1.0e+09 | 1.0e+09 | 9.2e+08 |
| TT2000 SIMD | 2.6e+09 | 2.9e+09 | 2.8e+09 | 2.8e+09 | 1.3e+09 |
| EPOCH scalar | 9.7e+08 | 1.1e+09 | 1.1e+09 | 1.1e+09 | 9.5e+08 |
| EPOCH SIMD | 2.6e+09 | 2.6e+09 | 2.7e+09 | 2.7e+09 | 1.4e+09 |
With AVX-512, TT2000 conversion peaked at ~8 billion epochs/s and EPOCH at ~14 billion epochs/s for L1/L2-resident data. With AVX2, SIMD runs 2.5–3× faster than scalar code. CDF_EPOCH conversions are exact since 0.14.0: they used to round to 256 ns. That made scalar EPOCH conversion about 2× slower (it was 2.3e+09 epochs/s on the 5800X), while the SIMD version, reworked to need no 64-bit integer conversions, is faster than the old scalar one. At 64M values (1 GB of data), every SIMD row drops to 1.2–1.5 billion epochs/s: the data no longer fits in cache.
Epochs/s, one row per CPU:
| Method | CPU | 1K | 64K | 1M | 64M |
|---|---|---|---|---|---|
| Branchless | 7840U/HS | 9.1e+07 | 9.4e+07 | 9.5e+07 | 1.0e+08 |
| Branchless | 5800X | 1.2e+08 | 1.1e+08 | 1.2e+08 | 1.2e+08 |
| Baseline | 7840U/HS | 2.4e+08 | 2.4e+08 | 2.4e+08 | 2.4e+08 |
| Baseline | 5800X | 3.1e+08 | 3.2e+08 | 3.2e+08 | 3.2e+08 |
| Operation | CPU | 1 KB | 16 KB | 64 KB | 1 MB |
|---|---|---|---|---|---|
| Deflate | 7840U/HS | 7.4e+08 | 6.9e+08 | 3.3e+08 | 2.7e+08 |
| Deflate | 5800X | 6.2e+08 | 3.8e+08 | 2.7e+08 | 2.6e+08 |
| Inflate | 7840U/HS | 1.8e+09 | 1.7e+09 | 1.8e+09 | 4.5e+08 |
| Inflate | 5800X | 1.3e+09 | 6.4e+08 | 4.4e+08 | 4.2e+08 |
| Roundtrip | 7840U/HS | 5.1e+08 | 5.0e+08 | 1.9e+08 | 1.7e+08 |
| Roundtrip | 5800X | 4.5e+08 | 2.0e+08 | 1.7e+08 | 1.7e+08 |
RLE inflate sustains ~1.8 GB/s on the 7840U/HS for data that fits in cache; on the 5800X it drops from 1.3 GB/s at 1 KB to 0.4 GB/s at 64 KB. At 1 MB both CPUs run at the same speed.
- Reading
- CDF versions 2.2 through 3.x
- Compressed files and variables (GZip, RLE)
- Row and column major
- Nested VXRs
- Lazy variable loading
- UTF-8 and ISO 8859-1 (Latin-1, auto-converted to UTF-8)
- In-memory loading (
std::vector<char>,char*, Pythonbytes) - DEC floating-point encoding (VAX, Alpha, Itanium)
- Pad values and sparse records
- Writing
- Uncompressed and compressed files/variables
- All numeric types, strings, datetime types
- Pad values
- General
- libdeflate for faster GZip
- SIMD time conversions (AVX512/AVX2/SSE2 with runtime dispatch)
- Leap-second handling
- Python bindings with GIL-free I/O
- Documentation
- NRV variables shape: PyCDFpp exposes the record count as the first dimension, so NRV variables will have shape
(0, ...)or(1, ...). - Reference invalidation: Python wrappers hold references into C++ containers. Adding or removing variables/attributes can invalidate them. Always re-fetch after mutation:
# UNSAFE - ref may dangle after add_variable var = cdf["B_GSM"] cdf.add_variable("new_var", values=np.zeros(10)) var.values # potential segfault # SAFE - re-fetch cdf.add_variable("new_var", values=np.zeros(10)) var = cdf["B_GSM"]
See the full documentation for more details.