Skip to content

Latest commit

 

History

704 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

GitHub License Documentation Status CPP20 PyPi Coverage Discover on MyBinder Try CDFpp Explorer

CDFpp (CDF++)

A modern, from-scratch C++20 implementation of NASA's CDF (Common Data Format) with full Python bindings.

📖 Documentation: quickstart, guides for reading and for producing ISTP-compliant files, cookbook, and C++ guide. The examples run live in your browser.

Why another CDF library? NASA's official C implementation has no multi-thread support (global shared state), an aging C89 interface, and a license incompatible with most Linux distribution policies. CDFpp solves all three: it is thread-safe, idiomatic C++20, and MIT-licensed.

Highlights

  • C++20 library, mostly headers — see the C++ guide for the compiler flags
  • Complete read/write support — CDF versions 2.2 through 3.x, row and column major, compressed files and variables (GZip, RLE)
  • Python bindings (pycdfpp) via pybind11 — zero-copy NumPy integration, GIL-free I/O
  • SIMD-accelerated time conversions — AVX512/AVX2/SSE2 runtime dispatch for TT2000, EPOCH, EPOCH16
  • Lazy loading — variable data is read on first access, not at file open
  • Fast — 5.6× to several hundred times faster than spacepy and cdflib at reading, 3.7× to 14× at writing compressed files (see below); SIMD time conversions at up to 8 billion epochs/s (TT2000, AVX-512)
  • Runs everywhere — Linux, Windows, macOS (x86_64 + ARM64), and WebAssembly (Pyodide / emscripten-forge)
  • In-browser app — CDFpp Explorer inspects, plots, and ISTP-validates CDF files entirely client-side, no install

Packages & CI

Linux x86_64 Linux aarch64 Windows x86_64 macOS x86_64 macOS ARM64 WASM (Pyodide)
Wheels
Tests

Also available on emscripten-forge for use in JupyterLite and other Emscripten-based environments.


CDFpp Explorer — CDF files in your browser

CDFpp compiles to WebAssembly, so the whole library runs client-side — no install, no server, and your files never leave your machine. CDFpp Explorer is a small web app built on it:

  • Inspect — browse variables and attributes grouped by ISTP VAR_TYPE, with live search and a value preview
  • Plot — line plots and spectrograms, ISTP-aware (DEPEND_0 time axis, DISPLAY_TYPE, SCALETYP, fill/valid masking), with CSV / JSON export
  • Validate — hand the file to AstraLint for ISTP conformance checking

👉 Open CDFpp Explorer

The WebAssembly wrapper lives in wacdfpp/ and ships TypeScript declarations (wacdfpp/cdfpp.d.ts) for embedding CDF read/write — with zero-copy typed arrays — in your own JavaScript app.


Installing

From PyPI

pip install pycdfpp

C++ library

See Adding CDFpp to your project in the C++ guide.

From source (Python wheel)

python -m build .
# wheel is in dist/

Quick start (Python)

Reading a CDF file

import pycdfpp

cdf = pycdfpp.load("my_data.cdf")

# Variables expose numpy arrays (zero-copy when possible)
data = cdf["variable_name"].values

# Global attributes
print(cdf.attributes["Project"][0])

# Variable attributes
print(cdf["variable_name"].attributes["UNITS"].value)

# Iterate over all variables
for name, var in cdf.items():
    print(f"{name}: shape={var.shape}, type={var.type}")

Loading from memory

import pycdfpp

# Load from any bytes-like object (useful with HTTP responses, S3, etc.)
with open("my_data.cdf", "rb") as f:
    cdf = pycdfpp.load(f.read())

Time conversions

CDFpp handles all three CDF time types (EPOCH, EPOCH16, TT2000) and converts them to numpy datetime64[ns] or Python datetime:

import pycdfpp
import numpy as np

cdf = pycdfpp.load("my_data.cdf")

# Convert any CDF time variable to numpy datetime64 (fast, ~2ns/element)
times = pycdfpp.to_datetime64(cdf["Epoch"])

# Or to Python datetime objects
times_dt = pycdfpp.to_datetime(cdf["Epoch"])

# Format as strings (e.g. for PDS4 compliance)
time_strings = pycdfpp.to_time_string(cdf["Epoch"], "%Y-%m-%dT%H:%M:%SZ")
# array([b'2020-02-01T00:00:00.000000000Z', ...])

# Convert Python/numpy times to CDF types
tt2000_values = pycdfpp.to_tt2000(np.array(['2020-01-01', '2020-06-15'], dtype='datetime64[ns]'))
epoch_values = pycdfpp.to_epoch(np.array(['2020-01-01', '2020-06-15'], dtype='datetime64[ns]'))

Creating and writing CDF files

import pycdfpp
import numpy as np
from datetime import datetime

cdf = pycdfpp.CDF()

# Add global attributes (each entry is a list of values)
cdf.add_attribute("Project", ["MyMission"])
cdf.add_attribute("PI_name", ["Jane Doe"])

# Add a time variable with TT2000 encoding
times = np.arange('2020-01-01', '2020-01-02', dtype='datetime64[h]').astype('datetime64[ns]')
cdf.add_variable("Epoch", values=times, data_type=pycdfpp.DataType.CDF_TIME_TT2000)

# Add a data variable with variable attributes
cdf.add_variable("B_GSM",
    values=np.random.randn(24, 3).astype(np.float32),
    attributes={
        "FIELDNAM": "Magnetic Field",
        "UNITS": "nT",
        "DEPEND_0": "Epoch",
    })

# Save to disk
pycdfpp.save(cdf, "output.cdf")

# Or save to memory (returns bytes)
data = pycdfpp.save(cdf)

Compressed CDF files

import pycdfpp
import numpy as np

cdf = pycdfpp.CDF()
cdf.add_variable("data", values=np.zeros(10000, dtype=np.float64))

# Whole-file GZip compression
cdf.compression = pycdfpp.CompressionType.gzip_compression
pycdfpp.save(cdf, "compressed.cdf")

# Or per-variable compression
cdf2 = pycdfpp.CDF()
cdf2.add_variable("data",
    values=np.zeros(10000, dtype=np.float64),
    compression=pycdfpp.CompressionType.gzip_compression)

Filtering variables

import pycdfpp

cdf = pycdfpp.load("large_file.cdf")

# Keep only specific variables and attributes (returns a new CDF)
filtered = cdf.filter(variables=["Epoch", "B_GSM"], attributes=["Project"])

# Filter with a regex pattern
filtered = cdf.filter(variables="B_.*", attributes=".*")

# Filter with a callable
filtered = cdf.filter(variables=lambda var: var.name.startswith("B_"))

Cloning variables between CDF files

import pycdfpp

src = pycdfpp.load("source.cdf")
dst = pycdfpp.CDF()

# Clone a variable (deep copy, including its attributes)
dst.add_variable(src["Epoch"])
dst.add_variable(src["B_GSM"])

NumPy buffer protocol

Variables implement the Python buffer protocol, so they work directly with NumPy and any library that accepts array-like objects:

import pycdfpp
import numpy as np

cdf = pycdfpp.load("my_data.cdf")

# Direct numpy array construction (zero-copy for numeric types)
arr = np.array(cdf["B_GSM"])

Quick start (C++)

Reading

#include "cdfpp/cdf-io/cdf-io.hpp"
#include <iostream>

int main()
{
    // cdf::io::load returns std::optional<CDF>
    if (auto cdf = cdf::io::load("my_data.cdf"))
    {
        for (const auto& [name, variable] : cdf->variables)
            std::cout << name << " shape: " << variable.shape() << "\n";

        for (const auto& [name, attribute] : cdf->attributes)
            std::cout << name << "\n";
    }
}

Writing

#include "cdfpp/cdf-io/cdf-io.hpp"

int main()
{
    cdf::CDF my_cdf;

    // Save to file (returns bool)
    cdf::io::save(my_cdf, "output.cdf");

    // Or save to memory (returns a vector<char>)
    auto data = cdf::io::save(my_cdf);
}

Loading from memory

#include "cdfpp/cdf-io/cdf-io.hpp"
#include <vector>

void process(const std::vector<char>& buffer)
{
    if (auto cdf = cdf::io::load(buffer.data(), buffer.size()))
    {
        // ...
    }
}

Benchmarks

Compared with spacepy and cdflib

Everyday tasks on real CDAWeb files, with each library used the way its documentation shows, including its fastest time conversion. All three return the same values.

Task Data pycdfpp spacepy.pycdf cdflib
Open a file, list variables, read all attributes MMS FPI electron distribution, 178 MB 0.6 ms 264 ms (407×) 9.9 ms (15×)
Read B and its time axis as datetime64 MMS FGM survey, 1.2 M points, gzip, TT2000 7.8 ms 3.77 s (483×) 306 ms (39×)
Read B and its time axis as datetime64 Wind MFI, 0.9 M points, CDF_EPOCH 3.6 ms 40.2 ms (11×) 13.3 s (3718×)
Read every variable of a file MMS FPI electron distribution, 178 MB, gzip 110 ms 960 ms (8.7×) 620 ms (5.6×)
Read every variable of a folder 23 CDAWeb files, 11 missions, 528 MB 391 ms 2.37 s (6.1×) 2.81 s (7.2×)
Same folder, 8 threads 23 CDAWeb files, 11 missions, 528 MB 161 ms not thread-safe 2.41 s (15×)
Write B and its time axis, gzip MMS FGM survey, 1.2 M points, 29 MB 46.0 ms 458 ms (10×) 172 ms (3.7×)
Write B and its time axis, uncompressed MMS FGM survey, 1.2 M points, 29 MB 14.1 ms 12.3 ms (0.9×) 17.4 ms (1.2×)
Write a particle distribution file, gzip MMS FPI electron distribution, 210 MB 275 ms 3.93 s (14×) 1.87 s (6.8×)

(N×) = N times longer than pycdfpp. Median of 5 runs, files in the page cache. AMD Ryzen 7 5800X (AVX2, no AVX-512), Python 3.13, pycdfpp 0.15.0, spacepy 0.7.0 (NASA CDF 3.9.0), cdflib 1.3.14.

Why is pycdfpp faster?

  • Opening a file only parses the headers, in C++. NASA's library, used by spacepy, hashes the whole file to check its MD5 checksum on every open. pycdfpp skips that check.
  • Time conversion runs in C++ with SIMD, straight to datetime64[ns]. For TT2000, spacepy creates one Python datetime per value. cdflib converts CDF_EPOCH in a Python loop.
  • Gzip blocks are decompressed on all cores at once (an MMS FPI distribution variable has 640 of them), with libdeflate, itself 1.5–1.7× faster than zlib on these files. Big buffers use 2 MB huge pages.
  • Threads work: pycdfpp releases the GIL while reading and decompressing. cdflib is mostly Python, so it holds the GIL. NASA's library keeps global state.
  • Writing compresses 256 KB blocks on all cores, with libdeflate. The other two compress with zlib, on one thread. Without compression there is little to gain. With copy=False, pycdfpp writes straight from the arrays, like spacepy. The rest of the gap is the time axis: pycdfpp converts it from datetime64 (1.3 ms), spacepy is given TT2000 integers.

Reading scales to about 2× with threads, and to 3.5 GB/s on big files; writing reaches 710 MB/s with gzip, where spacepy stays at 42 MB/s. See the scaling results.

Details and caveats in the performance page. Reproduce with benchmarks/python_libs/compare.py.

C++ micro-benchmarks

Release builds (-O3). Source code in benchmarks/.

SIMD time conversions

Converting CDF time types to nanoseconds since 1970 (epochs/s, higher is better, one thread). CDFpp picks the best instruction set the CPU has at run time (AVX-512, AVX2 or SSE2).

AMD Ryzen 7 7840U/HS laptop CPU (Zen 4, 5.1 GHz boost, 16 MB L3), AVX-512, measured with pycdfpp 0.13, before CDF_EPOCH conversions became exact (see below):

Conversion 64 1K 64K 1M 64M
TT2000 scalar 7.8e+08 8.7e+08 8.8e+08 8.6e+08 5.9e+08
TT2000 SIMD 2.5e+09 8.1e+09 5.3e+09 3.4e+09 1.5e+09
EPOCH scalar 2.2e+09 2.3e+09 2.2e+09 2.1e+09 1.1e+09
EPOCH SIMD 9.6e+09 1.4e+10 6.7e+09 3.8e+09 1.5e+09

AMD Ryzen 7 5800X desktop CPU (Zen 3, 32 MB L3), AVX2 (median of 3 runs):

Conversion 64 1K 64K 1M 64M
TT2000 scalar 8.4e+08 9.7e+08 1.0e+09 1.0e+09 9.2e+08
TT2000 SIMD 2.6e+09 2.9e+09 2.8e+09 2.8e+09 1.3e+09
EPOCH scalar 9.7e+08 1.1e+09 1.1e+09 1.1e+09 9.5e+08
EPOCH SIMD 2.6e+09 2.6e+09 2.7e+09 2.7e+09 1.4e+09

With AVX-512, TT2000 conversion peaked at ~8 billion epochs/s and EPOCH at ~14 billion epochs/s for L1/L2-resident data. With AVX2, SIMD runs 2.5–3× faster than scalar code. CDF_EPOCH conversions are exact since 0.14.0: they used to round to 256 ns. That made scalar EPOCH conversion about 2× slower (it was 2.3e+09 epochs/s on the 5800X), while the SIMD version, reworked to need no 64-bit integer conversions, is faster than the old scalar one. At 64M values (1 GB of data), every SIMD row drops to 1.2–1.5 billion epochs/s: the data no longer fits in cache.

Leap-second lookup

Epochs/s, one row per CPU:

Method CPU 1K 64K 1M 64M
Branchless 7840U/HS 9.1e+07 9.4e+07 9.5e+07 1.0e+08
Branchless 5800X 1.2e+08 1.1e+08 1.2e+08 1.2e+08
Baseline 7840U/HS 2.4e+08 2.4e+08 2.4e+08 2.4e+08
Baseline 5800X 3.1e+08 3.2e+08 3.2e+08 3.2e+08

RLE compression (bytes/s)

Operation CPU 1 KB 16 KB 64 KB 1 MB
Deflate 7840U/HS 7.4e+08 6.9e+08 3.3e+08 2.7e+08
Deflate 5800X 6.2e+08 3.8e+08 2.7e+08 2.6e+08
Inflate 7840U/HS 1.8e+09 1.7e+09 1.8e+09 4.5e+08
Inflate 5800X 1.3e+09 6.4e+08 4.4e+08 4.2e+08
Roundtrip 7840U/HS 5.1e+08 5.0e+08 1.9e+08 1.7e+08
Roundtrip 5800X 4.5e+08 2.0e+08 1.7e+08 1.7e+08

RLE inflate sustains ~1.8 GB/s on the 7840U/HS for data that fits in cache; on the 5800X it drops from 1.3 GB/s at 1 KB to 0.4 GB/s at 64 KB. At 1 MB both CPUs run at the same speed.


Features & roadmap

  • Reading
    • CDF versions 2.2 through 3.x
    • Compressed files and variables (GZip, RLE)
    • Row and column major
    • Nested VXRs
    • Lazy variable loading
    • UTF-8 and ISO 8859-1 (Latin-1, auto-converted to UTF-8)
    • In-memory loading (std::vector<char>, char*, Python bytes)
    • DEC floating-point encoding (VAX, Alpha, Itanium)
    • Pad values and sparse records
  • Writing
    • Uncompressed and compressed files/variables
    • All numeric types, strings, datetime types
    • Pad values
  • General
    • libdeflate for faster GZip
    • SIMD time conversions (AVX512/AVX2/SSE2 with runtime dispatch)
    • Leap-second handling
    • Python bindings with GIL-free I/O
    • Documentation

Caveats

  • NRV variables shape: PyCDFpp exposes the record count as the first dimension, so NRV variables will have shape (0, ...) or (1, ...).
  • Reference invalidation: Python wrappers hold references into C++ containers. Adding or removing variables/attributes can invalidate them. Always re-fetch after mutation:
    # UNSAFE - ref may dangle after add_variable
    var = cdf["B_GSM"]
    cdf.add_variable("new_var", values=np.zeros(10))
    var.values  # potential segfault
    
    # SAFE - re-fetch
    cdf.add_variable("new_var", values=np.zeros(10))
    var = cdf["B_GSM"]

See the full documentation for more details.

About

A modern C++ header only cdf library with Python bindings

Topics

Resources

Contributing

Stars

13 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages