Skip to content

Dataset storage

datashard

Available

A table format for datasets and blobs that does not need a database.

Storage for datasets, recordings and large files that keeps its history: records are appended, snapshots are taken, and a dataset can be read as it stood at any earlier moment. It takes the place of loose CSV and JSON dumps, hand-managed folders, and databases stood up only to hold files.

Who it is for. Python projects that persist recordings, datasets or analytical data and need versioned, readable storage.

Status
Available
Part of
Open source
Licence
Open source, Apache 2.0
Tags
PythonSnapshotsTime travelOpen formatDisk or object storage

What it does

  • Append-only, with snapshots

    Records are appended and snapshotted, so a dataset can be read as it stood at a point in time.

  • Disk or S3

    The same table format works on local disk or in object storage, so moving between them is a configuration change.

  • Parquet underneath

    Data is stored as Parquet, so any tool that reads Parquet can read it.

For documentation, a pilot or an integration, contact us.