Dataset storage
datashard
AvailableA table format for datasets and blobs that does not need a database.
Storage for datasets, recordings and large files that keeps its history: records are appended, snapshots are taken, and a dataset can be read as it stood at any earlier moment. It takes the place of loose CSV and JSON dumps, hand-managed folders, and databases stood up only to hold files.
Who it is for. Python projects that persist recordings, datasets or analytical data and need versioned, readable storage.
- Status
- Available
- Part of
- Open source
- Licence
- Open source, Apache 2.0
- Tags
- PythonSnapshotsTime travelOpen formatDisk or object storage
What it does
-
Append-only, with snapshots
Records are appended and snapshotted, so a dataset can be read as it stood at a point in time.
-
Disk or S3
The same table format works on local disk or in object storage, so moving between them is a configuration change.
-
Parquet underneath
Data is stored as Parquet, so any tool that reads Parquet can read it.
For documentation, a pilot or an integration, contact us.