Installation
Datacrafter requires Python 3.10 or newer.
The PyPI distribution is datacrafter-etl (the datacrafter name on
PyPI belongs to an unrelated project). The Python module and the CLI command
stay datacrafter.
Using pip (recommended)
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install datacrafter-etl
datacrafter version
From source
git clone https://github.com/apicrafter/datacrafter.git
cd datacrafter
pip install -e .
For tests and linters, install the development extras:
pip install -r requirements-dev.txt
pip install -e .
Optional destinations
File JSONL, BSON, and CSV destinations ship with the core package. Some destinations need extra Python packages:
| Destination / feature | Package |
|---|---|
file-parquet | pip install "datacrafter[parquet]" (pyarrow) |
.zst compression (sources and destinations) | pip install "datacrafter[compression]" (zstandard) |
mongodb | pymongo |
arangodb | python-arango |
couchdb | pycouchdb |
meilisearch | meilisearch |
APIBackuper extractors also need the apibackuper CLI (>=1.0.4) on PATH.
datacrafter check reports missing optional packages for the destination you
configured.
Requirements
- Python 3.10 or greater (CI tests 3.10–3.13)
- A writable project directory for
current/,output/, andstate.json