Extractor configuration
Use extractor: for one source or extractors: for a list. Each spec needs
type. File extractors typically set mode and method.
Modes
singlefile, api, code. Unknown modes fail validation.
File types
file-csv, file-json, file-jsonl, file-xml, file-xls, file-xlsx,
file-zip.
Methods: url (needs config.url), urlbypattern (needs prefix and
data_prefix).
force: true re-downloads into current/.
APIBackuper
extractor:
type: api
method: apibackuper
mode: api
Keep apibackuper.cfg under storage/.
Code extractor
extractor:
type: code
mode: code
config:
script: scripts/collect.py
The script must live under the project directory and expose collect(). It is
trusted input. See Security.
RSS and DCAT
extractors:
- name: feed
type: rss
config:
url: https://example.com/feed.xml
download_enclosures: false # true: fetch enclosure files into current/
- name: catalog
type: dcat
config:
url: https://example.com/catalog.json
download: false # true: fetch distribution files
format: csv # only distributions matching this format
Download options
URL downloads accept optional keys in config (they apply to url,
urlbypattern, and feed/catalog downloads):
config:
url: https://example.com/data.csv
timeout: 60 # request timeout in seconds (default: 30)
verify_tls: false # only for trusted endpoints; logs a warning
# aria2: true # hand the download to aria2 (invoked without a shell)
# aria2path: aria2c # explicit aria2 binary path
TLS verification is on by default. See Security.
Concepts: Extractors.