Tabalyst Scan
Overview, commands and pages.
Tabalyst ToolsAvailable
Tabalyst Scan is the analysis engine of Tabalyst. It reads a CSV, JSON, JSONL or Excel file once, as a stream, and writes one JSON document that describes every field. Tabalyst Report is built on it, and you can use it directly.
tabalyst scan customers.csv -o customers.scan.json
Scanned 3,000 records: 1 dataset, 34 fields.
Scan: customers.scan.json
tabalyst report --scan customers.scan.json
Analyzed 3,000 rows and 34 columns.
Report: customers.html
Real output of tabalyst 0.6.1, paths shortened.
The tool
For each field: how often it is present, its native types, missing values, frequencies, exact statistics, normalization variants and a technical type.
Numbers with decimal commas, dates, booleans, enumerations, email addresses, URLs, phone numbers, postal codes, currency amounts, percentages, quantities, UUIDs and IP addresses. Values of sensitive fields, such as email addresses, are masked by default.
The result is a JSON document. A report built from it does not read the source again, and Tabalyst refuses a scan whose source changed since it was written.
Scan reads a file as a stream, with memory bounded by configurable limits. Files of 16 MiB or more are analyzed by several worker processes, with the same result; --workers sets their number. Results are written atomically: an interrupted scan never leaves a partial file.
Get started
pip install tabalyst# Also updates it: Tabalyst changes often
tabalyst scan customers.csv -o customers.scan.json# Without -o or -d, the scan is stored for reuse in Tabalyst's local storage
tabalyst scan orders.json --collection "$.customers[]"# Same command; --collection chooses an array or a sheet when needed
tabalyst scan *.csv -d scans/# One scan document per source
tabalyst report --scan customers.scan.json# Without reading the source again
import tabalyst # One scan document per source, written in scans/batch = tabalyst.generate_scans(["*.csv"], output_dir="scans")print(batch.succeeded, len(batch.failures))
tabalyst report customers.csv reuses the stored scan when the source content and the scan settings are current, and replaces a stale one.
Real output
The scan document describes the source, then each dataset and each of its fields. Its format is versioned and documented.
{
"format": "tabalyst.scan",
"format_version": "0.1.0a",
"status": "complete",
"source": { ... },
"config": { ... },
"datasets": [{
"id": "rows",
"kind": "table",
"record_count": 3000,
"fields": [{
"name": "filiale",
"occurrences": 3000,
"native_types": { "string": 3000 },
"missing": { "count": 0, ... },
"values": { "cardinality": { ... }, ... }
}, ... ]
}]
}Synthetic insurance customers. Shortened: each field also holds its statistics and detector results.
Scope, not hype
format_version and format_revision before you rely on a field.The known limitations ↗ list the rest.
Next step
Overview, commands and pages.
Commands, sources, reuse, settings and the Python API.
The structure of the scan document.
tabalyst cache info and tabalyst cache clean.