CLI reference
Every command, option and exit code.
Tabalyst for developers
Tabalyst is an open-source, local-first toolkit. Install it with pip, point it at a CSV, Excel, JSON or JSONL file, and get an HTML report for you and a JSON profile for your code. No service to call, no telemetry.
tabalyst report basic.csv -d reports/
Error: Report output already exists. Use --force to replace it:
reports/basic.html
reports/basic.json
tabalyst report shop.json
2 collections of shop.json are equally plausible.
Nothing was analyzed.
Candidates:
$.customers[] (2 elements)
$.products[] (2 elements)
Real output of tabalyst 0.6.0, paths shortened.
Your angle
pip install tabalyst gives you the tabalyst command and the Python package. One command per file or per folder, and one report with its profile for each file.
Every run writes a JSON profile beside the HTML report: the same analysis, structured. The format is versioned with format_version and format_revision, and still experimental: check both before you rely on a field.
Exit codes tell an invalid command or a choice to make (2) from a failed analysis (1) and an unreadable input (4). A failed file does not stop the others, and nothing is replaced without --force.
Workflow
Tabalyst Report runs on your machine and needs Python 3.11 or later. The command line and the Python API do the same work.
pip install tabalyst# Also updates it: Tabalyst changes often
tabalyst report data.csv# data.html, data.json and executions.json, beside the source
tabalyst report *.csv -d reports/# One report and one profile per file. A failed file does not stop the others
import tabalyst # One file: the JSON profile comes back as a dictprofile = tabalyst.analyze("data.csv", "reports/data.html")print(profile["datasets"][0]["summary"]["row_count"]) # A batch: successes and failures are kept apartbatch = tabalyst.generate_reports(["*.csv"], output_dir="reports")print(batch.succeeded, len(batch.failures))
Expected failures raise tabalyst.TabalystError. Pass force=True to replace existing outputs on purpose. Settings go in a strict JSON file given with --config, or with config_path in Python; explicit arguments override it.
The toolkit chains. tabalyst inspect orders.json writes the choice of collection to a file you can edit. tabalyst scan writes the complete description of every field. tabalyst report --scan file.scan.json builds a report from a scan document without reading the source again.
Proof
The profile of the example report, shortened. The report that people read is rendered from it.
{
"format_version": "0.1.0a",
"format_revision": 11,
"source": {
"filename": "insurance-customers.csv",
"format": "csv",
"size_bytes": 827600,
"delimiter": ","
},
"datasets": [{
"id": "rows",
"summary": {
"row_count": 3000,
"column_count": 34,
"missing_count": 12532,
"duplicate_row_count": 8
},
"columns": [ ... ],
"issues": [ ... ]
}]
}Synthetic insurance customers, analyzed as they are. Shortened: some fields are left out.
Scope, not hype
format_version and format_revision.tabalyst.analyze() reads CSV files. For Excel, JSON and JSONL, use generate_reports(), the command line or scan().*.csv match the files of one folder, not its subfolders.The known limitations ↗ list the rest.
Next step
Tabalyst is open source under MPL-2.0. Read the source, report a problem or ask for a feature on GitHub.
Every command, option and exit code.
The public functions and their results.
Source code, issues and releases.
The published package.
Resources and integration notes in one place.
Your next pipeline
Install Tabalyst, run it on a file or a folder, and read the profile from your own code.