Tabalyst ToolsAvailable

Describe every field, once.

Tabalyst Scan is the analysis engine of Tabalyst. It reads a CSV, JSON, JSONL or Excel file once, as a stream, and writes one JSON document that describes every field. Tabalyst Report is built on it, and you can use it directly.

One scanA document per source
tabalyst scan customers.csv -o customers.scan.json
Scanned 3,000 records: 1 dataset, 34 fields.
Scan: customers.scan.json
Then a reportWithout reading the source again
tabalyst report --scan customers.scan.json
Analyzed 3,000 rows and 34 columns.
Report: customers.html

Real output of tabalyst 0.6.1, paths shortened.

The tool

What a scan records

  • 01 / EVERY FIELD

    Presence, types, gaps, values

    For each field: how often it is present, its native types, missing values, frequencies, exact statistics, normalization variants and a technical type.

  • 02 / DETECTORS

    What a value looks like

    Numbers with decimal commas, dates, booleans, enumerations, email addresses, URLs, phone numbers, postal codes, currency amounts, percentages, quantities, UUIDs and IP addresses. Values of sensitive fields, such as email addresses, are masked by default.

  • 03 / REUSABLE

    Read once, use many times

    The result is a JSON document. A report built from it does not read the source again, and Tabalyst refuses a scan whose source changed since it was written.

Scan reads a file as a stream, with memory bounded by configurable limits. Files of 16 MiB or more are analyzed by several worker processes, with the same result; --workers sets their number. Results are written atomically: an interrupted scan never leaves a partial file.

Get started

Scan a file, reuse the result

  1. 1 · Install
    pip install tabalyst

    # Also updates it: Tabalyst changes often

  2. 2 · A CSV file
    tabalyst scan customers.csv -o customers.scan.json

    # Without -o or -d, the scan is stored for reuse in Tabalyst's local storage

  3. 3 · JSON, JSONL, Excel
    tabalyst scan orders.json --collection "$.customers[]"

    # Same command; --collection chooses an array or a sheet when needed

  4. 4 · Many files
    tabalyst scan *.csv -d scans/

    # One scan document per source

  5. 5 · A report from a scan
    tabalyst report --scan customers.scan.json

    # Without reading the source again

tabalyst report customers.csv reuses the stored scan when the source content and the scan settings are current, and replaces a stale one.

Real output

The scan document

The scan document describes the source, then each dataset and each of its fields. Its format is versioned and documented.

customers.scan.jsonOne document, every field
{
  "format": "tabalyst.scan",
  "format_version": "0.1.0a",
  "status": "complete",
  "source": { ... },
  "config": { ... },
  "datasets": [{
    "id": "rows",
    "kind": "table",
    "record_count": 3000,
    "fields": [{
      "name": "filiale",
      "occurrences": 3000,
      "native_types": { "string": 3000 },
      "missing": { "count": 0, ... },
      "values": { "cardinality": { ... }, ... }
    }, ... ]
  }]
}

Synthetic insurance customers. Shortened: each field also holds its statistics and detector results.

Scope, not hype

What to know before you rely on it

  • The scan is an engine first: most people use it through Tabalyst Report. Use Scan directly when you want the full description of the fields, without a report.
  • An Excel sheet is read into memory, and a sheet above 1 GiB of XML is refused. Tabalyst publishes no benchmark: test your own files.
  • A scan written by another version of Tabalyst, or with other settings than the ones requested, is replaced or refused, never silently reused.
  • The scan format is still experimental: check format_version and format_revision before you rely on a field.

The known limitations ↗ list the rest.

Next step

Where to go next