Tabalyst Report · CSVAvailable

A CSV opens fine. Tabalyst shows what is in it.

Tabalyst Report is an open-source tool that turns a data file into an interactive HTML report and a JSON profile, on your machine. Point it at a CSV file or a whole folder of them, and see the types, gaps, ambiguous dates and duplicates before you build on the data.

One fileA report beside the source
tabalyst report insurance-customers.csv
Analyzed 3,000 rows and 34 columns.
Report: insurance-customers.html
A folderOne report per file
tabalyst report *.csv -d reports/
Report: basic.csv -> reports/basic.html
Report: insurance-customers.csv -> reports/insurance-customers.html
2 succeeded, 0 failed

Real output of tabalyst 0.6.0, paths shortened.

One clear chain · three steps

From raw file to a view you can use.

Analysis and presentation are separate. One analysis produces a JSON profile for code and a self-contained HTML report for people.

SOURCE FILE

Your CSV

The source is read locally. Original values remain visible in the bounded preview.

ANALYSIS ENGINE

Tabalyst Report

Types, missing values, duplicates, dates and patterns are profiled with explicit settings.

TWO DELIVERABLES

Human + machine

A structured profile for your tooling. An interactive report for exploration and sharing.

report.json is the canonical result. executions.json records successful analyses.
Tabalyst Report generates the structured JSON profile and the self-contained HTML report together.

Schematic rows, for illustration only.

Format

Where CSV files get tricky

  • 01 / TYPES

    Everything is text

    A column that looks like dates can mix 1978-02-30, 04/05/1983 and empty cells. A look at the first rows will not show it.

  • 02 / GAPS

    Empty cells and repeated rows

    Missing values and duplicate rows are easier to deal with before the file enters your import, analysis or pipeline.

  • 03 / EXPORTS

    Separators and encodings

    Spreadsheets and legacy systems export with a semicolon or cp1252. You say so once, with --delimiter and --encoding.

The idea

Simple. Clear. Smart.

Simple. Clear. Smart.
  • Simple

    One command: tabalyst report file.csv. It writes the HTML report, the JSON profile and the execution history beside the source.

  • Clear

    For every column: its inferred type, its missing values, its detected formats and its duplicates. An HTML report for people, a JSON profile for code.

  • Smart

    Detectors recognize dates, emails, phone numbers and more. A folder gives one report per file, and files of 16 MiB or more are analyzed by several workers on their own.

Outputs

What you get

Tabalyst Report writes three files beside the source: customers.html, an interactive report for people; customers.json, a structured profile for code; and executions.json, the history of successful runs. Add --details for one HTML page per column.

Proof

See a real CSV report

Tabalyst Report on a real CSVGenerated from the example report below
report.html report.json insurance-customers.csv Screenshots of a real Tabalyst report

TabalystReport

Explore the report

Report

Analysis

Data

Dataset overview: Size, completeness and the main signal to check first.Open in full report

See the whole-file signal first.

Columns: Inferred and semantic types per column. With issues: missing values or mixed type.Open in full report

Every column, its type and its issues.

Transformations: Occurrences changed by each normalization stage, and the spellings it groups. Raw preview values remain unchanged.Open in full report

What was normalized, column by column.

Numeric analysis: Range and distribution statistics for accepted numeric values.Open in full report

Range, mean and median of each numeric column.

Date analysis: Strict date parsing keeps ambiguous, invalid and non-date values separate. Ambiguous values are never resolved from the other values of the column.Open in full report

Ambiguous dates are flagged, never guessed.

String analysis: Length classes, fixed widths and representative values.Open in full report

The length profile of every text column.

Detectors and formats: What each detector recognized per column, with the formats it found. Primary: the interpretation shown as semantic type.Open in full report

What each detector recognized, and in which format.

Data sample: First 20 records with original row numbers and raw values.Open in full report

Raw values, exactly as they are in the file.

Analysis settings: Rules used for this analysis, so the result can be reproduced.Open in full report

Every rule used, for a reproducible analysis.

Explore the full report

A synthetic file of insurance customers, analyzed as it is. The report and the profile are the files Tabalyst Report wrote.

Source
CSV
Rows
3,000
Columns
34

Run it

Run it on your file

Tabalyst Report is open source and runs on your machine. It needs Python 3.11 or later.

  1. 1 · Install
    pip install tabalyst

    # Also updates it: Tabalyst changes often

  2. 2 · One file
    tabalyst report myfile.csv

    # myfile.html, myfile.json and executions.json, beside the source

  3. 3 · Several files
    tabalyst report *.csv -d reports/

    # One report and one profile per file, all in reports/. Files are not merged

Every option, including separator and encoding, is in the CSV guide ↗.

Scope, not hype

What to know before you start

  • A CSV has one header row. A row with too few or too many fields stops the analysis, unless you choose the tolerant error policy, which excludes and lists those rows.
  • The batch command reads the matching files of one folder, not its subfolders.
  • Tabalyst Report profiles the file. It does not clean it or check your domain rules.

The known limitations ↗ list the rest.

Your next CSV

Start with a clearer picture.

Install Tabalyst, point it at a CSV or a folder and open the reports.