Tabalyst for data professionals

Understand the dataset before analysis or modeling.

Before the first chart or the first model, see what the file holds: how complete it is, what each column contains, where values are rare or inconsistent. Tabalyst Report reads the whole file on your machine and gives you an interactive report to explore.

The overviewOne command, one report
tabalyst report insurance-customers.csv
Analyzed 3,000 rows and 34 columns.
Report: insurance-customers.html
A smaller fileA reproducible sample
tabalyst sample insurance-customers.csv --sample-method random --rows 1000 --seed 42
Sample created successfully.
Source rows:   3,000
Sample rows:   1,000

Real output of tabalyst 0.6.0, paths shortened. Some lines of the sample summary are left out.

Your angle

What you want to know first

  • 01 / OVERVIEW

    Completeness and shape first

    Rows, columns, missing cells, duplicate rows and the types found, on one screen, before you open a notebook. In the example report, 15 of 34 columns carry an observation.

  • 02 / DISTRIBUTIONS

    The values, not only the types

    Range, mean and median of numeric columns, string lengths, the frequency of each value in a column with few distinct values, and representative examples. Odd spellings and rare values stand out.

  • 03 / DOUBT

    Ambiguity is shown, not hidden

    An ambiguous date stays ambiguous: the report shows the evidence of the column and does not choose for you. Missing values are counted by kind: absent, null, empty, blank, or a marker such as N/A.

Workflow

From a file to a dataset you understand

The report is one HTML file. Open it in a browser, work through its sections, and keep the JSON profile if you want the same facts in code.

  1. 1 · Install
    pip install tabalyst

    # Also updates it: Tabalyst changes often

  2. 2 · Report a file
    tabalyst report data.csv

    # data.html to explore, data.json for code, beside the source

  3. 3 · Work on a sample
    tabalyst sample data.csv --sample-method random --rows 1000 --seed 42

    # data.sample.csv: a smaller CSV, the source is never modified

Sampling methods are first, last, random and stratified, which approximately preserves the distribution of a field you choose. Sampling reads CSV files only.

What each section of the report answers

  • Overview: the signals of the whole file.
  • Columns: the type, the missing values and the issues of every column.
  • Numeric, Strings: ranges, means and medians; the length profile of text columns.
  • Dates: the formats found, and which values are ambiguous or invalid.
  • Detectors: what was recognized, such as email addresses, phone numbers and postal codes.
  • Sample, Settings: the raw values as they are in the file, and every rule used.

Proof

Explore a real report

Tabalyst Report on a real fileGenerated from the example report
report.html report.json insurance-customers.csv Screenshots of a real Tabalyst report

TabalystReport

Explore the report

Report

Analysis

Data

Dataset overview: Size, completeness and the main signal to check first.Open in full report

See the whole-file signal first.

Columns: Inferred and semantic types per column. With issues: missing values or mixed type.Open in full report

Every column, its type and its issues.

Transformations: Occurrences changed by each normalization stage, and the spellings it groups. Raw preview values remain unchanged.Open in full report

What was normalized, column by column.

Numeric analysis: Range and distribution statistics for accepted numeric values.Open in full report

Range, mean and median of each numeric column.

Date analysis: Strict date parsing keeps ambiguous, invalid and non-date values separate. Ambiguous values are never resolved from the other values of the column.Open in full report

Ambiguous dates are flagged, never guessed.

String analysis: Length classes, fixed widths and representative values.Open in full report

The length profile of every text column.

Detectors and formats: What each detector recognized per column, with the formats it found. Primary: the interpretation shown as semantic type.Open in full report

What each detector recognized, and in which format.

Data sample: First 20 records with original row numbers and raw values.Open in full report

Raw values, exactly as they are in the file.

Analysis settings: Rules used for this analysis, so the result can be reproduced.Open in full report

Every rule used, for a reproducible analysis.

Explore the full report

The demo shows screenshots of a report on 3,000 synthetic insurance customers. The full report and its profile are the files Tabalyst Report wrote.

Scope, not hype

What a report does not tell you

  • A report counts what is in the file. A column that is mostly empty can be normal or a problem: that judgment is yours.
  • Tabalyst Report does not clean the file, fill gaps, model the data or test hypotheses.
  • The report and the profile may contain values of the source, even though sensitive fields such as email addresses are masked by default. Share them accordingly.
  • Sampling reads CSV files only.

The known limitations ↗ list the rest.

Next step

Where to go next

Tabalyst Report summarizes and profiles a file. Tabalyst Explore, which will filter and investigate it interactively, is still to come.

Coming soon Tabalyst Explore: guided investigation of your data

Use Tabalyst Report today

Your next dataset

See the data before you model it.

Install Tabalyst, run it on your file and open the report.