Tabalyst Report · Open source · CSV, Excel, JSON, JSONL

Your data looks fine. Look closer.

Tabalyst Report turns CSV, Excel, JSON and JSONL files into an interactive HTML report and a structured JSON profile. It reads the whole file locally and shows the missing values, mixed types and date ambiguity before they enter your import, analysis or pipeline.

pip install tabalyst
  • Local analysis
  • MPL-2.0 licensed
  • Python 3.11+
Open the full live report
report.html report.json insurance-customers.csv Screenshots of a real Tabalyst report

TabalystReport

Explore the report

Report

Analysis

Data

Dataset overview: Size, completeness and the main signal to check first.Open in full report

See the whole-file signal first.

Columns: Inferred and semantic types per column. With issues: missing values or mixed type.Open in full report

Every column, its type and its issues.

Transformations: Occurrences changed by each normalization stage, and the spellings it groups. Raw preview values remain unchanged.Open in full report

What was normalized, column by column.

Numeric analysis: Range and distribution statistics for accepted numeric values.Open in full report

Range, mean and median of each numeric column.

Date analysis: Strict date parsing keeps ambiguous, invalid and non-date values separate. Ambiguous values are never resolved from the other values of the column.Open in full report

Ambiguous dates are flagged, never guessed.

String analysis: Length classes, fixed widths and representative values.Open in full report

The length profile of every text column.

Detectors and formats: What each detector recognized per column, with the formats it found. Primary: the interpretation shown as semantic type.Open in full report

What each detector recognized, and in which format.

Data sample: First 20 records with original row numbers and raw values.Open in full report

Raw values, exactly as they are in the file.

Analysis settings: Rules used for this analysis, so the result can be reproduced.Open in full report

Every rule used, for a reproducible analysis.

Explore the full report

The first pass matters

“It opens” is not the same as “it is ready.”

Before an import, an analysis or a new pipeline, you need to know what the file actually contains. A manual skim rarely tells the whole story. The example below comes from a CSV; Excel, JSON and JSONL reports flag the same kinds of signals.

IN THE EXAMPLE / date_naissance
356

ambiguous dates.

They appear inside a column with 2,635 valid dates. A glance at the first few rows would miss the distinction.

Valid 2,635Ambiguous 356Invalid 3

What a skim misses

ILLUSTRATIVE ROWS
rowid_clientdate_naissanceprime_annuelleAnnotation
1CLI-104211961-03-311240.00
2CLI-1042204/05/1983980.50day or month first?
3CLI-104231105.00missing
4CLI-104241978-02-301310.00February 30th?
5CLI-104231105.00same as row 3
Flagged by a full-file profileMissing · ambiguous · invalid · duplicate, reported separately
01 / STRUCTURE

Types can be uncertain.

A column that looks like dates may mix formats, ambiguous values and impossible ones.

02 / QUALITY

Small gaps have consequences.

Missing values and duplicate rows are easier to handle before a file enters your workflow.

03 / PROCESS

A first look should be repeatable.

Replace a one-off spreadsheet scan with a documented analysis your tools can reuse.

One clear chain · three steps

From raw file to a view you can use.

Analysis and presentation are separate. One analysis produces a JSON profile for code and a self-contained HTML report for people.

SOURCE FILE

Your file

A CSV, Excel, JSON or JSONL file is read locally. Original values remain visible in the bounded preview.

ANALYSIS ENGINE

Tabalyst Report

Types, missing values, duplicates, dates and patterns are profiled with explicit settings, whatever the format.

TWO DELIVERABLES

Human + machine

A structured profile for your tooling. An interactive report for exploration and sharing.

report.json is the canonical result. executions.json records successful analyses.
Tabalyst Report generates the structured JSON profile and the self-contained HTML report together.

Schematic rows, for illustration only. The example is a CSV file.

Run it now

From CSV to your first Tabalyst Report in two commands.

Install or update the package, point Tabalyst Report at a file and open the generated report. One file or hundreds, the command is the same. Excel, JSON and JSONL work the same way: see Excel, JSON and JSONL. Use the configuration guide ↗ for separators, encodings and settings.

Beta, interfaces may still change · Python 3.11+ · Available on PyPI ↗

The same analysis, available from Python.

Use the public API when your file assessment belongs inside a script or pipeline. The return value is JSON-serializable, and the HTML report remains available for people to inspect.

SCHEMA

The JSON profile is versioned but still young. Check the documented schema version before depending on fields across upgrades.

Read the documentation ↗
  1. 1 · Install or update
    pip install --upgrade tabalyst

    # Also updates it: Tabalyst changes often

  2. 2 · One file
    tabalyst report data.csv

    # data.html, data.json and executions.json, beside data.csv

  3. 3 · One report per file
    tabalyst report *.csv -d reports/

    # One HTML report and JSON profile per file, all in reports/. Files are not merged.

Before the next step

Better inputs for the work you actually do.

Tabalyst Report gives you a first assessment of the file. You decide what to fix, reject or pass downstream.

01 / IMPORT

Before an ingestion job.

Inspect column types and missing values before mapping a partner export into your system.

02 / ANALYSIS

Before a notebook.

See the shape of an unfamiliar dataset before your first transformation or chart.

03 / AI WORKFLOW

Before a model sees it.

Review the file yourself and let your own tools consume the structured JSON profile where useful.

Scope, not hype. Tabalyst Report is a deterministic profiler, not an LLM. It does not generate AI recommendations or clean the file for you.

Four formats

Works with the data you already have.

The same two outputs for CSV, Excel, JSON and JSONL. Each format has its own page with a real report, its own command and its limits.

.csv

CSV

Profile an export column by column: types, gaps, ambiguous dates, duplicate rows.

.xlsx, .xlsm

Excel

Report one sheet or named table of a workbook, header found even under a title.

.json

JSON

Report a collection of records, with nested fields named by their path.

.jsonl, .ndjson

JSONL

Report one record per line, with invalid lines excluded, counted and listed.

Local-first · Built for big files

Stays on your machine, reads files as a stream.

Tabalyst Report analyzes the file where it is, with no upload and no account. It adds no telemetry or remote processing.

01 / LOCAL

No upload required

The Python package runs the analysis on your machine. The generated report may contain values from the source, so review it before sharing.

02 / STREAMING

Bounded memory

Reports read CSV, JSON and JSONL files as a stream, with memory bounded by configurable limits. An Excel workbook is read sheet by sheet into memory.

03 / HONEST

No benchmark claimed

Large files are a design goal, not a measured promise: no speed or maximum size is published. See the known limitations ↗.

More than reports

One toolkit, one engine.

Tabalyst Report is the first tool of the Tabalyst toolkit. The others share its engine and its command grammar. A tool is linked here once its page is published.

TABALYST FOR

Profile and inspect data from your CLI or Python workflows.

TABALYST FOR

Understand your sources before ETL, migration or integration.

TABALYST FOR

Make data problems visible before their downstream effects.

Open by design

Inspect the method. Keep control of the file.

The code is MPL-2.0-licensed, and the configuration and execution history explain how each result was produced. Install it from PyPI, read the source on GitHub, and follow the documentation.

01 / FORMATS

CSV, Excel, JSON and JSONL

One engine reads the four formats, and the profile records which one was read.

02 / OUTPUT

Readable output, structured result

An HTML report for people and a JSON profile for code.

03 / OPEN

Open source and reproducible

MPL-2.0-licensed code, effective settings and a history of successful runs.

Good to know

Questions before you start.

Something missing? Open an issue ↗

Which formats does Tabalyst Report read?

CSV, Excel (.xlsx and .xlsm), JSON and JSONL (.jsonl and .ndjson). Older .xls files are not read: save them as .xlsx.

Does Tabalyst Report upload my file?

No. The current Python package analyzes your file on your machine. The generated report may contain values from the source, so review it before sharing.

Do I need to write code?

You need Python and a terminal to run the current version. A Python API is also available. A hosted option is planned, but it is not available yet.

Can it report several files at once?

Yes. tabalyst report *.csv -d reports/ writes one HTML report and one JSON profile per file. The files are not merged, joined or compared.

Does it fix my file?

No. Tabalyst Report profiles a file and surfaces potential issues. It does not silently change your source or validate your domain-specific rules.

Can I use the report offline?

Yes. The analysis runs locally and the HTML report is a single self-contained file: it opens without an internet connection.

Can it handle very large files?

Reports read CSV, JSON and JSONL files as a stream, with memory bounded by configurable limits. An Excel workbook is read sheet by sheet into memory, and a sheet above 1 GiB of XML is refused. No speed or maximum file size is published.

Can I use the JSON in a pipeline?

Yes. The structured profile is designed for reuse in your own tools. Check the documented schema version when upgrading.

Your next file

Start with a clearer picture.

Install Tabalyst and generate your own report with Tabalyst Report, or open a real report before you begin.

pip install tabalyst
Open the live report ↗GitHub ↗