Tabalyst ToolsAvailable

A smaller file, same shape.

Tabalyst Sample creates a smaller CSV from a large one: the first rows, the last rows, a random draw or a draw that keeps the distribution of a field. The header, the column order and the raw cell values are kept, and the source is never modified.

ReproducibleA random draw with a seed
tabalyst sample customers.csv --sample-method random --rows 1000 --seed 42
Sample created successfully.
Method:        random
Source rows:   3,000
Sample rows:   1,000
Sample rate:   33.33%
Seed:          42
Output:        customers.sample.csv
Nothing overwrittenA second run stops
tabalyst sample customers.csv --sample-method random --rows 1000 --seed 42
Error: Sample output already exists. Use --force to replace it:
  customers.sample.csv

Real output of tabalyst 0.6.1, paths shortened.

The tool

Four methods, one rule: the source stays intact

  • 01 / METHODS

    first, last, random, stratified

    Keep the first or last rows, draw uniformly from the whole file, or draw so that the distribution of a chosen field is approximately preserved. Empty values form their own stratum.

  • 02 / SIZE

    A number of rows or a percentage

    Choose exactly one of --rows and --percent. Percentages round up to the next whole row, and a 100% sample holds every data row.

  • 03 / SAFE

    Reproducible, never destructive

    Add --seed and another run draws the same rows. An existing output is not replaced without --force, and an input file can never be overwritten.

Sample reads the CSV as a stream. Random and stratified draws keep only the requested sample, plus the counts of each stratum, in memory.

Get started

Draw a sample

  1. 1 · Install
    pip install tabalyst

    # Also updates it: Tabalyst changes often

  2. 2 · A random draw
    tabalyst sample customers.csv --sample-method random --rows 1000 --seed 42

    # Writes customers.sample.csv beside the source

  3. 3 · Keep a distribution
    tabalyst sample customers.csv --sample-method stratified --field province_code --rows 500 --seed 42

    # Approximately preserves the share of each province_code

  4. 4 · Several files
    tabalyst sample *.csv --sample-method random --percent 5 -d samples/

    # One sample per file, in samples/. -o names the output of a single file

Real output

What it prints

The sample is an ordinary CSV file. It has the same header and the same columns as the source, so it opens in any tool, and you can give it to Tabalyst Report to get a quick first look at a very large file.

On the 3,000 customers of the example, a stratified draw of 500 rows on province_code writes a file of 501 lines: the header and the 500 sampled rows.

Scope, not hype

What to know before you rely on it

  • Sample reads CSV files only. For Excel, JSON and JSONL, use Tabalyst Inspect and Tabalyst Report.
  • A stratified draw needs an existing field: an unknown name stops with an error. It approximates the distribution, it does not reproduce it exactly.
  • -o accepts a single input, and cannot be combined with -d.
  • A sample is a subset of the file, not a summary of it: to measure the whole file, run Tabalyst Report on the source.

The known limitations ↗ list the rest.

Next step

Where to go next