Tabalyst Sample
Methods, sizes, output locations and the Python API.
Tabalyst ToolsAvailable
Tabalyst Sample creates a smaller CSV from a large one: the first rows, the last rows, a random draw or a draw that keeps the distribution of a field. The header, the column order and the raw cell values are kept, and the source is never modified.
tabalyst sample customers.csv --sample-method random --rows 1000 --seed 42
Sample created successfully.
Method: random
Source rows: 3,000
Sample rows: 1,000
Sample rate: 33.33%
Seed: 42
Output: customers.sample.csv
tabalyst sample customers.csv --sample-method random --rows 1000 --seed 42
Error: Sample output already exists. Use --force to replace it:
customers.sample.csv
Real output of tabalyst 0.6.1, paths shortened.
The tool
Keep the first or last rows, draw uniformly from the whole file, or draw so that the distribution of a chosen field is approximately preserved. Empty values form their own stratum.
Choose exactly one of --rows and --percent. Percentages round up to the next whole row, and a 100% sample holds every data row.
Add --seed and another run draws the same rows. An existing output is not replaced without --force, and an input file can never be overwritten.
Sample reads the CSV as a stream. Random and stratified draws keep only the requested sample, plus the counts of each stratum, in memory.
Get started
pip install tabalyst# Also updates it: Tabalyst changes often
tabalyst sample customers.csv --sample-method random --rows 1000 --seed 42# Writes customers.sample.csv beside the source
tabalyst sample customers.csv --sample-method stratified --field province_code --rows 500 --seed 42# Approximately preserves the share of each province_code
tabalyst sample *.csv --sample-method random --percent 5 -d samples/# One sample per file, in samples/. -o names the output of a single file
import tabalyst result = tabalyst.sample_csv( "customers.csv", method="stratified", field="province_code", rows=500, seed=42, output="customers.sample.csv",)print(result.sample_rows)
Real output
The sample is an ordinary CSV file. It has the same header and the same columns as the source, so it opens in any tool, and you can give it to Tabalyst Report to get a quick first look at a very large file.
On the 3,000 customers of the example, a stratified draw of 500 rows on province_code writes a file of 501 lines: the header and the 500 sampled rows.
Scope, not hype
-o accepts a single input, and cannot be combined with -d.The known limitations ↗ list the rest.
Next step
Methods, sizes, output locations and the Python API.
Every command, option and exit code.