Sampling is without replacement
The data rows are shuffled with a Fisher–Yates process, then the first N rows are returned. A source row cannot appear twice in one sample. If N is larger than the available row count, the tool simply returns all rows.
A seed means “draw the same lottery again”
Use the same nonblank seed with the same CSV and conditions, and you can reproduce the same result. That is useful for audits or reviews where someone else needs the exact same set of records.
Leave the seed blank and the tool starts from fresh randomness, so blank does not mean a hidden fixed sample.
Random does not mean balanced by category
If 90% of your rows are group A and 10% are group B, a 10-row sample may contain one B row—but it can also contain none. This is simple random sampling, not stratified sampling.
The header stays; only data rows are shuffled
The first header row is always kept at the top. The data-row order is randomized. If original row order matters for later tracing, add an ID or row-number field before sampling.
A practical review workflow
For a 100,000-row file, start with a fresh 100-row sample to look for obvious problems. If you find something worth sharing and need others to inspect the same records, rerun with a fixed seed. That gives you both exploration and reproducibility.
Want the deeper explanation?