Free web tool

CSV deduplicate by key

Remove duplicate CSV rows by selected key columns, keeping the first or last occurrence.

Input & settings

How to use

Use this to reduce repeated customer, product, or order rows based on one or more key columns. You can keep the first or last occurrence, so it works best when you understand both the key definition and the current row order.

  1. Paste the CSV and enter one or more key column names. Separate multiple keys with commas.
  2. Choose whether to keep the first or last row for each key.
  3. Run the tool, review source/kept/removed counts, then download the deduplicated CSV as a new file.

When is this useful?

It is useful when the same email, SKU, or order appears more than once. The hardest part is not the removal algorithm—it is deciding what “the same record” means for your dataset.

The important thing to know

Keys can be composite, such as store_id,sku. A SKU might be duplicated globally but unique within each store. Choosing the key based on business identity changes which rows are considered duplicates.

Common traps

  • “First” and “last” refer to the current CSV row order. The tool does not inspect timestamps to decide which row is newest.
  • Key values are not automatically trimmed or case-normalized.
  • Multiple blank key values are treated as the same key.

Example

If a customer file is already sorted oldest-to-newest and later rows contain the latest data, using email + “keep last” can work. If the file is not reliably sorted, “last” does not mean “latest,” so sort or verify the order first.

How to read the result

“Duplicates removed” is source rows minus kept rows. With “keep last,” the value from the last occurrence is kept, but key output order follows when each key was first seen. If row order has business meaning, review that detail before using the output.