AI for Work
Build a Small CSV Duplicate Checker With Claude Code Without Touching the Original
Cihan's view: TRY Claude Code for a small, reviewable file utility; SKIP running generated scripts on production data before inspecting the diff and tests.
Evidence label: documentation-based builder lesson with illustrative code. The script below was not executed here and does not claim a reproduced Claude Code demo.
A duplicate-row problem is small enough to teach the useful parts of coding with AI: define identity, preserve the original, request a dry run, add a test, then review the change. Claude Code is documented as a coding assistant that can edit files, run commands and work across a codebase. Those powers are exactly why the boundary matters.
Prerequisites
You need Claude Code on a terminal, IDE, desktop or web surface, plus a local test folder and a supported Claude subscription or Anthropic Console account. The official overview lists those surfaces and access requirements. You need basic comfort opening a terminal.
Create a copy of the sample below. Never begin with a customer export or a folder containing secrets.
email,company,amount
alee@example.test,Northwind,1200
pat@example.test,Contoso,800
alee@example.test,Northwind,1200
mina@example.test,Fabrikam,450
pat@example.test,Contoso,800
Here, email + company + amount defines an exact duplicate. That identity rule is a business assumption, not something the model should guess.
The bounded task
-
Make a folder containing
sample.csv. Ask Claude Code to inspect the folder only. Do not grant access to unrelated directories. -
Give it this instruction:
Create a small Python utility named dedupe_csv.py for sample.csv.
Requirements:
- Never overwrite sample.csv.
- Read the input and write sample.deduped.csv.
- Treat email, company, and amount together as the row identity.
- Preserve the first occurrence and report how many exact duplicates were removed.
- Add a --dry-run option that prints the proposed output path and counts but writes no file.
- Reject a missing required column with a clear error.
- Add tests for two duplicates, no duplicates, a missing column, and a quoted comma in a company name.
- Explain the code before running it. Do not execute it until I approve the plan.
- Review before execution. Check that the code does not write to the input path, that the identity columns are explicit, and that the tests assert the original file remains unchanged. Approve a dry run first. Then compare the output row count with a hand calculation: the sample has five rows and two exact duplicate rows, so the proposed output should contain three rows. This is an acceptance check for the fictional sample, not a measured product result.
No-code alternative
If terminal setup is not appropriate, use a spreadsheet’s duplicate-removal or conditional-formatting feature on a copy. Select the three identity columns, flag repeated combinations, and review the flagged rows before deleting anything. The same rule still applies: preserve the source and verify the count manually.
Troubleshooting
- The tool chooses email alone: restate the composite identity and add a test where the same email has a different company or amount.
- The output overwrites the source: stop, restore the original from the copy, and require a separate output path.
- Quoted commas break the file: add the quoted-company test and use a CSV parser rather than splitting lines manually.
TRY / SKIP
TRY this for a small, non-sensitive CSV where the generated diff and tests can be reviewed. SKIP it for destructive cleanup, regulated records, or large production exports until a qualified owner has tested the utility on a safe copy.
Claude Code can connect to external tools through MCP, but this lesson does not need an integration. If you later add one, verify the server, scopes and first operation; read access is not permission for unattended writes. Claude Code overview · Claude Code MCP reference.