Duplicate File Finder — same bytes means duplicate, whatever the name says
Duplicate File Finder
Drop in a batch of files and this page finds the ones whose content is byte-identical,
in three passes: a size filter (files of different sizes can never be duplicates), a
SHA-256 grouping of the survivors, and a final full byte comparison as insurance.
You get each duplicate cluster with the space it wastes, checkboxes for the copies you
would delete, and a CSV checklist to take to your file manager. Two files named
photo.jpg with different bytes are not duplicates; two files with
different names and the same bytes are. Nothing is uploaded.
1. Pick files
Drop files here, or click to choose them
Up to 50 files, 200 MB each. The files stay on this machine.
No files selected.
2. Scan
3. Duplicate clusters
Clusters of identical files appear here after the scan. The first file of each cluster is pre-marked as the copy to keep.
A browser page cannot delete your files — and should not. The checklist
you export (name, size, SHA-256, cluster id) is meant to be applied in Explorer, Finder or
your file manager of choice, where deletions are visible and undoable.
Privacy: the whole scan runs locally; no file or digest leaves this tab.
What is a duplicate file?
A true duplicate is a file whose every byte equals another file's bytes. Not
“same name” (copy.jpg from two cameras are different photos), not “same
size” (a million files are 1,024 bytes), not “looks the same” (a re-saved
JPEG is a different byte stream). Only content equality counts — which is also why the
scan is immune to renamed copies, re-dated copies and files scattered across different folders.
How the three passes work
Pass
Compares
Cost
Why
1 — size
file size in bytes
instant (metadata)
different size can never be identical; typically eliminates most files with zero reading
2 — SHA-256
full-content digest
one read per survivor
same digest for different content is computationally infeasible
3 — bytes
the actual bytes
one extra read per pair
insurance: proves equality without relying on the hash at all
The size pre-pass is not an optimisation nicety — it is what makes content
comparison affordable. A folder of 5,000 photos may contain 50 clusters, but only a few
hundred files ever reach the hashing stage.
What it deliberately does not do
No near-duplicates: the same photo at two resolutions or two quality levels are
different bytes and are not flagged — that problem needs perceptual image hashing, which
is a different tool. No deletion: a web page deleting files on your disk would be a
security bug, not a feature; you get the checklist instead. No cloud: nothing is
uploaded, which also means nothing can be scanned without your files being read.
FAQ
Why is the first file marked “keep”? — some file in the cluster must
survive; the first by selection order is a deterministic, boring choice. Change any checkbox
before exporting the checklist. What is “wasted space”? — the size of every copy beyond the first in
its cluster, summed over all clusters: the bytes you would get back by deleting the extras. Can SHA-256 be wrong? — not in practice: a collision would need a deliberately
engineered pair and has never occurred for SHA-256. Pass 3 removes even that theoretical doubt.