Why matching by name and size finds the wrong files
Name matching fails in both directions. The same photo saved twice becomes IMG_4021.jpg and IMG_4021 (1).jpg — different names, identical bytes — while two entirely unrelated documents can both be called invoice.pdf. Add file size and the false matches drop, but they do not disappear: two different camera JPEGs from the same body frequently land within a few bytes of one another, and a great many small configuration files are exactly the same size by coincidence.
The direction that actually matters for safety is the false positive: a tool that calls two files duplicates when they merely look alike will eventually delete something irreplaceable. This is also why “similar image” matching deserves suspicion for anything automated. It is genuinely useful for triaging a photo library by hand, but a burst of ten nearly identical shots is ten different photographs, not nine duplicates, and no algorithm knows which frame has the eyes open.
The only definition that holds up is byte-for-byte identity. Two files are duplicates if their content is identical, regardless of name, folder, timestamp or extension. That is a question a hash answers, and it answers it the same way every time.
How content hashing works, and why it is fast
A cryptographic hash reduces any file to a fixed-length fingerprint — 64 hexadecimal characters for SHA-256. Identical content always produces an identical fingerprint, and for practical purposes two different files never collide. Comparing fingerprints is therefore a reliable substitute for comparing whole files, and it lets you compare thousands of files against each other instead of pairwise.
Hashing everything would mean reading every byte on the drive, so a competent implementation works in stages. First it groups files by exact size, because files of different sizes cannot be identical, and that step alone eliminates the overwhelming majority of candidates without reading any content. Then it hashes only a small prefix of the files that remain in multi-file groups. Only the survivors of that round get a full SHA-256. On a typical drive the full-read stage ends up touching a small fraction of the data.
That staging explains what a scan should feel like. A first pass over a few hundred thousand files is mostly metadata work and finishes in minutes; the slow part is the final hashing of large video files that genuinely share a size. If a tool takes hours and pins the disk while claiming to compare by content, it is probably hashing everything indiscriminately.
Doing it by hand with PowerShell
You do not need any software for a targeted check. The following one-liner groups files by length, discards unique sizes, hashes the rest with SHA-256, and prints only the groups where more than one file shares a hash. Point it at a specific folder — a whole drive works but will take a while.
- Get-ChildItem -Path 'C:\Users\You\Pictures' -Recurse -File | Group-Object Length | Where-Object { $_.Count -gt 1 } | ForEach-Object { $_.Group } | Get-FileHash -Algorithm SHA256 | Group-Object Hash | Where-Object { $_.Count -gt 1 } | ForEach-Object { $_.Group.Path }
- Add -Filter '*.jpg' to Get-ChildItem to restrict it to one file type
- Add | Out-File C:\Temp\dupes.txt at the end to save the list for review
- Run it in Windows Terminal or PowerShell; no administrator rights are needed for your own folders
- Use Get-FileHash -Algorithm SHA1 instead if you are hashing a very large video library and want it faster
Read the output before deleting anything
The command above deliberately prints paths and deletes nothing. That separation is the whole safety model: generate a list, read it, then act. Sort the output by folder and you will usually see the pattern immediately — a camera import folder duplicated into a dated backup folder, a project exported twice, a music library that exists both in Music and in an old external-drive copy that was pasted onto C:.
When you do delete, delete to the Recycle Bin rather than permanently, and do it in small batches with a pause to use the machine normally in between. The failure mode with duplicate cleanup is never a single wrong file; it is a confident bulk selection across folders whose relationship you had not understood. Keeping the Recycle Bin in the loop turns an expensive mistake into an inconvenience.
Where duplicates actually pile up
Real duplicates concentrate in a handful of predictable places, and scanning only those is both faster and far safer than scanning C: as a whole. Downloads is the champion: the same installer or PDF fetched three times, each with a numbered suffix. Photo folders come second, usually because a phone import ran twice or a card was copied before and after a rename.
- Downloads — repeat downloads with (1), (2) suffixes
- Pictures and phone import folders — an import run more than once
- Cloud folders — OneDrive and Dropbox conflict copies with a device name appended
- Desktop and “old desktop” folders left over from a Windows reinstall
- Messenger media folders — the same image saved by several chats
- Project export folders — renders and builds exported repeatedly
- An old external drive copied wholesale into a folder on C:
What must never be deduplicated
Duplicate content inside software installations is normal and intentional. Many applications ship the same runtime library into several folders on purpose, so each one loads a version it was tested against; deleting one copy because it matches another is how you get an application that starts fine and fails on a specific feature months later. The same logic applies to game asset packs and to development folders where identical dependencies appear in many project subfolders.
Windows itself is the sharpest edge. C:\Windows\WinSxS is full of files that appear duplicated but are hard links pointing at the same data on disk, so a naive scanner reports gigabytes of “duplicates” that free nothing when deleted and break servicing when they do. Restrict any duplicate scan to your own data folders, and treat a tool that offers to scan system folders as a reason to change tools.
- C:\Windows and everything under it, WinSxS above all
- C:\Program Files and C:\Program Files (x86)
- %APPDATA% and %LOCALAPPDATA% — application state and profiles
- C:\ProgramData — licences and application databases
- Game install folders and launcher content directories
- Development folders: node_modules, vendor, virtual environments, build output
- Anything on a drive you are not the only user of
Choosing which copy survives, and the traps
Once a group is confirmed identical, the surviving copy should be the one in the location you would look in first, which is usually the original folder rather than the backup, the Downloads copy or the one on the Desktop. Prefer the path with the shortest, most sensible name. If one copy sits in a synced cloud folder and one does not, keep the synced one only if that folder is genuinely your working location — otherwise you are handing your only copy to a service whose sync you might later reset.
Three traps deserve naming. Hard-linked files show up as separate paths that share the same data, so deleting one frees nothing and the other keeps working — harmless but confusing when the reported savings never materialize. Cloud placeholder files under OneDrive's Files On-Demand may not be present locally at all, and hashing them forces a download of the entire folder. And files inside a folder that something else manages — a backup set, a repository, a game library — are duplicates only from the file system's point of view; the program that owns them expects both.
For a repeatable version of this workflow, Kleaner PRO's duplicate finder compares by SHA-256 rather than by name, excludes system and program folders by default, groups matches with full paths and lets you keep the copy you choose before anything is removed. The PowerShell route above does the same job for a single folder, and is the right tool when you only want to check one place.
Questions and Answers
Is it safe to delete duplicate files automatically?
Not in system or program folders, where identical files are often intentional or hard-linked. In your own documents, photos and downloads it is safe provided the tool matched by content hash and you delete to the Recycle Bin.
Can two different files have the same SHA-256 hash?
Not in any practical sense — no SHA-256 collision has ever been produced. That is precisely why hashing is used for deduplication instead of comparing names, sizes or timestamps.
Why does deleting duplicates free less space than reported?
Usually hard links: several paths pointing at one copy of the data, so removing one path frees nothing. Cloud placeholder files can also be counted at full size while occupying almost nothing locally.
Are similar photos the same as duplicate photos?
No. Similar-image matching compares appearance, so a burst of nearly identical shots looks like duplicates while being ten distinct photographs. Use it to review by hand, never to delete in bulk.
Know what is included before you buy.
The one-time 30-minute trial covers core tools. PRO-labelled features stay locked until a paid license is activated.
Read next
Write to us: [email protected]