Exploring Dedupe with Python/AI
An exploration of Dedupe. A series of dedupe analysis with Python/AI
Exploring Dedupe with Python/AI
An exploration of Dedupe.
I have this data I have been collecting for a while. It is the NWS Marine forecast for the Potomac river near the Airport. It is a very structured text and I believe there is a high chance that it can be deduplicated for storage. Currently I store six of these a day in 1K or less text files.
ANZ535-220800
Tidal Potomac from Key Bridge to Indian Head
734 PM EDT Thu Aug 21 2025
SMALL CRAFT ADVISORY IN EFFECT UNTIL 8 PM EDT THIS EVENING
TONIGHT
N winds 5 to 10 kt. Waves 1 ft.
FRI
NE winds 5 kt. Waves less than 1 ft.
FRI NIGHT
SE winds around 5 kt. Waves 1 ft.
SAT
S winds 5 to 10 kt. Waves 1 ft.
SAT NIGHT
S winds 5 to 10 kt. Waves 1 ft.
SUN
S winds 5 kt. Waves less than 1 ft. A chance of showers.
SUN NIGHT
S winds around 5 kt. Waves 1 ft. A chance of showers.
MON
NW winds 5 to 10 kt. Waves 1 ft.
TUE
NW winds 5 to 10 kt. Waves 1 ft.
There are at least 4 types of dedupe:
- Word-level — easy to understand. High pointer overhead.
- Sentence or Phrase level — good for repetitive text, works on structured text
- Fixed-block chunking— best for larger datasets, that don’t have much change
- Variable-block chunking — hard to implement.

AI generated image of deduplication in Renaissance Italy.
So, lets do some analysis of our data to see what is the best fit.
you are an expert on deduplication methods and expert at python cli coding.
create a cli program that takes a glob of files such as *.txt.
Tabulates the duplicate words.
counts the bytes of each word as well.
Goal is to save space by using dedupe techiques
Final report of the program is to give a summary of top 50 unique words.
Potential dedupe factor and reduction in bytes for the data analyzed.
I will put the source code at the end and provide a github link
Here is the output
--- Deduplication Analysis Report ---
Number of Processed Files: 7551
-----------------------------------
--- Overall Statistics ---
Total Words Analyzed: 1,107,614
Unique Words Found: 924
-----------------------------------
--- Deduplication Potential ---
Original Size (sum of all word bytes): 3,645.96 KB
Deduplicated Size (unique words + pointers): 4,330.48 KB
Potential Bytes Saved: -684.52 KB
Potential Deduplication Factor: 0.84x
Potential Space Reduction: -18.77%
-----------------------------------
--- Top 50 Most Frequent Words ---
Rank Word Frequency Bytes/Word
---- ------------------- -------------- -----------
1 kt 78,468 2
2 1 66,593 1
3 winds 66,216 5
4 waves 66,102 5
5 ft 62,819 2
6 to 59,357 2
7 5 52,674 1
8 10 34,878 2
9 night 20,225 5
10 of 19,851 2
11 showers 19,166 7
12 around 16,504 6
13 a 16,463 1
14 chance 16,463 6
15 less 15,779 4
16 and 15,216 3
17 nw 13,385 2
18 with 12,813 4
19 gusts 12,796 5
20 s 12,061 1
21 tstms 11,772 5
22 in 11,526 2
23 than 10,310 4
24 sw 9,847 2
25 20 9,798 2
26 15 9,609 2
27 from 9,395 4
28 w 9,324 1
29 wed 8,283 3
30 tue 8,280 3
31 thu 8,270 3
32 fri 8,262 3
33 sat 8,244 3
34 mon 8,243 3
35 sun 8,236 3
36 the 7,885 3
37 anz535 7,551 6
38 tidal 7,551 5
39 potomac 7,551 7
40 key 7,551 3
41 bridge 7,551 6
42 indian 7,551 6
43 head 7,551 4
44 tonight 7,488 7
45 this 6,739 4
46 afternoon 6,438 9
47 n 6,266 1
48 edt 6,258 3
49 e 5,941 1
50 rain 5,769 4
-----------------------------------
The program assumes a 4 byte pointer — so the overhead of a word based method doesn’t work.
So, lets do a sentence based analysis
Prompt:
you are an expert on deduplication methods and expert at python cli coding.
create a cli program that takes a glob of files such as *.txt.
Tabulates the duplicate sentences.
counts the bytes of each sentence as well. a new line can be a sentence
Goal is to save space by using dedupe techiques.
Final report of the program is to give a summary of top 50 unique
sentences or phrases
Potential dedupe factor and reduction in bytes for the data analyzed.
in the report print number of files processed, but not names of files.
print bytes processed and other significant observations of the data
add a --limit parm where you specify all as 0 or as many top as you want.
Here is the run — we would achieve almost 8 to 1 compression assuming a 4 byte reference pointer.
--- Deduplication Analysis Report ---
Files Processed: 7551
Total Bytes Processed: 4,844,327 bytes
Total Sentences/Lines: 181,407
Unique Sentences/Lines: 15,205
--- Potential Savings ---
Potential Reduction: 4,239,096 bytes
Deduplication Factor: 8.00x
(This means the original data is ~8.00 times larger than the unique data)
--- Top 50 Unique Sentences/Phrases by Frequency ---
Rank | Count | Bytes (Single) | Bytes (Total) | Sentence/Phrase
--------------------------------------------------------------------------------
1 | 6,809 | 44 | 299,596 | "Tidal Potomac from Key Bridge to Indian Head"
2 | 6,050 | 7 | 42,350 | "TONIGHT"
3 | 4,682 | 3 | 14,046 | "THU"
4 | 4,679 | 3 | 14,037 | "FRI"
5 | 4,675 | 3 | 14,025 | "WED"
6 | 4,670 | 3 | 14,010 | "TUE"
7 | 4,668 | 3 | 14,004 | "SAT"
8 | 4,656 | 3 | 13,968 | "SUN"
9 | 4,653 | 3 | 13,959 | "MON"
10 | 2,934 | 64 | 187,776 | "Winds and waves higher and visibilities lower in and near tstms."
11 | 2,689 | 32 | 86,048 | "NW winds 5 to 10 kt. Waves 1 ft."
12 | 2,526 | 9 | 22,734 | "WED NIGHT"
13 | 2,525 | 9 | 22,725 | "TUE NIGHT"
14 | 2,519 | 9 | 22,671 | "THU NIGHT"
15 | 2,506 | 9 | 22,554 | "FRI NIGHT"
16 | 2,503 | 9 | 22,527 | "SAT NIGHT"
17 | 2,503 | 9 | 22,527 | "SUN NIGHT"
18 | 2,502 | 9 | 22,518 | "MON NIGHT"
19 | 1,869 | 5 | 9,345 | "TODAY"
20 | 1,797 | 32 | 57,504 | "S winds around 5 kt. Waves 1 ft."
21 | 1,705 | 31 | 52,855 | "S winds 5 to 10 kt. Waves 1 ft."
22 | 1,636 | 6 | 9,816 | "tstms."
23 | 1,498 | 14 | 20,972 | "THIS AFTERNOON"
24 | 1,472 | 35 | 51,520 | "S winds 5 kt. Waves less than 1 ft."
25 | 1,457 | 31 | 45,167 | "N winds 5 to 10 kt. Waves 1 ft."
26 | 1,393 | 32 | 44,576 | "SW winds 5 to 10 kt. Waves 1 ft."
27 | 1,387 | 31 | 42,997 | "W winds 5 to 10 kt. Waves 1 ft."
28 | 1,303 | 13 | 16,939 | "REST OF TODAY"
29 | 1,294 | 33 | 42,702 | "SW winds around 5 kt. Waves 1 ft."
30 | 1,087 | 33 | 35,871 | "NW winds around 5 kt. Waves 1 ft."
31 | 1,055 | 11 | 11,605 | "Waves 1 ft."
32 | 1,039 | 32 | 33,248 | "E winds around 5 kt. Waves 1 ft."
33 | 977 | 36 | 35,172 | "NW winds 5 kt. Waves less than 1 ft."
34 | 976 | 32 | 31,232 | "W winds around 5 kt. Waves 1 ft."
35 | 953 | 15 | 14,295 | "REST OF TONIGHT"
36 | 916 | 33 | 30,228 | "SE winds around 5 kt. Waves 1 ft."
37 | 898 | 35 | 31,430 | "N winds 5 kt. Waves less than 1 ft."
38 | 897 | 36 | 32,292 | "SW winds 5 kt. Waves less than 1 ft."
39 | 832 | 32 | 26,624 | "NE winds 5 to 10 kt. Waves 1 ft."
40 | 778 | 33 | 25,674 | "NE winds around 5 kt. Waves 1 ft."
41 | 750 | 32 | 24,000 | "N winds around 5 kt. Waves 1 ft."
42 | 742 | 45 | 33,390 | "Tidal Potomac from Key Bridge to Indian Head-"
43 | 733 | 35 | 25,655 | "W winds 5 kt. Waves less than 1 ft."
44 | 701 | 36 | 25,236 | "NE winds 5 kt. Waves less than 1 ft."
45 | 661 | 10 | 6,610 | "and tstms."
46 | 614 | 35 | 21,490 | "E winds 5 kt. Waves less than 1 ft."
47 | 593 | 9 | 5,337 | "OVERNIGHT"
48 | 577 | 31 | 17,887 | "E winds 5 to 10 kt. Waves 1 ft."
49 | 571 | 52 | 29,692 | "NW winds 5 to 10 kt with gusts to 20 kt. Waves 1 ft."
50 | 549 | 16 | 8,784 | "chance of tstms."
What about chunking?
Conversation with Gemini
you are an expert on deduplication methods and expert at python cli coding.
Create a cli program that takes a glob of files such as *.txt. Take input into fixed sized chunks specified by --chunk. size is in bytes
Goal is to save space by using dedupe techiques.
you are an expert on deduplication methods and expert at python cli coding.
create a cli program that takes a glob of files such as *.txt.
Tabulates the duplicate sentences.
counts the bytes of each sentence as well. a new line can be a sentence
Goal is to save space by using dedupe techiques.
Final report of the program is to give a summary of top 50 unique fixed size chunks
have a --limit parm where you specify 0 for all chunks, or number, but default to 50.
I did several runs
4096 — no duplicates at all
1024 — not worth it
256/128 — some, but not enough to be worth doing
[*] Starting analysis with chunk size: 1024 bytes
[*] Searching for files with pattern: /home/jon2allen/python/data/jon2allen/ANZ/data/*.txt
[*] Found 7551 files to process.
--- Deduplication Report ---
Summary:
- Total Files Analyzed: 7551
- Total Original Size: 4992.68 KB (5112501 bytes)
- Chunk Size: 1024 bytes
- Total Chunks Processed: 90
- Unique Chunks Found: 78
- Size After Dedupe: 78.00 KB (79872 bytes)
- Duplicate Chunks Found: 10
- Total Potential Savings: 12.00 KB (12288 bytes)
- Deduplication Ratio: 64.01:1
Top 10 Duplicate Chunks by Potential Savings:
----------------------------------------------------------------------
Rank | Chunk Hash (first 16 chars) | Count | Savings (Bytes)
----------------------------------------------------------------------
1 | 414e5a3533352d31... | 4 | 3072
2 | 414e5a3533352d30... | 2 | 1024
3 | 414e5a3533352d32... | 2 | 1024
4 | 414e5a3533352d30... | 2 | 1024
5 | 414e5a3533352d30... | 2 | 1024
6 | 414e5a3533352d30... | 2 | 1024
7 | 414e5a3533352d32... | 2 | 1024
8 | 414e5a3533352d30... | 2 | 1024
9 | 414e5a3533352d30... | 2 | 1024
10 | 414e5a3533352d32... | 2 | 1024
----------------------------------------------------------------------
[jon2allen@jons-bad-ass-fedora-server-37 dedupe]$ ./dedupe_analyzer_chunk.py --chunk 2048 --limit 50 '/home/jon2allen/python/data/jon2allen/ANZ/data/*.txt'
[*] Starting analysis with chunk size: 2048 bytes
[*] Searching for files with pattern: /home/jon2allen/python/data/jon2allen/ANZ/data/*.txt
[*] Found 7551 files to process.
--- Analysis Complete ---
No full-sized chunks found to analyze. Try a smaller chunk size.
[jon2allen@jons-bad-ass-fedora-server-37 dedupe]$ ./dedupe_analyzer_chunk.py --chunk 256 --limit 50 '/home/jon2allen/python/data/jon2allen/ANZ/data/*.txt'
[*] Starting analysis with chunk size: 256 bytes
[*] Searching for files with pattern: /home/jon2allen/python/data/jon2allen/ANZ/data/*.txt
[*] Found 7551 files to process.
--- Deduplication Report ---
Summary:
- Total Files Analyzed: 7551
- Total Original Size: 4992.68 KB (5112501 bytes)
- Chunk Size: 256 bytes
- Total Chunks Processed: 16157
- Unique Chunks Found: 13663
- Size After Dedupe: 3415.75 KB (3497728 bytes)
- Duplicate Chunks Found: 2418
- Total Potential Savings: 623.50 KB (638464 bytes)
- Deduplication Ratio: 1.46:1
Top 50 Duplicate Chunks by Potential Savings:
----------------------------------------------------------------------
Rank | Chunk Hash (first 16 chars) | Count | Savings (Bytes)
----------------------------------------------------------------------
1 | 7468616e20312066... | 5 | 1024
2 | 414e5a3533352d32... | 4 | 768
3 | 414e5a3533352d30... | 4 | 768
4 | 414e5a3533352d31... | 4 | 768
5 | 6573732e0a0a4652... | 4 | 768
6 | 0a0a5341540a5720... | 4 | 768
7 | 677573747320746f... | 4 | 768
8 | 414e5a3533352d30... | 4 | 768
9 | 30206b742e205761... | 4 | 768
10 | 414e5a3533352d30... | 4 | 768
11 | 414e5a3533352d31... | 4 | 768
12 | 30206b742e2e2e20... | 4 | 768
13 | 206368616e636520... | 4 | 768
14 | 414e5a3533352d30... | 4 | 768
15 | 414e5a3533352d32... | 4 | 768
16 | 206b742e2e2e2062... | 4 | 768
17 | 414e5a3533352d31... | 3 | 512
18 | 2077696e64732035... | 3 | 512
19 | 726f75732073686f... | 3 | 512
20 | 55204e494748540a... | 3 | 512
21 | 7769746820610a63... | 3 | 512
22 | 414e5a3533352d31... | 3 | 512
23 | 6d732e0a0a465249... | 3 | 512
24 | 6f203130206b742e... | 3 | 512
25 | 490a452077696e64... | 3 | 512
26 | 4e572077696e6473... | 3 | 512
27 | 414e5a3533352d32... | 3 | 512
28 | 66742e0a0a534154... | 3 | 512
29 | 696e647320352074... | 3 | 512
30 | 414e5a3533352d31... | 3 | 512
31 | 7468206775737473... | 3 | 512
32 | 494748540a572077... | 3 | 512
33 | 6b742e2057617665... | 3 | 512
34 | 4652490a45207769... | 3 | 512
35 | 414e5a3533352d31... | 3 | 512
36 | 20616e640a747374... | 3 | 512
37 | 35206b742e205761... | 3 | 512
38 | 414e5a3533352d30... | 3 | 512
39 | 414e5a3533352d30... | 3 | 512
40 | 4748540a53207769... | 3 | 512
41 | 414e5a3533352d31... | 3 | 512
42 | 203520746f203130... | 3 | 512
43 | 2e20576176657320... | 3 | 512
44 | 414e5a3533352d32... | 3 | 512
45 | 2077697468206775... | 3 | 512
46 | 206b742e2e2e2064... | 3 | 512
47 | 657320312066742e... | 3 | 512
48 | 3020746f20313520... | 3 | 512
49 | 742e204120636861... | 3 | 512
50 | 52490a5357207769... | 3 | 512
----------------------------------------------------------------------
[jon2allen@jons-bad-ass-fedora-server-37 dedupe]$ ./dedupe_analyzer_chunk.py --chunk 128 --limit 50 '/home/jon2allen/python/data/jon2allen/ANZ/data/*.txt'
[*] Starting analysis with chunk size: 128 bytes
[*] Searching for files with pattern: /home/jon2allen/python/data/jon2allen/ANZ/data/*.txt
[*] Found 7551 files to process.
--- Deduplication Report ---
Summary:
- Total Files Analyzed: 7551
- Total Original Size: 4992.68 KB (5112501 bytes)
- Chunk Size: 128 bytes
- Total Chunks Processed: 36183
- Unique Chunks Found: 30285
- Size After Dedupe: 3785.62 KB (3876480 bytes)
- Duplicate Chunks Found: 5614
- Total Potential Savings: 737.25 KB (754944 bytes)
- Deduplication Ratio: 1.32:1
Top 50 Duplicate Chunks by Potential Savings:
----------------------------------------------------------------------
Rank | Chunk Hash (first 16 chars) | Count | Savings (Bytes)
----------------------------------------------------------------------
1 | 7468616e20312066... | 5 | 512
2 | 646e696768742e0a... | 5 | 512
3 | 0a4652490a572077... | 5 | 512
4 | 2045445420544849... | 5 | 512
5 | 414e5a3533352d32... | 4 | 384
6 | 4f4e494748540a4e... | 4 | 384
7 | 2077696e64732035... | 4 | 384
8 | 53204556454e494e... | 4 | 384
9 | 414e5a3533352d30... | 4 | 384
10 | 6c65737320746861... | 4 | 384
11 | 6473203520746f20... | 4 | 384
12 | 2045445420544849... | 4 | 384
13 | 203235206b742e20... | 4 | 384
14 | 414e5a3533352d31... | 4 | 384
15 | 465445524e4f4f4e... | 4 | 384
16 | 6573732e0a0a4652... | 4 | 384
17 | 2057617665730a32... | 4 | 384
18 | 0a0a5341540a5720... | 4 | 384
19 | 35206b7420616674... | 4 | 384
20 | 677573747320746f... | 4 | 384
21 | 6265636f6d696e67... | 4 | 384
22 | 0a0a544f4e494748... | 4 | 384
23 | 414e5a3533352d30... | 4 | 384
24 | 616e20312066742e... | 4 | 384
25 | 30206b742e205761... | 4 | 384
26 | 2057617665732031... | 4 | 384
27 | 6c79207769746820... | 4 | 384
28 | 642035206b742e20... | 4 | 384
29 | 414e5a3533352d30... | 4 | 384
30 | 66742e0a0a544f4e... | 4 | 384
31 | 414e5a3533352d31... | 4 | 384
32 | 4944415920414654... | 4 | 384
33 | 30206b742e2e2e20... | 4 | 384
34 | 65726e6f6f6e2077... | 4 | 384
35 | 206368616e636520... | 4 | 384
36 | 572077696e647320... | 4 | 384
37 | 73203520746f2031... | 4 | 384
38 | 4553542054484953... | 4 | 384
39 | 414e5a3533352d30... | 4 | 384
40 | 4544542054484953... | 4 | 384
41 | 0a4e572077696e64... | 4 | 384
42 | 2045535420544849... | 4 | 384
43 | 414e5a3533352d32... | 4 | 384
44 | 4953204556454e49... | 4 | 384
45 | 206b742e2e2e2062... | 4 | 384
46 | 68696e6720746f20... | 4 | 384
47 | 6176657320312066... | 4 | 384
48 | 657373207468616e... | 3 | 256
49 | 20746f203130206b... | 3 | 256
50 | 6c6573732e205061... | 3 | 256
----------------------------------------------------------------------
For this NWS marine forecast data— it looks like the sentence/phrase is the best. The NWS marine forecast is somewhat structured. Here is one of the scripts.
[embed]
Here is the github repository
git clone https://github.com/jon2allen/dedupe.git
cd dedupe
./run.sh
I used 3 separate prompts for each script. So, the report format is not consistent. A lesson learned is maybe specify more about the report format. LLM may not do the same thing twice.
I will revisit and implement the sentence dedupe and see how it can work in practice.
A message from our Founder
Hey, Sunil here. I wanted to take a moment to thank you for reading until the end and for being a part of this community.
Did you know that our team run these publications as a volunteer effort to over 3.5m monthly readers? We don’t receive any funding, we do this to support the community. ❤️
If you want to show some love, please take a moment to follow me on LinkedIn, TikTok, **Instagram. You can also subscribe to our [weekly newsletter](https://newsletter.plainenglish.io/)**.
And before you go, don’t forget to clap and follow the writer️!
메타데이터
- post_id
- 494771e39ea3
- slug
- exploring-dedupe-with-python-ai-494771e39ea3
- url
- https://python.plainenglish.io/exploring-dedupe-with-python-ai-494771e39ea3
- canonical_url
- https://python.plainenglish.io/exploring-dedupe-with-python-ai-494771e39ea3
- author_url
- https://medium.com/@jallenswrx2016
- status
- ok
- fetched_at
- 2026-07-18 01:40:03