← Back to list

Exploring Dedupe with Python/AI

An exploration of Dedupe. A series of dedupe analysis with Python/AI

jon allen in Python in Plain English · 2025-08-22 02:48 · 0 claps · 11.5 min read
#python #dedupe #ai #comparison #prompting-technique
Open on Medium ↗
Wiki topics: FT · Fine-tuning & Adaptation PE · Prompt Engineering AI · AI · General

Exploring Dedupe with Python/AI

An exploration of Dedupe.

I have this data I have been collecting for a while. It is the NWS Marine forecast for the Potomac river near the Airport. It is a very structured text and I believe there is a high chance that it can be deduplicated for storage. Currently I store six of these a day in 1K or less text files.

ANZ535-220800
Tidal Potomac from Key Bridge to Indian Head
734 PM EDT Thu Aug 21 2025

SMALL CRAFT ADVISORY IN EFFECT UNTIL 8 PM EDT THIS EVENING

TONIGHT
N winds 5 to 10 kt. Waves 1 ft.

FRI
NE winds 5 kt. Waves less than 1 ft.

FRI NIGHT
SE winds around 5 kt. Waves 1 ft.

SAT
S winds 5 to 10 kt. Waves 1 ft.

SAT NIGHT
S winds 5 to 10 kt. Waves 1 ft.

SUN
S winds 5 kt. Waves less than 1 ft. A chance of showers.

SUN NIGHT
S winds around 5 kt. Waves 1 ft. A chance of showers.

MON
NW winds 5 to 10 kt. Waves 1 ft.

TUE
NW winds 5 to 10 kt. Waves 1 ft. 

There are at least 4 types of dedupe:

  1. Word-level — easy to understand. High pointer overhead.
  2. Sentence or Phrase level — good for repetitive text, works on structured text
  3. Fixed-block chunking— best for larger datasets, that don’t have much change
  4. Variable-block chunking — hard to implement.

AI generated image of deduplication in Renaissance Italy.

AI generated image of deduplication in Renaissance Italy.

So, lets do some analysis of our data to see what is the best fit.

you are an expert on deduplication methods and expert at python cli coding.

create a cli program that takes a glob of files such as *.txt. 
Tabulates the duplicate words.

counts the bytes of each word as well.

Goal is to save space by using dedupe techiques
Final report of the program is to give a summary of top 50 unique words.
Potential dedupe factor and reduction in bytes for the data analyzed.

I will put the source code at the end and provide a github link

Here is the output

--- Deduplication Analysis Report ---
Number of Processed Files: 7551
-----------------------------------

--- Overall Statistics ---
Total Words Analyzed: 1,107,614
Unique Words Found:   924
-----------------------------------

--- Deduplication Potential ---
Original Size (sum of all word bytes): 3,645.96 KB
Deduplicated Size (unique words + pointers): 4,330.48 KB
Potential Bytes Saved: -684.52 KB
Potential Deduplication Factor: 0.84x
Potential Space Reduction: -18.77%
-----------------------------------

--- Top 50 Most Frequent Words ---
Rank  Word                 Frequency       Bytes/Word  
----  -------------------  --------------  ----------- 
1     kt                   78,468          2           
2     1                    66,593          1           
3     winds                66,216          5           
4     waves                66,102          5           
5     ft                   62,819          2           
6     to                   59,357          2           
7     5                    52,674          1           
8     10                   34,878          2           
9     night                20,225          5           
10    of                   19,851          2           
11    showers              19,166          7           
12    around               16,504          6           
13    a                    16,463          1           
14    chance               16,463          6           
15    less                 15,779          4           
16    and                  15,216          3           
17    nw                   13,385          2           
18    with                 12,813          4           
19    gusts                12,796          5           
20    s                    12,061          1           
21    tstms                11,772          5           
22    in                   11,526          2           
23    than                 10,310          4           
24    sw                   9,847           2           
25    20                   9,798           2           
26    15                   9,609           2           
27    from                 9,395           4           
28    w                    9,324           1           
29    wed                  8,283           3           
30    tue                  8,280           3           
31    thu                  8,270           3           
32    fri                  8,262           3           
33    sat                  8,244           3           
34    mon                  8,243           3           
35    sun                  8,236           3           
36    the                  7,885           3           
37    anz535               7,551           6           
38    tidal                7,551           5           
39    potomac              7,551           7           
40    key                  7,551           3           
41    bridge               7,551           6           
42    indian               7,551           6           
43    head                 7,551           4           
44    tonight              7,488           7           
45    this                 6,739           4           
46    afternoon            6,438           9           
47    n                    6,266           1           
48    edt                  6,258           3           
49    e                    5,941           1           
50    rain                 5,769           4           
-----------------------------------

The program assumes a 4 byte pointer — so the overhead of a word based method doesn’t work.

So, lets do a sentence based analysis

Prompt:

you are an expert on deduplication methods and expert at python cli coding.

create a cli program that takes a glob of files such as *.txt. 
Tabulates the duplicate sentences.

counts the bytes of each sentence as well. a new line can be a sentence
Goal is to save space by using dedupe techiques.
Final report of the program is to give a summary of top 50 unique 
sentences or phrases
Potential dedupe factor and reduction in bytes for the data analyzed.

in the report print number of files processed, but not names of files.

print bytes processed and other significant observations of the data
add a --limit parm where you specify all as 0 or as many top as you want.

Here is the run — we would achieve almost 8 to 1 compression assuming a 4 byte reference pointer.

--- Deduplication Analysis Report ---
Files Processed: 7551
Total Bytes Processed: 4,844,327 bytes
Total Sentences/Lines: 181,407
Unique Sentences/Lines: 15,205

--- Potential Savings ---
Potential Reduction: 4,239,096 bytes
Deduplication Factor: 8.00x
(This means the original data is ~8.00 times larger than the unique data)

--- Top 50 Unique Sentences/Phrases by Frequency ---
Rank  | Count      | Bytes (Single)  | Bytes (Total)   | Sentence/Phrase
--------------------------------------------------------------------------------
1     | 6,809      | 44              | 299,596         | "Tidal Potomac from Key Bridge to Indian Head"
2     | 6,050      | 7               | 42,350          | "TONIGHT"
3     | 4,682      | 3               | 14,046          | "THU"
4     | 4,679      | 3               | 14,037          | "FRI"
5     | 4,675      | 3               | 14,025          | "WED"
6     | 4,670      | 3               | 14,010          | "TUE"
7     | 4,668      | 3               | 14,004          | "SAT"
8     | 4,656      | 3               | 13,968          | "SUN"
9     | 4,653      | 3               | 13,959          | "MON"
10    | 2,934      | 64              | 187,776         | "Winds and waves higher and visibilities lower in and near tstms."
11    | 2,689      | 32              | 86,048          | "NW winds 5 to 10 kt. Waves 1 ft."
12    | 2,526      | 9               | 22,734          | "WED NIGHT"
13    | 2,525      | 9               | 22,725          | "TUE NIGHT"
14    | 2,519      | 9               | 22,671          | "THU NIGHT"
15    | 2,506      | 9               | 22,554          | "FRI NIGHT"
16    | 2,503      | 9               | 22,527          | "SAT NIGHT"
17    | 2,503      | 9               | 22,527          | "SUN NIGHT"
18    | 2,502      | 9               | 22,518          | "MON NIGHT"
19    | 1,869      | 5               | 9,345           | "TODAY"
20    | 1,797      | 32              | 57,504          | "S winds around 5 kt. Waves 1 ft."
21    | 1,705      | 31              | 52,855          | "S winds 5 to 10 kt. Waves 1 ft."
22    | 1,636      | 6               | 9,816           | "tstms."
23    | 1,498      | 14              | 20,972          | "THIS AFTERNOON"
24    | 1,472      | 35              | 51,520          | "S winds 5 kt. Waves less than 1 ft."
25    | 1,457      | 31              | 45,167          | "N winds 5 to 10 kt. Waves 1 ft."
26    | 1,393      | 32              | 44,576          | "SW winds 5 to 10 kt. Waves 1 ft."
27    | 1,387      | 31              | 42,997          | "W winds 5 to 10 kt. Waves 1 ft."
28    | 1,303      | 13              | 16,939          | "REST OF TODAY"
29    | 1,294      | 33              | 42,702          | "SW winds around 5 kt. Waves 1 ft."
30    | 1,087      | 33              | 35,871          | "NW winds around 5 kt. Waves 1 ft."
31    | 1,055      | 11              | 11,605          | "Waves 1 ft."
32    | 1,039      | 32              | 33,248          | "E winds around 5 kt. Waves 1 ft."
33    | 977        | 36              | 35,172          | "NW winds 5 kt. Waves less than 1 ft."
34    | 976        | 32              | 31,232          | "W winds around 5 kt. Waves 1 ft."
35    | 953        | 15              | 14,295          | "REST OF TONIGHT"
36    | 916        | 33              | 30,228          | "SE winds around 5 kt. Waves 1 ft."
37    | 898        | 35              | 31,430          | "N winds 5 kt. Waves less than 1 ft."
38    | 897        | 36              | 32,292          | "SW winds 5 kt. Waves less than 1 ft."
39    | 832        | 32              | 26,624          | "NE winds 5 to 10 kt. Waves 1 ft."
40    | 778        | 33              | 25,674          | "NE winds around 5 kt. Waves 1 ft."
41    | 750        | 32              | 24,000          | "N winds around 5 kt. Waves 1 ft."
42    | 742        | 45              | 33,390          | "Tidal Potomac from Key Bridge to Indian Head-"
43    | 733        | 35              | 25,655          | "W winds 5 kt. Waves less than 1 ft."
44    | 701        | 36              | 25,236          | "NE winds 5 kt. Waves less than 1 ft."
45    | 661        | 10              | 6,610           | "and tstms."
46    | 614        | 35              | 21,490          | "E winds 5 kt. Waves less than 1 ft."
47    | 593        | 9               | 5,337           | "OVERNIGHT"
48    | 577        | 31              | 17,887          | "E winds 5 to 10 kt. Waves 1 ft."
49    | 571        | 52              | 29,692          | "NW winds 5 to 10 kt with gusts to 20 kt. Waves 1 ft."
50    | 549        | 16              | 8,784           | "chance of tstms."

What about chunking?


Conversation with Gemini
you are an expert on deduplication methods and expert at python cli coding.

Create a cli program that takes a glob of files such as *.txt. Take input into fixed sized chunks specified by --chunk. size is in bytes

Goal is to save space by using dedupe techiques.

you are an expert on deduplication methods and expert at python cli coding.

create a cli program that takes a glob of files such as *.txt. 
Tabulates the duplicate sentences.

counts the bytes of each sentence as well. a new line can be a sentence

Goal is to save space by using dedupe techiques.

Final report of the program is to give a summary of top 50 unique fixed size chunks

have a --limit parm where you specify 0 for all chunks, or number, but default to 50.

I did several runs

4096 — no duplicates at all

1024 — not worth it

256/128 — some, but not enough to be worth doing

[*] Starting analysis with chunk size: 1024 bytes
[*] Searching for files with pattern: /home/jon2allen/python/data/jon2allen/ANZ/data/*.txt

[*] Found 7551 files to process.

--- Deduplication Report ---

Summary:
  - Total Files Analyzed: 7551
  - Total Original Size: 4992.68 KB (5112501 bytes)
  - Chunk Size: 1024 bytes
  - Total Chunks Processed: 90
  - Unique Chunks Found: 78
  - Size After Dedupe: 78.00 KB (79872 bytes)
  - Duplicate Chunks Found: 10
  - Total Potential Savings: 12.00 KB (12288 bytes)
  - Deduplication Ratio: 64.01:1

Top 10 Duplicate Chunks by Potential Savings:
----------------------------------------------------------------------
Rank  | Chunk Hash (first 16 chars)    | Count      | Savings (Bytes)     
----------------------------------------------------------------------
1     | 414e5a3533352d31...            | 4          | 3072                
2     | 414e5a3533352d30...            | 2          | 1024                
3     | 414e5a3533352d32...            | 2          | 1024                
4     | 414e5a3533352d30...            | 2          | 1024                
5     | 414e5a3533352d30...            | 2          | 1024                
6     | 414e5a3533352d30...            | 2          | 1024                
7     | 414e5a3533352d32...            | 2          | 1024                
8     | 414e5a3533352d30...            | 2          | 1024                
9     | 414e5a3533352d30...            | 2          | 1024                
10    | 414e5a3533352d32...            | 2          | 1024                
----------------------------------------------------------------------
[jon2allen@jons-bad-ass-fedora-server-37 dedupe]$ ./dedupe_analyzer_chunk.py --chunk 2048 --limit 50 '/home/jon2allen/python/data/jon2allen/ANZ/data/*.txt'
[*] Starting analysis with chunk size: 2048 bytes
[*] Searching for files with pattern: /home/jon2allen/python/data/jon2allen/ANZ/data/*.txt

[*] Found 7551 files to process.

--- Analysis Complete ---
No full-sized chunks found to analyze. Try a smaller chunk size.
[jon2allen@jons-bad-ass-fedora-server-37 dedupe]$ ./dedupe_analyzer_chunk.py --chunk 256 --limit 50 '/home/jon2allen/python/data/jon2allen/ANZ/data/*.txt'
[*] Starting analysis with chunk size: 256 bytes
[*] Searching for files with pattern: /home/jon2allen/python/data/jon2allen/ANZ/data/*.txt

[*] Found 7551 files to process.

--- Deduplication Report ---

Summary:
  - Total Files Analyzed: 7551
  - Total Original Size: 4992.68 KB (5112501 bytes)
  - Chunk Size: 256 bytes
  - Total Chunks Processed: 16157
  - Unique Chunks Found: 13663
  - Size After Dedupe: 3415.75 KB (3497728 bytes)
  - Duplicate Chunks Found: 2418
  - Total Potential Savings: 623.50 KB (638464 bytes)
  - Deduplication Ratio: 1.46:1

Top 50 Duplicate Chunks by Potential Savings:
----------------------------------------------------------------------
Rank  | Chunk Hash (first 16 chars)    | Count      | Savings (Bytes)     
----------------------------------------------------------------------
1     | 7468616e20312066...            | 5          | 1024                
2     | 414e5a3533352d32...            | 4          | 768                 
3     | 414e5a3533352d30...            | 4          | 768                 
4     | 414e5a3533352d31...            | 4          | 768                 
5     | 6573732e0a0a4652...            | 4          | 768                 
6     | 0a0a5341540a5720...            | 4          | 768                 
7     | 677573747320746f...            | 4          | 768                 
8     | 414e5a3533352d30...            | 4          | 768                 
9     | 30206b742e205761...            | 4          | 768                 
10    | 414e5a3533352d30...            | 4          | 768                 
11    | 414e5a3533352d31...            | 4          | 768                 
12    | 30206b742e2e2e20...            | 4          | 768                 
13    | 206368616e636520...            | 4          | 768                 
14    | 414e5a3533352d30...            | 4          | 768                 
15    | 414e5a3533352d32...            | 4          | 768                 
16    | 206b742e2e2e2062...            | 4          | 768                 
17    | 414e5a3533352d31...            | 3          | 512                 
18    | 2077696e64732035...            | 3          | 512                 
19    | 726f75732073686f...            | 3          | 512                 
20    | 55204e494748540a...            | 3          | 512                 
21    | 7769746820610a63...            | 3          | 512                 
22    | 414e5a3533352d31...            | 3          | 512                 
23    | 6d732e0a0a465249...            | 3          | 512                 
24    | 6f203130206b742e...            | 3          | 512                 
25    | 490a452077696e64...            | 3          | 512                 
26    | 4e572077696e6473...            | 3          | 512                 
27    | 414e5a3533352d32...            | 3          | 512                 
28    | 66742e0a0a534154...            | 3          | 512                 
29    | 696e647320352074...            | 3          | 512                 
30    | 414e5a3533352d31...            | 3          | 512                 
31    | 7468206775737473...            | 3          | 512                 
32    | 494748540a572077...            | 3          | 512                 
33    | 6b742e2057617665...            | 3          | 512                 
34    | 4652490a45207769...            | 3          | 512                 
35    | 414e5a3533352d31...            | 3          | 512                 
36    | 20616e640a747374...            | 3          | 512                 
37    | 35206b742e205761...            | 3          | 512                 
38    | 414e5a3533352d30...            | 3          | 512                 
39    | 414e5a3533352d30...            | 3          | 512                 
40    | 4748540a53207769...            | 3          | 512                 
41    | 414e5a3533352d31...            | 3          | 512                 
42    | 203520746f203130...            | 3          | 512                 
43    | 2e20576176657320...            | 3          | 512                 
44    | 414e5a3533352d32...            | 3          | 512                 
45    | 2077697468206775...            | 3          | 512                 
46    | 206b742e2e2e2064...            | 3          | 512                 
47    | 657320312066742e...            | 3          | 512                 
48    | 3020746f20313520...            | 3          | 512                 
49    | 742e204120636861...            | 3          | 512                 
50    | 52490a5357207769...            | 3          | 512                 
----------------------------------------------------------------------
[jon2allen@jons-bad-ass-fedora-server-37 dedupe]$ ./dedupe_analyzer_chunk.py --chunk 128 --limit 50 '/home/jon2allen/python/data/jon2allen/ANZ/data/*.txt'
[*] Starting analysis with chunk size: 128 bytes
[*] Searching for files with pattern: /home/jon2allen/python/data/jon2allen/ANZ/data/*.txt

[*] Found 7551 files to process.

--- Deduplication Report ---

Summary:
  - Total Files Analyzed: 7551
  - Total Original Size: 4992.68 KB (5112501 bytes)
  - Chunk Size: 128 bytes
  - Total Chunks Processed: 36183
  - Unique Chunks Found: 30285
  - Size After Dedupe: 3785.62 KB (3876480 bytes)
  - Duplicate Chunks Found: 5614
  - Total Potential Savings: 737.25 KB (754944 bytes)
  - Deduplication Ratio: 1.32:1

Top 50 Duplicate Chunks by Potential Savings:
----------------------------------------------------------------------
Rank  | Chunk Hash (first 16 chars)    | Count      | Savings (Bytes)     
----------------------------------------------------------------------
1     | 7468616e20312066...            | 5          | 512                 
2     | 646e696768742e0a...            | 5          | 512                 
3     | 0a4652490a572077...            | 5          | 512                 
4     | 2045445420544849...            | 5          | 512                 
5     | 414e5a3533352d32...            | 4          | 384                 
6     | 4f4e494748540a4e...            | 4          | 384                 
7     | 2077696e64732035...            | 4          | 384                 
8     | 53204556454e494e...            | 4          | 384                 
9     | 414e5a3533352d30...            | 4          | 384                 
10    | 6c65737320746861...            | 4          | 384                 
11    | 6473203520746f20...            | 4          | 384                 
12    | 2045445420544849...            | 4          | 384                 
13    | 203235206b742e20...            | 4          | 384                 
14    | 414e5a3533352d31...            | 4          | 384                 
15    | 465445524e4f4f4e...            | 4          | 384                 
16    | 6573732e0a0a4652...            | 4          | 384                 
17    | 2057617665730a32...            | 4          | 384                 
18    | 0a0a5341540a5720...            | 4          | 384                 
19    | 35206b7420616674...            | 4          | 384                 
20    | 677573747320746f...            | 4          | 384                 
21    | 6265636f6d696e67...            | 4          | 384                 
22    | 0a0a544f4e494748...            | 4          | 384                 
23    | 414e5a3533352d30...            | 4          | 384                 
24    | 616e20312066742e...            | 4          | 384                 
25    | 30206b742e205761...            | 4          | 384                 
26    | 2057617665732031...            | 4          | 384                 
27    | 6c79207769746820...            | 4          | 384                 
28    | 642035206b742e20...            | 4          | 384                 
29    | 414e5a3533352d30...            | 4          | 384                 
30    | 66742e0a0a544f4e...            | 4          | 384                 
31    | 414e5a3533352d31...            | 4          | 384                 
32    | 4944415920414654...            | 4          | 384                 
33    | 30206b742e2e2e20...            | 4          | 384                 
34    | 65726e6f6f6e2077...            | 4          | 384                 
35    | 206368616e636520...            | 4          | 384                 
36    | 572077696e647320...            | 4          | 384                 
37    | 73203520746f2031...            | 4          | 384                 
38    | 4553542054484953...            | 4          | 384                 
39    | 414e5a3533352d30...            | 4          | 384                 
40    | 4544542054484953...            | 4          | 384                 
41    | 0a4e572077696e64...            | 4          | 384                 
42    | 2045535420544849...            | 4          | 384                 
43    | 414e5a3533352d32...            | 4          | 384                 
44    | 4953204556454e49...            | 4          | 384                 
45    | 206b742e2e2e2062...            | 4          | 384                 
46    | 68696e6720746f20...            | 4          | 384                 
47    | 6176657320312066...            | 4          | 384                 
48    | 657373207468616e...            | 3          | 256                 
49    | 20746f203130206b...            | 3          | 256                 
50    | 6c6573732e205061...            | 3          | 256                 
----------------------------------------------------------------------

For this NWS marine forecast data— it looks like the sentence/phrase is the best. The NWS marine forecast is somewhat structured. Here is one of the scripts.

[embed]

Here is the github repository

[embed]GitHub - jon2allen/dedupe: Python dedupe analysis tool Python dedupe analysis tool. Contribute to jon2allen/dedupe development by creating an account on GitHub.github.com

git clone https://github.com/jon2allen/dedupe.git
cd dedupe
./run.sh

I used 3 separate prompts for each script. So, the report format is not consistent. A lesson learned is maybe specify more about the report format. LLM may not do the same thing twice.

I will revisit and implement the sentence dedupe and see how it can work in practice.

A message from our Founder

Hey, Sunil here. I wanted to take a moment to thank you for reading until the end and for being a part of this community.

Did you know that our team run these publications as a volunteer effort to over 3.5m monthly readers? We don’t receive any funding, we do this to support the community. ❤️

If you want to show some love, please take a moment to follow me on LinkedIn, TikTok, **Instagram. You can also subscribe to our [weekly newsletter](https://newsletter.plainenglish.io/)**.

And before you go, don’t forget to clap and follow the writer️!


메타데이터
post_id
494771e39ea3
slug
exploring-dedupe-with-python-ai-494771e39ea3
url
https://python.plainenglish.io/exploring-dedupe-with-python-ai-494771e39ea3
canonical_url
https://python.plainenglish.io/exploring-dedupe-with-python-ai-494771e39ea3
author_url
https://medium.com/@jallenswrx2016
status
ok
fetched_at
2026-07-18 01:40:03