Cataloging Automation at the National Building Arts Center
Background: What is the NBAC?
Cataloging Automation at the National Building Arts Center

National Building Arts Center Publication Archive
Background: What is the NBAC?
This past summer I worked as an intern at the National Building Arts Center, an archive and study center in St. Louis that holds the largest collection of built environment artifacts in the United States. Alongside these artifacts, NBAC maintains a large archive and library that extensively covers how America was built: architectural design, materials manufacturing, labor history, patenting, geology, and a plethora of other practices that made these buildings possible. This archive supports authors working on book projects, historians tracing a manufacturer or material back to its source, and even university classes studying historic preservation. Many times, the only surviving documentation for how something was made is within a trade catalog that the NBAC holds.
The Problem: A Catalog Bottleneck
However, a collection such as this is only as findable as its catalog. An artifact or periodical sitting on a shelf but not in a database is incredibly hard for a researcher to know exists, much less utilize in their work. At the time of writing, the periodicals and books currently listed in the NBAC’s online catalog are a small minority compared to their overall collection. This ratio reflects the difficulty in cataloging an item by hand: pulling a book off the shelf, examining its bibliographic information, manually typing the fields into the cataloging software, putting the item back, and repeating. My work this summer aimed to streamline this process. Alongside Iovane Sakhelashvil, a fellow SLU student, and Alexis McKenzie Dyer, our technical lead and former professor, we created a cataloging pipeline that takes photos of books, uses deterministic code and generative AI to pull out important information, and produces a file that can be imported into cataloging software.

Step 1: Figuring Out Which Photo Is Which Book
The first technical obstacle we encountered was that the pipeline requires a way to determine which physical item a photo belongs to. Based on a dump of images, identifying whether an image was of one book or another would be both technically intensive and error-prone if the images were out of order or lacked enough information to identify the book using bibliographic details. This obstacle led to our idea of altering the physical picture-taking workflow to include a QR code for each item, representing a code in accordance with the NBAC’s ID convention. The ID convention is as follows:
A1.(Building #).(Bay #).(Shelf #)-(Item #), where the first book on shelf 1 of bay 1 in building 102 would be A1.102.1.1–001
NOTE: The Dewey Decimal System doesn't make sense for them for reasons we do not have time to discuss here.
In order to actualize this physical workflow problem, we developed a Python script that produces a PDF of QR codes based on which shelf you are about to photograph + roughly how many books are on said shelf. Now, we had a means for the future pipeline to properly associate an item with all of its respective images.

Step 2: Reading the Text with OCR
The next step in automating the extraction of an item’s metadata was performing OCR on the images themselves in order to extract text. Upon benchmarking the results of popular OCR libraries (PaddleOCR, EasyOCR, Tesseract), it became clear that Tesseract was the most performant option for both speed and accuracy.

Step 3: Cleaning Up the Images
However, even though Tesseract was the most performant, the OCR results were not accurate enough to be used in a meaningfully streamlined cataloging process. Errors in output included misspellings (a line being read as “THE ENQINGEER” rather than “THE ENGINEER”, for example), certain lines missing from the output entirely, and overall low OCR confidence scores. These errors warranted introducing preprocessing techniques such as grayscale conversion, Otsu thresholding, adaptive Gaussian thresholding, contour detection, bounding-box cropping, bicubic upscaling, deskewing, and multi-variant generation with confidence-based selection.
Grayscale conversion: removes color from the image to decrease the noise the OCR engine must go through, as color triples the amount of data sent and complicates defining a letter as dark text on a light background.
Otsu thresholding: Turns the gray image into pure black and white by looking at the spread of brightness values across the image and chooses the value that best separates “dark” from “light”. This results in more defined letter shapes.
Adaptive Gaussian thresholding: Same idea as Otsu thresholding but chooses a value for different parts of the image, allowing letters to be readable everywhere in the frame
Contour detection: Traces the outlines of shapes in the black and white image. This allows us to know which part of the image is a page, for example, and which is irrelevant noise.
Bounding box cropping: Draws a rectangle around the outline created via contour detection and removes irrelevant parts of the image. This prevents the OCR engine from outputting fake characters it mistakenly identifies from shadows, background shapes, etc.
Bicubic upscaling: Enlarges the image and presents the letters at the stroke width/spacing Tesseract expects. (Tesseract was trained on text the size of ~300 DPI scan, and pictures from our phones do not match that).
Deskewing: Rotates tilted text to horizontal text so that Tesseract reads characters in the correct order.
Multi-variant generation with confidence-based selection: The pipeline builds several combinations of the preprocessed image and keeps the variant with the highest confidence score, as these techniques do not perform equally on all images.

Raw Image

Preprocessed Image
Step 4: Fixing Typos with SymSpell
Once the OCR output is produced, there might still be character misrepresentations. In order to circumvent this, we implemented a Python library called SymSpell, which essentially works as a spell check for the text Tesseract produces. It contains a dictionary of English words with frequency counts. For a misread word, it finds the closest real words and picks the most common one. For example, OCR output like “libraly” would become “library” once passed through SymSpell.
Step 5: Sorting Text into Catalog Fields
Now that text was properly able to be extracted from the images of the publications, a data normalization step was required in order to take the text and fit it into the proper columns. These fields included metadata such as title, author, publication date, publisher, and ISBN. Data normalization was accomplished via LLM API calls (Google Gemini), where the prompt included the column specifications as well as the JSON body containing the text obtained via OCR. An example of what goes into the LLM stage and what comes out is as follows:
OCR output (shortened for length)
{ “item_id”: “A1.100.67.7–001”, “images”: [{ “image_name”: “0ee9d9f5-…jpg”, “lines”: [ { “text”: “A1.100.67.7–001”, “confidence”: 0.955 }, { “text”: “Copyright © 1962 by”, “confidence”: 0.884 }, { “text”: “Reinhold Publishing Corporation”, “confidence”: 0.978 }, { “text”: “Originally Published by”, “confidence”: 0.947 }, … more lines ], “confidence”: 0.925 }] }
LLM output
{ “item_id”: “A1.100.67.7–001”, “llm_response”: { “title”: “Engineering Alloys”, “subtitle”: null, “author”: null, “publisher”: “Reinhold Publishing Corporation”, “llm_confidence”: 0.7 }, “final_record”: { “title”: “Engineering Alloys”, “subtitle”: null, “author”: null, “publisher”: “Reinhold Publishing Corporation”, “publication_year”: 1962, “isbn”: null, “llm_confidence”: 0.7, “ocr_confidence”: 0.925 } }
Tying It All Together
Now that the core components of the pipeline were developed, the next focus was to create the orchestration layer between the stages.

The orchestration layer is built around two CSV files that act as the pipeline’s memory. The first one is image_status.csv, which contains one row per photo and stores what has actually been done to the image so far: the original filename, the image’s UUID, the item ID, how and if it was sorted, whether it has undergone OCR, whether the data has been normalized by the LLM, and a “needs review” flag. This flag is marked true if the OCR confidence is below the threshold, the LLM confidence is below the threshold, or the item is missing a required field such as author. Output.csv, on the other hand, contains one row per item and is essentially a draft of the final CSV which will be imported into the cataloging software. It holds the data dictionary-related fields as well as whether or not the entry needs human review. We decided to structure the pipeline this way so that the stages can be resumed and not require the entire pipeline to run in one setting. Each stage reads the CSV manifest and uses it to determine if a photo needs to undergo that respective stage.
A rough summary of each stage is as follows:
- Sort: Takes the folder of unorganized photos, decodes the QR code of each image, creates a UUID for each image, and moves the images into their respective per-item folder. Moves any images whose QR could not be decoded into an unidentified_photos folder for later human review.
- OCR: Performs the preprocessing and OCR, writing the text and confidence score into a per-item JSON file that sits alongside the images in their folder.
- Data normalization: Sends the JSON body of the OCR text and the column specifications to the LLM. ISBNs and publication years are extracted deterministically using regex. Creates output.csv where each item entry has a row, even if not ready for importing into the cataloging software.
- Processing: Reads output.csv, maps the pipeline’s column names onto CatalogIt’s import columns, and writes only valid entries (defined as having high enough OCR/LLM confidence and containing all required fields) to final_catalogit.csv.
Manual Sorting
It is unlikely that the automated pipeline will get every item right, so we built two command-line scripts for the cases that need human review. The first is the manual sorting script, which works through the unidentified_photos folder (photos in which the QR code could not be decoded by the pipeline) and allows the user to go through and identify the items properly. This is made possible by the script opening each photo in the computer’s photo viewer and asking the reviewer to read the item ID from the QR code sticker and type it into the command line. The photo is then moved into that item’s folder, just as the sorting stage would have done. If a photo is unusable (blurry, a random photo that somehow found its way in, etc.), the reviewer can also tell the script to discard it.

CLI for typing the ID from the image
Manual Review
The second script’s purpose is to manually review items that have been flagged as inadequate by the pipeline. For each item, it opens all of the associated photos at once and provides the reason that it was flagged (either low OCR confidence, low LLM confidence, or a missing required field). The reviewer is then able to enter each field and submit, writing the entry to the final output CSV.

Google Drive Synchronization
To allow for multiple people using the pipeline and sharing photos and data, the pipeline is connected to a Google Drive folder via the Drive API. This way, photos uploaded straight from a phone to the Drive folder can go through the entire pipeline on someone else’s machine without needing to be copied across devices. Authentication uses Google’s OAuth flow.
Packaging with Docker
The OCR project depends on more than just Python packages. Tesseract, the QR decoding library, and OpenCV all need system-level software. Installing each of these onto each staff member’s machine would be a very fragile and time-consuming approach, so we containerized the pipeline with Docker. The Docker image comes with every dependency already installed, so running the pipeline will behave the same on any machine. Sensitive files (like Google Drive credentials) are kept out of the image and are provided at runtime. The only parts that run outside of the container are the manual sorting and review scripts, as they need to open photos on the reviewer’s screen.
The Result
The result of our project is that the process of cataloging no longer includes grabbing a book, reading it, and manually entering the data fields into cataloging software. It is instead a much faster process of uploading images of the items, running the pipeline, and using the CLI to review a minority of entries that raised accuracy concerns. For a collection of tens of thousands of items, the amount of time to be saved is immense.
What I Learned
The most valuable thing I learned by working on this project is that the best solution to a problem is not always the most technical or impressive one. When faced with the issue of determining which book a photo belongs to, my first instinct was to solve it in the software itself. Approaches like matching OCR text across photos or clustering images by visual similarity both seemed interesting, but they also would have been fragile. Changing the photography workflow to include a QR code only took a couple of hours of coding and made the problem disappear entirely. This project taught me that rather than building a system smart enough to handle super unconstrained input, it can be easier to constrain the input itself. I expect this fact to be relevant in any of my future engineering work as well.
The next valuable lesson I took away from this project is the importance of benchmarking rather than assuming. For example, I had used OCR libraries in the past and initially implemented an approach that performed well for scanned PDFs in our pipeline. However, after running PaddleOCR, EasyOCR, and Tesseract against the same set of real photos from the collection, I was able to conclude that Tesseract by itself performed the best in every metric for our use case. This realization influenced design choices such as the multi-variant preprocessing approach. I learned I can’t assume one combination of techniques would be the best for every scenario.
Overall, I really enjoyed how tangible this work was, and knowing the NBAC staff would use and benefit from the project made me want to get it right. It was very valuable for me to gain experience in solving real-world problems with code where there is not only one right answer. I’m also grateful for the time Alexis put in every week to review our code and talk through decisions with us. Her advice shaped a lot of what’s in this article, and I learned a lot from those conversations.
메타데이터
- post_id
- 636866e8cd98
- slug
- cataloging-automation-at-the-national-building-arts-center-636866e8cd98
- url
- https://blog.newmathdata.com/cataloging-automation-at-the-national-building-arts-center-636866e8cd98
- canonical_url
- https://blog.newmathdata.com/cataloging-automation-at-the-national-building-arts-center-636866e8cd98
- author_url
- https://medium.com/@alecarnoldslu
- status
- ok
- fetched_at
- 2026-10-05 13:15:06