Transforming Wellcome’s Spanish Manuscripts
By Will Greenacre & Odalys Caballero
Transforming Wellcome’s Spanish Manuscripts
By Odalys Caballero Valdés and Will Greenacre

MS.1422, Libro de ruedas y diferentes curiosidades. https://wellcomecollection.org/works/a4u5fmc8
To read this article in Spanish, click here.
The library at Wellcome Collection has a unique collection of thousands of Western manuscripts, dating back to the formation of Henry Wellcome’s collections in the early 20th century, although the collection has been added to ever since. Although they are centred around the unifying theme of medicine, they cover a remarkable range of subjects including alchemy, astrology, botany, chemistry, mathematics, and magic, to name just a few. They are an extraordinarily rich resource, in a multitude of languages, and at Wellcome Collection we are continually seeking new and better ways to make them more discoverable and accessible to researchers. One of our recent projects has built on our work using the Text Encoding Initiative (TEI) for manuscript description, applying it specifically to our collection of Spanish manuscripts and using our knowledge of Spanish culture, history, and language to make these manuscripts more visible.
According to the creation date of the Spanish manuscripts, they were written during the colonization of America by Spain. America was discovered for Spain in 1492. From this date a new period started for the Indigenous people who were submitted to an unknown culture. The conquerors brought to the discovered land new food, a new religion, different customs but also new diseases. Many of the natives died not only because of the cruelty of the colonizers, but also because of these imported diseases. In fact, there are reports in the collection about how many deaths were associated to specific diseases at various moments in time.
The meeting of these two cultures and the health consequences, led to the urgent need for new medical solutions. Most of the Collection’s manuscripts are about the new remedies, based primarily on herbs, used to face these new illnesses.
Several experts in this field suggest that the two cultures fused with the aim of achieving common cures for the diseases that emerged. Conquerors and conquered also established agreement in the detection and name of the same disease, although the criteria of the colonizers always prevailed.
Increasing the accessibility and discoverability of the Wellcome Collection’s Spanish-language collection is a way to decolonize the medical resources knowledge of the pre- Columbian era that could be applied to modern medicine.
The bulk of our Spanish and Western manuscripts were originally described in printed catalogues, primarily those by S.A.J. Moorat and Robin Price. Although both of these catalogues are digitised and freely available online, they are dense and lengthy volumes, and individual items within them are often difficult to find. When the library established its online catalogue around the year 2000, the information in these catalogues was digitally transcribed and imported into a digital database using CALM software, and became part of what is now the archives and manuscripts catalogue, with records of our manuscripts and archive collections sitting side by side in the same system.
There are problems with this approach however, since archives and manuscripts are different in fundamental ways: archive collections are aggregations of material with a unique arrangement that forms part of the material’s history and context (see here for an example), whereas manuscripts are generally (but not always) single items — a recipe book or journal for example. But this is not always a simple distinction: a manuscript may contain multiple items with their own unique histories that may be better catalogued as an archive, or multiple items may have been purposely assembled to create a new manuscript. One of the complicating factors in managing our collections is where to draw the line between archives and manuscripts, so that items are accurately described so as to bring out their unique and distinctive qualities, in a way that makes them as discoverable and accessible as possible.
What is TEI and why use it?
TEI by its acronym, or Text Encoding Initiative, is a list of guidelines to standardize and represent texts in digital format through XML (extensible markup language) metalanguage encoding. It is used mostly in the social sciences and humanities with an emphasis on primary sources such as manuscripts, archival documents, and correspondence, among others.
Using TEI in connection with to the XML markup language allows us to:
• Integrate data and its exchange on a large scale.
• Focus on the function of the content and not just its appearance.
• Define custom tags and describe them with specific attributes.
• Use a wide variety of languages.
• Enable its application by any human sciences professional while other languages need expertise for their use and understanding.
• Configure scenarios to decode texts written in another software.
The use of digital technologies like TEI enables us to appraise and describe the huge volume and variety of our collections in a way that would not have been possible even just a few years ago. As a starting point, there are around 200 manuscripts in our archive cataloguing system that are primarily in Spanish, and so we decided to use these to test a novel experimental approach that involved extracting data from the cataloguing system and converting it into TEI descriptions, rather than using the printed catalogue information to create TEI descriptions from scratch. Our hope was to come up with a faster and more efficient approach to creating TEI files that would be potentially scalable to the thousands of other manuscripts that are currently catalogued in our archives and manuscripts database and would be better described using TEI.
Our first step was to extract the catalogue information from CALM in raw form, to serve as the basis for our TEI files. For this we made use of CALM’s export function, which allows records to be exported into a number of formats, including XML. Next, we created an XSL stylesheet that would enable the shapeless block of XML text extracted from CALM to be easily converted to match our TEI template, which creates a coherent description that is consistent with all our other TEI manuscript files. Now, with just a few clicks, the catalogue record for each manuscript can be quickly and easily converted into a TEI file and uploaded to our online repository, saving time compared with the lengthier process of transcribing information from the printed catalogues from scratch.
Benefits of TEI for Wellcome Collection manuscripts
The benefits of TEI in maintaining the Wellcome manuscript collection are undeniable. It not only facilitates an efficient set-up for encoding CALM-catalogued manuscripts and their migration to a more widely used language, but it is also outstanding at optimising the findability of these collections.
The current catalogue offers limited information which reduces the discoverability of each manuscript. There are no fields that provide more specific information about each manuscript, always useful for any research. However, TEI enriches the catalogue information, showing specific fields such as decoration, conditions, binding, script, provenance, etc.
An example of the above is the manuscript, “Mexico, 19th century: medical uses of plants growing in Yucatán”. The images below show clearly what is the physical condition of the manuscript and how it is bound, including its decoration.

WMS/Amer.17, Noticia de varias plantas y sus virtudes. https://wellcomecollection.org/works/t49xmm9t
In the following images, you can see the difference between the current catalogue and the information that will be displayed once TEI has been applied.

The specification of the water stain as a sign of deterioration does not appear explicitly in the information offered by the catalogue in the first image. It is only possible to access such information after cataloguing with TEI as shown in the second image. The same thing happens with the specifics of the decoration: manuscript binding with red cloth, decorated with gold initials, on its cover. Information on the decoration and condition of the manuscript, in this case, is a good supplement to any investigation, providing valuable information to a researcher.
Another important benefit of TEI is that it makes it easier to improve the quality of the information that we provide. It allows us to detect and correct errors based on the cataloguer’s lack of knowledge of the original language. In this case we see that the title assigned to the manuscript has been inferred from its chronology, and that it differs entirely from the real title.

What would happen if a user searches for the actual title when the meaning is very different from the assigned title?
We hope that the solutions to these and other concerns, informed by working in TEI, will be further clarified by a physical inventory of Wellcome’s manuscripts in the near future. This would be the beginning of a deeper investigation into the spectacular collection of Spanish manuscripts in America that the Wellcome Collection treasures.
메타데이터
- post_id
- dd36ff2447d
- slug
- transforming-wellcomes-spanish-manuscripts-dd36ff2447d
- url
- https://stacks.wellcomecollection.org/transforming-wellcomes-spanish-manuscripts-dd36ff2447d
- canonical_url
- https://stacks.wellcomecollection.org/transforming-wellcomes-spanish-manuscripts-dd36ff2447d
- author_url
- https://medium.com/@w.greenacre
- status
- ok
- fetched_at
- 2026-07-26 17:24:34