← Back to list

Sciences of Language in Product Design

Using techniques from linguistics in the discovery and mapping of information architecture, along with some foundations from philosophy and…

STRATEGY CX · 2026-01-21 20:37 · 55 claps · 12.6 min read
#ooux #product-design #linguistics #user-experience #words
Open on Medium ↗
Wiki topics: UX · UI/UX Design PRD · Product Design PHI · Philosophy LNG · Linguistics & Language 🔬 · Science · General 🏛️ · Architecture

Sciences of Language in Product Design

Using techniques from linguistics in the discovery and mapping of information architecture, along with some foundations from philosophy and psycholinguistics.

By Francisco Nunes, Product Design Lead at STCX

image: Jolygon/Istock

image: Jolygon/Istock

When I have conversations with other peers about upskilling and craft, I often end up explaining how my background helps me become a better designer. I had the chance to discuss this with the design community from Porto Alegre-RS during a class on OOUX. During that session, I explained how I incorporate techniques from linguistics into my design process.

While people ask their questions, I felt the need to explore some of the topics I shared earlier more deeply. This article aims to show how language sciences contribute to the design process and how they work with OOUX (Object-Oriented User Experience).

Photo taken during my talk on OOUX and language, held at PUC-RS by GUIX.

Photo taken during my talk on OOUX and language, held at PUC-RS by GUIX.

This essay will not cover all the stages of my process. Fortunately, we will focus on applying these methods within the problem space, specifically when we already understand our users well, such as during benchmarking and when seeking insights into the ecosystem in which the user experience occurs. We will also explore how to grasp the big picture of a system, the behavior of the product and service, and, importantly, how to deliver a Product Specification (PRD) that clearly communicates the structure of the product we aim to build.

From Words to Products

Digital products and systems do not exist outside or apart from linguistic reality; they are embedded within it and occur through it.

Common Crawl screenshot — text and HTML account for 99% of the share among all types of media on the web.

Common Crawl screenshot — text and HTML account for 99% of the share among all types of media on the web.

What would the internet be without text?

The Common Crawl (CC) is a massive archive (in petabytes) of web-tracking data that makes available the entire internet for research, development, and AI. This project performs monthly scans, creating datasets of millions of web pages that are accessible to anyone, making it a fundamental pillar for the advancement of artificial intelligence and research across various specialized areas of knowledge. If we closely examine CC Statistics, we’ll see that 99% of all internet content is words.

In an information era, the word-as-data provides insights into how people connect and behave online. The word is the interface of social interaction; it is the bridge between the internal and external psychism. It is not insignificant that the data (I also mean words, linguistic data) are the “new petroleum,” and that the way these data are organized, structured, and synthesized is key to designing an architecture of information in complex systems.

Discovery and word mapping

The design field already integrates linguistic techniques into its process, so this is not new hype content. In OOUX, during the discovery stage of our problem space, we use “Noun-Foraging.”

This occurs when I analyze a selected corpus of media content, such as texts, articles, and bibliographic sources, to identify the key nouns and naming structures that constitute the discourse of a particular product. I have used this type of analysis on official documentation (such as data structures, APIs, and internal wikis) as well as on interview materials (audio transcriptions) from a public health management system, event management applications, and streaming platforms.

The data structures and API documentation had a well-established taxonomy and a hierarchy of concepts and objects that reveal how the product’s structure was or could be organized, so it makes less sense to try a mapping method in these documents.

  • If these documents already define or suggest what each feature inside the system is, I usually refer to them as “data structured documents”.

The situation is very different when I had to analyze interview transcriptions and other documents, such as blog posts, articles, journal pieces, and repository texts. I had an experience analyzing a selected corpus from a customer feedback service called “ReclameAqui” and exporting files from the Customer Success Team with texts that end-users send to helpdesk channels.

  • These types of documents do not define or suggest what each feature in the system is, but how people perceive their own experience with the system; I usually call these “non-structured data documents”.

In my experience working with consultancies, I found structured data to be more common than non-structured data. This happened because our clients had already worked on it: discovery reports, wikis, or because they already had a clear vision of the product, or the product already exists.

Well, in this first category, structured data, it is easier to find and work with patterns because, generally, the discourse is ideologically aligned with business interests. In the non-structured data category, the discourse may not be aligned, which opens the possibility of testing by comparing structured data with discoveries from non-structured data analysis.

For both categories, I utilize two techniques:

1. Discourse Analysis: Understanding the Informational Ecosystem.

The first thing is the “discourse analysis,” a discipline that draws on linguistics, psychoanalysis, and Marxism (historical dialectical materialism) and shares methodological and epistemological origins with them. So, when faced with a text, I start mapping:

  • The effects of sense that is being produced;
  • The enunciator’s position within the discourse (where they speak from, why they speak what they do, and how they speak);
  • What are the predicates that this enunciator-subject produces?
  • For whom are they speaking?
  • About what they discuss.

By doing so, I gain a broad, qualitative understanding of the context, the system, and the journey; as I reason about the problem we need to solve, this process provides me with a good starting point for identifying the objects that will structure my product. This practice is valuable for product discovery because it allows us to see the scope of the information ecosystem in which the user experience happens.

As I have said, it is a reasoning process, so we don’t need to “document the analysis” itself. The reasoning process must serve as a cognitive exercise rather than a procedural design step. Most of the time, our DoD must be tabulations; I mean, you can still document it if you want, or if you’re not used to this research method, or simply don’t have a research repertoire.

2. Morphological Analysis: Gathering evidence for an Information Architecture

The second thing I do, and the first delivery of our discovery process, is Noun-Foraging, a technique applied pioneering in Product Design by Sophia Prater through OOUX. This technique was directly inspired by linguistics, and it is a morphological mapping of the text.

  • What are the main nouns that appear in the content?
  • How many times do these nouns appear in the content?
  • What adverbs follow these nouns?
  • What adjectives come along with these nouns?
  • Which categories do these nouns represent in the product?

This is a quantitative analysis of qualitative content, aimed at extracting relevant data to help us build the “product boxes.” Before, I had to manually structure a table, but in OOUX, Sophia recommended a tool called “Parts-of-speech.”

Today’s reality demands new tools, and for that, I use Cursor to create functions for parsing content and identifying morphological variables I want to extract. Then, I generate a CSV or JSON with tabulated results. This way, I can analyze large amounts of text more quickly.

The advantage is that I do not become dependent on a tool that can limit my analysis because of a paywall or retain my data. Many projects contain sensitive privacy data, so Cursor is a great tool for keeping everything locally.

Up to this point, completing the first step (discourse analysis) has enabled me to analyze results tabulation and determine whether, for example, a certain noun is eligible to become a structure object. It is only possible to do this review properly if we have done this analysis beforehand; without that, we can easily make mistakes with numbers or the semantic diversity of isolated words.

In the example below, I exported a morphological analysis (only nouns) from content related to a team-building events app. You can see that ‘event’ (br- evento) appears frequently, obviously, but other words that are not necessarily suitable to represent objects also appear frequently.

As you can see, the word ‘event’ appears most frequently; this is the main object (obviously) in this type of product.

As you can see, the word ‘event’ appears most frequently; this is the main object (obviously) in this type of product.

If we look closely, the words ‘solution’ and ‘product’ appear in the 2nd and 3rd positions, respectively. The word ‘product’ is how some users refer to ‘event.’ This happens due to synonym usage and because, when discussing their experiences, some of them refer to products as a subcategory of events: they create an event, and within that event, they hire products (games, workshops, surveys).

When it comes to the word ‘solution’, it was used to describe “the object of seek and desire.” The solution was what they expected the system to deliver, so here the word solution does not serve as an object of information architecture.

The words ‘audience’, ‘person’, ‘company’, and ‘user’ appear in sequence, and here these words can serve as both objects and attributes of these objects. For example, we can have an object called ‘user’ that has a type of attribute granting different access levels inside the system; ‘audience’ is a piece of metadata that can be used to feed the recommendation engine, ‘person’ is the user who will interact with the purchased product (like a game or survey), and ‘company’ is the client that uses the app to create a team-building event.

See, in order to understand the relevance and the type of semantic cluster of the words in product structure categories (the category column indicates what type of object a word can be clustered), I need to have an understanding of the context, and this understanding emerges in the discourse analysis phase.

Through this analysis, I could understand the product’s main objects, reflect on their relationships, and, furthermore, validate these connections through an Object Relationship Matrix.

Fundamental pillars

Questions that might come up right now could be: Does this whole thing have any fundamentals? How can a simple morphological analysis reveal important elements of the information architect’s product?

Well, let me start by asking you this: what is the material evidence of something that’s inside someone’s mind about a certain product? Where do we all run when we work with information?

Out of the 600 common words shared across many languages, around 90% are NOUNS. — From Sophia Prater in OOUX

The best way to understand how people think as a user of a product or service is to see how they use language to talk about it. The word is the core of meaning; even when the meaning changes, the informational material, in the form of data, remains the word.

Does this approach have some fundamentals? To answer this question, in this section, we will see what the sciences of language tell us.

In Philosophy

Kant — The Categories of Understanding and Judgment

In philosophy, if we take Kant as a starting point, when he discusses the Judgment Faculty, he explains that it is the Judgment, a form of cognition that allows a person to define a conscious mental representation of an object, a cognitive capacity that organizes sensory data, and works like this:

  • The Intuition (Kant calls Sensitivity — Sinnlichkeit) is the sensorial stimulus that provides the raw material.
  • The imagination (Kant calls it Einbildungskraft) occurs when a person organizes the stimulus into shapes, and also when they create an image (mental image).
  • And the autoconsciousness of this process occurs when a person can make a judgment, reason about the external (the world) and internal (the mind) experience.

When Kant writes about the Categories of Understanding, he organizes them around what are also linguistic categories. So we can understand that the enunciation of a judgment (an opinion about a product, a review, a complaint) about the experience is only possible through the linguistic system.

We can even go back to the Greeks, and we will still see the same explanation: the Socratic categories of thought are essentially the categories of the Greek linguistic system.

Considering how philosophy views this matter, it is perfectly possible to treat the word as research data to understand the information architecture and to use morphology as a method of analysis to get closer to people’s concepts and objects about a certain experience. Moreover, this can be transformed into a better information architecture and produce interaction.

Bakhtin and Volochinov — The word as ideological core

For Bakhtin and Volochinov, consciousness exists and takes shape in the signs created by an organized group through social relations. Words are formed by a mixture of ideological threads and serve as the foundation for all social interactions across different domains; the word is always the most accurate indicator of social transformation.

Here, we already have statements that serve as evidence or indicators of how interactions with a product or service occur. Since words are produced within social relations, this is where we can start to understand how these interactions happen.

The word is so closely connected to social interactions, that is, the way people organize themselves and the conditions under which they do so, that the authors have said a change in these forms produces a change in the sign (in the data, in the word).

In this moment of developing the theory, the authors discuss how signs (while ‘containers’ of meaning) are objects shaped by social interaction, and that every sign, including individuality, is social.

To Bakhtin and Volochinov, the sign and the social situation in which the sign is embedded are inextricably linked. What applies to the product design process here is that it is impossible to separate the semantic content of words from their ideological context (social, historical, situated within their own interaction system).

Considering the use of morphological analysis for information architecture, even when I reach the end of a word-discovery process with a table full of isolated words categorized by frequency, these words do not lose their connection to the real situation, their journey context, or their relationship with the informational ecosystem. This is, according to these authors, impossible from a linguistic philosophy perspective.

What I’m trying to convey with this complex explanation is that we shouldn’t worry about a possible “loss of context” in the information we analyze during our research, because any product meant for consumption can also be transformed into an ideological sign. Everything that is ideological has a semiotic value and references something outside of itself.

So, even if words are segmented, placed in a table, and categorized, they do not detach from their reality; the act of categorizing will force these words to return to their proper context because the design process will be influenced by the ideological bias present in it.

Some perspective from Psycholinguistics

In linguistics, Benveniste states that subjectivity arranges itself within language, with enunciation serving as the way subjectivity is expressed within the linguistic system. Chomsky, a leading figure in psycholinguistics, developed the Generative Grammar theory (or generativism), which posits the existence of an innate and universal linguistic faculty in the human being (universal grammar).

In this theory, he explains how our species has a “language tool,” an innate cognitive capacity for language acquisition (a close symbolic system — socially built) by a repertoire already present in a person’s mind (the open-ended, creative symbolic system of meaning), and this repertoire has a structure that organizes concepts like time, space, objects, etc.

The scientific community dedicated to studying language in neuroscience could identify the network of brain areas where this language cognitive capacity functions; three of them are important to understand here:

  1. Broca’s area, in the left frontal lobe: This area is responsible for the production of speech and expressive language.
  2. Wernicke’s area in the left temporal lobe: This area is responsible for the production of written and spoken language.

These two areas are connected through the arcuate fasciculus, which is a bundle of nerve fibers, and the result of this connection is the ability to coordinate the production and comprehension of speech.

  1. The angular gyrus, located near Wernicke’s area in the parietal lobe, assists in processing language, numbers, and memory. This region is crucial for linking the written form of words with their sounds, which we learn to master in our language system.

The closest we get to understanding how users of a product relate to reality is by understanding how they relate to language.

All the ways lead us to language as a window into the human mind in terms of symbolic systems, and to the tongue as a vessel for expressing that system.

What comes next?

Until now, you could see how we can use discourse analysis and morphological analysis as part of the discovery process in product design, with the goal of generating a set of information to build the information architecture of a digital product.

I could also show you some theories that support the use of these research methods in the design process. This theoretical corpus goes from the neurophysiological aspects of an individual to the philosophical and social dimensions.

Looking at language and tongue as a research vector is neither idealistic nor out of the ground. It is about to bring to the design process tools and a repertoire that is proper from the human and social sciences; and Design as a applied social science, can integrate these knowledge brilliantly.

In my next article, I will explain more practically how we can organize our data into structured objects and how to create a structure that serves as the foundation for the informational and navigational architecture of a product, also using the OOUX approach.

References

Below, you will see some references I used to integrate the science of language and OOUX, so you can deep dive into this knowledge yourself:

Articles about OOUX

Books

  • Problemas de linguística General I. Émile Benveniste; Siglo Veintinuno Editores, 17ª edición en español, 1993.
  • Noam Chomsky, 60 años de gramática generativa. Pasado, presente y futuro de la teoría lingüística; prólogo de Laura Kornfield. 1ª edición, Ciudad Autónoma de Buenos Aires: Editorial de la Facultad de Filosofía y Letras Universidade de Buenos Aires, 2016.
  • Marxismo e Filosofia da Linguagem, Mikhail Bakhtin e Valentin Volochínov; Hucitec editora, 16ª edição, 2014.

메타데이터
post_id
da998c77f83a
slug
sciences-of-language-in-product-design-da998c77f83a
url
https://medium.com/@strategycx/sciences-of-language-in-product-design-da998c77f83a
canonical_url
https://medium.com/@strategycx/sciences-of-language-in-product-design-da998c77f83a
author_url
https://medium.com/@strategycx
status
ok
fetched_at
2026-06-20 20:29:01