Main Corpus vs Reference Corpus
In corpus linguistics, we rarely study a corpus in isolation. Meaning emerges through comparison. This is why the distinction between a…
Main Corpus vs Reference Corpus
In corpus linguistics, we rarely study a corpus in isolation. Meaning emerges through comparison. This is why the distinction between a main and a reference corpus sits at the heart of nearly every keyness study, register analysis, or discourse investigation.
Main corpus (also: study, focus, or target corpus) is the body of texts under investigation the language you actually want to describe. Pakistani newspaper editorials, climate-change tweets, Jane Austen’s novels, your PhD interview transcripts: whatever you are analyzing is your main corpus.
Reference corpus (also: comparison or benchmark corpus) is the yardstick. It represents “general” or “normal” use of the same language variety, against which the distinctive features of your main corpus become visible. Treat it as the control group in an experiment.
A worked example close to home. Suppose you are investigating Pakistani English perhaps a corpus of Dawn and The News editorials. Which reference corpus you pick will shape what you find:
- Compare it against the British National Corpus (BNC) and items like load-shedding, lakh, biradari, Inshallah, suo motu, miscreants, the concerned authorities surface as keywords. You are now describing what makes Pakistani English distinct from British English.
- Compare the same corpus against the Corpus of Contemporary American English (COCA) and a slightly different set emerges, because American English itself differs from British English (lift/elevator, petrol/gas, autumn/fall). Your “distinctive” Pakistani items shift accordingly.
- Compare it against the Indian component of the International Corpus of English (ICE-India) and most of those South Asian features disappear from the keyword list, because they are shared across the subcontinent. What remains is what is distinctive to Pakistani English specifically — perhaps Urdu-origin borrowings, local political vocabulary, or particular collocations.
Notice the lesson: the same main corpus produces three different stories depending on the reference. This is why naming and justifying your reference corpus is a methodological obligation, not a footnote.
Other quick illustrations.
- Main: Trump’s campaign speeches → Reference: COCA spoken → reveals populist vocabulary and repetition patterns.
- Main: Shakespeare’s tragedies → Reference: Early Modern English drama corpus → isolates Shakespearean style from period style.
- Main: ESL learner essays (e.g. ICNALE Pakistan subset) → Reference: LOCNESS (native student writing) → reveals over- and underused features in learner English.
- Main: COVID-19 news coverage 2020 → Reference: general news 2015–2019 → reveals pandemic-era discourse shifts.
Why the distinction matters. Raw frequency tells you little. The is the most frequent word in almost every English corpus, and that fact is useless. What is useful is identifying which items behave unusually against a baseline. Reference corpora make that calculation possible.
On which aspects should the final decision rest? When choosing a reference corpus, weigh these seven factors deliberately:
- Research question. This dominates everything else. “Distinct from what?” must have a clear, defensible answer. Pakistani English vs British norms, vs American norms, and vs South Asian norms are three different questions.
- Variety of the language. Match the regional variety unless contrast across varieties is itself the point. Don’t benchmark Pakistani English against the BNC and then claim you have described Pakistani English in absolute terms — you have described it relative to British English.
- Mode. Spoken with spoken, written with written. Comparing transcribed interviews against newspaper prose will produce noise, not insight.
- Genre and register. A corpus of academic articles needs an academic or broad-written reference, not casual conversation. The closer the genre match, the more your keywords reflect content rather than register differences.
- Time period. Diachronic mismatches distort results. The BNC’s 1990s baseline will make any 2020s corpus look full of “new” vocabulary that is simply contemporary, not distinctive.
- Size. The reference should be substantially larger than the main corpus five times is a common rule of thumb so the baseline frequencies are statistically reliable.
- Representativeness and balance. A good reference samples widely across sub-genres, speakers, and sources, rather than being dominated by one publication or domain.
The easiest way to decide a four-step shortcut. When students freeze at this choice, I give them this:
- Write one sentence: “I want to know what is distinctive about compared to .” Fill in both blanks. The first is your main corpus; the second describes your reference.
- Match four labels: variety, mode, genre, and time period. If three or four match, you have a strong reference. If two or fewer match, look again.
- Check size: is the candidate reference roughly five times larger? If yes, proceed. If no, either find a bigger one or interpret results cautiously.
- Name it in your methodology and state honestly what it represents and what it does not.
If you only need a quick, defensible default: for general written English, use the BNC or COCA; for academic English, use BAWE or BNC-academic; for learner English, use LOCNESS; for web English, use EnTenTen. These are widely accepted, so reviewers rarely object provided they fit your research question.
Worst-case scenarios the mistakes that ruin a keyness study. Knowing the failure modes is as important as knowing the procedure:
- Mismatched variety. Studying Pakistani English against COCA and reporting cricket, monsoon, parliament, ministry as “distinctive Pakistani vocabulary.” They are distinctive only against American English; against any South Asian reference they vanish. The finding is an artefact of the reference, not a property of the data.
- Mismatched mode. Comparing a spoken corpus of classroom talk to written news. You will get uh, you know, I mean, like as keywords but those reflect speech, not your topic. Your real findings drown in register noise.
- Mismatched genre. Comparing legal judgments against general English and reporting plaintiff, hereinafter, whereas as keywords. True, but trivial you have rediscovered that legal English is legal English.
- Mismatched time period. A 2024 social-media corpus benchmarked against the BNC will throw up Covid, TikTok, lockdown, Zoom, selfie as “distinctive.” They are distinctive of the era, not of your data.
- Reference too small. A 50,000 word reference for a 200,000 word main corpus inverts the design. Low-frequency items in the reference become statistically unstable, and your keyness scores become unreliable.
- Reference dominated by one source. A “reference corpus” that is 80% one newspaper imports that newspaper’s vocabulary and ideology into your baseline. Whatever your main corpus does not share with that paper looks falsely distinctive.
- Using the main corpus as part of the reference. Including your study texts inside the reference (sometimes by accident with web-scraped data) inflates baseline frequencies and suppresses real keywords. Always check for overlap.
- Treating the reference as “neutral.” No corpus is neutral. The BNC over-represents 1990s middle-class British print media; COCA over-represents American mass media; EnTenTen over-represents searchable web content. Forgetting this leads to claims of universality that the data cannot support.
- Letting the reference drive the question. Choosing the BNC simply because it is convenient, then framing the research question around what the BNC happens to reveal. The question should choose the reference, not the reverse.
A useful diagnostic: if your top 20 keywords look like they describe a register (formality, mode, period) rather than your topic or variety, your reference is wrong.
Only two corpora? Your research question decides which is main and which is reference. Asking “what is distinctive about A?” makes A the main and B the reference. Reverse the question and the roles flip.
More than two corpora? Three sensible designs:
- Pairwise comparisons Pakistani vs British, Pakistani vs American, Pakistani vs Indian English, each answering a different question.
- One central reference, several main corpora medical, legal, and journalistic texts each compared against the BNC. Ideal for genre/register studies.
- Multi-way keyness, supported in #LancsBox and Sketch Engine, which highlights items distinctive to one sub-corpus across several simultaneously.
메타데이터
- post_id
- da8ea32158ce
- slug
- main-corpus-vs-reference-corpus-da8ea32158ce
- url
- https://medium.com/@shoaibtahir410/main-corpus-vs-reference-corpus-da8ea32158ce
- canonical_url
- https://medium.com/@shoaibtahir410/main-corpus-vs-reference-corpus-da8ea32158ce
- author_url
- https://medium.com/@shoaibtahir410
- status
- ok
- fetched_at
- 2026-06-09 15:37:30