← Back to list

Peer Review — Validation in Forensic Text Comparison: Issues and Opportunities

Abstract:

Journal of Landing Across Linguistic Foreground · 2026-03-06 06:26 · 0 claps · 17.8 min read
#case-study #ratio #style #linguistics #authors
Open on Medium ↗
Wiki topics: TLS · Design Tools & Workflow LIT · Literature & Writing LNG · Linguistics & Language 🔬 · Science · General 👗 · Fashion

Peer Review — Validation in Forensic Text Comparison: Issues and Opportunities

Abstract:

It has been argued in forensic science that the empirical validation of a forensic inference system or methodology should be performed by replicating the conditions of the case under investigation and using data relevant to the case. This study demonstrates that the above requirement for validation is also critical in forensic text comparison (FTC); otherwise, the trier-of-fact may be misled for their final decision. Two sets of simulated experiments are performed: one fulfilling the above validation requirement and the other overlooking it, using mismatch in topics as a case study. Likelihood ratios (LRs) are calculated via a Dirichlet-multinomial model, followed by logistic-regression calibration. The derived LRs are assessed by means of the log-likelihood-ratio cost, and they are visualized using Tippett plots. Following the experimental results, this paper also attempts to describe some of the essential research required in FTC by highlighting some central issues and challenges unique to textual evidence. Any deliberations on these issues and challenges will contribute to making a scientifically defensible and demonstrably reliable FTC available

Introduction and Background

The paper “Validation in Forensic Text Comparison: Issues and Opportunities,” published in 2024 in the journal Languages (Volume 9, Issue 2, Article 47), is authored by Shunichi Ishihara, Sonia Kulkarni, Michael Carne, Sabine Ehrhardt, and Andrea Nini. It addresses a critical gap in forensic linguistics, specifically in forensic text comparison (FTC), which involves analyzing texts to determine authorship for legal purposes, such as identifying the writer of threatening messages or disputed documents.

Forensic linguistics has long been used in courtrooms, drawing on expert opinions about linguistic styles (idiolects) to infer authorship. However, the field has faced criticism for lacking rigorous empirical validation, making it vulnerable to challenges under standards like those in the U.S. Daubert ruling or U.K. Forensic Science Regulator codes. The authors argue that FTC must align with broader forensic science principles: using quantitative measurements, statistical models, the likelihood ratio (LR) framework, and empirical validation. Validation here means testing the system’s reliability by replicating real-case conditions and using data relevant to those conditions. Ignoring this can mislead judges or juries (triers-of-fact) about the evidence’s strength.

The paper uses mismatched topics between questioned (unknown authorship) and known (attributed) texts as a case study, as topics influence writing style and are common in real cases (e.g., a ransom note vs. a diary). It demonstrates through experiments that improper validation leads to unreliable LRs, potentially causing miscarriages of justice. The study also highlights unique challenges in textual evidence, like its complexity (influenced by genre, emotion, etc.), and calls for future research to make FTC more defensible.

This work builds on prior studies in authorship attribution (e.g., Grant 2007; Nini 2023) and forensic voice comparison, emphasizing transparency, reproducibility, and bias resistance. At around 24 pages, it’s a dense academic piece, but its implications extend to improving forensic practices globally.

Key Concepts: Forensic Science Principles and LR Framework

The authors outline four pillars for scientific forensic inference, adapted from fields like voice or DNA analysis (Meuwly et al. 2017; Morrison 2014):

Quantitative Measurements: Texts are measured using features like word frequencies, avoiding subjective expert opinions. Statistical Models: These analyze data probabilistically. Likelihood Ratio (LR) Framework: This quantifies evidence strength without encroaching on the court’s ultimate decision (guilt/innocence). Empirical Validation: Systems must be tested under case-like conditions with relevant data.

The LR framework is central. An LR compares two hypotheses: the prosecution’s (Hp, e.g., “The defendant wrote the questioned text”) vs. the defense’s (Hd, e.g., “Someone else did”). It’s calculated as:

LR = P(E | Hp) / P(E | Hd)

Where E is the evidence (text similarities). An LR > 1 supports Hp; < 1 supports Hd; =1 is neutral. The farther from 1, the stronger the support. LRs update prior odds (pre-existing beliefs) to posterior odds via Bayes’ Theorem, but forensic experts only provide the LR, not posteriors, to avoid overstepping (Robertson et al. 2016).

Textual evidence is complex: Texts encode authorship (idiolect), group traits (e.g., gender, socio-economic background), and situational factors (e.g., topic, formality). Topics can skew analyses if mismatched, as seen in authorship challenges like PAN (Kestemont et al. 2020). The paper stresses that mismatches are case-specific, requiring tailored validation.

Validation requirements (Forensic Science Regulator 2021; Morrison 2022):

Requirement 1: Reflect case conditions (e.g., topic mismatch). Requirement 2: Use relevant data (e.g., similar topics, populations).

Failing these can produce misleading LRs, as shown in the experiments.

Database and Experimental Setup

The study uses the Amazon Authorship Verification Corpus (AAVC), a public dataset of 21,347 Amazon product reviews from 3,227 authors, each ~700–800 words (Halvani et al. 2017). Reviews are categorized into 17 topics (e.g., “Beauty,” “Movies and TV”). The corpus controls for length and genre but not for variables like device used or English variety.

To visualize topic differences, documents from the top eight topics are vectorized using BERT (a transformer-based language model) and plotted in 2D via t-SNE (van der Maaten and Hinton 2008). This reveals clusters: “Beauty” and “Movies and TV” are distant (high mismatch), while “Home and Kitchen” and “Electronics” overlap (low mismatch). “Grocery and Gourmet Food” vs. “Cell Phones and Accessories” is intermediate.

Four settings simulate mismatches:

Cross-topic 1: Beauty vs. Movies and TV (high mismatch, assumed case condition). Cross-topic 2: Grocery vs. Cell Phones (medium). Cross-topic 3: Home vs. Electronics (low). Any-topics: Random pairs.

For each, 1,776 same-author (SA) and 1,776 different-author (DA) pairs are generated, partitioned into six batches for cross-validation. Each batch has 296 SA/DA comparisons.

Texts are tokenized (using quanteda in R), treating punctuation as tokens and preserving case. Each document is represented as a bag-of-words vector with the 140 most frequent tokens (mostly function words like “the,” “I,” punctuation — topic-agnostic to minimize bias).

The pipeline for LR calculation:

Score Calculation: Use Dirichlet-multinomial model (Ishihara 2023) to compute similarity and typicality scores. Calibration: Convert scores to LRs via logistic regression (Morrison 2013).

Databases: Test (for performance assessment), Reference (for typicality stats), Calibration (for training logistic regression). Six-fold cross-validation ensures independence.

Two experiments test validation requirements:

Experiment 1: Varies all databases (matching vs. mismatching case condition). Experiment 2: Fixes Test to Cross-topic 1, varies Reference/Calibration (relevant vs. irrelevant data).

Performance is assessed via log-likelihood-ratio cost (C_llr), where <1 indicates useful info; lower is better. C_llr splits into discrimination (C_llr^min) and calibration (C_llr^cal) losses. Tippett plots visualize LR distributions.

Results and Analysis

Experiment 1 (Case Conditions): Validating under matched conditions (Cross-topic 1) yields mean C_llr = 0.781 (worst performance), reflecting high mismatch difficulty. Mismatched validations (Cross-topic 2: 0.656; Cross-topic 3: 0.528; Any-topics: 0.644) overestimate performance, misleading triers-of-fact. Discrimination drives differences (higher C_llr^min for high mismatches), while calibration is stable (low C_llr^cal ~0.04–0.06). This shows ignoring case conditions paints an unrealistically rosy picture.

Experiment 2 (Relevant Data): Using relevant data (all Cross-topic 1) gives C_llr = 0.781. Irrelevant data worsens it: Cross-topic 2 (0.928), Cross-topic 3 (1.290 >1, useless), Any-topics (0.979). Discrimination remains stable (C_llr^min ~0.74), but calibration degrades sharply (C_llr^cal up to 0.554 for Cross-topic 3), with more variability in mismatched cases.

Tippett plots confirm: Matched LRs are conservative (±3 log10 scale), while mismatched ones overestimate, shifting neutral points and amplifying contrary-to-fact LRs (risking false convictions/acquittals).

Results echo prior findings: Calibration is more sensitive to data issues than discrimination (Ishihara 2020; Hughes 2017). Adverse conditions (quantity/quality) hit calibration hardest.

Discussion, Challenges, and Opportunities

The study proves validation must mirror case conditions and use relevant data; otherwise, LRs mislead. Textual evidence’s complexity (idiolect + situational factors) makes mismatches highly variable, demanding case-specific databases.

Key challenges:

Identifying Conditions: What mismatches (e.g., topic, genre, medium) require validation? Topics are vague/hierarchical. Relevant Data/Population: Define “relevant” (e.g., similar topics, demographics). Empirical studies needed for guidelines. Data Quality/Quantity: Real cases limit data; small samples hurt reliability. Bayesian shrinkage could help (Brümmer and Swart 2014). Robust Features/Models: Seek topic-agnostic features (Halvani et al. 2020) or domain-adaptive models (Daumé 2009). Deep learning (Ishihara et al. 2022) or random-variation methods (Kocher and Savoy 2017) show promise. Synthesis/Engineering: Algorithmically select/compile data (Morrison et al. 2012) or generate texts (Brown et al. 2020).

The authors call for community consensus on protocols, learning from other forensics (e.g., voice guidelines, Drygajlo et al. 2015). While AAVC has limitations (e.g., uncontrolled variables, potential fakes), it’s realistic for forensics.

Peer Review

Overall Assessment

Recommendation: Accept with major revisions.

This is a timely, methodologically sound, and conceptually important contribution to forensic text comparison (FTC). The paper convincingly demonstrates — through controlled experiments — that empirical validation of LR-based FTC systems must reflect case-specific conditions (especially topic mismatch) and employ relevant background data. By showing how ignoring these principles leads to misleading performance metrics (particularly calibration failure), the authors highlight a critical vulnerability in current FTC practice. The work bridges forensic science standards (e.g., Forensic Science Regulator 2021; Morrison 2022) with authorship analysis, addressing a long-standing criticism of the field: lack of rigorous, case-relevant validation (Juola 2021; Ainsworth & Juola 2019).

Strengths include clear experimental design, appropriate use of the LR framework, transparent reporting, and thoughtful discussion of FTC-specific challenges. The results are compelling and have direct implications for admissibility and reliability in court. However, several methodological, interpretive, and structural issues require substantial revision before publication. These include over-reliance on a single corpus and feature set, limited generalizability claims, incomplete discussion of alternative approaches, and some presentational inconsistencies.

Strengths

Conceptual Clarity and Relevance

The core argument — that validation must satisfy Requirement 1 (reflect case conditions) and Requirement 2 (use relevant data) — is powerfully illustrated. The topic-mismatch case study is well-chosen: topic variation is ubiquitous in real cases (e.g., threatening letters vs. social-media posts by the same author), yet frequently under-addressed in validation protocols. The experiments neatly contrast “matched” vs. “mismatched” validation, showing that mismatched setups produce overly optimistic C_llr values (Experiment 1) or outright useless/deceptive LRs when irrelevant data are used for reference/calibration (Experiment 2). This directly supports the claim that triers-of-fact can be misled, aligning with broader forensic science concerns about overconfidence in unvalidated methods (PCAST 2016).

Methodological Rigor

Use of the Dirichlet-multinomial model + logistic-regression calibration is defensible and consistent with recent FTC work (Ishihara 2023). Six-fold cross-validation with mutually exclusive Test/Reference/Calibration partitions avoids data leakage. Performance assessment via C_llr (with discrimination/calibration decomposition) and Tippett plots follows best practice (Brümmer & du Preez 2006; Morrison 2013). Visualization of topic distributions via BERT + t-SNE provides intuitive justification for the chosen cross-topic pairings.

Transparency and Reproducibility

The authors provide the corpus link, feature list (Appendix A), partitioning scheme, and even code/data availability statement. This sets a high standard in a field where reproducibility remains patchy.

Forward-Looking Discussion

Section 7 identifies pressing research questions (relevant populations, data quantity/quality trade-offs, robust features, deep-learning alternatives) and calls for community guidelines/protocols. This positions the paper as both diagnostic and agenda-setting.

Major Concerns (Requiring Revision)

Corpus and Generalizability Limitations (Critical)

The reliance on AAVC (Amazon reviews) is acknowledged but understated in impact. All documents are product reviews (~700–800 words, same genre/register), limiting ecological validity for forensic texts (e.g., short threatening messages, suicide notes, forged wills, social-media rants). Topic categories are retailer-defined and partially overlapping/arbitrary (e.g., “Cell Phones” as distinct from “Electronics”). While t-SNE shows separation, real forensic topic mismatches are often subtler or multi-faceted (e.g., shifting from neutral to emotional register within the same broad topic). No control for confounding variables (device, auto-correction, fake reviews, dialectal English). These are noted but not quantified. Result: The finding that topic mismatch degrades performance is convincing within this domain, but extrapolation to broader FTC (especially short/heterogeneous texts) is premature.

Revision required: Explicitly qualify all claims as “within the domain of longer, review-style texts.” Add a dedicated subsection in Discussion quantifying how AAVC differs from typical forensic corpora (length, register, adversarial intent). Suggest (and ideally pilot) replication on a more forensic-like corpus (e.g., Enron email subsets, Reddit threads, or PAN cross-domain datasets with ground-truth authorship).

Feature Set and Model Sensitivity

The 140 most frequent tokens (mostly function words + punctuation) are deliberately topic-agnostic, which is sensible for isolating authorship effects. However: This choice biases toward function-word stylometry, known to be relatively robust to topic but weaker overall than modern embeddings or character n-grams in cross-domain settings (Halvani et al. 2020; Kestemont et al. 2020). No ablation or comparison to alternatives (e.g., BERT embeddings, POS tags, syntax trees, compression-based methods). The Dirichlet-multinomial model excels with discrete counts but may underperform compared to neural baselines on heterogeneous data.

Revision required: Include a brief sensitivity analysis or literature discussion showing how results might change with more powerful/robust representations. Acknowledge that the observed degradation might be partly model-specific; cite work showing topic-robust features (Halvani & Graner 2021) or domain-adaptation techniques (Daumé 2009).

Interpretation of Calibration vs. Discrimination

Experiment 2 shows that irrelevant reference data primarily harms calibration (C_llr^cal rises sharply) while discrimination (C_llr^min) remains stable. This is an important finding, echoing Hughes (2017) and Ishihara (2020). However, the paper does not sufficiently explore why calibration is more fragile. Is it because logistic regression struggles to map mismatched score distributions? Does the Dirichlet model produce inherently poorly calibrated scores under mismatch? Could fully Bayesian approaches (Brümmer & Swart 2014) mitigate this?

Revision required: Expand Section 6.2 with a deeper mechanistic explanation (e.g., score histograms before/after mismatch). Discuss Bayesian shrinkage as a potential mitigation for small/relevant populations.

Statistical Reporting and Effect Sizes

C_llr values are reported as means over six folds, but ranges (max/min) are wide in some conditions (e.g., Figure 10). Confidence intervals or bootstrap estimates would strengthen claims. No formal statistical tests compare conditions (e.g., paired t-test on C_llr across folds). Tippett plots pool all folds — useful for visualization but hides fold-to-fold variability.

Revision required: Add CIs and/or hypothesis tests for key comparisons.

Minor but Cumulative Issues

Some figures (e.g., Figure 3 t-SNE) are hard to interpret without color legends for all points. Table A1 lists tokens but not frequencies — consider adding top-20 with counts. The term “forensic text comparison” is justified but still uncommon; briefly contrast with “authorship attribution/verification.” References to “trier-of-fact” are repeated excessively — streamline.

Minor Revisions

Expand abstract to explicitly state key numerical findings (e.g., “C_llr rises from 0.78 to >1.2 when irrelevant data are used”). In Introduction, clarify that the paper tests system-level validation, not human-expert validation (Mayring 2020 footnote is good but buried). Discussion could engage more with recent critiques of stylometry in forensics (e.g., Cammarota 2024 on probabilistic gaps). Proofread for typos (e.g., “obliges us to do” → awkward phrasing; “cross-validated experiments were performed separately” repetition).

Conclusion

This manuscript makes a strong, empirically grounded case for case-specific, relevant-data validation in FTC — a message the field urgently needs. With the requested major revisions — particularly qualifying generalizability, adding sensitivity analyses, deepening mechanistic discussion, and improving statistical reporting — it would be a landmark contribution suitable for publication.

  • Ainsworth, Janet, and Patrick Juola. 2019. Who wrote this: Modern forensic authorship analysis as a model for valid forensic science. Washington University Law Review 96: 1159–89.
  • Aitken, Colin, and Franco Taroni. 2004. Statistics and the Evaluation of Evidence for Forensic Scientists, 2nd ed. Chichester: John Wiley & Sons.
  • Aitken, Colin, Paul Roberts, and Graham Jackson. 2010. Fundamentals of Probability and Statistical Evidence in Criminal Proceedings: Guidance for Judges, Lawyers, Forensic Scientists and Expert Witnesses. London: Royal Statistical Society. Available online: http://www.rss.org.uk/Images/PDF/influencing-change/rss-fundamentals-probability-statistical-evidence.pdf (accessed on 4 July 2011).
  • Association of Forensic Science Providers. 2009. Standards for the formulation of evaluative forensic science expert opinion. Science & Justice 49: 161–64. [CrossRef]
  • Ballantyne, Kaye, Joanna Bunford, Bryan Found, David Neville, Duncan Taylor, Gerhard Wevers, and Dean Catoggio. 2017. An Introductory Guide to Evaluative Reporting. Available online: https://www.anzpaa.org.au/forensic-science/our-work/projects/evaluative-reporting (accessed on 26 January 2022).
  • Benoit, Kenneth, Kohei Watanabe, Haiyan Wang, Paul Nulty, Adam Obeng, Stefan Müller, and Akitaka Matsuo. 2018. quanteda: An R package for the quantitative analysis of textual data. Journal of Open Source Software 3: 774. [CrossRef]
  • Boenninghoff, Benedikt, Steffen Hessler, Dorothea Kolossa, and Robert Nickel. 2019. Explainable authorship verification in social media via attention-based similarity learning. Paper presented at 2019 IEEE International Conference on Big Data, Los Angeles, CA, USA, December 9–12.
  • Brown, Tom, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems 33: 1877–901.
  • Brümmer, Niko, and Albert Swart. 2014. Bayesian calibration for forensic evidence reporting. Paper presented at Interspeech 2014, Singapore, September 14–18.
  • Brümmer, Niko, and Johan du Preez. 2006. Application-independent evaluation of speaker detection. Computer Speech and Language 20: 230–75. [CrossRef]
  • Coulthard, Malcolm, Alison Johnson, and David Wright. 2017. An Introduction to Forensic Linguistics: Language in Evidence, 2nd ed. Abingdon and Oxon: Routledge.
  • Coulthard, Malcolm, and Alison Johnson. 2010. The Routledge Handbook of Forensic Linguistics. Milton Park, Abingdon and Oxon: Routledge.
  • Daumé, Hal, III. 2009. Frustratingly easy domain adaptation. arXiv arXiv:0907.1815. [CrossRef]
  • Daumé, Hal, III, and Daniel Marcu. 2006. Domain adaptation for statistical classifiers. Journal of Artificial Intelligence Research 26: 101–26. [CrossRef]
  • Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. Paper presented at 17th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, June 9–12.
  • Doddington, George, Walter Liggett, Alvin Martin, Mark Przybocki, and Douglas Reynolds. 1998. SHEEP, GOATS, LAMBS and WOLVES: A statistical analysis of speaker performance in the NIST 1998 speaker recognition evaluation. Paper presented at the 5th International Conference on Spoken Language Processing, Sydney, Australia, November 30–December 4.
  • Drygajlo, Andrzej, Michael Jessen, Sefan Gfroerer, Isolde Wagner, Jos Vermeulen, and Tuija Niemi. 2015. Methodological Guidelines for Best Practice in Forensic Semiautomatic and Automatic Speaker Recognition (3866764421). Available online: http://enfsi.eu/wp-content/uploads/2016/09/guidelines_fasr_and_fsasr_0.pdf (accessed on 28 December 2016).
  • Evett, Ian, Graham Jackson, J. A. Lambert, and S. McCrossan. 2000. The impact of the principles of evidence interpretation on the structure and content of statements. Science & Justice 40: 233–39. [CrossRef]
  • Forensic Science Regulator. 2021. Forensic Science Regulator Codes of Practice and Conduct Development of Evaluative Opinions. Available online: https://assets.publishing.service.gov.uk/government/uploads/system/uploads/attachment_data/file/960051/FSR-C-118_Interpretation_Appendix_Issue_1_002.pdf (accessed on 18 March 2022).
  • Good, Irving. 1991. Weight of evidence and the Bayesian likelihood ratio. In The Use of Statistics in Forensic Science. Edited by Colin Aitken and David Stoney. Chichester: Ellis Horwood, pp. 85–106.
  • Grant, Tim. 2007. Quantifying evidence in forensic authorship analysis. International Journal of Speech, Language and the Law 14: 1–25. [CrossRef]
  • Grant, Tim. 2010. Text messaging forensics: Txt 4n6: Idiolect free authorship analysis? In The Routledge Handbook of Forensic Linguistics. Edited by Malcolm Coulthard and Alison Johnson. Milton Park, Abingdon and Oxon: Routledge, pp. 508–22.
  • Grant, Tim. 2022. The Idea of Progress in Forensic Authorship Analysis. Cambridge: Cambridge University Press.
  • Halvani, Oren, and Lukas Graner. 2021. POSNoise: An effective countermeasure against topic biases in authorship analysis. Paper presented at the 16th International Conference on Availability, Reliability and Security, Vienna, Austria, August 17–20.
  • Halvani, Oren, Christian Winter, and Lukas Graner. 2017. Authorship verification based on compression-models. arXiv arXiv:1706.00516. [CrossRef]
  • Halvani, Oren, Lukas Graner, and Roey Regev. 2020. Cross-Domain Authorship Verification Based on Topic Agnostic Features. Paper presented at CLEF (Working Notes), Thessaloniki, Greece, September 22–25.
  • Hicks, Tacha, Alex Biedermann, Jan de Koeijer, Franco Taroni, Christophe Champod, and Ian Evett. 2017. Reply to Morrison et al. (2016) Refining the relevant population in forensic voice comparison — A response to Hicks et al. ii (2015) The importance of distinguishing information from evidence/observations when formulating propositions. Science & Justice 57: 401–2. [CrossRef]
  • Hughes, Vincent. 2017. Sample size and the multivariate kernel density likelihood ratio: How many speakers are enough? Speech Communication 94: 15–29. [CrossRef]
  • Hughes, Vincent, and Paul Foulkes. 2015. The relevant population in forensic voice comparison: Effects of varying delimitations of social class and age. Speech Communication 66: 218–30. [CrossRef]
  • Ishihara, Shunichi. 2017. Strength of linguistic text evidence: A fused forensic text comparison system. Forensic Science International 278: 184–97. [CrossRef] [PubMed]
  • Ishihara, Shunichi. 2020. The influence of background data size on the performance of a score-based likelihood ratio system: A case of forensic text comparison. Paper presented at the 18th Workshop of the Australasian Language Technology Association, Online, January 14–15.
  • Ishihara, Shunichi. 2021. Score-based likelihood ratios for linguistic text evidence with a bag-of-words model. Forensic Science International 327: 110980. [CrossRef] [PubMed]
  • Ishihara, Shunichi. 2023. Weight of Authorship Evidence with Multiple Categories of Stylometric Features: A Multinomial-Based Discrete Model. Science & Justice 63: 181–99. [CrossRef]
  • Ishihara, Shunichi, and Michael Carne. 2022. Likelihood ratio estimation for authorship text evidence: An empirical comparison of score- and feature-based methods. Forensic Science International 334: 111268. [CrossRef] [PubMed]
  • Ishihara, Shunichi, Satoru Tsuge, Mitsuyuki Inaba, and Wataru Zaitsu. 2022. Estimating the strength of authorship evidence with a deep-learning-based approach. Paper presented at the 20th Annual Workshop of the Australasian Language Technology Association, Adelaide, Australia, December 14–16.
  • Juola, Patrick. 2021. Verifying authorship for forensic purposes: A computational protocol and its validation. Forensic Science International 325: 110824. [CrossRef]
  • Kafadar, Karen, Hal Stern, Maria Cuellar, James Curran, Mark Lancaster, Cedric Neumann, Christopher Saunders, Bruce Weir, and Sandy Zabell. 2019. American Statistical Association Position on Statistical Statements for Forensic Evidence. Available online: https://www.amstat.org/asa/files/pdfs/POL-ForensicScience.pdf (accessed on 5 May 2022).
  • Kestemont, Mike, Enrique Manjavacas, Ilia Markov, Janek Bevendorff, Matti Wiegmann, Efstathios Stamatatos, Martin Potthast, and Benno Stein. 2020. Overview of the cross-domain authorship verification task at PAN 2020. Paper presented at the CLEF 2020 Conference and Labs of the Evaluation Forum, Thessaloniki, Greece, September 9–12.
  • Kestemont, Mike, Enrique Manjavacas, Ilia Markov, Janek Bevendorff, Matti Wiegmann, Efstathios Stamatatos, Martin Potthast, and Benno Stein. 2021. Overview of the cross-domain authorship verification task at PAN 2021. Paper presented at the CLEF 2021 Conference and Labs of the Evaluation Forum, Bucharest, Romania, September 21–24.
  • Kestemont, Mike, Michael Tschuggnall, Efstathios Stamatatos, Walter Daelemans, Günther Specht, Benno Stein, and Martin Potthast. 2018. Overview of the author identification task at PAN-2018: Cross-domain authorship attribution and style change detection. Paper presented at the CLEF 2018 Conference and the Labs of the Evaluation Forum, Avignon, France, September 10–14.
  • Kirchhüebel, Christin, Georgina Brown, and Paul Foulkes. 2023. What does method validation look like for forensic voice comparison by a human expert? Science & Justice 63: 251–57. [CrossRef]
  • Kocher, Mirco, and Jacques Savoy. 2017. A simple and efficient algorithm for authorship verification. Journal of the Association for Information Science and Technology 68: 259–69. [CrossRef]
  • Koppel, Moshe, and Jonathan Schler. 2004. Authorship verification as a one-class classification problem. Paper presented at the 21st International Conference on Machine Learning, Banff, AB, Canada, July 4–8.
  • Koppel, Moshe, Shlomo Argamon, and Anat Rachel Shimoni. 2002. Automatically categorizing written texts by author gender. Literary and Linguistic Computing 17: 401–12. [CrossRef]
  • López-Monroy, Pastor, Manuel Montes-y-Gómez, Hugo Jair Escalante, Luis Villasenor-Pineda, and Efstathios Stamatatos. 2015. Discriminative subprofile-specific representations for author profiling in social media. Knowledge-Based Systems 89: 134–47. [CrossRef]
  • Lynch, Michael, and Ruth McNally. 2003. “Science”, “common sense”, and DNA evidence: A legal controversy about the public understanding of science. Public Understanding of Science 12: 83–103. [CrossRef]
  • Mayring, Philipp. 2020. Qualitative Content Analysis: Theoretical Foundation, Basic Procedures and Software Solution. Klagenfurt: Springer.
  • McMenamin, Gerald. 2001. Style markers in authorship studies. International Journal of Speech, Language and the Law 8: 93–97. [CrossRef]
  • McMenamin, Gerald. 2002. Forensic Linguistics: Advances in Forensic Stylistics. Boca Raton: CRC Press.
  • Menon, Rohith, and Yejin Choi. 2011. Domain independent authorship attribution without domain adaptation. Paper presented at International Conference Recent Advances in Natural Language Processing 2011, Hissar, Bulgaria, September 12–14.
  • Meuwly, Didier, Daniel Ramos, and Rudolf Haraksim. 2017. A guideline for the validation of likelihood ratio methods used for forensic evidence evaluation. Forensic Science International 276: 142–53. [CrossRef]
  • Morrison, Geoffrey. 2011. Measuring the validity and reliability of forensic likelihood-ratio systems. Science & Justice 51: 91–98. [CrossRef]
  • Morrison, Geoffrey. 2013. Tutorial on logistic-regression calibration and fusion: Converting a score to a likelihood ratio. Australian Journal of Forensic Sciences 45: 173–97. [CrossRef]
  • Morrison, Geoffrey. 2014. Distinguishing between forensic science and forensic pseudoscience: Testing of validity and reliability, and approaches to forensic voice comparison. Science & Justice 54: 245–56. [CrossRef]
  • Morrison, Geoffrey. 2018. The impact in forensic voice comparison of lack of calibration and of mismatched conditions between the known-speaker recording and the relevant-population sample recordings. Forensic Science International 283: E1–E7. [CrossRef]
  • Morrison, Geoffrey. 2022. Advancing a paradigm shift in evaluation of forensic evidence: The rise of forensic data science. Forensic Science International: Synergy 5: 100270. [CrossRef]
  • Morrison, Geoffrey, Ewald Enzinger, and Cuiling Zhang. 2016. Refining the relevant population in forensic voice comparison — A response to Hicks et al.ii (2015) The importance of distinguishing information from evidence/observations when formulating propositions. Science & Justice 56: 492–97. [CrossRef]
  • Morrison, Geoffrey, Ewald Enzinger, Vincent Hughes, Michael Jessen, Didier Meuwly, Cedric Neumann, Sigrid Planting, William Thompson, David van der Vloed, Rolf Ypma, et al. 2021. Consensus on validation of forensic voice comparison. Science & Justice 61: 299–309. [CrossRef]
  • Morrison, Geoffrey, Felipe Ochoa, and Tharmarajah Thiruvaran. 2012. Database selection for forensic voice comparison. Paper presented at Odyssey 2012, Singapore, June 25–28.
  • Murthy, Dhiraj, Sawyer Bowman, Alexander Gross, and Marisa McGarry. 2015. Do we Tweet differently from our mobile devices? A study of language differences on mobile and web-based Twitter platforms. Journal of Communication 65: 816–37. [CrossRef]
  • Nini, A. 2023. A Theory of Linguistic Individuality for Authorship Analysis. Cambridge: Cambridge University Press.
  • President’s Council of Advisors on Science and Technology (U.S.). 2016. Forensic Science in Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods. Available online: https://obamawhitehouse.archives.gov/sites/default/files/microsites/ostp/PCAST/pcast_forensic_science_report_final.pdf (accessed on 3 March 2017).
  • Ramos, Daniel, and Joaquin Gonzalez-Rodriguez. 2013. Reliable support: Measuring calibration of likelihood ratios. Forensic Science International 230: 156–69. [CrossRef] [PubMed]
  • Ramos, Daniel, Juan Maroñas, and Jose Almirall. 2021. Improving calibration of forensic glass comparisons by considering uncertainty in feature-based elemental data. Chemometrics and Intelligent Laboratory Systems 217: 104399. [CrossRef]
  • Ramos, Daniel, Rudolf Haraksim, and Didier Meuwly. 2017. Likelihood ratio data to report the validation of a forensic fingerprint evaluation method. Data in Brief 10: 75–92. [CrossRef]
  • Rivera-Soto, Rafael, Olivia Miano, Juanita Ordonez, Barry Chen, Aleem Khan, Marcus Bishop, and Nicholas Andrews. 2021. Learning universal authorship representations. Paper presented at the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, April 17.
  • Robertson, Bernard, Anthony Vignaux, and Charles Berger. 2016. Interpreting Evidence: Evaluating Forensic Science in the Courtroom, 2nd ed. Chichester: Wiley.
  • Stamatatos, Efstathios. 2009. A survey of modern authorship attribution methods. Journal of the American Society for Information Science and Technology 60: 538–56. [CrossRef]
  • van der Maaten, Laurens, and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9: 2579–605.
  • van Leeuwen, David, and Niko Brümmer. 2007. An introduction to application-independent evaluation of speaker recognition systems. In Speaker Classification I: Fundamentals, Features, and Methods. Edited by Christian Müller. Berlin/Heidelberg: Springer, pp. 330–53.
  • Willis, Sheila, Louise McKenna, Sean McDermott, Geraldine O’Donell, Aurélie Barrett, Birgitta Rasmusson, Tobias Höglund, Anders Nordgaard, Charles Berger, Marjan Sjerps, et al. 2015. Strengthening the Evaluation of Forensic Results Across Europe (STEOFRAE): ENFSI Guideline for Evaluative Reporting in Forensic Science. Available online: http://enfsi.eu/wp-content/uploads/2016/09/m1_guideline.pdf (accessed on 28 December 2018).
  • Yager, Neil, and Ted Dunstone. 2008. The biometric menagerie. IEEE Transactions on Pattern Analysis and Machine Intelligence 32: 220–30. [CrossRef]
  • Zhang, Chunxia, Xindong Wu, Zhendong Niu, and Wei Ding. 2014. Authorship identification from unstructured texts. Knowledge-Based Systems 66: 99–111. [CrossRef]

메타데이터
post_id
3c67f8dc8b39
slug
peer-review-validation-in-forensic-text-comparison-issues-and-opportunities-3c67f8dc8b39
url
https://medium.com/@jolalf/peer-review-validation-in-forensic-text-comparison-issues-and-opportunities-3c67f8dc8b39
canonical_url
https://medium.com/@jolalf/peer-review-validation-in-forensic-text-comparison-issues-and-opportunities-3c67f8dc8b39
author_url
https://medium.com/@jolalf
status
ok
fetched_at
2026-07-13 06:23:13