← Back to list

Profile: Matthew Purver — No Mic Podcast Scribed By Facelesslingjutsu

Matthew Purver is a Professor of Computational Linguistics at Queen Mary University of London, where he leads the Computational…

Journal of Landing Across Linguistic Foreground · 2026-01-27 13:27 · 0 claps · 19.3 min read
#linguistics #language #corpus-linguistics #corpus-intelligence #data-analysis
Open on Medium ↗
Wiki topics: LNG · Linguistics & Language EDU · Education & Learning 🎵 · Music & Audio

Profile: Matthew Purver — No Mic Podcast Scribed By Facelesslingjutsu

Matthew Purver is a Professor of Computational Linguistics at Queen Mary University of London, where he leads the Computational Linguistics Laboratory. He is primarily known for his extensive research into the semantics and pragmatics of conversational language, focusing on how technology can process and “make sense” of human dialogue, including face-to-face interactions and social media. His work spans several key areas of Natural Language Processing (NLP), such as incremental language understanding, the analysis of miscommunication (like clarification requests and self-repairs), and the application of computational models to mental health monitoring and news media analysis. Beyond his academic role, he is a Turing Fellow at the Alan Turing Institute and co-founded Chatterbox Labs, a company specializing in AI and natural language understanding.

While specific details regarding his personal upbringing and childhood remain relatively private, Matthew Purver’s early academic life established the foundation for his career in linguistics. He attended the University of Cambridge, where he earned an MPhil in Computer Speech and Language Processing, followed by a PhD in Computer Science from King’s College London in 2004. His doctoral research focused on the theory and practice of incremental dialogue processing, specifically how systems can understand and respond to human speech in real-time. This early period was marked by a shift from traditional computer science toward the intersection of logic, linguistics, and cognitive science, setting the stage for his later prominence in the field of computational semantics.

  • On the means for clarification in dialogue, 2001 — This early paper investigates methods for achieving clarification within dialogue systems.
  • SCoRE: A tool for searching the BNC, 2001 — This technical report introduces a tool designed for searching the British National Corpus.
  • Processing unknown words in a dialogue system, 2002 — This work addresses the challenge of handling out-of-vocabulary words in spoken dialogue systems.
  • Tonal active control in production on a large turbo-prop aircraft, 2002 — This conference paper deals with active noise control in an aviation context, a more technical application.
  • Answering clarification questions, 2003 — This paper focuses on how dialogue systems can effectively respond to user requests for clarification.
  • Experimenting with clarification in dialogue, 2003 — This cognitive science paper presents experimental work on the process of clarification in human-machine dialogue.
  • Incremental generation by incremental parsing, 2003 — This paper explores a method for generating language output step-by-step, aligned with incremental parsing.
  • Incremental generation by incremental parsing: Tactical generation in Dynamic Syntax, 2003 — This workshop paper applies incremental parsing to tactical language generation within the Dynamic Syntax framework.
  • The theory and use of clarification requests in dialogue, 2004 — This PhD thesis provides a comprehensive theoretical and practical examination of clarification requests in dialogue.
  • Clarie: The clarification engine, 2004 — This workshop paper introduces “Clarie,” an engine designed to handle clarification in dialogue.
  • Clarifying noun phrase semantics, 2004 — This journal article explores methods for resolving ambiguity in noun phrase semantics within dialogue.
  • Context-based incremental generation for dialogue, 2004 — This paper presents an approach to incremental language generation that is sensitive to dialogue context.
  • Incrementality and alignment in shared utterances, 2004 — This workshop paper discusses incremental processing and alignment in collaboratively constructed utterances.
  • Incremental parsing, or incremental grammar?, 2004 — This workshop paper debates the foundational mechanisms behind incremental language processing.
  • Information State Update: Semantics or Pragmatics?, 2004 — This workshop paper examines whether updating a dialogue system’s information state is a semantic or pragmatic process.
  • Meeting structure annotation: Data and tools, 2005 — This paper introduces data and tools for annotating the structure of multi-party meetings.
  • A Meeting Browser that Learns., 2005 — This symposium paper describes an adaptive meeting browser that improves through learning.
  • A multimodal discourse ontology for meeting understanding, 2005 — This workshop paper proposes an ontology to support the understanding of multimodal meeting discourse.
  • Collaborative and argumentative models of meeting discussions, 2005 — This workshop paper contrasts collaborative and argumentative models for analyzing meeting discussions.
  • Combining Confidence Scores with Contextual Features for Robust Multi-Device Dialogue, 2005 — This workshop paper combines confidence scores with context for robust dialogue across devices.
  • Ontology-Based Discourse Understanding for a Persistent Meeting Assistant., 2005 — This symposium paper describes an ontology-based approach for a long-term meeting assistant.
  • Unsupervised topic modelling for multi-party spoken discourse, 2006 — This conference paper applies unsupervised topic modeling to spoken multi-party conversations.
  • CHAT: a conversational helper for automotive tasks., 2006 — This paper introduces CHAT, a dialogue system designed for in-car assistance.
  • Clarie: Handling clarification requests in a dialogue system, 2006 — This journal article details the “Clarie” system and its approach to managing clarification requests.
  • Grammars as parsers: Meeting the dialogue challenge, 2006 — This journal article argues for grammar formalisms that can also function as incremental parsers to handle dialogue phenomena.
  • Robust interpretation in dialogue by combining confidence scores with contextual features., 2006 — This paper presents a method for robust dialogue interpretation by fusing confidence scores with contextual information.
  • Shallow discourse structure for action item detection, 2006 — This workshop paper uses shallow discourse structure to identify action items in meetings.
  • A probabilistic model of meetings that combines words and discourse features, 2008 — This journal article presents a probabilistic model integrating lexical and discourse features for meeting analysis.
  • Detecting and summarizing action items in multi-party dialogue, 2007 — This workshop paper focuses on automatically detecting and summarizing action items from meeting dialogues.
  • Automatic annotation of dialogue structure from simple user interaction, 2007 — This workshop paper explores methods for automatically annotating dialogue structure through user interaction.
  • CHAT to your destination, 2007 — This workshop paper further discusses the development and application of the CHAT in-car dialogue system.
  • Context and well-formedness: the dynamics of ellipsis, 2007 — This journal article examines how context influences the interpretation and grammaticality of elliptical constructions.
  • Detecting Action Items in Multi-party Meetings: Annotation and Initial Experiments, 2007 — This book chapter details annotation schemes and experiments for detecting action items.
  • Disambiguating between generic and referential “you” in dialog, 2007 — This conference paper tackles the disambiguation of the pronoun “you” in dialogue.
  • Integrating conversational move types in the grammar of conversation, 2007 — This book chapter proposes integrating types of conversational moves directly into a formal grammar of dialogue.
  • Meeting adjourned: off-line learning interfaces for automatic meeting understanding, 2008 — This conference paper presents interfaces for offline learning to improve automatic meeting understanding.
  • Modelling and detecting decisions in multi-party dialogue, 2008 — This workshop paper focuses on modeling and identifying decision points within multi-party conversations.
  • Resolving “you” in multi-party dialog, 2007 — This workshop paper addresses the resolution of the second-person pronoun in conversations with more than two participants.
  • The CALO meeting speech recognition and understanding system, 2008 — This workshop paper describes the speech recognition and understanding components of the CALO meeting assistant.
  • Grammar resources for modelling dialogue dynamically, 2009 — This journal article discusses grammar resources suitable for dynamic modeling of dialogue.
  • Cascaded lexicalised classifiers for second-person reference resolution, 2009 — This conference paper uses a cascade of classifiers to resolve references to “you” in dialogue.
  • Split utterances in dialogue: a corpus study, 2009 — This corpus study analyzes utterances that are split across multiple speakers in dialogue.
  • Does structural priming occur in ordinary conversation, 2010 — This paper investigates the occurrence of structural priming in everyday, unscripted conversation.
  • Feedback in conversation as incremental semantic update, 2010 (published 2015) — This paper conceptualizes conversational feedback as a process of incremental semantic updating.
  • IncrementalTurnProcessinginDialogue, 2010 — This work (likely a presentation or poster) focuses on the incremental processing of turns in dialogue.
  • Making a contribution: Processing clarification requests in dialogue, 2010 (published 2011) — This paper examines the cognitive processing involved in responding to clarification requests.
  • Splitting the ‘I’s and crossing the ‘You’s: Context, speech acts and grammar, 2010 — This workshop paper explores how context and speech acts influence grammatical constructs like pronoun use.
  • Structural divergence in dialogue, 2010 — This conference paper investigates instances where conversational participants’ understandings diverge.
  • The CALO meeting assistant system, 2010 — This journal article provides a comprehensive overview of the CALO meeting assistant system.
  • Tracking lexical and syntactic alignment in conversation, 2010 — This cognitive science paper studies how participants align their lexical and syntactic choices during conversation.
  • Dialogue management using scripts and combined confidence scores, 2011 — This patent describes a dialogue management technique that uses scripts and combined confidence metrics.
  • Incrementality and intention-recognition in utterance processing, 2011 — This journal article explores how incremental processing interacts with recognizing a speaker’s intentions.
  • Incremental semantic construction in a dialogue system, 2011 — This conference paper presents a model for building semantic representations incrementally within a dialogue system.
  • On incrementality in dialogue: Evidence from compound contributions, 2011 — This journal article uses compound contributions (utterances split across speakers) as evidence for incremental processing in dialogue.
  • Topic segmentation, 2011 — This book chapter reviews methods and approaches for segmenting discourse into topical units.
  • Finishing each other’s… responding to incomplete contributions in dialogue, 2012 — This cognitive science paper studies how people respond to utterances that are intentionally or unintentionally incomplete.
  • Experimenting with Distant Supervision for Emotion Classification, 2012 — This conference paper applies distant supervision techniques to the task of classifying emotion in text.
  • Helping the medicine go down: Repair and adherence in patient-clinician dialogues, 2012 — This workshop paper analyzes repair mechanisms and their relation to treatment adherence in medical dialogues.
  • Processing self-repairs in an incremental type-theoretic dialogue system, 2012 — This workshop paper describes how an incremental, type-theoretic dialogue system processes self-corrections.
  • Probabilistic grammar induction in an incremental semantic framework, 2012 — This workshop paper presents a method for inducing grammars probabilistically within an incremental semantic framework.
  • Tracking changes in ESG representation: Initial investigations in uk annual reports, 2012 (published 2022) — This workshop paper (published later) begins an analysis of how Environmental, Social, and Governance (ESG) terms appear in corporate reports.
  • Conversational interactions: Capturing dialogue dynamics, 2012 — This book chapter discusses computational methods for capturing the dynamic nature of conversational interaction.
  • Dylan: Parser for dynamic syntax, 2013 — This likely refers to the release or description of “Dylan,” a parser implementing Dynamic Syntax.
  • Incremental grammar induction from child-directed dialogue utterances, 2013 — This workshop paper explores inducing grammatical structures incrementally from child-directed speech.
  • Interactivity and user engagement in art presentation interfaces, 2013 (published 2016) — This book chapter (published later) examines interfaces for presenting art that foster interactivity and engagement.
  • Investigating topic modelling for therapy dialogue analysis, 2013 — This workshop paper applies topic modeling to analyze therapeutic dialogue.
  • Modelling expectation in the self-repair processing of annotat-, um, listeners, 2013 — This paper models how listeners form expectations that are updated during a speaker’s self-repair.
  • Presentation and communication of visual artworks in an interactive virtual environment, 2013 — This poster presents a system for interactive virtual exhibition of visual art.
  • Using conversation topics for predicting therapy outcomes in schizophrenia, 2013 — This journal article explores whether topics discussed in therapy can predict outcomes for schizophrenia patients.
  • A simple baseline for discriminating similar languages, 2014 — This workshop paper proposes a straightforward baseline method for distinguishing between closely related languages.
  • Divergence in Dialogue, 2014 — This journal article formally studies and quantifies moments of misunderstanding or divergence in conversation.
  • Evaluating neural word representations in tensor-based compositional settings, 2014 — This conference paper evaluates different neural word embeddings within tensor-based compositional semantic models.
  • Helping, I mean assessing psychiatric communication: An application of incremental self-repair detection, 2014 — This workshop paper applies self-repair detection to analyze communication in psychiatric settings.
  • Linguistic indicators of severity and progress in online text-based therapy for depression, 2014 — This workshop paper identifies linguistic features correlated with severity and progress in online therapy for depression.
  • Probabilistic type theory for incremental dialogue processing, 2014 — This workshop paper develops a probabilistic type theory to model incremental processing in dialogue.
  • Strongly Incremental Repair Detection, 2014 — This conference paper presents a method for detecting speech repairs in an strongly incremental (word-by-word) manner.
  • Twitter language use reflects psychological differences between democrats and republicans, 2015 — This journal article analyzes Twitter language to uncover psychological differences between US political groups.
  • Conceptualizing Creativity: From Distributional Semantics to Conceptual Spaces., 2015 — This conference paper links distributional semantic models to conceptual spaces as a model for creativity.
  • Ellipsis, 2015 — This handbook chapter provides a comprehensive overview of ellipsis phenomena from a dynamic syntax perspective.
  • From distributional semantics to conceptual spaces: A novel computational method for concept creation, 2015 — This journal article presents a computational model for creating new concepts by moving in distributional semantic space.
  • From distributional semantics to distributional pragmatics, 2015 — This workshop position paper argues for extending distributional semantics methods to model pragmatic phenomena.
  • How natural is argument in natural dialogue?, 2015 — This workshop paper investigates the frequency and nature of argumentative structures in everyday conversation.
  • Information-Theoretic Segmentation of Natural, 2015 — This conference abstract likely discusses information-theoretic approaches to segmenting natural signals (e.g., text, audio).
  • Predicting emotion labels for chinese microblog texts, 2015 — This book chapter focuses on predicting emotions in Chinese microblog (Weibo) posts.
  • Shifting opinions: experiments on agreement and disagreement in dialogue, 2015 — This likely refers to experimental work on how opinions shift during conversational agreement and disagreement.
  • Taking a stance: a corpus study of reported speech, 2015 — This corpus study analyzes how reported speech is used to convey stance or attitude.
  • The use of english colour terms in big data, 2015 — This conference paper uses large datasets to study the usage patterns of English color terms.
  • Better late than Now-or-Never: The case of interactive repair phenomena, 2016 — This commentary article discusses interactive repair phenomena in light of the “Now-or-Never” bottleneck theory of language processing.
  • Entraining IDyOT: timing in the information dynamics of thinking, 2016 — This journal article incorporates temporal dynamics into a computational model of thinking (IDyOT).
  • Modeling metaphor perception with distributional semantics vector space models., 2016 — This workshop paper uses distributional semantic models to simulate how metaphors are perceived and understood.
  • Process based evaluation of computer generated poetry, 2016 — This workshop paper proposes evaluating computer-generated poetry based on the creative process, not just the output.
  • Robust co-occurrence quantification for lexical distributional semantics, 2016 — This student research workshop paper improves methods for quantifying word co-occurrence in distributional semantics.
  • Verb phrase ellipsis using frobenius algebras in categorical compositional distributional semantics, 2016 — This workshop paper uses Frobenius algebras from category theory to model verb phrase ellipsis in compositional distributional semantics.
  • Words, concepts, and the geometry of analogy, 2016 — This preprint explores geometric relationships in vector spaces as a model for analogy.
  • A geometric method for detecting semantic coercion, 2017 — This conference paper presents a geometric approach for identifying when a word’s meaning is contextually coerced.
  • Incongruent headlines: Yet another way to mislead your readers, 2017 — This workshop paper studies how headlines that are incongruent with article content can mislead readers.
  • Opening Up and Closing Down Discussion: Experimenting with Epistemic Statusin Conversation, 2017 — This cognitive science paper experiments with how speakers signal their certainty or knowledge in conversation.
  • Probabilistic record type lattices for incremental reference processing, 2017 — This book chapter uses probabilistic type lattices to model how referents are incrementally resolved in dialogue.
  • Self-repetition in dialogue and monologue, 2018 — This workshop paper compares patterns of self-repetition in dialogue versus monologue.
  • A tensor-based vector space semantics for dynamic syntax, 2018 — This workshop paper integrates tensor-based vector space semantics with the Dynamic Syntax framework.
  • Automatic detection of narrative structure for high-level story representation, 2018 — This symposium paper presents a method for automatically detecting narrative structure in stories.
  • Computational models of miscommunication phenomena, 2018 — This journal article reviews and proposes computational models for various types of communicative breakdown.
  • Exploring semantic incrementality with dynamic syntax and vector space semantics, 2018 — This preprint explores incremental semantic composition using a combination of Dynamic Syntax and vector space models.Applying distributional semantics to enhance classifying emotions in arabic tweets, 2018 — This conference paper uses distributional semantics to improve emotion classification in Arabic tweets.
  • Detecting depression with word-level multimodal fusion., 2019 — This conference paper fuses word-level features from multiple modalities (audio, text) to detect depression from speech.
  • A corpus study on questions, responses and misunderstanding signals in conversations with Alzheimer’s patients, 2019 — This workshop paper analyzes question-response patterns and signals of misunderstanding in dialogues with Alzheimer’s patients.
  • Detecting summary-worthy sentences: the effect of discourse features, 2019 — This conference paper investigates how discourse features can help identify sentences important for summarization.
  • How furiously can colorless green ideas sleep? sentence acceptability in context, 2019 (published 2020) — This journal article studies how context affects the perceived acceptability of semantically anomalous sentences.
  • Re-Representing Metaphor: Modeling metaphor perception using dynamically contextual distributional semantics, 2019 — This journal article presents a dynamic contextual model of distributional semantics to simulate metaphor perception.
  • SemEval-2020 task 3: Graded word similarity in context, 2019 (published 2020) — This workshop paper describes a shared task on determining graded word similarity within specific contexts.
  • Affordance competition in dialogue: the case of syntactic universals, 2020 — This likely argues that syntactic universals emerge from a competition between production and comprehension affordances in dialogue.
  • A review of cross-domain text-to-SQL models, 2020 — This conference paper surveys models that translate natural language questions to SQL queries across different databases.
  • CoSimLex: A resource for evaluating graded word similarity in context, 2020 — This conference paper introduces a dataset for evaluating computational models on graded word similarity in context.
  • Completability vs (in) completeness, 2020 — This journal article discusses the theoretical distinction between utterances that are completable by an interlocutor versus those that are fundamentally incomplete.
  • Creative language generation in a society of engagement and reflection, 2020 — This computational creativity conference paper presents a multi-agent architecture for creative language generation.
  • Investigating the Semantic Wave in Tutorial Dialogues: An Annotation Scheme and Corpus Study on Analogy Components., 2020 — This likely presents a study on how analogies are built and elaborated in tutoring dialogues.
  • Temporal mental health dynamics on social media, 2020 — This preprint analyzes how mental health expressions change over time on social media platforms.
  • Alzheimer’s dementia recognition from spontaneous speech using disfluency and interactional features, 2021 — This journal article uses disfluency and turn-taking features to detect Alzheimer’s dementia from speech.
  • Alzheimer’s dementia recognition using acoustic, lexical, disfluency and speech pause features robust to noisy inputs, 2021 — This preprint focuses on dementia detection methods that are robust to audio noise, using a combination of features.
  • Analysing text-based messages sent between patients and therapists, 2021 — This patent describes a system for analyzing text-based therapeutic communication.
  • Characterising the IETF through the lens of RFC deployment, 2021 — This conference paper analyzes the Internet Engineering Task Force (IETF) by studying the deployment of its technical standards (RFCs).
  • Detecting alzheimer’s disease using interactional and acoustic features from spontaneous speech, 2021 — This conference paper explores the use of conversational interaction features alongside acoustic features for Alzheimer’s detection.
  • EMBEDDIA tools, datasets and challenges: Resources and hackathon contributions, 2021 — This workshop paper overviews resources and challenges from the EMBEDDIA project for less-resourced languages.
  • Evaluation of contextual embeddings on less-resourced languages, 2021 — This preprint evaluates modern contextual embeddings (like BERT) on languages with limited digital resources.
  • Exploring underexplored limitations of cross-domain text-to-sql generalization, 2021 — This preprint investigates specific, underexamined failure modes of text-to-SQL models when applied to new domains.
  • Incremental composition in distributional semantics, 2021 — This journal article presents a formal model for composing distributional word meanings incrementally, aligned with psycholinguistic theories.
  • Investigating cross-lingual training for offensive language detection, 2021 — This journal article explores methods for detecting offensive language across multiple languages using cross-lingual training.
  • Mitigating topic bias when detecting decisions in dialogue, 2021 — This workshop paper addresses the issue of topic bias in automated decision detection within dialogue transcripts.
  • Multi-modal fusion with gating using audio, lexical and disfluency features for Alzheimer’s dementia recognition from spontaneous speech, 2021 — This preprint uses a gating mechanism to fuse multi-modal features for dementia recognition.
  • Natural SQL: Making SQL easier to infer from natural language specifications, 2021 — This preprint introduces “Natural SQL,” a query language designed to be more easily learnable from natural language.
  • Not all comments are equal: Insights into comment moderation from a topic-aware model, 2021 — This conference paper shows that incorporating topic awareness improves models for comment moderation.
  • Rare-class dialogue act tagging for Alzheimer’s disease diagnosis, 2021 — This workshop paper applies dialogue act tagging, particularly for rare act types, to aid in Alzheimer’s disease diagnosis.
  • Towards robustness of text-to-SQL models against synonym substitution, 2021 — This preprint evaluates and improves the robustness of text-to-SQL models when faced with synonym changes in the input question.
  • Zero-shot cross-lingual content filtering: Offensive language and hate speech detection, 2021 — This workshop paper tackles detecting offensive content in multiple languages without task-specific training data for each language.
  • CoRAL: a context-aware Croatian abusive language dataset, 2022 — This conference paper presents a dataset for abusive language detection in Croatian that includes conversational context.
  • How do Encoder-only LMs Predict Closeness and Respect from Thai Conversations?, 2022 (published 2024) — This workshop paper investigates how language models can predict social dimensions like closeness and respect from Thai dialogue.
  • Knowledge informed sustainability detection from short financial texts, 2022 — This workshop paper uses external knowledge to improve the detection of sustainability-related topics in short financial texts.
  • Language and cognition as distributed process interactions, 2022 — This workshop paper argues for a view of language and cognition as emerging from distributed interactions among processes.
  • Measuring and improving compositional generalization in text-to-sql via component alignment, 2022 — This preprint addresses the challenge of compositional generalization in text-to-SQL models by aligning query components.
  • Misspelling semantics in Thai, 2022 — This preprint explores the semantic impact of common misspellings in the Thai language.
  • Re-appraising the Schema Linking for Text-to-SQL, 2022 (published 2023) — See 2023 entry.
  • The web we weave: Untangling the social graph of the IETF, 2022 — This conference paper maps and analyzes the social network structure within the Internet Engineering Task Force.
  • Tracking changes in ESG representation: Initial investigations in uk annual reports, 2022 — See 2012 entry.
  • A self-evaluating architecture for describing data, 2022 — This conference paper presents an AI architecture capable of generating and then evaluating its own descriptions of data.
  • JSI at SemEval-2022 Task 1: CODWOE-Reverse Dictionary: Monolingual and cross-lingual approaches, 2022 — This workshop paper describes a system for the reverse dictionary task (finding a word from its definition) in monolingual and cross-lingual settings.
  • Comparing neural sentence encoders for topic segmentation across domains: not your typical text similarity task, 2023 — This journal article evaluates different neural sentence encoders for the task of segmenting text into topics across various domains.
  • LEDA: a Large-Organization Email-Based Decision-Dialogue-Act Analysis Dataset, 2023 — This conference paper introduces a large dataset of organizational emails annotated for decision-related dialogue acts.
  • Lexicools at SemEval-2023 task 10: Sexism lexicon construction via XAI, 2023 — This workshop paper describes a system that uses Explainable AI (XAI) techniques to build a lexicon of sexist language.
  • Lon-eå at SemEval-2023 Task 11: A Comparison of Activation Functions for Soft and Hard Label Prediction, 2023 — This workshop paper compares activation functions in neural models for predicting both soft and hard labels in a shared task.
  • Multimodal topic segmentation of podcast shows with pre-trained neural encoders, 2023 — This conference paper uses pre-trained neural encoders on audio and text to segment podcasts into topics.
  • Re-appraising the Schema Linking for Text-to-SQL, 2023 — This conference paper critically re-examines and improves the schema linking component, a crucial step in text-to-SQL systems.
  • Tracing linguistic markers of influence in a large online organisation, 2023 — This conference paper analyzes email communication to identify linguistic markers associated with influence within a large organization.
  • Analysis of transfer learning for named entity recognition in south-Slavic languages, 2023 — This workshop paper investigates the effectiveness of transfer learning for named entity recognition in South Slavic languages.
  • Compared to Us, They Are…: An Exploration of Social Biases in English and Italian Language Models Using Prompting and Sentiment Analysis, 2023 — This conference paper uses prompting and sentiment analysis to uncover social biases in English and Italian language models.
  • Errare humanum est: What do RFC Errata say about Internet Standards?, 2023 — This conference paper analyzes errata reports for Internet RFCs to understand common errors in technical standards.
  • Exploring Pre-Trained Neural Audio Representations for Audio Topic Segmentation, 2023 — This conference paper evaluates various pre-trained neural network models for segmenting audio based on topic changes.
  • Lessons learnt from linear text segmentation: a fair comparison of architectural and sentence encoding strategies for successful segmentation, 2023 — This conference paper provides a comparative analysis of different model architectures and sentence encoders for linear text segmentation.
  • Multimodal Machine Learning in Mental Health: A Survey of Data, Algorithms, and Challenges, 2024 — This preprint surveys the field of multimodal machine learning as applied to mental health, covering data, methods, and open challenges.
  • Power and vulnerability: managing sensitive language in organizational communication, 2024 — This journal article explores how power dynamics and vulnerability are managed through language in organizational settings like the IETF.
  • A computational analysis of the dehumanisation of migrants from syria and ukraine in slovene news media, 2024 — This preprint computationally analyzes Slovene news media to compare portrayals of migrants from Syria and Ukraine, focusing on dehumanizing language.
  • A longitudinal multi-modal dataset for dementia monitoring and diagnosis, 2024 — This journal article introduces a longitudinal, multi-modal dataset designed for research into dementia monitoring and diagnosis.
  • Analysing and enhancing clarification strategies for ambiguous references in consumer service interactions, 2024 — This workshop paper studies and proposes improvements for clarification strategies used to resolve ambiguous references in customer service chats.
  • Comparing News Framing of Migration Crises using Zero-Shot Classification, 2024 — This workshop paper uses zero-shot classification to compare how news media in different countries frame migration crises.
  • Denoising labeled data for comment moderation using active learning, 2024 — This conference paper applies active learning to clean noisy labels in datasets used for training comment moderation systems.
  • EquiPrompt: Debiasing Diffusion Models via Iterative Bootstrapping in Chain of Thoughts., 2024 — This preprint proposes a method to reduce biases in text-to-image diffusion models by using a chain-of-thought reasoning process.
  • Recent trends in linear text segmentation: A survey, 2024 — This preprint provides a comprehensive survey of recent advances in segmenting linear text (e.g., documents, transcripts) into coherent blocks.
  • Scaling for fairness? analyzing model size, data composition, and multilinguality in vision-language bias, 2024 (published 2025) — See 2025 entry.
  • When cohesion lies in the embedding space: Embedding-based reference-free metrics for topic segmentation, 2024 — This conference paper proposes new metrics for topic segmentation that are reference-free and based on embedding space cohesion.
  • Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models, 2025 — This conference paper investigates whether multilingual vision-language models propagate or mitigate gender and racial biases across languages.
  • CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models, 2025 — This preprint proposes a reinforcement learning method that operates at test-time to improve the performance of on-device LLMs, guided by context.
  • Clarq-llm: A benchmark for models clarifying and requesting information in task-oriented dialog, 2025 (published 2024) — See 2024 entry.
  • Cost-Effective Attention Mechanisms for Low Resource Settings: Necessity & Sufficiency of Linear Transformations, 2025 (published 2024) — See 2024 entry.
  • Efficient solutions for an intriguing failure of llms: Long context window does not mean llms can analyze long sequences flawlessly, 2025 — This conference paper addresses specific failure modes of LLMs when processing long context windows, despite their technical capability to do so.
  • EquiPrompt: Debiasing Diffusion Models via Iterative Bootstrapping in Chain of Thoughts., 2025 — See 2024 entry.
  • Evaluating and explaining training strategies for zero-shot cross-lingual news sentiment analysis, 2025 (published 2024) — See 2024 entry.
  • FairJudge: MLLM Judging for Social Attributes and Prompt Image Alignment, 2025 — This preprint presents “FairJudge,” a method using Multimodal Large Language Models to assess and align generated images with social attributes specified in prompts.
  • FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification, 2025 — This preprint introduces a new benchmark for verifying facts at a fine-grained level within dialogue contexts.
  • Identifying Social Self in Text: A Machine Learning Study, 2025 — This conference paper applies machine learning to identify expressions related to an individual’s “social self” in text.
  • Improving factuality for dialogue response generation via graph-based knowledge augmentation, 2025 — This preprint enhances the factuality of dialogue responses by augmenting the model with knowledge from structured graphs.
  • Mono-and cross-lingual evaluation of representation language models on less-resourced languages, 2025 (published 2026) — This journal article comprehensively evaluates language representation models on less-resourced languages in both monolingual and cross-lingual setups.
  • Referential ambiguity and clarification requests: comparing human and LLM behaviour, 2025 — This workshop paper compares how humans and Large Language Models handle referential ambiguity and generate clarification requests.
  • Scaling for fairness? analyzing model size, data composition, and multilinguality in vision-language bias, 2025 — This preprint investigates how factors like model scale, training data composition, and multilingual training affect social bias in vision-language models.

메타데이터
post_id
8d888e1b6ca6
slug
profile-matthew-purver-no-mic-podcast-scribed-by-facelesslingjutsu-8d888e1b6ca6
url
https://medium.com/@jolalf/profile-matthew-purver-no-mic-podcast-scribed-by-facelesslingjutsu-8d888e1b6ca6
canonical_url
https://medium.com/@jolalf/profile-matthew-purver-no-mic-podcast-scribed-by-facelesslingjutsu-8d888e1b6ca6
author_url
https://medium.com/@jolalf
status
ok
fetched_at
2026-06-09 15:37:30