Questions before benchmarks
A model or dataset is a means to answer a question. We should be able to explain why the question matters before optimising a score.
Research
We work across four dimensions of trustworthy language technology: linguistic validity; knowledge representation and grounding; integrity under failure or attack; and real-world fit and consequences for communities, institutions, and the material environment.
While our framework is shared, the individual projects span a variety of research questions. Each faculty programme crosses the dimensions differently, with methods and standards matched to the problem.
Language, meaning & representation
Language technology inherits assumptions about structure, meaning, variation, and which languages count as normal. We use linguistic typology, formal semantics, sociolinguistics, multilingual representation analysis, transfer learning, and evaluation to make those assumptions visible.
This layer ranges from semantic parsing and language embeddings to Creole machine translation, tokenisation, and multilingual model analysis. We don’t simply look at a system’s ability to transfer between languages, but also investigate the important cultural and societal questions surrounding it.
Faculty: Johannes Bjerva · Heather Lent · Russa Biswas
Knowledge representation & grounding
Knowledge graphs are first-class objects of study: we investigate how to represent, type, complete, and reason over structured knowledge across languages and domains. We also study how that structure can support language models alongside evidence attribution, scientific fact-checking, hallucination evaluation, and the ways information changes through summaries and reports.
Structured knowledge and unstructured evidence are complementary: explicit graph paths expose relations, while documents and citations preserve context that cannot always be reduced to a graph. Across both, a claim should be traceable, checkable, and open to revision.
Faculty: Russa Biswas · Dustin Wright · Johannes Bjerva
Security, reliability & risk
Trustworthiness breaks in different ways. We study memorisation, embedding inversion, poisoning, privacy leakage, and security across languages alongside factual-consistency metrics, robustness under distribution shift, and uncertainty.
These are not interchangeable scores. Threat models, metrics, disclosure choices, and deployment constraints determine what evidence is needed—and who may bear the risk when a system fails.
Faculty: Johannes Bjerva · Heather Lent · Russa Biswas · Dustin Wright
Communities, domains & consequences
A system's success depends on whose language, task, evidence, values, and costs shape the research design. We work across community-centred and cross-cultural NLP, cultural heritage, science communication, education, finance, and health.
These real-world settings are crucial in exposing missing data, inappropriate tasks, distorted claims, lost epistemic diversity, institutional constraints, and environmental or social costs that laboratory evaluation can miss. Efficient modelling matters here, but it is not by itself a definition of sustainability.
Faculty: Heather Lent · Russa Biswas · Dustin Wright · Johannes Bjerva
Research principles
A model or dataset is a means to answer a question. We should be able to explain why the question matters before optimising a score.
Formal, empirical, multilingual, human-centred, and security work require different assumptions and standards. We make them explicit.
Relevant linguistic, semantic, technical, security, or domain expertise belongs in a project while the design can still change.
Languages, communities, threat models, domains, and material costs belong in the research design—not only in a limitations section.