Research

What does it take to earn trust?

We work across four dimensions of trustworthy language technology: linguistic validity; knowledge representation and grounding; integrity under failure or attack; and real-world fit and consequences for communities, institutions, and the material environment.

While our framework is shared, the individual projects span a variety of research questions. Each faculty programme crosses the dimensions differently, with methods and standards matched to the problem.

01

Language, meaning & representation

How do models represent linguistic diversity?

Language technology inherits assumptions about structure, meaning, variation, and which languages count as normal. We use linguistic typology, formal semantics, sociolinguistics, multilingual representation analysis, transfer learning, and evaluation to make those assumptions visible.

This layer ranges from semantic parsing and language embeddings to Creole machine translation, tokenisation, and multilingual model analysis. We don’t simply look at a system’s ability to transfer between languages, but also investigate the important cultural and societal questions surrounding it.

Current questions

  • Which linguistic structures support transfer—and when do they transmit failure?
  • How should low-resource and Creole language technology be evaluated?
  • What do semantics, typology, tokenisation, and sociolinguistic variation reveal?
  • How should languages and benchmarks be sampled for valid multilingual claims?

Faculty: Johannes Bjerva · Heather Lent · Russa Biswas

02

Knowledge representation & grounding

How should AI represent knowledge—and what should support its claims?

Knowledge graphs are first-class objects of study: we investigate how to represent, type, complete, and reason over structured knowledge across languages and domains. We also study how that structure can support language models alongside evidence attribution, scientific fact-checking, hallucination evaluation, and the ways information changes through summaries and reports.

Structured knowledge and unstructured evidence are complementary: explicit graph paths expose relations, while documents and citations preserve context that cannot always be reduced to a graph. Across both, a claim should be traceable, checkable, and open to revision.

Current questions

  • How can graph paths and structured knowledge ground generation?
  • How should systems attribute claims to long, unstructured evidence?
  • Which factuality metrics remain valid under controlled perturbations?
  • How do scientific claims change as they are summarised and communicated?

Faculty: Russa Biswas · Dustin Wright · Johannes Bjerva

03

Security, reliability & risk

What fails under attack, uncertainty, and distribution shift?

Trustworthiness breaks in different ways. We study memorisation, embedding inversion, poisoning, privacy leakage, and security across languages alongside factual-consistency metrics, robustness under distribution shift, and uncertainty.

These are not interchangeable scores. Threat models, metrics, disclosure choices, and deployment constraints determine what evidence is needed—and who may bear the risk when a system fails.

Current questions

  • Which vulnerabilities cross languages, scripts, and model families?
  • How do evaluation metrics behave under controlled change and distribution shift?
  • Can linguistic and formal methods expose rare or hidden failures?
  • What do harm minimisation and responsible disclosure require in NLP?

Faculty: Johannes Bjerva · Heather Lent · Russa Biswas · Dustin Wright

04

Communities, domains & consequences

What changes when language technology meets the world?

A system's success depends on whose language, task, evidence, values, and costs shape the research design. We work across community-centred and cross-cultural NLP, cultural heritage, science communication, education, finance, and health.

These real-world settings are crucial in exposing missing data, inappropriate tasks, distorted claims, lost epistemic diversity, institutional constraints, and environmental or social costs that laboratory evaluation can miss. Efficient modelling matters here, but it is not by itself a definition of sustainability.

Current questions

  • How should community priorities shape NLP tasks and datasets?
  • How can historical collections become connected knowledge without flattening their context?
  • How do model values and information collapse alter which knowledge remains accessible?
  • What would environmentally sustainable AI require beyond efficiency?

Faculty: Heather Lent · Russa Biswas · Dustin Wright · Johannes Bjerva

Research principles

How we decide what counts as progress

Questions before benchmarks

A model or dataset is a means to answer a question. We should be able to explain why the question matters before optimising a score.

Claims matched to evidence

Formal, empirical, multilingual, human-centred, and security work require different assumptions and standards. We make them explicit.

The right expertise, early

Relevant linguistic, semantic, technical, security, or domain expertise belongs in a project while the design can still change.

Context changes the method

Languages, communities, threat models, domains, and material costs belong in the research design—not only in a limitations section.