Natural Language Processing · Aalborg University

Language is where AI meets the world.

We study how AI represents language and knowledge, connect its outputs to evidence, and test whether systems remain reliable across languages, communities, threats, and real-world constraints.

Aalborg University's Copenhagen campus beside the harbour
Based in CopenhagenCollaborating globally
5faculty with distinct research programmes
18current researchers across career stages
8active funded projects
≈ DKK 38.5mgross current project volume

A shared intellectual framework

Four dimensions of trustworthy language technology

Our faculty pursue distinct programmes across four dimensions: linguistic validity; knowledge representation and grounding; integrity under failure or attack; and real-world fit and consequences for communities, institutions, and the material environment.

01 · Linguistic validity

Language, meaning & variation

We study linguistic structure, semantics, multilinguality, low-resource languages, and what model representations reveal about transfer and failure.

Explore language & representation
02 · Representation & grounding

Knowledge & grounding

We build and reason over knowledge graphs, connecting model outputs to source documents and other traceable evidence.

Explore knowledge & grounding
03 · System integrity

Security, reliability & risk

We study attacks, privacy, uncertainty, robustness, evaluation, and the conditions under which apparently capable systems break.

Explore reliability & risk
04 · Real-world fit

Communities & consequences

We examine how community priorities and institutional settings—including environmental and social costs—change what responsible success means.

Explore communities & consequences

Faculty

Distinct programmes, shared foundations

Each faculty member has an independent intellectual centre of gravity. Collaboration grows from genuine overlap across language, knowledge, evidence, reliability, and consequences—not from forcing every project into one agenda.

Portrait of Johannes Bjerva

Professor of Natural Language Processing · Group Leader

Johannes Bjerva

At AAU 2020 Now

Johannes studies how linguistic structure can make multilingual language technology more capable, secure, private, and trustworthy. He leads AAU NLP and the Copenhagen Section of AAU's Department of Computer Science.

  • Multilingual NLP
  • Linguistic grounding
  • LLM security & privacy
Portrait of Russa Biswas

Tenure-Track Assistant Professor · Co-director, AI:PAGE-Lab

Russa Biswas

At AAU Jun 2024 Now

Russa studies how explicit knowledge graphs and learned representations can strengthen one another, with applications in multilingual factuality and cultural heritage.

  • Knowledge representation
  • Factuality and LLM reasoning
  • Cultural Heritage
Portrait of Dustin Wright

Tenure-Track Assistant Professor

Dustin Wright

At AAU Feb 2026 Now

Dustin studies how AI transforms information, spanning evidence attribution, science communication, factuality and evaluation, epistemic diversity, efficient modelling, and system-level sustainability.

  • Reliable information
  • Science communication
  • Sustainable AI
Portrait of Heather Lent

Assistant Professor

Heather Lent

At AAU Sep 2022 Now

Heather develops community-grounded language technology for Creole and other lower-resourced languages, spanning multilingual evaluation, machine translation, data quality, security, and research ethics.

  • Creole & lower-resourced NLP
  • Multilingual evaluation
  • Responsible security
Portrait of Xikun Jiang

Assistant Professor · Joint affiliation with Formal Methods for Security & Privacy

Xikun Jiang

At AAU May 2026 Now

Xikun develops trustworthy and responsible AI, with a focus on privacy-preserving methods and AI verification for secure and reliable systems.

  • Trustworthy AI
  • Privacy-preserving ML
  • AI verification

Selected programmes

Long-horizon questions, properly resourced

Our projects connect fundamental research with concrete risks and settings, from multilingual model security to educational feedback and cultural heritage.

TRUST 2026–2030

Building TRUST in Text

Linguistically Motivated Language Model Detection

Develops linguistically grounded ways to detect poisoned or tampered language models and scalable safeguards across languages and domains.

  • Model integrity
  • Linguistic fingerprints
  • Explainable AI
PI
Johannes Bjerva
Funder
Independent Research Fund Denmark · Sapere Aude
Amount
DKK 6.19m
Partners
Stockholm University, NVIDIA
Official record
MML-RPL 2022–2026

Multilingual Modelling for Resource-Poor Languages

Explores how systematic linguistic similarities can extend useful language technology to languages with limited labelled data and digital resources.

  • Multilingual & low-resource NLP
  • Linguistic typology
PI
Johannes Bjerva
Funder
Carlsberg Foundation · additional Google support
Amount
DKK 5m + DKK 425k
Official record
FS-AIS 2026–2028

Formal Semantic Methods for AI Safety

Combines formal semantic analysis and neural interpretability to identify and explain rare multilingual failures in generative AI.

  • Formal semantics
  • AI safety
  • Interpretability
PI
Johannes Bjerva
Funder
Coefficient Giving · Technical AI Safety Research
Amount
DKK 2.4m
Official record

How we work

Setting people up for success

Excellent research depends on individual judgement and a healthy collective system: expertise should be visible, feedback routine, expectations clear, contributions recognised, and collaboration easier.

Our research culture
  1. Rigour Questions and evidence before venue or fashion.
  2. Responsibility Independence without isolation.
  3. Trust Open disagreement with respect and support.

Selected work

Research across the programme

A small selection spanning multilingual evaluation, language-model security, structured knowledge, factuality, and reliable information.

EMNLP 2026Conference paper2026

Few-shot Semantic Recovery Attacks on Image Embeddings

Yiyi Chen, Qiongkai Xu, and Johannes Bjerva

Introduces a few-shot black-box method for recovering captions and sensitive attributes from image embeddings, and tests the utility cost of current defences.

  • Security, Privacy & Safety
ACL 2026Conference paper2026

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP

Kushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather Lent, and Miryam de Lhoneux

Audits all non-English Wikipedia editions and shows why data quality—not language coverage alone—must be a first-class concern in multilingual NLP.

  • Multilingual & Lower-Resource NLP
  • Factuality, Reliability & Evaluation

News & milestones

What is happening now

People

Six new researchers join AAU NLP

Since 1 June, Ren Tao, Anahita Baninajjar, Dorielle Lonke, Anna Lackner, Davis Davalos-DeLosh, and Lena Pickartz have joined the group, strengthening our work in linguistics, trustworthy AI, and language-model security.

Publications

Five papers at EMNLP 2026

Four Main Conference papers span multilingual language modelling, vision-language models, embedding security, and epistemic diversity. They include Tao's first paper and the final PhD thesis papers for Marcell and Yiyi. A fifth contribution was accepted to the System Demonstrations track.

Projects

Two new programmes begin

TRUST and Formal Semantic Methods for AI Safety strengthen the group's work on model integrity, formal semantics, multilingual failure analysis, and explainable safeguards.

Work with us

Bring us a hard language problem.

We collaborate across linguistics, semantics, logic, machine learning, AI, security, privacy, HCI, and the domains where language technology has to work in practice.

Start a conversation