Research

Research that de-risks production AI

We publish because privacy-minded European organizations should not have to take our word for it. The benchmark is open, the leaderboard is live, and the papers are citable.

EuroPriv-Bench

National-ID leakage on real documents

Lower is better. A detector can score well on average and still miss the rare tokens that identify a person, which is the whole reason this measures re-identification risk and not just detection.

spaCy
89%
Best open detector
30%
kp-deid (ours)
0%
See the live leaderboard
Where we push

Every line of work has something you can check

EuroPriv-Bench & kp-deid

Our reference de-identifier leaks 0% of national IDs where the best open detector leaks 30%. Open benchmark, live leaderboard, 8 EU languages.

The first unified pan-European de-identification benchmark. It measures re-identification risk alongside detection, on one GDPR-aligned taxonomy, because a high detection score can still miss the rare tokens that actually identify someone.

See the live leaderboard

8

EU languages in EuroPriv-Bench

Synthetic Data Generation

Training data at scale without exposing a single real record.

Our IEEE Access survey covers techniques from prompt engineering through reinforcement learning, and it is the foundation the rest of the programme builds on.

Read it on IEEE Xplore

9M+

Synthetic training examples

$0.14 Cost per 1K samples

Efficient AI Systems

Models small enough to run on ordinary hardware, and on your own compute rather than someone else's.

Quantization, pruning and distillation, taken far enough that a capable model fits where a data-residency rule says it has to sit.

Models on Hugging Face

26M

Smallest model params

Multilingual NLP

Languages the large labs under-serve, treated as first-class.

The TinyFabulist line: three million synthetic English fables (TF1-EN-3M), English-Romanian literary translation resources published in Frontiers in Artificial Intelligence, and TF3-RO-50M, compact Romanian models trained from scratch on synthetic moral microfiction. Plus a comparative study of diacritic restoration.

Datasets on Hugging Face

180+

Citations of our research