
Research
Our research focuses on efficient, interpretable information extraction using encoder models. We also explore reasoning over knowledge graphs. Our goal is to create explainable technologies to expand human knowledge.

A lightweight router that scores every candidate model for a task without generating text, balancing speed, cost and quality on a per-request basis.

A suite of GLiNER models adapted for biomedicine, distilled from large language models and state of the art at zero-shot biomedical entity recognition.

A family of compact encoder guardrails that flag toxicity, jailbreaks and unsafe responses in real time, including edge variants under 100M parameters.

One model that extracts entities and the relations between them in a single pass, with arbitrary entity and relation types given at inference time.

A bi-encoder GLiNER that separates label and context encoding, recognising thousands of entity types at once with up to 130 times the throughput.

The GLiNER architecture adapted for sequence classification, matching the speed of embedding models while keeping zero-shot and few-shot flexibility.

A small encoder model that handles named entity recognition, question answering, summarisation and relation extraction in a single place.

mT5 transformers trained on 120 million PubChem molecules to translate between SMILES strings and IUPAC chemical names in both directions.

A comparison of transformers, LSTMs and statistical methods for spotting drug-induced liver injury in the literature, and where simple models win.

A review of where machine learning can realistically improve diagnostics, monitoring and risk prediction in healthcare, and what still blocks adoption.
[ Contact form ]
Let's build the future of open information together
Have questions, feedback, or collaboration ideas? Fill out the form - we'll get back to you soon.