Jakub's current research sits at the intersection of health informatics, natural language processing and machine learning, with a focus on turning routinely collected health data into meaningful insight for people living with multiple long-term conditions (MLTCs).
Natural Language Processing for Clinical Research
Jakub focuses on applying natural language processing (NLP) and pretrained language models to solve one of the biggest bottlenecks in observational health research: harmonising inconsistent clinical data across studies.
As health datasets grow in volume and diversity, comparing findings across cohorts requires aligning variables that are labelled and structured differently between studies. This work develops a semantics-aware, language-model-based approach to automate and scale this harmonisation process, making multi-cohort and multi-study clinical research faster and more reproducible.
Multimorbidity and the Lived Experience of Multiple Long-Term Conditions
Jakub investigates how multimorbidity (living with two or more long-term health conditions) is experienced, measured, and reported across health and social care systems in the UK.
Much of my research addresses a core problem in multimorbidity science: routine electronic health records (EHRs) capture diagnoses and contacts, but often fail to capture the lived burden of managing multiple conditions simultaneously. This theme includes work establishing large-scale data resources for this purpose, such as the SAIL MELD-B e-cohorts, which link population-scale linked data to study "burdensomeness" in both adults and children alongside analyses of variation in social care need reporting among GP practices, and methodological work using informatics techniques to cluster and profile the "burden space" of people under 65 with multiple conditions. A related paper examines why the human impact of MLTC frequently gets "lost in translation" between patient experience and what is recorded in routine data.
Machine Learning and Co-Design for Diabetes Management
Jakub applies machine learning to type 1 diabetes management and explores how patients, family carers, and clinicians can be meaningfully involved in designing the ML tools built to support them.
Work here spans both technical and participatory methods: developing machine learning models to predict glucose levels from continuous glucose monitoring (CGM) data, and using computational notebooks as a co-design tool to engage young adults with diabetes, their carers, and clinicians directly in shaping how these predictive models are built and used.