This is a project which is currently making use of HPC facilities at Newcastle University. It is active.
For further information about this project, please contact:
This research aims to develop a personalised multimodal deep learning model for healthcare diagnosis and monitoring. The model will integrate text, images, and tabular data to capture the dynamic evolution of health states. Using dementia as the primary study case, the project will combine brain imaging, clinical notes, cognitive assessments, and longitudinal follow-up data to build a model capable of early detection, disease staging, and progression monitoring. The datasets used include authorised, fully anonymised public resources such as ADNI, OASIS, DementiaBank, and the UK Biobank, as well as the Newcastle dementia datasets, which have received full ethical approval from the Faculty of Medical Sciences. Key challenges addressed include effective multimodal data fusion, personalised patient-specific representation learning, context-aware diagnostic reasoning, and the generation of interpretable outputs suitable for clinical use. The project will also establish a comprehensive evaluation framework that assesses interpretability, scalability, and real-world effectiveness, supporting future generalisation to additional healthcare applications such as sarcopenia.
This project will use Python-based workflows for data preprocessing, statistical analysis, numerical modelling, visualisation, and reproducible computational experiments. The main software environment will include Python, Jupyter notebooks or scripts, and open-source scientific computing libraries such as NumPy, pandas, SciPy, scikit-learn, Matplotlib, and PyTorch.
Workloads will be submitted and managed through the Slurm scheduler. The expected processing activities include batch data processing, parallel CPU-based computation, repeated model runs, parameter optimisation, machine-learning model training and evaluation, and analysis of large datasets.
The project will require both multi-core CPU resources and GPU resources. CPU resources will be used for preprocessing, data handling, statistical analysis, and general parallel workloads, while GPU resources will be used to accelerate PyTorch-based machine-learning and deep-learning model training, inference, and hyperparameter experiments.
Software environments, scripts, and job configurations will be managed in a reproducible way using version-controlled code and documented dependencies. No sensitive data, confidential project details, or personally identifiable information are included in this description.