How to apply
To apply, please send your CV and cover letter (summarising the most relevant skills and experience that you have for the position) to https://employmenthero.com/sg/jobs/position/cambridge-cares-lead-research-software-engineer-lv4yc/
Who are we
We are Cambridge CARES, the University of Cambridge's research centre in Singapore, established under the National Research Foundation (NRF) CREATE programme CAM.CREATE. Within Cambridge CARES, BloodCounts! is a collaborative research initiative bringing together A.STAR, Nanyang Technological University, National University of Singapore, National University Health System (NUHS) and SingHealth Hospitals, with the University of Cambridge and the company Sysmex, and many other partners in theUK, Belgium, The Gambia, Ghana, India, and the Netherlands.
BloodCounts! focusses on elevating the value of the complete blood count (known as full blood count in Singapore and the UK), the most common medical test globally, by developing new AI methodologies. In particular, we have access to population scale raw flow cytometry data that is used to generate the summary complete blood count report used by millions of healthcare workers on a daily basis. We have demonstrated that applying AI methods to this data allows for additional insights to be discovered, such as identifying markers for cancers, iron deficiency and causes of infection.
This National Research Foundation AI for Science (AI4S) supported BloodCounts! project will develop multi-modal foundation models of the blood using not only flow cytometry data but also cell imaging data. The foundation models will be developed using anonymised data from NUHS and SingHealth patients, with the aim of developing a population scale foundation model of blood that performs equitable for all ancestries in Singapore. The models developed will be applied to clinical studies for Stroke and Lung Cancer, with the aim to identify new markers that impact clinical treatment.
The BloodCounts!-AI4S project in Singapore will integrate into the wider BloodCounts! consortium, allowing for co-development of models and rapid testing of hypotheses in European, African, South Asian and East Asian populations. Researchers globally will be able to access our models and insights through a secure platform, APIs and open sourced code and models.
The BloodCounts! AI4S Team is led by Profs.Michael Roberts, Carola-Bibiane Schönlieb (Department of Applied Mathematics and Theoretical Physics (DAMTP), University of Cambridge, UK) and Lin Weisi (College of Computing & Data Science, Nanyang Technological University, Singapore). They are supported by Dr Nicholas Gleadall (University of Cambridge), Prof. Iain Bee Huat Tan (SingHealth), Prof. Hui Ji (NUS), Prof.Mickey Koh (St. George's Hospital, London ACTRIS, Singapore), Dr Hwee Kuan Lee (A.STAR), Prof. Parashkev Nachev (University College London Hospitals), Prof. Willem H Ouwehand (University of Cambridge), Dr Suthesh Sivapalaratnam (Barts Health, London) and Dr Chuen Wan Tan (SingHealth) alongside a wider group of collaborators.
The Lead Research Software Engineer will be responsible for the software and data infrastructure deliverables of the project, including the delivery of curated datasets for the clinical studies, all backend infrastructure and pipelines along with the deployment tools. They will work day-to-day with the Lead Machine Learning Researcher to deliver an ecosystem of tools required to perform rapid-iteration high-quality reproducible AI research within which all research will be performed. They will be the main point of contact for all queries around hardware and software.
Requirements:
- Substantial software engineering experience in a research, data-intensive or clinical context, typically gained over 5-10 years post-qualification (postdoctoral research or industrial equivalent).
- Expert proficiency in Python, with strong SQL, and sound general software engineering practice: version control (Git/GitHub or GitLab), automated testing, CI/CD (e.g. GitHub Actions, GitLab CI), code review and documentation.
- Expert understanding of coding assistant tools, e.g. Claude Code, alongside the limitations and guard rails required to ensure high control, safety and maintainability of deployed tools.
- Strong experience with automated software testing, including unit, integration, end-to-end, regression and performance testing, as well as automated validation of data pipelines.
- Experience designing, operating and troubleshooting distributed systems, such as federated learning infrastructure, across heterogeneous environments. This includes participant orchestration, network and authentication failures, inconsistent data or model states, distributed logging and tracing, reproducibility, fault recovery, and performance bottleneck diagnosis.
- Strong networking fundamentals, including HTTP/TLS, DNS, proxies, load balancing, firewalls, private networking and secure communication across cloud, on-premises and containerised environments.
- Demonstrable experience applying security-by-design throughout the software-development lifecycle, including threat modelling, least-privilege access, secure API design, dependency management and vulnerability remediation.
- Ability to rapidly prototype researcher-facing applications and data interfaces using JavaScript or TypeScript and a modern front-end framework.
- Experience designing and operating high-volume databases, such as relational (e.g. PostgreSQL) and, ideally, analytical/columnar stores for large datasets (e.g. DuckDB,ClickHouse, BigQuery).
- Experience working with large-data file formats (e.g. Parquet, Zarr, HDF5 DICOM or similar for imaging).
- Experience building reproducible data curation and ETL/ELT pipelines using workflow orchestration tools (e.g. Airflow, Prefect, Dagster, Nextflow or Snakemake) and, ideally, transformation frameworks such as dbt.
- Experience with containerisation and infrastructure tooling: Docker and an orchestrator (e.g.Kubernetes), HPC container run times (e.g. Apptainer/Singularity), and infrastructure-as-code (e.g. Terraform).
- Experience working across on-premises, cloud-based (e.g. AWS, Azure or GCP) and hybrid infrastructure.
- Demonstrable understanding of cybersecurity and the tools and practices that keep research secure when working with sensitive healthcare data (e.g. secrets management, access control, encryption at rest/in transit, audit logging).
- Experience structuring multi-modal, multi-format data, and awareness of FAIR data principles and their implications for the project.
- Experience in line management, with the ability to set technical direction for a team.
- Skilled in communicating technical requirements to varied stakeholders, including clinical and IT audiences.
- Excellent organisation, prioritisation and time-management skills, and the ability to work in an agile manner in a changing research landscape.
- A positive, collaborative, problem-solving approach.
Desirable criteria are:
- Experience deploying AI/ML models to end users (e.g. model-serving and API frameworks such as FastAPI), and a realistic appreciation of the challenges of deployment.
- Familiarity with Trusted Research Environments (TREs) and the Five Safes framework.
- Familiarity with commonly used Electronic Health Care system (e.g. Cerner, Epic) and medical common datamodels, e.g. OMOP.
- Experience with federated learning or other privacy-preserving / distributed-computation infrastructure (e.g. Flower, NVIDIA FLARE, OpenFL).
- Familiarity with ML frameworks (e.g. PyTorch) and the surrounding workflow tooling for reproducible ML, such as those for experiment tracking (e.g. MLflow, Weights & Biases) and data/model versioning (e.g. DVC).
- Experience in a research leadership position, including mentoring and team leadership.
- A track record of authoring or contributing to technical or software engineering publications.
- Awareness of governance requirements for healthcare data across multiple jurisdictions (e.g. Singapore PDPA/HBRA and UK GDPR).
Responsibilities:
- Responsible for the development of a high-volume database for clinical data storage.
- Responsible for the development of high-quality reproducible data curation pipelines.
- Responsible for delivery of clinical data aligned with OMOP.
- Responsible for approval and dissemination of SOPs describing the infrastructure to clinical and IT stakeholders.
- Lead on the deployment of the federated learning infrastructure, both hardware and software.
- Contribute to National Research Foundation reporting documents as required,
- Serves as a role model and mentor to other researchers in the team.
- Deep understanding of the main research challenges and innovations required for this project.
- Authoring manuscripts for a software engineering audience.
Please note that this post is mainly based in the CREATE Tower at NUS University Town, Singapore.
When is the position available and for how long
The position is available for an immediate start and is offered on a fixed-term contract of two years in the first instance, with the possibility of extension.