about

resources

events

contribute

republishing

☰

ΑΙhub.org

A deep learning model for identifying disease and risk factor biomarkers

by Linköping University

30 October 2023

Smoking leaves traces in the DNA

To test their model, the LiU researchers compared it with existing models. There are already existing models of the effects of smoking on the body, building on the fact that specific epigenetic changes reflect the effect of smoking on the functioning of the lungs. These traces remain in the DNA long after a person has quit smoking, and this type of model can identify whether someone is a current or former, or has never smoked. Other models can, based on epigenetic markers, estimate the chronological age of an individual, or group individuals according to whether they have a disease or are healthy.

The LiU researchers trained their autoencoder and then used the result to answer three different queries: age determination, smoker status and diagnosing the disease systemic lupus erythematosus, SLE.

“Our models not only enable us to classify individuals based on their epigenetic data. We found that our models can identify previously known epigenetic markers used in other models, but also new markers associated with the condition we’re examining. One example of this is that our model for smoking identifies markers associated with respiratory diseases, such as lung cancer, and DNA damage,” says David Martínez, PhD student at Linköping University.

David Martínez, PhD student. Photo: Thor Balkhed.

The objective of the autoencoder models is to enable compression of extremely complex biological data into a representation of the most relevant characteristics and patterns in data.

“We didn’t steer the model and had no hypotheses based on existing biological knowledge, but let the data speak for itself. When subsequently looking at what was happening in the autoencoder, we saw that data self-organised in a way similar to how it works in the body,” says Mika Gustafsson, professor of translational bioinformatics at Linköping University, who led the study now published in Briefings in Bioinformatics.

In the next step, the researchers can use the most important characteristics found by the autoencoder to create models able to classify a large amount of environment-related, individual-specific factors where there is not enough training data to train more complex AI models.

Interpretable AI models

Certain types of AI are sometimes likened to a black box that provides answers, but humans cannot see how the AI arrived at the answer. Mika Gustafsson and his colleagues however strive to create interpretable AI models that, so to speak, let the researchers peek under the lid of the “black box” to understand what is going on inside.

“We want to be able to understand what the model shows us about the biology behind disease and other conditions. Then we’ll see not only whether someone is ill or not, but, by interpreting data, we’ll also have a chance to learn why,” says Mika Gustafsson.

Mika Gustafsson, professor. Photo: Thor Balkhed.

This research was funded by, among others, the Swedish Research Council, the Wallenberg AI, Autonomous Systems and Software Program (WASP) and the SciLifeLab & Wallenberg National Program for Data-Driven Life Science (DDLS).

Read the research in full

NCAE: data-driven representations using a deep network-coherent DNA methylation autoencoder identify robust disease and risk factor signatures, David Martínez-Enguita, Sanjiv K. Dwivedi, Rebecka Jörnsten and Mika Gustafsson, Briefings in Bioinformatics, (2023).

Linköping University

AIhub is supported by:

RWDS Big Questions: how do we balance innovation and regulation in the world of AI?

Real World Data Science 06 Mar 2026

The panel explores the tensions, trade-offs and practical realities facing policymakers and data scientists alike.

Studying multiplicity: an interview with Prakhar Ganesh

Lucy Smith 05 Mar 2026

What is multiplicity, and what implications does it have for fairness, privacy and interpretability in real-world systems?

Top AI ethics and policy issues of 2025 and what to expect in 2026

AI Matters, Larry Medsker and Ella Scallan 04 Mar 2026

In the latest issue of AI Matters, a publication of ACM SIGAI, Larry Medsker summarised the year in AI ethics and policy, and looked ahead to 2026.

The greatest risk of AI in higher education isn’t cheating – it’s the erosion of learning itself

The Conversation 03 Mar 2026

Will AI hollow out the pipeline of students, researchers and faculty that is the basis of today’s universities?

Forthcoming machine learning and AI seminars: March 2026 edition

Lucy Smith 02 Mar 2026

A list of free-to-attend AI-related seminars that are scheduled to take place between 2 March and 30 April 2026.

monthly digest

AIhub monthly digest: February 2026 – collective decision making, multi-modal learning, and governing the rise of interactive AI

Lucy Smith 27 Feb 2026

Welcome to our monthly digest, where you can catch up with AI research, events and news from the month past.

The Good Robot podcast: the role of designers in AI ethics with Tomasz Hollanek

The Good Robot Podcast 26 Feb 2026

In this episode, Tomasz argues that design is central to AI ethics and explores the role designers should play in shaping ethical AI systems.

Reinforcement learning applied to autonomous vehicles: an interview with Oliver Chang

Lucy Smith 25 Feb 2026

In the third of our interviews with the 2026 AAAI Doctoral Consortium cohort, we hear from Oliver Chang.