← Back to list

What is Computational Biology?

Using computers to solve puzzles in nature? This is still a bit tricky for some to grasp.

Ian Lee · 2023-10-17 03:08 · 97 claps · 3.5 min read
#phd #computational-biology #bioinformatics #science #computer-science
Open on Medium ↗
Wiki topics: BIN · Bioinformatics 🔬 · Science · General 🏔️ · Outdoor & Adventure

What is Computational Biology?

Using computers to solve puzzles in nature? This is still a bit tricky for some to grasp.

Introducing others to my research field can be challenging at times. As a Ph.D. student majoring in computational biology, I often describe it as using computers to solve biological problems, but this can still be too abstract for most people. I would like to take this opportunity to explain what computational biology is and what a computational biologist does.

What is the difference between computational biology and bioinformatics?

The most frequently asked question is: what is the difference between computational biology and bioinformatics?

While these terms are often used interchangeably, they do have distinct differences. Bioinformatics primarily focuses on analyzing large datasets, particularly sequence data. On the other hand, computational biology is a broader term that encompasses various fields such as systems simulation, molecular dynamics, structural biology, and metabolites.

(Generated by DALL·E 3)

(Generated by DALL·E 3)

Interestingly, in German, computational biology is referred to as “bioinformatik”!

A number of factors contribute to the confusion between the terms, including the fact that one of the top journals in computational biology is entitled “Bioinformatics” and that in German for example, computer science is referred to as “informatik” and computational biology is referred to as “bioinformatik.” (Robert F. Murphy)

When it comes to skillsets, computational biologists lean more towards computer science, while bioinformatics leans more towards biology. An essential ability of a computational biologist is to frame biomedical problems as computational problems. For example, convert chemical structures into graph theory and DNA sequence assembly with De Bruijn graph.

Computational biology is becoming an essential subject

In my opinion, the current state of computational biology is comparable to molecular biology in the 1960s. Rapid development of new technologies is expected, and some of these will eventually become standardized experimental methods that biomedical professionals should be familiar with.

Significant advancements in research fields are the result of long-term accumulation and waiting for technological breakthroughs. With the emergence of next-generation sequencing and the maturation of this technology, accessing large volumes of high-quality sequencing data has become a fundamental step in research.

Currently, computational methods have permeated every field of biology. It is rare to find a publication in high-impact journals that does not include some form of computational analysis.

Exploring the Realms of Computational Biology Research

In this part, we will discuss various research focuses within computational biology and provide a brief introduction to some of them. While a more detailed discussion will be reserved for future articles, a subsequent article will provide a roadmap for newcomers, particularly those without a computational background.

Broadly speaking, computational biology can be divided into two interconnected sub-domains: the development of methods and tools, and exploratory research. The former often arises from specific technical challenges, such as improving the speed and accuracy of metagenome assembly. Once these robust tools are available, researchers can explore questions related to the species composition within the microbiome and its interaction with the host. Over time, this cycle repeats itself, with the development of new methods addressing the needs posed by newly collected data or new discoveries.

Key areas of focus in computational biology include metagenomics, metaproteomics, systems biology, biological simulation and modeling, cancer genomics, population genomics, functional genomics, epigenomics, biomedical imaging analysis, and computational neuroscience, among others. Below are some domains I am more familiar with:

Structural Biology

This domain has gained recent attention with the introduction of AlphaFold by Google DeepMind. AlphaFold utilizes the transformer generative model architecture to accurately predict 3D protein structures from peptide sequences. It extracts evolutionary features from multiple sequence alignments and predicts the 3D coordinates. While a detailed technical analysis of AlphaFold is beyond the scope of this discussion, it is important to recognize the significant progress made by the Rosetta community. The improvement in precision achieved by AlphaFold in forecasting 3D structures is truly remarkable. Recent advancements in this field also include Cryo-EM based protein structure construction, de novo protein design for specific functionalities, and docking interactions of protein-protein/protein-peptide/protein-ligand.

The field of protein language models, driven by ChatGPT and the rise of large language models, is a notable aspect of structural biology. Prominent models in this field include ESM-2 by Meta, PLM, AminoBert, and Prot-t5.

Single-cell RNA Sequencing (scRNAseq)

scRNAseq distinguishes itself from bulk RNA sequencing by its ability to uncover transcriptome dynamics. Unlike bulk sequencing, scRNAseq captures differentially expressed genes in individual cells, which enhances data resolution but also introduces the challenge of dealing with high-dimensional data. The increasing development of biologically relevant and mathematically robust pipelines for scRNAseq analysis reflects its growing popularity. Leading the way in this technology is 10X genomics, with Seurat and Scanpy being the predominant tools.

Proteome and Metabolome Analysis

Computational mass spectrometry has revolutionized the identification of proteins and metabolites from mass spectra, providing a comprehensive view of protein expression and interactions through cross-link mass spectrometry. The GNPS database, an extensive collection of mass spectrometry data, contains a wealth of metabolite datasets, although 98% of the spectra remain unannotated, earning the term “dark matter of metabolites”.

Conclusion

Hopefully this gives you a little flavor of what computational biology is all about!


메타데이터
post_id
20edf18611a8
slug
what-is-computational-biology-20edf18611a8
url
https://medium.com/@CompXBio/what-is-computational-biology-20edf18611a8
canonical_url
https://medium.com/@CompXBio/what-is-computational-biology-20edf18611a8
author_url
https://medium.com/@CompXBio
status
ok
fetched_at
2026-06-17 19:05:49