Sovereign accessJoin the waitlist →
Bioinformatics

What is bioinformatics? A practical guide for modern life-science R&D

A bioinformatics analysis can produce a convincing figure within a few hours. Six months later, another researcher may struggle to explain how that figure was made.

Roberto Honegger
Roberto Honegger
Founder, OVAITY14 min read
Connected bioinformatics analyses around a central scientific intelligence hub, including genomics, proteomics, statistics, and workflow panels

The topic at a glance

  • Bioinformatics turns biological measurements into evidence, but confidence depends on experimental design, statistical judgment, interpretation, and documentation.
  • Results are hard to reuse when data, code, parameters, and decisions live in disconnected systems.
  • OVAITY is building a Scientific Intelligence Platform that keeps the path from research question to decision traceable and reusable.

The raw data might still exist. The code may be somewhere in a repository. Parameters could be buried in a notebook, while the reasoning behind a key decision remains in a presentation, meeting note or message thread. The result survives, but much of the context around it has disappeared.

This is one of the less visible realities of modern bioinformatics. The field is usually described as the use of computational methods to collect, store and analyse biological data. That definition is accurate, but it does not capture the full work involved. Bioinformatics also depends on experimental design, statistical judgment, biological interpretation and a record of how the analysis led to a conclusion. [1]

As life-science research produces larger and more varied datasets, the connection between these elements matters as much as any individual analysis.

What bioinformatics means in practice

Bioinformatics combines biology and computing with statistics, data engineering and software development. It gives researchers ways to organise biological measurements and use them to investigate scientific questions.

The field covers the acquisition, storage, analysis and communication of biological information, including DNA and amino-acid sequences and the annotations associated with them. [1]

A bioinformatics analysis might help a research team examine:

  • which genes are more active in diseased tissue
  • whether a genetic variant is associated with a phenotype
  • how a treatment changes gene or protein expression
  • which biological pathways appear to be affected
  • whether an observed difference is likely to be real
  • whether a conclusion still holds when the analysis is repeated differently.

The final two questions are easy to underestimate. Producing a result is often simpler than deciding how much confidence to place in it.

A simple example: comparing treatment responders and non-responders

Consider a team investigating why some patients respond to a treatment while others do not.

The researchers collect tissue samples from both groups and perform RNA sequencing. The sequencing instrument produces millions of short reads. Those reads are measurements, not conclusions.

Before the team can interpret the data, it has to pass through several analytical stages. A typical RNA-sequencing workflow may include experimental design, quality control, read alignment, expression quantification, visualisation, differential-expression analysis and functional interpretation. [2]

The team might begin by checking sequencing quality and identifying samples that behave unexpectedly. The reads are then aligned to a reference genome or transcriptome, and the researchers estimate how strongly different genes are expressed in each sample.

They can compare responders with non-responders, rank the genes that differ between the groups and test whether particular biological pathways occur more often near the top of that ranking.

Suppose the analysis suggests that an immune-related pathway is associated with treatment response.

That finding may be interesting, but it remains conditional. It depends on the samples, the experimental design, the normalisation method, the statistical model and the thresholds chosen during the analysis. RNA-sequencing methods do not perform identically under every condition, and comparative evaluations have found meaningful differences between differential-expression approaches. [2,3]

A few influential samples might be driving the result. A different ranking method could place another pathway at the top. A batch effect might partially explain the separation between the groups.

A careful team therefore asks more than whether the result is statistically significant. It asks whether the pattern remains visible under reasonable changes to the analysis and whether another explanation fits the data just as well.

This is where bioinformatics becomes scientific work rather than data processing.

The main types of data used in bioinformatics

Bioinformatics covers a wide range of biological data. The boundaries between the individual fields are becoming less distinct because many projects combine several types of measurements.

Genomics, transcriptomics, proteomics and metabolomics examine different layers of a biological system: DNA, RNA, proteins and metabolites. Integrating these layers can help researchers investigate how molecular changes relate to a phenotype. [4]

Genomics

Genomics examines the complete genetic material of an organism rather than concentrating on a single gene.

Researchers use sequencing and bioinformatics to assemble genomes, identify variants, compare sequences and analyse the structure and function of genomic regions. [5]

The questions vary widely. A cancer-genomics project may look for mutations associated with tumour development or treatment response. A rare-disease study may search for variants that could explain a patient’s phenotype. Population-genetics research may compare variation across groups to study ancestry, selection or disease risk.

In each case, bioinformatics provides the methods required to process large amounts of sequence data and connect the observed variants to existing biological knowledge.

Transcriptomics

Transcriptomics focuses on RNA and gene-expression patterns.

RNA-sequencing data can show which genes are active in a sample and how their expression differs between conditions, tissues or cell types. The analysis commonly involves quality control, alignment or transcript-level mapping, quantification, normalisation and statistical comparison. [2]

Transcriptomic data is often used to study how cells respond to disease, treatment or environmental changes. It can also reveal groups of genes that move together and point towards pathways that may be involved in the observed response.

The result still needs interpretation. A change in RNA abundance does not automatically show that the corresponding protein has changed in the same way, nor does it prove that the gene caused the phenotype being studied.

Proteomics

Proteomics is the large-scale study of the proteins produced in an organism, tissue, cell or other biological context.

A proteome changes across cell types and over time. Proteomic experiments may examine protein abundance, localisation, turnover, post-translational modifications and interactions between proteins. [6]

Proteomic data can therefore provide information that transcriptomic data alone cannot supply. Two samples may contain similar amounts of RNA for a particular gene while differing in protein abundance, phosphorylation state or protein activity.

Bioinformatics helps researchers identify proteins from experimental measurements, compare their abundance and connect the observed changes to pathways, complexes and known molecular interactions.

Metabolomics

Metabolomics examines small molecules, known as metabolites, within cells, tissues, biofluids or organisms.

Metabolites include substrates and products of cellular metabolism. Their concentrations are influenced by genetic, environmental and physiological factors, which can make metabolomic measurements a relatively direct view of the biochemical state of a sample. [7]

The data can be difficult to interpret. A single measured feature may correspond to several possible molecules, and differences in sample handling can alter the observed profile. Bioinformatics methods help with feature detection, compound identification, statistical comparison and pathway-level interpretation.

Structural bioinformatics

Structural bioinformatics studies the three-dimensional form of proteins, DNA, RNA and their molecular complexes.

The Protein Data Bank stores experimental three-dimensional structure data for large biological molecules and makes those structures available for scientific analysis. [8]

Researchers can use structural information to investigate how a protein functions, where another molecule may bind or how a mutation could affect molecular stability. Structural analysis also contributes to protein engineering and parts of the drug-discovery process.

AI-based structure prediction has greatly increased the amount of structural information available. AlphaFold, for example, showed that neural-network methods could predict many protein structures with substantially improved accuracy. [9]

Predicted structures still require scientific judgment. Confidence varies across a model, proteins can adopt several conformations, and a static structure does not by itself explain the behaviour of a molecule inside a cell.

Single-cell and spatial biology

Traditional bulk measurements combine signals from many cells. A result may therefore describe the average of several distinct cell populations.

Single-cell methods measure individual cells and can reveal populations or biological states that would otherwise be hidden. Spatial methods retain information about where gene expression occurs within a tissue. The original spatial-transcriptomics method, for example, combined RNA sequencing with positional barcodes to preserve two-dimensional tissue information. [10]

These datasets are information-rich, but their analysis is demanding. Single-cell measurements can be sparse, technically variable and affected by batch effects. Comparisons of integration methods show that method selection can change how well biological variation is preserved and technical variation is removed. [11]

The analytical choices therefore need to remain visible, especially when cell populations or disease-associated states are defined through several successive processing steps.

What a bioinformatician actually does

A bioinformatician rarely spends the entire day running established tools.

The work often begins before the data exists. A bioinformatician may help define the comparison, evaluate the proposed sample size, plan how batches should be handled or identify a source of bias that would be difficult to correct later.

Once the data is available, the work may include:

  • checking sample quality and metadata
  • writing and reviewing code
  • selecting statistical methods
  • building computational workflows
  • integrating data from several experiments
  • creating figures and reports
  • discussing results with experimental scientists
  • recording assumptions, decisions and limitations.

Biological understanding matters throughout the process.

An analysis can be technically correct and still answer the wrong question. A statistically strong association may have little biological relevance. A plausible interpretation may disappear after an outlier is removed or an alternative threshold is applied.

Good bioinformatics therefore requires judgment. The analyst has to decide which methods fit the experiment, which conclusions the data supports and where the evidence remains weak.

Bioinformatics and computational biology

Bioinformatics and computational biology overlap heavily, and organisations do not use the two terms consistently.

One common distinction is that bioinformatics often concentrates on the collection, processing and analysis of biological data, while computational biology may place more emphasis on computational models and theories used to describe biological systems.

The National Library of Medicine defines computational biology broadly enough to include the manipulation of datasets, the development of computational techniques and the use of models to make biological discoveries or predictions. [12]

A pipeline that identifies differentially expressed genes would usually be described as bioinformatics. A mathematical model of how a signalling network changes over time might be described as computational biology.

In day-to-day research, the distinction is less tidy. The same project may involve data engineering, statistical analysis, simulation and biological modelling. The job title often depends more on the organisation than on a strict division between the fields.

How a bioinformatics workflow develops

A reliable analysis starts with the research question.

This sounds obvious, but some projects begin with a dataset and a broad instruction to “see what is in it.” Exploratory analysis can be useful. It becomes easier to interpret when the team has already agreed on the main comparison, the expected sources of variation and the decision the analysis is intended to inform.

Experimental design comes next.

Sample size, biological replicates, controls, randomisation and batch structure affect what can be learned later. RNA-sequencing guidance stresses that the experimental and analytical strategy should follow the biological question and account for variability and potential sources of bias before sequencing begins. [2]

Some problems cannot be repaired with a better algorithm.

The raw data then needs to be checked. Low-quality measurements, incorrect metadata, contamination, missing values or sample mix-ups can create patterns that look biological. Quality control is part of the scientific argument, not a routine step to complete and forget.

After preprocessing and normalisation, the main analysis begins. Researchers may use statistical tests, clustering, dimensionality reduction, pathway analysis or machine-learning models. The appropriate method depends on the data and the question.

At this point, the most obvious conclusion often receives the most attention. That is understandable, but it creates a risk. An analysis can gradually be adjusted until an expected or attractive pattern appears.

A stronger workflow actively tries to weaken its own conclusion.

The team may vary thresholds, compare statistical methods, change the ranking procedure, remove influential observations and inspect contradictory evidence. If the same interpretation survives these tests, confidence increases. If it does not, the uncertainty should remain visible.

Why bioinformatics results are difficult to reproduce

Reproducibility problems are often described as failures of code or documentation. The underlying issue is broader.

A computational result depends on a connected set of elements:

  • the scientific question
  • the dataset and its version
  • the sample metadata
  • the preprocessing steps
  • the software and computational environment
  • the executed code
  • the parameters and thresholds
  • the intermediate and final outputs
  • the interpretation of the result
  • the decisions made afterwards.

Research on computational reproducibility describes a similar analytical stack, in which data, software, workflows, notebooks and publications depend on metadata that preserves their meaning and relationships. [13]

In many projects, this chain is split across several systems.

Data may sit in object storage or on a shared drive. Code may be in GitHub, an analysis server or a local folder. Experimental details may be recorded in an electronic lab notebook. The final result may appear in a presentation, while the reason for choosing one method over another remains in the memory of the analyst.

This fragmentation becomes visible when a project changes hands.

A new researcher may find several versions of the same analysis. One is labelled “final,” another “final_v2,” and a third was used in the presentation. The outputs are still there, but it is unclear which dataset produced them, what changed between versions or why one result was accepted and another rejected.

Reproducibility therefore requires more than saving a script. A proposed framework for computational research includes code versioning, control of the computational environment, persistent data sharing, documentation and the connection of narrative explanations to the analysis itself. [14]

The problem is not always that the organisation has lost its data. It has lost the relationships between the data, the analysis and the scientific reasoning.

Provenance makes the analytical path visible

Provenance records how a result was produced.

The W3C provenance model describes the entities, activities and responsible agents involved in creating a piece of data or another digital object. That information can be used to assess its quality, reliability or trustworthiness. [15]

For a bioinformatics analysis, provenance might include:

  • which version of a dataset was used
  • which workflow was executed
  • which software versions were present
  • which parameters were supplied
  • which outputs were generated
  • who reviewed the result
  • which later decision relied on it.

These relationships should not have to be reconstructed manually from file names.

Standards such as Workflow Run RO-Crate show how workflow inputs, outputs, code and execution information can be packaged with structured provenance. [16]

This does not remove the need for scientific interpretation. It gives that interpretation a traceable foundation.

Where artificial intelligence fits

AI already supports several parts of bioinformatics.

Protein-structure prediction is the most visible example, but researchers are also evaluating large language models for bioinformatics code generation. BioCoder, for instance, was created as a benchmark containing bioinformatics-specific programming problems for assessing model performance in this domain. [9,17]

More autonomous systems can go further.

AutoBA is a research example of an AI agent designed to propose multi-omics analysis plans, generate code, execute that code and perform subsequent analytical steps with limited user input. [18]

These systems can automate several steps between a scientific question and an initial result. They can also make weak analysis easier to produce at scale.

Generated code may run without the researcher fully understanding each operation. An agent may select a method that is technically valid but poorly suited to the experiment. A polished written explanation can make an uncertain result appear more settled than it is.

For that reason, AI-assisted bioinformatics needs a visible analytical record.

Researchers should be able to inspect what code was executed, which data was accessed, which parameters were used and how the output led to a conclusion. They also need to know where human review occurred, which limitations were identified and which decisions were approved.

Autonomy without traceability is a poor fit for scientific work.

Bioinformatics is also a knowledge-management problem

Bioinformatics is usually discussed as an analytical discipline. Inside a research organisation, it is also a problem of continuity.

An analysis may be correct and well documented when it is performed. Months later, the wider project context can still be difficult to reconstruct. This is especially common when people leave, projects are paused or a team revisits an old result for a new study.

The useful knowledge is rarely contained in one file. It sits across experiments, sample metadata, datasets, code, figures, discussions and decisions.

Preserving those connections changes what a team can do with its previous work. Researchers can compare analyses across projects, understand why an approach failed and avoid repeating work that has already been done. AI systems can retrieve earlier evidence without treating every isolated document as equally reliable.

The FAIR principles provide a useful foundation for this kind of reuse. They call for research objects to be findable, accessible, interoperable and reusable. The principles apply to data as well as the algorithms, tools and workflows that produced it. [19]

Accessible does not mean publicly available without restriction. Sensitive data, intellectual property and confidential research still require permissions and governance. The goal is to make appropriate knowledge retrievable and usable by authorised people and systems without stripping away its ownership or context.

Diagram of a bioinformatics workflow connecting data types, analysis steps, provenance and decisions, reusable knowledge, and the wider life-science community
From raw biological data to reusable scientific knowledge: analysis only becomes durable when provenance, decisions, and context stay connected.

From individual analyses to reusable scientific knowledge

This is the problem we are working on at OVAITY.

We are building a Scientific Intelligence Platform for life-science R&D that connects projects, experiments, datasets, analyses, code and scientific decisions in one governed environment. A result should remain understandable after the person who produced it has moved on to something else.

Over time, scientific knowledge should also become accessible and actionable beyond a single team or organisation.

Where permissions, confidentiality and intellectual-property rules allow, researchers should be able to discover what has already been tested, inspect the evidence behind it and build on it rather than starting again. This does not require every research dataset to become public. It requires ways to share useful knowledge while preserving its provenance, ownership and context.

The useful unit of bioinformatics is not the final plot.

It is the traceable path from the research question to the experiment, the data, the analysis and the decision that followed. Making that path reusable across the life-science community would allow each experiment to contribute to more than the project in which it was first performed.

If your team is exploring AI-native research workflows but struggling with context, traceability, or project memory, we’d love to learn from you.

References

Sources

  1. National Human Genome Research Institute

    Bioinformatics
  2. Conesa, A., Madrigal, P., Tarazona, S. et al.

    A survey of best practices for RNA-seq data analysis

    Genome Biology · 2016

  3. Rapaport, F., Khanin, R., Liang, Y. et al.

    Comprehensive evaluation of differential gene expression analysis methods for RNA-seq data

    Genome Biology · 2013

  4. EMBL–European Bioinformatics Institute

    What is functional genomics?
  5. EMBL–European Bioinformatics Institute

    What is genomics?
  6. EMBL–European Bioinformatics Institute

    What is proteomics?
  7. EMBL–European Bioinformatics Institute

    What is metabolomics?
  8. Jumper, J., Evans, R., Pritzel, A. et al.

    Highly accurate protein structure prediction with AlphaFold

    Nature · 2021

  9. Ståhl, P. L., Salmén, F., Vickovic, S. et al.

    Visualization and analysis of gene expression in tissue sections by spatial transcriptomics

    Science · 2016

  10. Luecken, M. D., Büttner, M., Chaichoompu, K. et al.

    Benchmarking atlas-level data integration in single-cell genomics

    Nature Methods · 2022

  11. National Library of Medicine

    Computational Biology
  12. Leipzig, J., Nüst, D., Hoyt, C. T., Ram, K. and Greenberg, J.

    The role of metadata in reproducible computational research

    Patterns · 2021

  13. Ziemann, M., Poulain, P. and Bora, A.

    The five pillars of computational reproducibility: bioinformatics and beyond

    Briefings in Bioinformatics · 2023

  14. World Wide Web Consortium

    PROV-DM: The PROV Data Model

    W3C Recommendation · 2013

  15. Leo, S., Crusoe, M. R., Rodríguez-Navas, L. et al.

    Recording provenance of workflow runs with RO-Crate

    PLOS ONE · 2024

  16. Tang, X., Tran, A., Tan, J. et al.

    BioCoder: a benchmark for bioinformatics code generation with large language models

    Bioinformatics · 2024

  17. Zhou, J. et al.

    An AI Agent for Fully Automated Multi-Omic Analyses

    Advanced Science · 2024

  18. Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J. et al.

    The FAIR Guiding Principles for scientific data management and stewardship

    Scientific Data · 2016