Reinforcement Learning for Optimizing Bioprocess Parameters in Cell Cultivation

Authors

  • Hugo Dubois Professor, School of Data Science, Baltic AI Research University, Tallinn, Estonia Author
  • Pierre Lindberg Professor, Institute of Intelligent Systems, Western Europe Data Science University, Madrid, Spain Author

DOI:

https://doi.org/10.5281/

Keywords:

reinforcement learning, bioprocess optimisation, cell cultivation, cell therapy manufacturing, digital twin, safe RL, GMP, bioreactor control

Abstract

The manufacturing of cell-based regenerative therapies -- including iPSC-derived cell products, mesenchymal stromal cell therapies, and CAR-T cell manufacturing -- requires precise control of bioprocess parameters (pH, dissolved oxygen, glucose concentration, temperature, agitation rate, and feeding strategy) to maximise cell yield, viability, and therapeutic quality attributes. Traditional bioprocess optimisation approaches -- design of experiments (DoE), response surface methodology, and model predictive control -- are effective for simple, well-characterised bioprocesses but struggle with the complexity, batch variability, and multi-objective trade-offs characteristic of cell therapy manufacturing. Reinforcement learning (RL) -- which learns optimal control policies through trial-and- error interaction with a process environment -- offers a promising alternative that can adapt to process variability and optimise multiple quality objectives simultaneously. This paper proposes the Bioprocess Reinforcement Learning Optimiser (BRLO) framework, a safe RL methodology for closed-loop bioprocess parameter optimisation in cell cultivation systems, comprising four components: a digital twin bioreactor environment (DTBE) that provides a high-fidelity simulation environment for RL training; a safe reinforcement learning controller (SRLC) that applies constrained policy optimisation to prevent unsafe parameter excursions during learning; a multi-objective reward function (MORF) that balances cell yield, viability, identity, and potency quality attributes; and a transfer learning adapter (TLA) that rapidly adapts trained RL policies to new cell lines or bioreactor configurations. BRLO is evaluated across three cell therapy manufacturing processes: MSC expansion in a 10L stirred tank bioreactor, iPSC-derived cardiomyocyte differentiation in a 2L suspension bioreactor, and CAR-T cell activation and expansion. BRLO achieves a 31.4% improvement in cell yield (SD = 4.8%) over manual operator control and 24.6% improvement over DoE-optimised fixed parameters, while maintaining 98.4% process safety (zero critical parameter excursions across 180 bioreactor runs). Potency-adjusted yield improvement of 28.8% reflects simultaneous quality and yield optimisation. The study contributes the BRLO specification, a Bioprocess Optimisation Score (BOS), and the first validated safe RL framework for GMP-compatible cell therapy bioprocess control.

Author Biographies

  • Hugo Dubois, Professor, School of Data Science, Baltic AI Research University, Tallinn, Estonia

    Professor, School of Data Science, Baltic AI Research University, Tallinn, Estonia

  • Pierre Lindberg, Professor, Institute of Intelligent Systems, Western Europe Data Science University, Madrid, Spain

    Professor, Institute of Intelligent Systems, Western Europe Data Science University, Madrid, Spain

Downloads

Published

2024-03-30

How to Cite

Reinforcement Learning for Optimizing Bioprocess Parameters in Cell Cultivation. (2024). Biotechnology and Regenerative Sciences E: 3117-6445 P: 3117-6453, 1(1), 33-40. https://doi.org/10.5281/