Reinforcement Learning for Optimizing Bioprocess Parameters in Cell Cultivation
DOI:
https://doi.org/10.5281/Keywords:
reinforcement learning, bioprocess optimisation, cell cultivation, cell therapy manufacturing, digital twin, safe RL, GMP, bioreactor controlAbstract
The manufacturing of cell-based regenerative therapies -- including iPSC-derived cell products, mesenchymal stromal cell therapies, and CAR-T cell manufacturing -- requires precise control of bioprocess parameters (pH, dissolved oxygen, glucose concentration, temperature, agitation rate, and feeding strategy) to maximise cell yield, viability, and therapeutic quality attributes. Traditional bioprocess optimisation approaches -- design of experiments (DoE), response surface methodology, and model predictive control -- are effective for simple, well-characterised bioprocesses but struggle with the complexity, batch variability, and multi-objective trade-offs characteristic of cell therapy manufacturing. Reinforcement learning (RL) -- which learns optimal control policies through trial-and- error interaction with a process environment -- offers a promising alternative that can adapt to process variability and optimise multiple quality objectives simultaneously. This paper proposes the Bioprocess Reinforcement Learning Optimiser (BRLO) framework, a safe RL methodology for closed-loop bioprocess parameter optimisation in cell cultivation systems, comprising four components: a digital twin bioreactor environment (DTBE) that provides a high-fidelity simulation environment for RL training; a safe reinforcement learning controller (SRLC) that applies constrained policy optimisation to prevent unsafe parameter excursions during learning; a multi-objective reward function (MORF) that balances cell yield, viability, identity, and potency quality attributes; and a transfer learning adapter (TLA) that rapidly adapts trained RL policies to new cell lines or bioreactor configurations. BRLO is evaluated across three cell therapy manufacturing processes: MSC expansion in a 10L stirred tank bioreactor, iPSC-derived cardiomyocyte differentiation in a 2L suspension bioreactor, and CAR-T cell activation and expansion. BRLO achieves a 31.4% improvement in cell yield (SD = 4.8%) over manual operator control and 24.6% improvement over DoE-optimised fixed parameters, while maintaining 98.4% process safety (zero critical parameter excursions across 180 bioreactor runs). Potency-adjusted yield improvement of 28.8% reflects simultaneous quality and yield optimisation. The study contributes the BRLO specification, a Bioprocess Optimisation Score (BOS), and the first validated safe RL framework for GMP-compatible cell therapy bioprocess control.

