The design of new molecules with desired properties is in general a very difficult problem, involving heavy experimentation with high investment of resources and possible negative impact on the environment. The standard approach consists of iteration among formulation, synthesis, and testing cycles, which is a very long and laborious process. In this paper we address the so-called lead optimisation process by developing a new strategy to design experiments and modelling data, namely, the evolutionary model-based design for optimisation (EDO). This approach is developed on a very small set of experimental points, which change in relation to the response of the experimentation according to the principle of evolution and insights gained through statistical models. This new procedure is validated on a data set provided as test environment by Pickett et al. (2011), and the results are analysed and compared to the genetic algorithm optimisation (GAO) as a benchmark. The very good performance of the EDO approach is shown in its capacity to uncover the optimum value using a very limited set of experimental points, avoiding unnecessary experimentation. 1. Introduction Designing molecules with particular properties is usually a long and complex process, in which the nonlinearity of the model, the high number of variables with a leading role, and the categorical structure of these variables can make difficult modelling, experimentation, and analysis. In new drug discovery, a key phase concerns the generation of small molecules modulators of protein function, under the hypothesis that this activity can affect a particular disease state. Current practices rely on the screening of vast libraries of small molecules (often 1-2 million molecules) in order to identify a molecule that specifically inhibits or activates the protein function, commonly known as the lead molecule. The lead molecule interacts with the required target, but it generally lacks the other attributes needed for a drug candidate such as absorption, distribution, metabolism, and excretion (ADME). In order to achieve these attributes, retaining the interaction capacity with the target protein, the lead molecule must be modified. This transformation of the lead molecule is known as lead optimisation. Lead optimisation research involves long synthesis and testing cycles, analyses of the structure-activity relationships (SAR), and quantitative structure activity relationships (QSAR), which are currently the bottleneck of this process [1]. Under traditional approaches these analyses are conducted by
References
[1]
S. D. Pickett, D. V. S. Green, D. L. Hunt, D. A. Pardoe, and I. Hughes, “Automated lead optimization of MMP-12 inhibitors using a genetic algorithm,” ACS Medicinal Chemistry Letters, vol. 2, no. 1, pp. 28–33, 2011.
[2]
A. Z. Dudek, T. Arodz, and J. Gálvez, “Computational methods in developing quantitative structure-activity relationships (QSAR): a review,” Combinatorial Chemistry and High Throughput Screening, vol. 9, no. 3, pp. 213–228, 2006.
[3]
M. Butkiewicz, R. Mueller, D. Selic, E. Dawson, and J. Meiler, “Application of machine learning approaches on quantitative structure activity relationships,” in Proceedings of the IEEE Symposium on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB '09), pp. 255–262, April 2009.
[4]
J. C. Stalring, L. A. Carlsson, P. Almeida, and S. Boyer, “AZOrange—high performance open source machine learning for QSAR modeling in a graphical programming environment,” Journal of Cheminformatics, vol. 3, no. 28, 2011.
[5]
R. Cox, D. V. S. Green, C. N. Luscombe, N. Malcolm, and S. D. Pickett, “QSAR workbench: automating QSAR modeling to drive compound design,” Journal of Computer-Aided Molecular Design, vol. 27, no. 4, pp. 321–336, 2013.
[6]
D. E. Clark, Evolutionary Algorithms in Molecular Design, John Wiley and Sons, 2008.
[7]
T. Solmajer and J. Zupan, “Optimization algorithms and natural computing in drug discovery,” Drug Discovery Today, vol. 1, no. 3, pp. 247–252, 2004.
[8]
M. Forlin, I. Poli, D. De March, N. Packard, G. Gazzola, and R. Serra, “Evolutionary experiments for self-assembling amphiphilic systems,” Chemometrics and Intelligent Laboratory Systems, vol. 90, no. 2, pp. 153–160, 2008.
[9]
D. De March, M. Forlin, D. Slanzi, and I. Poli, “An evolutionary predictive approach to design high dimensional experiments,” in Proceedings of the Articial Life and Evolutionary Computation (WIVACE '08), R. Serra, I. Poli, and M. Villani, Eds., pp. 81–88, World Scientific, 2008.
[10]
R. Baragona, F. Battaglia, and I. Poli, “Evolutionary Statistical Procedures,” in Statistics and Computing, Springer, Berlin, Germany, 2011.
[11]
D. Ferrari, M. Borrotti, and D. De March, “Response improvement in complex experiments by co-information composite likelihood optimisation,” Statistics and Computing, pp. 1–13, 2013.
[12]
M. Borrotti and I. Poli, “Nave Bayes ant colony optimisation for experimental design,” in Synergies of Soft Computing and Statistics for Intelligent Data Analysis, Advances in Intelligent Systems and Computing, R. Kruse, M. R. Berthold, C. Moewes, et al., Eds., vol. 190, pp. 489–497, Springer, 2013.
[13]
L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.
[14]
A. Verikas, A. Gelzinis, and M. Bacauskiene, “Mining data with random forests: a survey and results of new tests,” Pattern Recognition, vol. 44, no. 2, pp. 330–349, 2011.
[15]
A. L. Boulesteix, S. Janitza, J. Kruppa, and I. R. Knig, “Overview of random forest method- ology and practical guidance with emphasis on computational biology and bioinformatics,” Wiley Interdisciplinary Reviews, vol. 2, no. 6, pp. 493–507, 2012.
[16]
A. Liaw and M. Wiener, “Classification and regression by random forest,” R News, vol. 2, no. 3, pp. 18–22, 2002.