全部 标题 作者
关键词 摘要

OALib Journal期刊
ISSN: 2333-9721
费用:99美元

查看量下载量

相关文章

更多...
ISRN Agronomy  2013 

Deterministic Imputation in Multienvironment Trials

DOI: 10.1155/2013/978780

Full-Text   Cite this paper   Add to My Lib

Abstract:

This paper proposes five new imputation methods for unbalanced experiments with genotype by-environment interaction ( ). The methods use cross-validation by eigenvector, based on an iterative scheme with the singular value decomposition (SVD) of a matrix. To test the methods, we performed a simulation study using three complete matrices of real data, obtained from interaction trials of peas, cotton, and beans, and introducing lack of balance by randomly deleting in turn 10%, 20%, and 40% of the values in each matrix. The quality of the imputations was evaluated with the additive main effects and multiplicative interaction model (AMMI), using the root mean squared predictive difference (RMSPD) between the genotypes and environmental parameters of the original data set and the set completed by imputation. The proposed methodology does not make any distributional or structural assumptions and does not have any restrictions regarding the pattern or mechanism of missing values. 1. Introduction In plant breeding, multienvironment trials are important for testing the general and specific adaptations of cultivars. A cultivar developed in different environments will show significant fluctuations of performance in production relative to other cultivars. These changes are influenced by different environmental conditions and are referred to as genotype-by-environment interactions, or . Often, experiments are unbalanced because several genotypes are not tested in some environments. A common way of analyzing this type of study is by imputing the missing values and then applying established procedures on the completed data matrix (observed + imputed), for example, the additive main effects and multiplicative interaction model—AMMI—or factorial regression [1–5]. An alternative approximation is to work with the incomplete data using a mixed model with estimates based on maximum likelihood [6]. Several imputation methods have been suggested in the literature to solve the problem of missing values. One of the first was made by Freeman [7], who suggested imputing the missing values iteratively by minimizing the residual sum of squares and doing the analysis on the completed table, reducing the degrees of freedom by the number of missing values. This work was developed by Gauch Jr. and Zobel [8], who made the imputations using the EM algorithm and the AMMI model or EM-AMMI. Some variants of this procedure using multivariate statistics (cluster analysis) were described in Godfrey et al. [9] and Godfrey [10]. Raju [11] proposed the EM-AMMI algorithm by treating the

References

[1]  H. G. Gauch Jr., “Statistical analysis of yield trials by AMMI and GGE,” Crop Science, vol. 46, no. 4, pp. 1488–1500, 2006.
[2]  F. A. van Eeuwijk, M. Malosetti, X. Yin, P. C. Struik, and P. Stam, “Statistical models for genotype by environment data: from conventional ANOVA models to eco-physiological QTL models,” Australian Journal of Agricultural Research, vol. 56, no. 9, pp. 883–894, 2005.
[3]  F. A. van Eeuwijk, M. Malosetti, and M. P. Boer, “Modelling the genetic basis of response curves underlying genotype environment interaction,” in Scale and Complexity in Plant Systems Research: Gene-Plant-Crop Relations, J. H. J. Spiertz, P. C. Struik, and H. H. van Laar, Eds., Wageningen UR Frontier Series, pp. 115–126, Springer, New York, NY, USA, 2007.
[4]  I. Romagosa, J. Voltas, M. Malosetti, and F. A. van Eeuwijk, “Interaccíon Genotipo por Ambiente,” in La Adaptación Ambiente y Los Estreses Abióticos en la Mejora Vegetal, C. M. Avila, S. G. Atienza, M. T. Moreno, and J. I. Cubero, Eds., pp. 107–136, Instituto de Investigación y Formación Agraria y Pesquera; Consejería de Agricultura y Pesca, 2008.
[5]  S. Arciniegas-Alarcón and C. T. S. Dias, “AMMI analysis with imputed data in genotype × environment interaction experiments in cotton,” Pesquisa Agropecuaria Brasileira, vol. 44, no. 11, pp. 1391–1397, 2009.
[6]  M. S. Kang, M. G. Balzarini, and J. L. L. Guerra, “Genotype-by-environment interaction,” in Genetic Analysis of Complex Traits Using SAS, A. M. Saxton, Ed., pp. 69–96, SAS Institute nc, Cary, NC, USA, 2004.
[7]  G. H. Freeman, “Analysis of interactions in incomplete two-ways tables,” Journal of Applied Statistics, vol. 24, no. 1, pp. 47–55, 1975.
[8]  H. G. Gauch Jr. and R. W. Zobel, “Imputing missing yield trial data,” Theoretical and Applied Genetics, vol. 79, no. 6, pp. 753–761, 1990.
[9]  A. J. R. Godfrey, G. R. Wood, S. Ganesalingam, M. A. Nichols, and C. G. Qiao, “Two-stage clustering in genotype-by-environment analyses with missing data,” Journal of Agricultural Science, vol. 139, no. 1, pp. 67–77, 2002.
[10]  A. J. R. Godfrey, Dealing with Sparsity in Genotype × Environment Analysis [Dissertation], Massey University, 2004.
[11]  B. M. K. Raju, “A study on AMMI model and its biplots,” Journal of the Indian Society of Agricultural Statistics, vol. 55, pp. 297–322, 2002.
[12]  J. Mandel, “The analysis of two-way tables with missing values,” Applied Statistics, vol. 42, pp. 85–93, 1993.
[13]  F. A. van Eeuwijk and P. M. Kroonenberg, “Multiplicative models for interaction in three-way ANOVA, with applications to plant breeding,” Biometrics, vol. 54, no. 4, pp. 1315–1333, 1998.
[14]  J. B. Denis, “Ajustements de modèles linéaires et bilinéaires sous contraintes linéaires avec données manquantes,” Revue de Statistique Appliquée, vol. 39, pp. 5–24, 1991.
[15]  T. Caliński, S. Czajka, J. B. Denis, and Z. Kaczmarek, “EM and ALS algorithms applied to estimation of missing data in series of variety trials,” Biuletyn Oceny Odmian, vol. 24-25, pp. 7–31, 1992.
[16]  J. B. Denis and C. P. Baril, “Sophisticated models with numerous missing values: the multiplicative interaction model as an example,” Biuletyn Oceny Odmian, vol. 24-25, pp. 33–45, 1992.
[17]  T. Caliński, S. Czajka, J. B. Denis, and Z. Kaczmarek, “Further study on estimating missing values in series of variety trials,” Biuletyn Oceny Odmian, vol. 30, pp. 7–38, 1999.
[18]  H. P. Piepho, “Methods for estimating missing genotype-location combinations in multilocation trials: an empirical comparison,” Informatik, Biometrie und Epidemiologie in Medizin und Biologie, vol. 26, pp. 335–349, 1995.
[19]  G. C. Bergamo, C. T. S. Dias, and W. J. Krzanowski, “Distribution-free multiple imputation in an interaction matrix through singular value decomposition,” Scientia Agricola, vol. 65, no. 4, pp. 422–427, 2008.
[20]  S. Arciniegas-Alarcón, Data Imputation in Trials with Genotype by Environment Interaction: An Application on Cotton Data [Dissertation], University of S?o Paulo, 2008.
[21]  S. Arciniegas-Alarcón and C. T. S. Dias, “Data imputation in trials with genotype by environment interaction: an application on cotton data,” Revista Brasileira de Biometria, vol. 27, pp. 125–138, 2009.
[22]  S. Arciniegas-Alarcón, M. García-Pe?a, C. T. S. Dias, and W. J. Krzanowski, “An alternative methodology for imputing missing data in trials with genotype-by-environment interaction,” Biometrical Letters, vol. 47, pp. 1–14, 2010.
[23]  B. M. K. Raju and V. K. Bhatia, “Bias in the estimates of sensitivity from incomplete GxE tables,” Journal of the Indian Society of Agricultural Statistics, vol. 56, pp. 177–189, 2003.
[24]  B. M. K. Raju, V. K. Bhatia, and V. V. Kumar, “Assessment of sensitivity with incomplete data,” Journal of the Indian Society of Agricultural Statistics, vol. 60, pp. 118–125, 2006.
[25]  B. M. K. Raju, V. K. Bhatia, and L. M. Bhar, “Assessing stability of crop varieties with incomplete data,” Journal of the Indian Society of Agricultural Statistics, vol. 63, pp. 139–149, 2009.
[26]  D. G. Pereira, J. T. Mexia, and P. C. Rodrigues, “Robustness of joint regression analysis,” Biometrical Letters, vol. 44, pp. 105–128, 2007.
[27]  P. C. Rodrigues, D. G. S. Pereira, and J. T. Mexia, “A comparison between joint regression analysis and the additive main and multiplicative interaction model: the robustness with increasing amounts of missing data,” Scientia Agricola, vol. 68, no. 6, pp. 679–705, 2011.
[28]  P. J. C. Rodrigues, New Strategies to Detect and Understand Genotype-by-Environment Interactions and QTL-by-Environment Interactions [Dissertation], Universidade Nova de Lisboa, 2012.
[29]  R. Bro, K. Kjeldahl, A. K. Smilde, and H. A. L. Kiers, “Cross-validation of component models: a critical look at current methods,” Analytical and Bioanalytical Chemistry, vol. 390, no. 5, pp. 1241–1251, 2008.
[30]  S. Wold, “Cross-validatory estimation of the number of components in factor and principal components models,” Technometrics, vol. 20, pp. 397–405, 1978.
[31]  H. T. Eastment and W. J. Krzanowski, “Cross-validatory choice of the number of components from a principal component analysis,” Technometrics, vol. 24, no. 1, pp. 73–77, 1982.
[32]  S. Arciniegas-Alarcón, M. García-Pe?a, and C. T. S. Dias, “Data imputation in trials with genotype × environment interaction,” Interciencia, vol. 36, pp. 444–449, 2011.
[33]  A. Smilde, R. Bro, and P. Geladi, Multi-Way Analysis with Applications in the Chemical Sciences, John Wiley and Sons, Chichester, UK, 2004.
[34]  W. J. Krzanowski, “Missing value imputation in multivariate data using the singular value decomposition of a matrix,” Biometrical Letters, vol. 25, pp. 31–39, 1988.
[35]  I. J. Good, “Applications of the singular decomposition of a matrix,” Technometrics, vol. 11, no. 4, pp. 823–831, 1969.
[36]  T. Hastie, R. Tibshirani, G. Sherlock, M. Eisen, P. Brown, and D. Botstein, “Imputing missing data for gene expression arrays,” Technical Report, Division of Biostatistics, Standford University, 1999.
[37]  D. Hedderley and L. Wakeling, “A comparison of imputation techniques for internal preference mapping, using Monte Carlo simulation,” Food Quality and Preference, vol. 6, no. 4, pp. 281–297, 1995.
[38]  J. Josse, J. Pagès, and F. Husson, “Multiple imputation in principal component analysis,” Advances in Data Analysis and Classification, vol. 5, no. 3, pp. 231–246, 2011.
[39]  C. T. S. Dias and W. J. Krzanowski, “Model selection and cross validation in additive main effect and multiplicative interaction models,” Crop Science, vol. 43, no. 3, pp. 865–873, 2003.
[40]  A. L. Bello, “Choosing among imputation techniques for incomplete multivariate data: a simulation study,” Communications in Statistics, vol. 22, pp. 853–877, 1993.
[41]  T. Caliński, S. Czajka, Z. Kaczmarek, P. Krajewski, and W. Pilarczyk, “Analyzing the genotype-by-environment interactions under a randomization-derived mixed model,” Journal of Agricultural, Biological, and Environmental Statistics, vol. 14, pp. 224–241, 2009.
[42]  F. J. C. Farias, Selection Index in Upland Cotton Cultivars [Dissertation], University of S?o Paulo, 2005.
[43]  F. Flores, M. T. Moreno, and J. I. Cubero, “A comparison of univariate and multivariate methods to analyze interaction,” Field Crops Research, vol. 56, no. 3, pp. 271–286, 1998.
[44]  A. L. Bello, “Imputation techniques in regression analysis: looking closely at their implementation,” Computational Statistics and Data Analysis, vol. 22, pp. 853–877, 1995.
[45]  R Development Core Team, R: A Language and Environment For Statistical Computing. R Foundation For Statistical Computing, Vienna, Austria, 2012, http://www.R-project.org/.
[46]  H. G. Gauch, “Model selection and validation for yield trials with interaction,” Biometrics, vol. 44, pp. 705–715, 1988.
[47]  H. G. Gauch, Statistical Analysis of Regional Yield Trials: AMMI Analysis of Factorial Designs, Elsevier, Amsterdam, The Netherlands, 1992.
[48]  K. R. Gabriel, “Le biplot-outil d'exploration de données multidimensionelles,” Journal de la Societe Francaise de Statistique, vol. 143, pp. 5–55, 2002.
[49]  C. T. S. Dias and W. J. Krzanowski, “Choosing components in the additive main effect and multiplicative interaction (AMMI) models,” Scientia Agricola, vol. 63, no. 2, pp. 169–175, 2006.
[50]  M. García-Pe?a and C. T. S. Dias, “Analysis of bivariate additive models with multiplicative interaction (AMMI),” Revista Brasileira de Biometria, vol. 27, pp. 586–602, 2009.
[51]  K. Hongyu, Empirical Distribution of Eigenvalues Associated with the Interaction Matrix of the AMMI Models by Non-Parametric Bootstrap Method [Dissertation], University of S?o Paulo, 2012.
[52]  P. Sprent and N. C. Smeeton, Applied Nonparametric Statistical Methods, Chapman and Hall, London, UK, 2001.
[53]  H. G. Gauch Jr., H.-P. Piepho, and P. Annicchiarico, “Statistical analysis of yield trials by AMMI and GGE: further considerations,” Crop Science, vol. 48, no. 3, pp. 866–889, 2008.
[54]  P. C. Rodrigues, S. Mejza, and J. T. Mexia, “Structuring genotype × environment interaction: an overview,” Bulletin of Plant Breeding and Acclimatization Institute, vol. 250, pp. 225–236, 2009.
[55]  J. L. Schafer and J. W. Graham, “Missing data: our view of the state of the art,” Psychological Methods, vol. 7, no. 2, pp. 147–177, 2002.
[56]  A. L. Bello, “A bootstrap method for using imputation techniques for data with missing values,” Biometrical Journal, vol. 36, pp. 453–464, 1994.
[57]  R. J. Little and D. B. Rubin, Statistical Analysis With Missing Data, John Wiley and Sons, New York, NY, USA, 2002.
[58]  H.-P. Piepho and J. M?hring, “Selection in cultivar trials: is it ignorable?” Crop Science, vol. 46, no. 1, pp. 192–201, 2006.

Full-Text

Contact Us

service@oalib.com

QQ:3279437679

WhatsApp +8615387084133