<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">umovest</journal-id><journal-title-group><journal-title xml:lang="ru">Статистика и Экономика</journal-title><trans-title-group xml:lang="en"><trans-title>Statistics and Economics</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">2500-3925</issn><publisher><publisher-name>Plekhanov Russian University of Economics</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.21686/2500-3925-2017-3-10-20</article-id><article-id custom-type="elpub" pub-id-type="custom">umovest-1092</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>МЕТОДОЛОГИЯ СТАТИСТИКИ</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>METHODOLOGY OF STATISTICS</subject></subj-group></article-categories><title-group><article-title>Вычисление истинного уровня значимости предикторов при проведении процедуры спецификации уравнения регрессии</article-title><trans-title-group xml:lang="en"><trans-title>Calculating the true level of predictors significance when carrying out the procedure of regression equation specification</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Моисеев</surname><given-names>Н. А.</given-names></name><name name-style="western" xml:lang="en"><surname>Moiseev</surname><given-names>Nikita A.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Кандидат экономических наук, доцент кафедры Математических методов в экономике</p></bio><bio xml:lang="en"><p>Cand. Sci. (Economics), Associate Professor of the Department of Mathematical Methods in Economics</p></bio><email xlink:type="simple">moiseev.na@rea.ru</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Российский экономический университет имени Г.В. Плеханова,</institution><country>Россия</country></aff><aff xml:lang="en"><institution>Plekhanov Russian University of Economics</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2017</year></pub-date><pub-date pub-type="epub"><day>25</day><month>04</month><year>2017</year></pub-date><volume>0</volume><issue>3</issue><fpage>10</fpage><lpage>20</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Моисеев Н.А., 2017</copyright-statement><copyright-year>2017</copyright-year><copyright-holder xml:lang="ru">Моисеев Н.А.</copyright-holder><copyright-holder xml:lang="en">Moiseev N.A.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://statecon.rea.ru/jour/article/view/1092">https://statecon.rea.ru/jour/article/view/1092</self-uri><abstract><p>Данная научная работа посвящена новому численному методу, вычисляющему несмещенные оценки p-значений для предикторов линейных регрессионных моделей с учетом числа потенциальных объясняющих переменных, их дисперсионно-ковариационной матрицы и степени ее неопределенности, основанной на числе рассматриваемых наблюдений. Такая поправка помогает ограничивать число ошибок 1-ого рода в научных исследованиях, значительно понижая число публикаций, декларирующих ложные зависимости в качестве истинных. Сравнительный анализ с такими существующими методами как поправка Бонферрони и поправка Шехата и Уайта явным образом демонстрирует их недостатки, особенно в случае, когда число потенциальных предикторов сравнимо с числом наблюдений. Также в процессе проведения сравнительного анализа было показано, что когда дисперсионно-ковариационная матрица набора потенциальных предикторов является диагональной, т.е. данные независимы, предложенная простая поправка является лучшим и самым легким в реализации методом для получения несмещенных корректировок традиционных p-значений. Однако, в случае присутствия сильно коррелированных данных простая поправка переоценивает истинные p-значения, что может приводить к ошибкам 2-ого рода. Также было выявлено, что исправленные p-значения зависят от числа наблюдений, числа потенциальных объясняющих переменных и выборочной дисперсионно-ковариационной матрицы. Например, если имеется только две потенциальных объясняющих переменных, конкурирующие за одну позицию в регрессионной модели, тогда, если они слабо коррелированы, исправленное p-значение будет ниже, чем в случае когда число наблюдений меньше и наоборот; если данные сильно коррелированы, случай с большим числом наблюдений будет показывать более низкое исправленное p-значение. С увеличением корреляции все поправки независимо от числа наблюдений стремятся к исходному p-значению. Данный феномен легко объяснить: с приближением коэффициента корреляции к единице две переменных практически линейно зависят друг от друга и в случае, если одна из них является значимой, то и другая почти наверняка будет демонстрировать такую же значимость. С другой стороны, если выборочная дисперсионно-ковариационная матрица стремится к диагональной и число наблюдений стремится к бесконечности, то предложенный численный метод будет возвращать поправки, близкие к простой поправке. В случае, когда число наблюдений много больше числа потенциальных предикторов, тогда поправка Шехата и Уайта дают примерно одинаковые поправки с предложенным численным методом. Однако, в намного более распространенных случаях, когда число наблюдений сравнимо с числом потенциальных предикторов, существующие методы демонстрируют достаточно значительные неточности. Когда число потенциальных предикторов больше доступного числа наблюдений, представляется невозможным рассчитать истинные p-значения. Вследствие этого рекомендуется не рассматривать такие наборы данных при построении регрессионных моделей, поскольку только выполнение вышеупомянутого условия обеспечивает расчет несмещенных корректировок p-значения. Предлагаемый метод полностью алгоритмизирован и может быть внедрен в любой пакет статистического анализа данных.</p></abstract><trans-abstract xml:lang="en"><p>The paper is devoted to a new randomization method that yields unbiased adjustments of p-values for linear regression models predictors by incorporating the number of potential explanatory variables, their variance-covariance matrix and its uncertainty, based on the number of observations. This adjustment helps to control type I errors in scientific studies, significantly decreasing the number of publications that report false relations to be authentic ones. Comparative analysis with such existing methods as Bonferroni correction and Shehata and White adjustments explicitly shows their imperfections, especially in case when the number of observations and the number of potential explanatory variables are approximately equal. Also during the comparative analysis it was shown that when the variance-covariance matrix of a set of potential predictors is diagonal, i.e. the data are independent, the proposed simple correction is the best and easiest way to implement the method to obtain unbiased corrections of traditional p-values. However, in the case of the presence of strongly correlated data, a simple correction overestimates the true pvalues, which can lead to type II errors. It was also found that the corrected p-values depend on the number of observations, the number of potential explanatory variables and the sample variance-covariance matrix. For example, if there are only two potential explanatory variables competing for one position in the regression model, then if they are weakly correlated, the corrected p-value will be lower than when the number of observations is smaller and vice versa; if the data are highly correlated, the case with a larger number of observations will show a lower corrected p-value. With increasing correlation, all corrections, regardless of the number of observations, tend to the original p-value. This phenomenon is easy to explain: as correlation coefficient tends to one, two variables almost linearly depend on each other, and in case if one of them is significant, the other will almost certainly show the same significance. On the other hand, if the sample variance-covariance matrix tends to be diagonal and the number of observations tends to infinity, the proposed numerical method will return corrections close to the simple correction. In the case when the number of observations is much greater than the number of potential predictors, then the Shehata and White corrections give approximately the same corrections with the proposed numerical method. However, in much more common cases, when the number of observations is comparable to the number of potential predictors, the existing methods demonstrate significant inaccuracies. When the number of potential predictors is greater than the available number of observations, it seems impossible to calculate the true p-values. Therefore, it is recommended not to consider such datasets when constructing regression models, since only the fulfillment of the above condition ensures calculation of unbiased p-value corrections. The proposed method is easy to program and can be integrated into any statistical software package.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>регрессионные модели</kwd><kwd>корректировка p-значений</kwd><kwd>значимость предикторов</kwd><kwd>численный метод</kwd><kwd>распределение Уишарта</kwd><kwd>дисперсионно-ковариационная матрица</kwd><kwd>преобразование Холецкого</kwd></kwd-group><kwd-group xml:lang="en"><kwd>regression models</kwd><kwd>p-value adjustment</kwd><kwd>significance of predictors</kwd><kwd>randomization method</kwd><kwd>Wishart distribution</kwd><kwd>variance-covariance matrix</kwd><kwd>Cholesky decomposition</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Akaike H. Information theory and an extension of the maximum likelihood principle. In: Petroc B., Csake F. (Eds.) Second International Symposium on Information Theory. 1973.</mixed-citation><mixed-citation xml:lang="en">Akaike H. Information theory and an extension of the maximum likelihood principle. In: Petroc B., Csake F. (Eds.) Second International Symposium on Information Theory. 1973.</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Akaike H. A Bayesian extension of the minimum AIC procedure of autoregressive model fitting // Biometrika. 1979. 66. P. 237–242.</mixed-citation><mixed-citation xml:lang="en">Akaike H. A Bayesian extension of the minimum AIC procedure of autoregressive model fitting // Biometrika. 1979. 66. P. 237–242.</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Bates J.M., Granger, C.W.J. The combination of forecasts // Operations Research Quarterly. 1969. 20. P. 451–468.</mixed-citation><mixed-citation xml:lang="en">Bates J.M., Granger, C.W.J. The combination of forecasts // Operations Research Quarterly. 1969. 20. P. 451–468.</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Buckland S.T., Burnham K.P., Augustin, N.H. Model selection: An integral part of inference // Biometrics. 1997. 53. P. 603–618.</mixed-citation><mixed-citation xml:lang="en">Buckland S.T., Burnham K.P., Augustin, N.H. Model selection: An integral part of inference // Biometrics. 1997. 53. P. 603–618.</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Canning F.L. 1959. Estimating load requirements in a job shop // Journal of Industrial Engineering. 1959. 10. P. 447.</mixed-citation><mixed-citation xml:lang="en">Canning F.L. 1959. Estimating load requirements in a job shop // Journal of Industrial Engineering. 1959. 10. P. 447.</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Derksen S., Keselman H.J. Backward, forward and stepwise automated subset selection algorithms: frequency of obtaining authentic and noise variables // British Journal of Mathematical and Statistical Psychology. 1992. 45. P. 265–282.</mixed-citation><mixed-citation xml:lang="en">Derksen S., Keselman H.J. Backward, forward and stepwise automated subset selection algorithms: frequency of obtaining authentic and noise variables // British Journal of Mathematical and Statistical Psychology. 1992. 45. P. 265–282.</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Hurvich C.M., Tsai C.L. The impact of model selection on inference in linear regression // The American Statistician. 1990. 44. 3. P. 214–217.</mixed-citation><mixed-citation xml:lang="en">Hurvich C.M., Tsai C.L. The impact of model selection on inference in linear regression // The American Statistician. 1990. 44. 3. P. 214–217.</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Kramer C.Y. Simplified computations for multiple regression // Industrial Quality Control. 1957. 13. 8. 8.</mixed-citation><mixed-citation xml:lang="en">Kramer C.Y. Simplified computations for multiple regression // Industrial Quality Control. 1957. 13. 8. 8.</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Larzelere R.E., Mulaik S.A. Single-sample tests for many correlations // Psychological Bulletin. 1977. 84. P. 557 – 569.</mixed-citation><mixed-citation xml:lang="en">Larzelere R.E., Mulaik S.A. Single-sample tests for many correlations // Psychological Bulletin. 1977. 84. P. 557 – 569.</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">Lovell M.C. Data mining. The Review of Economics and Statistics. 1983. 65. P. 1–12.</mixed-citation><mixed-citation xml:lang="en">Lovell M.C. Data mining. The Review of Economics and Statistics. 1983. 65. P. 1–12.</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">Miller A. J. Selection of subsets of regression variables (with discussion) // Journal of the Royal Statistical Society. 1984. A. 147. P. 389–425.</mixed-citation><mixed-citation xml:lang="en">Miller A. J. Selection of subsets of regression variables (with discussion) // Journal of the Royal Statistical Society. 1984. A. 147. P. 389–425.</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">Mittelhammer Ron C., Judge George G., Miller Douglas J. Econometric Foundations. Cambridge University Press. 2000. P. 73–74.</mixed-citation><mixed-citation xml:lang="en">Mittelhammer Ron C., Judge George G., Miller Douglas J. Econometric Foundations. Cambridge University Press. 2000. P. 73–74.</mixed-citation></citation-alternatives></ref><ref id="cit13"><label>13</label><citation-alternatives><mixed-citation xml:lang="ru">Moiseev N.A. Linear model averaging by minimizing mean-squared forecast error unbiased estimator // Model Assisted Statistics and Applications. 2016. Vol. 11, No 4, P. 325–338.</mixed-citation><mixed-citation xml:lang="en">Moiseev N.A. Linear model averaging by minimizing mean-squared forecast error unbiased estimator // Model Assisted Statistics and Applications. 2016. Vol. 11, No. 4, P. 325–338.</mixed-citation></citation-alternatives></ref><ref id="cit14"><label>14</label><citation-alternatives><mixed-citation xml:lang="ru">Shehata Yasser A., White Paul A Randomization Method to Control the Type I Error Rates in Best Subset Regression // Journal of Modern Applied Statistical Methods. 2008. 7. 2. P. 398–407.</mixed-citation><mixed-citation xml:lang="en">Shehata Yasser A., White Paul A Randomization Method to Control the Type I Error Rates in Best Subset Regression // Journal of Modern Applied Statistical Methods. 2008. 7. 2. P. 398–407.</mixed-citation></citation-alternatives></ref><ref id="cit15"><label>15</label><citation-alternatives><mixed-citation xml:lang="ru">Shibata Ritaei. Asymptotically efficient selection of the order of the model for estimating parameters of a linear process // Annals of Statistics. 1990. 8. Pp. 147–164.</mixed-citation><mixed-citation xml:lang="en">Shibata Ritaei. Asymptotically efficient selection of the order of the model for estimating parameters of a linear process // Annals of Statistics. 1990. 8. P. 147–164.</mixed-citation></citation-alternatives></ref><ref id="cit16"><label>16</label><citation-alternatives><mixed-citation xml:lang="ru">Shibata Ritaei. An optimal selection of regression variables // Biometrika. 1981. 68. P. 45–54.</mixed-citation><mixed-citation xml:lang="en">Shibata Ritaei. An optimal selection of regression variables // Biometrika. 1981. 68. P. 45–54.</mixed-citation></citation-alternatives></ref><ref id="cit17"><label>17</label><citation-alternatives><mixed-citation xml:lang="ru">Shibata Ritaei. Asymptotic mean efficiency of a selection of regression variables // Annals of the Institute of Statistical Mathematics. 1983. 35. P. 415–423.</mixed-citation><mixed-citation xml:lang="en">Shibata Ritaei. Asymptotic mean efficiency of a selection of regression variables // Annals of the Institute of Statistical Mathematics. 1983. 35. P. 415–423.</mixed-citation></citation-alternatives></ref><ref id="cit18"><label>18</label><citation-alternatives><mixed-citation xml:lang="ru">Wishart J. The generalized product moment distribution in samples from a normal multivariate population // Biometrica. 1928. 20A. P. 32–52.</mixed-citation><mixed-citation xml:lang="en">Wishart J. The generalized product moment distribution in samples from a normal multivariate population // Biometrica. 1928. 20A. P. 32–52.</mixed-citation></citation-alternatives></ref><ref id="cit19"><label>19</label><citation-alternatives><mixed-citation xml:lang="ru">Глазьев С. Проблемы прогнозирования макроэкономической динамики // Российский экономический журнал. 2001. № 3. C. 76–85; № 4. C. 12–22.</mixed-citation><mixed-citation xml:lang="en">Glaz’ev S. Problemy prognozirovaniya makroekonomicheskoi dinamiki // Rossiiskii ekonomicheskii zhurnal. 2001. № 3. P. 76–85; № 4. P. 12–22. (in Russ.)</mixed-citation></citation-alternatives></ref><ref id="cit20"><label>20</label><citation-alternatives><mixed-citation xml:lang="ru">Крыштановский А.О. Методы анализа временных рядов // Мониторинг общественного мнения: экономические и социальные перемены. 2000. № 2 (46). С. 44–51.</mixed-citation><mixed-citation xml:lang="en">Kryshtanovskii A.O. Metody analiza vremennykh ryadov // Monitoring obshchestvennogo mneniya: ekonomicheskie i sotsial’nye peremeny. 2000. № 2 (46). P. 44–51. (in Russ.)</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
