{"id":373,"date":"2019-03-12T08:05:35","date_gmt":"2019-03-12T08:05:35","guid":{"rendered":"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/?post_type=chapter&#038;p=373"},"modified":"2019-03-12T08:49:26","modified_gmt":"2019-03-12T08:49:26","slug":"chi-square-%cf%872distribution-and-tests-of-significance-based-on-%cf%872","status":"publish","type":"chapter","link":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/chapter\/chi-square-%cf%872distribution-and-tests-of-significance-based-on-%cf%872\/","title":{"rendered":"Chi Square (\u03c72)Distribution and tests of significance based on \u03c72"},"content":{"raw":"<div>\r\n\r\n\u00a0 \u00a0 1.\u00a0<strong>Introduction<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Chi Square <strong>(<\/strong>\u03c72)distribution is an important probability distribution which is used in a number of statistical tests of significance, most famous among which are \u03c72 test of independence and \u03c72 test of the goodness of fit. As a non-parametric, rank-based method, tests based on \u03c72-distribution do not make any assumptions on the normality of the distributions of populations from which the samples came from, and therefore, the test is suitable for the analysis of nominal or categorical samples. However, for the analysis of non-Gaussian data, better non-parametric tests are available (for example, Mann-Whitney U test to compare two unpaired groups, and Kruskal-Wallis test to compare means of three or more unpaired groups). When the outcome is binomial-especially for the analysis of 2x2 contingency tables-Fisher\u2019s exact test is preferred. To compare three or more unpaired groups, \u03c72 test of independence is still the best method. When we want to find the fit of an observed distribution (data) to a theoretically expected distribution (model),\u00a0\u03c72 test of the goodness of fit is performed. The main difference between \u03c72 test of independence and \u03c72 test of the goodness of fit is that while the former automatically calculates expected frequency from the input data of observed frequencies, the latter require input of expected frequencies derived from an explicit model (for example, Mendel\u2019s dihybrid cross ratio, or\u00a0Fisherian sex ratio).<\/p>\r\n&nbsp;\r\n\r\n<strong>2.\u00a0<\/strong><strong>Learning Outcome:<\/strong>\r\n\r\n<strong>\u00a0<\/strong>\r\n<p style=\"text-align: justify\">a. To learn about the properties of \u03c72-distribution and statistical tests of significance based upon \u03c72 distribution<\/p>\r\nb. To learn assumptions for \u03c72-tests\r\n\r\nc. To learn how \u03c72 test of independence is performed\r\n<p style=\"text-align: justify\">d. To learn how fisher\u2019s exact test is performed, which is a much better alternative to \u03c72 test of independence especially with small sample sizes<\/p>\r\ne. To learn how \u03c72 test of the goodness of fit is performed\r\n\r\n&nbsp;\r\n\r\n<strong>3.\u00a0<\/strong><strong>\u03c7<\/strong><strong>2<\/strong><strong> distribution<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">\u03c72 distribution (Chi-Square distribution) is a type of asymmetric continuous probability distribution with probabilities of every \u03c72 statistic known under the assumption of null hypothesis. It was first used\u00a0<span style=\"font-size: 1em;text-align: initial\">and described by Karl Pearson in 1900, one of the founding fathers of statistics and Population genetics.\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">This distribution enables us to calculate <\/span>P<span style=\"text-align: initial;font-size: 1em\"> value from a given \u03c72 statistic.<\/span><\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n<img class=\"aligncenter size-full wp-image-377\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-219.png\" alt=\"\" width=\"177\" height=\"56\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Distribution of these \u03c72 statistic under the assumption of null hypothesis plotted as in a probability histogram is called \u03c72 distribution. Like lognormal distribution and F-distribution, \u03c72 distribution is right-skewed (with a long tail towards right. The shape of \u03c72 -distribution depends only on <em>k<\/em> the shape parameter (degrees of freedom, df). \u03c72 distribution is a special case of more generalized gamma distribution.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-378\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-220.png\" alt=\"\" width=\"553\" height=\"405\" \/>\r\n\r\n&nbsp;\r\n\r\n<\/div>\r\n&nbsp;\r\n\r\n<strong style=\"text-align: initial;font-size: 1em\">4.\u00a0<\/strong><strong style=\"text-align: initial;font-size: 1em\">Tests of significance based on \u03c7<\/strong><strong style=\"text-align: initial;font-size: 1em\">2<\/strong><strong style=\"text-align: initial;font-size: 1em\"> distribution<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">\u03c72 distribution is used for two main tests; \u03c72 test of independence for the analysis of categorical data (to test whether two categorical variables are correlated), and \u03c72 test of the goodness of fit of an observed distribution (data) to a theoretically expected distribution (model). This distribution is also used for likelihood ratio test (and its variant, hierarchical LRT) used for model selection in molecular phylogenetics<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">5.\u00a0<\/strong><strong style=\"text-align: initial;font-size: 1em\">Assumptions for \u03c7<\/strong><strong style=\"text-align: initial;font-size: 1em\">2<\/strong><strong style=\"text-align: initial;font-size: 1em\"> tests<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">1.\u00a0\u00a0\u00a0\u00a0\u00a0 Random observations<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">2.\u00a0\u00a0\u00a0\u00a0\u00a0 Independent measurements<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">3.\u00a0\u00a0\u00a0\u00a0\u00a0 Accurate data<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Note that there are no explicit assumptions about the distributions of populations from which these samples are drawn, as \u03c72 is considered as a nonparametric test.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">6.<\/strong><span style=\"text-align: initial;font-size: 1em\">\u00a0<\/span><strong style=\"text-align: initial;font-size: 1em\">\u03c7<\/strong><strong style=\"text-align: initial;font-size: 1em\">2<\/strong><strong style=\"text-align: initial;font-size: 1em\"> test of independence<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">This test analyses two sets of information (two variables) for any association between them, therefore in <\/span><em style=\"text-align: initial;font-size: 1em\">sensu stricto<\/em><span style=\"text-align: initial;font-size: 1em\">, this test falls under multivariate statistics. In effect, this test compares two proportions like other methods for comparing proportions such as Relative Risk, Attributable risk, Odd\u2019s Ratio and so on. The null hypothesis is that there is no association between them (these variables are statistically independent), while <\/span>alternative<span style=\"text-align: initial;font-size: 1em\"> hypothesis is that there is an association between them (variables are statistically dependant).<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">The \u03c72 test statistic is computed as<\/span><\/p>\r\n\r\n<div>\r\n\r\n<img class=\"aligncenter size-full wp-image-379\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-221.png\" alt=\"\" width=\"108\" height=\"62\" \/>\r\n<p style=\"text-align: justify\">Where <em>f<\/em><em>o<\/em> is observed frequency and <em>f<\/em><em>e<\/em> is expected frequency. <em>f<\/em><em>e<\/em> can be calculated as (row total x column total)\/n<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Let us consider an example. Is there any association between income level and happiness level? To study, imagine we have done a questionnaire survey and obtained the results as given below:<\/p>\r\n\r\n<\/div>\r\n<img class=\"aligncenter size-full wp-image-380\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-222.png\" alt=\"\" width=\"710\" height=\"239\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">This is the result of <\/span>questionnaire<span style=\"text-align: initial;font-size: 1em\"> survey with <\/span>total<span style=\"text-align: initial;font-size: 1em\"> number of participants 2955, which is indicated in the table as the overall total. The first cell 272 means out of 615 rich participants, 272 responded that their happiness level is high. These numbers (highlighted in italics) are the observed frequencies (<\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">o<\/em><span style=\"text-align: initial;font-size: 1em\"> in our equation). Note that these numbers are only frequencies; measurement merely measures into three categories (high, middle and low), so the level of measurement here is nominal (categorical). Tables like these where exact measured values are entered <\/span>is<span style=\"text-align: initial;font-size: 1em\"> called contingency tables.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Let us first define our null hypothesis and alternative hypotheses:<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">H0: Income level happiness level are independent<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Ha: Income level and happiness level are <\/span>dependant<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">First<span style=\"text-align: initial;font-size: 1em\"> step in <\/span>\u03c72<span style=\"text-align: initial;font-size: 1em\"> test of independence is to calculate <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\">, the expected frequencies of all these cells. This is computed as:<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\">= (row total x column total)\/ overall total<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">For the first cell (where <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">o<\/em><em style=\"text-align: initial;font-size: 1em\">=272<\/em><span style=\"text-align: initial;font-size: 1em\">), row total is 615 and column total is 911. Plugging into the above equation,<\/span><\/p>\r\n<p style=\"text-align: justify\"><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\">= (615 x 911)\/2955 =189.6<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">It is better to calculate these values in a tabular format:<\/span><\/p>\r\n\r\n<div>\r\n\r\n<img class=\"size-full wp-image-381 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-223.png\" alt=\"\" width=\"432\" height=\"587\" \/>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n\u03c72 test statistic = 172.28\r\n<p style=\"text-align: justify\">Next step is to look up \u03c72 table to find \u03c72 critical value, for which we should know degree of freedom and significance level (which is 0.05). For \u03c72 tests df can be calculated by the following equation<\/p>\r\ndf = (No. of rows-1) x (No. of columns -1)\r\n\r\nRemember that these numbers means that of the actual data; totals or labels are excluded.\r\n\r\nDf= (3-1) x (3-1)\r\n\r\n=2 x 2\r\n\r\n=4\r\n\r\n<\/div>\r\n<img class=\"size-full wp-image-382 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-224.png\" alt=\"\" width=\"528\" height=\"340\" \/>\r\n<div>\r\n\r\n&nbsp;\r\n\r\n<\/div>\r\n&nbsp;\r\n<div>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">As table value (critical \u03c72, 9.488) is far less than our obtained \u03c72 test statistic (172.28), we can conclude that P&lt;0.05, we reject null hypothesis of independence of two variables and conclude that two variables are dependant, or associated. Moving towards right in the table, we can see that even at significance level 0.005, critical \u03c7214.86 is still far less than our obtained \u03c72 test statistic, therefore P value must be &lt;0.005<\/p>\r\n&nbsp;\r\n\r\nThere is no support for \u03c72 test in excel. An online calculator like the following can be used instead\r\n\r\n&nbsp;\r\n\r\n<a href=\"http:\/\/turner.faculty.swau.edu\/mathematics\/math241\/materials\/contablecalc\/\">http:\/\/turner.faculty.swau.edu\/mathematics\/math241\/materials\/contablecalc\/<\/a>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The \u03c72 statistic is sensitive to small cell sizes. Whenever any of your cell sizes are &lt;5, a slightly modified formula (Yates\u2019 correction for continuity) to calculate chi square should be used<\/p>\r\n&nbsp;\r\n\r\nModified Formula (0.5 is deducted from the absolute value of fo \u2013 fe before squaring)\r\n\r\n<img class=\"aligncenter size-full wp-image-383\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-225.png\" alt=\"\" width=\"162\" height=\"60\" \/>\r\n\r\nHowever, most statisticians agree that Yates correction overcorrects it.\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">When there are only two categories (like head and tail in coin flipping, or male or female in gender), to make inferences of one population the best option is <\/span>binomial<span style=\"text-align: initial;font-size: 1em\"> test, which <\/span>calculate<span style=\"text-align: initial;font-size: 1em\"> the exact probabilities using binomial equation. <\/span>Binomial<span style=\"text-align: initial;font-size: 1em\"> test is available at <\/span><a style=\"text-align: initial;font-size: 1em\" href=\"https:\/\/www.graphpad.com\/quickcalcs\/binomial1\/\">https:\/\/www.graphpad.com\/quickcalcs\/binomial1\/<\/a><span style=\"text-align: initial;font-size: 1em\"> P values inferred from <\/span>\u03c72<span style=\"text-align: initial;font-size: 1em\"> test are only approximations, not exact.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">7.\u00a0<\/strong><strong style=\"text-align: initial;font-size: 1em\">Fisher\u2019s exact test<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">For 2 x 2 contingency tables used frequently in case control studies, the best test is Fisher\u2019s exact test that can be found here:<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><a style=\"text-align: initial;font-size: 1em\" href=\"https:\/\/www.graphpad.com\/quickcalcs\/contingency1.cfm\">https:\/\/www.graphpad.com\/quickcalcs\/contingency1.cfm<\/a><\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n<img class=\"size-full wp-image-384 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-226.png\" alt=\"\" width=\"390\" height=\"154\" \/>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\nH0: Apixaban does not alter the risk of a recurrent thromboembolism\r\n\r\nHa: Apixaban alters the risk of a recurrent thromboembolism\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The values in the table indicate No. of patients (treated with placebo or apixaban, two rows) who already had thromboembolism and going on to have another thromboembolism during the study period (in the column \u201drecurrent\u201d) and those who do not have second episode of thromboembolism (in the column \u201cNo Recurrence\u201d). In 2 x 2 contingency tables like this, it is customary to enter groups as rows and outcomes as columns. Fisher\u2019s exact test uses the following formula which is simple and straightforward to understand:<\/p>\r\n<img class=\"size-full wp-image-385 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-227.png\" alt=\"\" width=\"426\" height=\"104\" \/>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Where a, b, c and d are values in 2 x 2 contingency table and n is the total number of values of the table. When the numbers become very large, calculation of factorials becomes mathematically unwieldy so that \u03c72 test is preferred.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Fisher\u2019s exact test for our above example returns a P value less than 0.0001, so the difference is very significant.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">8.<\/strong><span style=\"text-align: initial;font-size: 1em\">\u00a0<\/span><strong style=\"text-align: initial;font-size: 1em\">\u03c7<\/strong><strong style=\"text-align: initial;font-size: 1em\">2<\/strong><strong style=\"text-align: initial;font-size: 1em\"> test of the goodness of fit<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">\u03c72<span style=\"text-align: initial;font-size: 1em\"> test of the goodness of fit is used when we want to find the fit of an observed distribution (data) to a theoretically expected distribution (model). The main difference from the earlier test (independence) is that for <\/span>goodness<span style=\"text-align: initial;font-size: 1em\"> of fit the expected frequencies are derived from a theory or a mathematical model, while in the former, expected frequencies are calculated from the observed frequencies itself. Therefore, for <\/span>test<span style=\"text-align: initial;font-size: 1em\"> of independence, input data is only the observed (empirical) frequencies. In the case of <\/span>test<span style=\"text-align: initial;font-size: 1em\"> of goodness of fit, input data encompasses expected frequencies derived from theory in addition to the observed frequencies.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">The \u03c72 test statistic for <\/span>test<span style=\"text-align: initial;font-size: 1em\"> of goodness of fit is computed exactly as in <\/span>\u03c72<span style=\"text-align: initial;font-size: 1em\"> test of independence:<\/span><\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n<img class=\"aligncenter size-full wp-image-386\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-228.png\" alt=\"\" width=\"106\" height=\"61\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Where <em>f<\/em><em>o<\/em> is observed frequency and <em>f<\/em><em>e<\/em> is expected frequency. Only difference from \u03c72 test of independence is that <em>f<\/em><em>e<\/em> is not computed from <em>f<\/em><em>o<\/em> but from a model.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Let us consider an example. Out of total 556 pea plants, the famous Geneticist Gregor Mendel observed (dihybrid cross) four seed phenotypes in frequencies given below:<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-387 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-229.png\" alt=\"\" width=\"718\" height=\"207\" \/>\r\n\r\n<\/div>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">According to his famous law of independent <\/span>assortment<span style=\"text-align: initial;font-size: 1em\"> Mendel expected a certain ratio (9:3:3:1) of those phenotypes. This ratio is a model, <\/span>a theoretically expected proportions<span style=\"text-align: initial;font-size: 1em\">. Let us plot those expected proportions in this table as well:<\/span><\/p>\r\n\r\n<div>\r\n\r\n<img class=\"size-full wp-image-389 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-231.png\" alt=\"\" width=\"451\" height=\"226\" \/>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">To get Expected frequencies, all we have to do is to multiply each of the expected proportions with the total no. of plants (556). Note that total of expected frequencies add up to the total (556)<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-388 alignleft\" style=\"text-indent: 0px\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-230.png\" alt=\"\" width=\"714\" height=\"252\" \/>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\nQuestion is whether observed frequencies deviate significantly from the expected frequencies?\r\n\r\n&nbsp;\r\n\r\nLet us first define our null hypothesis and alternative hypotheses:\r\n\r\nH0: <em>f<\/em><em>o<\/em> = <em>f<\/em><em>e<\/em> (i.e, our data fits model well)\r\n\r\nHa: <em>f<\/em><em>o<\/em> \u2260 <em>f<\/em><em>e<\/em> (i.e, our data do not fits model)\r\n\r\n&nbsp;\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">Now let us complete the \u03c72 table<\/span>\r\n\r\n<\/div>\r\n<div>\r\n\r\n<img class=\"alignleft size-full wp-image-390\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-232.png\" alt=\"\" width=\"594\" height=\"348\" \/>\r\n\r\n<\/div>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n<div>\r\n\r\n&nbsp;\r\n\r\n\u03c72 test statistic = 0.47\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Next step is to look up \u03c72 table to find \u03c72 critical value, for which we should know degree of freedom and significance level (which is 0.05). For \u03c72 tests of goodness of fit, our data is grouped only in rows, not in columns. So df is (no. of rows \u2013 1)<\/p>\r\n&nbsp;\r\n\r\n4-1 = 3\r\n\r\n<img class=\"alignleft size-full wp-image-391\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-233.png\" alt=\"\" width=\"528\" height=\"340\" \/>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">As table value (critical \u03c72, 7.815) is far higher than our obtained \u03c72 test statistic (0.47), we can conclude that P&gt;0.05 and the results are not significant; we fail to reject <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis of equal frequencies. It means that our data is not significantly deviating from the model. A high P value means that the data fits model really well, and is desirable in <\/span>goodness<span style=\"text-align: initial;font-size: 1em\"> of fit analysis. Moving towards left in the table, we can see that even at significance level 0.9, critical \u03c72 0.58 is still less than our obtained \u03c72 test statistic, therefore P value must be &gt;0.9. Almost all of Mendel\u2019s data have unrealistically high P values that lead Fisher to doubt whether these values were real or not!<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Microsoft Excel <\/span>do<span style=\"text-align: initial;font-size: 1em\"> not support \u03c72 test; online calculators like the one below can do the job well.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><a style=\"text-align: initial;font-size: 1em\" href=\"http:\/\/www.socscistatistics.com\/tests\/goodnessoffit\/Default2.aspx\">http:\/\/www.socscistatistics.com\/tests\/goodnessoffit\/Default2.aspx<\/a><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">9.\u00a0<\/strong><strong style=\"text-align: initial;font-size: 1em\">Summary<\/strong><\/p>\r\n\r\n<ol>\r\n \t<li style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">\u03c72 test statistic can be computed by this formula: \u2211[( <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">o<\/em><span style=\"text-align: initial;font-size: 1em\"> - <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\"> )2\/ <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\"> ] Where <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">o<\/em><span style=\"text-align: initial;font-size: 1em\"> is observed frequency and <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em> is<span style=\"text-align: initial;font-size: 1em\"> expected frequency. Both variants of <\/span>\u03c72<span style=\"text-align: initial;font-size: 1em\"> test <\/span>uses<span style=\"text-align: initial;font-size: 1em\"> the same formula.<\/span><\/li>\r\n \t<li style=\"text-align: justify\">For \u03c72<span style=\"text-align: initial;font-size: 1em\"> test of independence, <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\"> is calculated from the table itself. For <\/span>\u03c72<span style=\"text-align: initial;font-size: 1em\"> test of the goodness of fit, <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\"> is derived from a theoretical model and user have to explicitly input the values to the table.<\/span><\/li>\r\n \t<li style=\"text-align: justify\">\u03c72 test of independence is used to find association between two categorical variables. However whenever a cell value is less than 5, this test should not be used. For paired (dependant) values, this test should not be used McNemar\u2019s test is used instead)<\/li>\r\n \t<li style=\"text-align: justify\">\u03c72 test of independence is also used to compare three or more unpaired groups of binomial data. For the analysis of 2 x 2 contingency table of binomial data, a better alternative is Fisher\u2019s Exact Test. For the fit of observed data to a theoretical model involving only one population, binomial<span style=\"text-align: initial;font-size: 1em\"> test is preferred as it returns <\/span>exact<span style=\"text-align: initial;font-size: 1em\"> P value.<\/span><\/li>\r\n \t<li style=\"text-align: justify\">For the fit of observed data to a theoretical model involving more than one population, \u03c72 test of the goodness of fit can be used.<\/li>\r\n<\/ol>\r\n<ol>\r\n \t<li>A web-based Chi-Square calculator for test of independance is available at <a href=\"http:\/\/turner.faculty.swau.edu\/mathematics\/math241\/materials\/contablecalc\/\">http:\/\/turner.faculty.swau.edu\/mathematics\/math241\/materials\/contablecalc\/<\/a><\/li>\r\n<\/ol>\r\n<ol start=\"2\">\r\n \t<li>For chi square test of goodness of fit, visit <a href=\"https:\/\/www.graphpad.com\/quickcalcs\/chisquared1.cfm\">https:\/\/www.graphpad.com\/quickcalcs\/chisquared1.cfm<\/a><\/li>\r\n<\/ol>\r\n<ol start=\"3\">\r\n \t<li>For binomial test, visit <a href=\"https:\/\/www.graphpad.com\/quickcalcs\/binomial1\/\">https:\/\/www.graphpad.com\/quickcalcs\/binomial1\/<\/a><\/li>\r\n<\/ol>\r\n<ol start=\"4\">\r\n \t<li>For fisher\u2019s exact test, visit <a href=\"https:\/\/www.graphpad.com\/quickcalcs\/contingency1.cfm\">https:\/\/www.graphpad.com\/quickcalcs\/contingency1.cfm<\/a><\/li>\r\n<\/ol>\r\n<\/div>","rendered":"<div>\n<p>\u00a0 \u00a0 1.\u00a0<strong>Introduction<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Chi Square <strong>(<\/strong>\u03c72)distribution is an important probability distribution which is used in a number of statistical tests of significance, most famous among which are \u03c72 test of independence and \u03c72 test of the goodness of fit. As a non-parametric, rank-based method, tests based on \u03c72-distribution do not make any assumptions on the normality of the distributions of populations from which the samples came from, and therefore, the test is suitable for the analysis of nominal or categorical samples. However, for the analysis of non-Gaussian data, better non-parametric tests are available (for example, Mann-Whitney U test to compare two unpaired groups, and Kruskal-Wallis test to compare means of three or more unpaired groups). When the outcome is binomial-especially for the analysis of 2&#215;2 contingency tables-Fisher\u2019s exact test is preferred. To compare three or more unpaired groups, \u03c72 test of independence is still the best method. When we want to find the fit of an observed distribution (data) to a theoretically expected distribution (model),\u00a0\u03c72 test of the goodness of fit is performed. The main difference between \u03c72 test of independence and \u03c72 test of the goodness of fit is that while the former automatically calculates expected frequency from the input data of observed frequencies, the latter require input of expected frequencies derived from an explicit model (for example, Mendel\u2019s dihybrid cross ratio, or\u00a0Fisherian sex ratio).<\/p>\n<p>&nbsp;<\/p>\n<p><strong>2.\u00a0<\/strong><strong>Learning Outcome:<\/strong><\/p>\n<p><strong>\u00a0<\/strong><\/p>\n<p style=\"text-align: justify\">a. To learn about the properties of \u03c72-distribution and statistical tests of significance based upon \u03c72 distribution<\/p>\n<p>b. To learn assumptions for \u03c72-tests<\/p>\n<p>c. To learn how \u03c72 test of independence is performed<\/p>\n<p style=\"text-align: justify\">d. To learn how fisher\u2019s exact test is performed, which is a much better alternative to \u03c72 test of independence especially with small sample sizes<\/p>\n<p>e. To learn how \u03c72 test of the goodness of fit is performed<\/p>\n<p>&nbsp;<\/p>\n<p><strong>3.\u00a0<\/strong><strong>\u03c7<\/strong><strong>2<\/strong><strong> distribution<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">\u03c72 distribution (Chi-Square distribution) is a type of asymmetric continuous probability distribution with probabilities of every \u03c72 statistic known under the assumption of null hypothesis. It was first used\u00a0<span style=\"font-size: 1em;text-align: initial\">and described by Karl Pearson in 1900, one of the founding fathers of statistics and Population genetics.\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">This distribution enables us to calculate <\/span>P<span style=\"text-align: initial;font-size: 1em\"> value from a given \u03c72 statistic.<\/span><\/p>\n<\/div>\n<div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-377\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-219.png\" alt=\"\" width=\"177\" height=\"56\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-219.png 177w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-219-65x21.png 65w\" sizes=\"auto, (max-width: 177px) 100vw, 177px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Distribution of these \u03c72 statistic under the assumption of null hypothesis plotted as in a probability histogram is called \u03c72 distribution. Like lognormal distribution and F-distribution, \u03c72 distribution is right-skewed (with a long tail towards right. The shape of \u03c72 -distribution depends only on <em>k<\/em> the shape parameter (degrees of freedom, df). \u03c72 distribution is a special case of more generalized gamma distribution.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-378\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-220.png\" alt=\"\" width=\"553\" height=\"405\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-220.png 553w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-220-300x220.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-220-65x48.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-220-225x165.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-220-350x256.png 350w\" sizes=\"auto, (max-width: 553px) 100vw, 553px\" \/><\/p>\n<p>&nbsp;<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p><strong style=\"text-align: initial;font-size: 1em\">4.\u00a0<\/strong><strong style=\"text-align: initial;font-size: 1em\">Tests of significance based on \u03c7<\/strong><strong style=\"text-align: initial;font-size: 1em\">2<\/strong><strong style=\"text-align: initial;font-size: 1em\"> distribution<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">\u03c72 distribution is used for two main tests; \u03c72 test of independence for the analysis of categorical data (to test whether two categorical variables are correlated), and \u03c72 test of the goodness of fit of an observed distribution (data) to a theoretically expected distribution (model). This distribution is also used for likelihood ratio test (and its variant, hierarchical LRT) used for model selection in molecular phylogenetics<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">5.\u00a0<\/strong><strong style=\"text-align: initial;font-size: 1em\">Assumptions for \u03c7<\/strong><strong style=\"text-align: initial;font-size: 1em\">2<\/strong><strong style=\"text-align: initial;font-size: 1em\"> tests<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">1.\u00a0\u00a0\u00a0\u00a0\u00a0 Random observations<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">2.\u00a0\u00a0\u00a0\u00a0\u00a0 Independent measurements<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">3.\u00a0\u00a0\u00a0\u00a0\u00a0 Accurate data<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Note that there are no explicit assumptions about the distributions of populations from which these samples are drawn, as \u03c72 is considered as a nonparametric test.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">6.<\/strong><span style=\"text-align: initial;font-size: 1em\">\u00a0<\/span><strong style=\"text-align: initial;font-size: 1em\">\u03c7<\/strong><strong style=\"text-align: initial;font-size: 1em\">2<\/strong><strong style=\"text-align: initial;font-size: 1em\"> test of independence<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">This test analyses two sets of information (two variables) for any association between them, therefore in <\/span><em style=\"text-align: initial;font-size: 1em\">sensu stricto<\/em><span style=\"text-align: initial;font-size: 1em\">, this test falls under multivariate statistics. In effect, this test compares two proportions like other methods for comparing proportions such as Relative Risk, Attributable risk, Odd\u2019s Ratio and so on. The null hypothesis is that there is no association between them (these variables are statistically independent), while <\/span>alternative<span style=\"text-align: initial;font-size: 1em\"> hypothesis is that there is an association between them (variables are statistically dependant).<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">The \u03c72 test statistic is computed as<\/span><\/p>\n<div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-379\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-221.png\" alt=\"\" width=\"108\" height=\"62\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-221.png 108w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-221-65x37.png 65w\" sizes=\"auto, (max-width: 108px) 100vw, 108px\" \/><\/p>\n<p style=\"text-align: justify\">Where <em>f<\/em><em>o<\/em> is observed frequency and <em>f<\/em><em>e<\/em> is expected frequency. <em>f<\/em><em>e<\/em> can be calculated as (row total x column total)\/n<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Let us consider an example. Is there any association between income level and happiness level? To study, imagine we have done a questionnaire survey and obtained the results as given below:<\/p>\n<\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-380\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-222.png\" alt=\"\" width=\"710\" height=\"239\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-222.png 710w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-222-300x101.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-222-65x22.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-222-225x76.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-222-350x118.png 350w\" sizes=\"auto, (max-width: 710px) 100vw, 710px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">This is the result of <\/span>questionnaire<span style=\"text-align: initial;font-size: 1em\"> survey with <\/span>total<span style=\"text-align: initial;font-size: 1em\"> number of participants 2955, which is indicated in the table as the overall total. The first cell 272 means out of 615 rich participants, 272 responded that their happiness level is high. These numbers (highlighted in italics) are the observed frequencies (<\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">o<\/em><span style=\"text-align: initial;font-size: 1em\"> in our equation). Note that these numbers are only frequencies; measurement merely measures into three categories (high, middle and low), so the level of measurement here is nominal (categorical). Tables like these where exact measured values are entered <\/span>is<span style=\"text-align: initial;font-size: 1em\"> called contingency tables.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Let us first define our null hypothesis and alternative hypotheses:<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">H0: Income level happiness level are independent<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Ha: Income level and happiness level are <\/span>dependant<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">First<span style=\"text-align: initial;font-size: 1em\"> step in <\/span>\u03c72<span style=\"text-align: initial;font-size: 1em\"> test of independence is to calculate <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\">, the expected frequencies of all these cells. This is computed as:<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\">= (row total x column total)\/ overall total<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">For the first cell (where <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">o<\/em><em style=\"text-align: initial;font-size: 1em\">=272<\/em><span style=\"text-align: initial;font-size: 1em\">), row total is 615 and column total is 911. Plugging into the above equation,<\/span><\/p>\n<p style=\"text-align: justify\"><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\">= (615 x 911)\/2955 =189.6<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">It is better to calculate these values in a tabular format:<\/span><\/p>\n<div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-381 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-223.png\" alt=\"\" width=\"432\" height=\"587\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-223.png 432w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-223-221x300.png 221w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-223-65x88.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-223-225x306.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-223-350x476.png 350w\" sizes=\"auto, (max-width: 432px) 100vw, 432px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>\u03c72 test statistic = 172.28<\/p>\n<p style=\"text-align: justify\">Next step is to look up \u03c72 table to find \u03c72 critical value, for which we should know degree of freedom and significance level (which is 0.05). For \u03c72 tests df can be calculated by the following equation<\/p>\n<p>df = (No. of rows-1) x (No. of columns -1)<\/p>\n<p>Remember that these numbers means that of the actual data; totals or labels are excluded.<\/p>\n<p>Df= (3-1) x (3-1)<\/p>\n<p>=2 x 2<\/p>\n<p>=4<\/p>\n<\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-382 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-224.png\" alt=\"\" width=\"528\" height=\"340\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-224.png 528w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-224-300x193.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-224-65x42.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-224-225x145.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-224-350x225.png 350w\" sizes=\"auto, (max-width: 528px) 100vw, 528px\" \/><\/p>\n<div>\n<p>&nbsp;<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<div>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">As table value (critical \u03c72, 9.488) is far less than our obtained \u03c72 test statistic (172.28), we can conclude that P&lt;0.05, we reject null hypothesis of independence of two variables and conclude that two variables are dependant, or associated. Moving towards right in the table, we can see that even at significance level 0.005, critical \u03c7214.86 is still far less than our obtained \u03c72 test statistic, therefore P value must be &lt;0.005<\/p>\n<p>&nbsp;<\/p>\n<p>There is no support for \u03c72 test in excel. An online calculator like the following can be used instead<\/p>\n<p>&nbsp;<\/p>\n<p><a href=\"http:\/\/turner.faculty.swau.edu\/mathematics\/math241\/materials\/contablecalc\/\">http:\/\/turner.faculty.swau.edu\/mathematics\/math241\/materials\/contablecalc\/<\/a><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The \u03c72 statistic is sensitive to small cell sizes. Whenever any of your cell sizes are &lt;5, a slightly modified formula (Yates\u2019 correction for continuity) to calculate chi square should be used<\/p>\n<p>&nbsp;<\/p>\n<p>Modified Formula (0.5 is deducted from the absolute value of fo \u2013 fe before squaring)<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-383\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-225.png\" alt=\"\" width=\"162\" height=\"60\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-225.png 162w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-225-65x24.png 65w\" sizes=\"auto, (max-width: 162px) 100vw, 162px\" \/><\/p>\n<p>However, most statisticians agree that Yates correction overcorrects it.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">When there are only two categories (like head and tail in coin flipping, or male or female in gender), to make inferences of one population the best option is <\/span>binomial<span style=\"text-align: initial;font-size: 1em\"> test, which <\/span>calculate<span style=\"text-align: initial;font-size: 1em\"> the exact probabilities using binomial equation. <\/span>Binomial<span style=\"text-align: initial;font-size: 1em\"> test is available at <\/span><a style=\"text-align: initial;font-size: 1em\" href=\"https:\/\/www.graphpad.com\/quickcalcs\/binomial1\/\">https:\/\/www.graphpad.com\/quickcalcs\/binomial1\/<\/a><span style=\"text-align: initial;font-size: 1em\"> P values inferred from <\/span>\u03c72<span style=\"text-align: initial;font-size: 1em\"> test are only approximations, not exact.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">7.\u00a0<\/strong><strong style=\"text-align: initial;font-size: 1em\">Fisher\u2019s exact test<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">For 2 x 2 contingency tables used frequently in case control studies, the best test is Fisher\u2019s exact test that can be found here:<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><a style=\"text-align: initial;font-size: 1em\" href=\"https:\/\/www.graphpad.com\/quickcalcs\/contingency1.cfm\">https:\/\/www.graphpad.com\/quickcalcs\/contingency1.cfm<\/a><\/p>\n<\/div>\n<div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-384 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-226.png\" alt=\"\" width=\"390\" height=\"154\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-226.png 390w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-226-300x118.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-226-65x26.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-226-225x89.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-226-350x138.png 350w\" sizes=\"auto, (max-width: 390px) 100vw, 390px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>H0: Apixaban does not alter the risk of a recurrent thromboembolism<\/p>\n<p>Ha: Apixaban alters the risk of a recurrent thromboembolism<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The values in the table indicate No. of patients (treated with placebo or apixaban, two rows) who already had thromboembolism and going on to have another thromboembolism during the study period (in the column \u201drecurrent\u201d) and those who do not have second episode of thromboembolism (in the column \u201cNo Recurrence\u201d). In 2 x 2 contingency tables like this, it is customary to enter groups as rows and outcomes as columns. Fisher\u2019s exact test uses the following formula which is simple and straightforward to understand:<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-385 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-227.png\" alt=\"\" width=\"426\" height=\"104\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-227.png 426w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-227-300x73.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-227-65x16.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-227-225x55.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-227-350x85.png 350w\" sizes=\"auto, (max-width: 426px) 100vw, 426px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Where a, b, c and d are values in 2 x 2 contingency table and n is the total number of values of the table. When the numbers become very large, calculation of factorials becomes mathematically unwieldy so that \u03c72 test is preferred.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Fisher\u2019s exact test for our above example returns a P value less than 0.0001, so the difference is very significant.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">8.<\/strong><span style=\"text-align: initial;font-size: 1em\">\u00a0<\/span><strong style=\"text-align: initial;font-size: 1em\">\u03c7<\/strong><strong style=\"text-align: initial;font-size: 1em\">2<\/strong><strong style=\"text-align: initial;font-size: 1em\"> test of the goodness of fit<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">\u03c72<span style=\"text-align: initial;font-size: 1em\"> test of the goodness of fit is used when we want to find the fit of an observed distribution (data) to a theoretically expected distribution (model). The main difference from the earlier test (independence) is that for <\/span>goodness<span style=\"text-align: initial;font-size: 1em\"> of fit the expected frequencies are derived from a theory or a mathematical model, while in the former, expected frequencies are calculated from the observed frequencies itself. Therefore, for <\/span>test<span style=\"text-align: initial;font-size: 1em\"> of independence, input data is only the observed (empirical) frequencies. In the case of <\/span>test<span style=\"text-align: initial;font-size: 1em\"> of goodness of fit, input data encompasses expected frequencies derived from theory in addition to the observed frequencies.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">The \u03c72 test statistic for <\/span>test<span style=\"text-align: initial;font-size: 1em\"> of goodness of fit is computed exactly as in <\/span>\u03c72<span style=\"text-align: initial;font-size: 1em\"> test of independence:<\/span><\/p>\n<\/div>\n<div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-386\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-228.png\" alt=\"\" width=\"106\" height=\"61\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-228.png 106w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-228-65x37.png 65w\" sizes=\"auto, (max-width: 106px) 100vw, 106px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Where <em>f<\/em><em>o<\/em> is observed frequency and <em>f<\/em><em>e<\/em> is expected frequency. Only difference from \u03c72 test of independence is that <em>f<\/em><em>e<\/em> is not computed from <em>f<\/em><em>o<\/em> but from a model.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Let us consider an example. Out of total 556 pea plants, the famous Geneticist Gregor Mendel observed (dihybrid cross) four seed phenotypes in frequencies given below:<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-387 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-229.png\" alt=\"\" width=\"718\" height=\"207\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-229.png 718w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-229-300x86.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-229-65x19.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-229-225x65.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-229-350x101.png 350w\" sizes=\"auto, (max-width: 718px) 100vw, 718px\" \/><\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">According to his famous law of independent <\/span>assortment<span style=\"text-align: initial;font-size: 1em\"> Mendel expected a certain ratio (9:3:3:1) of those phenotypes. This ratio is a model, <\/span>a theoretically expected proportions<span style=\"text-align: initial;font-size: 1em\">. Let us plot those expected proportions in this table as well:<\/span><\/p>\n<div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-389 alignleft\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-231.png\" alt=\"\" width=\"451\" height=\"226\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-231.png 451w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-231-300x150.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-231-65x33.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-231-225x113.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-231-350x175.png 350w\" sizes=\"auto, (max-width: 451px) 100vw, 451px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">To get Expected frequencies, all we have to do is to multiply each of the expected proportions with the total no. of plants (556). Note that total of expected frequencies add up to the total (556)<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-388 alignleft\" style=\"text-indent: 0px\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-230.png\" alt=\"\" width=\"714\" height=\"252\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-230.png 714w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-230-300x106.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-230-65x23.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-230-225x79.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-230-350x124.png 350w\" sizes=\"auto, (max-width: 714px) 100vw, 714px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>Question is whether observed frequencies deviate significantly from the expected frequencies?<\/p>\n<p>&nbsp;<\/p>\n<p>Let us first define our null hypothesis and alternative hypotheses:<\/p>\n<p>H0: <em>f<\/em><em>o<\/em> = <em>f<\/em><em>e<\/em> (i.e, our data fits model well)<\/p>\n<p>Ha: <em>f<\/em><em>o<\/em> \u2260 <em>f<\/em><em>e<\/em> (i.e, our data do not fits model)<\/p>\n<p>&nbsp;<\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">Now let us complete the \u03c72 table<\/span><\/p>\n<\/div>\n<div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignleft size-full wp-image-390\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-232.png\" alt=\"\" width=\"594\" height=\"348\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-232.png 594w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-232-300x176.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-232-65x38.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-232-225x132.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-232-350x205.png 350w\" sizes=\"auto, (max-width: 594px) 100vw, 594px\" \/><\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<div>\n<p>&nbsp;<\/p>\n<p>\u03c72 test statistic = 0.47<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Next step is to look up \u03c72 table to find \u03c72 critical value, for which we should know degree of freedom and significance level (which is 0.05). For \u03c72 tests of goodness of fit, our data is grouped only in rows, not in columns. So df is (no. of rows \u2013 1)<\/p>\n<p>&nbsp;<\/p>\n<p>4-1 = 3<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignleft size-full wp-image-391\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-233.png\" alt=\"\" width=\"528\" height=\"340\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-233.png 528w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-233-300x193.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-233-65x42.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-233-225x145.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-233-350x225.png 350w\" sizes=\"auto, (max-width: 528px) 100vw, 528px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">As table value (critical \u03c72, 7.815) is far higher than our obtained \u03c72 test statistic (0.47), we can conclude that P&gt;0.05 and the results are not significant; we fail to reject <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis of equal frequencies. It means that our data is not significantly deviating from the model. A high P value means that the data fits model really well, and is desirable in <\/span>goodness<span style=\"text-align: initial;font-size: 1em\"> of fit analysis. Moving towards left in the table, we can see that even at significance level 0.9, critical \u03c72 0.58 is still less than our obtained \u03c72 test statistic, therefore P value must be &gt;0.9. Almost all of Mendel\u2019s data have unrealistically high P values that lead Fisher to doubt whether these values were real or not!<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Microsoft Excel <\/span>do<span style=\"text-align: initial;font-size: 1em\"> not support \u03c72 test; online calculators like the one below can do the job well.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><a style=\"text-align: initial;font-size: 1em\" href=\"http:\/\/www.socscistatistics.com\/tests\/goodnessoffit\/Default2.aspx\">http:\/\/www.socscistatistics.com\/tests\/goodnessoffit\/Default2.aspx<\/a><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">9.\u00a0<\/strong><strong style=\"text-align: initial;font-size: 1em\">Summary<\/strong><\/p>\n<ol>\n<li style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">\u03c72 test statistic can be computed by this formula: \u2211[( <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">o<\/em><span style=\"text-align: initial;font-size: 1em\"> &#8211; <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\"> )2\/ <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\"> ] Where <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">o<\/em><span style=\"text-align: initial;font-size: 1em\"> is observed frequency and <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em> is<span style=\"text-align: initial;font-size: 1em\"> expected frequency. Both variants of <\/span>\u03c72<span style=\"text-align: initial;font-size: 1em\"> test <\/span>uses<span style=\"text-align: initial;font-size: 1em\"> the same formula.<\/span><\/li>\n<li style=\"text-align: justify\">For \u03c72<span style=\"text-align: initial;font-size: 1em\"> test of independence, <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\"> is calculated from the table itself. For <\/span>\u03c72<span style=\"text-align: initial;font-size: 1em\"> test of the goodness of fit, <\/span><em style=\"text-align: initial;font-size: 1em\">f<\/em><em style=\"text-align: initial;font-size: 1em\">e<\/em><span style=\"text-align: initial;font-size: 1em\"> is derived from a theoretical model and user have to explicitly input the values to the table.<\/span><\/li>\n<li style=\"text-align: justify\">\u03c72 test of independence is used to find association between two categorical variables. However whenever a cell value is less than 5, this test should not be used. For paired (dependant) values, this test should not be used McNemar\u2019s test is used instead)<\/li>\n<li style=\"text-align: justify\">\u03c72 test of independence is also used to compare three or more unpaired groups of binomial data. For the analysis of 2 x 2 contingency table of binomial data, a better alternative is Fisher\u2019s Exact Test. For the fit of observed data to a theoretical model involving only one population, binomial<span style=\"text-align: initial;font-size: 1em\"> test is preferred as it returns <\/span>exact<span style=\"text-align: initial;font-size: 1em\"> P value.<\/span><\/li>\n<li style=\"text-align: justify\">For the fit of observed data to a theoretical model involving more than one population, \u03c72 test of the goodness of fit can be used.<\/li>\n<\/ol>\n<ol>\n<li>A web-based Chi-Square calculator for test of independance is available at <a href=\"http:\/\/turner.faculty.swau.edu\/mathematics\/math241\/materials\/contablecalc\/\">http:\/\/turner.faculty.swau.edu\/mathematics\/math241\/materials\/contablecalc\/<\/a><\/li>\n<\/ol>\n<ol start=\"2\">\n<li>For chi square test of goodness of fit, visit <a href=\"https:\/\/www.graphpad.com\/quickcalcs\/chisquared1.cfm\">https:\/\/www.graphpad.com\/quickcalcs\/chisquared1.cfm<\/a><\/li>\n<\/ol>\n<ol start=\"3\">\n<li>For binomial test, visit <a href=\"https:\/\/www.graphpad.com\/quickcalcs\/binomial1\/\">https:\/\/www.graphpad.com\/quickcalcs\/binomial1\/<\/a><\/li>\n<\/ol>\n<ol start=\"4\">\n<li>For fisher\u2019s exact test, visit <a href=\"https:\/\/www.graphpad.com\/quickcalcs\/contingency1.cfm\">https:\/\/www.graphpad.com\/quickcalcs\/contingency1.cfm<\/a><\/li>\n<\/ol>\n<\/div>\n","protected":false},"author":3,"menu_order":20,"template":"","meta":{"pb_show_title":"on","pb_short_title":"","pb_subtitle":"","pb_authors":["dr-felix-bast"],"pb_section_license":""},"chapter-type":[],"contributor":[59],"license":[],"class_list":["post-373","chapter","type-chapter","status-publish","hentry","contributor-dr-felix-bast"],"part":3,"_links":{"self":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapters\/373","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapters"}],"about":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/wp\/v2\/types\/chapter"}],"author":[{"embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/wp\/v2\/users\/3"}],"version-history":[{"count":5,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapters\/373\/revisions"}],"predecessor-version":[{"id":393,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapters\/373\/revisions\/393"}],"part":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/parts\/3"}],"metadata":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapters\/373\/metadata\/"}],"wp:attachment":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/wp\/v2\/media?parent=373"}],"wp:term":[{"taxonomy":"chapter-type","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapter-type?post=373"},{"taxonomy":"contributor","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/wp\/v2\/contributor?post=373"},{"taxonomy":"license","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/wp\/v2\/license?post=373"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}