{"id":340,"date":"2019-03-12T06:07:12","date_gmt":"2019-03-12T06:07:12","guid":{"rendered":"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/?post_type=chapter&#038;p=340"},"modified":"2019-03-12T06:25:24","modified_gmt":"2019-03-12T06:25:24","slug":"statistical-hypothesis-testing-and-p-values","status":"publish","type":"chapter","link":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/chapter\/statistical-hypothesis-testing-and-p-values\/","title":{"rendered":"Statistical Hypothesis Testing and P Values"},"content":{"raw":"<div>\r\n\r\n&nbsp;\r\n\r\n<strong>1.\u00a0<\/strong><strong>Introduction<\/strong>\r\n\r\n<strong>\u00a0<\/strong>\r\n<p style=\"text-align: justify\">Statistical hypothesis testing is the current gold standard of scientific methodology and is a key concept in inferential statistics. This involves defining two contrasting hypotheses, calculating P value from the sample data and deciding the fate of the two hypotheses based on the P value and the threshold cut-offs significance level. Most common forms of errors in hypothesis testing are Type I and Type II errors. These errors can be minimized with appropriate choice of significance level. Most statisticians agree that P value-based significance testing is over rated, and a simpler representation of 95% CI is a better approach in most of the cases. A number of ways P values can be decreased, the so called P-hacking- is common across scientific disciplines and it construes a scientific misconduct.<\/p>\r\n<strong>\u00a0<\/strong>\r\n\r\n<strong>2.\u00a0<\/strong><strong>Learning Outcome:<\/strong>\r\n\r\n<strong>\u00a0<\/strong>\r\n\r\na)\u00a0\u00a0\u00a0\u00a0\u00a0 To learn concepts statistical hypothesis testing\r\n\r\nb)\u00a0\u00a0\u00a0\u00a0\u00a0 To learn significance levels and choosing the right significance level\r\n\r\nc)\u00a0\u00a0\u00a0\u00a0\u00a0 To learn about Type I, type II and type III errors\r\n\r\nd)\u00a0\u00a0\u00a0\u00a0\u00a0 To learn about P values and discern a number of P value fallacies\r\n\r\ne)\u00a0\u00a0\u00a0\u00a0\u00a0 To learn about statistical Power and False discovery Rate\r\n\r\nf)\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 To learn the relationship between Confidence Intervals and P value\r\n\r\n&nbsp;\r\n\r\n<strong>3.\u00a0<\/strong><strong>Statistical Hypothesis Testing<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Scientific methodology was developed with Francis Bacon\u2019s \u201cNovum Organum\u201d published in 1620. Subsequently, Karl Popper introduced concept of falsifiability that the scientific claims have to be independently verifiable, and Thomas Kuhn introduced concept of Paradigm Shift that the science progresses through unconventional \u2018revolution\u2019 of ideas. How do we distinguish science from pseudoscience? The current gold standard for experimental scientific disciplines is statistical hypothesis testing, which is very intuitive to comprehend with a simple analogy.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Let us design a very simple experiment that tests how \u2018lucky\u2019 you are at any particular moment by coin toss. Effectively by doing this coin toss experiment you are testing two contrasting hypotheses: 1) You\u00a0<span style=\"text-align: initial;font-size: 1em\">are not lucky and 2) you are lucky. There will be ten toss trials and even before the first coin is flipped, you decide that if you get 6 heads or more, you are deemed lucky. After flipping the coin for 10 times, you got 6 heads and 4 tails. Conclusion? Hypothesis No. 1 that \u2018you are not lucky\u2019 is rejected (beware of double negatives!) and conclude that you are lucky. Suppose you got only 5 heads, then your conclusion would be you \u2018fail to reject hypothesis 1- that you are not lucky\u2019 (a triple negative logic jargon to say that you are not lucky). If you get 5 heads, instead of concluding you are not lucky, you take a correction pen and change your original design statement from \u20186 heads or more\u2019 to \u20185 heads or more\u2019, and conclude that you are lucky! Of course, this is not fair; this is cheating and this form of misconduct is very common in scientific research. Another option is instead of concluding you are lucky with 5 heads in 10 toss trials, you increase no. of trials; you go ahead tossing coin 11th time, 12th and so on till you get 6 heads on <\/span>13th<span style=\"text-align: initial;font-size: 1em\"> time, and suddenly stop the experiment there to conclude that you are lucky. Again, you take a correction pen and change the original proposition from \u201cThere will be ten toss trials\u201d to \u201cThere will be thirteen toss trials\u201d. Of course, this is cheating too. What is so special about 6 heads? Nothing, really. It is our threshold value for crisply deciding success from failure.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Statistical hypothesis testing follows the following four steps sequentially:<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">1.\u00a0\u00a0\u00a0\u00a0\u00a0 Define alpha (\u03b1 )-threshold P value<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">2.\u00a0\u00a0\u00a0\u00a0\u00a0 Define <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis (H0)<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">3.\u00a0\u00a0\u00a0\u00a0\u00a0 Define <\/span>alternative<span style=\"text-align: initial;font-size: 1em\"> hypothesis (Ha)<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">4.\u00a0\u00a0\u00a0\u00a0\u00a0 Calculate P value, and decide the fate of hypotheses<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">First<span style=\"text-align: initial;font-size: 1em\"> step is defining the threshold alpha. This threshold is exactly like the threshold of 6 heads we used in our example. 6 heads out of 10 trials constitute 60% \u2018confidence,\u2019 so if you get 6 heads or more, you have 60% confidence to say you are lucky. However, experimental scientists want a lot more confidence in their conclusion to be independently verifiable (Karl Popper\u2019s falsifiability\u2019). Typically they choose a confidence level of 95% (or 9.5 out of 10). This would mean 5% tolerance <\/span>of<span style=\"text-align: initial;font-size: 1em\"> making incorrect mistakes. The tolerance for making incorrect mistakes is what is called threshold P value or alpha. P in P value <\/span>stand<span style=\"text-align: initial;font-size: 1em\"> for Probability, and it is always expressed as a number between 0 and 1. A 5% tolerance for error expressed percentage is 0.05 tolerance level expressed in probability (or P-value). In practice, the threshold value (called alpha) is almost always set to 0.05 (an arbitrary value that has been widely\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">adopted). Ideally, you should set this value based on the relative consequences of falsely finding a difference (False Positive) or missing a true difference (False Negative), more on this to be followed.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is usually the opposite of what you are trying to prove. In our first example, if you were trying to test the proposition that you are lucky. <\/span>Null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is that you are not lucky. If you are testing the efficacy of a drug, <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is that the drug is ineffective. If you are comparing <\/span>means<span style=\"text-align: initial;font-size: 1em\"> of two groups, <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is that their group means are equal. For clinical testing, <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is that the result is negative. For spam detection in <\/span>email<span style=\"text-align: initial;font-size: 1em\">, null is that email is not spam.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Alternative<span style=\"text-align: initial;font-size: 1em\"> hypothesis is the hypothesis what you are usually trying to prove. In our earlier examples, alternate hypotheses <\/span>is<span style=\"text-align: initial;font-size: 1em\"> that you are lucky, or drug is effective, or group means are not equal, or clinical test result is positive, or email is spam.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">After completing your experiment and after performing appropriate statistical tests, you get a P value. It is important to note that P values and statistical tests only assess the validity of our null hypothesis defined in step 2 above (they do not test the validity of our alternate hypothesis directly). Therefore, we can make conclusions only about the null hypothesis. If you get a P value (for example, 0.04) which is less than our predefined level of tolerance for errors (0.05) defined in step 1, you may conclude to \u2018reject the null hypothesis\u2019 and state that the results are statistically significant. If your obtained P value (for example, 0.06) is higher than significance level alpha, then the conclusion would be \u2018not to reject the null hypothesis\u2019 and that the results are not statistically significant. Remember that statistically significant does not mean scientifically or clinically significant. You cannot conclude that the null hypothesis is true. All you can do is conclude that you \u2018don't have sufficient evidence to reject the null hypothesis\u2019. This sort of arguments <\/span>are<span style=\"text-align: initial;font-size: 1em\"> called arguments by contradiction often used in logic and philosophy.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">An intuitive analogy is <\/span>the similar<span style=\"text-align: initial;font-size: 1em\"> situation in <\/span>jurisdiction<span style=\"text-align: initial;font-size: 1em\">. Here H0 is accused is innocent and Ha is accused is criminal. A juror (judge) can pronounce the verdict as guilty (reject the null hypothesis) or not guilty (do not reject the null hypothesis). She can never pronounce that the accused is criminal (accept the alternative hypothesis) or innocent (accept the null hypothesis). Also, the juror\u2019s conclusion is crisp and clear: guilty or not guilty. She cannot say that the accused is \u2018probably guilty\u2019 or \u2018I don\u2019t know\u2019, as no one likes uncertainty in <\/span>judgements<span style=\"text-align: initial;font-size: 1em\">. Similarly, in statistical hypothesis testing conclusion\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">is crisp and clear as it is based on the threshold significance level. A test that returns P value of 0.4999 should be deemed significant while <\/span>P<span style=\"text-align: initial;font-size: 1em\"> value of 0.5001 is not significant. Many statisticians <\/span>criticise<span style=\"text-align: initial;font-size: 1em\"> making this sort of crisp conclusions based on threshold P value and overly emphasizing the term significant in papers. Many investigators do not report the exact P value (like P=0.499, which is in fact very much borderline) and instead make ambiguous statements like \u2018results are significant (P&lt;0.05)\u2019.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">5.<\/strong><span style=\"text-align: initial;font-size: 1em\">\u00a0<\/span><strong style=\"text-align: initial;font-size: 1em\">Type 1 and 2 errors<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">As already stated, a significance level of 0.05 means 95% confidence level and tolerance of 5% error level. Even though this 5% level is very less, 5% of all results will be erroneous. Out of 100 statistical tests, 5 tests would lead to <\/span>erroneous<span style=\"text-align: initial;font-size: 1em\"> conclusion that the result is significant, while in <\/span>reality<span style=\"text-align: initial;font-size: 1em\"> the result would not be significant. For example, when you compare means of two groups, you got a P value &lt;0.05, therefore reject the null hypothesis and conclude that two means are significantly different, but in <\/span>reality<span style=\"text-align: initial;font-size: 1em\"> means are not different (null hypothesis of no difference in means is wrongly rejected). Another example is that you go for an ELISA blood test to see whether the test is positive or negative. The test report says positive (the null hypothesis that you have no HIV infection is rejected), but this result is erroneous. You will never know is it an error or not unless you repeat the test many times or you already know the reality (for example, in simulations). This type of errors where <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is wrongly rejected (incorrectly rejecting the null hypothesis) is called type-1 error.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">We can also have yet another kind of error. The test report there is no difference in two group means, but in <\/span>reality<span style=\"text-align: initial;font-size: 1em\"> there is <\/span>difference<span style=\"text-align: initial;font-size: 1em\">. A person is really HIV positive, but the test says he is HIV negative (\u2018false negative\u2019). A person is really guilty but the juror pronounces the verdict as not guilty. The email is truly spam but the spam filter says email as not spam. This type of error is called Type II error or False Negative, which is incorrectly not rejecting the null hypothesis. You've made a type II error when you wrongly do not reject <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis. There really is a difference (association, correlation) overall, but random sampling caused your data to not show a statistically significant difference. So your conclusion that the two groups are not really different is incorrect.<\/span><\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n<img class=\"aligncenter size-full wp-image-344\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-202.png\" alt=\"\" width=\"763\" height=\"204\" \/>\r\n\r\n<\/div>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">It is important to note that we will never know is it <\/span>error<span style=\"text-align: initial;font-size: 1em\"> or no error when we <\/span>analyse<span style=\"text-align: initial;font-size: 1em\"> the <\/span>sample,<span style=\"text-align: initial;font-size: 1em\"> unless the analysis is <\/span>simulation<span style=\"text-align: initial;font-size: 1em\"> in which the true population values and reality is known precisely, which is almost impossible when <\/span>analysing<span style=\"text-align: initial;font-size: 1em\"> the scientific experimental data.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">There is yet another type of errors, Type III or type S. This would happen when <\/span>direction<span style=\"text-align: initial;font-size: 1em\"> of effect goes opposite the direction of prediction (especially prone if you use one tail P values to be explained soon). An example of such an error happened in <\/span>Cardiac<span style=\"text-align: initial;font-size: 1em\"> Arrhythmia Suppression Trial (CAST). Anti-arrhythmia drugs (like atenolol to control atrial and ventricular fibrillations) were previously thought to either prolong the life of the patients or no effect, so the investigators chose to report one tail P value. They got a high one tail P value, so concluded the drug has \u2018no effect\u2019 on prolonging the life. <\/span>Actually<span style=\"text-align: initial;font-size: 1em\"> patients given drugs were four times more likely to die than placebo, so the effect went in <\/span>opposite<span style=\"text-align: initial;font-size: 1em\"> direction!<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">6.<\/strong><span style=\"text-align: initial;font-size: 1em\">\u00a0 <\/span><strong style=\"text-align: initial;font-size: 1em\">Choosing the level of significance<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">In the table above, quadrat denoted as [A] signifies false positives. Sum of quadrats [A] and [B] is the total of <\/span>first<span style=\"text-align: initial;font-size: 1em\"> row, \u2018null hypothesis is true\u2019. A\/(A+B) is \u03b1, level of significance. It is \u2018Fraction of false positives when <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is true\u2019. If the null hypothesis is true, what is the probability of incorrectly rejecting it? That probability is the significance level.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">As already explained, a significance level (\u03b1 ) of 0.05 (95% confidence) is chosen almost universally in experimental scientific disciplines, especially in biological and environmental sciences. This traditional significance level is sometimes called \u201ctwo sigma level\u201d, as 95% of all values are expected within two standard deviations from the sample mean in normal distributions (<\/span>central<span style=\"text-align: initial;font-size: 1em\"> limit theorem of statistics).\u00a0<\/span>Ideally<span style=\"text-align: initial;font-size: 1em\"> this level should be decided based upon the consequence of two main kinds of errors, type I and type II errors explained earlier.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Use low \u03b1 (like 0.01), to have less chance of making Type I error (False Positive), but beware, <\/span>chance<span style=\"text-align: initial;font-size: 1em\"> of making Type II error (False Negative) is increased! The consequence of setting \u03b1=0.01 would be more Type II errors (False Negatives). Examples of type II errors include more spam in <\/span>inbox<span style=\"text-align: initial;font-size: 1em\">, a criminal is set free, the test says negative for <\/span>HIV+<span style=\"text-align: initial;font-size: 1em\"> person, A doped athlete is set free, and an effective drug is declared as ineffective and abort drug development. If these sort of false negatives are tolerable, choosing low \u03b1 is justified. Use low \u03b1 (like 0.01) to minimize false positives and is ideal for situations when false negatives are tolerated like Criminal jurisdiction (\u201cit is better to let many guilty <\/span>person<span style=\"text-align: initial;font-size: 1em\"> go free than to falsely convict one innocent person\u201d), Spam Filter (it is better to have spam emails in inbox rather than automatically detecting as spam and deleting that \u2018article acceptance\u2019 or \u2018job selection\u2019 email!), and clinical trial for \u2018me-too\u2019 drug (it is better to abort an effective drug development than marketing an ineffective, useless drug when better alternatives are already available)<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Use high \u03b1 (like 0.1), to have less chance of making Type II error (False Negative), but beware, <\/span>chance<span style=\"text-align: initial;font-size: 1em\"> of making Type I error (False Positive) is increased! The consequence of setting \u03b1=0.1 would be more Type I errors (False Positives). ). Examples of type I errors include: Good email declared as spam and deleted, An innocent is punished, The test says positive for HIV- person, An innocent athlete is declared as doped and banned for life, and an ineffective drug is declared as effective and market it. Use high \u03b1 (like 0.1) to minimize false negatives and is ideal for situations when false positives are tolerated like Civil jurisdiction (it is better to punish many innocent than let one huge illegal corporation go free), and clinical trial for novel drug (it is better to market an ineffective, useless drug than aborting the development of an effective drug when no drugs are available in the market).<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Please be noted that traditional two sigma level of 0.05 is not universally followed in other disciplines, especially in particle physics. To conclude evidence of detecting a (previously known) particle, accepted level of significance is 3 sigma (\u03b1=0.003). To conclude evidence of the discovery of a new kind of particle (like in the paper describing the discovery of Higg\u2019s Boson), <\/span>accepted<span style=\"text-align: initial;font-size: 1em\"> level of significance is 5 sigma (\u03b1=0.0000003).<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">7.\u00a0\u00a0 P-values: Correct interpretations and fallacies<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Definition of P value starts with the proposition \u2018If the null hypothesis is true\u2019. The definition is \u201cIf the null hypothesis is true, what is the probability that random sampling would lead to <\/span>difference<span style=\"text-align: initial;font-size: 1em\"> as large as or larger than that observed in this study?\u201d<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Let us consider an example. A study compared <\/span>mean<span style=\"text-align: initial;font-size: 1em\"> of two <\/span>groups,<span style=\"text-align: initial;font-size: 1em\"> and got a P value of 0.03.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">That means if two population means are identical (i.e., if the null hypothesis is true), there is a 3% chance of observing a difference as large as or larger than what you observed. As that chance is very low 9only 3 out of 100), you conclude that the means are not identical. Alternatively, you can conclude that random sampling from identical populations would lead to a difference smaller than you observed in 97% of experiments, and larger than you observed in 3% of experiments. P value answers the question: \u201cIn an experiment of this size, if the populations really have the same mean, what is the probability of observing at least as large a difference between sample means as was, in fact, observed?\u201d<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">P values are indeed confusing to many scientists because <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is almost always false. Logic works <\/span>backwards<span style=\"text-align: initial;font-size: 1em\"> (Population to Sample, Logical Deduction akin to Probability) and is counterintuitive; the hypothesis testing works on the principle of argument by contradiction.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">There are two types of P values, two-tail <\/span>and<span style=\"text-align: initial;font-size: 1em\"> one tail. A two-tail P value indicates <\/span>critical<span style=\"text-align: initial;font-size: 1em\"> region extends to both extremes of distribution (significantly large or significantly small). A one-tail P value <\/span>extend<span style=\"text-align: initial;font-size: 1em\"> only to either of the two tails of <\/span>distribution<span style=\"text-align: initial;font-size: 1em\">. A difference between two can easily be spotted by looking at your null hypothesis. If the null hypothesis contains the symbol = explicitly or indirectly (like \u2018there are no differences between means\u2019 is equivalent to \u2018mean of group x = mean of group y\u2019), P value should be two-tail. If the null hypothesis <\/span>contain<span style=\"text-align: initial;font-size: 1em\"> the symbol &gt; or &lt; explicitly or indirectly (like \u201cthere is no fever\u201d is equivalent to \u201ctemperature \u226437\u00b0C\u201d or \u201cthere is no profit\u201d is equivalent to \u201cnet gain &gt; net loss\u201d), P value should be one-tail. A one tail P value explicitly make predictions of the directionality of effect as part of the experimental design. This predicted directionality might very well be wrong too, that lead to type III error as explained earlier. When in doubt, always stick with two-tail P values. One tail P values are always less than Two tail P values, so investigators like One tail P to prove their hypothesis! Some reviewers criticize any use of one-tail P values, no matter how well justification is.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">The following is a list of P value fallacies (incorrect interpretations):<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">1.\u00a0\u00a0\u00a0\u00a0\u00a0 The <\/span><em style=\"text-align: initial;font-size: 1em\">P<\/em><span style=\"text-align: initial;font-size: 1em\"> value is the probability of rejecting the null hypothesis<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">2.\u00a0\u00a0\u00a0\u00a0\u00a0 A high <\/span><em style=\"text-align: initial;font-size: 1em\">P<\/em><span style=\"text-align: initial;font-size: 1em\"> value proves that the null hypothesis is true.<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">3.\u00a0\u00a0\u00a0\u00a0\u00a0 1-<\/span><em style=\"text-align: initial;font-size: 1em\">P<\/em><span style=\"text-align: initial;font-size: 1em\"> is the probability that the results will hold up when the experiment is repeated<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">4.\u00a0\u00a0\u00a0\u00a0\u00a0 1-<\/span><em style=\"text-align: initial;font-size: 1em\">P<\/em><span style=\"text-align: initial;font-size: 1em\"> is the probability that the alternative hypothesis is true<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">5.\u00a0\u00a0\u00a0\u00a0\u00a0 The <\/span><em style=\"text-align: initial;font-size: 1em\">P<\/em><span style=\"text-align: initial;font-size: 1em\"> value is the probability that the null hypothesis is true<\/span><\/p>\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">6.\u00a0\u00a0\u00a0\u00a0\u00a0 <\/span><em style=\"text-align: initial;font-size: 1em\">P <\/em><span style=\"text-align: initial;font-size: 1em\">value is the probability that the result was due to sampling error<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Be noted that all six of the above propositions are wrong interpretations of P values.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">8.<\/strong><span style=\"text-align: initial;font-size: 1em\">\u00a0\u00a0\u00a0 <\/span><strong style=\"text-align: initial;font-size: 1em\">P-hacking<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Remember the coin toss experiment we described <\/span>in the beginning<span style=\"text-align: initial;font-size: 1em\"> to test how lucky a person is? Two ways the person can cheat are by changing the threshold value and increasing the no of trials. Similarly, experimental scientists unethically resort to many practices to decrease the P value (a low P value is desirable, as a P value lesser than the alpha (P&lt;0.05) would reject the null hypothesis and lead to a conclusion of statistical significance. Some of these tricks include:<\/span><\/p>\r\n\r\n<ol>\r\n \t<li style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Try different tests, parametric and nonparametric and picking the one with <\/span>lowest<span style=\"text-align: initial;font-size: 1em\"> P value\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">The statistical test that will be used (for example, unpaired t-test) should be clearly specified as part of the experimental design. It is not fair to try different tests and go with the one that gives <\/span>lowest<span style=\"text-align: initial;font-size: 1em\"> P value.<\/span><\/li>\r\n \t<li style=\"text-align: justify\">Dynamic sample size and stop when P is lower than threshold, Sample size should clearly be decided as part of the experimental design. It is not fair to change this till<span style=\"text-align: initial;font-size: 1em\"> P value is less than the threshold.<\/span><\/li>\r\n \t<li style=\"text-align: justify\">Slice and dice the dataset to get lowest P, Obviously an unethical practice. We should use all of the elements of our dataset without slicing and dicing.<\/li>\r\n \t<li style=\"text-align: justify\">Cherry-picking the dataset to get lowest P Obviously an unethical practice. We should use all of the elements of our dataset without picking certain values that affirms<span style=\"text-align: initial;font-size: 1em\"> our preconceived conclusions.<\/span><\/li>\r\n \t<li style=\"text-align: justify\">Play with outliers, remove one or more selectively to get desired<span style=\"text-align: initial;font-size: 1em\"> P<\/span><\/li>\r\n<\/ol>\r\n<div>\r\n<p style=\"text-align: justify\">\u00a0 \u00a0Obviously an unethical practice. We should use all of the elements of our dataset without removing any outliers manually. In case outliers need to be removed, a formal statistic test such as Grubb\u2019s test can be used for this.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">All of the above are cheating, and construes scientific misconduct. As already mentioned, statistical hypothesis testing based on P values is based upon the threshold significance level, which is nothing but an arbitrary value. Statistical significance is usually denoted by the symbol * in graphs. Over emphasizing the statistical significance and * is sometimes referred as \u2018stargazing\u2019. P values depend largely on sample size. It is easy to get very low P values with huge sample size with tiny effect, or tiny sample size with huge effect. P value answers the question, is there a significant evidence of difference? That question is different from \u201cIs there an evidence of significant difference?\u201d. P value does not answer that question. Difference between sample means might be tiny, or scientifically negligible, yet decision based on P value could be \u2018statistically significant\u2019. Add on to the woes, many studies have revealed that P values are not very well reproducible. Instead of P values, the current consensus among the statisticians is to used 95% Confidence Intervals.<\/p>\r\n&nbsp;\r\n\r\n<strong>9.\u00a0\u00a0 P-values and Confidence Intervals<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Confidence Intervals and P values are very well connected. A 95% Confidence Interval is similar to 0.05 significance level as already explained. The question is does the range of values defined by 95% Confidence Interval include our null hypothesis? For example, while comparing the difference in mean between two groups. Null hypothesis is that the group means are same, or the difference between two group means is zero. If 95% Confidence Intervals of difference between sample means does not include zero-our null Hypothesis, then P&lt;0.05 (inversely, if it includes zero, then P&gt;0.05) and indicate a statistically significant difference. Let us consider another example. A random sample consisting of 10 individuals volunteered to have their body temperatures measured to know whether they have normal body temperature or not. Mean body temperature was found to be 36.8\u00b0C. The width of 95% Confidence\u00a0<span style=\"text-align: initial;font-size: 1em\">Interval was found to be 0.3\u00b0C and therefore 95% CI about the sample mean ranges from (36.8-0.3) to (36.8+0.3) that is 36.5\u00b0C to 37.1\u00b0C. Our null hypothesis of having no fever is 37\u00b0C. Does our obtained 95% CI include the null hypothesis of 37\u00b0C? Yes, it does. So, P&gt;0.05 and we can conclude that there is no statistically significant difference from the null hypothesis.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">10. FDR and Statistical Power<\/strong><\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n<img class=\"aligncenter size-full wp-image-345\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-203.png\" alt=\"\" width=\"761\" height=\"202\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">We have already defined A\/(A+B) as \u03b1, level of significance. . It is \u2018Fraction of false positives when null hypothesis is true\u2019. A related term is False Discovery Rate (FDR), which is A\/A+C. FDR is defined as \u2018Fraction of false positives out of statistically significant results\u2019. If the result is statistically significant, what is the probability that null hypothesis is really true? Note that for both level of significance and FDR, numerator remains same,\u2019 fraction of false positives\u2019, but denominator is different. While significance level is based out of all results where null hypothesis is true (i.e., all results that in reality have no effect\/difference etc.), FDR is based out of results with a \u2018statistically significant\u2019 conclusion. FDR depends upon \u2018prior probability\u2019-the context of experiment akin to Bayesian Statistics. FDR and Prior Probability will be elaborated in the module of Bayesian Statistics. Another related term is \u2018statistical power\u2019 which is C\/C+D. It is \u2018fraction of statistically significant results out of all results where null hypothesis is false\u2019. The statistical power is the probability to obtain a statistically significant result assuming a certain effect size in population. Power depends on sample size, variability, choice of \u03b1 and hypothetical effect size. A high power is obtained when: 1) Large sample size is used, 2) Looking for a large effect (or large difference or a strong correlation etc), and 3) Data with little scatter (small variance and SD).<\/p>\r\n\r\n<\/div>\r\n<ol start=\"11\">\r\n \t<li><strong>Summary<\/strong><\/li>\r\n<\/ol>\r\n<p style=\"text-align: justify\">\u00a0 \u00a0a.\u00a0Statistical<span style=\"font-size: 1em\"> hypothesis is a <\/span>four step<span style=\"font-size: 1em\"> sequential process. Significance level should be defined first, then null and alternative hypotheses must be defined. Finally, based on obtained P value, the fate of two hypotheses is decided. It is cheating to change the decided significance <\/span>level,<span style=\"font-size: 1em\"> or changing the sample size and other numerous ways P values can be hacked.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"font-size: 1em\">b. Significance level should be chosen based upon tolerance limit to two types of errors; type I (false positives) and type II (false negatives)<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"font-size: 1em\">c. P value is defined as \u201cIf the null hypothesis is true, what is the probability that random sampling would lead to <\/span>difference<span style=\"font-size: 1em\"> as large as or larger than that observed in this study?\u201d It is more essential to know how P values are correctly interpreted rather than how P values are calculated.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"font-size: 1em\">d. A low P value does not mean differences are scientifically significant, or finding interesting or warrant further funding<\/span>. .<span style=\"font-size: 1em\"> P value answers the question, is there <\/span>a significant<span style=\"font-size: 1em\"> evidence of difference? Not \u201cIs there <\/span>an evidence<span style=\"font-size: 1em\"> of significant difference?\u201d.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"font-size: 1em\">e. If 95% CI does not include <\/span>null<span style=\"font-size: 1em\"> hypothesis, then P value must be &lt;0.05. If the range does include <\/span>null<span style=\"font-size: 1em\"> hypothesis, then P&gt;0.05. This connection is important.<\/span><\/p>\r\n&nbsp;\r\n\r\n<strong>Quadrant-III: Learn More\/ Web Resources \/ Supporting Materials:<\/strong>\r\n\r\n1. Brief review of concepts of Statistical Hypothesis testing at PennState:\r\n\r\nhttps:\/\/onlinecourses.science.psu.edu\/statprogram\/node\/138\r\n\r\n2. A good overview of hypothesis testing in GraphPad guide:\r\n\r\nhttps:\/\/www.graphpad.com\/guides\/prism\/7\/statistics\/index.htm?statistical_hypothesis_testing.htm\r\n\r\n3. P values- an introduction https:\/\/www.statsdirect.com\/help\/basics\/p_values.htm","rendered":"<div>\n<p>&nbsp;<\/p>\n<p><strong>1.\u00a0<\/strong><strong>Introduction<\/strong><\/p>\n<p><strong>\u00a0<\/strong><\/p>\n<p style=\"text-align: justify\">Statistical hypothesis testing is the current gold standard of scientific methodology and is a key concept in inferential statistics. This involves defining two contrasting hypotheses, calculating P value from the sample data and deciding the fate of the two hypotheses based on the P value and the threshold cut-offs significance level. Most common forms of errors in hypothesis testing are Type I and Type II errors. These errors can be minimized with appropriate choice of significance level. Most statisticians agree that P value-based significance testing is over rated, and a simpler representation of 95% CI is a better approach in most of the cases. A number of ways P values can be decreased, the so called P-hacking- is common across scientific disciplines and it construes a scientific misconduct.<\/p>\n<p><strong>\u00a0<\/strong><\/p>\n<p><strong>2.\u00a0<\/strong><strong>Learning Outcome:<\/strong><\/p>\n<p><strong>\u00a0<\/strong><\/p>\n<p>a)\u00a0\u00a0\u00a0\u00a0\u00a0 To learn concepts statistical hypothesis testing<\/p>\n<p>b)\u00a0\u00a0\u00a0\u00a0\u00a0 To learn significance levels and choosing the right significance level<\/p>\n<p>c)\u00a0\u00a0\u00a0\u00a0\u00a0 To learn about Type I, type II and type III errors<\/p>\n<p>d)\u00a0\u00a0\u00a0\u00a0\u00a0 To learn about P values and discern a number of P value fallacies<\/p>\n<p>e)\u00a0\u00a0\u00a0\u00a0\u00a0 To learn about statistical Power and False discovery Rate<\/p>\n<p>f)\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 To learn the relationship between Confidence Intervals and P value<\/p>\n<p>&nbsp;<\/p>\n<p><strong>3.\u00a0<\/strong><strong>Statistical Hypothesis Testing<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Scientific methodology was developed with Francis Bacon\u2019s \u201cNovum Organum\u201d published in 1620. Subsequently, Karl Popper introduced concept of falsifiability that the scientific claims have to be independently verifiable, and Thomas Kuhn introduced concept of Paradigm Shift that the science progresses through unconventional \u2018revolution\u2019 of ideas. How do we distinguish science from pseudoscience? The current gold standard for experimental scientific disciplines is statistical hypothesis testing, which is very intuitive to comprehend with a simple analogy.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Let us design a very simple experiment that tests how \u2018lucky\u2019 you are at any particular moment by coin toss. Effectively by doing this coin toss experiment you are testing two contrasting hypotheses: 1) You\u00a0<span style=\"text-align: initial;font-size: 1em\">are not lucky and 2) you are lucky. There will be ten toss trials and even before the first coin is flipped, you decide that if you get 6 heads or more, you are deemed lucky. After flipping the coin for 10 times, you got 6 heads and 4 tails. Conclusion? Hypothesis No. 1 that \u2018you are not lucky\u2019 is rejected (beware of double negatives!) and conclude that you are lucky. Suppose you got only 5 heads, then your conclusion would be you \u2018fail to reject hypothesis 1- that you are not lucky\u2019 (a triple negative logic jargon to say that you are not lucky). If you get 5 heads, instead of concluding you are not lucky, you take a correction pen and change your original design statement from \u20186 heads or more\u2019 to \u20185 heads or more\u2019, and conclude that you are lucky! Of course, this is not fair; this is cheating and this form of misconduct is very common in scientific research. Another option is instead of concluding you are lucky with 5 heads in 10 toss trials, you increase no. of trials; you go ahead tossing coin 11th time, 12th and so on till you get 6 heads on <\/span>13th<span style=\"text-align: initial;font-size: 1em\"> time, and suddenly stop the experiment there to conclude that you are lucky. Again, you take a correction pen and change the original proposition from \u201cThere will be ten toss trials\u201d to \u201cThere will be thirteen toss trials\u201d. Of course, this is cheating too. What is so special about 6 heads? Nothing, really. It is our threshold value for crisply deciding success from failure.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Statistical hypothesis testing follows the following four steps sequentially:<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">1.\u00a0\u00a0\u00a0\u00a0\u00a0 Define alpha (\u03b1 )-threshold P value<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">2.\u00a0\u00a0\u00a0\u00a0\u00a0 Define <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis (H0)<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">3.\u00a0\u00a0\u00a0\u00a0\u00a0 Define <\/span>alternative<span style=\"text-align: initial;font-size: 1em\"> hypothesis (Ha)<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">4.\u00a0\u00a0\u00a0\u00a0\u00a0 Calculate P value, and decide the fate of hypotheses<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">First<span style=\"text-align: initial;font-size: 1em\"> step is defining the threshold alpha. This threshold is exactly like the threshold of 6 heads we used in our example. 6 heads out of 10 trials constitute 60% \u2018confidence,\u2019 so if you get 6 heads or more, you have 60% confidence to say you are lucky. However, experimental scientists want a lot more confidence in their conclusion to be independently verifiable (Karl Popper\u2019s falsifiability\u2019). Typically they choose a confidence level of 95% (or 9.5 out of 10). This would mean 5% tolerance <\/span>of<span style=\"text-align: initial;font-size: 1em\"> making incorrect mistakes. The tolerance for making incorrect mistakes is what is called threshold P value or alpha. P in P value <\/span>stand<span style=\"text-align: initial;font-size: 1em\"> for Probability, and it is always expressed as a number between 0 and 1. A 5% tolerance for error expressed percentage is 0.05 tolerance level expressed in probability (or P-value). In practice, the threshold value (called alpha) is almost always set to 0.05 (an arbitrary value that has been widely\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">adopted). Ideally, you should set this value based on the relative consequences of falsely finding a difference (False Positive) or missing a true difference (False Negative), more on this to be followed.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is usually the opposite of what you are trying to prove. In our first example, if you were trying to test the proposition that you are lucky. <\/span>Null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is that you are not lucky. If you are testing the efficacy of a drug, <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is that the drug is ineffective. If you are comparing <\/span>means<span style=\"text-align: initial;font-size: 1em\"> of two groups, <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is that their group means are equal. For clinical testing, <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is that the result is negative. For spam detection in <\/span>email<span style=\"text-align: initial;font-size: 1em\">, null is that email is not spam.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Alternative<span style=\"text-align: initial;font-size: 1em\"> hypothesis is the hypothesis what you are usually trying to prove. In our earlier examples, alternate hypotheses <\/span>is<span style=\"text-align: initial;font-size: 1em\"> that you are lucky, or drug is effective, or group means are not equal, or clinical test result is positive, or email is spam.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">After completing your experiment and after performing appropriate statistical tests, you get a P value. It is important to note that P values and statistical tests only assess the validity of our null hypothesis defined in step 2 above (they do not test the validity of our alternate hypothesis directly). Therefore, we can make conclusions only about the null hypothesis. If you get a P value (for example, 0.04) which is less than our predefined level of tolerance for errors (0.05) defined in step 1, you may conclude to \u2018reject the null hypothesis\u2019 and state that the results are statistically significant. If your obtained P value (for example, 0.06) is higher than significance level alpha, then the conclusion would be \u2018not to reject the null hypothesis\u2019 and that the results are not statistically significant. Remember that statistically significant does not mean scientifically or clinically significant. You cannot conclude that the null hypothesis is true. All you can do is conclude that you \u2018don&#8217;t have sufficient evidence to reject the null hypothesis\u2019. This sort of arguments <\/span>are<span style=\"text-align: initial;font-size: 1em\"> called arguments by contradiction often used in logic and philosophy.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">An intuitive analogy is <\/span>the similar<span style=\"text-align: initial;font-size: 1em\"> situation in <\/span>jurisdiction<span style=\"text-align: initial;font-size: 1em\">. Here H0 is accused is innocent and Ha is accused is criminal. A juror (judge) can pronounce the verdict as guilty (reject the null hypothesis) or not guilty (do not reject the null hypothesis). She can never pronounce that the accused is criminal (accept the alternative hypothesis) or innocent (accept the null hypothesis). Also, the juror\u2019s conclusion is crisp and clear: guilty or not guilty. She cannot say that the accused is \u2018probably guilty\u2019 or \u2018I don\u2019t know\u2019, as no one likes uncertainty in <\/span>judgements<span style=\"text-align: initial;font-size: 1em\">. Similarly, in statistical hypothesis testing conclusion\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">is crisp and clear as it is based on the threshold significance level. A test that returns P value of 0.4999 should be deemed significant while <\/span>P<span style=\"text-align: initial;font-size: 1em\"> value of 0.5001 is not significant. Many statisticians <\/span>criticise<span style=\"text-align: initial;font-size: 1em\"> making this sort of crisp conclusions based on threshold P value and overly emphasizing the term significant in papers. Many investigators do not report the exact P value (like P=0.499, which is in fact very much borderline) and instead make ambiguous statements like \u2018results are significant (P&lt;0.05)\u2019.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">5.<\/strong><span style=\"text-align: initial;font-size: 1em\">\u00a0<\/span><strong style=\"text-align: initial;font-size: 1em\">Type 1 and 2 errors<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">As already stated, a significance level of 0.05 means 95% confidence level and tolerance of 5% error level. Even though this 5% level is very less, 5% of all results will be erroneous. Out of 100 statistical tests, 5 tests would lead to <\/span>erroneous<span style=\"text-align: initial;font-size: 1em\"> conclusion that the result is significant, while in <\/span>reality<span style=\"text-align: initial;font-size: 1em\"> the result would not be significant. For example, when you compare means of two groups, you got a P value &lt;0.05, therefore reject the null hypothesis and conclude that two means are significantly different, but in <\/span>reality<span style=\"text-align: initial;font-size: 1em\"> means are not different (null hypothesis of no difference in means is wrongly rejected). Another example is that you go for an ELISA blood test to see whether the test is positive or negative. The test report says positive (the null hypothesis that you have no HIV infection is rejected), but this result is erroneous. You will never know is it an error or not unless you repeat the test many times or you already know the reality (for example, in simulations). This type of errors where <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is wrongly rejected (incorrectly rejecting the null hypothesis) is called type-1 error.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">We can also have yet another kind of error. The test report there is no difference in two group means, but in <\/span>reality<span style=\"text-align: initial;font-size: 1em\"> there is <\/span>difference<span style=\"text-align: initial;font-size: 1em\">. A person is really HIV positive, but the test says he is HIV negative (\u2018false negative\u2019). A person is really guilty but the juror pronounces the verdict as not guilty. The email is truly spam but the spam filter says email as not spam. This type of error is called Type II error or False Negative, which is incorrectly not rejecting the null hypothesis. You&#8217;ve made a type II error when you wrongly do not reject <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis. There really is a difference (association, correlation) overall, but random sampling caused your data to not show a statistically significant difference. So your conclusion that the two groups are not really different is incorrect.<\/span><\/p>\n<\/div>\n<div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-344\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-202.png\" alt=\"\" width=\"763\" height=\"204\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-202.png 763w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-202-300x80.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-202-65x17.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-202-225x60.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-202-350x94.png 350w\" sizes=\"auto, (max-width: 763px) 100vw, 763px\" \/><\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">It is important to note that we will never know is it <\/span>error<span style=\"text-align: initial;font-size: 1em\"> or no error when we <\/span>analyse<span style=\"text-align: initial;font-size: 1em\"> the <\/span>sample,<span style=\"text-align: initial;font-size: 1em\"> unless the analysis is <\/span>simulation<span style=\"text-align: initial;font-size: 1em\"> in which the true population values and reality is known precisely, which is almost impossible when <\/span>analysing<span style=\"text-align: initial;font-size: 1em\"> the scientific experimental data.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">There is yet another type of errors, Type III or type S. This would happen when <\/span>direction<span style=\"text-align: initial;font-size: 1em\"> of effect goes opposite the direction of prediction (especially prone if you use one tail P values to be explained soon). An example of such an error happened in <\/span>Cardiac<span style=\"text-align: initial;font-size: 1em\"> Arrhythmia Suppression Trial (CAST). Anti-arrhythmia drugs (like atenolol to control atrial and ventricular fibrillations) were previously thought to either prolong the life of the patients or no effect, so the investigators chose to report one tail P value. They got a high one tail P value, so concluded the drug has \u2018no effect\u2019 on prolonging the life. <\/span>Actually<span style=\"text-align: initial;font-size: 1em\"> patients given drugs were four times more likely to die than placebo, so the effect went in <\/span>opposite<span style=\"text-align: initial;font-size: 1em\"> direction!<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">6.<\/strong><span style=\"text-align: initial;font-size: 1em\">\u00a0 <\/span><strong style=\"text-align: initial;font-size: 1em\">Choosing the level of significance<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">In the table above, quadrat denoted as [A] signifies false positives. Sum of quadrats [A] and [B] is the total of <\/span>first<span style=\"text-align: initial;font-size: 1em\"> row, \u2018null hypothesis is true\u2019. A\/(A+B) is \u03b1, level of significance. It is \u2018Fraction of false positives when <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is true\u2019. If the null hypothesis is true, what is the probability of incorrectly rejecting it? That probability is the significance level.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">As already explained, a significance level (\u03b1 ) of 0.05 (95% confidence) is chosen almost universally in experimental scientific disciplines, especially in biological and environmental sciences. This traditional significance level is sometimes called \u201ctwo sigma level\u201d, as 95% of all values are expected within two standard deviations from the sample mean in normal distributions (<\/span>central<span style=\"text-align: initial;font-size: 1em\"> limit theorem of statistics).\u00a0<\/span>Ideally<span style=\"text-align: initial;font-size: 1em\"> this level should be decided based upon the consequence of two main kinds of errors, type I and type II errors explained earlier.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Use low \u03b1 (like 0.01), to have less chance of making Type I error (False Positive), but beware, <\/span>chance<span style=\"text-align: initial;font-size: 1em\"> of making Type II error (False Negative) is increased! The consequence of setting \u03b1=0.01 would be more Type II errors (False Negatives). Examples of type II errors include more spam in <\/span>inbox<span style=\"text-align: initial;font-size: 1em\">, a criminal is set free, the test says negative for <\/span>HIV+<span style=\"text-align: initial;font-size: 1em\"> person, A doped athlete is set free, and an effective drug is declared as ineffective and abort drug development. If these sort of false negatives are tolerable, choosing low \u03b1 is justified. Use low \u03b1 (like 0.01) to minimize false positives and is ideal for situations when false negatives are tolerated like Criminal jurisdiction (\u201cit is better to let many guilty <\/span>person<span style=\"text-align: initial;font-size: 1em\"> go free than to falsely convict one innocent person\u201d), Spam Filter (it is better to have spam emails in inbox rather than automatically detecting as spam and deleting that \u2018article acceptance\u2019 or \u2018job selection\u2019 email!), and clinical trial for \u2018me-too\u2019 drug (it is better to abort an effective drug development than marketing an ineffective, useless drug when better alternatives are already available)<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Use high \u03b1 (like 0.1), to have less chance of making Type II error (False Negative), but beware, <\/span>chance<span style=\"text-align: initial;font-size: 1em\"> of making Type I error (False Positive) is increased! The consequence of setting \u03b1=0.1 would be more Type I errors (False Positives). ). Examples of type I errors include: Good email declared as spam and deleted, An innocent is punished, The test says positive for HIV- person, An innocent athlete is declared as doped and banned for life, and an ineffective drug is declared as effective and market it. Use high \u03b1 (like 0.1) to minimize false negatives and is ideal for situations when false positives are tolerated like Civil jurisdiction (it is better to punish many innocent than let one huge illegal corporation go free), and clinical trial for novel drug (it is better to market an ineffective, useless drug than aborting the development of an effective drug when no drugs are available in the market).<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Please be noted that traditional two sigma level of 0.05 is not universally followed in other disciplines, especially in particle physics. To conclude evidence of detecting a (previously known) particle, accepted level of significance is 3 sigma (\u03b1=0.003). To conclude evidence of the discovery of a new kind of particle (like in the paper describing the discovery of Higg\u2019s Boson), <\/span>accepted<span style=\"text-align: initial;font-size: 1em\"> level of significance is 5 sigma (\u03b1=0.0000003).<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">7.\u00a0\u00a0 P-values: Correct interpretations and fallacies<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Definition of P value starts with the proposition \u2018If the null hypothesis is true\u2019. The definition is \u201cIf the null hypothesis is true, what is the probability that random sampling would lead to <\/span>difference<span style=\"text-align: initial;font-size: 1em\"> as large as or larger than that observed in this study?\u201d<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Let us consider an example. A study compared <\/span>mean<span style=\"text-align: initial;font-size: 1em\"> of two <\/span>groups,<span style=\"text-align: initial;font-size: 1em\"> and got a P value of 0.03.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">That means if two population means are identical (i.e., if the null hypothesis is true), there is a 3% chance of observing a difference as large as or larger than what you observed. As that chance is very low 9only 3 out of 100), you conclude that the means are not identical. Alternatively, you can conclude that random sampling from identical populations would lead to a difference smaller than you observed in 97% of experiments, and larger than you observed in 3% of experiments. P value answers the question: \u201cIn an experiment of this size, if the populations really have the same mean, what is the probability of observing at least as large a difference between sample means as was, in fact, observed?\u201d<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">P values are indeed confusing to many scientists because <\/span>null<span style=\"text-align: initial;font-size: 1em\"> hypothesis is almost always false. Logic works <\/span>backwards<span style=\"text-align: initial;font-size: 1em\"> (Population to Sample, Logical Deduction akin to Probability) and is counterintuitive; the hypothesis testing works on the principle of argument by contradiction.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">There are two types of P values, two-tail <\/span>and<span style=\"text-align: initial;font-size: 1em\"> one tail. A two-tail P value indicates <\/span>critical<span style=\"text-align: initial;font-size: 1em\"> region extends to both extremes of distribution (significantly large or significantly small). A one-tail P value <\/span>extend<span style=\"text-align: initial;font-size: 1em\"> only to either of the two tails of <\/span>distribution<span style=\"text-align: initial;font-size: 1em\">. A difference between two can easily be spotted by looking at your null hypothesis. If the null hypothesis contains the symbol = explicitly or indirectly (like \u2018there are no differences between means\u2019 is equivalent to \u2018mean of group x = mean of group y\u2019), P value should be two-tail. If the null hypothesis <\/span>contain<span style=\"text-align: initial;font-size: 1em\"> the symbol &gt; or &lt; explicitly or indirectly (like \u201cthere is no fever\u201d is equivalent to \u201ctemperature \u226437\u00b0C\u201d or \u201cthere is no profit\u201d is equivalent to \u201cnet gain &gt; net loss\u201d), P value should be one-tail. A one tail P value explicitly make predictions of the directionality of effect as part of the experimental design. This predicted directionality might very well be wrong too, that lead to type III error as explained earlier. When in doubt, always stick with two-tail P values. One tail P values are always less than Two tail P values, so investigators like One tail P to prove their hypothesis! Some reviewers criticize any use of one-tail P values, no matter how well justification is.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">The following is a list of P value fallacies (incorrect interpretations):<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">1.\u00a0\u00a0\u00a0\u00a0\u00a0 The <\/span><em style=\"text-align: initial;font-size: 1em\">P<\/em><span style=\"text-align: initial;font-size: 1em\"> value is the probability of rejecting the null hypothesis<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">2.\u00a0\u00a0\u00a0\u00a0\u00a0 A high <\/span><em style=\"text-align: initial;font-size: 1em\">P<\/em><span style=\"text-align: initial;font-size: 1em\"> value proves that the null hypothesis is true.<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">3.\u00a0\u00a0\u00a0\u00a0\u00a0 1-<\/span><em style=\"text-align: initial;font-size: 1em\">P<\/em><span style=\"text-align: initial;font-size: 1em\"> is the probability that the results will hold up when the experiment is repeated<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">4.\u00a0\u00a0\u00a0\u00a0\u00a0 1-<\/span><em style=\"text-align: initial;font-size: 1em\">P<\/em><span style=\"text-align: initial;font-size: 1em\"> is the probability that the alternative hypothesis is true<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">5.\u00a0\u00a0\u00a0\u00a0\u00a0 The <\/span><em style=\"text-align: initial;font-size: 1em\">P<\/em><span style=\"text-align: initial;font-size: 1em\"> value is the probability that the null hypothesis is true<\/span><\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">6.\u00a0\u00a0\u00a0\u00a0\u00a0 <\/span><em style=\"text-align: initial;font-size: 1em\">P <\/em><span style=\"text-align: initial;font-size: 1em\">value is the probability that the result was due to sampling error<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Be noted that all six of the above propositions are wrong interpretations of P values.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">8.<\/strong><span style=\"text-align: initial;font-size: 1em\">\u00a0\u00a0\u00a0 <\/span><strong style=\"text-align: initial;font-size: 1em\">P-hacking<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Remember the coin toss experiment we described <\/span>in the beginning<span style=\"text-align: initial;font-size: 1em\"> to test how lucky a person is? Two ways the person can cheat are by changing the threshold value and increasing the no of trials. Similarly, experimental scientists unethically resort to many practices to decrease the P value (a low P value is desirable, as a P value lesser than the alpha (P&lt;0.05) would reject the null hypothesis and lead to a conclusion of statistical significance. Some of these tricks include:<\/span><\/p>\n<ol>\n<li style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Try different tests, parametric and nonparametric and picking the one with <\/span>lowest<span style=\"text-align: initial;font-size: 1em\"> P value\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">The statistical test that will be used (for example, unpaired t-test) should be clearly specified as part of the experimental design. It is not fair to try different tests and go with the one that gives <\/span>lowest<span style=\"text-align: initial;font-size: 1em\"> P value.<\/span><\/li>\n<li style=\"text-align: justify\">Dynamic sample size and stop when P is lower than threshold, Sample size should clearly be decided as part of the experimental design. It is not fair to change this till<span style=\"text-align: initial;font-size: 1em\"> P value is less than the threshold.<\/span><\/li>\n<li style=\"text-align: justify\">Slice and dice the dataset to get lowest P, Obviously an unethical practice. We should use all of the elements of our dataset without slicing and dicing.<\/li>\n<li style=\"text-align: justify\">Cherry-picking the dataset to get lowest P Obviously an unethical practice. We should use all of the elements of our dataset without picking certain values that affirms<span style=\"text-align: initial;font-size: 1em\"> our preconceived conclusions.<\/span><\/li>\n<li style=\"text-align: justify\">Play with outliers, remove one or more selectively to get desired<span style=\"text-align: initial;font-size: 1em\"> P<\/span><\/li>\n<\/ol>\n<div>\n<p style=\"text-align: justify\">\u00a0 \u00a0Obviously an unethical practice. We should use all of the elements of our dataset without removing any outliers manually. In case outliers need to be removed, a formal statistic test such as Grubb\u2019s test can be used for this.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">All of the above are cheating, and construes scientific misconduct. As already mentioned, statistical hypothesis testing based on P values is based upon the threshold significance level, which is nothing but an arbitrary value. Statistical significance is usually denoted by the symbol * in graphs. Over emphasizing the statistical significance and * is sometimes referred as \u2018stargazing\u2019. P values depend largely on sample size. It is easy to get very low P values with huge sample size with tiny effect, or tiny sample size with huge effect. P value answers the question, is there a significant evidence of difference? That question is different from \u201cIs there an evidence of significant difference?\u201d. P value does not answer that question. Difference between sample means might be tiny, or scientifically negligible, yet decision based on P value could be \u2018statistically significant\u2019. Add on to the woes, many studies have revealed that P values are not very well reproducible. Instead of P values, the current consensus among the statisticians is to used 95% Confidence Intervals.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>9.\u00a0\u00a0 P-values and Confidence Intervals<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Confidence Intervals and P values are very well connected. A 95% Confidence Interval is similar to 0.05 significance level as already explained. The question is does the range of values defined by 95% Confidence Interval include our null hypothesis? For example, while comparing the difference in mean between two groups. Null hypothesis is that the group means are same, or the difference between two group means is zero. If 95% Confidence Intervals of difference between sample means does not include zero-our null Hypothesis, then P&lt;0.05 (inversely, if it includes zero, then P&gt;0.05) and indicate a statistically significant difference. Let us consider another example. A random sample consisting of 10 individuals volunteered to have their body temperatures measured to know whether they have normal body temperature or not. Mean body temperature was found to be 36.8\u00b0C. The width of 95% Confidence\u00a0<span style=\"text-align: initial;font-size: 1em\">Interval was found to be 0.3\u00b0C and therefore 95% CI about the sample mean ranges from (36.8-0.3) to (36.8+0.3) that is 36.5\u00b0C to 37.1\u00b0C. Our null hypothesis of having no fever is 37\u00b0C. Does our obtained 95% CI include the null hypothesis of 37\u00b0C? Yes, it does. So, P&gt;0.05 and we can conclude that there is no statistically significant difference from the null hypothesis.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">10. FDR and Statistical Power<\/strong><\/p>\n<\/div>\n<div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-345\" src=\"http:\/\/esp14.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-203.png\" alt=\"\" width=\"761\" height=\"202\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-203.png 761w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-203-300x80.png 300w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-203-65x17.png 65w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-203-225x60.png 225w, https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-content\/uploads\/sites\/176\/2019\/03\/Untitled-203-350x93.png 350w\" sizes=\"auto, (max-width: 761px) 100vw, 761px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">We have already defined A\/(A+B) as \u03b1, level of significance. . It is \u2018Fraction of false positives when null hypothesis is true\u2019. A related term is False Discovery Rate (FDR), which is A\/A+C. FDR is defined as \u2018Fraction of false positives out of statistically significant results\u2019. If the result is statistically significant, what is the probability that null hypothesis is really true? Note that for both level of significance and FDR, numerator remains same,\u2019 fraction of false positives\u2019, but denominator is different. While significance level is based out of all results where null hypothesis is true (i.e., all results that in reality have no effect\/difference etc.), FDR is based out of results with a \u2018statistically significant\u2019 conclusion. FDR depends upon \u2018prior probability\u2019-the context of experiment akin to Bayesian Statistics. FDR and Prior Probability will be elaborated in the module of Bayesian Statistics. Another related term is \u2018statistical power\u2019 which is C\/C+D. It is \u2018fraction of statistically significant results out of all results where null hypothesis is false\u2019. The statistical power is the probability to obtain a statistically significant result assuming a certain effect size in population. Power depends on sample size, variability, choice of \u03b1 and hypothetical effect size. A high power is obtained when: 1) Large sample size is used, 2) Looking for a large effect (or large difference or a strong correlation etc), and 3) Data with little scatter (small variance and SD).<\/p>\n<\/div>\n<ol start=\"11\">\n<li><strong>Summary<\/strong><\/li>\n<\/ol>\n<p style=\"text-align: justify\">\u00a0 \u00a0a.\u00a0Statistical<span style=\"font-size: 1em\"> hypothesis is a <\/span>four step<span style=\"font-size: 1em\"> sequential process. Significance level should be defined first, then null and alternative hypotheses must be defined. Finally, based on obtained P value, the fate of two hypotheses is decided. It is cheating to change the decided significance <\/span>level,<span style=\"font-size: 1em\"> or changing the sample size and other numerous ways P values can be hacked.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"font-size: 1em\">b. Significance level should be chosen based upon tolerance limit to two types of errors; type I (false positives) and type II (false negatives)<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"font-size: 1em\">c. P value is defined as \u201cIf the null hypothesis is true, what is the probability that random sampling would lead to <\/span>difference<span style=\"font-size: 1em\"> as large as or larger than that observed in this study?\u201d It is more essential to know how P values are correctly interpreted rather than how P values are calculated.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"font-size: 1em\">d. A low P value does not mean differences are scientifically significant, or finding interesting or warrant further funding<\/span>. .<span style=\"font-size: 1em\"> P value answers the question, is there <\/span>a significant<span style=\"font-size: 1em\"> evidence of difference? Not \u201cIs there <\/span>an evidence<span style=\"font-size: 1em\"> of significant difference?\u201d.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"font-size: 1em\">e. If 95% CI does not include <\/span>null<span style=\"font-size: 1em\"> hypothesis, then P value must be &lt;0.05. If the range does include <\/span>null<span style=\"font-size: 1em\"> hypothesis, then P&gt;0.05. This connection is important.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p><strong>Quadrant-III: Learn More\/ Web Resources \/ Supporting Materials:<\/strong><\/p>\n<p>1. Brief review of concepts of Statistical Hypothesis testing at PennState:<\/p>\n<p>https:\/\/onlinecourses.science.psu.edu\/statprogram\/node\/138<\/p>\n<p>2. A good overview of hypothesis testing in GraphPad guide:<\/p>\n<p>https:\/\/www.graphpad.com\/guides\/prism\/7\/statistics\/index.htm?statistical_hypothesis_testing.htm<\/p>\n<p>3. P values- an introduction https:\/\/www.statsdirect.com\/help\/basics\/p_values.htm<\/p>\n","protected":false},"author":3,"menu_order":17,"template":"","meta":{"_acf_changed":false,"pb_show_title":"on","pb_short_title":"","pb_subtitle":"","pb_authors":["dr-felix-bast"],"pb_section_license":""},"chapter-type":[],"contributor":[59],"license":[],"class_list":["post-340","chapter","type-chapter","status-publish","hentry","contributor-dr-felix-bast"],"part":3,"_links":{"self":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapters\/340","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapters"}],"about":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/wp\/v2\/types\/chapter"}],"author":[{"embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/wp\/v2\/users\/3"}],"version-history":[{"count":4,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapters\/340\/revisions"}],"predecessor-version":[{"id":346,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapters\/340\/revisions\/346"}],"part":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/parts\/3"}],"metadata":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapters\/340\/metadata\/"}],"wp:attachment":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/wp\/v2\/media?parent=340"}],"wp:term":[{"taxonomy":"chapter-type","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/pressbooks\/v2\/chapter-type?post=340"},{"taxonomy":"contributor","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/wp\/v2\/contributor?post=340"},{"taxonomy":"license","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/esp14\/wp-json\/wp\/v2\/license?post=340"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}