{"id":104,"date":"2018-12-24T06:23:39","date_gmt":"2018-12-24T06:23:39","guid":{"rendered":"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/?post_type=chapter&#038;p=104"},"modified":"2018-12-24T06:39:17","modified_gmt":"2018-12-24T06:39:17","slug":"non-parametrics-hypothesis-testing-and-confidence-interval","status":"publish","type":"chapter","link":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/chapter\/non-parametrics-hypothesis-testing-and-confidence-interval\/","title":{"rendered":"Non-Parametrics Hypothesis Testing and Confidence Interval"},"content":{"raw":"<div>\r\n\r\n&nbsp;\r\n\r\n<strong>1 Nonparametric hypothesis testing and con dence in-terval: Why the need?<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The central idea of parametric (or classical) inference is the assumption regarding the un-derlying population. The entire theory of the parametric inference is developed under this assumption and consequently these procedures are valid as long as these assumptions are satis ed. For example, students t test is appropriate only when the underlying distribution is normal. But normal is not the only distribution having applications in real life. For example, in survival trials, the lifetime distributions are mostly exponential, gamma or Weibull, that is non normal. Therefore, inference based on t test in such a situation is often misleading. Therefore, parametric tests are useful only when the experimenter is su ciently con dent about the underlying distribution. But unfortunately, the available methods of identifying the underlying distribution is limited to standard distributions only.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Therefore, it will be better if hypothesis testing procedures can be developed with minimal assumptions about the underlying distribution(say, continuity of observations). Suppose there exists a statistic T , relevant to the testing problem such that the exact and\/or condi-tional and\/or asymptotic distribution of T under the null hypothesis is independent of the underlying distribution. Naturally, the signi cance level for the test based on T does not depend on the underlying distribution, that is, these tests are level robust. Tests based on T\u00a0 are commonly termed as nonparametric or distribution free.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">In real practice, T is often formed by taking only the signs or ranks of the actual obser-vations. But such a T does not take into account the full information contained in the individual observations and hence are often less e cient than their parametric counterparts. However, nonparametric tests are the valid choices as long as the validity of the parametric assumptions are questionable.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">As in parametric inference problems, the experimenter may be interested in constructing a con dence interval with coverage probability independent of the underlying distribution. However, in nonparametric inference, the unknown quantity of interest are mostly location\u00a0<span style=\"text-align: initial;font-size: 1em\">or scale parameters. As in the parametric counterpart, we can invert the acceptance region of a distribution free test to get a distribution free con dence interval with high coverage probability. These will be discussed later. However, if the unknown quantity of interest is a population quantile, we can develop a distribution free con dence interval based on the sample order statistics.\u00a0<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong><span style=\"text-align: initial;font-size: 1em\">2 Components of a nonparametric test<\/span><\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">In most of the problems of nonparametric inference, the underlying distribution is not spec-i ed except for continuity. Therefore the available methods(e.g. Neyman Pearson lemma, likelihood ratio method) of test construction are of no use and hence, we need to develop tests with intuitive appeal. In particular, if we can identify a distribution free statistic T for the problem, we can develop a distribution free test. However, for a meaningful development, we maintain the following sequence:<\/span><\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">1.\u00a0\u00a0 Model assumption(i.e. the minimal set of assumptions about the underlying distribu-tion)<\/p>\r\n<p style=\"text-align: justify\">2.\u00a0\u00a0 Hypothesis of interest(i.e. speci cation of the null and alternative hypotheses)<\/p>\r\n<p style=\"text-align: justify\">3.\u00a0\u00a0 Available tests for the problem(i.e. the usual parametric procedures for speci c prob-ability models)<\/p>\r\n<p style=\"text-align: justify\">4.\u00a0\u00a0 Suggesting a distribution free statistic(i.e. specifying some T )<\/p>\r\n<p style=\"text-align: justify\">5.\u00a0\u00a0 Justifying the form of critical region.<\/p>\r\n<p style=\"text-align: justify\">6.\u00a0\u00a0 Investigating unbiasedness and consistency of the suggested test and\u00a0 nally<\/p>\r\n<p style=\"text-align: justify\">7.\u00a0\u00a0 Providing a large sample test corresponding to the given test.<\/p>\r\n\r\n<\/div>\r\n&nbsp;\r\n\r\n<strong>3 Consistency of tests-Some basics<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">However, tests based on such a T are often far from being optimal and hence properties like maxmising power is not immediate. Consistency is, therefore, an important concern to measure the sensitivity of the test in large samples to a little departure from the null hypothesis. We provide below the notion of consistency of tests in little details.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Suppose X1; X2; ::; XN are iid observations from an unknown distribution G. We are inter-ested in testing H0 : G 2 0 against Ha : G 2 a, where 0( a) is the class of distributions speci ed by the null(alternative) hypothesis.<\/p>\r\n<p style=\"text-align: justify\">A sequence of tests f N g is said to be consistent if for every G 2\u00a0 a,<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-107 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-27.png\" alt=\"\" width=\"616\" height=\"474\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">However, asymptotic normality or asymptotic size condition is not immediate and hence we provide some simpler conditions. Assume that G is indexed by a real parameter , that is, G(x) = G(x; ) and the testing problem can be expressed as H0 : = 0 against Ha &gt; 0. Suppose a level test rejects the null hypothesis if SN ( 0) cN ( 0), where SN ( 0) is so constructed that<\/p>\r\n<img class=\"size-full wp-image-108 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-28.png\" alt=\"\" width=\"538\" height=\"538\" \/>\r\n\r\n<img class=\"size-full wp-image-109 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-29.png\" alt=\"\" width=\"623\" height=\"322\" \/>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">for varying n. Suppose we x 0 = 0 and = :2. It is easy to observe that the power for Test 1 remains very close to .05 for varying n whereas the power for Test 2 increases sharply. The same is observed for the assumed choices of . However, the rate increase for the power of Test 2 is higher for higher values of . This is expected, as the rst test is based on an\u00a0inconsistent estimator(i.e. X), i.e. , it does not concentrate around the true value for large ~ n. Consequently the power increases at a very slow rate for Test 1. On the other hand, X is consistent and hence for large n approaches the true parameter and consequently power increases for increasing values of n.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong>4\u00a0 Different nonparametric hypothesis testing problems<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Now we shall discuss the di erent types of hypotheses considered in nonparametric hypothesis testing together with their relevances. Based on the availability of data, hypothesis testing problems are either single sample or two sample or multi-sample. We, therefore, discuss hypotheses for each type of problems.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong><span style=\"font-size: 1em;text-align: initial\">4.1\u00a0 Single sample problems<\/span><\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Suppose X1; X2; ::; Xn are iid observations from an unknown distribution F . F is unknown but known to be continuous. Then depending on the requirement, we have the following di erent hypotheses.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong><span style=\"text-align: initial;font-size: 1em\">4.1.1 Problem of location<\/span><\/strong><\/p>\r\n\r\n<\/div>\r\n<div style=\"text-align: justify\">\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Suppose p 2 (0; 1) is a known quantity and let (F ) = p(F ) be the quantile of order p for F , that is F ( (F )) = p. Then the problem of location is to test H0 : (F ) = 0 against one of the alternatives Ha : (F ) &gt; 0 or Ha : (F ) &lt; 0 or Ha : (F ) 6= 0 for some known 0.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">For the above hypothesis, only the continuity of F is required. However, if we consider the same hypothesis with the added assumption of symmetry of F , it will be the problem location under symmetry. Since, the main concern in a problem of location or location under symmetry, is the location(e.g. median), the hypothesis is an analogue to the test of a location parameter in parametric counterpart.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong>4.1.2\u00a0 Goodness of t problem<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">In a goodness of t problem, the interest lies in investigating whether the sample comes from a speci ed distribution. Then the hypothesis of interest in a goodness of t problem can be described as H0 : F (x) = F0(x) for all x, where F0 is a completely known DF. The alternative hypothesis is naturally Ha : F (x) 6= F0(x) for at least one x. The problem of goodness of a t also arises in parametric inference after some known distribution is tted to a data.<\/p>\r\n&nbsp;\r\n\r\n<strong>4.2\u00a0 Two sample problems<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Suppose Xi; i = 1; 2; ::; n and Yj; j = 1; 2; ::; m are independent samples from unknown distributions F and G respectively. We only assume the continuity of observations from F and G. The basic hypothesis in any two sample problem is H0 : F (x) = G(x) for all x\u00a0<span style=\"text-align: initial;font-size: 1em\">against the usual one sided or two sided hypothesis. De ning 0 = f(F; G) : F (x) =\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">G(x)8xg, we can express the null hypothesis as H0\u00a0 : (F; G) 2 0. Depending on different\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">speci cations of 0 and a = f(F; G) : F (x) 6= G(x) for some x g, we have the following possible hypotheses.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">1. General\/ Homogeneity Alternative: Suppose the two underlying populations dif-fer in any manner(in location, scale or skewness). Then such an alternative hypothesis can be expressed as a = f(F; G) : F (x) 6= G(x) for some xg. The general hypothesis is also termed as Homogeneity alternative.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">2. Stochastic Alternative: A stochastic alternative is a restricted alternative, where a = f(F; G) : G(x) F (x) for all x with strict inequality for some xg:A<\/span><span style=\"text-align: initial;font-size: 1em\">ctually G(x) F (x) for all x implies that X observations are tend to be larger than\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">Y observations or, in other words, X is stochastically larger than Y. This is a general class of alternatives.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">3. Location Alternative: Suppose G(x) = F (x ), where 6= 0. That is, the un-derlying distributions di er only in location. Then F (x) &gt; G(x) or F (x) = G(x) or\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">F (x) &lt; G(x) according as &lt; 0 or = 0 or &gt; 0. Thus the null hypothesis can be re-stated as H0 : = 0 and the alternative is either Ha : &gt; 0 or Ha : &lt; 0 or Ha : 6= 0.\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">Clearly &gt; 0 implies a = f(F; G) : F is shifted to the right of G for some xg. This is a special case of a stochastic alternative, where G(x) = F (x ). Stochastic alter-native,in general, relates to the location alternative in a less restrictive sense because G(x) &gt; F (x) indicates larger Y observations and hence corresponds to larger location of Y observations.<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">4. Scale Alternative:\u00a0 Suppose G(x) = F ( x ) with\u00a0 \u00a0&gt; 0, that is, the two underlying\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">populations are assumed to di er only in scale. Since, F (x) R G(x) , R 1, the null hypothesis reduces to H0 : = 1. The alternative is either of Ha : &gt; 1 or Ha : &lt; 1<\/span><\/p>\r\n<img class=\"size-full wp-image-110 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-30.png\" alt=\"\" width=\"604\" height=\"346\" \/>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">or Ha : 6= 1. Scale alternative can also be viewed as a stochastic alternative, where G(x) = F ( x ).<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">It is worthwhile to mention that tests meant for general or stochastic alternative can also be used for location and scale alternatives. But they will be less e cient than the tests developed for the speci c alternative. A similar set of alternatives can be found corresponding to any multiple sample problems and will be discussed later considering speci c situations and hence are not discussed separately.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong>4.3 Paired sample problems<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Suppose (Xi; Yi)i = 1; 2; ::; n are sample observations from an unknown bivariate distribution F (x; y) with FX (x) and FY (y) as the marginal DF's. Then the following two hypotheses are of main interest:<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">1.\u00a0\u00a0 Problem of association: Here the interest lies in testing H0 : F (x; y) = FX (x)FY (y) for all (x,y)<\/p>\r\n\r\n<\/div>\r\n<p style=\"text-align: justify\"><img class=\"size-full wp-image-111 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-31.png\" alt=\"\" width=\"636\" height=\"432\" \/><\/p>\r\n\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">2.\u00a0\u00a0 Problem of location: Suppose X represents the response before a drug is admin-istered and Y denote that after the application of the drug. Then naturally Y are in uenced by X and hence X and Y are correlated. Then the natural objective in this situation is to determine whether the drug has any e ect. In statistical terms, this means X and Y are exchangeable. Thus the null hypothesis can be expressed as<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">H0 :\u00a0 X and Y are exchangeable.\u00a0D\u00a0H0 : (X; Y ) = (Y; X)\u00a0H0 : F (x; y) = F (y; x) 8 (x; y):<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">De ne D = Y X, Then under the null hypothesis distribution of D is symmetric about the origin. If X and Y di ers only in location under the alternative, then\u00a0<span style=\"font-size: 1em\">D = (D ), where is the location di erence. Naturally, under the null hypothesis, the distribution of D has median at the origin but under the alternative, the median becomes . Then the problem reduces to testing H0 : = 0 against all alternatives. However, the median of the distribution of the di erence is not always the di erence of the the marginal medians. If the marginal distributions and the distribution of the di erence are all symmetric, then median of the distribution of di erence and the di erence of the two medians coincide(see, Gibbons and Chakraborti, 2006, for details).<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">5 Distribution free con dence interval for quantiles<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"font-size: 1em\">Suppose Xi; i = 1; 2; ::; n are iid observations from a continuous but unknown DF F (x). The objective is to provide a con dence interval of the p th order quantile p satisfying F ( p) = p. Since p is a population quantile , it is natural to consider intervals based on the sample quantiles as a con dence interval. Thus we can start with X(r) (i.e. the sample nr th quantile) and X(s) (i.e. the sample ns th quantile). That is, we suggest to consider the interval (X(r); X(s)) with r &lt; s as a con dence interval for p. Now we shall show that the coverage probability of [X(r); X(s)] is independent of any F. Note that for any k, X(k) p , Z k; where Z has a binomial distribution with parameters n and success probability F ( p) = p. Thus<\/span><\/p>\r\n\r\n<\/div>\r\n<img class=\"size-full wp-image-112 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-32.png\" alt=\"\" width=\"469\" height=\"160\" \/>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Thus [X(r); X(s)] is a con dence interval for\u00a0 p with con dence coeient (n; r; s). Clearly, (n; r; s) is independent of any F and hence [X(r); X(s)] gives a distribution free con dence interval of p. However, in practice the con dence coe cient is set at least (1 ) with<\/p>\r\n\r\n<\/div>\r\n<img class=\"size-full wp-image-113 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-33.png\" alt=\"\" width=\"663\" height=\"269\" \/>\r\n\r\n&nbsp;\r\n\r\n&nbsp;","rendered":"<div>\n<p>&nbsp;<\/p>\n<p><strong>1 Nonparametric hypothesis testing and con dence in-terval: Why the need?<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The central idea of parametric (or classical) inference is the assumption regarding the un-derlying population. The entire theory of the parametric inference is developed under this assumption and consequently these procedures are valid as long as these assumptions are satis ed. For example, students t test is appropriate only when the underlying distribution is normal. But normal is not the only distribution having applications in real life. For example, in survival trials, the lifetime distributions are mostly exponential, gamma or Weibull, that is non normal. Therefore, inference based on t test in such a situation is often misleading. Therefore, parametric tests are useful only when the experimenter is su ciently con dent about the underlying distribution. But unfortunately, the available methods of identifying the underlying distribution is limited to standard distributions only.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Therefore, it will be better if hypothesis testing procedures can be developed with minimal assumptions about the underlying distribution(say, continuity of observations). Suppose there exists a statistic T , relevant to the testing problem such that the exact and\/or condi-tional and\/or asymptotic distribution of T under the null hypothesis is independent of the underlying distribution. Naturally, the signi cance level for the test based on T does not depend on the underlying distribution, that is, these tests are level robust. Tests based on T\u00a0 are commonly termed as nonparametric or distribution free.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In real practice, T is often formed by taking only the signs or ranks of the actual obser-vations. But such a T does not take into account the full information contained in the individual observations and hence are often less e cient than their parametric counterparts. However, nonparametric tests are the valid choices as long as the validity of the parametric assumptions are questionable.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">As in parametric inference problems, the experimenter may be interested in constructing a con dence interval with coverage probability independent of the underlying distribution. However, in nonparametric inference, the unknown quantity of interest are mostly location\u00a0<span style=\"text-align: initial;font-size: 1em\">or scale parameters. As in the parametric counterpart, we can invert the acceptance region of a distribution free test to get a distribution free con dence interval with high coverage probability. These will be discussed later. However, if the unknown quantity of interest is a population quantile, we can develop a distribution free con dence interval based on the sample order statistics.\u00a0<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong><span style=\"text-align: initial;font-size: 1em\">2 Components of a nonparametric test<\/span><\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">In most of the problems of nonparametric inference, the underlying distribution is not spec-i ed except for continuity. Therefore the available methods(e.g. Neyman Pearson lemma, likelihood ratio method) of test construction are of no use and hence, we need to develop tests with intuitive appeal. In particular, if we can identify a distribution free statistic T for the problem, we can develop a distribution free test. However, for a meaningful development, we maintain the following sequence:<\/span><\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">1.\u00a0\u00a0 Model assumption(i.e. the minimal set of assumptions about the underlying distribu-tion)<\/p>\n<p style=\"text-align: justify\">2.\u00a0\u00a0 Hypothesis of interest(i.e. speci cation of the null and alternative hypotheses)<\/p>\n<p style=\"text-align: justify\">3.\u00a0\u00a0 Available tests for the problem(i.e. the usual parametric procedures for speci c prob-ability models)<\/p>\n<p style=\"text-align: justify\">4.\u00a0\u00a0 Suggesting a distribution free statistic(i.e. specifying some T )<\/p>\n<p style=\"text-align: justify\">5.\u00a0\u00a0 Justifying the form of critical region.<\/p>\n<p style=\"text-align: justify\">6.\u00a0\u00a0 Investigating unbiasedness and consistency of the suggested test and\u00a0 nally<\/p>\n<p style=\"text-align: justify\">7.\u00a0\u00a0 Providing a large sample test corresponding to the given test.<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p><strong>3 Consistency of tests-Some basics<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">However, tests based on such a T are often far from being optimal and hence properties like maxmising power is not immediate. Consistency is, therefore, an important concern to measure the sensitivity of the test in large samples to a little departure from the null hypothesis. We provide below the notion of consistency of tests in little details.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Suppose X1; X2; ::; XN are iid observations from an unknown distribution G. We are inter-ested in testing H0 : G 2 0 against Ha : G 2 a, where 0( a) is the class of distributions speci ed by the null(alternative) hypothesis.<\/p>\n<p style=\"text-align: justify\">A sequence of tests f N g is said to be consistent if for every G 2\u00a0 a,<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-107 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-27.png\" alt=\"\" width=\"616\" height=\"474\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-27.png 616w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-27-300x231.png 300w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-27-65x50.png 65w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-27-225x173.png 225w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-27-350x269.png 350w\" sizes=\"auto, (max-width: 616px) 100vw, 616px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">However, asymptotic normality or asymptotic size condition is not immediate and hence we provide some simpler conditions. Assume that G is indexed by a real parameter , that is, G(x) = G(x; ) and the testing problem can be expressed as H0 : = 0 against Ha &gt; 0. Suppose a level test rejects the null hypothesis if SN ( 0) cN ( 0), where SN ( 0) is so constructed that<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-108 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-28.png\" alt=\"\" width=\"538\" height=\"538\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-28.png 538w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-28-150x150.png 150w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-28-300x300.png 300w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-28-65x65.png 65w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-28-225x225.png 225w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-28-350x350.png 350w\" sizes=\"auto, (max-width: 538px) 100vw, 538px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-109 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-29.png\" alt=\"\" width=\"623\" height=\"322\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-29.png 623w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-29-300x155.png 300w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-29-65x34.png 65w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-29-225x116.png 225w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-29-350x181.png 350w\" sizes=\"auto, (max-width: 623px) 100vw, 623px\" \/><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">for varying n. Suppose we x 0 = 0 and = :2. It is easy to observe that the power for Test 1 remains very close to .05 for varying n whereas the power for Test 2 increases sharply. The same is observed for the assumed choices of . However, the rate increase for the power of Test 2 is higher for higher values of . This is expected, as the rst test is based on an\u00a0inconsistent estimator(i.e. X), i.e. , it does not concentrate around the true value for large ~ n. Consequently the power increases at a very slow rate for Test 1. On the other hand, X is consistent and hence for large n approaches the true parameter and consequently power increases for increasing values of n.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong>4\u00a0 Different nonparametric hypothesis testing problems<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Now we shall discuss the di erent types of hypotheses considered in nonparametric hypothesis testing together with their relevances. Based on the availability of data, hypothesis testing problems are either single sample or two sample or multi-sample. We, therefore, discuss hypotheses for each type of problems.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong><span style=\"font-size: 1em;text-align: initial\">4.1\u00a0 Single sample problems<\/span><\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Suppose X1; X2; ::; Xn are iid observations from an unknown distribution F . F is unknown but known to be continuous. Then depending on the requirement, we have the following di erent hypotheses.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong><span style=\"text-align: initial;font-size: 1em\">4.1.1 Problem of location<\/span><\/strong><\/p>\n<\/div>\n<div style=\"text-align: justify\">\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Suppose p 2 (0; 1) is a known quantity and let (F ) = p(F ) be the quantile of order p for F , that is F ( (F )) = p. Then the problem of location is to test H0 : (F ) = 0 against one of the alternatives Ha : (F ) &gt; 0 or Ha : (F ) &lt; 0 or Ha : (F ) 6= 0 for some known 0.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">For the above hypothesis, only the continuity of F is required. However, if we consider the same hypothesis with the added assumption of symmetry of F , it will be the problem location under symmetry. Since, the main concern in a problem of location or location under symmetry, is the location(e.g. median), the hypothesis is an analogue to the test of a location parameter in parametric counterpart.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong>4.1.2\u00a0 Goodness of t problem<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In a goodness of t problem, the interest lies in investigating whether the sample comes from a speci ed distribution. Then the hypothesis of interest in a goodness of t problem can be described as H0 : F (x) = F0(x) for all x, where F0 is a completely known DF. The alternative hypothesis is naturally Ha : F (x) 6= F0(x) for at least one x. The problem of goodness of a t also arises in parametric inference after some known distribution is tted to a data.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>4.2\u00a0 Two sample problems<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Suppose Xi; i = 1; 2; ::; n and Yj; j = 1; 2; ::; m are independent samples from unknown distributions F and G respectively. We only assume the continuity of observations from F and G. The basic hypothesis in any two sample problem is H0 : F (x) = G(x) for all x\u00a0<span style=\"text-align: initial;font-size: 1em\">against the usual one sided or two sided hypothesis. De ning 0 = f(F; G) : F (x) =\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">G(x)8xg, we can express the null hypothesis as H0\u00a0 : (F; G) 2 0. Depending on different\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">speci cations of 0 and a = f(F; G) : F (x) 6= G(x) for some x g, we have the following possible hypotheses.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">1. General\/ Homogeneity Alternative: Suppose the two underlying populations dif-fer in any manner(in location, scale or skewness). Then such an alternative hypothesis can be expressed as a = f(F; G) : F (x) 6= G(x) for some xg. The general hypothesis is also termed as Homogeneity alternative.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">2. Stochastic Alternative: A stochastic alternative is a restricted alternative, where a = f(F; G) : G(x) F (x) for all x with strict inequality for some xg:A<\/span><span style=\"text-align: initial;font-size: 1em\">ctually G(x) F (x) for all x implies that X observations are tend to be larger than\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">Y observations or, in other words, X is stochastically larger than Y. This is a general class of alternatives.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">3. Location Alternative: Suppose G(x) = F (x ), where 6= 0. That is, the un-derlying distributions di er only in location. Then F (x) &gt; G(x) or F (x) = G(x) or\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">F (x) &lt; G(x) according as &lt; 0 or = 0 or &gt; 0. Thus the null hypothesis can be re-stated as H0 : = 0 and the alternative is either Ha : &gt; 0 or Ha : &lt; 0 or Ha : 6= 0.\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">Clearly &gt; 0 implies a = f(F; G) : F is shifted to the right of G for some xg. This is a special case of a stochastic alternative, where G(x) = F (x ). Stochastic alter-native,in general, relates to the location alternative in a less restrictive sense because G(x) &gt; F (x) indicates larger Y observations and hence corresponds to larger location of Y observations.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">4. Scale Alternative:\u00a0 Suppose G(x) = F ( x ) with\u00a0 \u00a0&gt; 0, that is, the two underlying\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">populations are assumed to di er only in scale. Since, F (x) R G(x) , R 1, the null hypothesis reduces to H0 : = 1. The alternative is either of Ha : &gt; 1 or Ha : &lt; 1<\/span><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-110 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-30.png\" alt=\"\" width=\"604\" height=\"346\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-30.png 604w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-30-300x172.png 300w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-30-65x37.png 65w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-30-225x129.png 225w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-30-350x200.png 350w\" sizes=\"auto, (max-width: 604px) 100vw, 604px\" \/><\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">or Ha : 6= 1. Scale alternative can also be viewed as a stochastic alternative, where G(x) = F ( x ).<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">It is worthwhile to mention that tests meant for general or stochastic alternative can also be used for location and scale alternatives. But they will be less e cient than the tests developed for the speci c alternative. A similar set of alternatives can be found corresponding to any multiple sample problems and will be discussed later considering speci c situations and hence are not discussed separately.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong>4.3 Paired sample problems<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Suppose (Xi; Yi)i = 1; 2; ::; n are sample observations from an unknown bivariate distribution F (x; y) with FX (x) and FY (y) as the marginal DF&#8217;s. Then the following two hypotheses are of main interest:<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">1.\u00a0\u00a0 Problem of association: Here the interest lies in testing H0 : F (x; y) = FX (x)FY (y) for all (x,y)<\/p>\n<\/div>\n<p style=\"text-align: justify\"><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-111 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-31.png\" alt=\"\" width=\"636\" height=\"432\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-31.png 636w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-31-300x204.png 300w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-31-65x44.png 65w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-31-225x153.png 225w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-31-350x238.png 350w\" sizes=\"auto, (max-width: 636px) 100vw, 636px\" \/><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">2.\u00a0\u00a0 Problem of location: Suppose X represents the response before a drug is admin-istered and Y denote that after the application of the drug. Then naturally Y are in uenced by X and hence X and Y are correlated. Then the natural objective in this situation is to determine whether the drug has any e ect. In statistical terms, this means X and Y are exchangeable. Thus the null hypothesis can be expressed as<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">H0 :\u00a0 X and Y are exchangeable.\u00a0D\u00a0H0 : (X; Y ) = (Y; X)\u00a0H0 : F (x; y) = F (y; x) 8 (x; y):<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">De ne D = Y X, Then under the null hypothesis distribution of D is symmetric about the origin. If X and Y di ers only in location under the alternative, then\u00a0<span style=\"font-size: 1em\">D = (D ), where is the location di erence. Naturally, under the null hypothesis, the distribution of D has median at the origin but under the alternative, the median becomes . Then the problem reduces to testing H0 : = 0 against all alternatives. However, the median of the distribution of the di erence is not always the di erence of the the marginal medians. If the marginal distributions and the distribution of the di erence are all symmetric, then median of the distribution of di erence and the di erence of the two medians coincide(see, Gibbons and Chakraborti, 2006, for details).<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">5 Distribution free con dence interval for quantiles<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"font-size: 1em\">Suppose Xi; i = 1; 2; ::; n are iid observations from a continuous but unknown DF F (x). The objective is to provide a con dence interval of the p th order quantile p satisfying F ( p) = p. Since p is a population quantile , it is natural to consider intervals based on the sample quantiles as a con dence interval. Thus we can start with X(r) (i.e. the sample nr th quantile) and X(s) (i.e. the sample ns th quantile). That is, we suggest to consider the interval (X(r); X(s)) with r &lt; s as a con dence interval for p. Now we shall show that the coverage probability of [X(r); X(s)] is independent of any F. Note that for any k, X(k) p , Z k; where Z has a binomial distribution with parameters n and success probability F ( p) = p. Thus<\/span><\/p>\n<\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-112 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-32.png\" alt=\"\" width=\"469\" height=\"160\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-32.png 469w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-32-300x102.png 300w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-32-65x22.png 65w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-32-225x77.png 225w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-32-350x119.png 350w\" sizes=\"auto, (max-width: 469px) 100vw, 469px\" \/><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Thus [X(r); X(s)] is a con dence interval for\u00a0 p with con dence coeient (n; r; s). Clearly, (n; r; s) is independent of any F and hence [X(r); X(s)] gives a distribution free con dence interval of p. However, in practice the con dence coe cient is set at least (1 ) with<\/p>\n<\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-113 aligncenter\" src=\"http:\/\/statp05.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-33.png\" alt=\"\" width=\"663\" height=\"269\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-33.png 663w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-33-300x122.png 300w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-33-65x26.png 65w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-33-225x91.png 225w, https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-content\/uploads\/sites\/130\/2018\/12\/Untitled-33-350x142.png 350w\" sizes=\"auto, (max-width: 663px) 100vw, 663px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n","protected":false},"author":3,"menu_order":7,"template":"","meta":{"_acf_changed":false,"pb_show_title":"on","pb_short_title":"","pb_subtitle":"","pb_authors":["prof-rahul-bhattacharya"],"pb_section_license":""},"chapter-type":[],"contributor":[58],"license":[],"class_list":["post-104","chapter","type-chapter","status-publish","hentry","contributor-prof-rahul-bhattacharya"],"part":3,"_links":{"self":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/pressbooks\/v2\/chapters\/104","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/pressbooks\/v2\/chapters"}],"about":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/wp\/v2\/types\/chapter"}],"author":[{"embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/wp\/v2\/users\/3"}],"version-history":[{"count":5,"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/pressbooks\/v2\/chapters\/104\/revisions"}],"predecessor-version":[{"id":116,"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/pressbooks\/v2\/chapters\/104\/revisions\/116"}],"part":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/pressbooks\/v2\/parts\/3"}],"metadata":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/pressbooks\/v2\/chapters\/104\/metadata\/"}],"wp:attachment":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/wp\/v2\/media?parent=104"}],"wp:term":[{"taxonomy":"chapter-type","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/pressbooks\/v2\/chapter-type?post=104"},{"taxonomy":"contributor","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/wp\/v2\/contributor?post=104"},{"taxonomy":"license","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/statp05\/wp-json\/wp\/v2\/license?post=104"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}