{"id":310,"date":"2019-04-24T11:31:03","date_gmt":"2019-04-24T11:31:03","guid":{"rendered":"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/?post_type=chapter&#038;p=310"},"modified":"2019-04-24T11:37:55","modified_gmt":"2019-04-24T11:37:55","slug":"multiple-linear-regression-analysis-simple-and-step-wise-regression","status":"publish","type":"chapter","link":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/chapter\/multiple-linear-regression-analysis-simple-and-step-wise-regression\/","title":{"rendered":"Multiple linear regression analysis ( simple and step- wise regression)"},"content":{"raw":"<div><span style=\"float: right\"><a href=\"https:\/\/youtu.be\/7RLoldYt7wg\" target=\"_blank\" rel=\"noopener\"><img src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"epgp books\" width=\"75px\" height=\"75px;\" \/><\/a>\r\n<\/span><\/div>\r\n<strong>\u00a0 \u00a0<\/strong>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>1) E text<\/strong>\r\n\r\n&nbsp;\r\n\r\n<strong>Multiple Linear Regression Analysis<\/strong>\r\n\r\n&nbsp;\r\n\r\n<strong>Regression Analysis<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Ultimate objective of any research is to understand the causes behind a process with objective to control it as per our requirement. Knowledge of the degree and direction of co relationship between a dependent and independent variable alone is not sufficient to help in this regard unless it gives the mathematical equation giving the relationship between them. Through this form of causal relationship (Cause and effect relationship ) we can evaluate the effect of different independent variablesalso known as determinants on the dependent variable and can also make projections of the dependent variable for some projected future values of independent variables etc.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Regression analysis helps us in identifying the above mentioned functional form of relationship. If we have only two variables to analyze, i.e. one dependent and one independent variable, the regression is known as bivariate regression. However, if the number of independent variables are two or more it is known as Multiple Regression analysis. Explanation of regression analysis becomes very complicated if we directly start from multiple regression analysis. It is easier to explain bivariate regression analysis first and explain Multiple Regression as its extension.<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>Bivariate Regression Analysis<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Functional form of relationship between to variables is best explained by taking a straight line on a graph paper having Y-coordinate and X-coordinate. If we take large number of points on the line and note their Y and X coordinates. The Y and X coordinates of these points will be related by a general functional form as : <strong>Y = a + b X +<\/strong> <strong>\u03b5<\/strong>. In any such equation of straight line <strong>a<\/strong> is known as intercept of the line i.e. the value at Y axis where the line will cut Y- axis (value of Y after putting X = 0.), <strong>b<\/strong>therate of change in Y with respect to X,is known as the slope of the line and<strong>\u03b5<\/strong> is known as the error and the relationship Y = a + b X. is known as linear relationship (as it relates to a line only). Thus if we refer to any specific straight line, we can specify it by fixing the values of <strong>a<\/strong> and <strong>b<\/strong>. However, most of the bivariate data from social sciences, though following a linear pattern of relationship, will not exactly fall on a straight line. On a scatter plot of such data the points may cluster around a straight line and give some deviations on either side of the line as indicated by<strong>\u03b5<\/strong>. In such cases the line around which the points cluster could be identified to give the relationship<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">between two variables X and Y. Here the values of Y and X will fall on the line and relationship between them will be exact. Actual relationship between values of Y and X falling around the line will only be approximate to it. The value of the coefficient of correlation will give the direction and the degree of such linear relationships. The equation of the line passing through the scatter plot will give the functional form of linear relationship between the two variables.<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>Principle of Least Square<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">As any number of lines can be drawn through a given scatter plot of a set of bivariate data, we have to develop a principle to select an optimal line. Statistical theory has developed the principle of least square to govern the choice of such an optimal line. Consider the following values of X and Y variables and their scatter plot.<\/p>\r\n&nbsp;\r\n<div>\r\n\r\n<strong>\u00a0 \u00a0 y<\/strong>\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 <strong>X<\/strong>\r\n\r\n20\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a02\r\n\r\n45\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a04\r\n\r\n65\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a07\r\n\r\n45\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a09\r\n\r\n70\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a012\r\n\r\n110\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 23\r\n\r\n89\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 15\r\n\r\n102\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a013\r\n\r\n110\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 18\r\n\r\n100\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 17\r\n\r\n<\/div>\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-311\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-142.png\" alt=\"\" width=\"372\" height=\"390\" \/>\r\n<p style=\"text-align: justify\">As is clear from the above graph, a line is drawn from the middle of the scatter in such a way that almost half of the points fall above it and half below it. Once the line is drawn, for every given value of X we have an estimated value of Y from the line apart from its given value. The difference between given value of Y and the estimated value of Y is known as the residual or error. For example in the above case for first given value of X=2 the given value is 20 whereas the estimated as given by line is 32. The difference is 32 \u2013 20 = + 12. In the second case for given value of X = 4 the given value of Y is 44 and the estimated value is 40 and the difference in the second case is 40 \u2013 44 = - 4. Likewise the residuals or errors for all other observations can also be calculated. For total error the sum of these residuals will be misleading as minus residuals will cancel the plus residuals. Residuals are, therefore, squared and then added to give the total error. If we draw different lines, we will get different values of the sum of the squares of residuals- one with every line. Ideally we will choose the line as optimal which will give us minimum value of the sum of the squares of residuals. This principle is known as <strong>Principle of Least Square<\/strong> and the line will be known as <strong>Regression Line. <\/strong>In practice we do not take several lines but take the help of Differential Calculus from mathematics which suggests that the slope \u201c b\u201d which is also known as the\u00a0\u00a0of such an optimum line will be <strong>:<\/strong><\/p>\r\n&nbsp;\r\n\r\n<strong>regression Coefficient<\/strong>\r\n\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-312\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-143.png\" alt=\"\" width=\"582\" height=\"65\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">After computing the regression coefficient the analysis also focuses on its magnitude. We test the hypothesis that whether the value of b is so very close to zero or not? In other words we test the hypothesis that can its actual value(or population value )may be considered as zero?<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">For testing such a hypothesis we use \u201ct\u201d test, according to which:<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Under the following assumptions:<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">1. The variable \u03b5 is a random variable distributed normally.\r\n2. Mean value of \u03b5is zero.\r\n3. Variance of \u03b5is constant for all value of X\r\n4. If there are more independent variables all are independent to each other i.e. there is no multi co-linearity among the independent variables.<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">A computed values of bwill follow \u201ct\u201d distribution with (n- 2) degrees of freedom i.e.<\/p>\r\n&nbsp;\r\n\r\nt = b \/ S. E. (b), where\r\n\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-313\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-144.png\" alt=\"\" width=\"635\" height=\"216\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">As estimated values of y are values found on the line there variations will always be less than the actual values of y. Objective of any regression line is always to get the estimated values of y giving variations as close to actual values of y as is possible. The ratio of Explained Some of square to total variations in y is therefore an indicator of the quality of any regression model and is known as coefficient of determination and denoted by R2 given by:<\/p>\r\n&nbsp;\r\n\r\n<strong>R<\/strong><strong>2<\/strong><strong> = Explained Sum of Squares\/ Total Sum of Squares.<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The value of R2 will vary between zero and unity. If a value is found to be say 0 .75 will mean that of the total variations in the dependent variable y 75 percent are being explained by the independent variable X chosen here.<\/p>\r\n<p style=\"text-align: justify\">Again to test the statistical significance of R2 we have F- ratio test as :<\/p>\r\n<p style=\"text-align: justify\">F = R2 (n-k)\/(1-R2)(k- 1)\u00a0 ( k-1, n-k), degrees of freedom<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>Example<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">From the data as given below we can compute the regression coefficient and the the intercept of the regression line as shown below :<\/p>\r\n&nbsp;\r\n<table class=\"aligncenter\" border=\"1\">\r\n<tbody>\r\n<tr>\r\n<td><strong>X<\/strong><\/td>\r\n<td><strong>y<\/strong><\/td>\r\n<td><strong>X<\/strong><strong>2<\/strong><\/td>\r\n<td><strong>Y<\/strong><strong>2<\/strong><\/td>\r\n<td><strong>YX<\/strong><\/td>\r\n<\/tr>\r\n<tr>\r\n<td>2<\/td>\r\n<td>20<\/td>\r\n<td>4<\/td>\r\n<td>400<\/td>\r\n<td>40<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>4<\/td>\r\n<td>45<\/td>\r\n<td>16<\/td>\r\n<td>2025<\/td>\r\n<td>90<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>7<\/td>\r\n<td>65<\/td>\r\n<td>49<\/td>\r\n<td>4225<\/td>\r\n<td>455<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>9<\/td>\r\n<td>45<\/td>\r\n<td>81<\/td>\r\n<td>2025<\/td>\r\n<td>405<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>12<\/td>\r\n<td>70<\/td>\r\n<td>144<\/td>\r\n<td>4900<\/td>\r\n<td>840<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>23<\/td>\r\n<td>110<\/td>\r\n<td>529<\/td>\r\n<td>12100<\/td>\r\n<td>2530<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>15<\/td>\r\n<td>89<\/td>\r\n<td>225<\/td>\r\n<td>7921<\/td>\r\n<td>1335<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>13<\/td>\r\n<td>102<\/td>\r\n<td>169<\/td>\r\n<td>10404<\/td>\r\n<td>1326<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>18<\/td>\r\n<td>110<\/td>\r\n<td>324<\/td>\r\n<td>12100<\/td>\r\n<td>1980<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>17<\/td>\r\n<td>100<\/td>\r\n<td>289<\/td>\r\n<td>10000<\/td>\r\n<td>1700<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong>Total\u00a0\u00a0 120<\/strong><\/td>\r\n<td><strong>756<\/strong><\/td>\r\n<td><strong>1830<\/strong><\/td>\r\n<td><strong>66100<\/strong><\/td>\r\n<td><strong>10701<\/strong><\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\nUsing the above formulas for regression coefficient b and the intercept a :\r\n\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-314\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-145.png\" alt=\"\" width=\"627\" height=\"53\" \/>\r\n<div>\r\n<p style=\"text-align: justify\">Regression analysis also requires calculation of standard error of b to carry out the statistical test of significance to assure that the regression coefficient \u201cb\u201d is not very close to zero.<\/p>\r\n&nbsp;\r\n\r\n<\/div>\r\n<img class=\"aligncenter size-full wp-image-315\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-146.png\" alt=\"\" width=\"619\" height=\"336\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">F value for d.f.= 1, 8 (2-1, 10-2)given in the table are 16.86 for 1% and 5.32 for 5% levels of significance respectively. Our calculated value is found to be significant at even at 1% level of significance. Thus we can say R2 is statistically significant from being zero. In other words we conclude that the independent variable is explaining the dependent variable in a substantial way ( not in any random way).<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>Multiple Linear Regression Analysis through Matrix Algebra<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Multiple Regression analysis is an extension of bivariate regression analysis, in which we have one dependent variable and two or more independent variables. In multiple regression analysis the regression model for p variables and n number of observations is;<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-316\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-147.png\" alt=\"\" width=\"546\" height=\"152\" \/>\r\n\r\n<img class=\"aligncenter size-full wp-image-317\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-148.png\" alt=\"\" width=\"489\" height=\"354\" \/>\r\n<p style=\"text-align: justify\">Where Y, B and <strong>\u03b5<\/strong> are column vectors of <em>p<\/em>x1 dimensions and X is a matrix of <em>nxp<\/em> dimensions. (here convention of writing row number first and column number next has been interchanged to suit to the equation formation. ). Applying the principle of least square to the above set of equations, we can get solution vector \u03b2 giving the estimated b1, b2, \u2026\u2026\u2026\u2026b<em>n<\/em> values of the constants of the regression line passing through the scatter of points as:<\/p>\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-318\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-149.png\" alt=\"\" width=\"578\" height=\"418\" \/>\r\n\r\n<img class=\"aligncenter size-full wp-image-319\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-150.png\" alt=\"\" width=\"507\" height=\"367\" \/>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">We can arrive at the same results using the method of normal equation from ordinary algebra. However, if the number of variables and number of observations are large as mostly is the case in geographical research it is quite useful to use matrix method as shown above. Advantage with matrix methods is that we can use some computer packages dealing in statistical analysis like SPSS, STATA and SAS etc.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Like bi-variate case we can also carry out the \u201ct\u201d test and work our out F- ratio for a multiple regression analysis also for which the values are given as:<\/p>\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-320\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-151.png\" alt=\"\" width=\"642\" height=\"432\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">significant at 1% level of significance. Thus the only regression coefficient b2 related to variable X1 has shown a significant effect on Y which is 1.7329 for a unit change in X1. Intercept b1 and the regression coefficient b3 related to variable X2 and the intercept are found to be insignificant even up to 5% level of significance.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-321\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-152.png\" alt=\"\" width=\"577\" height=\"29\" \/>\r\n\r\n&nbsp;\r\n<div>\r\n<p style=\"text-align: justify\">F value for d.f. = 2, 2 (3-1=2, 5-3=2) is 99.0 at 1% and 19.0 at 5% levels of significance respectively. The test of significance also reveals that the two independent variables X1 and X2 explain the variations in the dependent quite substantially.<\/p>\r\n\r\n<\/div>\r\n&nbsp;\r\n<div>\r\n\r\n<strong>\u00a0 \u00a0Stepwise Multiple Linear regression Analysis<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Stepwise regression analysis relates to evaluation of the relative efficiency of independent variables in explaining the dependent variable when the variables are added\/deleted to the model, one by one, in several steps. It has two approaches: Forward and backward. In forward stepwise regression, we start with one most important variable followed by the second, third \u2026\u2026.and the least important variables.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">In backward approach we start with all the variables and keep on excluding variables one by one and proceed in the reverse direction until we reach the optimal position.<\/p>\r\n\r\n<\/div>\r\n&nbsp;\r\n<div>\r\n\r\n<strong>\u00a0 Need for Stepwise regression analysis<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">One of the important assumption of the tests of significance of OLS regression analysis is that independent or explanatory variables are independent of each other. This assumption is known as assumption of absence of multi-collinearity. Sometimes this assumption is violated and variables are found to be collinear where independent variables show significant inter-correlation among themselves. In such cases there is some overlap in the explanatory power of the two or more collinear variables. Higher the relationship between the variables higher will be the overlap. For example if a variable is explaining 40 % of the dependent variable and another variable related to it is explaining 30 %. When these two variables are taken together, they may explain 70% of the dependent variable provided both the independent variables are independent. However, if they are collinear or correlated, they will explain less than 70%. Suppose these two variables together explain only 50 % of the dependent variable, it will be due to the fact that what second variable is contributing, 20 % of it has already been explained by the first variable. Second variable has now only 10 % contribution to explain in addition to first variable. In the absence of the first variable, however, second variable will have a higher explanatory power. In ordinary regression equation we will not have any idea about this complication arising due to the problem of multi-collinearity. Computer programmes have been develop to tackle this problem. Stepwise regression analysis is one such programme. In any regression model with some R2, if more variables are added the value of R2 will always increase, either the variable is\u00a0<span style=\"text-align: initial;font-size: 1em\">positively or negatively related to the dependent variable. In stepwise approach of a regression analysis, independent variables are sequentially added to the model one by one, untilthe criterion of variable addition is not met. The sequence starts in such a manner that in first step it gives regression line with one independent variable choosing from all the independent variables the one which gives maximum R2. In the second step it adds a one more independent variable to the model which adds maximum value to the existing R2. Likewisein third step one more variable is added and so on. Every time the addition to R2 due to new variable will be less than the previous value and the value of F-ratio will also change. In stepwise regression analysis we can fix a criterion of adding a new variable. Generally it is done by choosing the probability level of the changed value of F due to the addition of new variable. In most of the cases, if the probability exceeds the fixed limit say 0.05, the variable is not added to the model. R2 (adjusted) is designed in such a way that with the increase of every new variable it will decrease unless the new variable causes a significant increase in the value of R2. In step wise regression analysis, we keep on allowing the addition of new variables until R2 (Adjusted) increases. After few steps though R2 will continue to rise, R 2 (Adjusted) will start decreasing indicating the fact that addition to R2 is not big enough to be retained in the analysis.<\/span><\/p>\r\n\r\n<\/div>\r\n&nbsp;\r\n<div>\r\n\r\n<strong>\u00a0 \u00a0Example<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Declining sex ratio in India is a big concern of the society. There are a large number of factors behind it. In the following example , for the sake of simplicity in explanation, we have taken few of them and used a stepwise regression analysis to explain the variations in the \u201cSex Ratio\u201d in the 50 districts across Madhya Pradesh for 2011, with the help of the following variables.<\/p>\r\n&nbsp;\r\n\r\n1. Sex Ratio (Female per thousand male) (V1).\r\n\r\n2. Growth rate of population, 2001-11( in Percentage) (V2)\r\n\r\n3. Levels of Literacy ( in percentage) (V3)\r\n\r\n4. Population Density per square kilo meter of area (V4)\r\n\r\n5. Female work participation rate ( in percentage) (V5)\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Using SPSS when data( given in annexure) was subjected to the bivariate correlation and stepwise regression analysis, following results were obtained.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">First, it gives the inter-correlation matrix of each variable with other variables given in Table1, given below.It also identifies the level of significance at which these coefficients of correlation are significant.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The table shows that,the dependent variable Sex Ratio (V1)has anin-significant negative correlation coefficient with the variables: growth rate of population, 2001-11 (V2) and Levels of Literacy (V3) and a significant negative correlation with the variable\u00a0<span style=\"text-align: initial;font-size: 1em\">oLevels of Literacy (V3). It is found to be significant at 5% level of significance. The inter-correlation matrix also shows that the dependent variable has a strong positive relationship with the variable offemale work participation rate (V5), significant at 1 % level of significance.<\/span><\/p>\r\n\r\n<\/div>\r\n&nbsp;\r\n\r\n<strong>Note that :<\/strong>\r\n<ol>\r\n \t<li style=\"text-align: justify\">Two tail test means that a value of coefficient of correlation could be \u2260 0 i.e. it could be greater than or less than 0.<\/li>\r\n \t<li style=\"text-align: justify\">One tail test will mean that coefficient of correlation could be either &gt; 0 or &lt;0 i.e. either greater than 0 or less than 0.<\/li>\r\n \t<li style=\"text-align: justify\">The diagonal elements of an inter-correlation matrix are always 1, indicating the correlation of a variable with itself is perfect and coefficient of correlation will be ,therefore, 1.<\/li>\r\n<\/ol>\r\n<strong>\u00a0 \u00a0 Table 1<\/strong>\r\n\r\n&nbsp;\r\n\r\n<strong>Inter correlation Matrix<\/strong>\r\n\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-322\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-153.png\" alt=\"\" width=\"616\" height=\"381\" \/>\r\n\r\n&nbsp;\r\n<div>\r\n<p style=\"text-align: justify\">\u00a0 \u00a0 \u00a0*. Correlation is significant at the 0.05 level (2-tailed).<\/p>\r\n<p style=\"text-align: justify\">**. Correlation is significant at the 0.01 level (2-tailed).<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The inter-correlation matrix given above suggest reasonable justification for the choice of the explanatory or independent variables to explain the variations in the values of the dependent variable, Sex Ratio.<\/p>\r\n\r\n<\/div>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The above matrix of inter-correlations also suggest the overlap among the independent variables. The matrix shows inter-correlations among independent variables also. Fourth variable of the density of population (V4) has a strong positive relationship significant at 1% level of significance with the second variable of growth rate of population (V2) and third variable of literacy (V3) which has a strong significant positive relationship with the fifth variable of female work participation rate. Female work participation rate(V5) also has strong negative relationship with the density of population (V4).<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The inter-correlation among independent variables suggest that there exist some multicollinearity among them and an ordinary regression analysis will not be the optimal regression equation. A stepwise regression is likely to give better results by excluding the redundant variables and retaining only those which add a higher value to R2 as explained above. It will also give the order of the efficiency with which each of the independent variable explains the dependent variable of sex ratio.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Step wise regression analysis will give different models by adding independent variables one by one sequentially. The criterion to add the new variable is in terms of probability of its F value being less than 0.05 or 5% as is shown in Table 2 given below. We can also change the probability to 0.10 or 10% to allow more variables to enter into the analysis.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">In the present example the result given in Table 2 given below shows that only two variables; female work participation rate (V5) and density of population (V4) aresufficient to be retained in the multiple regression analysis. Other variables; population growth rate (V2) and level of literacy (V3) are not found to explain much the variations in the sex ratio of the districts of Madhya Pradesh. Their part of explanation is already explained by first two variable ; female work participation rate (V5).<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>Table 2<\/strong>\r\n\r\n&nbsp;\r\n\r\n<strong>Criteria forChoosing the Independent Variables<\/strong>\r\n\r\n&nbsp;\r\n\r\n<strong>Stepwise Regression Analysis<\/strong>\r\n\r\n<img class=\"aligncenter size-full wp-image-323\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-154.png\" alt=\"\" width=\"641\" height=\"202\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The summary results of the two steps of the regression model will follow in the computer output as given below in Table 3:<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>Table 3<\/strong>\r\n\r\n&nbsp;\r\n\r\n<strong>Model Summary<\/strong>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-324\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-155.png\" alt=\"\" width=\"558\" height=\"215\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Model 1 given by the first step shows that only one variable i.e. female work participation rate (V5) alone explains the sex ratio quite effectively. It explains 70.6 % variations of the sex ratio across 50 districts of Madhya Pradesh as per the data provided by the Census of India.The next variable which could be included in the model is the population density (V4) which could add to the explanatory power of the model only 3.6 % as the value of R2 could rise from 0.706 to 0.742 only. R2 (adjusted) could also rose from 0.700 to 0.731 only. Another two variables could not qualify the criterion of entering into the analysis du to multicollinearity.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Once the model is chosen, the main results follow. These include regression coefficient of the selected variables their standard errors, \u201ct-statistics\u201d and the level of their significance.These results are also given in Table 4 below.<\/p>\r\n&nbsp;\r\n\r\n<strong>Table 4<\/strong>\r\n\r\n&nbsp;\r\n\r\n<strong>Coefficients<\/strong>\r\n\r\n<img class=\"aligncenter size-full wp-image-325\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-156.png\" alt=\"\" width=\"579\" height=\"226\" \/>\r\n<div>\r\n<p style=\"text-align: justify\">Table 4 given above show the regression coefficients in unstandardized form as well as\u00a0<span style=\"text-align: initial;font-size: 1em\">in standardized form. Unstandardized coefficients relate to the data as provided in the computer input and also gives the value of the intercept as constant. Computer also converts the given data into their standard scores and give corresponding regression coefficients as standardized coefficients. The purpose is to bring the data to a standard form of zero mean and unit standard deviations. In standardized form when the mean of all the variables is zerointercept is not given as it also become zero. Standard error and \u2018t- statistics\u2019 of the regression coefficient, however, in both the cases remain the same.<\/span><\/p>\r\n\r\n<\/div>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The results of the above table how that as it is a unit change in the employment to female will promote an increase of 4.214 rise in the sex ratio. Whereas the density of population does not show much impact on the sex ratio. A change of one person per square km. will bring a change of only o.063 change in sex ratio.<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">It is important to note that female work participation rate varies with in a narrow range from 8.4 in Bhid to 52.9 in Dindori. Density of population has quite big range of variation from 855 in Bhopal to 94 in Annupur. These variations have been standardized by converting all the three variables into their standard scores. As a result the gap of 3.832 between the regression coefficients of unstandardized form (4.214 - 0.382 = 3.832 )is reduced to 0.734 between the same in unstandardized form ( 0.954 \u2013 0.220 ). The proportion of the two has also been reduce from 11.03 (= 4.214\/ 0.382 ) to 4.33 (= 0.954\/0.220).<\/p>\r\n&nbsp;\r\n\r\n<strong>Annexure I : Data for Stepwise Regression Analysis<\/strong>\r\n\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-326\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-157.png\" alt=\"\" width=\"622\" height=\"323\" \/>\r\n\r\n<img class=\"aligncenter size-full wp-image-327\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-158.png\" alt=\"\" width=\"628\" height=\"457\" \/>\r\n\r\n<img class=\"aligncenter size-full wp-image-328\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-159.png\" alt=\"\" width=\"623\" height=\"461\" \/>\r\n\r\n<img class=\"aligncenter size-full wp-image-329\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-160.png\" alt=\"\" width=\"627\" height=\"492\" \/>\r\n\r\n<img class=\"aligncenter size-full wp-image-330\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-161.png\" alt=\"\" width=\"624\" height=\"76\" \/>\r\n\r\n&nbsp;\r\n\r\n<span style=\"text-align: initial;text-indent: 1em;font-size: 1em\">\r\nSource: Census of India 2011<\/span>\r\n<ol>\r\n \t<li>Sex Ratio (Female per thousand male) (V1).<\/li>\r\n \t<li>Growth rate of population, 2001-11( in Percentage) (V2)<\/li>\r\n \t<li>Levels of Literacy ( in percentage) (V3)<\/li>\r\n \t<li>Population Density per square kilo meter of area (V4)<\/li>\r\n \t<li>Female work participation rate ( in percentage) (V5)<\/li>\r\n<\/ol>\r\n&nbsp;\r\n<table>\r\n<tbody>\r\n<tr>\r\n<td><strong>you can view video on Multiple linear regression analysis ( simple and step- wise regression)<\/strong><\/td>\r\n<td><a href=\"https:\/\/youtu.be\/7RLoldYt7wg\" target=\"_blank\" rel=\"noopener\"><img class=\"alignnone wp-image-120\" src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"\" width=\"36\" height=\"36\" \/><\/a><\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n&nbsp;\r\n<div class=\"textbox learning-objectives\">\r\n<h3>References<\/h3>\r\n<ul>\r\n \t<li style=\"text-align: justify\">Aslam Mahmood (1993). Statistical Methods in Geographical Research, Rajesh Publications, New Delhi.<\/li>\r\n \t<li style=\"text-align: justify\">Johnston J. (1972) Mc Graw Hills, pp 8<\/li>\r\n \t<li style=\"text-align: justify\">David Harvey (1969), Explanation in Geography, Edward Arnold London.<\/li>\r\n \t<li style=\"text-align: justify\">Koutsoyiannis A(1973). Theory of Econometric Mcmillan pp 225 \u2013 49.<\/li>\r\n \t<li style=\"text-align: justify\">Retherford R.D. and Choe M.K.( 1973) Statistical Methods For Causal Ananlysis. John Wiley &amp; Sons, INC.<\/li>\r\n \t<li style=\"text-align: justify\">Wooldridge J.M. Introduction to Econometrics: A Modern Approach (2009) Cengage Learning , India Pvt. Ltd. (Indian Edition).<\/li>\r\n<\/ul>\r\n<\/div>","rendered":"<div><span style=\"float: right\"><a href=\"https:\/\/youtu.be\/7RLoldYt7wg\" target=\"_blank\" rel=\"noopener\"><img decoding=\"async\" src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"epgp books\" width=\"75px\" height=\"75px;\" \/><\/a><br \/>\n<\/span><\/div>\n<p><strong>\u00a0 \u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>1) E text<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><strong>Multiple Linear Regression Analysis<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><strong>Regression Analysis<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Ultimate objective of any research is to understand the causes behind a process with objective to control it as per our requirement. Knowledge of the degree and direction of co relationship between a dependent and independent variable alone is not sufficient to help in this regard unless it gives the mathematical equation giving the relationship between them. Through this form of causal relationship (Cause and effect relationship ) we can evaluate the effect of different independent variablesalso known as determinants on the dependent variable and can also make projections of the dependent variable for some projected future values of independent variables etc.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Regression analysis helps us in identifying the above mentioned functional form of relationship. If we have only two variables to analyze, i.e. one dependent and one independent variable, the regression is known as bivariate regression. However, if the number of independent variables are two or more it is known as Multiple Regression analysis. Explanation of regression analysis becomes very complicated if we directly start from multiple regression analysis. It is easier to explain bivariate regression analysis first and explain Multiple Regression as its extension.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Bivariate Regression Analysis<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Functional form of relationship between to variables is best explained by taking a straight line on a graph paper having Y-coordinate and X-coordinate. If we take large number of points on the line and note their Y and X coordinates. The Y and X coordinates of these points will be related by a general functional form as : <strong>Y = a + b X +<\/strong> <strong>\u03b5<\/strong>. In any such equation of straight line <strong>a<\/strong> is known as intercept of the line i.e. the value at Y axis where the line will cut Y- axis (value of Y after putting X = 0.), <strong>b<\/strong>therate of change in Y with respect to X,is known as the slope of the line and<strong>\u03b5<\/strong> is known as the error and the relationship Y = a + b X. is known as linear relationship (as it relates to a line only). Thus if we refer to any specific straight line, we can specify it by fixing the values of <strong>a<\/strong> and <strong>b<\/strong>. However, most of the bivariate data from social sciences, though following a linear pattern of relationship, will not exactly fall on a straight line. On a scatter plot of such data the points may cluster around a straight line and give some deviations on either side of the line as indicated by<strong>\u03b5<\/strong>. In such cases the line around which the points cluster could be identified to give the relationship<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">between two variables X and Y. Here the values of Y and X will fall on the line and relationship between them will be exact. Actual relationship between values of Y and X falling around the line will only be approximate to it. The value of the coefficient of correlation will give the direction and the degree of such linear relationships. The equation of the line passing through the scatter plot will give the functional form of linear relationship between the two variables.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Principle of Least Square<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">As any number of lines can be drawn through a given scatter plot of a set of bivariate data, we have to develop a principle to select an optimal line. Statistical theory has developed the principle of least square to govern the choice of such an optimal line. Consider the following values of X and Y variables and their scatter plot.<\/p>\n<p>&nbsp;<\/p>\n<div>\n<p><strong>\u00a0 \u00a0 y<\/strong>\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 <strong>X<\/strong><\/p>\n<p>20\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a02<\/p>\n<p>45\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a04<\/p>\n<p>65\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a07<\/p>\n<p>45\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a09<\/p>\n<p>70\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a012<\/p>\n<p>110\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 23<\/p>\n<p>89\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 15<\/p>\n<p>102\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a013<\/p>\n<p>110\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 18<\/p>\n<p>100\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 17<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-311\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-142.png\" alt=\"\" width=\"372\" height=\"390\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-142.png 372w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-142-286x300.png 286w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-142-65x68.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-142-225x236.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-142-350x367.png 350w\" sizes=\"auto, (max-width: 372px) 100vw, 372px\" \/><\/p>\n<p style=\"text-align: justify\">As is clear from the above graph, a line is drawn from the middle of the scatter in such a way that almost half of the points fall above it and half below it. Once the line is drawn, for every given value of X we have an estimated value of Y from the line apart from its given value. The difference between given value of Y and the estimated value of Y is known as the residual or error. For example in the above case for first given value of X=2 the given value is 20 whereas the estimated as given by line is 32. The difference is 32 \u2013 20 = + 12. In the second case for given value of X = 4 the given value of Y is 44 and the estimated value is 40 and the difference in the second case is 40 \u2013 44 = &#8211; 4. Likewise the residuals or errors for all other observations can also be calculated. For total error the sum of these residuals will be misleading as minus residuals will cancel the plus residuals. Residuals are, therefore, squared and then added to give the total error. If we draw different lines, we will get different values of the sum of the squares of residuals- one with every line. Ideally we will choose the line as optimal which will give us minimum value of the sum of the squares of residuals. This principle is known as <strong>Principle of Least Square<\/strong> and the line will be known as <strong>Regression Line. <\/strong>In practice we do not take several lines but take the help of Differential Calculus from mathematics which suggests that the slope \u201c b\u201d which is also known as the\u00a0\u00a0of such an optimum line will be <strong>:<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><strong>regression Coefficient<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-312\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-143.png\" alt=\"\" width=\"582\" height=\"65\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-143.png 582w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-143-300x34.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-143-65x7.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-143-225x25.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-143-350x39.png 350w\" sizes=\"auto, (max-width: 582px) 100vw, 582px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">After computing the regression coefficient the analysis also focuses on its magnitude. We test the hypothesis that whether the value of b is so very close to zero or not? In other words we test the hypothesis that can its actual value(or population value )may be considered as zero?<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">For testing such a hypothesis we use \u201ct\u201d test, according to which:<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Under the following assumptions:<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">1. The variable \u03b5 is a random variable distributed normally.<br \/>\n2. Mean value of \u03b5is zero.<br \/>\n3. Variance of \u03b5is constant for all value of X<br \/>\n4. If there are more independent variables all are independent to each other i.e. there is no multi co-linearity among the independent variables.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">A computed values of bwill follow \u201ct\u201d distribution with (n- 2) degrees of freedom i.e.<\/p>\n<p>&nbsp;<\/p>\n<p>t = b \/ S. E. (b), where<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-313\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-144.png\" alt=\"\" width=\"635\" height=\"216\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-144.png 635w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-144-300x102.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-144-65x22.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-144-225x77.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-144-350x119.png 350w\" sizes=\"auto, (max-width: 635px) 100vw, 635px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">As estimated values of y are values found on the line there variations will always be less than the actual values of y. Objective of any regression line is always to get the estimated values of y giving variations as close to actual values of y as is possible. The ratio of Explained Some of square to total variations in y is therefore an indicator of the quality of any regression model and is known as coefficient of determination and denoted by R2 given by:<\/p>\n<p>&nbsp;<\/p>\n<p><strong>R<\/strong><strong>2<\/strong><strong> = Explained Sum of Squares\/ Total Sum of Squares.<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The value of R2 will vary between zero and unity. If a value is found to be say 0 .75 will mean that of the total variations in the dependent variable y 75 percent are being explained by the independent variable X chosen here.<\/p>\n<p style=\"text-align: justify\">Again to test the statistical significance of R2 we have F- ratio test as :<\/p>\n<p style=\"text-align: justify\">F = R2 (n-k)\/(1-R2)(k- 1)\u00a0 ( k-1, n-k), degrees of freedom<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Example<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">From the data as given below we can compute the regression coefficient and the the intercept of the regression line as shown below :<\/p>\n<p>&nbsp;<\/p>\n<table class=\"aligncenter\">\n<tbody>\n<tr>\n<td><strong>X<\/strong><\/td>\n<td><strong>y<\/strong><\/td>\n<td><strong>X<\/strong><strong>2<\/strong><\/td>\n<td><strong>Y<\/strong><strong>2<\/strong><\/td>\n<td><strong>YX<\/strong><\/td>\n<\/tr>\n<tr>\n<td>2<\/td>\n<td>20<\/td>\n<td>4<\/td>\n<td>400<\/td>\n<td>40<\/td>\n<\/tr>\n<tr>\n<td>4<\/td>\n<td>45<\/td>\n<td>16<\/td>\n<td>2025<\/td>\n<td>90<\/td>\n<\/tr>\n<tr>\n<td>7<\/td>\n<td>65<\/td>\n<td>49<\/td>\n<td>4225<\/td>\n<td>455<\/td>\n<\/tr>\n<tr>\n<td>9<\/td>\n<td>45<\/td>\n<td>81<\/td>\n<td>2025<\/td>\n<td>405<\/td>\n<\/tr>\n<tr>\n<td>12<\/td>\n<td>70<\/td>\n<td>144<\/td>\n<td>4900<\/td>\n<td>840<\/td>\n<\/tr>\n<tr>\n<td>23<\/td>\n<td>110<\/td>\n<td>529<\/td>\n<td>12100<\/td>\n<td>2530<\/td>\n<\/tr>\n<tr>\n<td>15<\/td>\n<td>89<\/td>\n<td>225<\/td>\n<td>7921<\/td>\n<td>1335<\/td>\n<\/tr>\n<tr>\n<td>13<\/td>\n<td>102<\/td>\n<td>169<\/td>\n<td>10404<\/td>\n<td>1326<\/td>\n<\/tr>\n<tr>\n<td>18<\/td>\n<td>110<\/td>\n<td>324<\/td>\n<td>12100<\/td>\n<td>1980<\/td>\n<\/tr>\n<tr>\n<td>17<\/td>\n<td>100<\/td>\n<td>289<\/td>\n<td>10000<\/td>\n<td>1700<\/td>\n<\/tr>\n<tr>\n<td><strong>Total\u00a0\u00a0 120<\/strong><\/td>\n<td><strong>756<\/strong><\/td>\n<td><strong>1830<\/strong><\/td>\n<td><strong>66100<\/strong><\/td>\n<td><strong>10701<\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>Using the above formulas for regression coefficient b and the intercept a :<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-314\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-145.png\" alt=\"\" width=\"627\" height=\"53\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-145.png 627w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-145-300x25.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-145-65x5.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-145-225x19.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-145-350x30.png 350w\" sizes=\"auto, (max-width: 627px) 100vw, 627px\" \/><\/p>\n<div>\n<p style=\"text-align: justify\">Regression analysis also requires calculation of standard error of b to carry out the statistical test of significance to assure that the regression coefficient \u201cb\u201d is not very close to zero.<\/p>\n<p>&nbsp;<\/p>\n<\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-315\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-146.png\" alt=\"\" width=\"619\" height=\"336\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-146.png 619w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-146-300x163.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-146-65x35.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-146-225x122.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-146-350x190.png 350w\" sizes=\"auto, (max-width: 619px) 100vw, 619px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">F value for d.f.= 1, 8 (2-1, 10-2)given in the table are 16.86 for 1% and 5.32 for 5% levels of significance respectively. Our calculated value is found to be significant at even at 1% level of significance. Thus we can say R2 is statistically significant from being zero. In other words we conclude that the independent variable is explaining the dependent variable in a substantial way ( not in any random way).<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Multiple Linear Regression Analysis through Matrix Algebra<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Multiple Regression analysis is an extension of bivariate regression analysis, in which we have one dependent variable and two or more independent variables. In multiple regression analysis the regression model for p variables and n number of observations is;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-316\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-147.png\" alt=\"\" width=\"546\" height=\"152\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-147.png 546w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-147-300x84.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-147-65x18.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-147-225x63.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-147-350x97.png 350w\" sizes=\"auto, (max-width: 546px) 100vw, 546px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-317\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-148.png\" alt=\"\" width=\"489\" height=\"354\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-148.png 489w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-148-300x217.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-148-65x47.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-148-225x163.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-148-350x253.png 350w\" sizes=\"auto, (max-width: 489px) 100vw, 489px\" \/><\/p>\n<p style=\"text-align: justify\">Where Y, B and <strong>\u03b5<\/strong> are column vectors of <em>p<\/em>x1 dimensions and X is a matrix of <em>nxp<\/em> dimensions. (here convention of writing row number first and column number next has been interchanged to suit to the equation formation. ). Applying the principle of least square to the above set of equations, we can get solution vector \u03b2 giving the estimated b1, b2, \u2026\u2026\u2026\u2026b<em>n<\/em> values of the constants of the regression line passing through the scatter of points as:<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-318\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-149.png\" alt=\"\" width=\"578\" height=\"418\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-149.png 578w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-149-300x217.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-149-65x47.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-149-225x163.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-149-350x253.png 350w\" sizes=\"auto, (max-width: 578px) 100vw, 578px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-319\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-150.png\" alt=\"\" width=\"507\" height=\"367\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-150.png 507w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-150-300x217.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-150-65x47.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-150-225x163.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-150-350x253.png 350w\" sizes=\"auto, (max-width: 507px) 100vw, 507px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">We can arrive at the same results using the method of normal equation from ordinary algebra. However, if the number of variables and number of observations are large as mostly is the case in geographical research it is quite useful to use matrix method as shown above. Advantage with matrix methods is that we can use some computer packages dealing in statistical analysis like SPSS, STATA and SAS etc.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Like bi-variate case we can also carry out the \u201ct\u201d test and work our out F- ratio for a multiple regression analysis also for which the values are given as:<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-320\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-151.png\" alt=\"\" width=\"642\" height=\"432\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-151.png 642w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-151-300x202.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-151-65x44.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-151-225x151.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-151-350x236.png 350w\" sizes=\"auto, (max-width: 642px) 100vw, 642px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">significant at 1% level of significance. Thus the only regression coefficient b2 related to variable X1 has shown a significant effect on Y which is 1.7329 for a unit change in X1. Intercept b1 and the regression coefficient b3 related to variable X2 and the intercept are found to be insignificant even up to 5% level of significance.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-321\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-152.png\" alt=\"\" width=\"577\" height=\"29\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-152.png 577w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-152-300x15.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-152-65x3.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-152-225x11.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-152-350x18.png 350w\" sizes=\"auto, (max-width: 577px) 100vw, 577px\" \/><\/p>\n<p>&nbsp;<\/p>\n<div>\n<p style=\"text-align: justify\">F value for d.f. = 2, 2 (3-1=2, 5-3=2) is 99.0 at 1% and 19.0 at 5% levels of significance respectively. The test of significance also reveals that the two independent variables X1 and X2 explain the variations in the dependent quite substantially.<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<div>\n<p><strong>\u00a0 \u00a0Stepwise Multiple Linear regression Analysis<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Stepwise regression analysis relates to evaluation of the relative efficiency of independent variables in explaining the dependent variable when the variables are added\/deleted to the model, one by one, in several steps. It has two approaches: Forward and backward. In forward stepwise regression, we start with one most important variable followed by the second, third \u2026\u2026.and the least important variables.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In backward approach we start with all the variables and keep on excluding variables one by one and proceed in the reverse direction until we reach the optimal position.<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<div>\n<p><strong>\u00a0 Need for Stepwise regression analysis<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">One of the important assumption of the tests of significance of OLS regression analysis is that independent or explanatory variables are independent of each other. This assumption is known as assumption of absence of multi-collinearity. Sometimes this assumption is violated and variables are found to be collinear where independent variables show significant inter-correlation among themselves. In such cases there is some overlap in the explanatory power of the two or more collinear variables. Higher the relationship between the variables higher will be the overlap. For example if a variable is explaining 40 % of the dependent variable and another variable related to it is explaining 30 %. When these two variables are taken together, they may explain 70% of the dependent variable provided both the independent variables are independent. However, if they are collinear or correlated, they will explain less than 70%. Suppose these two variables together explain only 50 % of the dependent variable, it will be due to the fact that what second variable is contributing, 20 % of it has already been explained by the first variable. Second variable has now only 10 % contribution to explain in addition to first variable. In the absence of the first variable, however, second variable will have a higher explanatory power. In ordinary regression equation we will not have any idea about this complication arising due to the problem of multi-collinearity. Computer programmes have been develop to tackle this problem. Stepwise regression analysis is one such programme. In any regression model with some R2, if more variables are added the value of R2 will always increase, either the variable is\u00a0<span style=\"text-align: initial;font-size: 1em\">positively or negatively related to the dependent variable. In stepwise approach of a regression analysis, independent variables are sequentially added to the model one by one, untilthe criterion of variable addition is not met. The sequence starts in such a manner that in first step it gives regression line with one independent variable choosing from all the independent variables the one which gives maximum R2. In the second step it adds a one more independent variable to the model which adds maximum value to the existing R2. Likewisein third step one more variable is added and so on. Every time the addition to R2 due to new variable will be less than the previous value and the value of F-ratio will also change. In stepwise regression analysis we can fix a criterion of adding a new variable. Generally it is done by choosing the probability level of the changed value of F due to the addition of new variable. In most of the cases, if the probability exceeds the fixed limit say 0.05, the variable is not added to the model. R2 (adjusted) is designed in such a way that with the increase of every new variable it will decrease unless the new variable causes a significant increase in the value of R2. In step wise regression analysis, we keep on allowing the addition of new variables until R2 (Adjusted) increases. After few steps though R2 will continue to rise, R 2 (Adjusted) will start decreasing indicating the fact that addition to R2 is not big enough to be retained in the analysis.<\/span><\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<div>\n<p><strong>\u00a0 \u00a0Example<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Declining sex ratio in India is a big concern of the society. There are a large number of factors behind it. In the following example , for the sake of simplicity in explanation, we have taken few of them and used a stepwise regression analysis to explain the variations in the \u201cSex Ratio\u201d in the 50 districts across Madhya Pradesh for 2011, with the help of the following variables.<\/p>\n<p>&nbsp;<\/p>\n<p>1. Sex Ratio (Female per thousand male) (V1).<\/p>\n<p>2. Growth rate of population, 2001-11( in Percentage) (V2)<\/p>\n<p>3. Levels of Literacy ( in percentage) (V3)<\/p>\n<p>4. Population Density per square kilo meter of area (V4)<\/p>\n<p>5. Female work participation rate ( in percentage) (V5)<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Using SPSS when data( given in annexure) was subjected to the bivariate correlation and stepwise regression analysis, following results were obtained.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">First, it gives the inter-correlation matrix of each variable with other variables given in Table1, given below.It also identifies the level of significance at which these coefficients of correlation are significant.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The table shows that,the dependent variable Sex Ratio (V1)has anin-significant negative correlation coefficient with the variables: growth rate of population, 2001-11 (V2) and Levels of Literacy (V3) and a significant negative correlation with the variable\u00a0<span style=\"text-align: initial;font-size: 1em\">oLevels of Literacy (V3). It is found to be significant at 5% level of significance. The inter-correlation matrix also shows that the dependent variable has a strong positive relationship with the variable offemale work participation rate (V5), significant at 1 % level of significance.<\/span><\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p><strong>Note that :<\/strong><\/p>\n<ol>\n<li style=\"text-align: justify\">Two tail test means that a value of coefficient of correlation could be \u2260 0 i.e. it could be greater than or less than 0.<\/li>\n<li style=\"text-align: justify\">One tail test will mean that coefficient of correlation could be either &gt; 0 or &lt;0 i.e. either greater than 0 or less than 0.<\/li>\n<li style=\"text-align: justify\">The diagonal elements of an inter-correlation matrix are always 1, indicating the correlation of a variable with itself is perfect and coefficient of correlation will be ,therefore, 1.<\/li>\n<\/ol>\n<p><strong>\u00a0 \u00a0 Table 1<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><strong>Inter correlation Matrix<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-322\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-153.png\" alt=\"\" width=\"616\" height=\"381\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-153.png 616w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-153-300x186.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-153-65x40.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-153-225x139.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-153-350x216.png 350w\" sizes=\"auto, (max-width: 616px) 100vw, 616px\" \/><\/p>\n<p>&nbsp;<\/p>\n<div>\n<p style=\"text-align: justify\">\u00a0 \u00a0 \u00a0*. Correlation is significant at the 0.05 level (2-tailed).<\/p>\n<p style=\"text-align: justify\">**. Correlation is significant at the 0.01 level (2-tailed).<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The inter-correlation matrix given above suggest reasonable justification for the choice of the explanatory or independent variables to explain the variations in the values of the dependent variable, Sex Ratio.<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The above matrix of inter-correlations also suggest the overlap among the independent variables. The matrix shows inter-correlations among independent variables also. Fourth variable of the density of population (V4) has a strong positive relationship significant at 1% level of significance with the second variable of growth rate of population (V2) and third variable of literacy (V3) which has a strong significant positive relationship with the fifth variable of female work participation rate. Female work participation rate(V5) also has strong negative relationship with the density of population (V4).<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The inter-correlation among independent variables suggest that there exist some multicollinearity among them and an ordinary regression analysis will not be the optimal regression equation. A stepwise regression is likely to give better results by excluding the redundant variables and retaining only those which add a higher value to R2 as explained above. It will also give the order of the efficiency with which each of the independent variable explains the dependent variable of sex ratio.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Step wise regression analysis will give different models by adding independent variables one by one sequentially. The criterion to add the new variable is in terms of probability of its F value being less than 0.05 or 5% as is shown in Table 2 given below. We can also change the probability to 0.10 or 10% to allow more variables to enter into the analysis.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In the present example the result given in Table 2 given below shows that only two variables; female work participation rate (V5) and density of population (V4) aresufficient to be retained in the multiple regression analysis. Other variables; population growth rate (V2) and level of literacy (V3) are not found to explain much the variations in the sex ratio of the districts of Madhya Pradesh. Their part of explanation is already explained by first two variable ; female work participation rate (V5).<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Table 2<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><strong>Criteria forChoosing the Independent Variables<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><strong>Stepwise Regression Analysis<\/strong><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-323\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-154.png\" alt=\"\" width=\"641\" height=\"202\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-154.png 641w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-154-300x95.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-154-65x20.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-154-225x71.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-154-350x110.png 350w\" sizes=\"auto, (max-width: 641px) 100vw, 641px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The summary results of the two steps of the regression model will follow in the computer output as given below in Table 3:<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Table 3<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><strong>Model Summary<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-324\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-155.png\" alt=\"\" width=\"558\" height=\"215\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-155.png 558w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-155-300x116.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-155-65x25.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-155-225x87.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-155-350x135.png 350w\" sizes=\"auto, (max-width: 558px) 100vw, 558px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Model 1 given by the first step shows that only one variable i.e. female work participation rate (V5) alone explains the sex ratio quite effectively. It explains 70.6 % variations of the sex ratio across 50 districts of Madhya Pradesh as per the data provided by the Census of India.The next variable which could be included in the model is the population density (V4) which could add to the explanatory power of the model only 3.6 % as the value of R2 could rise from 0.706 to 0.742 only. R2 (adjusted) could also rose from 0.700 to 0.731 only. Another two variables could not qualify the criterion of entering into the analysis du to multicollinearity.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Once the model is chosen, the main results follow. These include regression coefficient of the selected variables their standard errors, \u201ct-statistics\u201d and the level of their significance.These results are also given in Table 4 below.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Table 4<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><strong>Coefficients<\/strong><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-325\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-156.png\" alt=\"\" width=\"579\" height=\"226\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-156.png 579w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-156-300x117.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-156-65x25.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-156-225x88.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-156-350x137.png 350w\" sizes=\"auto, (max-width: 579px) 100vw, 579px\" \/><\/p>\n<div>\n<p style=\"text-align: justify\">Table 4 given above show the regression coefficients in unstandardized form as well as\u00a0<span style=\"text-align: initial;font-size: 1em\">in standardized form. Unstandardized coefficients relate to the data as provided in the computer input and also gives the value of the intercept as constant. Computer also converts the given data into their standard scores and give corresponding regression coefficients as standardized coefficients. The purpose is to bring the data to a standard form of zero mean and unit standard deviations. In standardized form when the mean of all the variables is zerointercept is not given as it also become zero. Standard error and \u2018t- statistics\u2019 of the regression coefficient, however, in both the cases remain the same.<\/span><\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The results of the above table how that as it is a unit change in the employment to female will promote an increase of 4.214 rise in the sex ratio. Whereas the density of population does not show much impact on the sex ratio. A change of one person per square km. will bring a change of only o.063 change in sex ratio.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">It is important to note that female work participation rate varies with in a narrow range from 8.4 in Bhid to 52.9 in Dindori. Density of population has quite big range of variation from 855 in Bhopal to 94 in Annupur. These variations have been standardized by converting all the three variables into their standard scores. As a result the gap of 3.832 between the regression coefficients of unstandardized form (4.214 &#8211; 0.382 = 3.832 )is reduced to 0.734 between the same in unstandardized form ( 0.954 \u2013 0.220 ). The proportion of the two has also been reduce from 11.03 (= 4.214\/ 0.382 ) to 4.33 (= 0.954\/0.220).<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Annexure I : Data for Stepwise Regression Analysis<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-326\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-157.png\" alt=\"\" width=\"622\" height=\"323\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-157.png 622w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-157-300x156.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-157-65x34.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-157-225x117.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-157-350x182.png 350w\" sizes=\"auto, (max-width: 622px) 100vw, 622px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-327\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-158.png\" alt=\"\" width=\"628\" height=\"457\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-158.png 628w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-158-300x218.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-158-65x47.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-158-225x164.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-158-350x255.png 350w\" sizes=\"auto, (max-width: 628px) 100vw, 628px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-328\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-159.png\" alt=\"\" width=\"623\" height=\"461\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-159.png 623w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-159-300x222.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-159-65x48.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-159-225x166.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-159-350x259.png 350w\" sizes=\"auto, (max-width: 623px) 100vw, 623px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-329\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-160.png\" alt=\"\" width=\"627\" height=\"492\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-160.png 627w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-160-300x235.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-160-65x51.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-160-225x177.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-160-350x275.png 350w\" sizes=\"auto, (max-width: 627px) 100vw, 627px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-330\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-161.png\" alt=\"\" width=\"624\" height=\"76\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-161.png 624w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-161-300x37.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-161-65x8.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-161-225x27.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-161-350x43.png 350w\" sizes=\"auto, (max-width: 624px) 100vw, 624px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p><span style=\"text-align: initial;text-indent: 1em;font-size: 1em\"><br \/>\nSource: Census of India 2011<\/span><\/p>\n<ol>\n<li>Sex Ratio (Female per thousand male) (V1).<\/li>\n<li>Growth rate of population, 2001-11( in Percentage) (V2)<\/li>\n<li>Levels of Literacy ( in percentage) (V3)<\/li>\n<li>Population Density per square kilo meter of area (V4)<\/li>\n<li>Female work participation rate ( in percentage) (V5)<\/li>\n<\/ol>\n<p>&nbsp;<\/p>\n<table>\n<tbody>\n<tr>\n<td><strong>you can view video on Multiple linear regression analysis ( simple and step- wise regression)<\/strong><\/td>\n<td><a href=\"https:\/\/youtu.be\/7RLoldYt7wg\" target=\"_blank\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-120\" src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"\" width=\"36\" height=\"36\" \/><\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<div class=\"textbox learning-objectives\">\n<h3>References<\/h3>\n<ul>\n<li style=\"text-align: justify\">Aslam Mahmood (1993). Statistical Methods in Geographical Research, Rajesh Publications, New Delhi.<\/li>\n<li style=\"text-align: justify\">Johnston J. (1972) Mc Graw Hills, pp 8<\/li>\n<li style=\"text-align: justify\">David Harvey (1969), Explanation in Geography, Edward Arnold London.<\/li>\n<li style=\"text-align: justify\">Koutsoyiannis A(1973). Theory of Econometric Mcmillan pp 225 \u2013 49.<\/li>\n<li style=\"text-align: justify\">Retherford R.D. and Choe M.K.( 1973) Statistical Methods For Causal Ananlysis. John Wiley &amp; Sons, INC.<\/li>\n<li style=\"text-align: justify\">Wooldridge J.M. Introduction to Econometrics: A Modern Approach (2009) Cengage Learning , India Pvt. Ltd. (Indian Edition).<\/li>\n<\/ul>\n<\/div>\n","protected":false},"author":12,"menu_order":19,"template":"","meta":{"pb_show_title":"on","pb_short_title":"","pb_subtitle":"","pb_authors":["prof-aslam-mahmood"],"pb_section_license":""},"chapter-type":[],"contributor":[58],"license":[],"class_list":["post-310","chapter","type-chapter","status-publish","hentry","contributor-prof-aslam-mahmood"],"part":3,"_links":{"self":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapters\/310","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapters"}],"about":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/wp\/v2\/types\/chapter"}],"author":[{"embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/wp\/v2\/users\/12"}],"version-history":[{"count":6,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapters\/310\/revisions"}],"predecessor-version":[{"id":336,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapters\/310\/revisions\/336"}],"part":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/parts\/3"}],"metadata":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapters\/310\/metadata\/"}],"wp:attachment":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/wp\/v2\/media?parent=310"}],"wp:term":[{"taxonomy":"chapter-type","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapter-type?post=310"},{"taxonomy":"contributor","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/wp\/v2\/contributor?post=310"},{"taxonomy":"license","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/wp\/v2\/license?post=310"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}