{"id":277,"date":"2019-04-24T09:38:29","date_gmt":"2019-04-24T09:38:29","guid":{"rendered":"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/?post_type=chapter&#038;p=277"},"modified":"2019-04-24T09:42:11","modified_gmt":"2019-04-24T09:42:11","slug":"introduction-to-bivariate-linear-regression-analysis","status":"publish","type":"chapter","link":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/chapter\/introduction-to-bivariate-linear-regression-analysis\/","title":{"rendered":"Introduction to Bivariate Linear Regression analysis"},"content":{"raw":"<div><span style=\"float: right\"><a href=\"https:\/\/youtu.be\/EoAzZXSEchc\" target=\"_blank\" rel=\"noopener\"><img src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"epgp books\" width=\"75px\" height=\"75px;\" \/><\/a>\r\n<\/span><\/div>\r\n<div><strong>\u00a0 \u00a0<\/strong><\/div>\r\n<div><\/div>\r\n<div><\/div>\r\n<div>\r\n\r\n<strong>\u00a0 \u00a0(1) <\/strong><strong>E- Text:<\/strong>\r\n\r\n&nbsp;\r\n\r\n<strong>Linear Regression Analysis<\/strong>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">In any given system some of the variables do not vary on their own. Their behavior depends on the outcome of some the other variables. For example, values of the agricultural production of an area will depend on its average annual rainfall along with some other factors also, whereas the values of the agricultural production will not affect the values of the average annual rainfall.Understanding such cause and effect relationshipin a system has been the basic concern of a scientific enquiry. The understanding of such cause and effect relationship of a phenomenon helps us in understanding its nature and will help us in controllingit and also in making predictions ofits future.<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>Empirical Relationship<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The study of any relationships will have two partstheoretical as well as empirical. Theoretical forms of relationship are based on the logical form of the inter-relationships amongdifferent variables. Empirical relationships are based on the co-variations of the values of these variables. These co-variations in the values of the variables can be verified from the real world data by plotting them on a scatter plot.Basic objective of the observation of an empirical relationship is to validate or invalidate an existing form of a theoretical relationship. Any empirical analysis, therefore, requires a theoretical framework and an empirical analysis without a theory will be a futile exercise.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The variables which affect the outcome of the values of any variable are known as ;\u201cindependent variables\u201d and the variable which is being affected by other variables is known as ; \u201cdependent variable\u201d. In the above example; agricultural production will be the dependent variable and the average annual rainfall of the area over different years will be the independent variable.It is to be noted that in some other exercise of meteorology,for example, the average annual rainfall may become dependent variable and atmospheric pressure, temperature etc. may become independent variables, and so on. Statistical methods of the study of relationship through correlation or\/and regression are based on empirical methods only and are not substitute of the theory behind these relationships.<\/p>\r\n&nbsp;\r\n\r\n<\/div>\r\n<div>\r\n<p style=\"text-align: justify\">\u00a0 \u00a0The relationship between the values of two different variables can be easily shown In a scatter plotby plotting the values of the dependent variable on Y- axis and the values of the independent variables on X-axis of a graph paper. In a scatter plot we observe the degree and direction of covariations between the values two variablesand measure its strength quantitatively through either Karl Pearson\u2019s or Spearman\u2019s coefficients of correlations.<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>Bivariate Linear Regression Analysis<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Knowledge of the degree and direction of relationship between dependent and independent variables alone is not sufficient in a cause and effect analysis unless it gives a mathematical form of the relationship between the two variables. Through the mathematical form we can evaluate the effect of different independent variables(also known as determinants) on the dependent variable with the help of their coefficients and can also make future projections of the dependent variable for some projected future values of independent variables by working out the estimated future values of the dependent variable etc.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">A visible relationship between the values of the dependent and independent variable can be converted into a nearest form of a line or a curve of prescribed mathematical form of relationshipswhich is the basis of a regression analysis. The simplest of all these mathematical forms is the equation of a straight line which is Y = a + bX. In a straight line all the points will fall on a line with either an upward or a downward slope indicated by the value of \u201cb\u201d also known as regression coefficient. The line when extended sufficiently will also cut Y-axis and X-axis. The value at which the line cuts the Y-axis is \u201ca\u201d, known as intercept ( value of Y-when X=0 ).<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The magnitude of the value of \u201cb\u201d the regression coefficient will indicate the change in the values of the variable Y for a unit change in the values of the variable X or simply the rate of change in Y with respect to X. If the sign of \u201cb\u201d is negative it shows the decline in Y with respect to X and a positive value of regression coefficient \u201cb\u201d will indicate a positive change in the values of Y with the change in the values of the variable X.From a given set of bivariate data for several observations, linear regression analysis provides us the basis to find out the mathematical form of relationship ( Y = a + bX ) closest to the related scatter plot with the help of the \u2018Principle of Least Square. Since we use the equation of a line to approximate the form of relationship between the two variables , it will be known as linear regression analysis . If we have only two variables to analyze, i.e. one dependent and one independent variable, the regression is known as a bivariate linear regression. However, if the number of independent variables aretwo or more it is known as multiple linear regression analysis. Explanation of regression analysis will be greatly facilitated if we start from bivariate linear regression analysis and then extend these concepts to multiple linear regression analysis.<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>Principle of Least Square Bivariate Case<\/strong>\r\n\r\n<\/div>\r\n<p style=\"text-align: justify\">From a scatter plot we identify a straight line by fixing the values of <strong>a<\/strong> and <strong>b<\/strong>. However, most of the data from social sciences, will not exactly fall on a straight line. On a scatter plot of such data the points may cluster around a straight line and give some deviationsalso (error) on either side of the line represented by <strong>\u03b5<\/strong>. If the values of a and b are identified and we substitute an actual ly given value of X of the independent variable the equation will give us an estimated value of the corresponding value of the dependent variable Y\u0302. The difference between the actual valve of Y and the estimated value i.e. (Y - Y\u0302 ) = <strong>\u03b5<\/strong>is the error term. The total error in any exercise will be the sum of all such deviations. Summation of these error,however, will have the problem that negative error will cancel out the positive error which hide the efficiency of the estimates. It is, therefore, suggested to take sum of the squares of the errors i.e. \u2211 (Y - Y\u0302 )2 .Theoretically, in the absence of any guideline, we can select any number of lines and each one of them will give a sum of the squares of the errors. Statistical theory has developed the ,\u201dprinciple of least square\u201d which suggest that out of all such lines choose the one which gives the least value of the above sum of squares. Thus fitting a regression line from a given set of correlated bivariate data amounts to identifying the corresponding values of the regression coefficient \u201cb\u201d and the intercept \u201ca\u201d of the line of best fit. Using the mathematics of calculus it also suggest that a straight line whose values of \u201cb\u201d and are calculated in the following manner will give this least sum of squares.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-278\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-121.png\" alt=\"\" width=\"372\" height=\"66\" \/>\r\n\r\n&nbsp;\r\n<div>\r\n\r\n\u00a0 \u00a0 \u00a0the intercept will be:\r\n\r\n&nbsp;\r\n\r\na= y \u2013 b x .\r\n\r\n&nbsp;\r\n\r\nA regression line with the above values of the regression coefficient; \u201cb\u201d and\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">the intercept; \u201ca\u201d will be known as the least square regression line.This principle is known as <strong>Principle of Least Square <\/strong>and the line will be known as<strong> Regression Line of Y on X. <\/strong>Similarly if we choose X as dependent variable (and Y as independent variable) and minimize the sum of the squares parallel to X-axis, the line will be known asregression line of X on Y. The word regression here has been used to negate progress on its own. Here it implies that the movement of Y will not progress on its own, it will be regressed by the movement of X.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The values of the regression coefficient \u201cb \u201cand the intercept \u201ca\u201d will indicate the average position rather than the actual as there are positive and negative deviations with each point.<\/p>\r\n&nbsp;\r\n\r\n<strong>Example<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Consider the following values of productivity of wheat Y (in 00 kgs.\/hectare) and average annual rainfall X (in cms.) of ten areas of a region. To test the hypothesis that wheat productivity in the region depends on the average annual rainfall of the area and find out the rate of change in wheat production for a unit change in average annual rainfall.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong>Solution<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">To show the relationship between wheat productivity Y and average annual rainfall graphically in the data given in the table, we have to prepare a scatter plot on a graph paper as shown below.<\/p>\r\n&nbsp;\r\n\r\n<\/div>\r\n<img class=\"aligncenter size-full wp-image-279\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-122.png\" alt=\"\" width=\"367\" height=\"57\" \/>\r\n\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-280\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-123.png\" alt=\"\" width=\"487\" height=\"466\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">It is clear from the above graph that there exist a positive relationship between the two variables. To fit a regression line, we use the principle of least square and find out the values of the regression coefficient \u201cb\u201d and the intercept \u201ca\u201d using the above formulas which require the following table for calculations.<\/p>\r\n&nbsp;\r\n\r\n<strong>Table 1:Computations required forregression analysis.<\/strong>\r\n\r\n&nbsp;\r\n<table class=\"aligncenter\" border=\"1\">\r\n<tbody>\r\n<tr>\r\n<td><\/td>\r\n<td><strong>X<\/strong><\/td>\r\n<td><strong>Y<\/strong><\/td>\r\n<td><strong>X<\/strong><strong>2<\/strong><\/td>\r\n<td><strong>Y<\/strong><strong>2<\/strong><\/td>\r\n<td><strong>XY<\/strong><\/td>\r\n<\/tr>\r\n<tr>\r\n<td><\/td>\r\n<td>40<\/td>\r\n<td>5<\/td>\r\n<td>1600<\/td>\r\n<td>25<\/td>\r\n<td>200<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><\/td>\r\n<td>45<\/td>\r\n<td>13<\/td>\r\n<td>2025<\/td>\r\n<td>169<\/td>\r\n<td>585<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><\/td>\r\n<td>50<\/td>\r\n<td>11<\/td>\r\n<td>2500<\/td>\r\n<td>121<\/td>\r\n<td>550<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><\/td>\r\n<td>50<\/td>\r\n<td>15<\/td>\r\n<td>2500<\/td>\r\n<td>225<\/td>\r\n<td>750<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><\/td>\r\n<td>60<\/td>\r\n<td>20<\/td>\r\n<td>3600<\/td>\r\n<td>400<\/td>\r\n<td>1200<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><\/td>\r\n<td>60<\/td>\r\n<td>13<\/td>\r\n<td>3600<\/td>\r\n<td>169<\/td>\r\n<td>780<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><\/td>\r\n<td>65<\/td>\r\n<td>18<\/td>\r\n<td>4225<\/td>\r\n<td>324<\/td>\r\n<td>1170<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><\/td>\r\n<td>70<\/td>\r\n<td>20<\/td>\r\n<td>4900<\/td>\r\n<td>400<\/td>\r\n<td>1400<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><\/td>\r\n<td>71<\/td>\r\n<td>25<\/td>\r\n<td>5041<\/td>\r\n<td>625<\/td>\r\n<td>1775<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><\/td>\r\n<td>75<\/td>\r\n<td>25<\/td>\r\n<td>5625<\/td>\r\n<td>625<\/td>\r\n<td>1875<\/td>\r\n<\/tr>\r\n<tr>\r\n<td><strong>Total<\/strong><\/td>\r\n<td><strong>586<\/strong><\/td>\r\n<td><strong>165<\/strong><\/td>\r\n<td><strong>35616<\/strong><\/td>\r\n<td><strong>3083<\/strong><\/td>\r\n<td><strong>10285<\/strong><\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-281\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-124.png\" alt=\"\" width=\"390\" height=\"128\" \/>\r\n<p style=\"text-align: justify\">The value of the regression coefficient \u201cb\u201d is found to be .482(00) kg.\/ht and the intercept is \u2013 11.75 (00) kg\/ht. To elaborate it further, the meaning of b = 0.482(00) is that there exists a positive relationship between wheat productivity and average annual rainfall in the area and data given above shows the tendency of an average rise in wheat production equal to 0.482x100= 48.2 kgs. per hectaredue to every unit centimeter rise in the average annual rainfall in the area. Interpretation of a positive value of the intercept relates to the value of the dependent variable Y when the value of the independent variable is at its lowest level i.e. X= 0. However a negative value of intercept a at X= 0 directly does not mean anything. At the best we can see at what value of X the value of Y will be zero from the regression line in the following way:<\/p>\r\n&nbsp;\r\n\r\n0 = - 11.75 + 0.482 X\r\n\r\n&nbsp;\r\n<div>\r\n\r\n\u00a0 \u00a0 \u00a0X = 24.34 Cm.\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The above equation also suggest that on the average production of wheat require the thresh hold value of average annual rainfall at least 24.34 cm. Another point suggested by the results of the analysis is that on the average the wheat production is likely to stat only after the average annual rainfall of 24.34 cm. has already taken place.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The equation of the fitted regression line is found to be Y = - 11.781 + 0.482 X. Once the algebraic equation of the least square regression line is fitted we can estimate the values of wheat productivity (Y) , for every given value of average annual rainfall (X). The difference between given value of Y and the estimated value of Y\u0302 is known as the residual or the error<strong>\u03b5<\/strong>. For example in the above case for first given value of average annual rainfall X=40 the given value of wheat production is 5(00) kg\/ht. Whereas the estimated value as given by line is Y = - 11.781 + 0.4826*40 = -11.781 +19.304 = 7.523(00) kgs. The difference is 5(00) \u2013 7.523(00) = - 2.523(00). In the second case (ignoring 00 and unit for the time being) for given value of X = 45 the given value of Y is 13 and the estimated value is 8.717 giving the difference in the second case as 13 \u2013 8.717 = 4.283. Likewise the residuals or errors for all other observations can also be calculated.<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">In any regression analysis computations of the value of the regression coefficient \u201cb\u201d and intercepts alone are not sufficient. We have to evaluate its magnitude also by testing the null hypothesis that : \u201cIs the value of b very close to zero or not?\u201d. We proceed to interpret the value of the two parameters only when the null hypothesis is rejected i.e. the value of b is not very close to zero. In such a case we call the value of \u201ca\u201d or \u201cb\u201d being statistically significant. In other words we test the hypothesis that; can its actual value (also known as population value ) may be considered as zero? For the purpose of the test of significance of the regression coefficient \u201cb\u201d following \u201ct\u201d \u2013test is carried out under the following assumptions:<\/p>\r\n\r\n<ol>\r\n \t<li style=\"text-align: justify\">The variable <strong>\u03b5<\/strong> is a random variable distributed normally.<\/li>\r\n \t<li style=\"text-align: justify\">Mean value of <strong>\u03b5<\/strong>is zero.<\/li>\r\n \t<li style=\"text-align: justify\">Variance of <strong>\u03b5<\/strong>is constant for all value of X ( condition of hetroscedasicity)<\/li>\r\n \t<li style=\"text-align: justify\">If there are more independent variables all are independent to each other i.e. there is no multi co-linearity among the independent variables.<\/li>\r\n<\/ol>\r\n<p style=\"text-align: justify\">If the above assumptions are met the computed value of <strong>b<\/strong>will give the \u201ct\u201d given below which willfollow <strong>\u201ct\u201d<\/strong> distribution with (n- 2) degrees of freedom.<\/p>\r\n<strong>t = b \/ S. E. (b), <\/strong>where\r\n\r\n<\/div>\r\n<img class=\"aligncenter size-full wp-image-282\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-125.png\" alt=\"\" width=\"660\" height=\"286\" \/>\r\n<div>\r\n<p style=\"text-align: justify\">As estimated values of y are values found on the line there variations will always be less than the actual values of y. Objective of any regression line is always to get the estimated values of y giving variations as close to actual values of y as is possible. The ratio of Explained Some of square to total variations in y is therefore an indicator of the quality of any regression model and is known as coefficient of determination and denoted by R2 given by:<\/p>\r\n&nbsp;\r\n\r\n<strong>R<\/strong><strong>2<\/strong><strong> = Explained Sum of Squares\/ Total Sum of Squares.<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The value of R2 will vary between zero and unity. If a value is found to be say 0 .75 will mean that of the total variations in the dependent variable y 75 percent are being explained by the independent variable X chosen here<\/p>\r\n&nbsp;\r\n\r\nAgain to test the statistical significance of R2 we have F- ratio test as :\r\n\r\n&nbsp;\r\n\r\n<strong>F = R<\/strong><strong>2<\/strong><strong> (n-k)\/(1-R<\/strong><strong>2<\/strong><strong>)(k- 1)<\/strong>\u00a0 \u00a0with ( k-1, n-k), degrees of freedom\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Use of the results of a regression analysis for any comparative research will be valid only when the estimated parameters are found to be statistically significant. Regression\u00a0<span style=\"text-align: initial;font-size: 1em\">analysis, therefore, also requires calculation of standard error of b to carry out the\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">statistical test of significance to assure that the regression coefficient \u201cb\u201d is not very close to\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">zero or in other words is not statistically insignificant and the explanatory power of the\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">regression model is also statistically significant. Such calculations required the following:<\/span><\/p>\r\n\r\n<\/div>\r\n<img class=\"aligncenter size-full wp-image-283\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-126.png\" alt=\"\" width=\"510\" height=\"244\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The tabulated value of \u201ct\u201d for 8 degrees of freedom is 2.31 at 5% level of significance and 3.36 at 1% level of significance respectively. The calculated value (6.025) is found to be significant even at 1% level of significance i.e. the regression coefficient is found to be statistically significant.<\/p>\r\n&nbsp;\r\n\r\nR2 = 296.12\/360.5= 0.821 and\r\n\r\n&nbsp;\r\n\r\nF = F = R2 (n-k)\/(1-R2)(k- 1) = 0.821(10 -2)\/ (1- 0.821)(2-1)=36.69\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">F value for d.f.= 1, 8 (2-1, 10-2)given in the table are 16.86 for 1% and 5.32 for 5% levels of significance respectively. Our calculated value is found to be significant at even at 1% level of significance. Thus we can say R2 is statistically significant from being zero. In other words we conclude that the independent variable is explaining the dependent variable in a substantial way ( not in any random way).<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">After carrying out the tests of significance and R2, one can verify the correspondence between actual values of the dependent variable and the valueestimated from the regression equation. The estimated values are obtained by putting the given values of X in the regression equation, as has been shown earlier while explaining principle of least square. It is interesting to observe that for the given values of Y the estimated values show good correspondence and thereby strengthening the confidence in the regression model. Since estimated values are found to be close to the actual values, the model can be used to make forecasting for the future values of Y for any future value of the independent variable X. the regression model can also be used for interpolating any missing values between the given values of the dependent variable Y.<\/p>\r\n&nbsp;\r\n<table class=\"aligncenter\" border=\"1\">\r\n<tbody>\r\n<tr>\r\n<td><strong>X<\/strong><\/td>\r\n<td><strong>Y<\/strong><\/td>\r\n<td>\u0302<strong>Y<\/strong><\/td>\r\n<td>\u0302<strong>(Y- \u0302Y)<\/strong><\/td>\r\n<\/tr>\r\n<tr>\r\n<td>40<\/td>\r\n<td>5<\/td>\r\n<td>7.523<\/td>\r\n<td>-2.523<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>45<\/td>\r\n<td>13<\/td>\r\n<td>9.936<\/td>\r\n<td>3.064<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>50<\/td>\r\n<td>11<\/td>\r\n<td>12.349<\/td>\r\n<td>-1.349<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>50<\/td>\r\n<td>15<\/td>\r\n<td>12.349<\/td>\r\n<td>2.651<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>60<\/td>\r\n<td>20<\/td>\r\n<td>17.175<\/td>\r\n<td>2.825<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>60<\/td>\r\n<td>13<\/td>\r\n<td>17.175<\/td>\r\n<td>-4.175<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>65<\/td>\r\n<td>18<\/td>\r\n<td>19.588<\/td>\r\n<td>-1.588<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>70<\/td>\r\n<td>20<\/td>\r\n<td>22.001<\/td>\r\n<td>-2.001<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>71<\/td>\r\n<td>25<\/td>\r\n<td>22.4836<\/td>\r\n<td>2.5164<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>75<\/td>\r\n<td>25<\/td>\r\n<td>24.414<\/td>\r\n<td>0.586<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n&nbsp;\r\n\r\n&nbsp;\r\n<table>\r\n<tbody>\r\n<tr>\r\n<td><strong>you can view video on Introduction to Bivariate Linear Regression analysis<\/strong><\/td>\r\n<td><a href=\"https:\/\/youtu.be\/EoAzZXSEchc\" target=\"_blank\" rel=\"noopener\"><img class=\"alignnone wp-image-120\" src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"\" width=\"36\" height=\"36\" \/><\/a><\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n&nbsp;\r\n\r\n&nbsp;","rendered":"<div><span style=\"float: right\"><a href=\"https:\/\/youtu.be\/EoAzZXSEchc\" target=\"_blank\" rel=\"noopener\"><img decoding=\"async\" src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"epgp books\" width=\"75px\" height=\"75px;\" \/><\/a><br \/>\n<\/span><\/div>\n<div><strong>\u00a0 \u00a0<\/strong><\/div>\n<div><\/div>\n<div><\/div>\n<div>\n<p><strong>\u00a0 \u00a0(1) <\/strong><strong>E- Text:<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><strong>Linear Regression Analysis<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In any given system some of the variables do not vary on their own. Their behavior depends on the outcome of some the other variables. For example, values of the agricultural production of an area will depend on its average annual rainfall along with some other factors also, whereas the values of the agricultural production will not affect the values of the average annual rainfall.Understanding such cause and effect relationshipin a system has been the basic concern of a scientific enquiry. The understanding of such cause and effect relationship of a phenomenon helps us in understanding its nature and will help us in controllingit and also in making predictions ofits future.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Empirical Relationship<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The study of any relationships will have two partstheoretical as well as empirical. Theoretical forms of relationship are based on the logical form of the inter-relationships amongdifferent variables. Empirical relationships are based on the co-variations of the values of these variables. These co-variations in the values of the variables can be verified from the real world data by plotting them on a scatter plot.Basic objective of the observation of an empirical relationship is to validate or invalidate an existing form of a theoretical relationship. Any empirical analysis, therefore, requires a theoretical framework and an empirical analysis without a theory will be a futile exercise.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The variables which affect the outcome of the values of any variable are known as ;\u201cindependent variables\u201d and the variable which is being affected by other variables is known as ; \u201cdependent variable\u201d. In the above example; agricultural production will be the dependent variable and the average annual rainfall of the area over different years will be the independent variable.It is to be noted that in some other exercise of meteorology,for example, the average annual rainfall may become dependent variable and atmospheric pressure, temperature etc. may become independent variables, and so on. Statistical methods of the study of relationship through correlation or\/and regression are based on empirical methods only and are not substitute of the theory behind these relationships.<\/p>\n<p>&nbsp;<\/p>\n<\/div>\n<div>\n<p style=\"text-align: justify\">\u00a0 \u00a0The relationship between the values of two different variables can be easily shown In a scatter plotby plotting the values of the dependent variable on Y- axis and the values of the independent variables on X-axis of a graph paper. In a scatter plot we observe the degree and direction of covariations between the values two variablesand measure its strength quantitatively through either Karl Pearson\u2019s or Spearman\u2019s coefficients of correlations.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Bivariate Linear Regression Analysis<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Knowledge of the degree and direction of relationship between dependent and independent variables alone is not sufficient in a cause and effect analysis unless it gives a mathematical form of the relationship between the two variables. Through the mathematical form we can evaluate the effect of different independent variables(also known as determinants) on the dependent variable with the help of their coefficients and can also make future projections of the dependent variable for some projected future values of independent variables by working out the estimated future values of the dependent variable etc.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">A visible relationship between the values of the dependent and independent variable can be converted into a nearest form of a line or a curve of prescribed mathematical form of relationshipswhich is the basis of a regression analysis. The simplest of all these mathematical forms is the equation of a straight line which is Y = a + bX. In a straight line all the points will fall on a line with either an upward or a downward slope indicated by the value of \u201cb\u201d also known as regression coefficient. The line when extended sufficiently will also cut Y-axis and X-axis. The value at which the line cuts the Y-axis is \u201ca\u201d, known as intercept ( value of Y-when X=0 ).<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The magnitude of the value of \u201cb\u201d the regression coefficient will indicate the change in the values of the variable Y for a unit change in the values of the variable X or simply the rate of change in Y with respect to X. If the sign of \u201cb\u201d is negative it shows the decline in Y with respect to X and a positive value of regression coefficient \u201cb\u201d will indicate a positive change in the values of Y with the change in the values of the variable X.From a given set of bivariate data for several observations, linear regression analysis provides us the basis to find out the mathematical form of relationship ( Y = a + bX ) closest to the related scatter plot with the help of the \u2018Principle of Least Square. Since we use the equation of a line to approximate the form of relationship between the two variables , it will be known as linear regression analysis . If we have only two variables to analyze, i.e. one dependent and one independent variable, the regression is known as a bivariate linear regression. However, if the number of independent variables aretwo or more it is known as multiple linear regression analysis. Explanation of regression analysis will be greatly facilitated if we start from bivariate linear regression analysis and then extend these concepts to multiple linear regression analysis.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Principle of Least Square Bivariate Case<\/strong><\/p>\n<\/div>\n<p style=\"text-align: justify\">From a scatter plot we identify a straight line by fixing the values of <strong>a<\/strong> and <strong>b<\/strong>. However, most of the data from social sciences, will not exactly fall on a straight line. On a scatter plot of such data the points may cluster around a straight line and give some deviationsalso (error) on either side of the line represented by <strong>\u03b5<\/strong>. If the values of a and b are identified and we substitute an actual ly given value of X of the independent variable the equation will give us an estimated value of the corresponding value of the dependent variable Y\u0302. The difference between the actual valve of Y and the estimated value i.e. (Y &#8211; Y\u0302 ) = <strong>\u03b5<\/strong>is the error term. The total error in any exercise will be the sum of all such deviations. Summation of these error,however, will have the problem that negative error will cancel out the positive error which hide the efficiency of the estimates. It is, therefore, suggested to take sum of the squares of the errors i.e. \u2211 (Y &#8211; Y\u0302 )2 .Theoretically, in the absence of any guideline, we can select any number of lines and each one of them will give a sum of the squares of the errors. Statistical theory has developed the ,\u201dprinciple of least square\u201d which suggest that out of all such lines choose the one which gives the least value of the above sum of squares. Thus fitting a regression line from a given set of correlated bivariate data amounts to identifying the corresponding values of the regression coefficient \u201cb\u201d and the intercept \u201ca\u201d of the line of best fit. Using the mathematics of calculus it also suggest that a straight line whose values of \u201cb\u201d and are calculated in the following manner will give this least sum of squares.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-278\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-121.png\" alt=\"\" width=\"372\" height=\"66\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-121.png 372w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-121-300x53.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-121-65x12.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-121-225x40.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-121-350x62.png 350w\" sizes=\"auto, (max-width: 372px) 100vw, 372px\" \/><\/p>\n<p>&nbsp;<\/p>\n<div>\n<p>\u00a0 \u00a0 \u00a0the intercept will be:<\/p>\n<p>&nbsp;<\/p>\n<p>a= y \u2013 b x .<\/p>\n<p>&nbsp;<\/p>\n<p>A regression line with the above values of the regression coefficient; \u201cb\u201d and<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">the intercept; \u201ca\u201d will be known as the least square regression line.This principle is known as <strong>Principle of Least Square <\/strong>and the line will be known as<strong> Regression Line of Y on X. <\/strong>Similarly if we choose X as dependent variable (and Y as independent variable) and minimize the sum of the squares parallel to X-axis, the line will be known asregression line of X on Y. The word regression here has been used to negate progress on its own. Here it implies that the movement of Y will not progress on its own, it will be regressed by the movement of X.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The values of the regression coefficient \u201cb \u201cand the intercept \u201ca\u201d will indicate the average position rather than the actual as there are positive and negative deviations with each point.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Example<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Consider the following values of productivity of wheat Y (in 00 kgs.\/hectare) and average annual rainfall X (in cms.) of ten areas of a region. To test the hypothesis that wheat productivity in the region depends on the average annual rainfall of the area and find out the rate of change in wheat production for a unit change in average annual rainfall.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong>Solution<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">To show the relationship between wheat productivity Y and average annual rainfall graphically in the data given in the table, we have to prepare a scatter plot on a graph paper as shown below.<\/p>\n<p>&nbsp;<\/p>\n<\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-279\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-122.png\" alt=\"\" width=\"367\" height=\"57\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-122.png 367w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-122-300x47.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-122-65x10.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-122-225x35.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-122-350x54.png 350w\" sizes=\"auto, (max-width: 367px) 100vw, 367px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-280\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-123.png\" alt=\"\" width=\"487\" height=\"466\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-123.png 487w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-123-300x287.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-123-65x62.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-123-225x215.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-123-350x335.png 350w\" sizes=\"auto, (max-width: 487px) 100vw, 487px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">It is clear from the above graph that there exist a positive relationship between the two variables. To fit a regression line, we use the principle of least square and find out the values of the regression coefficient \u201cb\u201d and the intercept \u201ca\u201d using the above formulas which require the following table for calculations.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Table 1:Computations required forregression analysis.<\/strong><\/p>\n<p>&nbsp;<\/p>\n<table class=\"aligncenter\">\n<tbody>\n<tr>\n<td><\/td>\n<td><strong>X<\/strong><\/td>\n<td><strong>Y<\/strong><\/td>\n<td><strong>X<\/strong><strong>2<\/strong><\/td>\n<td><strong>Y<\/strong><strong>2<\/strong><\/td>\n<td><strong>XY<\/strong><\/td>\n<\/tr>\n<tr>\n<td><\/td>\n<td>40<\/td>\n<td>5<\/td>\n<td>1600<\/td>\n<td>25<\/td>\n<td>200<\/td>\n<\/tr>\n<tr>\n<td><\/td>\n<td>45<\/td>\n<td>13<\/td>\n<td>2025<\/td>\n<td>169<\/td>\n<td>585<\/td>\n<\/tr>\n<tr>\n<td><\/td>\n<td>50<\/td>\n<td>11<\/td>\n<td>2500<\/td>\n<td>121<\/td>\n<td>550<\/td>\n<\/tr>\n<tr>\n<td><\/td>\n<td>50<\/td>\n<td>15<\/td>\n<td>2500<\/td>\n<td>225<\/td>\n<td>750<\/td>\n<\/tr>\n<tr>\n<td><\/td>\n<td>60<\/td>\n<td>20<\/td>\n<td>3600<\/td>\n<td>400<\/td>\n<td>1200<\/td>\n<\/tr>\n<tr>\n<td><\/td>\n<td>60<\/td>\n<td>13<\/td>\n<td>3600<\/td>\n<td>169<\/td>\n<td>780<\/td>\n<\/tr>\n<tr>\n<td><\/td>\n<td>65<\/td>\n<td>18<\/td>\n<td>4225<\/td>\n<td>324<\/td>\n<td>1170<\/td>\n<\/tr>\n<tr>\n<td><\/td>\n<td>70<\/td>\n<td>20<\/td>\n<td>4900<\/td>\n<td>400<\/td>\n<td>1400<\/td>\n<\/tr>\n<tr>\n<td><\/td>\n<td>71<\/td>\n<td>25<\/td>\n<td>5041<\/td>\n<td>625<\/td>\n<td>1775<\/td>\n<\/tr>\n<tr>\n<td><\/td>\n<td>75<\/td>\n<td>25<\/td>\n<td>5625<\/td>\n<td>625<\/td>\n<td>1875<\/td>\n<\/tr>\n<tr>\n<td><strong>Total<\/strong><\/td>\n<td><strong>586<\/strong><\/td>\n<td><strong>165<\/strong><\/td>\n<td><strong>35616<\/strong><\/td>\n<td><strong>3083<\/strong><\/td>\n<td><strong>10285<\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-281\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-124.png\" alt=\"\" width=\"390\" height=\"128\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-124.png 390w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-124-300x98.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-124-65x21.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-124-225x74.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-124-350x115.png 350w\" sizes=\"auto, (max-width: 390px) 100vw, 390px\" \/><\/p>\n<p style=\"text-align: justify\">The value of the regression coefficient \u201cb\u201d is found to be .482(00) kg.\/ht and the intercept is \u2013 11.75 (00) kg\/ht. To elaborate it further, the meaning of b = 0.482(00) is that there exists a positive relationship between wheat productivity and average annual rainfall in the area and data given above shows the tendency of an average rise in wheat production equal to 0.482&#215;100= 48.2 kgs. per hectaredue to every unit centimeter rise in the average annual rainfall in the area. Interpretation of a positive value of the intercept relates to the value of the dependent variable Y when the value of the independent variable is at its lowest level i.e. X= 0. However a negative value of intercept a at X= 0 directly does not mean anything. At the best we can see at what value of X the value of Y will be zero from the regression line in the following way:<\/p>\n<p>&nbsp;<\/p>\n<p>0 = &#8211; 11.75 + 0.482 X<\/p>\n<p>&nbsp;<\/p>\n<div>\n<p>\u00a0 \u00a0 \u00a0X = 24.34 Cm.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The above equation also suggest that on the average production of wheat require the thresh hold value of average annual rainfall at least 24.34 cm. Another point suggested by the results of the analysis is that on the average the wheat production is likely to stat only after the average annual rainfall of 24.34 cm. has already taken place.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The equation of the fitted regression line is found to be Y = &#8211; 11.781 + 0.482 X. Once the algebraic equation of the least square regression line is fitted we can estimate the values of wheat productivity (Y) , for every given value of average annual rainfall (X). The difference between given value of Y and the estimated value of Y\u0302 is known as the residual or the error<strong>\u03b5<\/strong>. For example in the above case for first given value of average annual rainfall X=40 the given value of wheat production is 5(00) kg\/ht. Whereas the estimated value as given by line is Y = &#8211; 11.781 + 0.4826*40 = -11.781 +19.304 = 7.523(00) kgs. The difference is 5(00) \u2013 7.523(00) = &#8211; 2.523(00). In the second case (ignoring 00 and unit for the time being) for given value of X = 45 the given value of Y is 13 and the estimated value is 8.717 giving the difference in the second case as 13 \u2013 8.717 = 4.283. Likewise the residuals or errors for all other observations can also be calculated.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In any regression analysis computations of the value of the regression coefficient \u201cb\u201d and intercepts alone are not sufficient. We have to evaluate its magnitude also by testing the null hypothesis that : \u201cIs the value of b very close to zero or not?\u201d. We proceed to interpret the value of the two parameters only when the null hypothesis is rejected i.e. the value of b is not very close to zero. In such a case we call the value of \u201ca\u201d or \u201cb\u201d being statistically significant. In other words we test the hypothesis that; can its actual value (also known as population value ) may be considered as zero? For the purpose of the test of significance of the regression coefficient \u201cb\u201d following \u201ct\u201d \u2013test is carried out under the following assumptions:<\/p>\n<ol>\n<li style=\"text-align: justify\">The variable <strong>\u03b5<\/strong> is a random variable distributed normally.<\/li>\n<li style=\"text-align: justify\">Mean value of <strong>\u03b5<\/strong>is zero.<\/li>\n<li style=\"text-align: justify\">Variance of <strong>\u03b5<\/strong>is constant for all value of X ( condition of hetroscedasicity)<\/li>\n<li style=\"text-align: justify\">If there are more independent variables all are independent to each other i.e. there is no multi co-linearity among the independent variables.<\/li>\n<\/ol>\n<p style=\"text-align: justify\">If the above assumptions are met the computed value of <strong>b<\/strong>will give the \u201ct\u201d given below which willfollow <strong>\u201ct\u201d<\/strong> distribution with (n- 2) degrees of freedom.<\/p>\n<p><strong>t = b \/ S. E. (b), <\/strong>where<\/p>\n<\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-282\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-125.png\" alt=\"\" width=\"660\" height=\"286\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-125.png 660w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-125-300x130.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-125-65x28.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-125-225x98.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-125-350x152.png 350w\" sizes=\"auto, (max-width: 660px) 100vw, 660px\" \/><\/p>\n<div>\n<p style=\"text-align: justify\">As estimated values of y are values found on the line there variations will always be less than the actual values of y. Objective of any regression line is always to get the estimated values of y giving variations as close to actual values of y as is possible. The ratio of Explained Some of square to total variations in y is therefore an indicator of the quality of any regression model and is known as coefficient of determination and denoted by R2 given by:<\/p>\n<p>&nbsp;<\/p>\n<p><strong>R<\/strong><strong>2<\/strong><strong> = Explained Sum of Squares\/ Total Sum of Squares.<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The value of R2 will vary between zero and unity. If a value is found to be say 0 .75 will mean that of the total variations in the dependent variable y 75 percent are being explained by the independent variable X chosen here<\/p>\n<p>&nbsp;<\/p>\n<p>Again to test the statistical significance of R2 we have F- ratio test as :<\/p>\n<p>&nbsp;<\/p>\n<p><strong>F = R<\/strong><strong>2<\/strong><strong> (n-k)\/(1-R<\/strong><strong>2<\/strong><strong>)(k- 1)<\/strong>\u00a0 \u00a0with ( k-1, n-k), degrees of freedom<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Use of the results of a regression analysis for any comparative research will be valid only when the estimated parameters are found to be statistically significant. Regression\u00a0<span style=\"text-align: initial;font-size: 1em\">analysis, therefore, also requires calculation of standard error of b to carry out the\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">statistical test of significance to assure that the regression coefficient \u201cb\u201d is not very close to\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">zero or in other words is not statistically insignificant and the explanatory power of the\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">regression model is also statistically significant. Such calculations required the following:<\/span><\/p>\n<\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-283\" src=\"http:\/\/geop01.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/225\/2019\/04\/1-126.png\" alt=\"\" width=\"510\" height=\"244\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-126.png 510w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-126-300x144.png 300w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-126-65x31.png 65w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-126-225x108.png 225w, https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-content\/uploads\/sites\/225\/2019\/04\/1-126-350x167.png 350w\" sizes=\"auto, (max-width: 510px) 100vw, 510px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The tabulated value of \u201ct\u201d for 8 degrees of freedom is 2.31 at 5% level of significance and 3.36 at 1% level of significance respectively. The calculated value (6.025) is found to be significant even at 1% level of significance i.e. the regression coefficient is found to be statistically significant.<\/p>\n<p>&nbsp;<\/p>\n<p>R2 = 296.12\/360.5= 0.821 and<\/p>\n<p>&nbsp;<\/p>\n<p>F = F = R2 (n-k)\/(1-R2)(k- 1) = 0.821(10 -2)\/ (1- 0.821)(2-1)=36.69<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">F value for d.f.= 1, 8 (2-1, 10-2)given in the table are 16.86 for 1% and 5.32 for 5% levels of significance respectively. Our calculated value is found to be significant at even at 1% level of significance. Thus we can say R2 is statistically significant from being zero. In other words we conclude that the independent variable is explaining the dependent variable in a substantial way ( not in any random way).<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">After carrying out the tests of significance and R2, one can verify the correspondence between actual values of the dependent variable and the valueestimated from the regression equation. The estimated values are obtained by putting the given values of X in the regression equation, as has been shown earlier while explaining principle of least square. It is interesting to observe that for the given values of Y the estimated values show good correspondence and thereby strengthening the confidence in the regression model. Since estimated values are found to be close to the actual values, the model can be used to make forecasting for the future values of Y for any future value of the independent variable X. the regression model can also be used for interpolating any missing values between the given values of the dependent variable Y.<\/p>\n<p>&nbsp;<\/p>\n<table class=\"aligncenter\">\n<tbody>\n<tr>\n<td><strong>X<\/strong><\/td>\n<td><strong>Y<\/strong><\/td>\n<td>\u0302<strong>Y<\/strong><\/td>\n<td>\u0302<strong>(Y- \u0302Y)<\/strong><\/td>\n<\/tr>\n<tr>\n<td>40<\/td>\n<td>5<\/td>\n<td>7.523<\/td>\n<td>-2.523<\/td>\n<\/tr>\n<tr>\n<td>45<\/td>\n<td>13<\/td>\n<td>9.936<\/td>\n<td>3.064<\/td>\n<\/tr>\n<tr>\n<td>50<\/td>\n<td>11<\/td>\n<td>12.349<\/td>\n<td>-1.349<\/td>\n<\/tr>\n<tr>\n<td>50<\/td>\n<td>15<\/td>\n<td>12.349<\/td>\n<td>2.651<\/td>\n<\/tr>\n<tr>\n<td>60<\/td>\n<td>20<\/td>\n<td>17.175<\/td>\n<td>2.825<\/td>\n<\/tr>\n<tr>\n<td>60<\/td>\n<td>13<\/td>\n<td>17.175<\/td>\n<td>-4.175<\/td>\n<\/tr>\n<tr>\n<td>65<\/td>\n<td>18<\/td>\n<td>19.588<\/td>\n<td>-1.588<\/td>\n<\/tr>\n<tr>\n<td>70<\/td>\n<td>20<\/td>\n<td>22.001<\/td>\n<td>-2.001<\/td>\n<\/tr>\n<tr>\n<td>71<\/td>\n<td>25<\/td>\n<td>22.4836<\/td>\n<td>2.5164<\/td>\n<\/tr>\n<tr>\n<td>75<\/td>\n<td>25<\/td>\n<td>24.414<\/td>\n<td>0.586<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<table>\n<tbody>\n<tr>\n<td><strong>you can view video on Introduction to Bivariate Linear Regression analysis<\/strong><\/td>\n<td><a href=\"https:\/\/youtu.be\/EoAzZXSEchc\" target=\"_blank\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-120\" src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"\" width=\"36\" height=\"36\" \/><\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n","protected":false},"author":12,"menu_order":17,"template":"","meta":{"pb_show_title":"on","pb_short_title":"","pb_subtitle":"","pb_authors":["prof-aslam-mahmood"],"pb_section_license":""},"chapter-type":[],"contributor":[58],"license":[],"class_list":["post-277","chapter","type-chapter","status-publish","hentry","contributor-prof-aslam-mahmood"],"part":3,"_links":{"self":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapters\/277","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapters"}],"about":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/wp\/v2\/types\/chapter"}],"author":[{"embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/wp\/v2\/users\/12"}],"version-history":[{"count":6,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapters\/277\/revisions"}],"predecessor-version":[{"id":289,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapters\/277\/revisions\/289"}],"part":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/parts\/3"}],"metadata":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapters\/277\/metadata\/"}],"wp:attachment":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/wp\/v2\/media?parent=277"}],"wp:term":[{"taxonomy":"chapter-type","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/pressbooks\/v2\/chapter-type?post=277"},{"taxonomy":"contributor","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/wp\/v2\/contributor?post=277"},{"taxonomy":"license","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/geop01\/wp-json\/wp\/v2\/license?post=277"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}