{"id":351,"date":"2018-10-31T10:09:32","date_gmt":"2018-10-31T10:09:32","guid":{"rendered":"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/?post_type=chapter&#038;p=351"},"modified":"2018-10-31T10:45:38","modified_gmt":"2018-10-31T10:45:38","slug":"linear-regression-simple-linear-regression-model-with-least-square-method","status":"publish","type":"chapter","link":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/chapter\/linear-regression-simple-linear-regression-model-with-least-square-method\/","title":{"rendered":"Linear Regression: Simple Linear Regression Model with Least Square Method"},"content":{"raw":"<div>\r\n\r\n&nbsp;\r\n\r\n<strong>Learning Objectives<\/strong>\r\n\r\n&nbsp;\r\n\r\nAfter this module the students will be able to:\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">1\u00a0 Clearly define the meaning of simple linear regression.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">2\u00a0 Differentiate between the correlation and regression also state the advantages of simple linear regression.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">3 Understand types of regression model and the assumptions<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">4 Establish the simple linear regression equation by least square method and determine the values of regression coefficients.<\/p>\r\n&nbsp;\r\n\r\n5 Clearly define the meaning of simple linear regression.\r\n\r\n&nbsp;\r\n\r\n6 Differentiate between the correlation and regression also state the advantages of simple linear regression.\r\n\r\n&nbsp;\r\n\r\n7\u00a0 Understand types of regression model and the assumptions\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">8 Establish the simple linear regression equation by least square method and determine the values of regression coefficients.<\/p>\r\n&nbsp;\r\n\r\n<strong>1. Introduction<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Correlation coefficient and covariance define a linear relationship between the variables but they are unable to state anything about the casual relationship or in other words with correlation coefficient and covariance we cannot say which variable is a cause and which one is an effect. Through the regression analysis we try to accomplish the causal relationship between the variables. The term regression basically means \u2018moving backwards\u2019.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Hence with the help of linear regression analysis we establish the causal relationship between the linearly related variables. In causal relationship there is a <strong><em>response variable<\/em><\/strong> which is influenced by other variable called <strong><em>explanatory variables<\/em><\/strong>. Some scholars have also given different names to both response variable and explanatory variable, like other names of response variable are dependent variable, predicted variable and explained variable etc. while explanatory variables are also called as independent variable, predictor variable and control variable etc.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">In regression model we try establish the complete cause and effect relationship between the variables but in most of the cases response variable does not depend on only one explanatory variable there could be more than one variable that could have either direct or indirect influence of response variable. To understand this better let us take an example of a fast moving consumer goods (FMCG) manufacturing company that wants to do research on the buying preferences of people for its new product that is going to be launched. For that they collect data of the housewives as they think that they are the decision maker for all the products that are used in kitchen but in real situation there may be the choice of children in the family that could affect the product purchase not only children there may be many other factors that should be considered like husbands preference, family income or the effect of opinion leaders etc. Hence there could be many explanatory variables that can influence the response variable.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">2. Difference between Correlation and Regression<\/strong><\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">In Statistics correlation and regression both are used to define the relationship between the variables. Correlation tells the degree of relationship between the variables where regression goes one step further and tells the cause and effect relationship between the variables. In other words Correlation measures the degree to which the two variables are related whereas regression is a method of describing the relationship between two variables. Below are the basic differences between correlation and regression-<\/p>\r\n&nbsp;\r\n<ol>\r\n \t<li style=\"text-align: justify\">A statistical measure that establishes the relationship between two variables is called correlation and regression establishes the value of one variable called dependent variable for a given value of another variable called independent variable.<\/li>\r\n \t<li style=\"text-align: justify\">Correlation represents the linear relationship between the two variables whereas regression fits the best line and estimates one variable on the basis of other variable.<\/li>\r\n \t<li style=\"text-align: justify\">Correlation does not state anything about dependent and independent variables and it is symmetrical in nature for example if x and y are the two variables there is no difference in the correlation of x and y and y and x in contrary regression relationship between x and y is different from y and x.<\/li>\r\n \t<li style=\"text-align: justify\">Correlation indicates the relationship between the two variables whereas regression denotes the impact of a unit change in independent variable on the dependent variable.<\/li>\r\n \t<li style=\"text-align: justify\">Correlation finds a numerical value that expresses relationship between two variables unlike regression that predicts the future value of dependent variable for given value of independent variable.<\/li>\r\n<\/ol>\r\n&nbsp;\r\n\r\n<strong>3.\u00a0 <\/strong><strong>Advantages of Regression Analysis<\/strong>\r\n\r\n&nbsp;\r\n\r\nFollowing are the advantages of the regression analysis-\r\n\r\n&nbsp;\r\n<ol>\r\n \t<li style=\"text-align: justify\">Establishes the relationship between variables<strong>-<\/strong> Regression analysis establishes relationship between the response (dependent) variable and the explanatory (independent) variable.<\/li>\r\n \t<li style=\"text-align: justify\">Determines the error- Regression analysis measure the standard error of estimates to measure the variability, as in regression line we establish a relationship line which fits all the values of x and y or in other words all the value of x and y should fall on the regression line and standard error estimates is equal to zero but it hardly happens. when all the variables either fall on the line or are very close to the line is called a good regression relationship.<\/li>\r\n \t<li style=\"text-align: justify\">Suits in case of large sample size- In case of large sample size (n \u2265 30) then interval estimation for predicting the value of a dependent variable based on standard error of estimate is considered to be acceptable by changing the values of either x or y. the magnitude of r2 remains the same regardless of the values of the two variables.<\/li>\r\n \t<li style=\"text-align: justify\">Predicting the future- In regression analysis we predict the future value of response variable based on the given value of explanatory variable.<\/li>\r\n<\/ol>\r\n&nbsp;\r\n\r\n<strong>4.\u00a0 <\/strong><strong>Types of Regression Models<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">There are two methods of studying the regression model, one is simple <strong>regression model<\/strong> where response variable completely depend upon the only one explanatory variable and other is <strong>multiple regression<\/strong> <strong>model <\/strong>where response variable is influenced by more than one variable.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">In those cases where the response variable completely depends upon the one explanatory variable, this relationship between the variables is called <\/span><strong style=\"text-align: initial;font-size: 1em\"><em>deterministic<\/em>.<\/strong><span style=\"text-align: initial;font-size: 1em\"> For example-<\/span><span style=\"text-align: initial;font-size: 1em\">y=1.5x<\/span><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Here the response variable y completely depends upon one explanatory variable x and no error is allowed while predicting the values of y. This is often seen in physical science concepts.<\/span><\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">But usually, the relationship between explanatory and response variable is inexact i.e. <strong><em>stochastic<\/em><\/strong>. It is due to the omission of relevant factors which are sometimes immeasurable, that influence the response variable.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">In Simple Regression model we judge the variability of response variable that is solely dependent upon only one explanatory variable. As the fundamental assumption in case of simple regression model is the expected value of y lies on a straight line, let us assume the response variable is denoted by y and explanatory variable is denoted by xi, then-y= \u03b20 +\u03b21xi<\/p>\r\n&nbsp;\r\n\r\nWhere \u03b20 and \u03b21 are unknown intercepts and slope parameters respectively.\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">In simple linear regression model \u03b20 +\u03b21xi is the deterministic component of the regression model, which tells the expected value of y for a given value of x.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">If the slop parameter \u03b21 is positive (\u03b21&gt;0) the relationship between x and y is positive, if the slop parameter \u03b21 is negative (\u03b21&lt;0) the relationship between x and y is negative and if the slop parameter \u03b21=0 there is no relationship between x and y.<\/p>\r\n&nbsp;\r\n\r\nGraphically we can represent positive, negative and no relationship of linear regression model as below-\r\n\r\n&nbsp;\r\n\r\n<img class=\"aligncenter size-full wp-image-354\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-176.png\" alt=\"\" width=\"674\" height=\"198\" \/>\r\n<p style=\"text-align: justify\">As we have discussed before actual value of response variable may defer from expected value hence we add \u03b5 (epsilon) as random error in deterministic component. Hence the sample regression model is defined as-<\/p>\r\n&nbsp;\r\n\r\ny= \u03b20 + \u03b21xi + \u03b5\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">where y and x are dependent variable and independent variable respectively and \u03b5 is random error.<\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n\r\n<strong>6.\u00a0 <\/strong><strong>Assumptions for a Simple Linear Regression model<\/strong>\r\n\r\n<strong>\u00a0<\/strong>\r\n<ol>\r\n \t<li style=\"text-align: justify\">There should be a linear relationship between two variables x and y whereas x is called dependent variable and y is independent variable. This relationship can be described by linear regression equation-<\/li>\r\n<\/ol>\r\n<img class=\"aligncenter size-full wp-image-355\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-177.png\" alt=\"\" width=\"106\" height=\"38\" \/>\r\n<p style=\"text-align: justify\">Where \u03b5 represents the difference between the expected value and actual value of response variable y for a given value of explanatory variable x.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">2. The set of expected values of response variables y for a given value of explanatory variable x are normally distributed. The mean of these normally distributed values fall on the line of regression.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">3. The dependent variable y is a continuous random variable, whereas values of the independent variables x are fixed and not random.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">4. The sampling error associated with the expected value of the response variable is assumed to be an independent random variable distributed normally with constant standard deviation. The amount of error in the value of response variable maybe different in successive observation.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">5. The standard deviation and variance of expected values of the response variable about the regression line are constant for all the values of the explanatory variable.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">6. Regression cannot have the symmetrical value of variables means the response variable y and explanatory variable x cannot be interchanged for a same regression line equation.<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>7.\u00a0 <\/strong><strong>Estimation: The Least Square Method (LS Method)<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">This method is also known as ordinary least square method (OLS method).The OLS method identifies that line which fits best for the given data. This is called the \u2018line of best fit\u2019 and is determined by identifying the line out of all of the probable lines which results in the least difference between the observed data points and the line.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Fig. -1 indicates that whenever a straight line is drawn passing through to a data, there will be some variations among the line values and the real observed values. Here, one is, fascinated about the vertical differences between this line and the real data. This line is used to make predictions about the values of Y (Dependent Variable) from different values of the X (Independent Variable). In context of regression, these differences are known as <em>residuals<\/em> and not as <em>deviations<\/em> (but actually both of them are same).<\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n<img class=\"aligncenter size-full wp-image-356\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-178.png\" alt=\"\" width=\"365\" height=\"216\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Fig-1 represents a scatter plot of any data where a line is projecting the general tendency. The vertical arrows are representing the gap (differences or residuals) between the actual data and the line.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">As with the mean, values of the variables fall both above and below the line resulting in positive as well as negative differences. Thus, in case, these positive and negative differences are added, they will annul each other out .To overcome this challenge, the differences are squared before summing them. This squared differences offer a estimate about the \u2018wellness\u2019 with which any particular line fits the data,i.e.in case the square of differences are huge the line is not a representative of the data but in case the value of differences is small, that line is assumed to be a representative.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Thus to find the \u2018line of best fit\u2019 we find the sum of SS (squared differences) of <strong><em>all the likely lines<\/em><\/strong> for the given values data and then compare them. The line with the least value of SS represents the required line, i.e. line of best fit. In reality, this tedious process needs not to be followed as this can be attained by using the method of OLS. It does so with the help of mathematical method used for finding maxima and minima. This procedure is used to find the line that minimizes the sum of squared differences. This line of best fit is a regression line.<\/p>\r\n&nbsp;\r\n\r\n<strong>8. Mathematical Explanation<\/strong>\r\n\r\n&nbsp;\r\n\r\nSuppose a sample of n pairs of observation (x1,y1), (x2,y2).........,(xn,yn) is taken from a population to\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">which we want to study to estimate the values of regression coefficient \u03b20 and \u03b21. The estimate values of \u03b20 and \u03b21 should result in a straight line where most pairs of observations fall very close to it. Such a straight line is referred to as \u2018best fitted\u2019 (least squares or estimates) regression line.<\/p>\r\n&nbsp;\r\n\r\nRewriting equation as follows-\r\n\r\n&nbsp;\r\n\r\nyi= \u03b20 + \u03b21xi + <em>e<\/em>i\r\n\r\nor\u00a0<em>e<\/em><em>i<\/em><em>\u00a0 <\/em>= yi \u2013 (\u03b20 + \u03b21xi)\r\n\r\nTo minimize\r\n\r\n<em>e<\/em><em>i<\/em><em> = <\/em>L =\u2211\u00a0 =1\u00a0\u00a0 2 =\u00a0 \u2211\u00a0 =1{yi \u2013 (\u03b20 + \u03b21xi)}2\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">let b0 and b1 be the least squares estimators of \u03b20 and \u03b21 respectively. After simplifying this equation, we get<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">\u2211\u00a0 =1 yi = n b0 + b1 \u2211\u00a0 =1<\/span>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">This equation is called the least squares normal equation. The values of least squares estimators b0 and b1 can be obtained by solving this equation. Hence the fitted or estimated regression line is given by<\/p>\r\n<img class=\"aligncenter size-full wp-image-357\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-179.png\" alt=\"\" width=\"183\" height=\"68\" \/>\r\n\r\n&nbsp;\r\n\r\nWhere-\u00a0ei = L =\u2211 =1 2 = \u2211 =1{yi \u2013 (\u03b20 + \u03b21xi)}2\r\n\r\nlet b0 and b1 be the least squares estimators of \u03b20 and \u03b21 respectively. After simplifying this equation, we get<img class=\"aligncenter size-full wp-image-358\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-180.png\" alt=\"\" width=\"106\" height=\"40\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">We use this equation to make the expectation of y for a given value of x since the expected value may be different from the actual value we take difference of both expected value and actual value and is generally represented by residual <em>e.<\/em><\/p>\r\n<img class=\"aligncenter size-full wp-image-359\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-181.png\" alt=\"\" width=\"96\" height=\"34\" \/>\r\n\r\n&nbsp;\r\n\r\n<strong>9. Regression Coefficients<\/strong>\r\n\r\n&nbsp;\r\n\r\nTo estimate values of population parameter \u03b20 and \u03b21,\u00a0 the estimated simple linear regression equation is\r\n\r\n\u0302 <strong>= b<\/strong><strong>0<\/strong> <strong>+ b<\/strong><strong>1<\/strong><strong>x<\/strong>\r\n\r\nwhere \u0302\u00a0 \u00a0 (y hat ) is estimated average value of response variable y for a given value of explanatory\r\n<p style=\"text-align: justify\">variable x. b0 or a is y-intercept that represents average value of \u00a0\u0305 and b1 or b is the slop of regression line that represents the expected change in the value of y for unit change in the value of x.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">a and b are also called as intercept and regression coefficient respectively. To determine the value of y at any given point of x we must calculate the values of a and b. After getting the values of a and b we can easily determine the value of response variable that is y at any given value of explanatory variable that is x.<\/p>\r\n&nbsp;\r\n\r\nThe regression coefficient \u2018b\u2019 is also denoted as-\r\n\r\n&nbsp;\r\n\r\nbyx that means regression coefficient of y on x and can be represented as y = a + bx\r\n\r\n&nbsp;\r\n\r\nbxy that means regression coefficient of x on y and can be represented as\r\n\r\n&nbsp;\r\n\r\nx\u00a0 = a + by\r\n\r\n&nbsp;\r\n\r\n<strong>Calculation of a and b<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\"><span style=\"font-size: 1em;text-align: initial\">For solving mathematical problems of regression analysis we need to calculate the intercept a and regression coefficient b.<\/span><\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">With a little algebra and differential calculus it can be shown that the following two equations, if solved simultaneously, will yield values of the parameters a and b such that the least squares requirement is fulfilled-<\/p>\r\n&nbsp;\r\n\r\n\u2211Y = Na + b\u2211X\r\n\r\n&nbsp;\r\n\r\n\u2211XY = a \u2211X + b\u2211X2\r\n\r\n&nbsp;\r\n\r\nThese equations are usually called the normal equations. N is the total pair of observed pairs of values.\r\n\r\n&nbsp;\r\n\r\n<strong>10. Properties of Regression coefficients<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">1.\u00a0 The correlation coefficient is the geometric mean of two regression coefficients byx and bxy i.e., <strong>r = \u221a (b<\/strong><strong>yx<\/strong><strong> X b<\/strong><strong>xy<\/strong><strong>)<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">2.\u00a0 If the one regression coefficient is greater than one, then other regression coefficient must be less than one because the value of correlation coefficient cannot be more than one.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">3.\u00a0 Both regression coefficient must have the same sign either +ve or \u2013ve. This property abolishes the case of opposite coefficient may be less than one.<\/p>\r\n&nbsp;\r\n\r\n4.\u00a0 Both the correlation coefficient and the two regression coefficient will have the same sign.\r\n\r\n&nbsp;\r\n\r\n5.\u00a0\u00a0 The arithmetic mean of regression coefficient byx and bxy is more than or equal to the correlation coefficient.\r\n\r\n&nbsp;\r\n\r\n<strong>11.\u00a0 <\/strong><strong>Limitations of Simple Linear Regression model<\/strong>\r\n\r\n<strong>\u00a0<\/strong>\r\n<ol>\r\n \t<li>With the help of Simple linear regression model so far, we\u2019ve only been able to examine the relationship between two variables.<\/li>\r\n \t<li>In many instances, we believe that more than one independent variable is correlated with the dependent variable.<\/li>\r\n \t<li>Multiple linear regressions provides is a tool that allows us to examine the relationship between 2 or more explanatory variable and a response variable.<\/li>\r\n<\/ol>\r\n&nbsp;\r\n\r\n<strong>12.Self-Check Questions:<\/strong>\r\n\r\n&nbsp;\r\n\r\n<strong>Illustration 1:<\/strong>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">Five Randomly selected students took a math aptitude test along with Statistics grade.<\/span>\r\n\r\n<span style=\"font-size: 1em;text-align: initial\">In the table below, the X column shows scores on the aptitude test. Similarly, the Y column shows statistics grades.<\/span>\r\n\r\n<\/div>\r\n<div>\r\n<table class=\"aligncenter\" style=\"width: 60%\" border=\"1\">\r\n<tbody>\r\n<tr>\r\n<td>Student<\/td>\r\n<td>X<\/td>\r\n<td>Y<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>1<\/td>\r\n<td>95<\/td>\r\n<td>85<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>2<\/td>\r\n<td>85<\/td>\r\n<td>95<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>3<\/td>\r\n<td>80<\/td>\r\n<td>70<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>4<\/td>\r\n<td>70<\/td>\r\n<td>65<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>5<\/td>\r\n<td>60<\/td>\r\n<td>70<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>SUM<\/td>\r\n<td>390<\/td>\r\n<td>385<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<ul>\r\n \t<li>\u00a0Find out the best fit regression line based on math aptitude test?<\/li>\r\n \t<li>If a student got 80 marks in the aptitude test find out his grade in statistics?<\/li>\r\n<\/ul>\r\n&nbsp;\r\n\r\nAnswer:\r\n\r\n&nbsp;\r\n\r\ni)\u00a0 The linear regression equation Yon X is \u0177 = b0 + b1x .\r\n\r\n&nbsp;\r\n\r\nStep 1: To find out best fit regression line we need to solve for b0 and b1. And the normal equations are \u2013\r\n\r\n&nbsp;\r\n\r\n\u2211Y = Na + b\u2211X\r\n\r\n&nbsp;\r\n\r\n\u2211XY = a \u2211X + b\u2211X2\r\n\r\n&nbsp;\r\n\r\nStep 2: Calculate the values of X2,Y2 and XY\r\n<table class=\"aligncenter\" style=\"width: 60%\" border=\"1\">\r\n<tbody>\r\n<tr>\r\n<td>Student<\/td>\r\n<td>X<\/td>\r\n<td>Y<\/td>\r\n<td>X2<\/td>\r\n<td>Y2<\/td>\r\n<td>XY<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>1<\/td>\r\n<td>95<\/td>\r\n<td>85<\/td>\r\n<td>9025<\/td>\r\n<td>7225<\/td>\r\n<td>8075<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>2<\/td>\r\n<td>85<\/td>\r\n<td>95<\/td>\r\n<td>7225<\/td>\r\n<td>9025<\/td>\r\n<td>8075<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>3<\/td>\r\n<td>80<\/td>\r\n<td>70<\/td>\r\n<td>6400<\/td>\r\n<td>4900<\/td>\r\n<td>5600<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>4<\/td>\r\n<td>70<\/td>\r\n<td>65<\/td>\r\n<td>4900<\/td>\r\n<td>4225<\/td>\r\n<td>4550<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>5<\/td>\r\n<td>60<\/td>\r\n<td>70<\/td>\r\n<td>3600<\/td>\r\n<td>4900<\/td>\r\n<td>4200<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>SUM<\/td>\r\n<td>390<\/td>\r\n<td>385<\/td>\r\n<td>31150<\/td>\r\n<td>30275<\/td>\r\n<td>30500<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/div>\r\n<span style=\"text-align: initial;font-size: 1em\">Step 3: Substitute the values in the normal equation-385 = 5a + b390\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">30500= 390b +31150b<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">Step 4:\u00a0 after solving these equations we get-<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">a=26.768<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">b= 0.644<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">Therefore, the regression equation is: \u0177 = 26.768 + 0.644x.<\/span>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">ii)\u00a0 Now once we have linear regression relationship equation between both the variables the aptitude test and the statistic\u2019s grade. We can easily predict the statistic\u2019s grade for any given value of aptitude test by just substituting the values in the found regression equation-<\/p>\r\n&nbsp;\r\n\r\n\u0177 = 26.768 + 0.644x = 26.768 + 0.644 * 80 = 26.768 + 51.52 = 78.288\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Caution: it not recommended to use the values of the independent variable that are outside the range of values used to create the equation. That is called <strong>extrapolation<\/strong>, and it can produce unreasonable estimates.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">In this example, the aptitude test scores used to create the regression equation ranged from 60 to 95. Therefore, only use values inside that range to estimate statistics grades. Using values outside that range (less than 60 or greater than 95) is problematic.<\/p>\r\n&nbsp;\r\n\r\n<strong>Illustration 2:<\/strong>\r\n\r\n&nbsp;\r\n\r\nCalculate the regression equation Y on X and X on Y from the following table -\r\n<table class=\"aligncenter\" style=\"border-collapse: collapse;width: 60%\" border=\"1\">\r\n<tbody>\r\n<tr>\r\n<td style=\"width: 16.6667%\">X<\/td>\r\n<td style=\"width: 16.6667%\">1<\/td>\r\n<td style=\"width: 16.6667%\">2<\/td>\r\n<td style=\"width: 16.6667%\">3<\/td>\r\n<td style=\"width: 16.6667%\">4<\/td>\r\n<td style=\"width: 16.6667%\">5<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"width: 16.6667%\">Y<\/td>\r\n<td style=\"width: 16.6667%\">2<\/td>\r\n<td style=\"width: 16.6667%\">5<\/td>\r\n<td style=\"width: 16.6667%\">3<\/td>\r\n<td style=\"width: 16.6667%\">8<\/td>\r\n<td style=\"width: 16.6667%\">7<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/div>\r\nAnswer: Calculate the values of X2, Y2, XY, \u2211X2,\u2211Y2 and \u2211 XY\r\n<table class=\"aligncenter\" style=\"border-collapse: collapse;width: 60%\" border=\"1\">\r\n<tbody>\r\n<tr>\r\n<td style=\"width: 20%\">X<\/td>\r\n<td style=\"width: 20%\">Y<\/td>\r\n<td style=\"width: 20%\">X<sup>2<\/sup><\/td>\r\n<td style=\"width: 20%\">Y<sup>2<\/sup><\/td>\r\n<td style=\"width: 20%\">XY<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"width: 20%\">1<\/td>\r\n<td style=\"width: 20%\">2<\/td>\r\n<td style=\"width: 20%\">1<\/td>\r\n<td style=\"width: 20%\">4<\/td>\r\n<td style=\"width: 20%\">2<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"width: 20%\">2<\/td>\r\n<td style=\"width: 20%\">5<\/td>\r\n<td style=\"width: 20%\">4<\/td>\r\n<td style=\"width: 20%\">25<\/td>\r\n<td style=\"width: 20%\">10<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"width: 20%\">3<\/td>\r\n<td style=\"width: 20%\">3<\/td>\r\n<td style=\"width: 20%\">9<\/td>\r\n<td style=\"width: 20%\">9<\/td>\r\n<td style=\"width: 20%\">9<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"width: 20%\">4<\/td>\r\n<td style=\"width: 20%\">8<\/td>\r\n<td style=\"width: 20%\">16<\/td>\r\n<td style=\"width: 20%\">64<\/td>\r\n<td style=\"width: 20%\">32<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<div>\r\n<table class=\"aligncenter\" style=\"width: 60%\" border=\"1\">\r\n<tbody>\r\n<tr>\r\n<td>5<\/td>\r\n<td>7<\/td>\r\n<td>25<\/td>\r\n<td>49<\/td>\r\n<td>35<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>\u2211X= 15<\/td>\r\n<td>\u2211Y= 25<\/td>\r\n<td>\u2211 X2= 55<\/td>\r\n<td>\u2211 Y2= 151<\/td>\r\n<td>\u2211 XY= 88<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n&nbsp;\r\n\r\ni) Step 1: For regression equations Y on X, Y= a + bX and the normal equations are-\r\n\r\n&nbsp;\r\n\r\n\u2211Y = Na + b\u2211X\r\n\r\n&nbsp;\r\n\r\n\u2211XY = a \u2211X + b\u2211X2\r\n\r\n&nbsp;\r\n\r\nStep 2: Put the values from the table in to the equations-\r\n\r\n&nbsp;\r\n\r\n25 = 5a +15b\r\n\r\n&nbsp;\r\n\r\n88 = 15a + 55b\r\n\r\n&nbsp;\r\n\r\nStep 3: After solving both the equations we get-\r\n\r\n&nbsp;\r\n\r\na= 1.10 and b = 1.3\r\n\r\n&nbsp;\r\n\r\nStep 4: Hence the required regression equation of Y on X is given by-\r\n\r\n&nbsp;\r\n\r\nY = 1.10 + 1.30X\r\n\r\n&nbsp;\r\n\r\nii) For regression equations Y on X, Y= a + bX and the normal equations are\r\n\r\n&nbsp;\r\n\r\n\u2211X = Na + b \u2211Y\r\n\r\n&nbsp;\r\n\r\n\u2211XY = a \u2211Y + b \u2211Y2\r\n\r\n&nbsp;\r\n\r\nStep 1: Substituting the values we get \u2013\r\n\r\n&nbsp;\r\n\r\n15= 5a + 25b\r\n\r\n&nbsp;\r\n\r\n88 = 25a + 151b\r\n\r\n&nbsp;\r\n\r\nStep 2: After solving both the equations we get-\r\n\r\n&nbsp;\r\n\r\na= 0.5 and b = 0.5\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n\r\nStep 3: Hence the required equation of X on Y is \u2013\r\n\r\n&nbsp;\r\n\r\nX = 0.5 + 0.5Y\r\n\r\n&nbsp;\r\n\r\n<strong>Illustration 3:<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">In a research it has been found the demand for automobiles in a city depends mainly, if not entirely, upon the number of families living in that city, below table shows for the sales of automobiles in the five cities for the year 2003 and the number of families living in those cities.<\/p>\r\n\r\n<table class=\"aligncenter\" style=\"width: 60%\" border=\"1\">\r\n<tbody>\r\n<tr>\r\n<td>City<\/td>\r\n<td>No of families in lakhs (X)<\/td>\r\n<td>Sale of Automobiles in 000\u2019s(Y)<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>A<\/td>\r\n<td>6<\/td>\r\n<td>9<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>B<\/td>\r\n<td>2<\/td>\r\n<td>11<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>C<\/td>\r\n<td>10<\/td>\r\n<td>5<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>D<\/td>\r\n<td>4<\/td>\r\n<td>8<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>E<\/td>\r\n<td>8<\/td>\r\n<td>7<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n&nbsp;\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Find the best fit linear regression equation and calculate the sales for the year 2006 for city A which is estimated to have 9 lakhs families assuming that the same relationship holds true. Also find out error and standard error between estimated value of Y and actual value of Y.<\/p>\r\n&nbsp;\r\n\r\nAnswer:\r\n\r\n&nbsp;\r\n\r\nRegression equation of Y on X is Y = a + bX\r\n\r\n&nbsp;\r\n\r\nTo determine the values of a and b, we shall solve the normal equation-\r\n\r\n&nbsp;\r\n\r\n\u2211Y = Na + b\u2211X\r\n\r\n&nbsp;\r\n\r\n\u2211XY = a \u2211X + b\u2211X2\r\n\r\n&nbsp;\r\n<table class=\"aligncenter\" style=\"width: 60%\" border=\"1\">\r\n<tbody>\r\n<tr>\r\n<td>X<\/td>\r\n<td>Y<\/td>\r\n<td>XY<\/td>\r\n<td>X2<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>6<\/td>\r\n<td>9<\/td>\r\n<td>54<\/td>\r\n<td>36<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>2<\/td>\r\n<td>11<\/td>\r\n<td>22<\/td>\r\n<td>4<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>10<\/td>\r\n<td>5<\/td>\r\n<td>50<\/td>\r\n<td>100<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>4<\/td>\r\n<td>8<\/td>\r\n<td>32<\/td>\r\n<td>16<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>8<\/td>\r\n<td>7<\/td>\r\n<td>56<\/td>\r\n<td>64<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>\u2211X= 30<\/td>\r\n<td>\u2211Y= 40<\/td>\r\n<td>\u2211XY= 214<\/td>\r\n<td>\u2211X2=220<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\ni) Step 1: Substituting the values in to normal equations 40= 5a + 30b\r\n\r\n<span style=\"font-size: 1em;text-align: initial\">214 = 30a + 220b<\/span>\r\n\r\n&nbsp;\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">Step 2: After solving both the equations we get<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">a= 11.9\u00a0 b = -0.65<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">and the regression equation is<\/span>\r\n\r\n&nbsp;\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">Y = 11.9 + (-0.65)X<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">ii) If X= 9<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">Y= 11.9 -0.65*9 Y= 6.05<\/span>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Hence the sales for the year 2006 for city A which is estimated to have 9 lakhs families assuming that the same relationship holds true is 6050 automobiles.<\/p>\r\n&nbsp;\r\n\r\niii)<img class=\"aligncenter size-full wp-image-360\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-182.png\" alt=\"\" width=\"586\" height=\"141\" \/>\r\n<div><\/div>\r\n<div>Error = \u221a 2 = \u221a3.1 =1.76<\/div>\r\n<div>Standard error =\u221a 2 =\u221a3.1 = 0.78<\/div>\r\n&nbsp;\r\n\r\n<strong>13. Summary<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">To summarize this model we can say that through correlation we establish the degree of relationship between the two variables whereas through regression analysis we try to establish a cause and effect relationship between the variables. Regression analysis establish a linear relationship between an independent variable that is already known, normally called explanatory variable and a dependent variable that is unknown, normally called as response variable.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Through the regression analysis we not only predict the future value of response variable for any given value of explanatory variable but also determine the error between the expected values and the estimated values of response variable.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The linear regression equation that completely depends upon one independent variable is called the simple regression equation and is represented by \u00a0\u0302 =b0 + b1x + \u03b5, whereas the extension of this is a multiple regression equation that depends upon more than one independent variable. The best fit\u00a0<span style=\"text-align: initial;font-size: 1em\">regression line is assumed to have all the points of both the dependent variable and the independent variable; this is an ideal situation where residual is zero.<\/span><\/p>\r\n\r\n<\/div>\r\n<p style=\"text-align: center\"><strong>Learn More:<\/strong><\/p>\r\n\r\n<ol>\r\n \t<li>Sharma, J K (2014), Business Statistics, S Chand &amp; Company, N Delhi.<\/li>\r\n \t<li>Bajpai, N (2010) Business Statistics, Pearson, N Delhi.<\/li>\r\n \t<li>Trevor Hastie, Robert Tibshirani, Jerome Friedman (2009), The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd Edition, Springer.<\/li>\r\n \t<li>Darrell Huff (2010), How to Lie with Statistics, W. W. Norton, California.<\/li>\r\n \t<li>K.R. Gupta (2012), Practical Statistics, Atlantic Publishers &amp; Distributors (P) Ltd., N. Delhi.<\/li>\r\n<\/ol>","rendered":"<div>\n<p>&nbsp;<\/p>\n<p><strong>Learning Objectives<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>After this module the students will be able to:<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">1\u00a0 Clearly define the meaning of simple linear regression.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">2\u00a0 Differentiate between the correlation and regression also state the advantages of simple linear regression.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">3 Understand types of regression model and the assumptions<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">4 Establish the simple linear regression equation by least square method and determine the values of regression coefficients.<\/p>\n<p>&nbsp;<\/p>\n<p>5 Clearly define the meaning of simple linear regression.<\/p>\n<p>&nbsp;<\/p>\n<p>6 Differentiate between the correlation and regression also state the advantages of simple linear regression.<\/p>\n<p>&nbsp;<\/p>\n<p>7\u00a0 Understand types of regression model and the assumptions<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">8 Establish the simple linear regression equation by least square method and determine the values of regression coefficients.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>1. Introduction<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Correlation coefficient and covariance define a linear relationship between the variables but they are unable to state anything about the casual relationship or in other words with correlation coefficient and covariance we cannot say which variable is a cause and which one is an effect. Through the regression analysis we try to accomplish the causal relationship between the variables. The term regression basically means \u2018moving backwards\u2019.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Hence with the help of linear regression analysis we establish the causal relationship between the linearly related variables. In causal relationship there is a <strong><em>response variable<\/em><\/strong> which is influenced by other variable called <strong><em>explanatory variables<\/em><\/strong>. Some scholars have also given different names to both response variable and explanatory variable, like other names of response variable are dependent variable, predicted variable and explained variable etc. while explanatory variables are also called as independent variable, predictor variable and control variable etc.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In regression model we try establish the complete cause and effect relationship between the variables but in most of the cases response variable does not depend on only one explanatory variable there could be more than one variable that could have either direct or indirect influence of response variable. To understand this better let us take an example of a fast moving consumer goods (FMCG) manufacturing company that wants to do research on the buying preferences of people for its new product that is going to be launched. For that they collect data of the housewives as they think that they are the decision maker for all the products that are used in kitchen but in real situation there may be the choice of children in the family that could affect the product purchase not only children there may be many other factors that should be considered like husbands preference, family income or the effect of opinion leaders etc. Hence there could be many explanatory variables that can influence the response variable.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong style=\"text-align: initial;font-size: 1em\">2. Difference between Correlation and Regression<\/strong><\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In Statistics correlation and regression both are used to define the relationship between the variables. Correlation tells the degree of relationship between the variables where regression goes one step further and tells the cause and effect relationship between the variables. In other words Correlation measures the degree to which the two variables are related whereas regression is a method of describing the relationship between two variables. Below are the basic differences between correlation and regression-<\/p>\n<p>&nbsp;<\/p>\n<ol>\n<li style=\"text-align: justify\">A statistical measure that establishes the relationship between two variables is called correlation and regression establishes the value of one variable called dependent variable for a given value of another variable called independent variable.<\/li>\n<li style=\"text-align: justify\">Correlation represents the linear relationship between the two variables whereas regression fits the best line and estimates one variable on the basis of other variable.<\/li>\n<li style=\"text-align: justify\">Correlation does not state anything about dependent and independent variables and it is symmetrical in nature for example if x and y are the two variables there is no difference in the correlation of x and y and y and x in contrary regression relationship between x and y is different from y and x.<\/li>\n<li style=\"text-align: justify\">Correlation indicates the relationship between the two variables whereas regression denotes the impact of a unit change in independent variable on the dependent variable.<\/li>\n<li style=\"text-align: justify\">Correlation finds a numerical value that expresses relationship between two variables unlike regression that predicts the future value of dependent variable for given value of independent variable.<\/li>\n<\/ol>\n<p>&nbsp;<\/p>\n<p><strong>3.\u00a0 <\/strong><strong>Advantages of Regression Analysis<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>Following are the advantages of the regression analysis-<\/p>\n<p>&nbsp;<\/p>\n<ol>\n<li style=\"text-align: justify\">Establishes the relationship between variables<strong>&#8211;<\/strong> Regression analysis establishes relationship between the response (dependent) variable and the explanatory (independent) variable.<\/li>\n<li style=\"text-align: justify\">Determines the error- Regression analysis measure the standard error of estimates to measure the variability, as in regression line we establish a relationship line which fits all the values of x and y or in other words all the value of x and y should fall on the regression line and standard error estimates is equal to zero but it hardly happens. when all the variables either fall on the line or are very close to the line is called a good regression relationship.<\/li>\n<li style=\"text-align: justify\">Suits in case of large sample size- In case of large sample size (n \u2265 30) then interval estimation for predicting the value of a dependent variable based on standard error of estimate is considered to be acceptable by changing the values of either x or y. the magnitude of r2 remains the same regardless of the values of the two variables.<\/li>\n<li style=\"text-align: justify\">Predicting the future- In regression analysis we predict the future value of response variable based on the given value of explanatory variable.<\/li>\n<\/ol>\n<p>&nbsp;<\/p>\n<p><strong>4.\u00a0 <\/strong><strong>Types of Regression Models<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">There are two methods of studying the regression model, one is simple <strong>regression model<\/strong> where response variable completely depend upon the only one explanatory variable and other is <strong>multiple regression<\/strong> <strong>model <\/strong>where response variable is influenced by more than one variable.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">In those cases where the response variable completely depends upon the one explanatory variable, this relationship between the variables is called <\/span><strong style=\"text-align: initial;font-size: 1em\"><em>deterministic<\/em>.<\/strong><span style=\"text-align: initial;font-size: 1em\"> For example-<\/span><span style=\"text-align: initial;font-size: 1em\">y=1.5x<\/span><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Here the response variable y completely depends upon one explanatory variable x and no error is allowed while predicting the values of y. This is often seen in physical science concepts.<\/span><\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">But usually, the relationship between explanatory and response variable is inexact i.e. <strong><em>stochastic<\/em><\/strong>. It is due to the omission of relevant factors which are sometimes immeasurable, that influence the response variable.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In Simple Regression model we judge the variability of response variable that is solely dependent upon only one explanatory variable. As the fundamental assumption in case of simple regression model is the expected value of y lies on a straight line, let us assume the response variable is denoted by y and explanatory variable is denoted by xi, then-y= \u03b20 +\u03b21xi<\/p>\n<p>&nbsp;<\/p>\n<p>Where \u03b20 and \u03b21 are unknown intercepts and slope parameters respectively.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In simple linear regression model \u03b20 +\u03b21xi is the deterministic component of the regression model, which tells the expected value of y for a given value of x.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">If the slop parameter \u03b21 is positive (\u03b21&gt;0) the relationship between x and y is positive, if the slop parameter \u03b21 is negative (\u03b21&lt;0) the relationship between x and y is negative and if the slop parameter \u03b21=0 there is no relationship between x and y.<\/p>\n<p>&nbsp;<\/p>\n<p>Graphically we can represent positive, negative and no relationship of linear regression model as below-<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-354\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-176.png\" alt=\"\" width=\"674\" height=\"198\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-176.png 674w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-176-300x88.png 300w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-176-65x19.png 65w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-176-225x66.png 225w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-176-350x103.png 350w\" sizes=\"auto, (max-width: 674px) 100vw, 674px\" \/><\/p>\n<p style=\"text-align: justify\">As we have discussed before actual value of response variable may defer from expected value hence we add \u03b5 (epsilon) as random error in deterministic component. Hence the sample regression model is defined as-<\/p>\n<p>&nbsp;<\/p>\n<p>y= \u03b20 + \u03b21xi + \u03b5<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">where y and x are dependent variable and independent variable respectively and \u03b5 is random error.<\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p><strong>6.\u00a0 <\/strong><strong>Assumptions for a Simple Linear Regression model<\/strong><\/p>\n<p><strong>\u00a0<\/strong><\/p>\n<ol>\n<li style=\"text-align: justify\">There should be a linear relationship between two variables x and y whereas x is called dependent variable and y is independent variable. This relationship can be described by linear regression equation-<\/li>\n<\/ol>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-355\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-177.png\" alt=\"\" width=\"106\" height=\"38\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-177.png 106w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-177-65x23.png 65w\" sizes=\"auto, (max-width: 106px) 100vw, 106px\" \/><\/p>\n<p style=\"text-align: justify\">Where \u03b5 represents the difference between the expected value and actual value of response variable y for a given value of explanatory variable x.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">2. The set of expected values of response variables y for a given value of explanatory variable x are normally distributed. The mean of these normally distributed values fall on the line of regression.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">3. The dependent variable y is a continuous random variable, whereas values of the independent variables x are fixed and not random.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">4. The sampling error associated with the expected value of the response variable is assumed to be an independent random variable distributed normally with constant standard deviation. The amount of error in the value of response variable maybe different in successive observation.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">5. The standard deviation and variance of expected values of the response variable about the regression line are constant for all the values of the explanatory variable.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">6. Regression cannot have the symmetrical value of variables means the response variable y and explanatory variable x cannot be interchanged for a same regression line equation.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>7.\u00a0 <\/strong><strong>Estimation: The Least Square Method (LS Method)<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">This method is also known as ordinary least square method (OLS method).The OLS method identifies that line which fits best for the given data. This is called the \u2018line of best fit\u2019 and is determined by identifying the line out of all of the probable lines which results in the least difference between the observed data points and the line.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Fig. -1 indicates that whenever a straight line is drawn passing through to a data, there will be some variations among the line values and the real observed values. Here, one is, fascinated about the vertical differences between this line and the real data. This line is used to make predictions about the values of Y (Dependent Variable) from different values of the X (Independent Variable). In context of regression, these differences are known as <em>residuals<\/em> and not as <em>deviations<\/em> (but actually both of them are same).<\/p>\n<\/div>\n<div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-356\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-178.png\" alt=\"\" width=\"365\" height=\"216\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-178.png 365w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-178-300x178.png 300w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-178-65x38.png 65w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-178-225x133.png 225w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-178-350x207.png 350w\" sizes=\"auto, (max-width: 365px) 100vw, 365px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Fig-1 represents a scatter plot of any data where a line is projecting the general tendency. The vertical arrows are representing the gap (differences or residuals) between the actual data and the line.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">As with the mean, values of the variables fall both above and below the line resulting in positive as well as negative differences. Thus, in case, these positive and negative differences are added, they will annul each other out .To overcome this challenge, the differences are squared before summing them. This squared differences offer a estimate about the \u2018wellness\u2019 with which any particular line fits the data,i.e.in case the square of differences are huge the line is not a representative of the data but in case the value of differences is small, that line is assumed to be a representative.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Thus to find the \u2018line of best fit\u2019 we find the sum of SS (squared differences) of <strong><em>all the likely lines<\/em><\/strong> for the given values data and then compare them. The line with the least value of SS represents the required line, i.e. line of best fit. In reality, this tedious process needs not to be followed as this can be attained by using the method of OLS. It does so with the help of mathematical method used for finding maxima and minima. This procedure is used to find the line that minimizes the sum of squared differences. This line of best fit is a regression line.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>8. Mathematical Explanation<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>Suppose a sample of n pairs of observation (x1,y1), (x2,y2)&#8230;&#8230;&#8230;,(xn,yn) is taken from a population to<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">which we want to study to estimate the values of regression coefficient \u03b20 and \u03b21. The estimate values of \u03b20 and \u03b21 should result in a straight line where most pairs of observations fall very close to it. Such a straight line is referred to as \u2018best fitted\u2019 (least squares or estimates) regression line.<\/p>\n<p>&nbsp;<\/p>\n<p>Rewriting equation as follows-<\/p>\n<p>&nbsp;<\/p>\n<p>yi= \u03b20 + \u03b21xi + <em>e<\/em>i<\/p>\n<p>or\u00a0<em>e<\/em><em>i<\/em><em>\u00a0 <\/em>= yi \u2013 (\u03b20 + \u03b21xi)<\/p>\n<p>To minimize<\/p>\n<p><em>e<\/em><em>i<\/em><em> = <\/em>L =\u2211\u00a0 =1\u00a0\u00a0 2 =\u00a0 \u2211\u00a0 =1{yi \u2013 (\u03b20 + \u03b21xi)}2<\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">let b0 and b1 be the least squares estimators of \u03b20 and \u03b21 respectively. After simplifying this equation, we get<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">\u2211\u00a0 =1 yi = n b0 + b1 \u2211\u00a0 =1<\/span><\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">This equation is called the least squares normal equation. The values of least squares estimators b0 and b1 can be obtained by solving this equation. Hence the fitted or estimated regression line is given by<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-357\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-179.png\" alt=\"\" width=\"183\" height=\"68\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-179.png 183w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-179-65x24.png 65w\" sizes=\"auto, (max-width: 183px) 100vw, 183px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p>Where-\u00a0ei = L =\u2211 =1 2 = \u2211 =1{yi \u2013 (\u03b20 + \u03b21xi)}2<\/p>\n<p>let b0 and b1 be the least squares estimators of \u03b20 and \u03b21 respectively. After simplifying this equation, we get<img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-358\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-180.png\" alt=\"\" width=\"106\" height=\"40\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-180.png 106w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-180-65x25.png 65w\" sizes=\"auto, (max-width: 106px) 100vw, 106px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">We use this equation to make the expectation of y for a given value of x since the expected value may be different from the actual value we take difference of both expected value and actual value and is generally represented by residual <em>e.<\/em><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-359\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-181.png\" alt=\"\" width=\"96\" height=\"34\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-181.png 96w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-181-65x23.png 65w\" sizes=\"auto, (max-width: 96px) 100vw, 96px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p><strong>9. Regression Coefficients<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>To estimate values of population parameter \u03b20 and \u03b21,\u00a0 the estimated simple linear regression equation is<\/p>\n<p>\u0302 <strong>= b<\/strong><strong>0<\/strong> <strong>+ b<\/strong><strong>1<\/strong><strong>x<\/strong><\/p>\n<p>where \u0302\u00a0 \u00a0 (y hat ) is estimated average value of response variable y for a given value of explanatory<\/p>\n<p style=\"text-align: justify\">variable x. b0 or a is y-intercept that represents average value of \u00a0\u0305 and b1 or b is the slop of regression line that represents the expected change in the value of y for unit change in the value of x.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">a and b are also called as intercept and regression coefficient respectively. To determine the value of y at any given point of x we must calculate the values of a and b. After getting the values of a and b we can easily determine the value of response variable that is y at any given value of explanatory variable that is x.<\/p>\n<p>&nbsp;<\/p>\n<p>The regression coefficient \u2018b\u2019 is also denoted as-<\/p>\n<p>&nbsp;<\/p>\n<p>byx that means regression coefficient of y on x and can be represented as y = a + bx<\/p>\n<p>&nbsp;<\/p>\n<p>bxy that means regression coefficient of x on y and can be represented as<\/p>\n<p>&nbsp;<\/p>\n<p>x\u00a0 = a + by<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Calculation of a and b<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><span style=\"font-size: 1em;text-align: initial\">For solving mathematical problems of regression analysis we need to calculate the intercept a and regression coefficient b.<\/span><\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">With a little algebra and differential calculus it can be shown that the following two equations, if solved simultaneously, will yield values of the parameters a and b such that the least squares requirement is fulfilled-<\/p>\n<p>&nbsp;<\/p>\n<p>\u2211Y = Na + b\u2211X<\/p>\n<p>&nbsp;<\/p>\n<p>\u2211XY = a \u2211X + b\u2211X2<\/p>\n<p>&nbsp;<\/p>\n<p>These equations are usually called the normal equations. N is the total pair of observed pairs of values.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>10. Properties of Regression coefficients<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">1.\u00a0 The correlation coefficient is the geometric mean of two regression coefficients byx and bxy i.e., <strong>r = \u221a (b<\/strong><strong>yx<\/strong><strong> X b<\/strong><strong>xy<\/strong><strong>)<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">2.\u00a0 If the one regression coefficient is greater than one, then other regression coefficient must be less than one because the value of correlation coefficient cannot be more than one.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">3.\u00a0 Both regression coefficient must have the same sign either +ve or \u2013ve. This property abolishes the case of opposite coefficient may be less than one.<\/p>\n<p>&nbsp;<\/p>\n<p>4.\u00a0 Both the correlation coefficient and the two regression coefficient will have the same sign.<\/p>\n<p>&nbsp;<\/p>\n<p>5.\u00a0\u00a0 The arithmetic mean of regression coefficient byx and bxy is more than or equal to the correlation coefficient.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>11.\u00a0 <\/strong><strong>Limitations of Simple Linear Regression model<\/strong><\/p>\n<p><strong>\u00a0<\/strong><\/p>\n<ol>\n<li>With the help of Simple linear regression model so far, we\u2019ve only been able to examine the relationship between two variables.<\/li>\n<li>In many instances, we believe that more than one independent variable is correlated with the dependent variable.<\/li>\n<li>Multiple linear regressions provides is a tool that allows us to examine the relationship between 2 or more explanatory variable and a response variable.<\/li>\n<\/ol>\n<p>&nbsp;<\/p>\n<p><strong>12.Self-Check Questions:<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><strong>Illustration 1:<\/strong><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">Five Randomly selected students took a math aptitude test along with Statistics grade.<\/span><\/p>\n<p><span style=\"font-size: 1em;text-align: initial\">In the table below, the X column shows scores on the aptitude test. Similarly, the Y column shows statistics grades.<\/span><\/p>\n<\/div>\n<div>\n<table class=\"aligncenter\" style=\"width: 60%\">\n<tbody>\n<tr>\n<td>Student<\/td>\n<td>X<\/td>\n<td>Y<\/td>\n<\/tr>\n<tr>\n<td>1<\/td>\n<td>95<\/td>\n<td>85<\/td>\n<\/tr>\n<tr>\n<td>2<\/td>\n<td>85<\/td>\n<td>95<\/td>\n<\/tr>\n<tr>\n<td>3<\/td>\n<td>80<\/td>\n<td>70<\/td>\n<\/tr>\n<tr>\n<td>4<\/td>\n<td>70<\/td>\n<td>65<\/td>\n<\/tr>\n<tr>\n<td>5<\/td>\n<td>60<\/td>\n<td>70<\/td>\n<\/tr>\n<tr>\n<td>SUM<\/td>\n<td>390<\/td>\n<td>385<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<ul>\n<li>\u00a0Find out the best fit regression line based on math aptitude test?<\/li>\n<li>If a student got 80 marks in the aptitude test find out his grade in statistics?<\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n<p>Answer:<\/p>\n<p>&nbsp;<\/p>\n<p>i)\u00a0 The linear regression equation Yon X is \u0177 = b0 + b1x .<\/p>\n<p>&nbsp;<\/p>\n<p>Step 1: To find out best fit regression line we need to solve for b0 and b1. And the normal equations are \u2013<\/p>\n<p>&nbsp;<\/p>\n<p>\u2211Y = Na + b\u2211X<\/p>\n<p>&nbsp;<\/p>\n<p>\u2211XY = a \u2211X + b\u2211X2<\/p>\n<p>&nbsp;<\/p>\n<p>Step 2: Calculate the values of X2,Y2 and XY<\/p>\n<table class=\"aligncenter\" style=\"width: 60%\">\n<tbody>\n<tr>\n<td>Student<\/td>\n<td>X<\/td>\n<td>Y<\/td>\n<td>X2<\/td>\n<td>Y2<\/td>\n<td>XY<\/td>\n<\/tr>\n<tr>\n<td>1<\/td>\n<td>95<\/td>\n<td>85<\/td>\n<td>9025<\/td>\n<td>7225<\/td>\n<td>8075<\/td>\n<\/tr>\n<tr>\n<td>2<\/td>\n<td>85<\/td>\n<td>95<\/td>\n<td>7225<\/td>\n<td>9025<\/td>\n<td>8075<\/td>\n<\/tr>\n<tr>\n<td>3<\/td>\n<td>80<\/td>\n<td>70<\/td>\n<td>6400<\/td>\n<td>4900<\/td>\n<td>5600<\/td>\n<\/tr>\n<tr>\n<td>4<\/td>\n<td>70<\/td>\n<td>65<\/td>\n<td>4900<\/td>\n<td>4225<\/td>\n<td>4550<\/td>\n<\/tr>\n<tr>\n<td>5<\/td>\n<td>60<\/td>\n<td>70<\/td>\n<td>3600<\/td>\n<td>4900<\/td>\n<td>4200<\/td>\n<\/tr>\n<tr>\n<td>SUM<\/td>\n<td>390<\/td>\n<td>385<\/td>\n<td>31150<\/td>\n<td>30275<\/td>\n<td>30500<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><span style=\"text-align: initial;font-size: 1em\">Step 3: Substitute the values in the normal equation-385 = 5a + b390\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">30500= 390b +31150b<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">Step 4:\u00a0 after solving these equations we get-<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">a=26.768<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">b= 0.644<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">Therefore, the regression equation is: \u0177 = 26.768 + 0.644x.<\/span><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">ii)\u00a0 Now once we have linear regression relationship equation between both the variables the aptitude test and the statistic\u2019s grade. We can easily predict the statistic\u2019s grade for any given value of aptitude test by just substituting the values in the found regression equation-<\/p>\n<p>&nbsp;<\/p>\n<p>\u0177 = 26.768 + 0.644x = 26.768 + 0.644 * 80 = 26.768 + 51.52 = 78.288<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Caution: it not recommended to use the values of the independent variable that are outside the range of values used to create the equation. That is called <strong>extrapolation<\/strong>, and it can produce unreasonable estimates.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In this example, the aptitude test scores used to create the regression equation ranged from 60 to 95. Therefore, only use values inside that range to estimate statistics grades. Using values outside that range (less than 60 or greater than 95) is problematic.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Illustration 2:<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>Calculate the regression equation Y on X and X on Y from the following table &#8211;<\/p>\n<table class=\"aligncenter\" style=\"border-collapse: collapse;width: 60%\">\n<tbody>\n<tr>\n<td style=\"width: 16.6667%\">X<\/td>\n<td style=\"width: 16.6667%\">1<\/td>\n<td style=\"width: 16.6667%\">2<\/td>\n<td style=\"width: 16.6667%\">3<\/td>\n<td style=\"width: 16.6667%\">4<\/td>\n<td style=\"width: 16.6667%\">5<\/td>\n<\/tr>\n<tr>\n<td style=\"width: 16.6667%\">Y<\/td>\n<td style=\"width: 16.6667%\">2<\/td>\n<td style=\"width: 16.6667%\">5<\/td>\n<td style=\"width: 16.6667%\">3<\/td>\n<td style=\"width: 16.6667%\">8<\/td>\n<td style=\"width: 16.6667%\">7<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Answer: Calculate the values of X2, Y2, XY, \u2211X2,\u2211Y2 and \u2211 XY<\/p>\n<table class=\"aligncenter\" style=\"border-collapse: collapse;width: 60%\">\n<tbody>\n<tr>\n<td style=\"width: 20%\">X<\/td>\n<td style=\"width: 20%\">Y<\/td>\n<td style=\"width: 20%\">X<sup>2<\/sup><\/td>\n<td style=\"width: 20%\">Y<sup>2<\/sup><\/td>\n<td style=\"width: 20%\">XY<\/td>\n<\/tr>\n<tr>\n<td style=\"width: 20%\">1<\/td>\n<td style=\"width: 20%\">2<\/td>\n<td style=\"width: 20%\">1<\/td>\n<td style=\"width: 20%\">4<\/td>\n<td style=\"width: 20%\">2<\/td>\n<\/tr>\n<tr>\n<td style=\"width: 20%\">2<\/td>\n<td style=\"width: 20%\">5<\/td>\n<td style=\"width: 20%\">4<\/td>\n<td style=\"width: 20%\">25<\/td>\n<td style=\"width: 20%\">10<\/td>\n<\/tr>\n<tr>\n<td style=\"width: 20%\">3<\/td>\n<td style=\"width: 20%\">3<\/td>\n<td style=\"width: 20%\">9<\/td>\n<td style=\"width: 20%\">9<\/td>\n<td style=\"width: 20%\">9<\/td>\n<\/tr>\n<tr>\n<td style=\"width: 20%\">4<\/td>\n<td style=\"width: 20%\">8<\/td>\n<td style=\"width: 20%\">16<\/td>\n<td style=\"width: 20%\">64<\/td>\n<td style=\"width: 20%\">32<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<div>\n<table class=\"aligncenter\" style=\"width: 60%\">\n<tbody>\n<tr>\n<td>5<\/td>\n<td>7<\/td>\n<td>25<\/td>\n<td>49<\/td>\n<td>35<\/td>\n<\/tr>\n<tr>\n<td>\u2211X= 15<\/td>\n<td>\u2211Y= 25<\/td>\n<td>\u2211 X2= 55<\/td>\n<td>\u2211 Y2= 151<\/td>\n<td>\u2211 XY= 88<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p>i) Step 1: For regression equations Y on X, Y= a + bX and the normal equations are-<\/p>\n<p>&nbsp;<\/p>\n<p>\u2211Y = Na + b\u2211X<\/p>\n<p>&nbsp;<\/p>\n<p>\u2211XY = a \u2211X + b\u2211X2<\/p>\n<p>&nbsp;<\/p>\n<p>Step 2: Put the values from the table in to the equations-<\/p>\n<p>&nbsp;<\/p>\n<p>25 = 5a +15b<\/p>\n<p>&nbsp;<\/p>\n<p>88 = 15a + 55b<\/p>\n<p>&nbsp;<\/p>\n<p>Step 3: After solving both the equations we get-<\/p>\n<p>&nbsp;<\/p>\n<p>a= 1.10 and b = 1.3<\/p>\n<p>&nbsp;<\/p>\n<p>Step 4: Hence the required regression equation of Y on X is given by-<\/p>\n<p>&nbsp;<\/p>\n<p>Y = 1.10 + 1.30X<\/p>\n<p>&nbsp;<\/p>\n<p>ii) For regression equations Y on X, Y= a + bX and the normal equations are<\/p>\n<p>&nbsp;<\/p>\n<p>\u2211X = Na + b \u2211Y<\/p>\n<p>&nbsp;<\/p>\n<p>\u2211XY = a \u2211Y + b \u2211Y2<\/p>\n<p>&nbsp;<\/p>\n<p>Step 1: Substituting the values we get \u2013<\/p>\n<p>&nbsp;<\/p>\n<p>15= 5a + 25b<\/p>\n<p>&nbsp;<\/p>\n<p>88 = 25a + 151b<\/p>\n<p>&nbsp;<\/p>\n<p>Step 2: After solving both the equations we get-<\/p>\n<p>&nbsp;<\/p>\n<p>a= 0.5 and b = 0.5<\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p>Step 3: Hence the required equation of X on Y is \u2013<\/p>\n<p>&nbsp;<\/p>\n<p>X = 0.5 + 0.5Y<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Illustration 3:<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In a research it has been found the demand for automobiles in a city depends mainly, if not entirely, upon the number of families living in that city, below table shows for the sales of automobiles in the five cities for the year 2003 and the number of families living in those cities.<\/p>\n<table class=\"aligncenter\" style=\"width: 60%\">\n<tbody>\n<tr>\n<td>City<\/td>\n<td>No of families in lakhs (X)<\/td>\n<td>Sale of Automobiles in 000\u2019s(Y)<\/td>\n<\/tr>\n<tr>\n<td>A<\/td>\n<td>6<\/td>\n<td>9<\/td>\n<\/tr>\n<tr>\n<td>B<\/td>\n<td>2<\/td>\n<td>11<\/td>\n<\/tr>\n<tr>\n<td>C<\/td>\n<td>10<\/td>\n<td>5<\/td>\n<\/tr>\n<tr>\n<td>D<\/td>\n<td>4<\/td>\n<td>8<\/td>\n<\/tr>\n<tr>\n<td>E<\/td>\n<td>8<\/td>\n<td>7<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Find the best fit linear regression equation and calculate the sales for the year 2006 for city A which is estimated to have 9 lakhs families assuming that the same relationship holds true. Also find out error and standard error between estimated value of Y and actual value of Y.<\/p>\n<p>&nbsp;<\/p>\n<p>Answer:<\/p>\n<p>&nbsp;<\/p>\n<p>Regression equation of Y on X is Y = a + bX<\/p>\n<p>&nbsp;<\/p>\n<p>To determine the values of a and b, we shall solve the normal equation-<\/p>\n<p>&nbsp;<\/p>\n<p>\u2211Y = Na + b\u2211X<\/p>\n<p>&nbsp;<\/p>\n<p>\u2211XY = a \u2211X + b\u2211X2<\/p>\n<p>&nbsp;<\/p>\n<table class=\"aligncenter\" style=\"width: 60%\">\n<tbody>\n<tr>\n<td>X<\/td>\n<td>Y<\/td>\n<td>XY<\/td>\n<td>X2<\/td>\n<\/tr>\n<tr>\n<td>6<\/td>\n<td>9<\/td>\n<td>54<\/td>\n<td>36<\/td>\n<\/tr>\n<tr>\n<td>2<\/td>\n<td>11<\/td>\n<td>22<\/td>\n<td>4<\/td>\n<\/tr>\n<tr>\n<td>10<\/td>\n<td>5<\/td>\n<td>50<\/td>\n<td>100<\/td>\n<\/tr>\n<tr>\n<td>4<\/td>\n<td>8<\/td>\n<td>32<\/td>\n<td>16<\/td>\n<\/tr>\n<tr>\n<td>8<\/td>\n<td>7<\/td>\n<td>56<\/td>\n<td>64<\/td>\n<\/tr>\n<tr>\n<td>\u2211X= 30<\/td>\n<td>\u2211Y= 40<\/td>\n<td>\u2211XY= 214<\/td>\n<td>\u2211X2=220<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>i) Step 1: Substituting the values in to normal equations 40= 5a + 30b<\/p>\n<p><span style=\"font-size: 1em;text-align: initial\">214 = 30a + 220b<\/span><\/p>\n<p>&nbsp;<\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">Step 2: After solving both the equations we get<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">a= 11.9\u00a0 b = -0.65<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">and the regression equation is<\/span><\/p>\n<p>&nbsp;<\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">Y = 11.9 + (-0.65)X<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">ii) If X= 9<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">Y= 11.9 -0.65*9 Y= 6.05<\/span><\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Hence the sales for the year 2006 for city A which is estimated to have 9 lakhs families assuming that the same relationship holds true is 6050 automobiles.<\/p>\n<p>&nbsp;<\/p>\n<p>iii)<img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-360\" src=\"http:\/\/mgmtp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/81\/2018\/10\/2-182.png\" alt=\"\" width=\"586\" height=\"141\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-182.png 586w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-182-300x72.png 300w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-182-65x16.png 65w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-182-225x54.png 225w, https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-content\/uploads\/sites\/81\/2018\/10\/2-182-350x84.png 350w\" sizes=\"auto, (max-width: 586px) 100vw, 586px\" \/><\/p>\n<div><\/div>\n<div>Error = \u221a 2 = \u221a3.1 =1.76<\/div>\n<div>Standard error =\u221a 2 =\u221a3.1 = 0.78<\/div>\n<p>&nbsp;<\/p>\n<p><strong>13. Summary<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">To summarize this model we can say that through correlation we establish the degree of relationship between the two variables whereas through regression analysis we try to establish a cause and effect relationship between the variables. Regression analysis establish a linear relationship between an independent variable that is already known, normally called explanatory variable and a dependent variable that is unknown, normally called as response variable.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Through the regression analysis we not only predict the future value of response variable for any given value of explanatory variable but also determine the error between the expected values and the estimated values of response variable.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The linear regression equation that completely depends upon one independent variable is called the simple regression equation and is represented by \u00a0\u0302 =b0 + b1x + \u03b5, whereas the extension of this is a multiple regression equation that depends upon more than one independent variable. The best fit\u00a0<span style=\"text-align: initial;font-size: 1em\">regression line is assumed to have all the points of both the dependent variable and the independent variable; this is an ideal situation where residual is zero.<\/span><\/p>\n<\/div>\n<p style=\"text-align: center\"><strong>Learn More:<\/strong><\/p>\n<ol>\n<li>Sharma, J K (2014), Business Statistics, S Chand &amp; Company, N Delhi.<\/li>\n<li>Bajpai, N (2010) Business Statistics, Pearson, N Delhi.<\/li>\n<li>Trevor Hastie, Robert Tibshirani, Jerome Friedman (2009), The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd Edition, Springer.<\/li>\n<li>Darrell Huff (2010), How to Lie with Statistics, W. W. Norton, California.<\/li>\n<li>K.R. Gupta (2012), Practical Statistics, Atlantic Publishers &amp; Distributors (P) Ltd., N. Delhi.<\/li>\n<\/ol>\n","protected":false},"author":3,"menu_order":32,"template":"","meta":{"pb_show_title":"on","pb_short_title":"","pb_subtitle":"","pb_authors":["dr-deependra-sharma"],"pb_section_license":""},"chapter-type":[],"contributor":[59],"license":[],"class_list":["post-351","chapter","type-chapter","status-publish","hentry","contributor-dr-deependra-sharma"],"part":3,"_links":{"self":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/pressbooks\/v2\/chapters\/351","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/pressbooks\/v2\/chapters"}],"about":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/wp\/v2\/types\/chapter"}],"author":[{"embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/wp\/v2\/users\/3"}],"version-history":[{"count":4,"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/pressbooks\/v2\/chapters\/351\/revisions"}],"predecessor-version":[{"id":362,"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/pressbooks\/v2\/chapters\/351\/revisions\/362"}],"part":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/pressbooks\/v2\/parts\/3"}],"metadata":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/pressbooks\/v2\/chapters\/351\/metadata\/"}],"wp:attachment":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/wp\/v2\/media?parent=351"}],"wp:term":[{"taxonomy":"chapter-type","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/pressbooks\/v2\/chapter-type?post=351"},{"taxonomy":"contributor","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/wp\/v2\/contributor?post=351"},{"taxonomy":"license","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/mgmtp15\/wp-json\/wp\/v2\/license?post=351"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}