{"id":92,"date":"2018-08-24T10:04:31","date_gmt":"2018-08-24T10:04:31","guid":{"rendered":"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/?post_type=chapter&#038;p=92"},"modified":"2019-01-02T06:01:41","modified_gmt":"2019-01-02T06:01:41","slug":"probability","status":"publish","type":"chapter","link":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/chapter\/probability\/","title":{"rendered":"Probability"},"content":{"raw":"<div><span style=\"float: right\"><a href=\"https:\/\/youtu.be\/f1-z06nJlR8\" target=\"_blank\" rel=\"noopener\"><img src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"epgp books\" width=\"75px\" height=\"75px;\" \/><\/a>\r\n<\/span><\/div>\r\n<div>\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\nWelcome to the e-PG Pathshala Lecture Series on Machine Learning.\r\n\r\n&nbsp;\r\n\r\n<strong>Learning Objectives:<\/strong>\r\n\r\n&nbsp;\r\n\r\nThe learning objectives of this module are as follows:\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">\u2022\u00a0 To understand the definitions of Basics of Probability for Machine Learning<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0 To understand the mathematics for Machine Learning<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0 To know the concept of Bayes Theorem and supervised Machine learning<\/p>\r\n<p style=\"text-align: justify\">\u2022 To design the Machine Learning problem using Bayes Theorem<\/p>\r\n&nbsp;\r\n\r\n<strong>6.1 \u201c7 Giants\u201d of Data<\/strong>\r\n\r\n&nbsp;\r\n\r\nData and analysis of data play a very important role in machine learning. These are the seven so called giants of data\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">1. Basic Statistics \u2013 This includes basic statistical measures such as counts, contingency table, mean, median, variance, range queries (SQL queries) etc..<\/p>\r\n<p style=\"text-align: justify\">2. Generalized N-body Problems which include kernel summations, clustering, spatial correlations, etc.<\/p>\r\n<p style=\"text-align: justify\">3. Graph Theoretic Problems which include betweeness, centrality, commute distance, graphical model inference<\/p>\r\n<p style=\"text-align: justify\">4. Optimization- general methods of Optimization<\/p>\r\n<p style=\"text-align: justify\">5. Linear Algebraic Problems which include Linear algebra, Principal component analysis, Gaussian Process Regression, Mainfold Learning<\/p>\r\n<p style=\"text-align: justify\">6. Integration using Bayesian Inference<\/p>\r\n<p style=\"text-align: justify\">7. Alignment problem which include BLAST in genomics, string matching, phylogenies, SLAM, cross-match and some other operations specific to bioinformatics<\/p>\r\n\r\n<\/div>\r\n<p style=\"text-align: justify\">However in this module we will be discussing only the basic aspects of probability and how these concepts are connected to machine learning.<\/p>\r\n&nbsp;\r\n\r\n<strong>6.2 Two views of Probability<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">When we talk about probability there are basically two aspects that we consider. One is the classical interpretation where we describe the frequency of outcomes in random experiments. The other is Bayesian viewpoint or subjective interpretation of probability where we describe the degree of belief about a particular event.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Before we go further let us first define probability. <strong>Probability<\/strong> is defined as the chance that an uncertain event will occur (always between 0 and 1). In this context <strong>Sample Space<\/strong> is defined as the collection of all possible events. Example of a sample space for all the 6 faces of a die is shown in Figure 6.1<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-93 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-49.png\" alt=\"\" width=\"442\" height=\"107\" \/>\r\n\r\n<strong>6.3 Types of Probability<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">There are three approaches to assessing the probability of an uncertain event.<\/p>\r\n<p style=\"text-align: justify\">They are defined below.<\/p>\r\n<p style=\"text-align: justify\"><strong>6.3.1 Apriori classical probability: <\/strong>the probability of an event is based on<strong> prior knowledge of the process involved<\/strong>.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-94 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-50.png\" alt=\"\" width=\"585\" height=\"88\" \/>\r\n\r\n<strong>Priori classical probability \u2013 Example 1<\/strong>\r\n\r\n&nbsp;\r\n\r\nFind the probability of selecting a face card (Jack, Queen, or King) from a standard deck of 52 cards<strong>.<\/strong>\r\n\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-95 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-51.png\" alt=\"\" width=\"523\" height=\"113\" \/>\r\n<p style=\"text-align: justify\"><strong>6.3.2 Empirical classical probability<\/strong>isthe probability of an event is based on<strong> observed data<\/strong>.<\/p>\r\n&nbsp;\r\n\r\nempirical classical probability\r\n\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-96 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-52.png\" alt=\"\" width=\"572\" height=\"84\" \/>\r\n\r\n<strong>Empirical classical probability \u2013 Example 2<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Find the probability of selecting a male taking statistics from the population described in the following table (Figure 6.2):<\/p>\r\n<img class=\"size-full wp-image-97 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-53.png\" alt=\"\" width=\"557\" height=\"305\" \/>\r\n<div>\r\n\r\n&nbsp;\r\n\r\n<strong>6.3.3\u00a0\u00a0 Subjective probability: <\/strong>the probability of an event is<strong> determined by an individual<\/strong>\r\n\r\n&nbsp;\r\n\r\n<strong>6.4 Events \u2013 Sample Space<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Associated with probability is the concept of an event. Event is defined as each possible type of occurrence or outcome. There are basically different concepts associated with an event.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0<strong>Simple event <\/strong>is an outcome from a sample space with one characteristic. ex. A red card from a deck of cards<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0 <strong>Complement of an event A <\/strong>(denoted A\/) is defined as all outcomes that are not part of event A<\/p>\r\nex. All cards that are not diamonds\r\n\r\n<\/div>\r\n<ul>\r\n \t<li style=\"text-align: justify\"><strong>Joint event<\/strong>Involves two or more characteristics simultaneously. ex. An ace that is also red from a deck of cards<\/li>\r\n \t<li style=\"text-align: justify\"><strong>Mutually exclusive event<\/strong>are eventsthat cannot occur together (simultaneously).<\/li>\r\n<\/ul>\r\n<ol>\r\n \t<li>B = having a boy; G = having a girl<\/li>\r\n<\/ol>\r\n<ul>\r\n \t<li style=\"text-align: justify\"><strong>Collectively exhaustive events<\/strong>\u2013here<strong> o<\/strong>ne of the events must occur from the set of events covers the entire sample space<\/li>\r\n<\/ul>\r\n&nbsp;\r\n\r\n<strong>Example 3<\/strong>\r\n\r\n&nbsp;\r\n\r\nA = aces; B = black cards; C = diamonds; D = hearts\r\n\r\n&nbsp;\r\n\r\nEvents A, B, C and D are <strong>collectively exhaustive<\/strong> (but not mutually exclusive) Events B, C and D are <strong>collectively exhaustive<\/strong> and also mutually exclusive\r\n\r\n&nbsp;\r\n\r\n<strong>6.5<\/strong>\u00a0\u00a0<strong>Visualizing Events in Sample Space<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">There are basically two ways in which events can be visualized in the sample space namely contingency table and tree diagrams.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Figure 6.3 and Figure 6.4 shows an example of black and red aces as events in a sample space which is a pack of 52 cards represented as contingency table and tree diagrams repectively.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-98 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-54.png\" alt=\"\" width=\"599\" height=\"464\" \/>\r\n\r\n<strong>6.6\u00a0\u00a0\u00a0 Simple vs. Joint Probability<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong>Simple (Marginal) Probability <\/strong>refers to the probability of a simple event.ex.<\/p>\r\n<p style=\"text-align: justify\">P(King). The probability of an event A given as P(A) is given below:<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-99 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-55.png\" alt=\"\" width=\"445\" height=\"75\" \/>\r\n<p style=\"text-align: justify\"><strong>Joint Probability <\/strong>refers to the probability of an occurrence of two or more events.ex. P(King and Spade)<\/p>\r\n<img class=\"size-full wp-image-100 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-56.png\" alt=\"\" width=\"533\" height=\"105\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The contingency table for joint probability is gen in Figure 6.5. In this figure we have taken the example of two sets of events <strong>A<\/strong><strong>1<\/strong> ,<strong>A<\/strong><strong>2<\/strong>and <strong>B<\/strong><strong>1<\/strong>, <strong>B<\/strong><strong>2<\/strong>. When events A and B occur together we have joint probability but the total is the marginal probability of events A and B.<\/p>\r\n<img class=\"size-full wp-image-101 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-57.png\" alt=\"\" width=\"497\" height=\"331\" \/>\r\n\r\n<strong>Probability Marginalization<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Let us now explain the concept of marginalization. Consider the probability of X irrespective of Y.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-102 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-58.png\" alt=\"\" width=\"264\" height=\"74\" \/>\r\n<p style=\"text-align: justify\">The number of instances in column j is the sum of instances in each cell of that column and is as given below:<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-103 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-59.png\" alt=\"\" width=\"477\" height=\"225\" \/>\r\n\r\n<strong>Sum and Product Rules<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">In general, we\u2019ll refer to a distribution over a random variable as p(X) and a distribution evaluated at a particular value as p(x). Two important rukes of probability include the sum and product rules.<\/p>\r\n<img class=\"size-full wp-image-104 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-60.png\" alt=\"\" width=\"468\" height=\"123\" \/>\r\n\r\n<strong>6.7 Conditional Probability and Bayes Theorem<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Another very important concept that we need to understand is the concept of conditional probability. Consider only instances where X = xj.. The fraction of these instances where Y = yi is the conditional probability. In other words the The probability of y given x is as given below<\/p>\r\n<img class=\"size-full wp-image-105 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-61.png\" alt=\"\" width=\"598\" height=\"229\" \/>\r\n<p style=\"text-align: justify\">The conditional probability can also be with respect to B given A. and is as given below:<\/p>\r\n<img class=\"size-full wp-image-106 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-62.png\" alt=\"\" width=\"571\" height=\"112\" \/>\r\n\r\nIn the above definitions\r\n\r\n&nbsp;\r\n\r\nP(A and B) = joint probability of A and B\r\n\r\nP(A) = marginal probability of A\r\n\r\nP(B) = marginal probability of B\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Based on the concept of conditional probability we go on to discuss a very important probability theorem which is used extensively in machine learning.<\/p>\r\n&nbsp;\r\n\r\n<strong>6.7.1 Bayes\u2019 Theorem<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Bayes\u2019 Theorem is used to revise previously calculated probabilities based on new information. This theorem was developed by Thomas Bayes in the 18th Century. It is an extension of conditional probability and is given by<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-107 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-63.png\" alt=\"\" width=\"573\" height=\"163\" \/>\r\n<div>\r\n\r\n&nbsp;\r\n\r\n<strong>6.7.2 Bayes\u2019 Theorem \u2013 Example 4<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">A drilling company has estimated a 40% chance of striking oil for their new well. A detailed test has been scheduled for more information. Historically, 60% of successful wells have had detailed tests, and 20% of unsuccessful wells have had detailed tests. Given that this well has been scheduled for a detailed test, what is the probability that the well will be successful?<\/p>\r\n\r\n<\/div>\r\n<ul>\r\n \t<li style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Let\u00a0S = successful well U = unsuccessful well\u00a0<\/span><\/li>\r\n \t<li style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">P(S) = .4 , P(U) = .6\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">(prior probabilities)<\/span><\/li>\r\n \t<li style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">D<\/span><span style=\"text-align: initial;font-size: 1em\">efine the detailed test event as D\u00a0<\/span><\/li>\r\n \t<li style=\"text-align: justify\"><span style=\"font-size: 1em;text-align: initial\">Conditional probabilities: P(D|S) =0.6\u00a0\u00a0\u00a0\u00a0\u00a0 P(D|U) =0.2<\/span><\/li>\r\n<\/ul>\r\n&nbsp;\r\n\r\n<strong>Goal: To find\u00a0\u00a0 P(S|D) using bayes Theorem<\/strong>\r\n\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-108 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-64.png\" alt=\"\" width=\"532\" height=\"204\" \/>\r\n<p style=\"text-align: justify\">Given the detailed test, the revised probability of a successful well has risen to .667 from the original estimate of 0.4. The given probabilities can be represented using a contingency table (Figure 6.6)<\/p>\r\n<img class=\"size-full wp-image-109 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-65.png\" alt=\"\" width=\"568\" height=\"255\" \/>\r\n\r\n<strong>6.7.3 Interpretation of Bayes Rule<\/strong>\r\n\r\n<img class=\"size-full wp-image-110 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-66.png\" alt=\"\" width=\"559\" height=\"224\" \/>\r\n\r\nThe three terms (Figure 6.7) of the Bayes Theorem are as follows:\r\n<ul>\r\n \t<li><strong>Prior<\/strong>: Information we have before observation.<\/li>\r\n \t<li><strong>Posterior<\/strong>: The distribution of Y after observing X<\/li>\r\n \t<li><strong>Likelihood: <\/strong>The likelihood of observing X given Y<\/li>\r\n<\/ul>\r\n&nbsp;\r\n\r\n<strong>6.9 Na\u00efve Bayesian Classification<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Na\u00efve Bayes Classifier is based on the Bayes Theorem. Given an instance of an entity we want to find its Class. In other words we want to find the posterior probability of the Class C given the sample data di. In order to find this class we need prior probability that is the frequency of items of Class C in the complete dataset and Likelihood which the probability that given the class C the data item di is likely to occur. If i-th attribute is <strong>categorical<\/strong>P(di|C) is estimated as the relative freq of samples having value di as i-th attribute in class C. If i-th attribute is <strong>continuous<\/strong>P(di|C) is estimated through a Gaussian density function. It is computationally easy in both cases to find this likelihood.<\/p>\r\n&nbsp;\r\n\r\n<strong>6.9.1 Theoretical basis:<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Now given the data items x1, x2, \u2026\u2026\u2026\u2026xn, we want to find a class C* that maximizes the posterior probability. We replace the posterior probability by prior probability and likelihood using Bayes Theorem,<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-111 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-67.png\" alt=\"\" width=\"656\" height=\"314\" \/>\r\n\r\n<strong>6.9.2 Maximum A Posteriori (MAP) Hypothesis and Maximum Likelihood<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Our goal is to find the most probable hypothesis <em>h<\/em> from a set of candidate hypotheses <em>H<\/em> given the observed data <em>D<\/em>. In other words we want to find hypothesis h which gives the maximum posterior probability. In Figure 6.8 the<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-112 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-68.png\" alt=\"\" width=\"561\" height=\"149\" \/>\r\n<p style=\"text-align: justify\">Posterior probability of the hypothesis h given the Data D is replaced by likelihood (P(D\/h)-probability of the hypothesis given the Data) and the prior of the hypothesis (P(h)) by applying the Bayes Theorem. If every hypothesis in <em>H<\/em> is equally probable a priori, we only need to consider the likelihood of the data <em>D <\/em>given<em> h<\/em>,<em> P(D|h). <\/em>In this case hMAP becomes the Maximum Likelihood,<\/p>\r\n&nbsp;\r\n\r\n<strong>6.9.3 Bayes Classifiers<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">For describingthe Bayes classifier we assume that training set consists of instances of different classes described <strong><em>cj<\/em><\/strong> as conjunctions of attributes values. Now the task is to classify a new instance <strong><em>d<\/em><\/strong>based on a tuple of attribute values into one of the classes <strong>cj<\/strong><strong>\u00ce<\/strong><strong>C.<\/strong> Thekey idea <strong>is to<\/strong> assign the most probable class\u00a0\u00a0using Bayes Theorem. The hypothesis we need here is the class C given the set of attribute values.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-113 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-69.png\" alt=\"\" width=\"395\" height=\"96\" \/>\r\n\r\n<strong>6.9.3.1 Example 5 \u2013 Text Classification<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">We consider a typical natural language processing application namely text classification to illustrate the method. In order to be able to perform all calculations, we will use an example with extremely small documents.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-114 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-70.png\" alt=\"\" width=\"454\" height=\"287\" \/>\r\n<p style=\"text-align: center\"><strong>Figure 6.9(a) Example 5 - Set of Documents<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">As given in Fig 6.9 (a), we are given 4training documents, and the class of each document is also given. Given a new document (called test) we want to find it\u2019s class.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">We need nonzero probabilities for all words, even words that don't exist. For this we just count every word one time more than it actually occurs (Figure 6.9(b)).Since we are only concerned with relative probabilities, this inaccuracy should not be a problem.<\/p>\r\n<p style=\"text-align: justify\">The first step in the calculation (Figure 6.9(c)) is to find the prior probability of positive and negative in the whole corpus (here it is a toy size is4).One additional document is used as test).<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-115 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-71.png\" alt=\"\" width=\"457\" height=\"181\" \/>\r\n<p style=\"text-align: justify\">Now the words in the test document are <em>good, cheat and lousy<\/em>. We need to find the likelihood for each of these words when the class is positive and when it is negative.<\/p>\r\n&nbsp;\r\n\r\n<strong>Calculations for the Example:<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">First we need to find prior probability. Here we have 3 positive and 1 negative document.<\/p>\r\n<p style=\"text-align: justify\">Therefore Prior Probability of positive and negative are P(pos) =3\/4 and P(neg) =1\/4.<\/p>\r\n<p style=\"text-align: justify\">Now we have to find Likelihood:<\/p>\r\n<p style=\"text-align: justify\">Likelihood of good in positive documents &amp; negative documents P(good\/pos) = 5 (number of goods in positive documents) +1 divided by total number of words in positive documents (8) + vocabulary (6)<\/p>\r\n<p style=\"text-align: justify\">= 6\/14 =3\/7<\/p>\r\n<p style=\"text-align: justify\">Similarly P(good\/neg) = (1+1)\/(3+6) =2\/9<\/p>\r\n<p style=\"text-align: justify\">P(cheat\/pos) = (0+1)\/(8+6) = 1\/14 and P(cheat\/neg) = (1+1)\/(3+6) = 2\/9 P(lousy\/pos) = (0+1)\/(8+6) = 1\/14 and p(lousy\/neg) = (1+1)\/(3+6) = 2\/9<\/p>\r\n<p style=\"text-align: justify\">Now we can calculate the probability of the test document being positive P(Pos\/D5) = P(POS) *( P(good\/POS) * P(cheat\/POS) * P(lousy\/POS)<\/p>\r\n\r\n<ul style=\"text-align: justify\">\r\n \t<li>= \u00be * 3\/7 * 1\/14 *1\/14 = 0.0003 Similarly P(neg\/D5) = \u00bc * 2\/9 *2\/9*2\/9 = 0.0001<\/li>\r\n<\/ul>\r\n<p style=\"text-align: justify\">Now we know that probability of positive is higher and so test document is classified aspositive.<\/p>\r\n<p style=\"text-align: justify\">Bayes Theorem in general and Na\u00efve Bayes in particular have been used in many other natural language processing applications such as Part of Speech<\/p>\r\n&nbsp;\r\n\r\nTagging, Statistical Spell Checking, Automatic Speech Recognition, Probabilistic Parsing and Statistical Machine Translation.\r\n\r\n&nbsp;\r\n\r\n<strong>6.9.3.2 Na\u00efve Bayes Classification - Example 6<\/strong>\r\n\r\n&nbsp;\r\n\r\nNow let us consider another example to illustrate Na\u00efve Bayes classification.\r\n\r\n&nbsp;\r\n\r\nWe use the equation given below to find the colour.\r\n\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-116 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-72.png\" alt=\"\" width=\"664\" height=\"434\" \/>\r\n<p style=\"text-align: justify\">Next we find the likelihood probabilities of red and blue given each of the attributes independently.<\/p>\r\n<img class=\"size-full wp-image-117 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-73.png\" alt=\"\" width=\"497\" height=\"73\" \/>\r\n<p style=\"text-align: justify\">Finally we find the posterior probabilities for each of the category and we find the category to be red since that category has maximum posterior probability<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-118 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-74.png\" alt=\"\" width=\"537\" height=\"96\" \/>\r\n<p style=\"text-align: justify\">The independence hypothesis that assuming that the attributes are independent of each other makes computation possible, yields optimal classifiers when satisfiedbut is seldom satisfied in practice, as attributes (variables) are often correlated.Bayesian networksthat combine Bayesian reasoning with causal relationships between attributes attempts to overcome this limitation. We will study Na\u00efve Bayes Classifier in detail in another module.<\/p>\r\n&nbsp;\r\n\r\n<strong>Summary<\/strong>\r\n<ul>\r\n \t<li>Described the basics of probability<\/li>\r\n \t<li>Discussed the concepts of Bayes Theorem<\/li>\r\n \t<li>Explained the application of Bayes theorem to Machine Learning<\/li>\r\n<\/ul>\r\n<table>\r\n<tbody>\r\n<tr>\r\n<td><strong>you can view video on Probability<\/strong><\/td>\r\n<td><a href=\"https:\/\/youtu.be\/f1-z06nJlR8\" target=\"_blank\" rel=\"noopener\"><img class=\"alignnone wp-image-120\" src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"\" width=\"36\" height=\"36\" \/><\/a><\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n\r\n<strong>Web Links<\/strong>\r\n<ul>\r\n \t<li style=\"text-align: justify\"><em>http:\/\/www.intmath.com\/counting-probability\/11-probability-distributions-concepts.php<\/em><\/li>\r\n \t<li style=\"text-align: justify\"><em>https:\/\/people.richland.edu\/james\/lecture\/m170\/ch06-prb.html<\/em><\/li>\r\n \t<li style=\"text-align: justify\"><em>http:\/\/www.stats.gla.ac.uk\/steps\/glossary\/probability_distributions.html<\/em><\/li>\r\n \t<li style=\"text-align: justify\"><em>https:\/\/en.wikipedia.org\/wiki\/Naive_Bayes_classifier<\/em><\/li>\r\n \t<li style=\"text-align: justify\"><em>https:\/\/en.wikipedia.org\/wiki\/Thomas_Bayes<\/em><\/li>\r\n<\/ul>","rendered":"<div><span style=\"float: right\"><a href=\"https:\/\/youtu.be\/f1-z06nJlR8\" target=\"_blank\" rel=\"noopener\"><img decoding=\"async\" src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"epgp books\" width=\"75px\" height=\"75px;\" \/><\/a><br \/>\n<\/span><\/div>\n<div>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>Welcome to the e-PG Pathshala Lecture Series on Machine Learning.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Learning Objectives:<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>The learning objectives of this module are as follows:<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0 To understand the definitions of Basics of Probability for Machine Learning<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0 To understand the mathematics for Machine Learning<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0 To know the concept of Bayes Theorem and supervised Machine learning<\/p>\n<p style=\"text-align: justify\">\u2022 To design the Machine Learning problem using Bayes Theorem<\/p>\n<p>&nbsp;<\/p>\n<p><strong>6.1 \u201c7 Giants\u201d of Data<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>Data and analysis of data play a very important role in machine learning. These are the seven so called giants of data<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">1. Basic Statistics \u2013 This includes basic statistical measures such as counts, contingency table, mean, median, variance, range queries (SQL queries) etc..<\/p>\n<p style=\"text-align: justify\">2. Generalized N-body Problems which include kernel summations, clustering, spatial correlations, etc.<\/p>\n<p style=\"text-align: justify\">3. Graph Theoretic Problems which include betweeness, centrality, commute distance, graphical model inference<\/p>\n<p style=\"text-align: justify\">4. Optimization- general methods of Optimization<\/p>\n<p style=\"text-align: justify\">5. Linear Algebraic Problems which include Linear algebra, Principal component analysis, Gaussian Process Regression, Mainfold Learning<\/p>\n<p style=\"text-align: justify\">6. Integration using Bayesian Inference<\/p>\n<p style=\"text-align: justify\">7. Alignment problem which include BLAST in genomics, string matching, phylogenies, SLAM, cross-match and some other operations specific to bioinformatics<\/p>\n<\/div>\n<p style=\"text-align: justify\">However in this module we will be discussing only the basic aspects of probability and how these concepts are connected to machine learning.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>6.2 Two views of Probability<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">When we talk about probability there are basically two aspects that we consider. One is the classical interpretation where we describe the frequency of outcomes in random experiments. The other is Bayesian viewpoint or subjective interpretation of probability where we describe the degree of belief about a particular event.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Before we go further let us first define probability. <strong>Probability<\/strong> is defined as the chance that an uncertain event will occur (always between 0 and 1). In this context <strong>Sample Space<\/strong> is defined as the collection of all possible events. Example of a sample space for all the 6 faces of a die is shown in Figure 6.1<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-93 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-49.png\" alt=\"\" width=\"442\" height=\"107\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-49.png 442w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-49-300x73.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-49-65x16.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-49-225x54.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-49-350x85.png 350w\" sizes=\"auto, (max-width: 442px) 100vw, 442px\" \/><\/p>\n<p><strong>6.3 Types of Probability<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">There are three approaches to assessing the probability of an uncertain event.<\/p>\n<p style=\"text-align: justify\">They are defined below.<\/p>\n<p style=\"text-align: justify\"><strong>6.3.1 Apriori classical probability: <\/strong>the probability of an event is based on<strong> prior knowledge of the process involved<\/strong>.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-94 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-50.png\" alt=\"\" width=\"585\" height=\"88\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-50.png 585w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-50-300x45.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-50-65x10.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-50-225x34.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-50-350x53.png 350w\" sizes=\"auto, (max-width: 585px) 100vw, 585px\" \/><\/p>\n<p><strong>Priori classical probability \u2013 Example 1<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>Find the probability of selecting a face card (Jack, Queen, or King) from a standard deck of 52 cards<strong>.<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-95 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-51.png\" alt=\"\" width=\"523\" height=\"113\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-51.png 523w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-51-300x65.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-51-65x14.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-51-225x49.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-51-350x76.png 350w\" sizes=\"auto, (max-width: 523px) 100vw, 523px\" \/><\/p>\n<p style=\"text-align: justify\"><strong>6.3.2 Empirical classical probability<\/strong>isthe probability of an event is based on<strong> observed data<\/strong>.<\/p>\n<p>&nbsp;<\/p>\n<p>empirical classical probability<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-96 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-52.png\" alt=\"\" width=\"572\" height=\"84\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-52.png 572w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-52-300x44.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-52-65x10.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-52-225x33.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-52-350x51.png 350w\" sizes=\"auto, (max-width: 572px) 100vw, 572px\" \/><\/p>\n<p><strong>Empirical classical probability \u2013 Example 2<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Find the probability of selecting a male taking statistics from the population described in the following table (Figure 6.2):<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-97 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-53.png\" alt=\"\" width=\"557\" height=\"305\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-53.png 557w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-53-300x164.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-53-65x36.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-53-225x123.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-53-350x192.png 350w\" sizes=\"auto, (max-width: 557px) 100vw, 557px\" \/><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p><strong>6.3.3\u00a0\u00a0 Subjective probability: <\/strong>the probability of an event is<strong> determined by an individual<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><strong>6.4 Events \u2013 Sample Space<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Associated with probability is the concept of an event. Event is defined as each possible type of occurrence or outcome. There are basically different concepts associated with an event.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0<strong>Simple event <\/strong>is an outcome from a sample space with one characteristic. ex. A red card from a deck of cards<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0 <strong>Complement of an event A <\/strong>(denoted A\/) is defined as all outcomes that are not part of event A<\/p>\n<p>ex. All cards that are not diamonds<\/p>\n<\/div>\n<ul>\n<li style=\"text-align: justify\"><strong>Joint event<\/strong>Involves two or more characteristics simultaneously. ex. An ace that is also red from a deck of cards<\/li>\n<li style=\"text-align: justify\"><strong>Mutually exclusive event<\/strong>are eventsthat cannot occur together (simultaneously).<\/li>\n<\/ul>\n<ol>\n<li>B = having a boy; G = having a girl<\/li>\n<\/ol>\n<ul>\n<li style=\"text-align: justify\"><strong>Collectively exhaustive events<\/strong>\u2013here<strong> o<\/strong>ne of the events must occur from the set of events covers the entire sample space<\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n<p><strong>Example 3<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>A = aces; B = black cards; C = diamonds; D = hearts<\/p>\n<p>&nbsp;<\/p>\n<p>Events A, B, C and D are <strong>collectively exhaustive<\/strong> (but not mutually exclusive) Events B, C and D are <strong>collectively exhaustive<\/strong> and also mutually exclusive<\/p>\n<p>&nbsp;<\/p>\n<p><strong>6.5<\/strong>\u00a0\u00a0<strong>Visualizing Events in Sample Space<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">There are basically two ways in which events can be visualized in the sample space namely contingency table and tree diagrams.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Figure 6.3 and Figure 6.4 shows an example of black and red aces as events in a sample space which is a pack of 52 cards represented as contingency table and tree diagrams repectively.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-98 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-54.png\" alt=\"\" width=\"599\" height=\"464\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-54.png 599w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-54-300x232.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-54-65x50.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-54-225x174.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-54-350x271.png 350w\" sizes=\"auto, (max-width: 599px) 100vw, 599px\" \/><\/p>\n<p><strong>6.6\u00a0\u00a0\u00a0 Simple vs. Joint Probability<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong>Simple (Marginal) Probability <\/strong>refers to the probability of a simple event.ex.<\/p>\n<p style=\"text-align: justify\">P(King). The probability of an event A given as P(A) is given below:<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-99 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-55.png\" alt=\"\" width=\"445\" height=\"75\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-55.png 445w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-55-300x51.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-55-65x11.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-55-225x38.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-55-350x59.png 350w\" sizes=\"auto, (max-width: 445px) 100vw, 445px\" \/><\/p>\n<p style=\"text-align: justify\"><strong>Joint Probability <\/strong>refers to the probability of an occurrence of two or more events.ex. P(King and Spade)<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-100 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-56.png\" alt=\"\" width=\"533\" height=\"105\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-56.png 533w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-56-300x59.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-56-65x13.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-56-225x44.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-56-350x69.png 350w\" sizes=\"auto, (max-width: 533px) 100vw, 533px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The contingency table for joint probability is gen in Figure 6.5. In this figure we have taken the example of two sets of events <strong>A<\/strong><strong>1<\/strong> ,<strong>A<\/strong><strong>2<\/strong>and <strong>B<\/strong><strong>1<\/strong>, <strong>B<\/strong><strong>2<\/strong>. When events A and B occur together we have joint probability but the total is the marginal probability of events A and B.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-101 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-57.png\" alt=\"\" width=\"497\" height=\"331\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-57.png 497w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-57-300x200.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-57-65x43.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-57-225x150.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-57-350x233.png 350w\" sizes=\"auto, (max-width: 497px) 100vw, 497px\" \/><\/p>\n<p><strong>Probability Marginalization<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Let us now explain the concept of marginalization. Consider the probability of X irrespective of Y.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-102 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-58.png\" alt=\"\" width=\"264\" height=\"74\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-58.png 264w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-58-65x18.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-58-225x63.png 225w\" sizes=\"auto, (max-width: 264px) 100vw, 264px\" \/><\/p>\n<p style=\"text-align: justify\">The number of instances in column j is the sum of instances in each cell of that column and is as given below:<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-103 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-59.png\" alt=\"\" width=\"477\" height=\"225\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-59.png 477w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-59-300x142.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-59-65x31.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-59-225x106.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-59-350x165.png 350w\" sizes=\"auto, (max-width: 477px) 100vw, 477px\" \/><\/p>\n<p><strong>Sum and Product Rules<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In general, we\u2019ll refer to a distribution over a random variable as p(X) and a distribution evaluated at a particular value as p(x). Two important rukes of probability include the sum and product rules.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-104 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-60.png\" alt=\"\" width=\"468\" height=\"123\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-60.png 468w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-60-300x79.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-60-65x17.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-60-225x59.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-60-350x92.png 350w\" sizes=\"auto, (max-width: 468px) 100vw, 468px\" \/><\/p>\n<p><strong>6.7 Conditional Probability and Bayes Theorem<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Another very important concept that we need to understand is the concept of conditional probability. Consider only instances where X = xj.. The fraction of these instances where Y = yi is the conditional probability. In other words the The probability of y given x is as given below<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-105 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-61.png\" alt=\"\" width=\"598\" height=\"229\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-61.png 598w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-61-300x115.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-61-65x25.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-61-225x86.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-61-350x134.png 350w\" sizes=\"auto, (max-width: 598px) 100vw, 598px\" \/><\/p>\n<p style=\"text-align: justify\">The conditional probability can also be with respect to B given A. and is as given below:<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-106 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-62.png\" alt=\"\" width=\"571\" height=\"112\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-62.png 571w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-62-300x59.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-62-65x13.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-62-225x44.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-62-350x69.png 350w\" sizes=\"auto, (max-width: 571px) 100vw, 571px\" \/><\/p>\n<p>In the above definitions<\/p>\n<p>&nbsp;<\/p>\n<p>P(A and B) = joint probability of A and B<\/p>\n<p>P(A) = marginal probability of A<\/p>\n<p>P(B) = marginal probability of B<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Based on the concept of conditional probability we go on to discuss a very important probability theorem which is used extensively in machine learning.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>6.7.1 Bayes\u2019 Theorem<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Bayes\u2019 Theorem is used to revise previously calculated probabilities based on new information. This theorem was developed by Thomas Bayes in the 18th Century. It is an extension of conditional probability and is given by<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-107 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-63.png\" alt=\"\" width=\"573\" height=\"163\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-63.png 573w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-63-300x85.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-63-65x18.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-63-225x64.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-63-350x100.png 350w\" sizes=\"auto, (max-width: 573px) 100vw, 573px\" \/><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p><strong>6.7.2 Bayes\u2019 Theorem \u2013 Example 4<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">A drilling company has estimated a 40% chance of striking oil for their new well. A detailed test has been scheduled for more information. Historically, 60% of successful wells have had detailed tests, and 20% of unsuccessful wells have had detailed tests. Given that this well has been scheduled for a detailed test, what is the probability that the well will be successful?<\/p>\n<\/div>\n<ul>\n<li style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">Let\u00a0S = successful well U = unsuccessful well\u00a0<\/span><\/li>\n<li style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">P(S) = .4 , P(U) = .6\u00a0<\/span><span style=\"text-align: initial;font-size: 1em\">(prior probabilities)<\/span><\/li>\n<li style=\"text-align: justify\"><span style=\"text-align: initial;font-size: 1em\">D<\/span><span style=\"text-align: initial;font-size: 1em\">efine the detailed test event as D\u00a0<\/span><\/li>\n<li style=\"text-align: justify\"><span style=\"font-size: 1em;text-align: initial\">Conditional probabilities: P(D|S) =0.6\u00a0\u00a0\u00a0\u00a0\u00a0 P(D|U) =0.2<\/span><\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n<p><strong>Goal: To find\u00a0\u00a0 P(S|D) using bayes Theorem<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-108 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-64.png\" alt=\"\" width=\"532\" height=\"204\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-64.png 532w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-64-300x115.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-64-65x25.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-64-225x86.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-64-350x134.png 350w\" sizes=\"auto, (max-width: 532px) 100vw, 532px\" \/><\/p>\n<p style=\"text-align: justify\">Given the detailed test, the revised probability of a successful well has risen to .667 from the original estimate of 0.4. The given probabilities can be represented using a contingency table (Figure 6.6)<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-109 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-65.png\" alt=\"\" width=\"568\" height=\"255\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-65.png 568w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-65-300x135.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-65-65x29.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-65-225x101.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-65-350x157.png 350w\" sizes=\"auto, (max-width: 568px) 100vw, 568px\" \/><\/p>\n<p><strong>6.7.3 Interpretation of Bayes Rule<\/strong><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-110 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-66.png\" alt=\"\" width=\"559\" height=\"224\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-66.png 559w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-66-300x120.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-66-65x26.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-66-225x90.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-66-350x140.png 350w\" sizes=\"auto, (max-width: 559px) 100vw, 559px\" \/><\/p>\n<p>The three terms (Figure 6.7) of the Bayes Theorem are as follows:<\/p>\n<ul>\n<li><strong>Prior<\/strong>: Information we have before observation.<\/li>\n<li><strong>Posterior<\/strong>: The distribution of Y after observing X<\/li>\n<li><strong>Likelihood: <\/strong>The likelihood of observing X given Y<\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n<p><strong>6.9 Na\u00efve Bayesian Classification<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Na\u00efve Bayes Classifier is based on the Bayes Theorem. Given an instance of an entity we want to find its Class. In other words we want to find the posterior probability of the Class C given the sample data di. In order to find this class we need prior probability that is the frequency of items of Class C in the complete dataset and Likelihood which the probability that given the class C the data item di is likely to occur. If i-th attribute is <strong>categorical<\/strong>P(di|C) is estimated as the relative freq of samples having value di as i-th attribute in class C. If i-th attribute is <strong>continuous<\/strong>P(di|C) is estimated through a Gaussian density function. It is computationally easy in both cases to find this likelihood.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>6.9.1 Theoretical basis:<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Now given the data items x1, x2, \u2026\u2026\u2026\u2026xn, we want to find a class C* that maximizes the posterior probability. We replace the posterior probability by prior probability and likelihood using Bayes Theorem,<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-111 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-67.png\" alt=\"\" width=\"656\" height=\"314\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-67.png 656w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-67-300x144.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-67-65x31.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-67-225x108.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-67-350x168.png 350w\" sizes=\"auto, (max-width: 656px) 100vw, 656px\" \/><\/p>\n<p><strong>6.9.2 Maximum A Posteriori (MAP) Hypothesis and Maximum Likelihood<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Our goal is to find the most probable hypothesis <em>h<\/em> from a set of candidate hypotheses <em>H<\/em> given the observed data <em>D<\/em>. In other words we want to find hypothesis h which gives the maximum posterior probability. In Figure 6.8 the<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-112 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-68.png\" alt=\"\" width=\"561\" height=\"149\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-68.png 561w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-68-300x80.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-68-65x17.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-68-225x60.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-68-350x93.png 350w\" sizes=\"auto, (max-width: 561px) 100vw, 561px\" \/><\/p>\n<p style=\"text-align: justify\">Posterior probability of the hypothesis h given the Data D is replaced by likelihood (P(D\/h)-probability of the hypothesis given the Data) and the prior of the hypothesis (P(h)) by applying the Bayes Theorem. If every hypothesis in <em>H<\/em> is equally probable a priori, we only need to consider the likelihood of the data <em>D <\/em>given<em> h<\/em>,<em> P(D|h). <\/em>In this case hMAP becomes the Maximum Likelihood,<\/p>\n<p>&nbsp;<\/p>\n<p><strong>6.9.3 Bayes Classifiers<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">For describingthe Bayes classifier we assume that training set consists of instances of different classes described <strong><em>cj<\/em><\/strong> as conjunctions of attributes values. Now the task is to classify a new instance <strong><em>d<\/em><\/strong>based on a tuple of attribute values into one of the classes <strong>cj<\/strong><strong>\u00ce<\/strong><strong>C.<\/strong> Thekey idea <strong>is to<\/strong> assign the most probable class\u00a0\u00a0using Bayes Theorem. The hypothesis we need here is the class C given the set of attribute values.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-113 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-69.png\" alt=\"\" width=\"395\" height=\"96\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-69.png 395w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-69-300x73.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-69-65x16.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-69-225x55.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-69-350x85.png 350w\" sizes=\"auto, (max-width: 395px) 100vw, 395px\" \/><\/p>\n<p><strong>6.9.3.1 Example 5 \u2013 Text Classification<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">We consider a typical natural language processing application namely text classification to illustrate the method. In order to be able to perform all calculations, we will use an example with extremely small documents.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-114 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-70.png\" alt=\"\" width=\"454\" height=\"287\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-70.png 454w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-70-300x190.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-70-65x41.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-70-225x142.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-70-350x221.png 350w\" sizes=\"auto, (max-width: 454px) 100vw, 454px\" \/><\/p>\n<p style=\"text-align: center\"><strong>Figure 6.9(a) Example 5 &#8211; Set of Documents<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">As given in Fig 6.9 (a), we are given 4training documents, and the class of each document is also given. Given a new document (called test) we want to find it\u2019s class.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">We need nonzero probabilities for all words, even words that don&#8217;t exist. For this we just count every word one time more than it actually occurs (Figure 6.9(b)).Since we are only concerned with relative probabilities, this inaccuracy should not be a problem.<\/p>\n<p style=\"text-align: justify\">The first step in the calculation (Figure 6.9(c)) is to find the prior probability of positive and negative in the whole corpus (here it is a toy size is4).One additional document is used as test).<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-115 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-71.png\" alt=\"\" width=\"457\" height=\"181\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-71.png 457w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-71-300x119.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-71-65x26.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-71-225x89.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-71-350x139.png 350w\" sizes=\"auto, (max-width: 457px) 100vw, 457px\" \/><\/p>\n<p style=\"text-align: justify\">Now the words in the test document are <em>good, cheat and lousy<\/em>. We need to find the likelihood for each of these words when the class is positive and when it is negative.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Calculations for the Example:<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">First we need to find prior probability. Here we have 3 positive and 1 negative document.<\/p>\n<p style=\"text-align: justify\">Therefore Prior Probability of positive and negative are P(pos) =3\/4 and P(neg) =1\/4.<\/p>\n<p style=\"text-align: justify\">Now we have to find Likelihood:<\/p>\n<p style=\"text-align: justify\">Likelihood of good in positive documents &amp; negative documents P(good\/pos) = 5 (number of goods in positive documents) +1 divided by total number of words in positive documents (8) + vocabulary (6)<\/p>\n<p style=\"text-align: justify\">= 6\/14 =3\/7<\/p>\n<p style=\"text-align: justify\">Similarly P(good\/neg) = (1+1)\/(3+6) =2\/9<\/p>\n<p style=\"text-align: justify\">P(cheat\/pos) = (0+1)\/(8+6) = 1\/14 and P(cheat\/neg) = (1+1)\/(3+6) = 2\/9 P(lousy\/pos) = (0+1)\/(8+6) = 1\/14 and p(lousy\/neg) = (1+1)\/(3+6) = 2\/9<\/p>\n<p style=\"text-align: justify\">Now we can calculate the probability of the test document being positive P(Pos\/D5) = P(POS) *( P(good\/POS) * P(cheat\/POS) * P(lousy\/POS)<\/p>\n<ul style=\"text-align: justify\">\n<li>= \u00be * 3\/7 * 1\/14 *1\/14 = 0.0003 Similarly P(neg\/D5) = \u00bc * 2\/9 *2\/9*2\/9 = 0.0001<\/li>\n<\/ul>\n<p style=\"text-align: justify\">Now we know that probability of positive is higher and so test document is classified aspositive.<\/p>\n<p style=\"text-align: justify\">Bayes Theorem in general and Na\u00efve Bayes in particular have been used in many other natural language processing applications such as Part of Speech<\/p>\n<p>&nbsp;<\/p>\n<p>Tagging, Statistical Spell Checking, Automatic Speech Recognition, Probabilistic Parsing and Statistical Machine Translation.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>6.9.3.2 Na\u00efve Bayes Classification &#8211; Example 6<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>Now let us consider another example to illustrate Na\u00efve Bayes classification.<\/p>\n<p>&nbsp;<\/p>\n<p>We use the equation given below to find the colour.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-116 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-72.png\" alt=\"\" width=\"664\" height=\"434\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-72.png 664w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-72-300x196.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-72-65x42.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-72-225x147.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-72-350x229.png 350w\" sizes=\"auto, (max-width: 664px) 100vw, 664px\" \/><\/p>\n<p style=\"text-align: justify\">Next we find the likelihood probabilities of red and blue given each of the attributes independently.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-117 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-73.png\" alt=\"\" width=\"497\" height=\"73\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-73.png 497w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-73-300x44.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-73-65x10.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-73-225x33.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-73-350x51.png 350w\" sizes=\"auto, (max-width: 497px) 100vw, 497px\" \/><\/p>\n<p style=\"text-align: justify\">Finally we find the posterior probabilities for each of the category and we find the category to be red since that category has maximum posterior probability<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-118 aligncenter\" src=\"http:\/\/csp15.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-74.png\" alt=\"\" width=\"537\" height=\"96\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-74.png 537w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-74-300x54.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-74-65x12.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-74-225x40.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-content\/uploads\/sites\/65\/2018\/08\/Untitled-74-350x63.png 350w\" sizes=\"auto, (max-width: 537px) 100vw, 537px\" \/><\/p>\n<p style=\"text-align: justify\">The independence hypothesis that assuming that the attributes are independent of each other makes computation possible, yields optimal classifiers when satisfiedbut is seldom satisfied in practice, as attributes (variables) are often correlated.Bayesian networksthat combine Bayesian reasoning with causal relationships between attributes attempts to overcome this limitation. We will study Na\u00efve Bayes Classifier in detail in another module.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Summary<\/strong><\/p>\n<ul>\n<li>Described the basics of probability<\/li>\n<li>Discussed the concepts of Bayes Theorem<\/li>\n<li>Explained the application of Bayes theorem to Machine Learning<\/li>\n<\/ul>\n<table>\n<tbody>\n<tr>\n<td><strong>you can view video on Probability<\/strong><\/td>\n<td><a href=\"https:\/\/youtu.be\/f1-z06nJlR8\" target=\"_blank\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-120\" src=\"http:\/\/epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/2018\/11\/download.png\" alt=\"\" width=\"36\" height=\"36\" \/><\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><strong>Web Links<\/strong><\/p>\n<ul>\n<li style=\"text-align: justify\"><em>http:\/\/www.intmath.com\/counting-probability\/11-probability-distributions-concepts.php<\/em><\/li>\n<li style=\"text-align: justify\"><em>https:\/\/people.richland.edu\/james\/lecture\/m170\/ch06-prb.html<\/em><\/li>\n<li style=\"text-align: justify\"><em>http:\/\/www.stats.gla.ac.uk\/steps\/glossary\/probability_distributions.html<\/em><\/li>\n<li style=\"text-align: justify\"><em>https:\/\/en.wikipedia.org\/wiki\/Naive_Bayes_classifier<\/em><\/li>\n<li style=\"text-align: justify\"><em>https:\/\/en.wikipedia.org\/wiki\/Thomas_Bayes<\/em><\/li>\n<\/ul>\n","protected":false},"author":3,"menu_order":6,"template":"","meta":{"pb_show_title":"on","pb_short_title":"","pb_subtitle":"","pb_authors":[],"pb_section_license":""},"chapter-type":[],"contributor":[],"license":[],"class_list":["post-92","chapter","type-chapter","status-publish","hentry"],"part":3,"_links":{"self":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/pressbooks\/v2\/chapters\/92","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/pressbooks\/v2\/chapters"}],"about":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/wp\/v2\/types\/chapter"}],"author":[{"embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/wp\/v2\/users\/3"}],"version-history":[{"count":6,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/pressbooks\/v2\/chapters\/92\/revisions"}],"predecessor-version":[{"id":462,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/pressbooks\/v2\/chapters\/92\/revisions\/462"}],"part":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/pressbooks\/v2\/parts\/3"}],"metadata":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/pressbooks\/v2\/chapters\/92\/metadata\/"}],"wp:attachment":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/wp\/v2\/media?parent=92"}],"wp:term":[{"taxonomy":"chapter-type","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/pressbooks\/v2\/chapter-type?post=92"},{"taxonomy":"contributor","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/wp\/v2\/contributor?post=92"},{"taxonomy":"license","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp15\/wp-json\/wp\/v2\/license?post=92"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}