{"id":53,"date":"2018-07-05T11:12:20","date_gmt":"2018-07-05T11:12:20","guid":{"rendered":"http:\/\/lisp7.epgpbooks.inflibnet.ac.in\/?post_type=chapter&#038;p=53"},"modified":"2018-08-01T05:29:15","modified_gmt":"2018-08-01T05:29:15","slug":"evaluation-and-measurement-of-information-retrieval-system","status":"publish","type":"chapter","link":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/chapter\/evaluation-and-measurement-of-information-retrieval-system\/","title":{"rendered":"Evaluation and measurement of Information Retrieval System"},"content":{"raw":"<div>\r\n\r\n&nbsp;\r\n\r\n<strong>I.\u00a0 <\/strong><strong>Objectives<\/strong>\r\n\r\n&nbsp;\r\n\r\nThe objectives of this module are to:\r\n\r\n&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 Introduce the need for evaluating information retrieval systems.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Introduce different points of view of evaluation study of IR systems.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Familiarize the reader about different factors which can affect the performance of IR systems.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Enlist different criteria to evaluate information retrieval systems.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Introduce the measures of precision and recall.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Introduce various retrieval tests and experimental tools which will help in evaluating the effectiveness of IR systems.\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>II.\u00a0\u00a0 Learning Outcomes\u00a0<\/strong>\r\n\r\n&nbsp;\r\n\r\nAfter reading this Module:\r\n\r\n&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 The reader will gain the knowledge of evaluation study and its benefits in IR.\r\n\r\n\u2022\u00a0\u00a0\u00a0 The readers will enrich their knowledge about various evaluation criteria and the importance of user oriented evaluation criteria for evaluating the IR systems.\r\n\r\n\u2022\u00a0\u00a0\u00a0 The reader will gain the knowledge of relationship between the recall and precision, fall out and generality.\r\n\r\n\u2022\u00a0\u00a0\u00a0 The reader will also learn about different evaluation tests for evaluation of information retrieval system.\r\n\r\n<\/div>\r\n&nbsp;\r\n<div>\r\n\r\n&nbsp;\r\n\r\n<strong>III.\u00a0\u00a0 Structure\u00a0<\/strong>\r\n\r\n&nbsp;\r\n\r\n1.\u00a0 Introduction\r\n\r\n2.\u00a0 Need for Evaluation\r\n\r\n3.\u00a0 Different Evaluation Criteria\r\n\r\n4.\u00a0 Evaluation of Outcome\r\n\r\n4.1\u00a0 Recall and Precision\r\n\r\n4.2\u00a0 Fallout and generality\r\n\r\n4.3\u00a0 Limitations of recall and precision\r\n\r\n5.\u00a0 Types of Evaluation Experiments\r\n\r\n5.1\u00a0 Cranfield Tests\r\n\r\n5.2\u00a0 MEDLARS\r\n\r\n5.3\u00a0 SMART Retrieval Experiment\r\n\r\n5.4\u00a0 The Stairs Project\r\n\r\n5.5\u00a0 TREC: The Text Retrieval Conference\r\n\r\n6.\u00a0 Summary\r\n\r\n7.\u00a0 References\r\n\r\n<\/div>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong style=\"text-align: initial;font-size: 1em\">1.\u00a0 Introduction\u00a0<\/strong>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Evaluation is a process of feedback against an investment in time, energy, money, knowledge and intelligence. Evaluation usually results in indicators that gauge the usefulness of systems and services. Information storage and retrieval systems are evaluated from viewpoints such as users, economy, coverage, hardware, software, man-power, environmental conditions, etc.<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>2.\u00a0 Need for Evaluation\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Evaluation studies investigate the degree to which the stated goals or expectations have been achieved or the degree to which these can be achieved. Keen (1971) gives three major purposes of evaluating an information retrieval system as follows:<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">\u2022\u00a0 \u00a0The need for measures with which to make merit comparison within a single test situation. In other words, evaluation studies are conducted to compare the merits (or demerits) of two or more systems<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0 The need for measures with which to make comparisons between results obtained in different test situations, and<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 The need for assessing the merit of real-life system.<\/p>\r\n<p style=\"text-align: justify\">Swanson (1971) states that evaluation studies have one or more of the following purposes:<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 To assess a set of goals, a program plan, or a design prior to implementation.<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 To \u00a0determine \u00a0whether \u00a0and \u00a0how \u00a0well \u00a0goals \u00a0or \u00a0performance \u00a0expectations \u00a0are \u00a0being fulfilled.<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 To determine specific reasons for successes and failures.<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 To uncover principles underlying a successful program.<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 To explore techniques for increasing program effectiveness.<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0 To establish a foundation of further research on the reasons for the relative success of alternative techniques, and<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0 To improve the means employed for attaining objectives or to redefine sub goals or goals in view of research findings.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong>3.\u00a0 Different Evaluation Criteria\u00a0<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">An evaluation study can be conducted from two different points of view. When it is conducted from managerial point of view, the evaluation study is called management-oriented; conducted from users' point of view it is called a user-oriented evaluation study. Many information scientists advocate that an evaluation of information retrieval system should always be user- oriented, i.e. evaluators should pay more attention to those factors that can provide improved service to the users.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Cleverdon (1962) says that a user oriented evaluation should try to answer the following questions which are quite relevant in modern context too:<\/p>\r\n\r\n<\/div>\r\n<div style=\"text-align: justify\">\r\n\r\n&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 To what extent does the system meet both the expressed and latent needs of its users' community?\r\n\r\n\u2022\u00a0\u00a0\u00a0 What are the reasons for the failure of the system to meet the users' needs?\r\n\r\n\u2022\u00a0\u00a0\u00a0 What is the cost-effectiveness of the searches made by the users themselves as against those made by the intermediaries?\r\n\r\n\u2022\u00a0\u00a0\u00a0 What basic changes are required to improve the output?\r\n\r\n\u2022\u00a0\u00a0\u00a0 Can the costs be reduced while maintaining the same level of performance?\r\n\r\n\u2022\u00a0\u00a0\u00a0 What would be the possible effect if some new services were introduced or an existing service were withdrawn?\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">As with any other system, we expect the best possible performance at the least cost from an information retrieval system. We can thus identify two major factors, performance and cost. Now, if we try to determine how we measure the performance of an information retrieval system we have to go back to the question of its basic objective. We know that the system is intended to retrieve all relevant documents from a collection. The system, therefore, should retrieve relevant and only relevant items. One also needs to assess how economically a system performs. Calculations of costs of an information retrieval system are not quite easy as it involves quite a number of indirect methods of calculation of costs. Lancaster (1979) lists the following major factors to be taken into consideration for the cost calculation:<\/p>\r\n&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 cost incurred per search\r\n\r\n\u2022\u00a0\u00a0\u00a0 users' efforts involved\r\n\r\n\u2022\u00a0\u00a0\u00a0 in learning how the system works\r\n\r\n\u2022\u00a0\u00a0\u00a0 in actual use\r\n\r\n\u2022\u00a0\u00a0\u00a0 in getting the documents through back-up document delivery system\r\n\r\n\u2022\u00a0\u00a0\u00a0 in retrieving information from the retrieved documents, and\r\n\r\n\u2022\u00a0\u00a0\u00a0 users' time\r\n\r\n\u2022\u00a0\u00a0\u00a0 from submission of query to the retrieval of references\r\n\r\n\u2022\u00a0\u00a0\u00a0 From \u00a0submission \u00a0of \u00a0query \u00a0to \u00a0the \u00a0retrieval \u00a0of \u00a0documents \u00a0and \u00a0the \u00a0actual information.\r\n\r\n&nbsp;\r\n\r\nA number of studies have been conducted so far to determine the cost of information retrieval system and subsystems.\r\n\r\n&nbsp;\r\n\r\nIn 1966, Cleverdon identified six criteria for the evaluation of an information retrieval system. These are:\r\n\r\n&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 <strong>recall<\/strong>, i.e., the ability of the system to present all the relevant items\r\n\r\n\u2022\u00a0\u00a0\u00a0 <strong>precision<\/strong>, i.e., the ability of the system to present only those item that are relevant\r\n\r\n\u2022\u00a0\u00a0\u00a0 <strong>time lag<\/strong>, i.e. the average interval between the time the search request is made and the time an answer is provided\r\n\r\n\u2022\u00a0\u00a0\u00a0 <strong>effort<\/strong>, intellectual as well as physical, required from the user in obtaining answers to the search requests\r\n\r\n\u2022\u00a0\u00a0\u00a0 <strong>form of presentation <\/strong>of the search output, which affects the user\u2019s ability to make use of the retrieved items, and\r\n\r\n<span style=\"text-align: justify;font-size: 1em\">\u2022\u00a0\u00a0\u00a0 <\/span><strong style=\"text-align: justify;font-size: 1em\">coverage of the collection<\/strong><span style=\"text-align: justify;font-size: 1em\">, i.e. the extent to which the system includes relevant matter. Vickery (1970) identifies six criteria, grouped into two sets as follows:<\/span>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong>Set 1<\/strong><\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 Coverage = the proportion of the total potentially useful literature that has been analysed<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 recall \u2013 the proportion of such references that are retrieved in a search, and<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 Response time \u2013 the average time needed to obtain a response from the system.<\/p>\r\n<p style=\"text-align: justify\">These three criteria are related to the availability of information, while the following three are related to the selectivity of output.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong>Set 2<\/strong><\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 precision \u2013 the ability of the system to screen out irrelevant references<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0 \u00a0usability \u00a0\u2013 \u00a0the \u00a0value \u00a0of \u00a0the \u00a0references \u00a0retrieved, \u00a0in \u00a0terms \u00a0of \u00a0such \u00a0factors \u00a0as \u00a0their reliability, comprehensibility, currency, etc., and<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0 \u00a0Presentation \u2013 the form in which search results are presented to the user. In 1971, Lancaster proposed five evaluation criteria:<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 coverage of the system<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 ability of the system to retrieve wanted items (i.e. recall);<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 ability of the system to avoid retrieval of unwanted items (i.e. precision)<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 the response time of the system, and<\/p>\r\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 The amount of effort required by the user.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">All these factors are related to the system parameters, and thus in order to identify the role played by each of the performance criteria mentioned above, each must be tagged with one or more system parameters. Salton and McGill (1983) identified the various parameters of an information retrieval system as related to each of five evaluation criteria:<\/p>\r\n&nbsp;\r\n<table class=\"aligncenter\" style=\"height: 196px\" border=\"1\" width=\"706\">\r\n<tbody>\r\n<tr>\r\n<td style=\"width: 23.0625px\"><strong>No<\/strong><\/td>\r\n<td style=\"width: 136.063px\"><strong>Evaluation<\/strong>\r\n\r\n<strong>Criteria<\/strong><\/td>\r\n<td style=\"width: 504.063px\">&nbsp;\r\n\r\n<strong>System Parameters<\/strong><\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"width: 23.0625px\">1<\/td>\r\n<td style=\"width: 136.063px\">Recall and\u00a0precision<\/td>\r\n<td style=\"width: 504.063px\">&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 Indexing exhaustively \u2013 Recall tends to increase the exhaustively of indexing terms.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Term specificity \u2013 Precision increases with the specificity of the index terms\r\n\r\n\u2022\u00a0\u00a0\u00a0 Indexing language \u2013 Availability of measures of recognition of synonyms, terms relations, etc., which improve recall.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Query formulation \u2013 Ability to formulate an accurate search request\r\n\r\n\u2022\u00a0\u00a0\u00a0 Search strategy \u2013 Ability of the user or intermediary to formulate an adequate search strategy.<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"width: 23.0625px\">2<\/td>\r\n<td style=\"width: 136.063px\">Response time<\/td>\r\n<td style=\"width: 504.063px\">&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 Organization of stored documents.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Type of query.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Location of information centre.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Frequency of receiving user\u2019s queries.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Size of the collection.<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/div>\r\n<div>\r\n<table class=\"aligncenter\" style=\"height: 315px\" border=\"1\" width=\"706\">\r\n<tbody>\r\n<tr>\r\n<td style=\"width: 22.0625px\">3<\/td>\r\n<td style=\"width: 136.063px\">User effort<\/td>\r\n<td style=\"width: 505.063px\">&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 Accessibility of the system.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Availability of guidance by system personnel.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Volume of retrieved items.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Facilities for interaction with the system.<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"width: 22.0625px\">4<\/td>\r\n<td style=\"width: 136.063px\">Form of\u00a0presentation<\/td>\r\n<td style=\"width: 505.063px\">&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 Type of display device.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Nature\u00a0\u00a0 of\u00a0\u00a0 output\u00a0\u00a0 \u2013\u00a0\u00a0\u00a0 bibliographic\u00a0\u00a0\u00a0 reference, abstract, or full text.<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"width: 22.0625px\">5<\/td>\r\n<td style=\"width: 136.063px\">Collection coverage<\/td>\r\n<td style=\"width: 505.063px\">&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 Type of input device and type and size of storage device.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Depth of subject analysis.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Nature of users\u2019 demand\r\n\r\n\u2022\u00a0\u00a0\u00a0 Physical forms of documents.<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n&nbsp;\r\n\r\n<strong>4.\u00a0 Evaluation of Outcome\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Some of the performance criteria mentioned above can be measured easily. For example, the parameters related to the collation coverage, and form of presentation is related to policy matters, and thus is defined by the system managers beforehand. However, the two other criteria, recall and precision, cannot be measured so easily.<\/p>\r\n&nbsp;\r\n\r\n<strong>4.1\u00a0 Recall and Precision\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The term recall refers to a measure of whether or not a particular item is retrieved or the extent to which the retrieval of wanted items occurs. Whenever a user puts his \/ her query, it is the responsibility of the system to retrieve all those items that are relevant to the given query. However, in reality it may not be possible to retrieve all the relevant items from a collection, especially when the collection is large. Thus, a system may be able to retrieve a proportion of the total relevant documents in response to a given query. The performance of a system is often measured by recall ratio, which denotes the percentage of relevant items retrieved in a given situation.<\/p>\r\n&nbsp;\r\n\r\nThe general formula for calculation of recall and precision may be stated as: Recall = Number of relevant items retrieved X 100\r\n\r\n&nbsp;\r\n\r\nTotal number of relevant items in the collection\r\n\r\n&nbsp;\r\n\r\nPrecision = Number of relevant items retrieved \u00a0X 100 Total number of items retrieved\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Recall thus relates to the ability of the system to retrieve relevant documents, and precision\u00a0<span style=\"font-size: 1em;text-align: initial\">relates to its ability not to retrieve non-relevant documents. The ideal system attempts to achieve 100% recall and 100 % precision, i.e. it attempts to retrieve all the relevant documents and relevant documents only. However, this is not possible in practice because as the level of recall increases, precision tends to decrease. They are inversely proportional. Following example shows the relationship between recall and precision for a given search.<\/span><\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Let us suppose that in a given situation a system retrieves 'a+b' number of documents, out of which 'a' documents are relevant, and 'b' documents are non-relevant. Say, for example, 'c+d' documents are left in the collection after the search has been conducted. This number will be quite large, because it represents the whole collection minus the retrieved documents. Out of the 'c+d' number, let\u2019s say, 'c' documents are relevant to the query but could not be retrieved, and 'd' documents are not relevant and thus have been correctly rejected. For a large collection the value of 'd' will be quite large in comparison to c because it represents all the non-relevant documents minus those that have been retrieved wrongly (here b). Lancaster suggests that these statistics can be represented as in the following table.<\/p>\r\n&nbsp;\r\n<table class=\"aligncenter\" style=\"height: 56px\" border=\"1\">\r\n<tbody>\r\n<tr style=\"height: 14px\">\r\n<td style=\"height: 14px;width: 107.063px;text-align: center\"><\/td>\r\n<td style=\"height: 14px;width: 76.0625px\"><strong>Relevant<\/strong><\/td>\r\n<td style=\"height: 14px;width: 103.063px\"><strong>Not-Relevant<\/strong><\/td>\r\n<td style=\"height: 14px;width: 66.0625px\"><strong>Total<\/strong><\/td>\r\n<\/tr>\r\n<tr style=\"height: 14px\">\r\n<td style=\"height: 14px;width: 107.063px\"><strong>Retrieved<\/strong><\/td>\r\n<td style=\"height: 14px;width: 76.0625px\">a (hits)<\/td>\r\n<td style=\"height: 14px;width: 103.063px\">b (noise)<\/td>\r\n<td style=\"height: 14px;width: 66.0625px\">a+b<\/td>\r\n<\/tr>\r\n<tr style=\"height: 14px\">\r\n<td style=\"height: 14px;width: 107.063px\"><strong>Not Retrieved<\/strong><\/td>\r\n<td style=\"height: 14px;width: 76.0625px\">c (misses)<\/td>\r\n<td style=\"height: 14px;width: 103.063px\">d (rejected)<\/td>\r\n<td style=\"height: 14px;width: 66.0625px\">c+d<\/td>\r\n<\/tr>\r\n<tr style=\"height: 14px\">\r\n<td style=\"height: 14px;width: 107.063px\"><strong>Total<\/strong><\/td>\r\n<td style=\"height: 14px;width: 76.0625px\">a+c<\/td>\r\n<td style=\"height: 14px;width: 103.063px\">b+d<\/td>\r\n<td style=\"height: 14px;width: 66.0625px\">a+b+c+d<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<p style=\"text-align: center\">Table 1: Recall \u2013 Precision Matrix<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">So as per the above table, 'a' denotes the \u2018hit\u2019 and 'b' denotes the \u2018noise\u2019. Now, out of the remaining 'c+d' documents, the system misses 'c' documents that should have been retrieved, but it correctly rejects 'd' documents that are not relevant to the given query. The recall and precision ratio in this case can be calculated as<\/p>\r\n&nbsp;\r\n\r\nR = [ a \/ ( a + c ) ] X 100 P = [ a \/ ( a + b ) ] X 100\r\n\r\n&nbsp;\r\n\r\nRecently, the theory of the \u2018inverse relationship between precision and recall\u2019 has been questioned by Fugmann (1993). By several examples, he has shown that:\r\n\r\n&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 An increasing in precision is by no means always accompanied by a corresponding decrease in recall, and\r\n\r\n\u2022\u00a0\u00a0\u00a0 An increase in recall is by no means observed to have always in its wake a decrease in precision.\r\n\r\n&nbsp;\r\n\r\n<strong>4.2\u00a0 Fallout and generality\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">While conducting a search, it is quite likely that some non-relevant items could be retrieved in a given search. This is often termed as the fallout ratio. At the same time, the proportion of relevant documents in the collection for a given query is called the generality ratio. Recall, precision, fallout, and generality ratios have been represented by Salton (1971) as shown in Tables 8.1 and Table 8.2. Thus from Table 8.1, the cut-off can be determined by the following formula:<\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n\r\nCut-off = [ ( a + b } \/ ( a + b + c + d ) ]\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Van Rijsbergen proposes that recall and precision can be combined in a single measure, called effectiveness, or E, which is a weighted combination of precision and recall where the lower the E value, the greater is the effectiveness. If recall and precision are represented by R and P respectively, the value of E can be measured through the following formula.<\/p>\r\n&nbsp;\r\n\r\nE = 100 X [ 1 - ( 1 + \u03b22) PR \/ (\u03b22 P + R) ]\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Where \u03b2 is used to reflect the relative importance of recall and precision to the user ( 0 &lt; \u03b2 &lt; \u221e ); \u03b2 = 0.5 corresponds to attaching half as much importance to recall as precision.<\/p>\r\n&nbsp;\r\n<table class=\"aligncenter\" border=\"1\">\r\n<tbody>\r\n<tr>\r\n<td><strong>Symbol<\/strong><\/td>\r\n<td><strong>Evaluation<\/strong>\r\n\r\n<strong>Measure<\/strong><\/td>\r\n<td><strong>Formula<\/strong><\/td>\r\n<td><strong>Explanation<\/strong><\/td>\r\n<\/tr>\r\n<tr>\r\n<td>R<\/td>\r\n<td>Recall<\/td>\r\n<td>a\/(a+c)<\/td>\r\n<td>Proportion of relevant items retrieved<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>P<\/td>\r\n<td>Precision<\/td>\r\n<td>a\/(a+b)<\/td>\r\n<td>Proportion \u00a0of \u00a0retrieved \u00a0items \u00a0that \u00a0are\r\n\r\nrelevant<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>F<\/td>\r\n<td>Fallout<\/td>\r\n<td>b\/(b+d)<\/td>\r\n<td>Proportion\u00a0\u00a0\u00a0\u00a0 of\u00a0\u00a0\u00a0\u00a0 non-relevant\u00a0\u00a0\u00a0\u00a0 items\r\n\r\nretrieved<\/td>\r\n<\/tr>\r\n<tr>\r\n<td>G<\/td>\r\n<td>Ganerality<\/td>\r\n<td>(a+c)\/(a+b+c+d)<\/td>\r\n<td>Proportion of relevant items per query<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<p style=\"text-align: center\">Table 2: Retrieval Measures<\/p>\r\n&nbsp;\r\n\r\n<strong>4.3\u00a0 <\/strong><strong>Limitations of recall and precision<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong>Number of documents to be retrieved: <\/strong>Different users may want different levels of recall, like a person going to prepare a state-of-the-art-report (SOTAR) on a topic would like to have all the items so the he \/ she will go for a high recall. Whereas, the user wanting to know \u2018something\u2019 about a given topic will prefer to have \u2018a few items\u2019, and thus will not require a high recall. Here, the major problem is that users very often are unable to specify exactly how many items they want to be retrieved.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Recall assumes that all relevant items have the same value, which is not always true. The retrieved items may have different degrees of relevance and this may vary from user to user and even from time to time to the same user. Both recall and precision depend largely on the relevance judgments of the user. The judgement is quite subjective and there may be different degrees of the retrieved output.<\/p>\r\n&nbsp;\r\n\r\n<strong>5.\u00a0 Types of Evaluation Experiments\u00a0<\/strong>\r\n\r\n&nbsp;\r\n\r\nLancaster (1979) identifies five major steps involved in the evaluation of an information retrieval system, which are\r\n\r\n&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 designing the scope of evaluation\r\n\r\n\u2022\u00a0\u00a0\u00a0 designing the evaluation program\r\n\r\n\u2022\u00a0\u00a0\u00a0 execution of the evaluation\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">\u2022\u00a0\u00a0\u00a0 analysis and interpretation of results, and<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">\u2022\u00a0\u00a0\u00a0 Modifying the system in the light of the evaluation results.<\/span>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\"><strong>Step 1<\/strong><\/p>\r\n<p style=\"text-align: justify\">In the step 1, an evaluation study is conducted to determine the level of performance of the given system. The evaluation would also point out the weaknesses of the system and the reasons for the same. This stage is where planning of evaluation is done, setting its purpose and scope, choosing appropriate methods, costing and staff are all decided.<\/p>\r\n&nbsp;\r\n\r\n<strong>Step 2<\/strong>\r\n<p style=\"text-align: justify\">In step 2 following the\u00a0 \u00a0basic objectives set and the proposed plans, the parameters for data collection are worked out. The evaluator must identify the points on which data are to be collected and a methodology is proposed. Data collection plans are outlined. Also the plans for exploiting data and required manipulations are decided upon.<\/p>\r\n&nbsp;\r\n\r\n<strong>Step 3<\/strong>\r\n<p style=\"text-align: justify\">Step 3 deals with execution of the evaluation. \u00a0This is time consuming. The data collectors have to collect data according to the plan and methods prescribed in the previous stage. Any alterations required to the proposed plan due to constraints at data collection stage must be communicated to the evaluator by the data collection personnel. Mutually they can agree upon required adjustments and execute them.<\/p>\r\n&nbsp;\r\n\r\n<strong>Step 4<\/strong>\r\n<p style=\"text-align: justify\">Step4 is about analysis and interpretation of the data. The success and impact of evaluation mainly depends on the accuracy of the interpretations. Data is manipulated and analysed according to aimed objective and results are obtained. These results are interpreted in the elight of the objectives of the evaluation.<\/p>\r\n&nbsp;\r\n\r\n<strong>Step 5<\/strong>\r\n<p style=\"text-align: justify\">Finally, feedback based on the results and interpretation of evaluation is given to the retrieval system so that it may be modified accordingly.<\/p>\r\n&nbsp;\r\n\r\n<strong>5.1\u00a0 Cranfield Tests\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The first extensive evaluation of information retrieval systems was undertaken at Cranfield, UK, under the direction of C W Cleverdon, and is known as the Cranfield 1 project. The first Cranfield Study began in 1957 and was reported by Cleverdon in 1962.The project was designed to compare the effectiveness of four indexing systems, viz.<\/p>\r\n&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 An alphabetical subject catalogue based on a subject heading list.\r\n\r\n\u2022\u00a0\u00a0\u00a0 A \u00a0UDC \u00a0classified \u00a0catalogue \u00a0with \u00a0alphabetical \u00a0chain \u00a0index \u00a0to \u00a0the \u00a0class \u00a0headings constructed.\r\n\r\n\u2022\u00a0\u00a0\u00a0 A catalogue based on a faceted classification and an alphabetical index to the class headings.\r\n\r\n\u2022\u00a0\u00a0\u00a0 A catalogue compiled by the uniterm coordinate index.\r\n\r\n<\/div>\r\n&nbsp;\r\n\r\n<strong><span style=\"text-align: initial;font-size: 1em\">System Parameters\u00a0<\/span><\/strong>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The study was involved 18,000 indexed items and 1200 search topics. The documents, half of which were research reports and half periodical articles, were chosen equally from the general field of aeronautics and the specialized field of high-speed aerodynamics.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Three indexes were chosen \u2013 one with subject knowledge, one with indexing experience and one straight from library school having neither subject background nor indexing experience.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Each indexer was asked to index each source document five times, spending 2, 4, 8, 12 and 16 minutes per document. One hundred source documents thus gave rise to a set of 6000 indexed items (100 documents X 3 indexers X 4 systems X 5 times ). Each of these 6000 items was tested in three places, and therefore the system worked on altogether 18,000 (6000 X 3 phases) indexed items. The test was conducted in three phases with a view to find out whether the level of performance increased with increasing experience of the system personnel.<\/p>\r\n&nbsp;\r\n\r\n<strong>Significance\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The results of the Cranfield 1 test contradicted the general belief regarding the nature of information retrieval system in many ways. The test proved that the performance of a system does not depend on the experience and subject background of the indexer. It showed that systems where documents are organized by faceted classification scheme perform poorly in comparison to the alphabetical index and uniterm system. It identified the major factors that affect the performance of retrieval systems, and developed for the first time the methodologies that could be applied successfully in evaluating information retrieval systems. Moreover, it also proved that recall and precision is the two most important parameters for determining the performance of information retrieval systems and that these two parameters are related inversely to each other.<\/p>\r\n&nbsp;\r\n\r\nThe following findings of Cranfield are significant\r\n\r\n&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 Indexing times over 4 minutes gave no real improvement in performance\r\n\r\n\u2022\u00a0\u00a0\u00a0 A high quality of indexing could be obtained from non-technical indexers\r\n\r\n\u2022\u00a0\u00a0\u00a0 The system operated at a recall rate of 70-90% and precision rate of 8 \u2013 20 %\r\n\r\n\u2022\u00a0\u00a0\u00a0 A 1 % improvement in precision could be achieved at the cost of 3% loss in recall\r\n\r\n\u2022\u00a0\u00a0\u00a0 Recall and precision were inversely related to each other\r\n\r\n\u2022\u00a0\u00a0\u00a0 All four indexing methods gave a broadly similar performance\r\n\r\n&nbsp;\r\n\r\n<strong>Cranfield test 2\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The second Cranfield test was a controlled experiment that attempted to assess the effects of the components of index languages on the performance of retrieval systems. This study tried to assess the effect by varying each factor, while keeping the others constant. Altogether 29 index languages formed by the combination of concepts were tested on 1400 documents. The test was conducted on a collection of 1400 reports and articles collected from the field of high speed aerodynamics and aircraft structures.<\/p>\r\n\r\n<\/div>\r\n<div>\r\n\r\n&nbsp;\r\n\r\n<strong>5.2\u00a0 MEDLARS\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Performance of the Medical Literature Analysis and Retrieval System (MEDLARS) of the US National Library of Medicine was assessed during August 1966 to July 1967. The test was conducted on the operational database of MEDLARS, a database of biomedical articles, index entries being drawn from the MeSH. The objective of the MEDLARS test was to evaluate the existing MEDLARS system and to find out the ways for improvisation. The document collection available on the MEDLARS service at the time of the test consisted of about 7,00,00 items.<\/p>\r\n&nbsp;\r\n\r\n21 user groups were selected from the users community that would -\r\n\r\n-\u00a0 supply some test questions;\r\n\r\n-\u00a0 cover all kinds of subjects in the requests; and\r\n\r\n-\u00a0 cover all categories of users.\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The user group so selected provided 302 search requests. Each query was formulated in terms of MeSH by the system operator and searches were conducted. After completion of a search the sample output was sent to the users for relevant assessment. Photocopies of the articles, rather than mere reference, were supplied for the relevance assessment. The user was asked to mark each retrieved item using the following scales:<\/p>\r\n&nbsp;\r\n\r\nH1 \u2013 of major value; H2 \u2013 of minor value; W1 \u2013 of no value; W2 \u2013 value unknown.\r\n\r\n&nbsp;\r\n\r\nPrecision of the searches were calculated from these figures with the following formula:\r\n\r\nPrecision ration = ((H1 + H2)\/L) X 100\r\n\r\n(Where L is the number of sample items retrieved)\r\n\r\n&nbsp;\r\n\r\n<strong>5.3\u00a0 SMART Retrieval Experiment\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">The SMART System was designed in 1964, largely as an experimental tool for the evaluation of the effectiveness of many different types of analysis and search procedures. Salton(1971) characterizes the system through the following steps of its function.<\/p>\r\n&nbsp;\r\n\r\nOverview of the functioning of SMART Retrieval System It is used to\r\n\r\n\u2022\u00a0\u00a0\u00a0 take documents and search queries posed in English;\r\n\r\n\u2022\u00a0\u00a0\u00a0 perform a fully automatic content analysis of texts;\r\n\r\n\u2022\u00a0\u00a0\u00a0 match analysed search statements and contents of documents;\r\n\r\n\u2022\u00a0\u00a0\u00a0 Retrieve the stored items which are most similar to the queries.\r\n\r\n&nbsp;\r\n\r\nA number of methods were adopted for automatic content analysis of documents, like \u2013\r\n\r\n<\/div>\r\n&nbsp;\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">-\u00a0 Word suffix cut-off methods;<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">-\u00a0 Thesaurus look-up procedures;<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">-\u00a0 phrase generation methods;<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">-\u00a0 Statistical term associations;<\/span>\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">-\u00a0 Hierarchical term expansion; and so on.<\/span>\r\n<div>\r\n\r\n&nbsp;\r\n\r\nThe following evaluation measures were generated by the SMART System:\r\n\r\n&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 A recall \u2013 precision graph reflecting the average precision value at ten discrete recall points - from a recall of 0.1 to a recall of 1.0 in intervals of 0.1.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Two \u00a0global \u00a0measures, \u00a0known \u00a0as \u00a0normalized \u00a0recall \u00a0and \u00a0normalized \u00a0precision, \u00a0which together reflect the overall performance level of the system.\r\n\r\n\u2022\u00a0\u00a0\u00a0 Two simplified global measures, known as rank recall and log precision, respectively.\r\n\r\n&nbsp;\r\n\r\n<strong>5.4\u00a0 The Stairs Project\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">In 1985, Blair and Maron (1985) published a report on a large scale experiment aimed at evaluating the retrieval effectiveness of a full-text search and retrieval system. This is known as STAIRS (Storage and Information Retrieval System) Study.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The database examined in the STAIRS study consisted of nearly 40,000 documents, representing roughly 350,000 pages of hard copy text used in the defence of a large corporate lawsuit. One important feature of STAIRS was that the lawyers who were to use the system for litigation support stipulated that they must be able to retrieve 75% of all the documents relevant to a given request. The major objective of the STAIRS evaluation was to assess how well the system could retrieve all the documents (and only those) relevant to a given request and measures of recall and precision were used for this purpose.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">In STAIRS Project, precision was calculated by dividing the total number of \u2018vital\u2019, \u2018satisfactory\u2019, and \u2018marginally relevant\u2019 documents by the total number of documents retrieved. For the calculation of recall, a sampling technique was adopted. Random samples were taken and these were evaluated by the lawyers. The total number of relevant documents that existed in these subsets was estimated.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">One of the reasons for the failure of STAIRS, as stated was, \u201cIt was impossibly difficult for users to predict the exact words, their combinations and phrases used in all or most of the relevant documents and only in those documents\".<\/p>\r\n&nbsp;\r\n\r\nNow, we will see the last information retrieval system evaluation experiment in this module, i.e.\r\n\r\n&nbsp;\r\n\r\n<strong>5.5\u00a0 TREC: The Text Retrieval Conference\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">In 1991, the US Defence Advanced Research Projects Agency (DARPA), funded the TREC experiments, to be run by the National Institute of Science and Technology (NIST) in order to enable information retrieval researchers to scale up from small collections of data to larger experiments.<\/p>\r\n\r\n<\/div>\r\n&nbsp;\r\n\r\n<span style=\"text-align: initial;font-size: 1em\">The objectives of the TREC experiments have been to<\/span>\r\n<div>\r\n\r\n&nbsp;\r\n\r\n\u2022\u00a0\u00a0\u00a0 Encourage retrieval research based on large test collections\r\n\r\n\u2022\u00a0\u00a0\u00a0 Increase communication among industry, academia and government\r\n\r\n\u2022\u00a0\u00a0\u00a0 Speed the transfer of technology from research labs to commercial products\r\n\r\n\u2022\u00a0\u00a0\u00a0 Increase the availability of appropriate evaluation techniques for use by industry and academia.\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">TREC is an annual workshop series aimed at building the infrastructure required for extracting relevant information from large volumes of electronic documents. The first TREC Conference was held in 1992 and research into improved text-retrieval methodologies has continued ever since. The document collection of TREC is known as the TIPSTER collection. TIPSTER reflects the diversity of subject matter, word choice, literary style, formats, and so on. Its primary collection is in gigabytes with over a million documents.<\/p>\r\n&nbsp;\r\n\r\n<strong>6.\u00a0 Summary\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">With this, we have completed the discussion on various aspects of information retrieval evaluation experiments and are now ready to conclude this module by saying that different users may want different levels of recall, like a person going to prepare a state-of-the-art-report (SOTAR) on a topic would like to have all the items so he \/ she will go for a high recall. On the other hand, the user wanting to know \u2018something\u2019 about a given topic will prefer to have \u2018a few items\u2019, and thus will not require a high recall. Here, the major problem is that users very often are unable to specify exactly how many items they want to be retrieved.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Recall assumes that all relevant items have the same value, which is not always true. The retrieved items may have different degrees of relevance and this may vary from user to user and even from time to time to the same user. Both recall and precision depend largely on the relevance judgments of the user. The judgement is quite subjective and there may be different degrees of the retrieved output.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Finally to conclude, one can say that the performance evaluation is crucial at many stages in information retrieval system development. At the end of development process, it is significant to show that the final retrieval system achieves an acceptable level of performance and that it represents a significant improvement over existing retrieval systems. To evaluate a retrieval system, there is need to estimate the future performance of the system.<\/p>\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n&nbsp;\r\n\r\n<strong>7.\u00a0 References\u00a0<\/strong>\r\n\r\n&nbsp;\r\n<ol>\r\n \t<li>Blair, D C and Maron, M E., An evaluation of retrieval effectiveness for a full-text document retrieval system. Communications of the ACM, 28(3), 1985. 289-99 pp.<\/li>\r\n \t<li>Cleverdon, \u00a0Cyril \u00a0W. \u00a01967. \u00a0The \u00a0Cranfield \u00a0tests \u00a0on \u00a0index \u00a0language \u00a0devices. \u00a0Aslib Proceedings 19, no. 6:173-194<\/li>\r\n<\/ol>\r\n<\/div>\r\n<ol start=\"3\">\r\n \t<li>Cleverdon, C W., Report on the first step of an investigation into comparative efficiency of indexing systems, Cranfield, College of Aeronautics, 1960<\/li>\r\n \t<li>Cleverdon, C W., Report on the testing and analysis of an investigation into comparative efficiency of indexing systems, Cranfield, College of Aeronautics, 1962<\/li>\r\n \t<li>Fungmann, R., Subject analysis and indexing: theoretical foundation and practical advice by Robert Fungman, Frankfurt, IdeksVerlag, 1993<\/li>\r\n \t<li>Keen, E. M., \u2018Evaluation parameters\u2019. In: In: Salta G. (ed.), The SMART Retrieval System: experiments in automatic document processing. Englewood Cliffs, New Jersey: Prentice Hall, 1971, 74-111 pp.<\/li>\r\n \t<li>Lancaster, F W. Information Retrieval Systems: Characteristics, testing and evaluation, 2ndedn, New York, John Wiley, 1979.<\/li>\r\n \t<li>Salton, Gerald. 1971. Relevance feedback and the optimization of retrieval effectiveness. Ghap \u00a015 \u00a0in \u00a0The \u00a0SMART \u00a0Retrieval \u00a0System \u00a0\u2013 \u00a0Experiments \u00a0in Automatic \u00a0Document processing. Ed. G Salton. Pp. 324-336. Englewood Cliffs, New Jersey: Prentice Hall<\/li>\r\n \t<li>Salton, Gerald, \u2018The SMART Project: Status report and plan\u2019. In: Salta G. (ed.), The SMART Retrieval System: experiments in automatic document processing. Englewood Cliffs, New Jersey: Prentice Hall, 1971, 03-11 pp.<\/li>\r\n \t<li>Salton, Gerald. 1971. Relevance feedback and the optimization of retrieval effectiveness. Ghap \u00a015 \u00a0in \u00a0The \u00a0SMART \u00a0Retrieval \u00a0System \u00a0\u2013 \u00a0Experiments \u00a0in Automatic \u00a0Document processing. Ed. G Salton. Pp. 324-336. Englewood Cliffs, New Jersey: Prentice Hall<\/li>\r\n \t<li>Salton, G. \u2018Another look at automatic text-retrieval systems\u2019, communications of the ACM, 29(7), 1986, 648-656 pp.<\/li>\r\n \t<li>Salton, G. and McGill, M. J., Introduction to modern information retrieval, Auckland, McGraw-Hill, 1983.<\/li>\r\n \t<li>Swanson, D R. \u2018Some unexplained aspects of the Cranfield tests of indexing language performance\u2019, Library Quarterly, 41, 1971, 223-228 pp.<\/li>\r\n \t<li>TREC. Available at <a href=\"http:\/\/trec.nist.gov\/pubs.html\">http:\/\/trec.nist.gov\/pubs.html<\/a><\/li>\r\n \t<li>van Rijsbergen, C. J., Information retrieval, 2nd ed., London, Butterworth, 1979<\/li>\r\n \t<li>Vickery, B. C. Techniques of information retrieval. London, Butterworth, 1970.<\/li>\r\n<\/ol>","rendered":"<div>\n<p>&nbsp;<\/p>\n<p><strong>I.\u00a0 <\/strong><strong>Objectives<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>The objectives of this module are to:<\/p>\n<p>&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Introduce the need for evaluating information retrieval systems.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Introduce different points of view of evaluation study of IR systems.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Familiarize the reader about different factors which can affect the performance of IR systems.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Enlist different criteria to evaluate information retrieval systems.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Introduce the measures of precision and recall.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Introduce various retrieval tests and experimental tools which will help in evaluating the effectiveness of IR systems.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>II.\u00a0\u00a0 Learning Outcomes\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>After reading this Module:<\/p>\n<p>&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 The reader will gain the knowledge of evaluation study and its benefits in IR.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 The readers will enrich their knowledge about various evaluation criteria and the importance of user oriented evaluation criteria for evaluating the IR systems.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 The reader will gain the knowledge of relationship between the recall and precision, fall out and generality.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 The reader will also learn about different evaluation tests for evaluation of information retrieval system.<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<div>\n<p>&nbsp;<\/p>\n<p><strong>III.\u00a0\u00a0 Structure\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>1.\u00a0 Introduction<\/p>\n<p>2.\u00a0 Need for Evaluation<\/p>\n<p>3.\u00a0 Different Evaluation Criteria<\/p>\n<p>4.\u00a0 Evaluation of Outcome<\/p>\n<p>4.1\u00a0 Recall and Precision<\/p>\n<p>4.2\u00a0 Fallout and generality<\/p>\n<p>4.3\u00a0 Limitations of recall and precision<\/p>\n<p>5.\u00a0 Types of Evaluation Experiments<\/p>\n<p>5.1\u00a0 Cranfield Tests<\/p>\n<p>5.2\u00a0 MEDLARS<\/p>\n<p>5.3\u00a0 SMART Retrieval Experiment<\/p>\n<p>5.4\u00a0 The Stairs Project<\/p>\n<p>5.5\u00a0 TREC: The Text Retrieval Conference<\/p>\n<p>6.\u00a0 Summary<\/p>\n<p>7.\u00a0 References<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong style=\"text-align: initial;font-size: 1em\">1.\u00a0 Introduction\u00a0<\/strong><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Evaluation is a process of feedback against an investment in time, energy, money, knowledge and intelligence. Evaluation usually results in indicators that gauge the usefulness of systems and services. Information storage and retrieval systems are evaluated from viewpoints such as users, economy, coverage, hardware, software, man-power, environmental conditions, etc.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>2.\u00a0 Need for Evaluation\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Evaluation studies investigate the degree to which the stated goals or expectations have been achieved or the degree to which these can be achieved. Keen (1971) gives three major purposes of evaluating an information retrieval system as follows:<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0 \u00a0The need for measures with which to make merit comparison within a single test situation. In other words, evaluation studies are conducted to compare the merits (or demerits) of two or more systems<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0 The need for measures with which to make comparisons between results obtained in different test situations, and<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 The need for assessing the merit of real-life system.<\/p>\n<p style=\"text-align: justify\">Swanson (1971) states that evaluation studies have one or more of the following purposes:<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 To assess a set of goals, a program plan, or a design prior to implementation.<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 To \u00a0determine \u00a0whether \u00a0and \u00a0how \u00a0well \u00a0goals \u00a0or \u00a0performance \u00a0expectations \u00a0are \u00a0being fulfilled.<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 To determine specific reasons for successes and failures.<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 To uncover principles underlying a successful program.<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 To explore techniques for increasing program effectiveness.<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0 To establish a foundation of further research on the reasons for the relative success of alternative techniques, and<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0 To improve the means employed for attaining objectives or to redefine sub goals or goals in view of research findings.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong>3.\u00a0 Different Evaluation Criteria\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">An evaluation study can be conducted from two different points of view. When it is conducted from managerial point of view, the evaluation study is called management-oriented; conducted from users&#8217; point of view it is called a user-oriented evaluation study. Many information scientists advocate that an evaluation of information retrieval system should always be user- oriented, i.e. evaluators should pay more attention to those factors that can provide improved service to the users.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Cleverdon (1962) says that a user oriented evaluation should try to answer the following questions which are quite relevant in modern context too:<\/p>\n<\/div>\n<div style=\"text-align: justify\">\n<p>&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 To what extent does the system meet both the expressed and latent needs of its users&#8217; community?<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 What are the reasons for the failure of the system to meet the users&#8217; needs?<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 What is the cost-effectiveness of the searches made by the users themselves as against those made by the intermediaries?<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 What basic changes are required to improve the output?<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Can the costs be reduced while maintaining the same level of performance?<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 What would be the possible effect if some new services were introduced or an existing service were withdrawn?<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">As with any other system, we expect the best possible performance at the least cost from an information retrieval system. We can thus identify two major factors, performance and cost. Now, if we try to determine how we measure the performance of an information retrieval system we have to go back to the question of its basic objective. We know that the system is intended to retrieve all relevant documents from a collection. The system, therefore, should retrieve relevant and only relevant items. One also needs to assess how economically a system performs. Calculations of costs of an information retrieval system are not quite easy as it involves quite a number of indirect methods of calculation of costs. Lancaster (1979) lists the following major factors to be taken into consideration for the cost calculation:<\/p>\n<p>&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 cost incurred per search<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 users&#8217; efforts involved<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 in learning how the system works<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 in actual use<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 in getting the documents through back-up document delivery system<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 in retrieving information from the retrieved documents, and<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 users&#8217; time<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 from submission of query to the retrieval of references<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 From \u00a0submission \u00a0of \u00a0query \u00a0to \u00a0the \u00a0retrieval \u00a0of \u00a0documents \u00a0and \u00a0the \u00a0actual information.<\/p>\n<p>&nbsp;<\/p>\n<p>A number of studies have been conducted so far to determine the cost of information retrieval system and subsystems.<\/p>\n<p>&nbsp;<\/p>\n<p>In 1966, Cleverdon identified six criteria for the evaluation of an information retrieval system. These are:<\/p>\n<p>&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 <strong>recall<\/strong>, i.e., the ability of the system to present all the relevant items<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 <strong>precision<\/strong>, i.e., the ability of the system to present only those item that are relevant<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 <strong>time lag<\/strong>, i.e. the average interval between the time the search request is made and the time an answer is provided<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 <strong>effort<\/strong>, intellectual as well as physical, required from the user in obtaining answers to the search requests<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 <strong>form of presentation <\/strong>of the search output, which affects the user\u2019s ability to make use of the retrieved items, and<\/p>\n<p><span style=\"text-align: justify;font-size: 1em\">\u2022\u00a0\u00a0\u00a0 <\/span><strong style=\"text-align: justify;font-size: 1em\">coverage of the collection<\/strong><span style=\"text-align: justify;font-size: 1em\">, i.e. the extent to which the system includes relevant matter. Vickery (1970) identifies six criteria, grouped into two sets as follows:<\/span><\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong>Set 1<\/strong><\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 Coverage = the proportion of the total potentially useful literature that has been analysed<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 recall \u2013 the proportion of such references that are retrieved in a search, and<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 Response time \u2013 the average time needed to obtain a response from the system.<\/p>\n<p style=\"text-align: justify\">These three criteria are related to the availability of information, while the following three are related to the selectivity of output.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong>Set 2<\/strong><\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 precision \u2013 the ability of the system to screen out irrelevant references<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0 \u00a0usability \u00a0\u2013 \u00a0the \u00a0value \u00a0of \u00a0the \u00a0references \u00a0retrieved, \u00a0in \u00a0terms \u00a0of \u00a0such \u00a0factors \u00a0as \u00a0their reliability, comprehensibility, currency, etc., and<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0 \u00a0Presentation \u2013 the form in which search results are presented to the user. In 1971, Lancaster proposed five evaluation criteria:<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 coverage of the system<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 ability of the system to retrieve wanted items (i.e. recall);<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 ability of the system to avoid retrieval of unwanted items (i.e. precision)<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 the response time of the system, and<\/p>\n<p style=\"text-align: justify\">\u2022\u00a0\u00a0\u00a0 The amount of effort required by the user.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">All these factors are related to the system parameters, and thus in order to identify the role played by each of the performance criteria mentioned above, each must be tagged with one or more system parameters. Salton and McGill (1983) identified the various parameters of an information retrieval system as related to each of five evaluation criteria:<\/p>\n<p>&nbsp;<\/p>\n<table class=\"aligncenter\" style=\"height: 196px; width: 706px;\">\n<tbody>\n<tr>\n<td style=\"width: 23.0625px\"><strong>No<\/strong><\/td>\n<td style=\"width: 136.063px\"><strong>Evaluation<\/strong><\/p>\n<p><strong>Criteria<\/strong><\/td>\n<td style=\"width: 504.063px\">&nbsp;<\/p>\n<p><strong>System Parameters<\/strong><\/td>\n<\/tr>\n<tr>\n<td style=\"width: 23.0625px\">1<\/td>\n<td style=\"width: 136.063px\">Recall and\u00a0precision<\/td>\n<td style=\"width: 504.063px\">&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Indexing exhaustively \u2013 Recall tends to increase the exhaustively of indexing terms.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Term specificity \u2013 Precision increases with the specificity of the index terms<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Indexing language \u2013 Availability of measures of recognition of synonyms, terms relations, etc., which improve recall.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Query formulation \u2013 Ability to formulate an accurate search request<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Search strategy \u2013 Ability of the user or intermediary to formulate an adequate search strategy.<\/td>\n<\/tr>\n<tr>\n<td style=\"width: 23.0625px\">2<\/td>\n<td style=\"width: 136.063px\">Response time<\/td>\n<td style=\"width: 504.063px\">&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Organization of stored documents.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Type of query.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Location of information centre.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Frequency of receiving user\u2019s queries.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Size of the collection.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<div>\n<table class=\"aligncenter\" style=\"height: 315px; width: 706px;\">\n<tbody>\n<tr>\n<td style=\"width: 22.0625px\">3<\/td>\n<td style=\"width: 136.063px\">User effort<\/td>\n<td style=\"width: 505.063px\">&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Accessibility of the system.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Availability of guidance by system personnel.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Volume of retrieved items.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Facilities for interaction with the system.<\/td>\n<\/tr>\n<tr>\n<td style=\"width: 22.0625px\">4<\/td>\n<td style=\"width: 136.063px\">Form of\u00a0presentation<\/td>\n<td style=\"width: 505.063px\">&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Type of display device.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Nature\u00a0\u00a0 of\u00a0\u00a0 output\u00a0\u00a0 \u2013\u00a0\u00a0\u00a0 bibliographic\u00a0\u00a0\u00a0 reference, abstract, or full text.<\/td>\n<\/tr>\n<tr>\n<td style=\"width: 22.0625px\">5<\/td>\n<td style=\"width: 136.063px\">Collection coverage<\/td>\n<td style=\"width: 505.063px\">&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Type of input device and type and size of storage device.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Depth of subject analysis.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Nature of users\u2019 demand<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Physical forms of documents.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>4.\u00a0 Evaluation of Outcome\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Some of the performance criteria mentioned above can be measured easily. For example, the parameters related to the collation coverage, and form of presentation is related to policy matters, and thus is defined by the system managers beforehand. However, the two other criteria, recall and precision, cannot be measured so easily.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>4.1\u00a0 Recall and Precision\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The term recall refers to a measure of whether or not a particular item is retrieved or the extent to which the retrieval of wanted items occurs. Whenever a user puts his \/ her query, it is the responsibility of the system to retrieve all those items that are relevant to the given query. However, in reality it may not be possible to retrieve all the relevant items from a collection, especially when the collection is large. Thus, a system may be able to retrieve a proportion of the total relevant documents in response to a given query. The performance of a system is often measured by recall ratio, which denotes the percentage of relevant items retrieved in a given situation.<\/p>\n<p>&nbsp;<\/p>\n<p>The general formula for calculation of recall and precision may be stated as: Recall = Number of relevant items retrieved X 100<\/p>\n<p>&nbsp;<\/p>\n<p>Total number of relevant items in the collection<\/p>\n<p>&nbsp;<\/p>\n<p>Precision = Number of relevant items retrieved \u00a0X 100 Total number of items retrieved<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Recall thus relates to the ability of the system to retrieve relevant documents, and precision\u00a0<span style=\"font-size: 1em;text-align: initial\">relates to its ability not to retrieve non-relevant documents. The ideal system attempts to achieve 100% recall and 100 % precision, i.e. it attempts to retrieve all the relevant documents and relevant documents only. However, this is not possible in practice because as the level of recall increases, precision tends to decrease. They are inversely proportional. Following example shows the relationship between recall and precision for a given search.<\/span><\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Let us suppose that in a given situation a system retrieves &#8216;a+b&#8217; number of documents, out of which &#8216;a&#8217; documents are relevant, and &#8216;b&#8217; documents are non-relevant. Say, for example, &#8216;c+d&#8217; documents are left in the collection after the search has been conducted. This number will be quite large, because it represents the whole collection minus the retrieved documents. Out of the &#8216;c+d&#8217; number, let\u2019s say, &#8216;c&#8217; documents are relevant to the query but could not be retrieved, and &#8216;d&#8217; documents are not relevant and thus have been correctly rejected. For a large collection the value of &#8216;d&#8217; will be quite large in comparison to c because it represents all the non-relevant documents minus those that have been retrieved wrongly (here b). Lancaster suggests that these statistics can be represented as in the following table.<\/p>\n<p>&nbsp;<\/p>\n<table class=\"aligncenter\" style=\"height: 56px\">\n<tbody>\n<tr style=\"height: 14px\">\n<td style=\"height: 14px;width: 107.063px;text-align: center\"><\/td>\n<td style=\"height: 14px;width: 76.0625px\"><strong>Relevant<\/strong><\/td>\n<td style=\"height: 14px;width: 103.063px\"><strong>Not-Relevant<\/strong><\/td>\n<td style=\"height: 14px;width: 66.0625px\"><strong>Total<\/strong><\/td>\n<\/tr>\n<tr style=\"height: 14px\">\n<td style=\"height: 14px;width: 107.063px\"><strong>Retrieved<\/strong><\/td>\n<td style=\"height: 14px;width: 76.0625px\">a (hits)<\/td>\n<td style=\"height: 14px;width: 103.063px\">b (noise)<\/td>\n<td style=\"height: 14px;width: 66.0625px\">a+b<\/td>\n<\/tr>\n<tr style=\"height: 14px\">\n<td style=\"height: 14px;width: 107.063px\"><strong>Not Retrieved<\/strong><\/td>\n<td style=\"height: 14px;width: 76.0625px\">c (misses)<\/td>\n<td style=\"height: 14px;width: 103.063px\">d (rejected)<\/td>\n<td style=\"height: 14px;width: 66.0625px\">c+d<\/td>\n<\/tr>\n<tr style=\"height: 14px\">\n<td style=\"height: 14px;width: 107.063px\"><strong>Total<\/strong><\/td>\n<td style=\"height: 14px;width: 76.0625px\">a+c<\/td>\n<td style=\"height: 14px;width: 103.063px\">b+d<\/td>\n<td style=\"height: 14px;width: 66.0625px\">a+b+c+d<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p style=\"text-align: center\">Table 1: Recall \u2013 Precision Matrix<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">So as per the above table, &#8216;a&#8217; denotes the \u2018hit\u2019 and &#8216;b&#8217; denotes the \u2018noise\u2019. Now, out of the remaining &#8216;c+d&#8217; documents, the system misses &#8216;c&#8217; documents that should have been retrieved, but it correctly rejects &#8216;d&#8217; documents that are not relevant to the given query. The recall and precision ratio in this case can be calculated as<\/p>\n<p>&nbsp;<\/p>\n<p>R = [ a \/ ( a + c ) ] X 100 P = [ a \/ ( a + b ) ] X 100<\/p>\n<p>&nbsp;<\/p>\n<p>Recently, the theory of the \u2018inverse relationship between precision and recall\u2019 has been questioned by Fugmann (1993). By several examples, he has shown that:<\/p>\n<p>&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 An increasing in precision is by no means always accompanied by a corresponding decrease in recall, and<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 An increase in recall is by no means observed to have always in its wake a decrease in precision.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>4.2\u00a0 Fallout and generality\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">While conducting a search, it is quite likely that some non-relevant items could be retrieved in a given search. This is often termed as the fallout ratio. At the same time, the proportion of relevant documents in the collection for a given query is called the generality ratio. Recall, precision, fallout, and generality ratios have been represented by Salton (1971) as shown in Tables 8.1 and Table 8.2. Thus from Table 8.1, the cut-off can be determined by the following formula:<\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p>Cut-off = [ ( a + b } \/ ( a + b + c + d ) ]<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Van Rijsbergen proposes that recall and precision can be combined in a single measure, called effectiveness, or E, which is a weighted combination of precision and recall where the lower the E value, the greater is the effectiveness. If recall and precision are represented by R and P respectively, the value of E can be measured through the following formula.<\/p>\n<p>&nbsp;<\/p>\n<p>E = 100 X [ 1 &#8211; ( 1 + \u03b22) PR \/ (\u03b22 P + R) ]<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Where \u03b2 is used to reflect the relative importance of recall and precision to the user ( 0 &lt; \u03b2 &lt; \u221e ); \u03b2 = 0.5 corresponds to attaching half as much importance to recall as precision.<\/p>\n<p>&nbsp;<\/p>\n<table class=\"aligncenter\">\n<tbody>\n<tr>\n<td><strong>Symbol<\/strong><\/td>\n<td><strong>Evaluation<\/strong><\/p>\n<p><strong>Measure<\/strong><\/td>\n<td><strong>Formula<\/strong><\/td>\n<td><strong>Explanation<\/strong><\/td>\n<\/tr>\n<tr>\n<td>R<\/td>\n<td>Recall<\/td>\n<td>a\/(a+c)<\/td>\n<td>Proportion of relevant items retrieved<\/td>\n<\/tr>\n<tr>\n<td>P<\/td>\n<td>Precision<\/td>\n<td>a\/(a+b)<\/td>\n<td>Proportion \u00a0of \u00a0retrieved \u00a0items \u00a0that \u00a0are<\/p>\n<p>relevant<\/td>\n<\/tr>\n<tr>\n<td>F<\/td>\n<td>Fallout<\/td>\n<td>b\/(b+d)<\/td>\n<td>Proportion\u00a0\u00a0\u00a0\u00a0 of\u00a0\u00a0\u00a0\u00a0 non-relevant\u00a0\u00a0\u00a0\u00a0 items<\/p>\n<p>retrieved<\/td>\n<\/tr>\n<tr>\n<td>G<\/td>\n<td>Ganerality<\/td>\n<td>(a+c)\/(a+b+c+d)<\/td>\n<td>Proportion of relevant items per query<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p style=\"text-align: center\">Table 2: Retrieval Measures<\/p>\n<p>&nbsp;<\/p>\n<p><strong>4.3\u00a0 <\/strong><strong>Limitations of recall and precision<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong>Number of documents to be retrieved: <\/strong>Different users may want different levels of recall, like a person going to prepare a state-of-the-art-report (SOTAR) on a topic would like to have all the items so the he \/ she will go for a high recall. Whereas, the user wanting to know \u2018something\u2019 about a given topic will prefer to have \u2018a few items\u2019, and thus will not require a high recall. Here, the major problem is that users very often are unable to specify exactly how many items they want to be retrieved.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Recall assumes that all relevant items have the same value, which is not always true. The retrieved items may have different degrees of relevance and this may vary from user to user and even from time to time to the same user. Both recall and precision depend largely on the relevance judgments of the user. The judgement is quite subjective and there may be different degrees of the retrieved output.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>5.\u00a0 Types of Evaluation Experiments\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p>Lancaster (1979) identifies five major steps involved in the evaluation of an information retrieval system, which are<\/p>\n<p>&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 designing the scope of evaluation<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 designing the evaluation program<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 execution of the evaluation<\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">\u2022\u00a0\u00a0\u00a0 analysis and interpretation of results, and<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">\u2022\u00a0\u00a0\u00a0 Modifying the system in the light of the evaluation results.<\/span><\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\"><strong>Step 1<\/strong><\/p>\n<p style=\"text-align: justify\">In the step 1, an evaluation study is conducted to determine the level of performance of the given system. The evaluation would also point out the weaknesses of the system and the reasons for the same. This stage is where planning of evaluation is done, setting its purpose and scope, choosing appropriate methods, costing and staff are all decided.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Step 2<\/strong><\/p>\n<p style=\"text-align: justify\">In step 2 following the\u00a0 \u00a0basic objectives set and the proposed plans, the parameters for data collection are worked out. The evaluator must identify the points on which data are to be collected and a methodology is proposed. Data collection plans are outlined. Also the plans for exploiting data and required manipulations are decided upon.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Step 3<\/strong><\/p>\n<p style=\"text-align: justify\">Step 3 deals with execution of the evaluation. \u00a0This is time consuming. The data collectors have to collect data according to the plan and methods prescribed in the previous stage. Any alterations required to the proposed plan due to constraints at data collection stage must be communicated to the evaluator by the data collection personnel. Mutually they can agree upon required adjustments and execute them.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Step 4<\/strong><\/p>\n<p style=\"text-align: justify\">Step4 is about analysis and interpretation of the data. The success and impact of evaluation mainly depends on the accuracy of the interpretations. Data is manipulated and analysed according to aimed objective and results are obtained. These results are interpreted in the elight of the objectives of the evaluation.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Step 5<\/strong><\/p>\n<p style=\"text-align: justify\">Finally, feedback based on the results and interpretation of evaluation is given to the retrieval system so that it may be modified accordingly.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>5.1\u00a0 Cranfield Tests\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The first extensive evaluation of information retrieval systems was undertaken at Cranfield, UK, under the direction of C W Cleverdon, and is known as the Cranfield 1 project. The first Cranfield Study began in 1957 and was reported by Cleverdon in 1962.The project was designed to compare the effectiveness of four indexing systems, viz.<\/p>\n<p>&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 An alphabetical subject catalogue based on a subject heading list.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 A \u00a0UDC \u00a0classified \u00a0catalogue \u00a0with \u00a0alphabetical \u00a0chain \u00a0index \u00a0to \u00a0the \u00a0class \u00a0headings constructed.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 A catalogue based on a faceted classification and an alphabetical index to the class headings.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 A catalogue compiled by the uniterm coordinate index.<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p><strong><span style=\"text-align: initial;font-size: 1em\">System Parameters\u00a0<\/span><\/strong><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The study was involved 18,000 indexed items and 1200 search topics. The documents, half of which were research reports and half periodical articles, were chosen equally from the general field of aeronautics and the specialized field of high-speed aerodynamics.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Three indexes were chosen \u2013 one with subject knowledge, one with indexing experience and one straight from library school having neither subject background nor indexing experience.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Each indexer was asked to index each source document five times, spending 2, 4, 8, 12 and 16 minutes per document. One hundred source documents thus gave rise to a set of 6000 indexed items (100 documents X 3 indexers X 4 systems X 5 times ). Each of these 6000 items was tested in three places, and therefore the system worked on altogether 18,000 (6000 X 3 phases) indexed items. The test was conducted in three phases with a view to find out whether the level of performance increased with increasing experience of the system personnel.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Significance\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The results of the Cranfield 1 test contradicted the general belief regarding the nature of information retrieval system in many ways. The test proved that the performance of a system does not depend on the experience and subject background of the indexer. It showed that systems where documents are organized by faceted classification scheme perform poorly in comparison to the alphabetical index and uniterm system. It identified the major factors that affect the performance of retrieval systems, and developed for the first time the methodologies that could be applied successfully in evaluating information retrieval systems. Moreover, it also proved that recall and precision is the two most important parameters for determining the performance of information retrieval systems and that these two parameters are related inversely to each other.<\/p>\n<p>&nbsp;<\/p>\n<p>The following findings of Cranfield are significant<\/p>\n<p>&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Indexing times over 4 minutes gave no real improvement in performance<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 A high quality of indexing could be obtained from non-technical indexers<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 The system operated at a recall rate of 70-90% and precision rate of 8 \u2013 20 %<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 A 1 % improvement in precision could be achieved at the cost of 3% loss in recall<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Recall and precision were inversely related to each other<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 All four indexing methods gave a broadly similar performance<\/p>\n<p>&nbsp;<\/p>\n<p><strong>Cranfield test 2\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The second Cranfield test was a controlled experiment that attempted to assess the effects of the components of index languages on the performance of retrieval systems. This study tried to assess the effect by varying each factor, while keeping the others constant. Altogether 29 index languages formed by the combination of concepts were tested on 1400 documents. The test was conducted on a collection of 1400 reports and articles collected from the field of high speed aerodynamics and aircraft structures.<\/p>\n<\/div>\n<div>\n<p>&nbsp;<\/p>\n<p><strong>5.2\u00a0 MEDLARS\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Performance of the Medical Literature Analysis and Retrieval System (MEDLARS) of the US National Library of Medicine was assessed during August 1966 to July 1967. The test was conducted on the operational database of MEDLARS, a database of biomedical articles, index entries being drawn from the MeSH. The objective of the MEDLARS test was to evaluate the existing MEDLARS system and to find out the ways for improvisation. The document collection available on the MEDLARS service at the time of the test consisted of about 7,00,00 items.<\/p>\n<p>&nbsp;<\/p>\n<p>21 user groups were selected from the users community that would &#8211;<\/p>\n<p>&#8211;\u00a0 supply some test questions;<\/p>\n<p>&#8211;\u00a0 cover all kinds of subjects in the requests; and<\/p>\n<p>&#8211;\u00a0 cover all categories of users.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The user group so selected provided 302 search requests. Each query was formulated in terms of MeSH by the system operator and searches were conducted. After completion of a search the sample output was sent to the users for relevant assessment. Photocopies of the articles, rather than mere reference, were supplied for the relevance assessment. The user was asked to mark each retrieved item using the following scales:<\/p>\n<p>&nbsp;<\/p>\n<p>H1 \u2013 of major value; H2 \u2013 of minor value; W1 \u2013 of no value; W2 \u2013 value unknown.<\/p>\n<p>&nbsp;<\/p>\n<p>Precision of the searches were calculated from these figures with the following formula:<\/p>\n<p>Precision ration = ((H1 + H2)\/L) X 100<\/p>\n<p>(Where L is the number of sample items retrieved)<\/p>\n<p>&nbsp;<\/p>\n<p><strong>5.3\u00a0 SMART Retrieval Experiment\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The SMART System was designed in 1964, largely as an experimental tool for the evaluation of the effectiveness of many different types of analysis and search procedures. Salton(1971) characterizes the system through the following steps of its function.<\/p>\n<p>&nbsp;<\/p>\n<p>Overview of the functioning of SMART Retrieval System It is used to<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 take documents and search queries posed in English;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 perform a fully automatic content analysis of texts;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 match analysed search statements and contents of documents;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Retrieve the stored items which are most similar to the queries.<\/p>\n<p>&nbsp;<\/p>\n<p>A number of methods were adopted for automatic content analysis of documents, like \u2013<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">&#8211;\u00a0 Word suffix cut-off methods;<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">&#8211;\u00a0 Thesaurus look-up procedures;<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">&#8211;\u00a0 phrase generation methods;<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">&#8211;\u00a0 Statistical term associations;<\/span><\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">&#8211;\u00a0 Hierarchical term expansion; and so on.<\/span><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p>The following evaluation measures were generated by the SMART System:<\/p>\n<p>&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 A recall \u2013 precision graph reflecting the average precision value at ten discrete recall points &#8211; from a recall of 0.1 to a recall of 1.0 in intervals of 0.1.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Two \u00a0global \u00a0measures, \u00a0known \u00a0as \u00a0normalized \u00a0recall \u00a0and \u00a0normalized \u00a0precision, \u00a0which together reflect the overall performance level of the system.<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Two simplified global measures, known as rank recall and log precision, respectively.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>5.4\u00a0 The Stairs Project\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In 1985, Blair and Maron (1985) published a report on a large scale experiment aimed at evaluating the retrieval effectiveness of a full-text search and retrieval system. This is known as STAIRS (Storage and Information Retrieval System) Study.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The database examined in the STAIRS study consisted of nearly 40,000 documents, representing roughly 350,000 pages of hard copy text used in the defence of a large corporate lawsuit. One important feature of STAIRS was that the lawyers who were to use the system for litigation support stipulated that they must be able to retrieve 75% of all the documents relevant to a given request. The major objective of the STAIRS evaluation was to assess how well the system could retrieve all the documents (and only those) relevant to a given request and measures of recall and precision were used for this purpose.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In STAIRS Project, precision was calculated by dividing the total number of \u2018vital\u2019, \u2018satisfactory\u2019, and \u2018marginally relevant\u2019 documents by the total number of documents retrieved. For the calculation of recall, a sampling technique was adopted. Random samples were taken and these were evaluated by the lawyers. The total number of relevant documents that existed in these subsets was estimated.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">One of the reasons for the failure of STAIRS, as stated was, \u201cIt was impossibly difficult for users to predict the exact words, their combinations and phrases used in all or most of the relevant documents and only in those documents&#8221;.<\/p>\n<p>&nbsp;<\/p>\n<p>Now, we will see the last information retrieval system evaluation experiment in this module, i.e.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>5.5\u00a0 TREC: The Text Retrieval Conference\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In 1991, the US Defence Advanced Research Projects Agency (DARPA), funded the TREC experiments, to be run by the National Institute of Science and Technology (NIST) in order to enable information retrieval researchers to scale up from small collections of data to larger experiments.<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p><span style=\"text-align: initial;font-size: 1em\">The objectives of the TREC experiments have been to<\/span><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Encourage retrieval research based on large test collections<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Increase communication among industry, academia and government<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Speed the transfer of technology from research labs to commercial products<\/p>\n<p>\u2022\u00a0\u00a0\u00a0 Increase the availability of appropriate evaluation techniques for use by industry and academia.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">TREC is an annual workshop series aimed at building the infrastructure required for extracting relevant information from large volumes of electronic documents. The first TREC Conference was held in 1992 and research into improved text-retrieval methodologies has continued ever since. The document collection of TREC is known as the TIPSTER collection. TIPSTER reflects the diversity of subject matter, word choice, literary style, formats, and so on. Its primary collection is in gigabytes with over a million documents.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>6.\u00a0 Summary\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">With this, we have completed the discussion on various aspects of information retrieval evaluation experiments and are now ready to conclude this module by saying that different users may want different levels of recall, like a person going to prepare a state-of-the-art-report (SOTAR) on a topic would like to have all the items so he \/ she will go for a high recall. On the other hand, the user wanting to know \u2018something\u2019 about a given topic will prefer to have \u2018a few items\u2019, and thus will not require a high recall. Here, the major problem is that users very often are unable to specify exactly how many items they want to be retrieved.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Recall assumes that all relevant items have the same value, which is not always true. The retrieved items may have different degrees of relevance and this may vary from user to user and even from time to time to the same user. Both recall and precision depend largely on the relevance judgments of the user. The judgement is quite subjective and there may be different degrees of the retrieved output.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Finally to conclude, one can say that the performance evaluation is crucial at many stages in information retrieval system development. At the end of development process, it is significant to show that the final retrieval system achieves an acceptable level of performance and that it represents a significant improvement over existing retrieval systems. To evaluate a retrieval system, there is need to estimate the future performance of the system.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><strong>7.\u00a0 References\u00a0<\/strong><\/p>\n<p>&nbsp;<\/p>\n<ol>\n<li>Blair, D C and Maron, M E., An evaluation of retrieval effectiveness for a full-text document retrieval system. Communications of the ACM, 28(3), 1985. 289-99 pp.<\/li>\n<li>Cleverdon, \u00a0Cyril \u00a0W. \u00a01967. \u00a0The \u00a0Cranfield \u00a0tests \u00a0on \u00a0index \u00a0language \u00a0devices. \u00a0Aslib Proceedings 19, no. 6:173-194<\/li>\n<\/ol>\n<\/div>\n<ol start=\"3\">\n<li>Cleverdon, C W., Report on the first step of an investigation into comparative efficiency of indexing systems, Cranfield, College of Aeronautics, 1960<\/li>\n<li>Cleverdon, C W., Report on the testing and analysis of an investigation into comparative efficiency of indexing systems, Cranfield, College of Aeronautics, 1962<\/li>\n<li>Fungmann, R., Subject analysis and indexing: theoretical foundation and practical advice by Robert Fungman, Frankfurt, IdeksVerlag, 1993<\/li>\n<li>Keen, E. M., \u2018Evaluation parameters\u2019. In: In: Salta G. (ed.), The SMART Retrieval System: experiments in automatic document processing. Englewood Cliffs, New Jersey: Prentice Hall, 1971, 74-111 pp.<\/li>\n<li>Lancaster, F W. Information Retrieval Systems: Characteristics, testing and evaluation, 2ndedn, New York, John Wiley, 1979.<\/li>\n<li>Salton, Gerald. 1971. Relevance feedback and the optimization of retrieval effectiveness. Ghap \u00a015 \u00a0in \u00a0The \u00a0SMART \u00a0Retrieval \u00a0System \u00a0\u2013 \u00a0Experiments \u00a0in Automatic \u00a0Document processing. Ed. G Salton. Pp. 324-336. Englewood Cliffs, New Jersey: Prentice Hall<\/li>\n<li>Salton, Gerald, \u2018The SMART Project: Status report and plan\u2019. In: Salta G. (ed.), The SMART Retrieval System: experiments in automatic document processing. Englewood Cliffs, New Jersey: Prentice Hall, 1971, 03-11 pp.<\/li>\n<li>Salton, Gerald. 1971. Relevance feedback and the optimization of retrieval effectiveness. Ghap \u00a015 \u00a0in \u00a0The \u00a0SMART \u00a0Retrieval \u00a0System \u00a0\u2013 \u00a0Experiments \u00a0in Automatic \u00a0Document processing. Ed. G Salton. Pp. 324-336. Englewood Cliffs, New Jersey: Prentice Hall<\/li>\n<li>Salton, G. \u2018Another look at automatic text-retrieval systems\u2019, communications of the ACM, 29(7), 1986, 648-656 pp.<\/li>\n<li>Salton, G. and McGill, M. J., Introduction to modern information retrieval, Auckland, McGraw-Hill, 1983.<\/li>\n<li>Swanson, D R. \u2018Some unexplained aspects of the Cranfield tests of indexing language performance\u2019, Library Quarterly, 41, 1971, 223-228 pp.<\/li>\n<li>TREC. Available at <a href=\"http:\/\/trec.nist.gov\/pubs.html\">http:\/\/trec.nist.gov\/pubs.html<\/a><\/li>\n<li>van Rijsbergen, C. J., Information retrieval, 2nd ed., London, Butterworth, 1979<\/li>\n<li>Vickery, B. C. Techniques of information retrieval. London, Butterworth, 1970.<\/li>\n<\/ol>\n","protected":false},"author":4,"menu_order":7,"template":"","meta":{"pb_show_title":"on","pb_short_title":"","pb_subtitle":"","pb_authors":["dr-nanaji-shewale"],"pb_section_license":""},"chapter-type":[],"contributor":[62],"license":[],"class_list":["post-53","chapter","type-chapter","status-publish","hentry","contributor-dr-nanaji-shewale"],"part":3,"_links":{"self":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/pressbooks\/v2\/chapters\/53","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/pressbooks\/v2\/chapters"}],"about":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/wp\/v2\/types\/chapter"}],"author":[{"embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/wp\/v2\/users\/4"}],"version-history":[{"count":4,"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/pressbooks\/v2\/chapters\/53\/revisions"}],"predecessor-version":[{"id":119,"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/pressbooks\/v2\/chapters\/53\/revisions\/119"}],"part":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/pressbooks\/v2\/parts\/3"}],"metadata":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/pressbooks\/v2\/chapters\/53\/metadata\/"}],"wp:attachment":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/wp\/v2\/media?parent=53"}],"wp:term":[{"taxonomy":"chapter-type","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/pressbooks\/v2\/chapter-type?post=53"},{"taxonomy":"contributor","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/wp\/v2\/contributor?post=53"},{"taxonomy":"license","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/lisp7\/wp-json\/wp\/v2\/license?post=53"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}