{"id":139,"date":"2018-07-18T10:50:19","date_gmt":"2018-07-18T10:50:19","guid":{"rendered":"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/?post_type=chapter&#038;p=139"},"modified":"2018-08-03T09:02:50","modified_gmt":"2018-08-03T09:02:50","slug":"pipelining-mips-implementation","status":"publish","type":"chapter","link":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/chapter\/pipelining-mips-implementation\/","title":{"rendered":"Pipelining \u2013 MIPS Implementation"},"content":{"raw":"&nbsp;\r\n<p style=\"text-align: justify\">The objectives of this module are to discuss the basics of pipelining and discuss the implementation of the MIPS pipeline.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">In the previous module, we discussed the drawbacks of a single cycle implementation. We observed that the longest delay determines the clock period and it is not feasible to vary period for different instructions. This violates the design principle of making the common case fast. One way of overcoming this problem is to go in for a pipelined implementation. We shall now discuss the basics of pipelining. Pipelining is a particularly effective way of organizing parallel activity in a computer system. The basic idea is very simple. It is frequently encountered in manufacturing plants, where pipelining is commonly known as an assembly line operation. By laying the production process out in an assembly line, products at various stages can be worked on simultaneously. You must have noticed that in an automobile assembly line, you will find that one car\u2019s chassis will be fitted when some other car\u2019s door is getting fixed and some other car\u2019s body is getting painted. All these are independent activities, taking place in parallel. This process is also referred to as pipelining, because, as in a pipeline, new inputs are accepted at one end and previously accepted inputs appear as outputs at the other end. As yet another real world example, Consider the case of doing a laundry. Assume that Ann, Brian, Cathy and Daveeach have one load of clothes to wash, dry, and fold and that the washer takes 30 minutes, dryer takes 40 minutes and the folder takes 20 minutes. Sequential laundry takes 6 hours for 4 loads. On the other hand, if they learned pipelining, how long would the laundry take? It takes only 3.5 hours for 4 loads! For four loads, you get a Speedup = 6\/3.5 = 1.7. If you work the washing machine non-stop, you get a Speedup = 110n\/40n + 70 \u2248 3 = number of stages.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">To apply the concept of instruction execution in pipeline, it is required to break the instruction execution into different tasks. Each task will be executed in different processing elements of the CPU. As we know that there are two distinct phases of instruction execution: one is instruction fetch and the other one is instruction execution. Therefore, the processor executes a program by fetching and executing instructions, one after another. The cycle time \u03c4 of an instruction pipeline is the time needed to advance a set of instructions one stage through the pipeline. The cycle time can be determined as<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-141 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-79.png\" alt=\"\" width=\"294\" height=\"28\" \/>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">where \u03c4m = maximum stage delay (delay through the stage which experiences the largest delay) , <em>k<\/em> = number of stages in the instruction pipeline, <em>d<\/em> = the time delay of a\u00a0<span style=\"text-align: initial;font-size: 1em\">latch needed to advance signals and data from one stage to the next. Now suppose that <\/span><em style=\"text-align: initial;font-size: 1em\">n<\/em><span style=\"text-align: initial;font-size: 1em\"> instructions are processed and these instructions are executed one after another. The total time required Tk to execute all <\/span><em style=\"text-align: initial;font-size: 1em\">n<\/em><span style=\"text-align: initial;font-size: 1em\"> instructions is<\/span><\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-143 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-80.png\" alt=\"\" width=\"173\" height=\"30\" \/>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">In general, let the instruction execution be divided into five stages as fetch, decode, execute, memory access and write back, denoted by Fi, Di, Ei, Mi and Wi. Execution of a program consists of a sequence of these steps. When the first instruction\u2019s decode happens, the second instruction\u2019s fetch is done. When the pipeline is filled, you see that there are five different activities taking place in parallel. All these activities are overlapped. Five instructions are in progress at any given time. This means that five distinct hardware units are needed. These units must be capable of performing their tasks simultaneously and without interfering with one another. Information is passed from one unit to the next through a storage buffer. As an instruction progresses through the pipeline, all the information needed by the stages downstream must be passed along.<\/p>\r\n&nbsp;\r\n\r\nIf all stages are balanced, i.e., all take the same time,\r\n\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-144 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-81.png\" alt=\"\" width=\"549\" height=\"85\" \/>\r\n<div>\r\n<p style=\"text-align: justify\">If the stages are not balanced, speedup will be less. Observe that the speedup is due to increased throughput and the latency (time for each instruction) does not decrease.<\/p>\r\n&nbsp;\r\n\r\nThe basic features of pipelining are:\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">\u2022 Pipelining does not help latency of single task, it only helps throughput of entire workload<\/p>\r\n\u2022 Pipeline rate is limited by the slowest pipeline stage\r\n\r\n\u2022 Multiple tasks operate simultaneously\r\n\r\n\u2022 It exploits parallelism among instructions in a sequential instruction stream\r\n\r\n\u2022 Unbalanced lengths of pipe stages reduces speedup\r\n<p style=\"text-align: justify\">\u2022 Time to \u201cfill\u201d pipeline and time to \u201cdrain\u201d it reduces speedup<\/p>\r\n<p style=\"text-align: justify\">\u2022 Ideally the speedup is equal to the number of stages and the CPI is 1<\/p>\r\n&nbsp;\r\n\r\nLet us consider the MIPS pipeline with five stages, with one step per stage:\r\n\r\n&nbsp;\r\n\r\n\u2022 IF: Instruction fetch from memory\r\n\r\n\u2022 ID: Instruction decode &amp; register read\r\n\r\n\u2022 EX: Execute operation or calculate address\r\n\r\n\u2022 MEM: Access memory operand\r\n\r\n\u2022 WB: Write result back to register\r\n\r\n<\/div>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Consider the details given in Figure 10.1. Assume that it takes 100ps for a register read or write and 200ps for all other stages. Let us calculate the speedup obtained by pipelining.<\/p>\r\n\r\n<\/div>\r\n<img class=\"size-full wp-image-145 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-82.png\" alt=\"\" width=\"672\" height=\"259\" \/>\r\n<div>\r\n\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-146 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-83.png\" alt=\"\" width=\"624\" height=\"428\" \/>\r\n<p style=\"text-align: center\"><strong>Figure 10.2<\/strong><\/p>\r\n\r\n<\/div>\r\n&nbsp;\r\n<p style=\"text-align: justify\">For a non pipelined implementation it takes 800ps for each instruction and for a pipelined implementation it takes only 200ps.<\/p>\r\n&nbsp;\r\n\r\nObserve that the MIPS ISA is designed in such a way that it is suitable for pipelining.\r\n\r\n&nbsp;\r\n\r\nFigure 10.3 shows the MIPS pipeline implementation.\r\n\r\n&nbsp;\r\n\r\n\u2013\u00a0\u00a0 All instructions are 32-bits\r\n<ul>\r\n \t<li>Easier to fetch and decode in one cycle<\/li>\r\n \t<li>Comparatively, the x86 ISA: 1- to 17-byte instructions<\/li>\r\n<\/ul>\r\n\u2013\u00a0\u00a0 Few and regular instruction formats\r\n<ul>\r\n \t<li>Can decode and read registers in one step<\/li>\r\n<\/ul>\r\n\u2013\u00a0\u00a0 Load\/store addressing\r\n<ul>\r\n \t<li>Can calculate address in 3rd stage, access memory in 4th stage<\/li>\r\n<\/ul>\r\n\u2013\u00a0\u00a0 Alignment of memory operands\r\n<ul>\r\n \t<li>Memory access takes only one cycle<\/li>\r\n<\/ul>\r\n<img class=\"size-full wp-image-147 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-84.png\" alt=\"\" width=\"649\" height=\"466\" \/>\r\n\r\n<img class=\"size-full wp-image-148 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-85.png\" alt=\"\" width=\"638\" height=\"358\" \/>\r\n<div>\r\n\r\n&nbsp;\r\n<p style=\"text-align: justify\">Figure 10.4 shows how buffers are introduced between the stages. This is mandatory. Each stage takes in data from that buffer, processes it and write into the next buffer. Also note that as an instruction moves down the pipeline from one buffer to the next, its relevant information also moves along with it. For example, during clock cycle 4, the information in the buffers is as follows:<\/p>\r\n&nbsp;\r\n<ul>\r\n \t<li style=\"text-align: justify\">Buffer IF\/ID holds instruction I4, which was fetched in cycle 4<\/li>\r\n \t<li style=\"text-align: justify\">Buffer ID\/EX holds the decoded instruction and both the source operands for instruction I3. This is the information produced by the decoding hardware in cycle 3.<\/li>\r\n \t<li style=\"text-align: justify\">Buffer EX\/MEM holds the executed result of I2. The buffer also holds the information needed for the write step of instruction I2. Even though it is not needed by the execution stage, this information must be passed on to the next stage and further down to the Write back stage in the following clock cycle to enable that stage to perform the required Write operation.<\/li>\r\n \t<li style=\"text-align: justify\">Buffer MEM\/WB holds the data fetched from memory (for a load) for I1, and for the arithmetic and logical operations, the results produced by the execution unit and the destination information for instruction I1 are just passed.<\/li>\r\n<\/ul>\r\n<p style=\"text-align: justify\"><\/p>\r\n<p style=\"text-align: justify\">We shall look at the single-clock-cycle diagrams for the load &amp; store instructions of the MIPS ISA. Figure 10.4 shows the instruction fetch for a load \/ store instruction. Observe that the PC is used to fetch the instruction, it is written into the IF\/ID buffer and the PC is incremented by 4. Figure 10.5 shows the next stage of ID. The instruction is decoded, the register file is read and the operands are written into the ID\/EX buffer. Note that the entire information of the instruction including the destination register is written into the ID\/EX buffer. The highlights in the figure show the resources involved. Figure 10.6 shows the execution stage. The base register\u2019s contents and the sign extended displacement are fed to the ALU, the addition operation is initiated and the ALU\u00a0<span style=\"text-align: initial;font-size: 1em\">calculates the memory address. This effective address is stored in the EX\/MEM buffer. Also the destination register\u2019s information is passed from the ID\/EX buffer to the EX\/MEM buffer. Next, the memory access happens and the read data is written into the MEM\/WB buffer. The destination register\u2019s information is passed from the EX\/MEM buffer to the MEM\/WB buffer. This is illustrated in Figure 10.7. The write back happens in the last stage. The data read from the data memory is written into the destination register specified in the instruction. This is shown in Figure 10.8. The destination register information is passed on from the MEM\/WB memory backwards to the register file, along with the data to be written. The datapath is shown in Figure 10.9.<\/span><\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-149 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-86.png\" alt=\"\" width=\"630\" height=\"356\" \/>\r\n\r\n<\/div>\r\n<img class=\"size-full wp-image-150 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-87.png\" alt=\"\" width=\"677\" height=\"371\" \/>\r\n\r\n<img class=\"size-full wp-image-151 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-88.png\" alt=\"\" width=\"644\" height=\"370\" \/>\r\n\r\n<img class=\"size-full wp-image-152 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-89.png\" alt=\"\" width=\"688\" height=\"393\" \/>\r\n\r\n<img class=\"size-full wp-image-153 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-90.png\" alt=\"\" width=\"668\" height=\"366\" \/>\r\n\r\n<img class=\"size-full wp-image-154 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-91.png\" alt=\"\" width=\"649\" height=\"340\" \/>\r\n<p style=\"text-align: justify\">For a store instruction, the effective address calculation is the same as that of load. But when it comes to the memory access stage, store performs a memory write. The effective address is passed on from the execution stage to the memory stage, the data read from the register file is passed from the ID\/EX buffer to the EX\/MEM buffer and taken from there. The store instruction completes with this memory stage. There is no write back for the store instruction.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">While discussing the cycle-by-cycle flow of instructions through the pipelined datapath, we can look at the following options:<\/p>\r\n&nbsp;\r\n<ul>\r\n \t<li>\u00a0\u201cSingle-clock-cycle\u201d pipeline diagram<\/li>\r\n \t<li>Shows pipeline usage in a single cycle<\/li>\r\n \t<li>Highlight resources used<\/li>\r\n \t<li>o \u201cmulti-clock-cycle\u201d diagram<\/li>\r\n \t<li>Graph of operation over time<\/li>\r\n<\/ul>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The multi-clock-cycle pipeline diagram showing the resource utilization is given in Figure 10.10. It can be seen that the Instruction memory is used in eth first stage, The register file is used in the second stage, the ALU in the third stage, the data memory in the fourth stage and the register file in the fifth stage again.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-155 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-92.png\" alt=\"\" width=\"565\" height=\"524\" \/>\r\n<p style=\"text-align: center\"><strong>Figure 10.11<\/strong><\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">The multi-cycle diagram showing the activities happening in each clock cycle is given in Figure 10.11.<\/p>\r\n&nbsp;\r\n<p style=\"text-align: justify\">Now, having discussed the pipelined implementation of the MIPS architecture, we need to discuss the generation of control signals. The pipelined implementation of MIPS, along with the control signals is given in Figure 10.12.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-156 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-93.png\" alt=\"\" width=\"715\" height=\"393\" \/>\r\n<p style=\"text-align: justify\">All the control signals indicated are not required at the same time. Different control signals are required at different stages of the pipeline. But the decision about the generation of the various control signals is done at the second stage, when the instruction is decoded. Therefore, just as the data flows from one stage to another as the instruction moves from one stage to another, the control signals also pass on from one buffer to another and are utilized at the appropriate instants. This is shown in Figure 10.13. The control signals for the execution stage are used in that stage. The control signals needed for the memory stage and the write back stage move along with that instruction to the next stage. The memory related control signals are used in the next stage, whereas, the write back related control signals move from there to the next stage and used when the instruction performs the write back operation.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-157 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-94.png\" alt=\"\" width=\"646\" height=\"418\" \/>\r\n<p style=\"text-align: justify\">The complete pipeline implementation, along with the control signals used at the various stages is given in Figure 10.14.<\/p>\r\n&nbsp;\r\n\r\n<img class=\"size-full wp-image-158 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-95.png\" alt=\"\" width=\"667\" height=\"379\" \/>\r\n<p style=\"text-align: justify\">To summarize, we have discussed the basics of pipelining in this module. We have made the following observations about pipelining.<\/p>\r\n\r\n<ul>\r\n \t<li style=\"text-align: justify\">Pipelining is overlapped execution of instructions<\/li>\r\n \t<li>Latency is the same, but throughput\u00a0 improves<\/li>\r\n \t<li style=\"text-align: justify\">Pipeline rate limited by slowest pipeline stage<\/li>\r\n \t<li>Potential speedup = Number of pipe stages<\/li>\r\n<\/ul>\r\n&nbsp;\r\n<p style=\"text-align: justify\">We have discussed about the implementation of pipelining in the MIPS architecture. We have shown the implementation of the various buffers, the data flow and the control flow for a pipelined implementation of the MIPS architecture.<\/p>\r\n\r\n<div>\r\n\r\n&nbsp;\r\n\r\n<strong>Web Links \/ Supporting Materials<\/strong>\r\n<ul>\r\n \t<li>Computer Organization and Design \u2013 The Hardware \/ Software Interface, David A. Patterson and John L. Hennessy, 4th.Edition, Morgan Kaufmann, Elsevier, 2009.<\/li>\r\n \t<li>Computer Organization, Carl Hamacher, Zvonko Vranesic and Safwat Zaky, 5th.Edition, McGraw- Hill Higher Education, 2011.<\/li>\r\n<\/ul>\r\n<\/div>","rendered":"<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The objectives of this module are to discuss the basics of pipelining and discuss the implementation of the MIPS pipeline.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In the previous module, we discussed the drawbacks of a single cycle implementation. We observed that the longest delay determines the clock period and it is not feasible to vary period for different instructions. This violates the design principle of making the common case fast. One way of overcoming this problem is to go in for a pipelined implementation. We shall now discuss the basics of pipelining. Pipelining is a particularly effective way of organizing parallel activity in a computer system. The basic idea is very simple. It is frequently encountered in manufacturing plants, where pipelining is commonly known as an assembly line operation. By laying the production process out in an assembly line, products at various stages can be worked on simultaneously. You must have noticed that in an automobile assembly line, you will find that one car\u2019s chassis will be fitted when some other car\u2019s door is getting fixed and some other car\u2019s body is getting painted. All these are independent activities, taking place in parallel. This process is also referred to as pipelining, because, as in a pipeline, new inputs are accepted at one end and previously accepted inputs appear as outputs at the other end. As yet another real world example, Consider the case of doing a laundry. Assume that Ann, Brian, Cathy and Daveeach have one load of clothes to wash, dry, and fold and that the washer takes 30 minutes, dryer takes 40 minutes and the folder takes 20 minutes. Sequential laundry takes 6 hours for 4 loads. On the other hand, if they learned pipelining, how long would the laundry take? It takes only 3.5 hours for 4 loads! For four loads, you get a Speedup = 6\/3.5 = 1.7. If you work the washing machine non-stop, you get a Speedup = 110n\/40n + 70 \u2248 3 = number of stages.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">To apply the concept of instruction execution in pipeline, it is required to break the instruction execution into different tasks. Each task will be executed in different processing elements of the CPU. As we know that there are two distinct phases of instruction execution: one is instruction fetch and the other one is instruction execution. Therefore, the processor executes a program by fetching and executing instructions, one after another. The cycle time \u03c4 of an instruction pipeline is the time needed to advance a set of instructions one stage through the pipeline. The cycle time can be determined as<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-141 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-79.png\" alt=\"\" width=\"294\" height=\"28\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-79.png 294w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-79-65x6.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-79-225x21.png 225w\" sizes=\"auto, (max-width: 294px) 100vw, 294px\" \/><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">where \u03c4m = maximum stage delay (delay through the stage which experiences the largest delay) , <em>k<\/em> = number of stages in the instruction pipeline, <em>d<\/em> = the time delay of a\u00a0<span style=\"text-align: initial;font-size: 1em\">latch needed to advance signals and data from one stage to the next. Now suppose that <\/span><em style=\"text-align: initial;font-size: 1em\">n<\/em><span style=\"text-align: initial;font-size: 1em\"> instructions are processed and these instructions are executed one after another. The total time required Tk to execute all <\/span><em style=\"text-align: initial;font-size: 1em\">n<\/em><span style=\"text-align: initial;font-size: 1em\"> instructions is<\/span><\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-143 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-80.png\" alt=\"\" width=\"173\" height=\"30\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-80.png 173w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-80-65x11.png 65w\" sizes=\"auto, (max-width: 173px) 100vw, 173px\" \/><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">In general, let the instruction execution be divided into five stages as fetch, decode, execute, memory access and write back, denoted by Fi, Di, Ei, Mi and Wi. Execution of a program consists of a sequence of these steps. When the first instruction\u2019s decode happens, the second instruction\u2019s fetch is done. When the pipeline is filled, you see that there are five different activities taking place in parallel. All these activities are overlapped. Five instructions are in progress at any given time. This means that five distinct hardware units are needed. These units must be capable of performing their tasks simultaneously and without interfering with one another. Information is passed from one unit to the next through a storage buffer. As an instruction progresses through the pipeline, all the information needed by the stages downstream must be passed along.<\/p>\n<p>&nbsp;<\/p>\n<p>If all stages are balanced, i.e., all take the same time,<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-144 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-81.png\" alt=\"\" width=\"549\" height=\"85\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-81.png 549w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-81-300x46.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-81-65x10.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-81-225x35.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-81-350x54.png 350w\" sizes=\"auto, (max-width: 549px) 100vw, 549px\" \/><\/p>\n<div>\n<p style=\"text-align: justify\">If the stages are not balanced, speedup will be less. Observe that the speedup is due to increased throughput and the latency (time for each instruction) does not decrease.<\/p>\n<p>&nbsp;<\/p>\n<p>The basic features of pipelining are:<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">\u2022 Pipelining does not help latency of single task, it only helps throughput of entire workload<\/p>\n<p>\u2022 Pipeline rate is limited by the slowest pipeline stage<\/p>\n<p>\u2022 Multiple tasks operate simultaneously<\/p>\n<p>\u2022 It exploits parallelism among instructions in a sequential instruction stream<\/p>\n<p>\u2022 Unbalanced lengths of pipe stages reduces speedup<\/p>\n<p style=\"text-align: justify\">\u2022 Time to \u201cfill\u201d pipeline and time to \u201cdrain\u201d it reduces speedup<\/p>\n<p style=\"text-align: justify\">\u2022 Ideally the speedup is equal to the number of stages and the CPI is 1<\/p>\n<p>&nbsp;<\/p>\n<p>Let us consider the MIPS pipeline with five stages, with one step per stage:<\/p>\n<p>&nbsp;<\/p>\n<p>\u2022 IF: Instruction fetch from memory<\/p>\n<p>\u2022 ID: Instruction decode &amp; register read<\/p>\n<p>\u2022 EX: Execute operation or calculate address<\/p>\n<p>\u2022 MEM: Access memory operand<\/p>\n<p>\u2022 WB: Write result back to register<\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Consider the details given in Figure 10.1. Assume that it takes 100ps for a register read or write and 200ps for all other stages. Let us calculate the speedup obtained by pipelining.<\/p>\n<\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-145 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-82.png\" alt=\"\" width=\"672\" height=\"259\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-82.png 672w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-82-300x116.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-82-65x25.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-82-225x87.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-82-350x135.png 350w\" sizes=\"auto, (max-width: 672px) 100vw, 672px\" \/><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-146 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-83.png\" alt=\"\" width=\"624\" height=\"428\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-83.png 624w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-83-300x206.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-83-65x45.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-83-225x154.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-83-350x240.png 350w\" sizes=\"auto, (max-width: 624px) 100vw, 624px\" \/><\/p>\n<p style=\"text-align: center\"><strong>Figure 10.2<\/strong><\/p>\n<\/div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">For a non pipelined implementation it takes 800ps for each instruction and for a pipelined implementation it takes only 200ps.<\/p>\n<p>&nbsp;<\/p>\n<p>Observe that the MIPS ISA is designed in such a way that it is suitable for pipelining.<\/p>\n<p>&nbsp;<\/p>\n<p>Figure 10.3 shows the MIPS pipeline implementation.<\/p>\n<p>&nbsp;<\/p>\n<p>\u2013\u00a0\u00a0 All instructions are 32-bits<\/p>\n<ul>\n<li>Easier to fetch and decode in one cycle<\/li>\n<li>Comparatively, the x86 ISA: 1- to 17-byte instructions<\/li>\n<\/ul>\n<p>\u2013\u00a0\u00a0 Few and regular instruction formats<\/p>\n<ul>\n<li>Can decode and read registers in one step<\/li>\n<\/ul>\n<p>\u2013\u00a0\u00a0 Load\/store addressing<\/p>\n<ul>\n<li>Can calculate address in 3rd stage, access memory in 4th stage<\/li>\n<\/ul>\n<p>\u2013\u00a0\u00a0 Alignment of memory operands<\/p>\n<ul>\n<li>Memory access takes only one cycle<\/li>\n<\/ul>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-147 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-84.png\" alt=\"\" width=\"649\" height=\"466\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-84.png 649w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-84-300x215.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-84-65x47.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-84-225x162.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-84-350x251.png 350w\" sizes=\"auto, (max-width: 649px) 100vw, 649px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-148 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-85.png\" alt=\"\" width=\"638\" height=\"358\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-85.png 638w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-85-300x168.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-85-65x36.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-85-225x126.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-85-350x196.png 350w\" sizes=\"auto, (max-width: 638px) 100vw, 638px\" \/><\/p>\n<div>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Figure 10.4 shows how buffers are introduced between the stages. This is mandatory. Each stage takes in data from that buffer, processes it and write into the next buffer. Also note that as an instruction moves down the pipeline from one buffer to the next, its relevant information also moves along with it. For example, during clock cycle 4, the information in the buffers is as follows:<\/p>\n<p>&nbsp;<\/p>\n<ul>\n<li style=\"text-align: justify\">Buffer IF\/ID holds instruction I4, which was fetched in cycle 4<\/li>\n<li style=\"text-align: justify\">Buffer ID\/EX holds the decoded instruction and both the source operands for instruction I3. This is the information produced by the decoding hardware in cycle 3.<\/li>\n<li style=\"text-align: justify\">Buffer EX\/MEM holds the executed result of I2. The buffer also holds the information needed for the write step of instruction I2. Even though it is not needed by the execution stage, this information must be passed on to the next stage and further down to the Write back stage in the following clock cycle to enable that stage to perform the required Write operation.<\/li>\n<li style=\"text-align: justify\">Buffer MEM\/WB holds the data fetched from memory (for a load) for I1, and for the arithmetic and logical operations, the results produced by the execution unit and the destination information for instruction I1 are just passed.<\/li>\n<\/ul>\n<p style=\"text-align: justify\">\n<p style=\"text-align: justify\">We shall look at the single-clock-cycle diagrams for the load &amp; store instructions of the MIPS ISA. Figure 10.4 shows the instruction fetch for a load \/ store instruction. Observe that the PC is used to fetch the instruction, it is written into the IF\/ID buffer and the PC is incremented by 4. Figure 10.5 shows the next stage of ID. The instruction is decoded, the register file is read and the operands are written into the ID\/EX buffer. Note that the entire information of the instruction including the destination register is written into the ID\/EX buffer. The highlights in the figure show the resources involved. Figure 10.6 shows the execution stage. The base register\u2019s contents and the sign extended displacement are fed to the ALU, the addition operation is initiated and the ALU\u00a0<span style=\"text-align: initial;font-size: 1em\">calculates the memory address. This effective address is stored in the EX\/MEM buffer. Also the destination register\u2019s information is passed from the ID\/EX buffer to the EX\/MEM buffer. Next, the memory access happens and the read data is written into the MEM\/WB buffer. The destination register\u2019s information is passed from the EX\/MEM buffer to the MEM\/WB buffer. This is illustrated in Figure 10.7. The write back happens in the last stage. The data read from the data memory is written into the destination register specified in the instruction. This is shown in Figure 10.8. The destination register information is passed on from the MEM\/WB memory backwards to the register file, along with the data to be written. The datapath is shown in Figure 10.9.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-149 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-86.png\" alt=\"\" width=\"630\" height=\"356\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-86.png 630w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-86-300x170.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-86-65x37.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-86-225x127.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-86-350x198.png 350w\" sizes=\"auto, (max-width: 630px) 100vw, 630px\" \/><\/p>\n<\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-150 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-87.png\" alt=\"\" width=\"677\" height=\"371\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-87.png 677w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-87-300x164.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-87-65x36.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-87-225x123.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-87-350x192.png 350w\" sizes=\"auto, (max-width: 677px) 100vw, 677px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-151 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-88.png\" alt=\"\" width=\"644\" height=\"370\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-88.png 644w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-88-300x172.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-88-65x37.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-88-225x129.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-88-350x201.png 350w\" sizes=\"auto, (max-width: 644px) 100vw, 644px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-152 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-89.png\" alt=\"\" width=\"688\" height=\"393\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-89.png 688w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-89-300x171.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-89-65x37.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-89-225x129.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-89-350x200.png 350w\" sizes=\"auto, (max-width: 688px) 100vw, 688px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-153 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-90.png\" alt=\"\" width=\"668\" height=\"366\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-90.png 668w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-90-300x164.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-90-65x36.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-90-225x123.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-90-350x192.png 350w\" sizes=\"auto, (max-width: 668px) 100vw, 668px\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-154 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-91.png\" alt=\"\" width=\"649\" height=\"340\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-91.png 649w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-91-300x157.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-91-65x34.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-91-225x118.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-91-350x183.png 350w\" sizes=\"auto, (max-width: 649px) 100vw, 649px\" \/><\/p>\n<p style=\"text-align: justify\">For a store instruction, the effective address calculation is the same as that of load. But when it comes to the memory access stage, store performs a memory write. The effective address is passed on from the execution stage to the memory stage, the data read from the register file is passed from the ID\/EX buffer to the EX\/MEM buffer and taken from there. The store instruction completes with this memory stage. There is no write back for the store instruction.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">While discussing the cycle-by-cycle flow of instructions through the pipelined datapath, we can look at the following options:<\/p>\n<p>&nbsp;<\/p>\n<ul>\n<li>\u00a0\u201cSingle-clock-cycle\u201d pipeline diagram<\/li>\n<li>Shows pipeline usage in a single cycle<\/li>\n<li>Highlight resources used<\/li>\n<li>o \u201cmulti-clock-cycle\u201d diagram<\/li>\n<li>Graph of operation over time<\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The multi-clock-cycle pipeline diagram showing the resource utilization is given in Figure 10.10. It can be seen that the Instruction memory is used in eth first stage, The register file is used in the second stage, the ALU in the third stage, the data memory in the fourth stage and the register file in the fifth stage again.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-155 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-92.png\" alt=\"\" width=\"565\" height=\"524\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-92.png 565w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-92-300x278.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-92-65x60.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-92-225x209.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-92-350x325.png 350w\" sizes=\"auto, (max-width: 565px) 100vw, 565px\" \/><\/p>\n<p style=\"text-align: center\"><strong>Figure 10.11<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">The multi-cycle diagram showing the activities happening in each clock cycle is given in Figure 10.11.<\/p>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">Now, having discussed the pipelined implementation of the MIPS architecture, we need to discuss the generation of control signals. The pipelined implementation of MIPS, along with the control signals is given in Figure 10.12.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-156 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-93.png\" alt=\"\" width=\"715\" height=\"393\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-93.png 715w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-93-300x165.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-93-65x36.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-93-225x124.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-93-350x192.png 350w\" sizes=\"auto, (max-width: 715px) 100vw, 715px\" \/><\/p>\n<p style=\"text-align: justify\">All the control signals indicated are not required at the same time. Different control signals are required at different stages of the pipeline. But the decision about the generation of the various control signals is done at the second stage, when the instruction is decoded. Therefore, just as the data flows from one stage to another as the instruction moves from one stage to another, the control signals also pass on from one buffer to another and are utilized at the appropriate instants. This is shown in Figure 10.13. The control signals for the execution stage are used in that stage. The control signals needed for the memory stage and the write back stage move along with that instruction to the next stage. The memory related control signals are used in the next stage, whereas, the write back related control signals move from there to the next stage and used when the instruction performs the write back operation.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-157 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-94.png\" alt=\"\" width=\"646\" height=\"418\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-94.png 646w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-94-300x194.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-94-65x42.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-94-225x146.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-94-350x226.png 350w\" sizes=\"auto, (max-width: 646px) 100vw, 646px\" \/><\/p>\n<p style=\"text-align: justify\">The complete pipeline implementation, along with the control signals used at the various stages is given in Figure 10.14.<\/p>\n<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-158 aligncenter\" src=\"http:\/\/csp2.epgpbooks.inflibnet.ac.in\/wp-content\/uploads\/sites\/46\/2018\/07\/2-95.png\" alt=\"\" width=\"667\" height=\"379\" srcset=\"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-95.png 667w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-95-300x170.png 300w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-95-65x37.png 65w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-95-225x128.png 225w, https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-content\/uploads\/sites\/46\/2018\/07\/2-95-350x199.png 350w\" sizes=\"auto, (max-width: 667px) 100vw, 667px\" \/><\/p>\n<p style=\"text-align: justify\">To summarize, we have discussed the basics of pipelining in this module. We have made the following observations about pipelining.<\/p>\n<ul>\n<li style=\"text-align: justify\">Pipelining is overlapped execution of instructions<\/li>\n<li>Latency is the same, but throughput\u00a0 improves<\/li>\n<li style=\"text-align: justify\">Pipeline rate limited by slowest pipeline stage<\/li>\n<li>Potential speedup = Number of pipe stages<\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify\">We have discussed about the implementation of pipelining in the MIPS architecture. We have shown the implementation of the various buffers, the data flow and the control flow for a pipelined implementation of the MIPS architecture.<\/p>\n<div>\n<p>&nbsp;<\/p>\n<p><strong>Web Links \/ Supporting Materials<\/strong><\/p>\n<ul>\n<li>Computer Organization and Design \u2013 The Hardware \/ Software Interface, David A. Patterson and John L. Hennessy, 4th.Edition, Morgan Kaufmann, Elsevier, 2009.<\/li>\n<li>Computer Organization, Carl Hamacher, Zvonko Vranesic and Safwat Zaky, 5th.Edition, McGraw- Hill Higher Education, 2011.<\/li>\n<\/ul>\n<\/div>\n","protected":false},"author":2,"menu_order":10,"template":"","meta":{"_acf_changed":false,"pb_show_title":"on","pb_short_title":"","pb_subtitle":"","pb_authors":["dr-a-p-shanthi"],"pb_section_license":""},"chapter-type":[],"contributor":[58],"license":[],"class_list":["post-139","chapter","type-chapter","status-publish","hentry","contributor-dr-a-p-shanthi"],"part":3,"_links":{"self":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/pressbooks\/v2\/chapters\/139","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/pressbooks\/v2\/chapters"}],"about":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/wp\/v2\/types\/chapter"}],"author":[{"embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/wp\/v2\/users\/2"}],"version-history":[{"count":4,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/pressbooks\/v2\/chapters\/139\/revisions"}],"predecessor-version":[{"id":447,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/pressbooks\/v2\/chapters\/139\/revisions\/447"}],"part":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/pressbooks\/v2\/parts\/3"}],"metadata":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/pressbooks\/v2\/chapters\/139\/metadata\/"}],"wp:attachment":[{"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/wp\/v2\/media?parent=139"}],"wp:term":[{"taxonomy":"chapter-type","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/pressbooks\/v2\/chapter-type?post=139"},{"taxonomy":"contributor","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/wp\/v2\/contributor?post=139"},{"taxonomy":"license","embeddable":true,"href":"https:\/\/ebooks.inflibnet.ac.in\/csp2\/wp-json\/wp\/v2\/license?post=139"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}