Computer Engineering DISCUSSION

Why does a five-stage pipeline not make a processor five times faster than single-cycle?

Started by sarmadrasheed pipeline speedupdata hazardcontrol hazardforwardingcycles per instruction
5 replies 248 views 6 participants
Latest activity · 30 Sep 2026

Why does a five-stage pipeline not make a processor five times faster than single-cycle?

sarmadrasheed Computer Engineering Forum
#1

In my architecture course we split a single-cycle RISC datapath into the classic five stages: fetch, decode, execute, memory and write-back. The stage delays are 250, 150, 200, 300 and 100 ps, 1000 ps in total, and I expected the pipelined version to run five times faster since five instructions are in progress at once.

The model answer gives a speedup below three. Where does the rest go, and how do hazards enter the calculation?

Community replies 5

Re: Why does a five-stage pipeline not make a processor five times faster than single-cycle?

#2

The first loss is in the clock period. Every stage gets the same clock, so the period is set by the slowest stage, not the average: 300 ps for the memory stage. The pipeline registers between stages add their own setup and clock-to-output delay; with 20 ps of register overhead the period is 320 ps.

The best possible speedup is therefore 1000 / 320 = 3.1, not 5. You only reach a factor equal to the number of stages if the stages are perfectly balanced and the registers are free. Here the 100 ps write-back stage wastes two thirds of its cycle.

Re: Why does a five-stage pipeline not make a processor five times faster than single-cycle?

#3

The second loss is that the pipeline cannot always finish one instruction per cycle. Data hazards arise when an instruction needs a result that an earlier one has not yet written back. Forwarding paths from the later stages to the execute-stage inputs remove most of them at no time cost. The exception is a load followed immediately by an instruction that uses the loaded value: the data only exists at the end of the memory stage, one cycle too late, so the pipeline inserts one bubble.

Control hazards come from branches. If the branch is resolved in the execute stage, the two instructions fetched behind it are wrong whenever the prediction was wrong and must be flushed: a two-cycle penalty.

Re: Why does a five-stage pipeline not make a processor five times faster than single-cycle?

#4

Put those into cycles per instruction. Suppose 20 percent of instructions are loads and 40 percent of those are followed by a dependent instruction: 0.20 × 0.40 × 1 = 0.08 stall cycles per instruction. Suppose 15 percent are branches, 30 percent of them mispredicted, at 2 cycles each: 0.15 × 0.30 × 2 = 0.09. The average is CPI = 1 + 0.08 + 0.09 = 1.17.

Time per instruction is then 1.17 × 320 ps = 374 ps, and the speedup over the 1000 ps single-cycle design is 1000 / 374 = 2.7. The percentages are illustrative, and your course will have its own, but the method is the same: speedup = single-cycle period / (pipelined period × CPI).

Re: Why does a five-stage pipeline not make a processor five times faster than single-cycle?

#5

One thing the number hides: pipelining improves throughput, not the latency of an individual instruction. Each instruction now takes 5 × 320 = 1600 ps from fetch to write-back, longer than the 1000 ps it took before. The gain comes entirely from overlapping instructions.

The five-stage design also depends on a few structural choices to avoid resource conflicts: separate instruction and data memories (or caches) so fetch and memory access can happen in the same cycle, and a register file that writes in the first half of the cycle and reads in the second, so a value being written back is visible to the instruction in decode.

Re: Why does a five-stage pipeline not make a processor five times faster than single-cycle?

#6

The same arithmetic explains why real processors did not keep adding stages indefinitely. Cutting the logic into more, shorter stages shortens the clock period, but the register overhead per stage stays fixed and becomes a larger share of each cycle, and the branch misprediction penalty grows with the distance between fetch and branch resolution. High-performance cores with pipelines of around 15 to 20 stages lose that many cycles on a mispredict, which is why they invest so heavily in branch predictors that are right well over 95 percent of the time.

TEP COMMUNITY