In my architecture course we split a single-cycle RISC datapath into the classic five stages: fetch, decode, execute, memory and write-back. The stage delays are 250, 150, 200, 300 and 100 ps, 1000 ps in total, and I expected the pipelined version to run five times faster since five instructions are in progress at once.
The model answer gives a speedup below three. Where does the rest go, and how do hazards enter the calculation?




