The model
A processor pipeline overlaps stages of multiple instructions, much like an assembly line. Pipelining aims to increase instruction throughput; it does not necessarily reduce the latency of one instruction. The benefit depends on keeping stages supplied with independent work and resolving control decisions quickly.
A concrete walk-through
If instruction B needs a value that instruction A has not produced yet, B has a data dependency. Forwarding may deliver a result directly between stages; otherwise the processor may insert a stall. A branch creates a control dependency, so the pipeline may speculate on a path and later discard work if the prediction was wrong.
Costs and failure cases
Pipeline diagrams are simplified models: real processors can issue multiple instructions, execute out of order, and retire results in program order. Dependencies constrain parallelism even when many functional units are available. A cache miss can leave dependent instructions waiting for data.
Check your understanding
Consider a sequence where each instruction adds the result of the previous one. Why does a wide processor not necessarily execute the whole sequence at once? What change could expose independent work?