This trace follows the actual state transitions behind the companion Why Memory Locality Matters. It describes a common execution path; implementation details can vary, so keep the contract separate from the mechanism.
Step 1: Access an address and cache line
Processors use a hierarchy of storage because small, nearby storage can be accessed faster than large, distant storage. A cache keeps copies of recently or predictably used memory blocks. Programs benefit when accesses show temporal locality, reusing data, or spatial locality, touching nearby addresses.
Step 2: Check the nearest cache
A row-major matrix traversal that increments the inner column index typically consumes adjacent elements, while stepping through a column may jump by an entire row stride. The values are mathematically identical, but the second access pattern may fetch cache lines that contain mostly unused data.
Step 3: Fetch on a miss
A miss fetches a block from a slower level, so contiguous access can make useful neighboring values arrive together; a large stride may waste most of each fetched line.
At this point, record the state that changed and check the invariant before advancing. If the operation repeats, make clear which values persist and which are recomputed.
Step 4: Choose a line to replace
A cache hit avoids a slower level; a miss fetches a block and may evict another useful block. Capacity, associativity, and line size shape conflict behavior. Caches improve typical access time but make performance depend on working-set size and access order.
Step 5: Reuse nearby or recent data
Rewrite a matrix transpose loop so reads are contiguous. Then reason about the writes: why can optimizing one side still leave the other side with poor locality?
The trace is complete when the result satisfies the stated contract. Compare this model with the concrete runtime or system you are studying before making a performance claim.