Performance is a property of a workload on a system under stated conditions. Before changing code, define the workload, the outcome that matters, and the environment in which you will measure it. Otherwise a benchmark can reward a change that does not help real users.
Pick a metric that matches the question
Throughput counts completed work per unit time. Latency measures the time for an operation. Utilization describes how busy a resource is; saturation describes queued demand that cannot be served immediately. A service can have moderate average latency but unacceptable tail latency for a fraction of requests, so inspect distributions and percentiles rather than only a mean.
Find the bottleneck
Start with an end-to-end measurement, then use profiles and traces to locate where time is spent. A CPU profile points to sampled execution time; an allocation profile exposes memory churn; a distributed trace shows time across service boundaries. A high CPU percentage alone does not identify the slow code, and a database query that looks expensive in isolation may not dominate a user request.
Change one cause and repeat
Control input size, warm-up, concurrency, caching, and background load. Compare before and after with repeated samples and report uncertainty. A faster microbenchmark can make the full system slower if it increases allocations, network calls, contention, or operational complexity.
For memory-bound loops, the CS:APP Cache Lab demonstrates how cache misses change runtime. For database work, use query plans and actual rows rather than assuming an index is always faster.