The model
Profiling attributes resource use to parts of a running program; benchmarking compares defined workloads under controlled conditions. A useful optimization starts with a hypothesis about the bottleneck and gathers evidence at the level where the cost occurs.
A concrete walk-through
If a CPU profile shows most samples inside JSON parsing, tuning database indexes will not fix the measured CPU bottleneck. A benchmark should use representative data, warm-up behavior, stable machine conditions, and enough repetitions to show variation rather than one favorable run.
Costs and failure cases
Instrumentation adds overhead and can distort timing, so use the least intrusive tool that answers the question. Microbenchmarks isolate operations but may fail to predict full-system behavior. Preserve correctness checks and compare end-to-end latency after any local speedup.
Check your understanding
A new implementation is twice as fast in a microbenchmark but the endpoint latency is unchanged. Name three possible reasons and the next measurement you would collect.