Journal

Metric menus for distributed workloads

Earth from space

A metric menu is a deliberate short list — usually seven or fewer gauges — that a team agrees will drive decisions for a quarter. Everything else stays available for forensics but does not appear on the shared board.

For edge workloads we often recommend: edge queue time, origin handler time, cache hit ratio by route family, error rate by fault domain, saturation of the hottest worker pool, sampling drop ratio, and a single user checkpoint delay. That set covers visibility without inviting every engineer to pin a personal favorite.

Menus expire. After a major topology change — new PoP, new CDN partner, new payment worker — schedule a thirty-minute menu review. Retire one metric for every metric you add. The Visibility Lab treats this as a ritual, not a bureaucracy.

If finance asks why ingest costs rose, show the sampling drop ratio beside the rare-path capture rate. Budget conversations stay calmer when both numbers sit on the same pastel card.

← All articles