Linux Operations
Linux High Load Debugging Guide
Interpret Linux load average with runnable tasks, uninterruptible I/O, CPU saturation and cgroup pressure instead of treating load as CPU percentage. This reference is written for developers who need practical validation behavior, reviewable rules and safe examples rather than copied snippets with no explanation.
Recommended workflow
| Step | Why it matters |
|---|---|
| Normalize the signal | Compare load with CPU count, traffic and the incident time window. |
| Separate runnable and blocked work | Use process states and wait channels to distinguish CPU demand from I/O stalls. |
| Inspect pressure | Review CPU, memory and I/O pressure stall information plus cgroup throttling. |
| Trace the owner | Connect hot or blocked tasks to services, deployments and dependent storage. |
Starter snippet
uptime; vmstat 1; ps -eo state,pid,ppid,comm,wchan:32 --sort=stateReview checks
- Capture short time-series samples.
- Check steal time on virtual machines.
- Compare host and container limits.
- Record queue depth before restarting workloads.
Common mistakes
- Assuming load 10 is always severe.
- Sorting only by CPU usage.
- Restarting before identifying blocked kernel waits.
Validation should help users correct input while protecting systems from bad data. Keep syntax checks, product policy, security review and deliverability checks separate.
Related Formalint references
Continue with Linux Admin Command Guide, Linux Disk Space Debugging Guide, Application Health Check Guide.