Linux Operations
Linux OOM Killer Debugging Guide
Investigate Linux out-of-memory kills with kernel evidence, cgroup limits, working-set growth, swap behavior and application allocation signals. This reference is written for developers who need practical validation behavior, reviewable rules and safe examples rather than copied snippets with no explanation.
Recommended workflow
| Step | Why it matters |
|---|---|
| Prove the kill | Use kernel messages and exit status to distinguish OOM from application termination. |
| Identify the boundary | Check host memory, container cgroup limits and systemd MemoryMax separately. |
| Measure the working set | Compare resident memory, cache, anonymous pages and growth over time. |
| Connect to application behavior | Correlate deploys, load, queues and allocation profiles with the pressure window. |
Starter snippet
journalctl -k -b | grep -Ei 'out of memory|killed process|oom'Review checks
- Record memory.current and memory.events for cgroups.
- Keep emergency operating-system headroom.
- Set bounded queues and caches.
- Validate alerting before the kill threshold.
Common mistakes
- Adding swap as the only fix.
- Looking only at free memory after the process died.
- Raising limits without finding unbounded growth.
Validation should help users correct input while protecting systems from bad data. Keep syntax checks, product policy, security review and deliverability checks separate.
Related Formalint references
Continue with Docker Container Logs Guide, Java Memory Debugging Guide, Prometheus Alert Rule Debugging.