When a production server suddenly kills a process, the first assumption is often simple: the server ran out of memory.
That assumption can be misleading.
Linux can invoke the Out-Of-Memory (OOM) killer even when monitoring dashboards show apparently available RAM. The reason is that Linux memory management is more complicated than looking at one number such as free memory. Allocation constraints, memory zones, cgroups, fragmentation, reserved memory, and kernel behavior can all contribute to an OOM event.
Recently, I investigated an OOM incident where the server appeared to have enough RAM. The kernel logs told a different story.
The First Clue: The Kernel Log
The investigation started with the usual commands:
dmesg -T | grep -i -E “oom|out of memory|killed process”
On systems using systemd, the same information can often be found with:
journalctl -k | grep -i -E “oom|out of memory|killed process”
The important part of the output looked similar to:
Out of memory: Killed process 18472 (java)
total-vm:12582912kB, anon-rss:7340032kB
At first glance, this appears straightforward: Java consumed too much memory.
But the kernel log contains much more information than just the name of the killed process. It can reveal the state of the system at the moment the allocation failed.
That distinction matters.
RAM Usage Is Not the Same as Allocation Availability
One of the biggest mistakes during OOM investigations is treating total free RAM as the only relevant metric.
Linux separates memory into several categories, including anonymous memory, page cache, kernel memory, slab allocations, and memory belonging to different zones.
A server might report:
MemTotal: 32768000 kB
MemFree: 1200000 kB
MemAvailable: 6500000 kB
It would be tempting to conclude that the machine has several gigabytes available and therefore cannot possibly experience an OOM.
But an individual allocation may still fail.
Linux does not simply ask, “How much RAM is free?” It must determine whether the required pages can be allocated under the current memory-management constraints.
Understanding the OOM Killer
The OOM killer is a recovery mechanism.
When the kernel cannot satisfy a memory allocation, it attempts to reclaim memory. It may reclaim page cache, swap pages, or other reclaimable memory. If those efforts are insufficient, the kernel can select a process to terminate.
The selected process is not necessarily the process that caused the original memory pressure.
This is a critical point.
Suppose a database process creates sustained memory pressure and a web application happens to have a high OOM score. The web application may ultimately be killed even though the database generated much of the pressure.
Therefore, this message:
Killed process 18472 (java)
does not automatically mean:
Java caused the OOM.
It means:
The kernel selected Java as a victim when it needed to recover memory.
The Importance of cgroups
Another common explanation is that the server itself wasn’t out of memory—the workload was.
Modern Linux systems frequently use cgroups to impose memory limits on services and containers.
For example, a host might have 64 GB of RAM while a container is limited to 4 GB.
From the host’s perspective:
64 GB RAM
may be perfectly healthy.
Inside the container, however:
Memory limit: 4 GB
Memory usage: 4 GB
can trigger an OOM event.
This is particularly common with Docker, Kubernetes, systemd services, and other workload-management systems.
When investigating an OOM event, always determine whether the kernel reported a global OOM or a cgroup-related OOM.
For systemd services, checking the configured memory limits can help:
systemctl show my-service | grep -i memory
For containers, inspect the container’s configured memory limits and current usage.
Don’t Ignore Swap
Another misconception is that having swap automatically prevents OOM conditions.
Swap can provide additional virtual memory, but it isn’t unlimited emergency RAM.
Check the current situation with:
free -h
swapon –show
If swap is disabled or already heavily utilized, the kernel has fewer options when physical memory becomes constrained.
However, enabling swap blindly isn’t a complete solution. Excessive swapping can introduce severe latency and I/O pressure, particularly for databases and latency-sensitive applications.
Swap should be considered part of the memory-management strategy rather than an automatic cure for every OOM event.
Look at What Happened Before the Kill
Kernel logs tell us what happened, but historical system metrics tell us why.
Useful metrics include:
- RAM usage
- MemAvailable
- swap usage
- major page faults
- disk I/O
- process RSS
- container memory usage
- cgroup memory events
- application-level memory metrics
Tools such as vmstat, sar, pidstat, and Prometheus-based monitoring can help reconstruct the minutes leading up to the event.
For example:
vmstat 1
can provide a quick view of memory pressure, swapping, and system activity.
Process-level information can be inspected with:
ps aux –sort=-%mem | head
But remember that a snapshot taken after the OOM event may be misleading. The offending process may already have been killed, leaving little evidence behind.
Check the OOM Score
Linux maintains an OOM score for processes that influences victim selection.
You can inspect a process with:
cat /proc/<PID>/oom_score
and its adjustment value with:
cat /proc/<PID>/oom_score_adj
A higher effective OOM score generally makes a process more likely to be selected.
This can explain why a seemingly important service disappeared while another memory-hungry process survived.
It also explains why process priority and memory consumption should be considered separately.
Conclusion
The most important lesson from this incident was simple:
An OOM event does not necessarily mean the server had zero free RAM.
The kernel makes allocation decisions based on a much broader picture: reclaimable memory, allocation constraints, zones, swap, cgroups, and process characteristics.
A reliable investigation therefore follows the evidence instead of starting with the assumption that “RAM was full.”
Start with the kernel log. Identify the exact OOM event. Determine whether it was global or constrained to a cgroup. Examine the selected process, its OOM score, memory limits, swap state, and historical resource metrics.
Most importantly, distinguish between the process that was killed and the process or condition that created the memory pressure.
That distinction can turn an apparently mysterious production failure into a measurable, explainable Linux memory-management problem.
The next time a monitoring dashboard says, “There was still RAM available,” don’t dismiss the OOM event.

