Home LinuxDiagnosing OOM Events When There Is No Swap

Diagnosing OOM Events When There Is No Swap

by Anjali Sindhu
Diagnosing OOM Events When There Is No Swap

Running a Linux system without swap can provide predictable performance, especially in environments where disk-based memory is undesirable. However, removing swap also means that the operating system has fewer options when physical RAM becomes exhausted. Once available memory reaches a critical level, Linux may trigger an Out-Of-Memory (OOM) event and terminate one or more processes to recover memory.

An unexpected OOM event can be confusing because the application that gets terminated is not necessarily the only process responsible for memory pressure. Diagnosing the problem requires looking at system memory, individual processes, kernel messages, and the timing of the failure. This article explains a practical approach to investigating OOM events on systems that operate without swap.

Understanding What Happens During an OOM Event

Every running process requires memory for its code, data, libraries, buffers, and other resources. The Linux kernel normally manages this memory dynamically. When available RAM becomes critically low, the kernel attempts to reclaim memory through mechanisms such as dropping filesystem caches.

If memory cannot be reclaimed quickly enough and there is no swap space available, the kernel may reach a point where it cannot satisfy new memory allocation requests. This is when the OOM mechanism can become active.

The kernel evaluates running processes and selects a process to terminate according to several factors. Therefore, it is not accurate to assume that the process using the largest amount of RAM will automatically be killed. The OOM decision depends on the process’s memory usage, its adjustment settings, and other kernel considerations.

Start by Checking Memory Availability

The first step in troubleshooting is to determine whether the machine actually ran out of physical memory.

A simple command such as:

free -h

provides an overview of total, used, available, and cached memory.

Pay particular attention to the available value rather than simply looking at the amount labelled “free.” Linux uses unused RAM for filesystem caching, and that memory can often be reclaimed when applications need it.

You can also inspect detailed memory information with:

cat /proc/meminfo

This exposes values such as MemTotal, MemAvailable, Cached, Buffers, and other memory-related statistics. Comparing these values over time can help reveal whether memory consumption gradually increased before the OOM event.

Examine Kernel Messages

Kernel logs are among the most useful sources of information after an OOM event. The kernel normally records details about the memory shortage and the process selected for termination.

On systems using systemd, try:

journalctl -k

You can search for OOM-related entries with:

journalctl -k | grep -i oom

Another useful command is:

dmesg | grep -i -E ‘oom|out of memory|killed process’

Depending on the Linux distribution and logging configuration, the exact messages may differ.

These records can reveal when the OOM condition occurred, which process was terminated, and sometimes how much memory different processes were consuming. This information is considerably more useful than simply checking which application disappeared.

Identify the Memory-Hungry Processes

After confirming that an OOM event occurred, the next step is to identify applications that were consuming significant amounts of memory.

The top command provides a live overview:

top

For a more convenient interface, htop can also be useful if it is installed.

Look at columns such as RES, which represents resident memory currently held in RAM. A process with continuously increasing resident memory may indicate a memory leak or uncontrolled workload growth.

However, memory usage should be interpreted carefully. Some memory may be shared between processes, meaning that simply adding the displayed memory values can produce misleading totals.

Check for Gradual Memory Growth

Not every OOM event is caused by a sudden spike. In many cases, an application slowly consumes additional memory until the machine eventually reaches its limit.

Monitoring tools can help identify this pattern. Even a basic periodic command can provide useful evidence:

watch -n 5 free -h

For production systems, dedicated monitoring solutions can record memory utilization over hours or days. Historical graphs are particularly valuable because they show whether memory usage is steadily climbing, fluctuating with workload, or suddenly increasing.

A gradual upward trend may point toward a memory leak, an expanding cache, excessive concurrency, or a workload that requires more resources than the server provides.

Investigate Containers and Services

Modern Linux systems often run applications inside containers. In such environments, an OOM event may be related to a container memory limit rather than the total physical RAM of the host.

For Docker-based systems, inspect container statistics with:

docker stats

Container orchestration platforms can also impose memory limits on individual workloads. A container may therefore encounter an OOM condition even when the host still has some available memory.

This distinction is important. Always determine whether the OOM event occurred at the host level, inside a container, or within another resource-controlled environment.

Examine the OOM Score

Linux uses an OOM scoring mechanism when selecting processes during memory pressure. The value can be inspected through the process information directory.

For example:

cat /proc/<PID>/oom_score

There is also an adjustment value:

cat /proc/<PID>/oom_score_adj

These values can influence how likely a process is to be selected during an OOM situation. This is useful when investigating why a particular process was terminated instead of another process with apparently similar memory usage.

Look Beyond RAM Usage

A complete investigation should not focus exclusively on one application’s RAM consumption. Memory pressure can involve several components, including filesystem caches, shared memory, kernel memory, memory mappings, and multiple applications consuming moderate amounts of RAM simultaneously.

It is also useful to check whether a service experienced a workload change shortly before the incident. A sudden increase in users, requests, database activity, background jobs, or file processing can explain why memory requirements increased.

Preventing Future OOM Events

Once the cause has been identified, prevention becomes the next priority. Possible solutions include reducing application memory usage, fixing memory leaks, limiting concurrency, increasing physical RAM, or configuring appropriate resource limits.

If the system is intentionally designed without swap, monitoring becomes even more important because there is less room for temporary memory pressure. Alerts should be configured before RAM reaches a critical level so administrators have an opportunity to investigate.

Conclusion

Diagnosing an OOM event on a Linux system without swap requires looking beyond the process that was terminated. The killed process is often only the final result of overall memory pressure, which may have been caused by a gradual memory leak, sudden workload increase, excessive concurrency, container limits, or insufficient physical RAM. Reviewing free, /proc/meminfo, kernel logs, process memory usage, and container statistics can help identify what happened before and during the OOM event.

Once the underlying cause is identified, administrators can take appropriate preventive measures, such as optimizing application memory usage, fixing memory leaks, adjusting workloads or resource limits, increasing physical memory, and improving monitoring. On systems intentionally configured without swap, proactive memory monitoring and alerts are especially important because there is less capacity to absorb temporary increases in memory demand. A combination of system-level and application-level analysis can help prevent recurring OOM events and maintain reliable system performance.

Facing issues?

Our technical support
engineers can solve it.

Contact Us today!
guy server checkup

You may also like

Leave a Comment