Home AWSAWS EC2 High Load Average: Linux vs Application vs Infrastructure Causes

AWS EC2 High Load Average: Linux vs Application vs Infrastructure Causes

by Anjali Sindhu
AWS EC2 High Load Average: Linux vs Application vs Infrastructure Causes

A high load average on an AWS EC2 instance can be an early warning that the server is under pressure. Administrators often see a high load value and immediately assume that the CPU is overloaded. However, load average does not directly measure CPU utilization.

A server can have a high load average while CPU usage remains relatively low. This can happen when processes are waiting for disk I/O, memory resources, network operations, or other system resources.

Understanding what load average actually represents is essential for troubleshooting EC2 performance problems. The cause may exist at the Linux operating-system level, inside an application, or within the underlying AWS infrastructure and resource configuration.

What Is Load Average?

On Linux, load average represents the average number of processes that are either actively running or waiting for certain resources, particularly CPU or uninterruptible I/O activity.

You can view the load average with:

up-time

For example:

load average: 8.20, 7.50, 6.90

These three values generally represent the average load over the last 1, 5, and 15 minutes.

The important point is that the number must be interpreted relative to the number of available CPU cores.

An instance with 8 vCPUs can generally handle a larger runnable workload than an instance with 2 vCPUs. Therefore, a load average of 4 means something very different on those two systems.

Load average should always be evaluated alongside CPU usage, I/O wait, memory pressure, and application behavior.

Linux-Level Causes of High Load

The first category to investigate is the operating system itself.

Run:

up-time

and:

top

The top command provides a useful overview of CPU utilization, running processes, memory, and load average.

Pay particular attention to the CPU breakdown. If %wa is high, the system may be spending significant time waiting for I/O rather than executing CPU instructions.

You can also examine the process state:

ps -eo state,pid,cmd | grep ‘^D’

Processes in the D state are typically waiting for an uninterruptible kernel operation, commonly related to I/O.

Disk I/O Pressure

Storage problems are one of the most common reasons a high load average does not correspond to high CPU usage.

Use:

iostat -xz 1

if sysstat is installed.

Look for high device utilization, increased latency, or substantial read/write activity.

Potential causes include:

  • Large backup operations
  • Database workloads
  • Log processing
  • File synchronization
  • Heavy application writes
  • Storage reaching its performance limits

If processes are waiting for storage operations, load average can rise even though the CPU has available capacity.

Memory Pressure

Insufficient memory can also contribute to system slowdowns.

Check memory with:

free -h

and inspect swap:

swap on –show

When physical memory becomes constrained, the system may begin using swap. Heavy swapping can increase I/O activity and cause processes to spend more time waiting.

Memory pressure should therefore be investigated alongside load average rather than treating the load number independently.

Application-Level Causes

The next area to investigate is the application itself.

An application can generate a high load average through excessive requests, inefficient operations, or too many concurrent workers.

For example, a web application may suddenly receive a large increase in traffic. Each request may not consume significant CPU individually, but hundreds or thousands of simultaneous requests can create a substantial queue of active or waiting processes.

Use:

ps aux –sort=-%cpu | head

to identify CPU-intensive processes.

Also examine memory consumption:

ps aux –sort=-%mem | head

The process consuming the most CPU or memory is not necessarily the root cause, but it provides a useful starting point.

Database Workloads

Databases are another common source of high load.

A database server may experience increased load because of:

  • Long-running queries
  • Missing indexes
  • Excessive concurrent connections
  • Large table scans
  • Lock contention
  • Heavy reporting operations
  • Backup activity

A poorly optimized query can cause many processes to wait for disk operations or locks. This may increase load average without producing consistently high CPU utilization.

Review database activity and logs when the load spike occurs. Compare the timing with application requests and scheduled tasks.

Too Many Workers or Processes

Application servers often use multiple workers to handle concurrent requests.

If worker limits are configured too aggressively, an EC2 instance can have many processes competing for memory, CPU, storage, or other resources.

For example, a PHP-FPM configuration with a large number of workers may appear reasonable on a high-memory server but create significant resource pressure on a smaller instance.

Check the number of active processes:

ps -e –no-headers | wc -l

Then identify which services are creating them.

Reducing unnecessary concurrency can sometimes improve performance without changing the EC2 instance type.

AWS Infrastructure Causes

The third category is the AWS infrastructure and instance configuration.

EBS Performance

For instances using Amazon EBS, review the attached volume’s performance characteristics.

Depending on the workload and volume type, the instance may encounter limitations involving IOPS, throughput, or burst-related performance.

A storage bottleneck can cause applications to wait for disk operations, resulting in a high load average.

Review relevant EBS metrics in CloudWatch and compare them with the time of the load spike.

Instance Capacity

An instance may simply be too small for its workload.

For example, an application that previously handled 500 requests per minute may struggle after traffic increases significantly.

Review CPU, memory, network, and storage metrics together before deciding whether the instance needs to be resized.

Scaling the instance can address genuine capacity limitations, but it should not be used as a substitute for fixing an application-level memory leak, inefficient query, or runaway process.

Network Dependencies

An application can also appear slow while waiting for remote services.

For example, an EC2-hosted application may depend on a database, API, cache, or another internal service. Network latency or connectivity problems can cause application processes to remain active while waiting for responses.

Review application logs and request timing to determine whether delays occur locally or between services.

How to Troubleshoot High Load Average

A structured investigation is more effective than reacting to the load number alone.

Start with:

up-time

Then check:

top

Follow with:

free -h

and:

iostat -xz 1

if available.

Next, identify resource-heavy processes:

ps aux –sort=-%cpu | head

ps aux –sort=-%mem | head

Then correlate the findings with CloudWatch metrics for CPU, EBS, network traffic, and other available instance metrics.

The objective is to determine whether processes are primarily running, waiting for storage, waiting for memory, or blocked by an application dependency.

High Load Does Not Always Mean High CPU

One of the most important troubleshooting principles is that load average and CPU utilization measure different aspects of system behavior.

A high load average with high CPU usage can indicate that the instance has a CPU-bound workload.

A high load average with low CPU usage may instead point toward I/O, memory pressure, blocked processes, database activity, or another bottleneck.

This distinction prevents administrators from unnecessarily upgrading CPU resources when the actual problem is storage or memory.

Final Thoughts

High load average on an AWS EC2 instance should be treated as a starting point for investigation, not as a diagnosis.

Linux-level issues such as disk I/O, swap activity, and blocked processes can increase load. Application-level problems such as inefficient queries, excessive workers, and traffic spikes can create additional pressure. Infrastructure limitations involving EBS performance, instance capacity, or network dependencies can also contribute.

Facing issues?

Our technical support
engineers can solve it.

Contact Us today!
guy server checkup

You may also like

Leave a Comment