A production server can look completely healthy from the outside and still be experiencing a serious kernel problem.
The server may still respond to some network requests. SSH might work intermittently. Monitoring may report that the machine is online. Yet applications can become extremely slow, CPUs can remain busy, and kernel log messages can reveal that a CPU has been stuck executing kernel code for an unusually long period.
This can be confusing during an incident because a soft lockup is not the same thing as a server crash or complete system hang.
The key is to understand what Linux is actually reporting.
The First Clue: “CPU Stuck”
The investigation usually starts with the kernel log.
You might see something similar to:
watchdog: BUG: soft lockup – CPU#3 stuck for 23s! [kworker/3:1:1234]
Another useful command is:
dmesg -T | grep -i “soft lockup”
On systems using systemd:
journalctl -k | grep -i “soft lockup”
The message tells us several important things.
First, the kernel watchdog detected that a CPU had not made sufficient scheduling progress. Second, it identifies the CPU involved. Finally, it usually identifies the task that was running when the watchdog detected the problem.
But the process shown in the message isn’t necessarily the root cause.
That distinction is important.
What Is a Soft Lockup?
A soft lockup occurs when a CPU spends too much time executing kernel code without allowing normal scheduling activity to occur.
Linux has watchdog mechanisms that periodically check whether CPUs are making progress. If a CPU appears to be stuck for longer than the configured threshold, the kernel reports a soft lockup.
The system hasn’t necessarily crashed.
In fact, many soft-lockup incidents leave the server partially functional.
You may still be able to:
- Ping the server
- Connect through SSH
- Read some files
- Access certain applications
- Collect monitoring data
At the same time, other workloads can experience severe delays.
That’s why the phrase “the server is down” can be misleading during this type of incident.
The machine may be alive, but the kernel is not making progress normally.
Don’t Blame the Reported Process Immediately
Consider a message like:
soft lockup – CPU#2 stuck for 31s!
[kworker/2:2:987]
It is tempting to conclude that kworker caused the problem.
That’s not necessarily true.
kworker threads execute deferred kernel work on behalf of many kernel subsystems. A kworker thread appearing in the message can indicate that some kernel operation was taking too long, but it doesn’t immediately identify the underlying cause.
Possible causes include:
- A problematic kernel module
- Driver issues
- Excessive kernel workload
- Storage problems
- Network driver activity
- CPU starvation
- Hardware problems
- Virtualization-related contention
- Kernel bugs
The stack trace following the warning is often much more useful than the process name itself.
Read the Stack Trace
A soft-lockup report commonly includes a kernel stack trace.
For example:
Modules linked in: …
CPU: 2 PID: 987 Comm: kworker/2:2
Call Trace:
<IRQ>
…
some_kernel_function
another_kernel_function
…
The call trace provides clues about where the CPU was spending its time.
Look for patterns.
If the trace repeatedly points into a storage driver, investigate storage hardware, firmware, and the relevant kernel module.
If it points into networking code, investigate network drivers, packet processing, interrupts, and traffic levels.
If virtualization-related functions appear repeatedly, examine CPU scheduling and host-level contention.
The goal isn’t simply to identify the function at the top of the stack. You want to understand why the CPU remained there for so long.
Check CPU Pressure
A soft lockup can sometimes be associated with extreme CPU pressure.
Start with:
uptime
Then inspect CPU activity:
top
or:
mpstat -P ALL 1
Pay particular attention to whether one CPU is behaving differently from the others.
A system with 32 CPUs might show normal average utilization while one CPU is effectively saturated.
That can be hidden by aggregate CPU graphs.
Per-CPU monitoring is therefore important when investigating kernel lockups.
Interrupts Matter Too
Kernel work is frequently triggered by hardware interrupts and deferred interrupt processing.
Inspect interrupt activity with:
cat /proc/interrupts
Look for unusually high interrupt counts or an imbalance between CPUs.
Network-intensive systems, storage workloads, and poorly distributed interrupt processing can create significant CPU pressure.
Tools such as mpstat, sar, and perform can provide additional visibility.
The important question is:
Was the CPU busy doing useful application work, or was it spending an unusual amount of time handling kernel activity?
Investigate Hardware and Drivers
Not every soft lockup is caused by software.
Hardware instability can produce symptoms that eventually appear as kernel problems.
Check the kernel log for related messages:
dmesg -T | grep -i -E “error|fail|timeout|reset|hardware”
Storage devices might report timeouts or resets.
Network devices might report link or driver errors.
PCIe devices may produce hardware-related warnings.
If the soft lockup occurs repeatedly around the same device or driver, that correlation is worth investigating.
Also check whether the issue started after a kernel, firmware, driver, or hardware change.
Virtual Machines Need a Different Perspective
A virtual machine can experience symptoms that look like kernel lockups even when the guest kernel isn’t the original source of the problem.
CPU contention on the virtualization host can affect guest scheduling.
This is why virtualized environments should be investigated at two levels:
Inside the guest:
top
mpstat -P ALL 1
dmesg -T
At the host or hypervisor level:
- CPU contention
- Steal time
- Oversubscription
- Host hardware errors
- Other noisy workloads
For a VM, guest-level CPU usage alone doesn’t always tell the whole story.
Check Whether It Is Reproducible
A single soft-lockup warning is different from hundreds occurring every few minutes.
Look at the timeline:
journalctl -k –since “1 hour ago”
Determine whether the warnings occur:
- During backups
- During high network traffic
- During storage-intensive workloads
- After application deployments
- After kernel updates
- At a particular time of day
Correlation is extremely useful.
If every lockup happens when a particular workload starts, that gives you a much stronger investigation path than simply knowing that a CPU became stuck.
Conclusion
A Linux soft lockup doesn’t necessarily mean the server is dead.
It means the kernel watchdog detected that a CPU failed to make expected scheduling progress for too long.
The correct response is therefore not immediately to reboot the machine and move on.
Start with the kernel logs. Capture the complete stack trace. Identify the affected CPU and task. Examine per-CPU utilization, interrupts, drivers, storage, networking, virtualization, and recent system changes.
Most importantly, distinguish between the task reported by the watchdog and the underlying condition that caused the CPU to become stuck.

