What Happens to Your Website When Your VPS Runs Out of RAM?
It's not a switch that flips at 100% — it's a chain of kernel reclaim, swapping, and OOM decisions that can kill MySQL while your server stays pingable.
“Server ran out of RAM” is one of those phrases that gets thrown around without much precision. What actually happens depends on which kernel mechanism kicks in first, what your OOM killer decides to sacrifice, and whether you're on Linux or Windows. This post walks through the mechanics, not the folklore.
The wrong mental model
The common assumption looks like this:
VPS has 4 GB RAM → applications consume 4 GB → RAM hits 100% → server crashes.
That's not how Linux memory management works, and it's not really how Windows works either. Both operating systems treat "free RAM" as wasted RAM. If you have unused memory sitting idle, the kernel will use it for disk caching, because reading from cache is faster than reading from disk. Seeing Mem: 3.8Gi used / 4.0Gi total tells you almost nothing about whether you're in trouble.
What "used" memory actually means on Linux
Run free -h on any Ubuntu or AlmaLinux box that's been up for more than an hour and you'll see most of the RAM reported as used, with a chunk of that under buff/cache. That cache is reclaimable. It's file data the kernel has kept around opportunistically — MySQL data files, PHP opcode-adjacent reads, static assets Apache served — in case something asks for them again.
The real sequence when memory gets tight is:
Applications request more memory.
Linux has been using spare RAM for page cache.
As allocations increase, memory pressure rises.
The kernel reclaims cache pages first, because they cost nothing to drop (they can be re-read from disk).
If reclaiming cache isn't enough, inactive anonymous memory may get pushed to swap, if swap exists.
Processes attempting fresh allocations start stalling while the kernel scrambles to free pages.
If none of that keeps up, the OOM killer picks a process and kills it.
Notice how many steps happen before anything actually dies. A lot of "my server is out of RAM" reports are really just step 3 or 4 — the system is fine, it's just using memory the way it's designed to.
kswapd and direct reclaim
Linux has background kernel threads, most commonly referenced as kswapd, whose job is to keep a buffer of free memory above certain watermarks (min, low, high). When free memory drops below the low watermark, kswapd wakes up and starts scanning for reclaimable pages — cache first, then less-recently-used anonymous pages if swap is configured.
If kswapd can't reclaim fast enough — because the rate of new allocations outpaces background reclamation — a process trying to allocate memory can be forced into direct reclaim. That means the process itself pauses mid-syscall while the kernel synchronously frees pages on its behalf. This is the point where your website gets slow. Requests that normally complete in 50ms start taking 2-3 seconds, not because the CPU is busy, but because processes are blocked waiting on memory reclamation rather than executing.
This is worth repeating because it's counterintuitive: the slowdown you feel usually happens before any process is killed. A gradually degrading response time under load is often a memory pressure symptom, not a CPU or network one.
Memory pressure causes latency
Imagine PHP asks for another page of memory. Under contention, the practical sequence looks like this:
Linux: "I don't currently have an easy page free. Give me a moment while I find something reclaimable."
PHP waits. Another process makes a request — it waits too. MySQL needs memory for a sort buffer — it waits as well. Nobody is deadlocked, everybody is just queued behind the same reclaim work.
Requests that previously took 80ms can suddenly take 500ms, 2 seconds, 8 seconds, or in bad cases 30 seconds. The system isn't technically dead — CPU might even look idle-ish, since processes are blocked on reclaim rather than computing. But from the outside:
the website feels painfully slow
the API starts timing out
the WordPress admin hangs on save
even SSH sessions feel delayed
Linux exposes this exact phenomenon through Pressure Stall Information (PSI). You can read it directly:
cat /proc/pressure/memoryPSI measures how much time tasks spend stalled on a resource — CPU, memory or I/O — rather than just how much of that resource is allocated. The kernel documentation for PSI explicitly ties sustained memory pressure to latency spikes, throughput loss, thrashing and, eventually, OOM kills. For diagnosing "why does the site feel slow right now", PSI is arguably a better early-warning signal than a bare RAM percentage, because it tells you time lost to contention rather than bytes allocated.
Then comes swap
Suppose your VPS has 8 GB RAM and 4 GB of swap configured. Linux can move memory pages that aren't actively needed out of RAM and onto disk, freeing up physical RAM for something that needs it more urgently right now. Conceptually:
RAM (before)
PHP page A
MySQL page B
inactive Python page
unused process page
hot application page
↓ swap out inactive pages ↓
RAM (after) SSD / swap
PHP page A inactive Python page
MySQL page B unused process page
FREE
FREE
hot application pageNow another application can use those freed physical RAM pages. How aggressively this happens is influenced by the kernel's vm.swappiness setting — a low value tells Linux to prefer reclaiming cache and only swap as a last resort; a higher value makes it more willing to swap out anonymous memory earlier.
Swap saves availability — but destroys performance if heavily used
RAM latency and NVMe storage latency are radically different, and on a VPS the storage backing swap is typically network-attached SSD over iSCSI rather than a local drive. Even very fast NVMe cannot behave like DRAM for random, small, latency-sensitive accesses. So once the active working set exceeds physical RAM, you get this pattern:
application needs page
↓
page is in swap
↓
read from SSD
↓
put into RAM
↓
perhaps evict another page
↓
application continuesIf the page that just got evicted is needed again immediately afterwards, you're into repeated read-evict-read cycles:
read another page
↓
evict something
↓
read it back
↓
evict something
↓
read it backThis is swap thrashing. The machine can appear completely frozen — load average climbing, nothing responding — despite technically still running and technically not out of memory.
Systemd's own OOM documentation notes that swap buys the system more time to react to memory pressure, whereas a machine with no swap at all can hit a severe out-of-memory state far more abruptly, with less warning. So the blanket rule of "never put swap on a VPS" is overly simplistic. A small amount of swap can be a genuinely useful emergency buffer that keeps things limping along long enough for you (or an autoscaler) to react. It's just not a substitute for having enough RAM in the first place, and heavy sustained swap use is itself a performance incident, not a fix.
Eventually Linux reaches real OOM
Eventually the kernel can reach a state where none of the above buys any more room:
physical RAM is effectively exhausted
useful reclaimable cache is exhausted
swap is exhausted, or unavailable
a memory allocation genuinely cannot be satisfied
At that point Linux has two options: let the entire machine sit stuck indefinitely trying to satisfy allocations it can't fulfil, or kill something and recover its memory immediately. It chooses the second. That's the OOM killer, and it's worth walking through what it actually picks and why — which the next section covers using a concrete stack.
Walking through a real stack: Apache, PHP-FPM, MySQL, Redis on Ubuntu 24.04, 8 GB RAM
Stage 1 — normal operation
Active memory: ~5 GB. Cache: ~2 GB. Available: ~3 GB. Swap: barely touched.
Apache is accepting connections, PHP-FPM workers are running requests, MySQL's InnoDB buffer pool is holding hot table data in RAM, and Linux is using the remaining headroom for filesystem cache. Nothing unusual.
Stage 2 — traffic spikes, reclaim begins
More concurrent requests mean more PHP-FPM child processes:
php-fpm 110 MB
php-fpm 95 MB
php-fpm 130 MB
php-fpm 87 MB
...Each new worker needs its own heap. MySQL is also opening more connections and allocating per-connection buffers (sort buffers, join buffers, etc. — separate from the shared buffer pool). Linux doesn't panic. It looks at the ~2 GB of cache and starts discarding pages that can be safely re-read from disk later. You'll watch buff/cache drop from something like 3 GB to 500 MB over a short window. This, on its own, is not damage — it's the kernel doing its job.
Stage 3 — reclaim can't keep up
If PHP-FPM keeps spawning workers (check your pm.max_children setting — this is usually the actual root cause, not "not enough RAM") and MySQL's buffer pool plus per-connection memory keeps growing, cache eventually runs out of slack to give up. At this point:
Swap usage starts climbing, if swap is configured, which is slow (especially on network-attached storage — see above).
Processes doing allocations enter direct reclaim and stall.
Response times spike even though CPU usage might look moderate.
Load average climbs, mostly from processes in uninterruptible sleep waiting on I/O, not from CPU-bound work.
Stage 4 — the OOM killer
If none of the above frees enough memory, the kernel's out-of-memory killer selects a process to terminate, based on an oom_score that weighs memory footprint, priority hints (oom_score_adj), and other heuristics. It does not necessarily kill the process using the most RAM outright, and it does not distinguish between something trivial and something critical unless you've configured it to.
In our example stack, that could be:
mysqld — usually the largest single consumer, often the OOM killer's first target by raw score. If MySQL dies, your entire site returns database connection errors while Apache and PHP keep running perfectly fine and answering health checks.
php-fpm master or workers — pages start returning 502/504 from Apache/Nginx because there's nothing to hand the request to.
redis-server — sessions or cache reads start failing, often silently, if your app doesn't handle Redis errors gracefully.
Check dmesg or journalctl -k after an incident — an OOM kill leaves a clear trail: Out of memory: Killed process 1234 (mysqld), along with a memory breakdown at the time of the kill.
The uptime trap
Here's the part that catches people out: your uptime monitor pings the server, ICMP or TCP on port 22, and it comes back fine. The VPS is alive. SSH works. But mysqld is dead, Apache is serving 500s for every PHP request, and your actual website has been down for twenty minutes before anyone notices — because "the server is up" and "the website works" are different claims, and most monitoring only checks the first one.
If you're monitoring anything in production, check the HTTP response of the actual page, not just whether the host responds to a ping.
AlmaLinux and other RHEL-family distros
The underlying kernel mechanism — kswapd, reclaim, watermarks, OOM killer — is identical to Ubuntu, since it's the same Linux kernel behaviour, not a distro-specific feature. What differs is tooling and defaults:
AlmaLinux/RHEL commonly ships with earlyoom or systemd-oomd disabled by default in some configurations, whereas newer Ubuntu releases (20.04+) enable
systemd-oomdby default, which acts before the kernel's own OOM killer, using pressure stall information (PSI) rather than waiting for a hard memory exhaustion event.SELinux (enforcing by default on AlmaLinux) doesn't affect memory reclaim itself, but can complicate diagnosis if a service restart triggers permission-denied errors that look unrelated to the original OOM event.
Package defaults for MySQL/MariaDB buffer pool sizing differ slightly between distros' repo versions, which affects how much headroom you have before Stage 3 kicks in.
Practically: the failure mode is the same story on both. What changes is which watchdog (kernel OOM killer vs. systemd-oomd vs. earlyoom) intervenes first and how aggressively.
What happens on Windows
Windows doesn't have a direct equivalent of the Linux OOM killer, and this is a genuine architectural difference, not just terminology.
Windows uses a unified memory manager where the pagefile is central to the design, far more so than Linux swap, which many production Linux servers run with minimal or no swap at all. When physical RAM runs low, Windows aggressively pages out memory to the pagefile rather than killing processes outright. The system will keep paging until it's genuinely exhausted both RAM and pagefile space.
The practical symptom on a Windows Server VPS running IIS, SQL Server, or similar is different from Linux: instead of a sudden process kill, you typically see severe, sustained slowness as the disk (or virtual disk) thrashes under paging I/O. Eventually you may see:
Application pools in IIS becoming unresponsive and getting recycled by the Application Pool's own idle/memory limits (if configured), which is an IIS-level safeguard, not an OS-level kill.
SQL Server failing to allocate memory for new connections, throwing out-of-memory exceptions in its own error log rather than being killed by the OS.
In genuinely severe exhaustion, Windows can terminate processes via its own low-memory resource notification API, but this is much less commonly triggered in practice than Linux's OOM killer, because Windows tends to page hard before it gets there.
So: Linux tends to fail fast (a process gets killed, service drops, you can restart it in seconds). Windows tends to fail slow (grinding pagefile thrashing that degrades everything gradually, and can take longer to resolve because there's no single "here's what got killed" log entry to point at).
Swap makes this worse, not better, on constrained VPS storage
Swap gives the kernel more room to delay OOM kills, but on a VPS where storage is SSD-backed but network-attached (which is how most cloud VPS storage, including sharded iSCSI-backed volumes, is architected), swapping introduces network round-trips into what's meant to be a fast memory operation. Heavy swapping under memory pressure can make a server dramatically slower than simply hitting the OOM killer sooner would have. Many operators deliberately run with small or no swap on latency-sensitive web workloads for exactly this reason, accepting a faster, cleaner OOM kill over a long, thrashing decline.
What actually fixes this
None of the above changes the underlying fix, which is capacity planning, not kernel tuning:
Cap
pm.max_childrenin PHP-FPM based on actual average worker memory usage, not guesswork.Set MySQL's
innodb_buffer_pool_sizedeliberately, leaving headroom for connection buffers and OS cache — don't let it default to a fixed fraction that assumes it owns the whole box.Set
oom_score_adjfor critical processes if you want to influence which one gets sacrificed first (lower priority for things you can afford to lose, like a batch job, higher protection for mysqld).Monitor actual available memory and PSI (
/proc/pressure/memoryon modern kernels) rather than just "% used", since used-but-cached is not a problem.If you're consistently hitting reclaim under normal traffic, the fix is more RAM, not more swap.
If you're resizing to add RAM on a VPS with detachable storage, you don't need to rebuild the box from scratch — you can move volumes and IPs to a larger instance rather than migrating data manually. We cover the mechanics of that in How Iotamine Cloud VPS's Detachable Volumes, IPs Actually Work.
Summary
Running low on RAM isn't a binary crash event on Linux — it's a multi-stage process of cache reclaim, background scanning, direct reclaim stalls, latency spikes visible in PSI, then swap, and only then an OOM kill, if it gets that far. Windows takes a different path, leaning on pagefile thrashing rather than process termination. In both cases, "the server is up" tells you nothing about whether the site is actually serving requests. Monitor the application layer, not just the host.
FAQ
Does 100% memory usage on Linux mean the server is about to crash?
No. Linux uses spare RAM for disk cache, so high 'used' memory often just reflects reclaimable cache, not application demand. Check available memory and swap activity, not the raw used percentage.
Which process gets killed first when a Linux VPS runs out of memory?
The OOM killer picks based on an oom_score that weighs memory footprint and configured priority (oom_score_adj). On a typical LAMP stack, mysqld is often the largest consumer and a common target, but it's not guaranteed.
Why does my website go down but the server still responds to ping?
ICMP/SSH checks only confirm the host is up, not that services like MySQL or PHP-FPM are running. An OOM kill can take down a single service while the OS keeps running normally.
Does Windows have an OOM killer like Linux?
Not directly. Windows relies heavily on the pagefile and tends to page aggressively under memory pressure, causing slowdown through disk thrashing rather than killing a process outright, though severe exhaustion can still trigger process termination via low-memory notifications.
Should I add swap to prevent OOM kills on a cloud VPS?
It depends on your storage. On network-backed storage like sharded iSCSI, heavy swapping adds network latency to memory operations, which can make performance worse than a fast OOM kill. Many operators prefer minimal swap for latency-sensitive web workloads.

