Skip to content
All articles
7 min read

Nginx 502 Bad Gateway: How to Find and Fix the Cause

MD Rakibul Islam RakibMD Rakibul Islam RakibFull-stack developer, DevOps & Linux engineer
Nginx 502 Bad Gateway: How to Find and Fix the Cause

A 502 Bad Gateway means Nginx couldn't get a valid response from your app. Check the Nginx error log: it names the cause, usually a crashed app or wrong port.

A 502 is one of the most common server emergencies, and it's usually fixed in minutes once you look in the right place. The key thing to understand is that Nginx is almost never the problem. Nginx is a reverse proxy: it receives the visitor's request and forwards it to your application (Node.js, Next.js, PHP-FPM, Python, a Docker container). A 502 is Nginx telling you "I tried to talk to the app behind me and it didn't answer properly." This guide shows how to find out why in one command, then fix each cause. It's the same routine I use on my own servers, where Nginx sits in front of a Next.js app and a NestJS API run by PM2.

Key takeaways

  • Read the Nginx error log first: sudo tail -n 50 /var/log/nginx/error.log. The message tells you which of the causes below you have.
  • "Connection refused" means nothing is listening on the upstream port: the app crashed, didn't start, or listens on a different port.
  • "Permission denied" on a socket usually means PHP-FPM or a Unix socket owned by the wrong user.
  • "Upstream prematurely closed connection" means the app crashed mid-request, often out of memory.
  • 502 is not 504. 504 means the app answered too slowly; 502 means it didn't answer or answered with something invalid.

Step 1: Confirm it's a real 502 from your server

Check from your laptop, not the server, and look at the response headers:

curl -sI https://example.com | head -5

If you see server: nginx with HTTP/2 502, it's your Nginx. If you see server: cloudflare, Cloudflare couldn't reach your origin server at all, which is a different problem (server down, firewall, or wrong origin IP). My website down checklist covers that path.

Step 2: Read the error log

sudo tail -n 50 /var/log/nginx/error.log
# or follow it live while you reload the page:
sudo tail -f /var/log/nginx/error.log

If your site has its own log (error_log inside the server block), check that file instead. grep -r error_log /etc/nginx/ shows where they go. Each 502 line includes the upstream address Nginx tried, such as upstream: "http://127.0.0.1:3000/". Match the message to one of the causes below.

visitor Nginxrunning app :3000crashed / not listening 111: refused 502 Nginx is fine. The thing behind it isn't.
Nginx accepts the request and forwards it to the app. When the app doesn't answer, Nginx has nothing to return, so it sends the visitor a 502. The fix is almost always on the app side.

Cause 1: "connect() failed (111: Connection refused)"

The most common one. Nothing is listening on the address Nginx forwards to. Check:

sudo ss -tlnp | grep -E ':3000|:5000|:8000'   # use your app's port
pm2 status                                     # Node apps under PM2
sudo systemctl status myapp                    # apps run by systemd
docker ps                                      # containers

If nothing is listening, the app is stopped or crash-looping. Read its own logs (pm2 logs myapp --lines 100, journalctl -u myapp -n 100, docker logs myapp). The usual reasons I find: a missing environment variable after a deploy, a failed database connection, a build that never finished, or the server rebooted and the app wasn't set to start on boot (pm2 startup plus pm2 save, or systemctl enable).

If something is listening but on a different port or address, fix the mismatch. A classic is the app listening on localhost resolving to IPv6 ::1 while Nginx connects to 127.0.0.1. Use the explicit IP in proxy_pass http://127.0.0.1:3000; and make the app bind to the same address.

Cause 2: "connect() to unix:/run/php/... failed (13: Permission denied)" or "(2: No such file)"

Nginx is using a Unix socket, usually for PHP-FPM, and can't open it. "No such file" means PHP-FPM isn't running or the socket path is wrong (it includes the PHP version, for example php8.3-fpm.sock, which changes when PHP is upgraded). "Permission denied" means the socket's owner doesn't match the Nginx user:

ls -l /run/php/
sudo systemctl status php8.3-fpm
grep -E '^(listen|listen.owner|listen.group|listen.mode)' /etc/php/8.3/fpm/pool.d/www.conf

Set listen.owner = www-data and listen.group = www-data (the user Nginx runs as on Ubuntu), point fastcgi_pass at the socket that actually exists, then sudo systemctl restart php8.3-fpm.

Cause 3: "upstream prematurely closed connection"

The app accepted the request, then died or closed the connection before answering. Often it's being killed for using too much memory. Check for the kernel's OOM killer:

sudo journalctl -k --since "1 hour ago" | grep -i -E 'out of memory|oom'
free -h
pm2 monit

If the OOM killer is hitting your app, add swap as a safety net, cap memory with PM2's --max-memory-restart, and find the leak or heavy request. Large uploads and image processing are common triggers. Also check the app log for an unhandled exception on that route.

Cause 4: "upstream sent too big header while reading response header"

The app answered, but its response headers (often big cookies or auth tokens) didn't fit Nginx's buffer. Raise the buffers in the location block:

proxy_buffer_size 16k;
proxy_buffers 8 16k;
proxy_busy_buffers_size 32k;

For PHP, the equivalents are fastcgi_buffer_size and fastcgi_buffers. Then test and reload: sudo nginx -t && sudo systemctl reload nginx.

Cause 5: "no live upstreams"

You use an upstream block with several servers and Nginx has marked all of them as failed. Fix the backends first (cause 1 or 3). Nginx retries them automatically after fail_timeout, 10 seconds by default.

Cause 6: SELinux blocking the connection (RHEL, Rocky, Alma)

On Red Hat family systems, "Permission denied" connecting to a TCP port can be SELinux, not file permissions. Check with sudo ausearch -m avc -ts recent, and allow Nginx to make network connections with sudo setsebool -P httpd_can_network_connect 1. Ubuntu and Debian don't use SELinux by default.

How to stop 502s from coming back

  • Auto-restart the app with PM2 or a systemd Restart=always, and make sure it starts on boot.
  • Health-check every deploy and roll back automatically if the app doesn't answer. That's what my GitHub Actions deploy with rollback does, so a bad release never leaves the site on a 502.
  • Watch memory and disk. A full disk causes crashes too; see fixing "No space left on device".
  • Add an external uptime check so you hear about a 502 before your customers do.
  • Show a friendly error page with error_page 502 /502.html; so visitors see a message instead of the plain Nginx page.

My Next.js on a VPS guide shows a complete Nginx + PM2 setup built this way.

Frequently asked questions

What causes a 502 Bad Gateway error in Nginx?

Nginx couldn't get a valid response from the application behind it. The usual causes are a crashed or stopped app, Nginx pointing at the wrong port or socket, socket permission problems, the app being killed for using too much memory, or response headers too large for Nginx's buffers.

Is a 502 error my fault or the hosting provider's?

On your own VPS it's almost always the application or its configuration, not the provider. On shared or managed hosting it can be their backend, so check their status page and contact support with the time of the error.

What's the difference between 502 and 504?

502 Bad Gateway means the app refused the connection, crashed or sent an invalid response. 504 Gateway Timeout means the app accepted the request but didn't finish within Nginx's proxy_read_timeout, 60 seconds by default.

Will restarting Nginx fix a 502?

Rarely, because Nginx is usually fine. Restarting the application often brings the site back, but the 502 will return unless you find out why it stopped, so always read its logs.

How do I fix 502 Bad Gateway with Cloudflare?

If the 502 page is Cloudflare's, Cloudflare couldn't reach your server. Check the server is up, ports 80 and 443 are open to Cloudflare, the DNS record points at the right IP, and the SSL mode matches your origin certificate.

Site showing 502 right now?

I fix 502 errors and other server outages on Linux, usually the same day, and then set up the restarts, health checks and monitoring so they don't come back. See my Linux system admin services or contact me now with your domain and the last lines of your error log.

MD Rakibul Islam Rakib

Written by

MD Rakibul Islam Rakib

Full-stack developer, DevOps engineer and Linux system administrator with 5+ years of production experience. I deploy, harden and fix servers and web apps for clients worldwide, and everything in this article runs on real servers I manage, including this site.

  • Nginx 502 Bad Gateway
  • 502 bad gateway fix
  • Nginx error log
  • upstream connection refused
  • PM2
  • PHP-FPM
  • Linux server

Keep reading

Certbot Renewal Failed? Fix an Expired Let's Encrypt SSL
DevOpsOct 6, 2026

Certbot Renewal Failed? Fix an Expired Let's Encrypt SSL

If Certbot renewal failed, run sudo certbot renew --dry-run to see the error. Usual causes: port 80 blocked, a redirect hiding the challenge, or wrong DNS. An expired certificate takes a site down as completely as a crashed server: browsers show a full-page "Your connection is not private" warning and most visitors leave. It's also become more likely. Let's Encrypt stopped sending expiry reminder emails on June 4, 2025, so when automatic renewal quietly breaks, nobody is warned any more. And certificates are getting shorter: Let's Encrypt's default lifetime drops from 90 to 64 days in February 2027 and to 45 days in February 2028. Renewal has to work every time. Here is how to find why it failed, fix it, and get warned next time. Key takeaways Start with sudo certbot renew --dry-run . It simulates renewal and prints the real error without using up rate limits. Most failures are the HTTP-01 challenge: port 80 closed, /.well-known/acme-challenge/ redirected or blocked, or DNS pointing somewhere else. Check the timer is running: snap.certbot.renew.timer for snap installs, certbot.timer for apt installs. Don't have both. Reload Nginx after renewal with a deploy hook, or it keeps serving the old certificate. Monitor expiry yourself. Let's Encrypt no longer emails you. Step 1: Get the site back right now If the certificate has already expired, try a forced renewal first and read the output: sudo certbot certificates # list certs, domains and expiry dates sudo certbot renew --cert-name example.com --force-renewal sudo systemctl reload nginx If that works, the site is back and you can find out why automatic renewal broke. If it fails, the error message tells you which of the causes below you're dealing with. Don't retry in a loop: Let's Encrypt has rate limits on failed validations, and --dry-run uses the staging environment, which doesn't count against production limits. Step 2: Run a dry run and read the error sudo certbot renew --dry-run sudo tail -n 100 /var/log/letsencrypt/letsencrypt.log The renewal config for each certificate is in /etc/letsencrypt/renewal/example.com.conf . The authenticator line ( nginx , webroot or standalone ) tells you how Certbot proves you own the domain, which matters for the fixes below. With HTTP-01, Let's Encrypt fetches a token from your server over plain HTTP on port 80. If a firewall, redirect, wrong DNS record or proxy gets in the way, validation fails and the certificate isn't renewed. Cause 1: Port 80 is closed Errors mention "Timeout during connect" or "Connection refused". HTTP-01 validation always starts on port 80, even if your site is HTTPS only. Check your firewall and your cloud provider's firewall: sudo ufw status | grep -E '80|Nginx' sudo ufw allow 'Nginx Full' # opens 80 and 443 curl -I http://example.com/.well-known/acme-challenge/test # from another machine Closing port 80 "for security" is a very common reason renewals stop. Keep it open and redirect everything except the challenge path to HTTPS. Cause 2: A redirect or rule hides the challenge Errors like "Invalid response" or a 404 for /.well-known/acme-challenge/ mean Let's Encrypt reached your server but didn't get the token. Typical culprits are a catch-all redirect to another domain, an app route that swallows every URL, basic auth on the whole site, or a "deny all dotfiles" rule like location ~ /\. . Serve the challenge explicitly in the port 80 server block: server { listen 80; server_name example.com www.example.com; location /.well-known/acme-challenge/ { root /var/www/letsencrypt; } location / { return 301 https://$host$request_uri; } } If you use this webroot, make sure the renewal config matches: authenticator = webroot and webroot_path = /var/www/letsencrypt . Or switch back to the Nginx plugin with sudo certbot --nginx -d example.com -d www.example.com . Cause 3: Standalone authenticator while Nginx runs "Could not bind TCP port 80 because it is already in use" means the certificate was first issued with --standalone (Certbot runs its own temporary web server) and Nginx now holds port 80. Reissue it with the Nginx or webroot method so renewals don't need port 80 free: sudo certbot certonly --nginx -d example.com -d www.example.com Cause 4: DNS points somewhere else Let's Encrypt validates every name on the certificate. If www or an old subdomain now points at a different server, a CDN or nowhere, renewal of the whole certificate fails. Check each name: for d in example.com www.example.com; do echo "$d: $(dig +short $d A) $(dig +short $d AAAA)"; done Watch out for a stale AAAA (IPv6) record : Let's Encrypt tries IPv6 first when an AAAA record exists, so an old AAAA record pointing at a server that no longer exists can break validation even when the A record is correct. Remove names you don't use from the certificate with sudo certbot --nginx --cert-name example.com -d example.com -d www.example.com , listing only the live ones. If the site is behind Cloudflare's proxy, HTTP-01 usually still works as long as the challenge path isn't blocked, or use a DNS-01 plugin. Cause 5: The renewal timer isn't running systemctl list-timers | grep -i certbot Snap installs (the method Certbot recommends) use snap.certbot.renew.timer ; apt installs use certbot.timer . If neither is listed, renewal never runs. If you switched install methods, you may have two Certbots or none. Pick one, remove the other ( sudo apt remove certbot if you use the snap), and check that which certbot points where you expect. Cause 6: Renewed, but the site still shows the old certificate Nginx loads certificates at startup and on reload. If renewal succeeded but browsers still see the expired one, Nginx wasn't reloaded. The Nginx plugin does this for you; with webroot, add a deploy hook so it happens after every successful renewal: sudo sh -c 'printf "#!/bin/sh\nsystemctl reload nginx\n" > /etc/letsencrypt/renewal-hooks/deploy/reload-nginx.sh' sudo chmod +x /etc/letsencrypt/renewal-hooks/deploy/reload-nginx.sh Also confirm Nginx points at /etc/letsencrypt/live/example.com/fullchain.pem and privkey.pem , not at a copied file that never updates. Step 3: Monitor expiry yourself Since the reminder emails are gone, add your own check. The SSL expiry script in my Bash scripts for DevOps post warns 14 days ahead; run it daily from cron or a systemd timer, or use any uptime service with certificate monitoring. With 45-day certificates coming, renewal should happen at about two thirds of the lifetime, which Certbot handles automatically as long as the timer runs. Expired certificates are one of the first things to check when a website is down , and my Next.js on a VPS guide shows a full Nginx + Certbot setup. Frequently asked questions Why did my Let's Encrypt certificate not renew automatically? Usually the renewal timer isn't running, port 80 is blocked, a redirect or app route hides the /.well-known/acme-challenge/ path, or a DNS record on the certificate points elsewhere. sudo certbot renew --dry-run shows which. Does Let's Encrypt still send expiration emails? No. Let's Encrypt ended its expiration notification emails on June 4, 2025. Use your own monitoring or a third-party service to get warned before a certificate expires. How long are Let's Encrypt certificates valid? 90 days by default today. Let's Encrypt plans to shorten the default to 64 days in February 2027 and 45 days in February 2028, so renewal must be fully automated. Do I need port 80 open if my site is HTTPS only? Yes, for HTTP-01 validation, which is what Certbot uses by default. Keep port 80 open, serve the challenge path, and redirect everything else to HTTPS. DNS-01 validation is the alternative if port 80 can't be opened. How do I renew an SSL certificate manually with Certbot? Run sudo certbot renew for all certificates, or sudo certbot renew --cert-name example.com --force-renewal for one, then reload Nginx. Fix the underlying cause too, or it will fail again at the next renewal. SSL expired and site showing warnings? I fix expired and failing SSL certificates quickly, then make renewal reliable with correct Nginx config, hooks and monitoring. See my DevOps services or contact me with your domain.

Read article →
Linux Server Hacked? What to Do in the First Hour
Linux System AdminOct 6, 2026

Linux Server Hacked? What to Do in the First Hour

If your Linux server is hacked, isolate it, snapshot it, find how the attacker got in, rotate every secret, then rebuild a fresh server from clean backups. The worst thing you can do with a hacked server is what most people try first: kill the strange process, delete the suspicious file, and carry on. Attackers almost always leave more than one way back in: a cron job, an SSH key, a systemd service, a modified binary. Whether the attacker installed a crypto miner, a spam sender or defaced the website, the reliable recovery is the same: contain, understand, rotate, rebuild. Here is what to do in the first hour, in order. Key takeaways Don't panic-delete. Contain first, and keep evidence so you can find out how they got in. Isolate with the cloud firewall (allow only your IP) rather than powering off, so memory and running processes can still be inspected. Assume every secret on the box is stolen: SSH keys, database passwords, API keys, .env files, payment and email credentials. Rebuild, don't clean. You can't prove a compromised system is clean. Deploy a fresh server and restore data from a backup taken before the break-in. Fix the hole before going live , or the new server is compromised within days. Signs your server has been hacked CPU at 100% from a process you don't recognise, often a crypto miner with a random or innocent-looking name. Your provider suspends the server or emails you about abuse, port scanning or outgoing spam. New user accounts, unknown keys in authorized_keys , or logins from strange countries in last . Spam pages, redirects or injected JavaScript on your website. Unknown cron jobs, systemd services or files in /tmp , /var/tmp or /dev/shm . Exposed services: an open database, Redis or Docker API reachable from the internet. Step 1: Contain it (first 10 minutes) Cut the attacker off without destroying evidence. In your provider's dashboard, apply a cloud firewall that allows only SSH from your own IP and blocks everything else, including outgoing traffic if the panel allows it. This stops the miner talking to its pool, the spam from going out, and the attacker's remote shell, while the server keeps running. Put up a maintenance page on another server or via your CDN if you need customers to see something. Tell your team not to log in or "fix" things on the box. Step 2: Snapshot it for evidence Take a disk snapshot in the provider panel before you change anything. It costs a little and it lets you investigate later at your own pace, or hand it to a specialist. If personal data may have been accessed, you may have legal reporting duties (for example under GDPR), and you'll need to know what was touched. Containment means cutting the attacker's connections while keeping the server running for investigation. A provider-level firewall does that without trusting anything on the compromised machine. Step 3: Look for how they got in (and what they left) Commands on a compromised box can be lied to by a rootkit, so treat what you see as clues, not proof. Still, these usually reveal a lot: # who logged in, and failed attempts last -a | head -30 sudo lastb -a | head -20 sudo journalctl -u ssh --since "7 days ago" | grep -E 'Accepted|Invalid' | tail -50 # what is running and listening ps auxf --sort=-%cpu | head -20 sudo ss -tunap # persistence: cron, systemd, keys, users for u in $(cut -d: -f1 /etc/passwd); do sudo crontab -l -u "$u" 2>/dev/null | sed "s/^/$u: /"; done ls -la /etc/cron.* /var/spool/cron/crontabs 2>/dev/null systemctl list-units --type=service --state=running sudo find /etc/systemd /lib/systemd -name '*.service' -mtime -14 sudo find / -name authorized_keys -exec ls -la {} \; 2>/dev/null awk -F: '$3 == 0' /etc/passwd # should only print root # recently changed files and temp dirs sudo find / -xdev -type f -mtime -3 -not -path '/proc/*' 2>/dev/null | grep -v -E '^/(var/log|var/lib/apt)' | head -50 ls -la /tmp /var/tmp /dev/shm On Debian and Ubuntu, sudo debsums -c (from the debsums package) lists system files that differ from the packaged versions, which catches replaced binaries. Copy what you find, with timestamps, into a notes file on your laptop. The most common ways in: Weak or leaked SSH password on root or a default user. Outdated web app or plugin : WordPress plugins, old frameworks, a forgotten admin panel. Exposed services : a database, Redis, Docker API or dashboard open to the internet with no password. Docker-published ports even bypass UFW; see Docker bypasses UFW . Leaked secrets : a .env file committed to a public repo or served by the web server. Malicious dependencies pulled into the app, including AI-hallucinated package names ( slopsquatting ). Step 4: Rotate every secret Do this from a clean machine, assuming anything stored on the server was copied: SSH keys that were on the server, and any key the server used to reach other systems (deploy keys, Git, backups). Database passwords and app user passwords stored in config or .env files. Third-party API keys: payment, email, SMS, cloud storage, AI providers. Check their dashboards for unusual usage. Admin passwords for the website, hosting panel, DNS and registrar, plus session secrets so existing logins are invalidated. Check other servers that trusted this one. Attackers move sideways using keys they find. Step 5: Rebuild on a fresh server Create a new server from a clean OS image, harden it first (my Ubuntu 24.04 hardening checklist : key-only SSH, firewall, automatic security updates, no exposed databases), then: Deploy the application from Git, not by copying files from the old server, and update it and its dependencies. Restore data from a backup taken before the compromise. Check uploads and database content for injected scripts or rogue admin accounts. Close the hole you found in step 3. Switch DNS to the new server, then delete the old one once you no longer need it for investigation. This is where good backups pay off. If you don't have off-server backups, set them up on the new server today with my Linux backup guide . Step 6: Watch and follow up Monitor logins, CPU and outgoing traffic on the new server for a few weeks. Ask Google to review the site in Search Console if it was flagged for malware or spam. Tell affected users if their data may have been accessed, and meet any legal obligations. Write down what happened and what you changed. Frequently asked questions How do I know if my Linux server is hacked? Common signs are unexplained 100% CPU, unknown processes or users, new SSH keys, strange cron jobs or services, logins from unfamiliar IPs, spam on your website, and abuse reports from your provider. Can I just delete the malware and keep using the server? It's risky. Attackers usually install several ways back in, and a rootkit can hide files and processes. Rebuilding on a fresh server from clean backups is the only way to be confident it's clean. Should I shut down a hacked server? Isolate it with your provider's firewall instead of shutting it down, if you can. That stops the damage while keeping running processes and memory available for investigation. Take a snapshot either way. What is a crypto miner on my server? It's malware that uses your CPU to mine cryptocurrency for the attacker. It's one of the most common results of a compromised server, and it often arrives through weak SSH passwords or exposed services. How do I stop my server from being hacked again? Use SSH keys only, keep the OS and apps updated, expose only ports 22, 80 and 443, never put databases or admin panels on the public internet, keep secrets out of Git, and monitor the server. Think your server has been hacked? I handle compromised Linux servers: containment, finding the entry point, rotating secrets and rebuilding a clean, hardened server with your data restored. See my Linux system admin services or contact me now . Don't wait; every hour gives the attacker more time.

Read article →
Linux "No Space Left on Device": Find and Free Disk Space
Linux System AdminOct 6, 2026

Linux "No Space Left on Device": Find and Free Disk Space

"No space left on device" means a filesystem is out of blocks or inodes. Run df -h and df -i, find big folders with du, then clear logs, Docker and caches. A full disk is one of the most common reasons a server looks "broken". The symptoms rarely say "disk full": the database refuses writes, uploads fail, the app crashes, deploys break halfway, SSH logins hang, or Nginx won't reload. Then df shows 100%. The good news is that it's usually fixable in a few minutes without losing anything important, as long as you delete the right things. This is the order I work in on Ubuntu servers, from the safest cleanup to the riskiest. Key takeaways Check both blocks and inodes: df -h for space, df -i for file count. Either at 100% gives the same error. Find the hog before deleting: du -xh / --max-depth=2 | sort -h | tail shows where the space went. The usual suspects: logs and the systemd journal, Docker images and container logs, old releases and backups, package caches. If du and df disagree, a process is holding a deleted file open. lsof +L1 finds it. Never delete files inside database data directories to make room. Clear something else first, then fix the database properly. Step 1: Confirm which filesystem is full df -h -x tmpfs -x devtmpfs -x squashfs df -i -x tmpfs -x devtmpfs -x squashfs Look at the Use% and IUse% columns. Note the mount point: / full is the common case, but /boot , /var or a separate data volume can fill on their own. If IUse% is 100% while space is free, you have millions of tiny files (jump to step 6). Step 2: Find what's using the space sudo du -xh / --max-depth=2 2>/dev/null | sort -h | tail -20 sudo du -xh /var --max-depth=3 2>/dev/null | sort -h | tail -20 -x stays on one filesystem, so it doesn't wander into mounted volumes. Drill into whatever is biggest. If you prefer an interactive view, sudo apt install ncdu then sudo ncdu -x / lets you browse and delete with arrow keys. On a 100% full disk, apt itself may fail, so free a little space with the steps below first. Step 3: Clear logs and the systemd journal The journal can grow to gigabytes. Shrink it safely: journalctl --disk-usage sudo journalctl --vacuum-size=500M Make the limit permanent with SystemMaxUse=500M in /etc/systemd/journald.conf , then sudo systemctl restart systemd-journald . For big files in /var/log , delete old rotated logs ( *.gz , *.1 ). For a huge log that is still being written, empty it instead of deleting it , or the app keeps writing to the deleted file and the space never comes back: sudo find /var/log -type f \( -name '*.gz' -o -name '*.[0-9]' \) -delete sudo truncate -s 0 /var/log/nginx/access.log PM2 logs live in ~/.pm2/logs and grow forever by default. pm2 flush empties them, and the pm2-logrotate module keeps them in check. Most of a full disk is usually things that can be regenerated: Docker images, the journal, old logs and releases. Clearing those frees the space without touching your data. Step 4: Clean up Docker On servers running containers, /var/lib/docker is the most common culprit: old images from every deploy, stopped containers, build cache and unbounded container logs. docker system df docker system prune -f # stopped containers, unused networks, dangling images, build cache docker image prune -a -f # also images not used by any container sudo du -sh /var/lib/docker/containers/*/*-json.log | sort -h | tail -5 Don't add --volumes unless you're sure: volumes hold database data. To stop container logs growing forever, set a limit in /etc/docker/daemon.json and restart Docker (it applies to newly created containers): { "log-driver": "json-file", "log-opts": { "max-size": "50m", "max-file": "3" } } Step 5: Packages, kernels, snaps and old releases sudo apt-get clean # downloaded .deb cache sudo apt-get autoremove --purge # old kernels and unused packages snap list --all | awk '/disabled/{print $1, $3}' Snap keeps old revisions of each package; remove disabled ones with sudo snap remove NAME --revision=REV , and keep fewer with sudo snap set system refresh.retain=2 . A full /boot is almost always old kernels, which autoremove clears. Then check your own files: old deploy releases, database dumps, uploaded files in temp folders, node_modules copies and forgotten .tar.gz backups in home directories. My deploy script keeps only the newest five releases and 30 database dumps for exactly this reason; see Bash scripts for DevOps . Backups belong off the server anyway, as in my Linux backup guide . Step 6: Out of inodes If df -i shows 100%, find the folder with the most files: sudo du -x --inodes / --max-depth=3 2>/dev/null | sort -n | tail -15 Typical causes are PHP session files, a mail queue, cache folders or a script that creates a file per request. Delete the old ones by age, for example sudo find /var/lib/php/sessions -type f -mtime +2 -delete , and fix whatever creates them. Step 7: df says full, du says not (deleted but open files) When you delete a file that a process still has open, the space isn't freed until the process closes it. df counts it; du can't see it. sudo lsof +L1 | sort -k7 -n | tail Restart the process holding the file (often a logging app, Nginx or a database client), and the space comes back immediately. Step 8: Check the reserved space and the volume size ext4 reserves 5% of the disk for root by default, so normal users see "full" before root does. On a large data-only volume you can reduce it with sudo tune2fs -m 1 /dev/sdb1 ; leave it on the root filesystem. If the disk really is too small for your data, resize the volume in your provider's panel, then grow the partition and filesystem ( growpart and resize2fs ). Take a snapshot first. Stop it from happening again Alert at 80-85% , not 100%. The disk alert script in my Bash scripting post does this in ten lines. Cap every log: journald SystemMaxUse , Docker max-size , logrotate for app logs, pm2-logrotate. Clean up on deploy: keep the last few releases and images, delete the rest automatically. Move bulky data such as uploads and backups to object storage instead of the root disk. A full disk often shows up first as a 502 Bad Gateway or a site that's down for no clear reason, so check df -h early. Frequently asked questions What does "No space left on device" mean in Linux? The filesystem you're writing to has no free blocks, or no free inodes (file entries) left. Check both with df -h and df -i ; either one at 100% produces this error. Is it safe to delete files in /var/log? Old rotated logs ( .gz and numbered files) are safe to delete. For logs that are still being written, empty them with truncate -s 0 instead of deleting, so the space is actually freed. Why is my disk still full after deleting files? A running process probably still has the deleted file open. Find it with sudo lsof +L1 and restart that process to release the space. Is docker system prune safe? docker system prune removes stopped containers, unused networks, dangling images and build cache, which is safe on most servers. Avoid --volumes unless you're certain, because volumes usually contain database data. How do I find large files on Linux? Use sudo du -xh / --max-depth=2 | sort -h | tail to find big folders, or sudo find / -xdev -type f -size +500M to list individual large files. ncdu gives an interactive view. Server out of disk right now? I recover full servers safely, without touching your data, then set up log limits, cleanup and alerts so it doesn't happen again. See my Linux system admin services or contact me with the output of df -h .

Read article →