Skip to content
All articles
7 min read

Postgres "Too Many Clients Already": How to Fix It

MD Rakibul Islam RakibMD Rakibul Islam RakibFull-stack developer, DevOps & Linux engineer
Postgres "Too Many Clients Already": How to Fix It

"Too many clients already" means Postgres hit max_connections. Find who holds connections in pg_stat_activity, shrink app pools, and only then raise the limit.

The error usually arrives all at once: the API starts failing every request with FATAL: sorry, too many clients already, or remaining connection slots are reserved for roles with the SUPERUSER attribute on Postgres 16 and later. Restarting the app makes it go away for a while, then it's back. It's almost never "too many users". It's too many connection pools, each holding more connections than you think. My own API runs NestJS with Prisma 7 and Postgres, so the Node examples below are the ones I actually use.

Key takeaways

  • Count connections, don't guess: pg_stat_activity shows who holds each one and whether it's active, idle or stuck "idle in transaction".
  • Do the math: total connections = app processes × pool size + workers + cron jobs + your admin tools. PM2 cluster mode and replicas multiply it.
  • Prisma 7 with driver adapters defaults to 10 connections per client; Prisma 6 defaulted to CPU cores × 2 + 1. Set it explicitly.
  • Kill stuck sessions with timeouts (idle_in_transaction_session_timeout) instead of by hand every night.
  • Raise max_connections last. Each connection costs memory. A pooler like PgBouncer scales better.

Step 1: Get in and see who holds the connections

Postgres keeps a few slots (superuser_reserved_connections, 3 by default) for superusers exactly for this moment. Connect as one:

sudo -u postgres psql
# in Docker:
docker exec -it db psql -U postgres
SHOW max_connections;
SELECT count(*) FROM pg_stat_activity;

SELECT usename, application_name, client_addr, state, count(*)
FROM pg_stat_activity
WHERE backend_type = 'client backend'
GROUP BY 1, 2, 3, 4
ORDER BY 5 DESC;

This tells you which user, app and host hold the slots and in what state:

  • Mostly idle from one app: pools that are too big, or too many processes each with their own pool.
  • idle in transaction: code that opened a transaction and never committed or rolled back. These hold locks too and are the most dangerous.
  • Many active: slow queries piling up. More connections won't help; faster queries will.

Step 2: Free slots right now

To get the site back while you fix the cause, end sessions stuck in a transaction for more than a few minutes:

SELECT pg_terminate_backend(pid), now() - state_change AS stuck_for, left(query, 60)
FROM pg_stat_activity
WHERE state = 'idle in transaction'
  AND now() - state_change > interval '5 minutes';

Restarting the app also releases its connections. Both are first aid; the next steps stop it coming back.

4 processes × pool of 25 = 100 connections proc 1: 25 proc 2: 25 proc 3: 25 proc 4: 25 97 normal slots: used 3 reservedsuperuser only Postgres max_connections = 100 too many clients
Four app processes with a pool of 25 each fill all 100 slots. The last few are reserved for superusers, so the next connection from the app, a cron job or a migration is refused. The fix is smaller pools, not more slots.

Step 3: Do the connection math

Write it down for every client of the database:

API:        4 PM2 cluster instances × pool 10  = 40
Workers:    2 queue workers × pool 10          = 20
Next.js:    2 instances × pool 10              = 20
Cron jobs, migrations, admin tools             ~ 10
                                       total  ~ 90  (of 97 usable)

This is how a small site hits 100 without much traffic. The usual multipliers:

  • PM2 cluster mode, Docker replicas and Kubernetes pods: every process has its own pool.
  • Serverless functions: each concurrent function instance can open its own connections. Use a pooler or your provider's pooled connection string.
  • Creating a new client per request instead of one shared client per process. Each new client opens a new pool.
  • Hot reload in development: every reload creates another client unless you keep one on globalThis.

Step 4: Set pool sizes explicitly

Prisma 7 (driver adapters)

In Prisma 7 the pool belongs to the driver. With @prisma/adapter-pg it's a pg pool, default 10 connections per client:

import { PrismaClient } from "@prisma/client";
import { PrismaPg } from "@prisma/adapter-pg";

const adapter = new PrismaPg({
  connectionString: process.env.DATABASE_URL,
  max: 5, // per process; multiply by your instance count
});
export const prisma = new PrismaClient({ adapter });

Prisma 6 and earlier

The pool was built into the query engine and sized from CPU cores. Set it in the URL: postgresql://…/app?connection_limit=5. Upgrading from 6 to 7 changes the default, so check this after the upgrade.

node-postgres, TypeORM, Sequelize, Knex

All have a pool max option. The pg default is 10. Create one pool per process at startup and reuse it everywhere.

A good starting point: keep the total of all pools under about 80% of max_connections, leaving room for migrations, backups and you.

Step 5: Add timeouts so leaks clean themselves up

ALTER DATABASE app SET idle_in_transaction_session_timeout = '60s';
ALTER DATABASE app SET statement_timeout = '30s';
-- Postgres 14+: close sessions idle for a long time (careful with poolers)
ALTER ROLE report_user SET idle_session_timeout = '10min';

New sessions pick these up. idle_in_transaction_session_timeout is the important one: it ends the session that forgot to commit, which also releases its locks.

Step 6: Raise max_connections or add PgBouncer

If the math says you really need more connections, you can raise the limit. It needs a restart, and each connection uses memory, so a small VPS can trade a connection error for the OOM killer (see the Linux OOM killer guide):

ALTER SYSTEM SET max_connections = 200;
-- then: sudo systemctl restart postgresql

For many app processes or serverless, a pooler is the better answer. PgBouncer in transaction mode lets hundreds of clients share a few dozen real connections. Some features (session-level prepared statements, SET, advisory locks) behave differently through it, so check your ORM's PgBouncer notes for your version before switching.

While you're in the database config: if Postgres runs in Docker with ports: "5432:5432", it's probably reachable from the internet despite UFW. Read why Docker bypasses UFW.

Frequently asked questions

What does "sorry, too many clients already" mean in Postgres?

Every connection slot allowed by max_connections is in use, so Postgres refuses new connections. It usually means app connection pools are too large or connections are leaking, not that you have too many users.

What does "remaining connection slots are reserved" mean?

All normal slots are used and only the few reserved for superusers are left. You can still log in as the postgres superuser to investigate. Postgres 16 changed the wording to "reserved for roles with the SUPERUSER attribute".

What is Prisma's default connection pool size?

In Prisma 7 with driver adapters such as @prisma/adapter-pg, the pool comes from the driver and defaults to 10 connections. In Prisma 6, the default was the number of physical CPU cores times two plus one, set with connection_limit in the URL.

Should I just increase max_connections?

Only after you've sized the pools and fixed leaks. Each connection is a separate process that uses memory, so raising the limit a lot on a small server can cause memory problems. A pooler such as PgBouncer scales further.

How do I find idle in transaction connections?

Query pg_stat_activity for rows where state is 'idle in transaction' and look at now() minus state_change. Set idle_in_transaction_session_timeout so Postgres ends them automatically.

Database refusing connections right now?

I fix Postgres connection, performance and pooling problems in Node.js and NestJS apps, and size them so the next traffic spike doesn't take the API down. See my web development services or contact me now with the error and the output of the pg_stat_activity query above.

MD Rakibul Islam Rakib

Written by

MD Rakibul Islam Rakib

Full-stack developer, DevOps engineer and Linux system administrator with 5+ years of production experience. I deploy, harden and fix servers and web apps for clients worldwide, and everything in this article runs on real servers I manage, including this site.

  • Postgres too many clients already
  • remaining connection slots are reserved
  • Prisma connection pool
  • max_connections
  • PgBouncer
  • pg_stat_activity
  • Node.js

Keep reading

Crawled - Currently Not Indexed: How to Fix It
Website DesignOct 6, 2026

Crawled - Currently Not Indexed: How to Fix It

"Crawled - currently not indexed" means Google fetched your page but chose not to index it. Rule out noindex, canonical and soft 404 issues, then improve it. You published the page, submitted the sitemap, waited, and Search Console still lists it under "Why pages aren't indexed". For a business site that means the service page or article you're counting on gets zero search traffic. The two statuses people hit most, "Crawled - currently not indexed" and "Discovered - currently not indexed", have different causes and different fixes. I went through this on this website: a stray canonical tag in the root layout was quietly pointing pages like Privacy and Terms at the homepage, and the canonical URL used the non-www domain that redirects. Here's the checklist I now use, technical blocks first, quality second. Key takeaways "Discovered" = not crawled yet. Google knows the URL but hasn't fetched it, often because it looks low priority or the server is slow. "Crawled" = fetched and not kept. Usually a quality or duplication judgement, but rule out technical causes first. Use URL Inspection's live test on one affected page. It shows the canonical Google chose, any noindex, and the HTML Google actually rendered. Internal links are the strongest lever you control: link to the page from your homepage, service pages and related posts. Request indexing once, after a real fix. Re-requesting an unchanged page doesn't help. What the two statuses mean According to Google's Page indexing report help , "Crawled - currently not indexed" means the page was crawled but not indexed, and it may or may not be indexed later. "Discovered - currently not indexed" means Google found the URL but postponed crawling it, typically to avoid overloading the site. Neither is an error in the technical sense. They're Google saying "not yet" or "not worth it". Your job is to remove any reason that's your fault and make the page clearly worth indexing. Pages move from discovered to crawled to indexed. "Discovered - currently not indexed" stops before Google fetches the page; "Crawled - currently not indexed" stops after the fetch, when Google decides whether to keep it. Only indexed pages can appear in results. Step 1: Inspect one affected URL In Search Console, paste the URL into the top search bar (URL Inspection), then click Test live URL . Check: Indexing allowed? If it says no, there's a noindex in a meta tag or an X-Robots-Tag header. User-declared vs Google-selected canonical. If they differ, Google thinks another URL is the main version. View tested page → HTML. Is your main content in the HTML Google rendered? If it's empty or a loading spinner, JavaScript rendering is the problem. Page fetch and HTTP status: must be successful with a 200. From a terminal you can check the same signals quickly: curl -sI https://www.example.com/services/web-design | grep -i -E '^HTTP|x-robots-tag|location' curl -s https://www.example.com/services/web-design | grep -o -i -E '<link rel="canonical"[^>]*>|<meta name="robots"[^>]*>' Step 2: Fix the technical causes Wrong canonical The most common one I find, and the one that hit this site. A canonical set once in a shared layout or theme makes every page point at the homepage, so Google treats them as duplicates. Each page should declare its own full, final URL as canonical: correct protocol, correct www or non-www, no redirect. If the canonical points at a URL that redirects, fix it to the destination. Accidental noindex Staging settings copied to production, a CMS checkbox ("discourage search engines"), or an SEO plugin rule on a category. Remove it, then request indexing. Soft 404 and thin pages Pages that return 200 but say "no results", "coming soon" or have a few lines of text look like empty pages. Real missing pages should return 404; thin pages need real content or should be merged. Content only visible after JavaScript If the HTML Google receives is an empty shell, indexing is slower and less reliable. Render important pages on the server. Next.js server components and static generation do this by default. Duplicate URLs The same content at ?page=1 , with tracking parameters, with and without a trailing slash, or under both http and https. Pick one, redirect or canonicalise the rest, and only list the main URL in the sitemap. Step 3: Give Google a reason to index it If the technical checks pass and it's still "Crawled - currently not indexed", Google doesn't see enough value yet. What moves it: Answer the query in the first paragraph and cover it more completely than the pages already ranking. The same structure gets pages quoted by AI answers; see how to get cited by ChatGPT and AI Overviews . Add first-hand detail: your real process, prices you can stand behind, screenshots, the problems you actually solved. That's what generic pages lack. Merge near-duplicates. Five thin location or service pages with swapped city names are weaker than one strong page. Link to it internally from the homepage, the relevant service page and related articles, with descriptive anchor text. An orphan page tells Google it isn't important to you either. Step 4: For "Discovered", make crawling easier Speed up the server. Google slows crawling when responses are slow or error out. A slow TTFB on a cheap host or an overloaded VPS shows up here. Cut low-value URLs: filter and sort parameters, tag pages, internal search results. Block or noindex them so crawl attention goes to pages that matter. Keep the sitemap clean: only canonical, indexable, 200-status URLs, with accurate lastmod dates. Link from pages that are already crawled often, such as the homepage and blog index. Step 5: Request indexing and validate After a real change, use Request indexing in URL Inspection for the most important pages (there's a daily limit, so prioritise). For a fix that affects many URLs, open the issue in the Page indexing report and click Validate fix . Google's help says validation typically takes up to about two weeks. For Bing, which also feeds ChatGPT search, submit changed URLs through IndexNow; this site pings it automatically after every publish. Lots of pages dropping out at once, or the whole site missing, is a different problem: start with the website down checklist to rule out outages, robots.txt and server errors. If your site was built on a drag-and-drop builder and you're fighting its SEO limits, my comparison of AI website builders vs a custom site may help you decide what to do next. Frequently asked questions How long does it take for "Crawled - currently not indexed" to resolve? There's no fixed time. Some pages are indexed within days of improvements and internal links; others never are if Google keeps judging them low value. If nothing changes after a few weeks, improve or merge the page rather than waiting. What's the difference between "Discovered" and "Crawled - currently not indexed"? Discovered means Google knows the URL but hasn't fetched it yet. Crawled means Google fetched it and decided not to index it for now. The first is usually about crawl priority and server speed, the second about duplication or content quality. Does requesting indexing again help? Not for an unchanged page. Request indexing once after you've fixed something or substantially improved the content, then give Google time to re-evaluate. Can a wrong canonical tag stop pages being indexed? Yes. If a page's canonical points to another URL, Google may treat it as a duplicate and index only the other one. Every page should point its canonical at its own final URL unless it really is a duplicate. Should I delete pages that aren't indexed? Delete or merge pages that have no value for visitors, such as empty tag pages or near-duplicate service pages. Pages that matter for your business should be improved and linked, not deleted. Your key pages still not on Google? I audit and fix the technical SEO of business websites, from canonicals and rendering to internal links and page content, and build sites that get indexed and cited from day one. See my website design services or contact me with your domain and a screenshot of the Page indexing report.

Read article →
Linux OOM Killer Killed Your App? Find and Fix It
Linux System AdminOct 6, 2026

Linux OOM Killer Killed Your App? Find and Fix It

If your app dies with no error, the Linux OOM killer probably ended it. Run journalctl -k | grep -i "killed process" to confirm, then add swap and cap memory. The symptom is confusing: the app stops, there's no stack trace, the app log just ends, and Nginx starts returning 502s. Often it happens at the busiest time of day, or during a deploy. The reason is that the kernel ran out of memory and picked a process to kill so the whole server wouldn't freeze. Your app, usually the biggest process on a small VPS, was the pick. My own sites run a Next.js frontend and a NestJS API on one Ubuntu server, and this is the checklist I use to confirm it and keep it from happening. Key takeaways Confirm it in the kernel log: journalctl -k | grep -i -E 'killed process|out of memory' . The app's own log won't show it. The OOM killer picks the process with the highest score, which is mostly the one using the most memory. That's why it's usually your app or database. Add swap on small VPSs. It turns a sudden kill into a slowdown you can see and fix. Cap memory where it grows: Node's --max-old-space-size , PM2's max_memory_restart , Docker or systemd memory limits. Don't build on a tiny production server. next build or npm install on a 1-2 GB box is a classic trigger. Build in CI. Step 1: Confirm the OOM killer did it sudo journalctl -k --since "24 hours ago" | grep -i -E 'killed process|out of memory|oom-kill' # or, if the journal isn't persistent: sudo dmesg -T | grep -i -E 'killed process|out of memory' A hit looks like this (example): Out of memory: Killed process 4321 (node) total-vm:2412344kB, anon-rss:1489012kB, ... The name in brackets is what died, and anon-rss is roughly how much RAM it was using. If you see Memory cgroup out of memory instead, a container or systemd memory limit was hit rather than the whole server. In Docker that shows up as exit code 137 with OOMKilled: true ; my post on containers that keep restarting covers that side. Step 2: See where the memory goes free -h ps aux --sort=-%mem | head -n 10 systemd-cgtop -m # memory per service pm2 monit # if you use PM2 Look at the available column in free -h , not free . Linux uses spare RAM as disk cache, so "free" is always low on a healthy server. If available sits near zero and swap is 0B or full, you have your answer. Then ask what grew. A process that climbs steadily for hours is a leak. One that spikes is a heavy request: image processing, a big export, a query that loads a whole table into memory, or a build. When an app keeps growing and there's no free RAM or swap left, the kernel kills the process with the highest OOM score, which is usually the largest one. The app disappears without writing an error, and the site goes down. Step 3: Get the site back Restart the app ( pm2 restart app , systemctl restart app or docker compose up -d ) and check Nginx stops returning 502s. If it was killed without a process manager to bring it back, set one up now: PM2 with pm2 startup and pm2 save , or Restart=always in a systemd unit. Then fix the cause, or it will happen again at the next peak. Step 4: Add swap (if you have none) Many VPS images ship with no swap. A modest swap file gives the kernel somewhere to put idle pages and buys time before anything is killed: sudo fallocate -l 2G /swapfile sudo chmod 600 /swapfile sudo mkswap /swapfile sudo swapon /swapfile echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab # prefer RAM, use swap as a safety net echo 'vm.swappiness=10' | sudo tee /etc/sysctl.d/99-swappiness.conf sudo sysctl --system Check disk space first ( df -h / ); a swap file on a nearly full disk causes a different outage, see "No space left on device" . Swap is a safety net, not more RAM. If the server lives in swap all day, it needs a bigger plan or a smaller app. Step 5: Cap memory where it grows Node.js and PM2 Node's heap limit stops one process from taking everything, and PM2 restarts it cleanly before the kernel has to step in: // ecosystem.config.js module.exports = { apps: [{ name: "web", script: "node_modules/next/dist/bin/next", args: "start -p 3000", node_args: "--max-old-space-size=768", max_memory_restart: "900M", }], }; A restart at 900 MB is a quick, logged blip; an OOM kill is an outage with no explanation. Keep the restart threshold above the heap limit so the heap limit catches normal growth first. If it restarts every hour, you've confirmed a leak and the restart buys you time to find it. Docker and systemd # docker-compose.yml services: api: mem_limit: 768m # systemd: sudo systemctl edit app [Service] MemoryHigh=700M MemoryMax=800M A limit means only that container or service is killed, not whatever the kernel picks across the whole server. Databases Postgres and MySQL are sized at startup. On a small shared server, check shared_buffers and work_mem (Postgres) or innodb_buffer_pool_size (MySQL). A buffer pool tuned for a dedicated 8 GB server will starve everything else on a 2 GB VPS. Each extra database connection also costs memory, another reason not to run hundreds of them (see Postgres "too many clients" ). Step 6: Stop building on the production server Installing dependencies and building a Next.js app can briefly use more memory than the running site. On a 1-2 GB server that alone triggers the OOM killer, and it can take the live site with it. Build in CI and ship the finished release, as in my GitHub Actions deploy with auto-rollback and the Next.js on a VPS guide . Protect the processes that must survive You can tell the kernel to avoid certain services. In a systemd unit, OOMScoreAdjust=-500 makes it a much less likely target. Use it sparingly, for things like the database. Setting everything to "never kill" just means the kernel has nothing left to pick and the whole server hangs. Frequently asked questions How do I know if the OOM killer killed my process? Search the kernel log with journalctl -k or dmesg -T for "Killed process" or "Out of memory". The line names the process, its PID and how much memory it used when it was killed. Why does the OOM killer pick my app and not something else? It kills the process with the highest OOM score, which mostly follows how much memory a process uses. On a small server the main app or the database is usually the largest process, so it's the first target. How much swap should a VPS have? For a small web server, 1-2 GB of swap is a sensible safety net, with swappiness set low so RAM is preferred. Swap should absorb short spikes; if it's in constant use, the server needs more RAM or the app needs less. Is it safe to disable the OOM killer? No. Without it, a server that runs out of memory can freeze completely and need a hard reboot. Set memory limits and protect key services instead. Why does my server run out of memory during deploys? Installing packages and building the app use a lot of memory on top of the running site. Build in CI or on a separate machine and copy the finished build to the server. Server running out of memory? I find what's eating the RAM on Linux servers, fix the leak or the config, and set up limits, swap and alerts so the next spike doesn't take the site down. See my Linux system admin services or contact me now with the "Killed process" line from your kernel log.

Read article →
Docker Container Keeps Restarting? How to Fix It
DevOpsOct 6, 2026

Docker Container Keeps Restarting? How to Fix It

A Docker container that keeps restarting is crashing on start. Run docker logs and docker inspect to get its exit code; that tells you the cause in seconds. docker ps shows Restarting (1) 4 seconds ago , the site is down, and restarting it again changes nothing. The restart policy is doing its job: the process inside dies, Docker starts it again, it dies again. The fix is never in the restart, it's in why the process exits. I run my own API and its database in Docker Compose, and this is the routine I use, in the order that finds the cause fastest. Key takeaways Logs survive restarts: docker logs --tail 100 <name> shows the crash from the last run even mid-loop. The exit code narrows it down: 1 = app error, 127 = command not found, 137 = killed (often out of memory), 139 = segfault, 0 = the process simply finished. Pause the loop to debug: docker update --restart=no <name> , or open a shell with a different entrypoint. The usual causes: missing env vars, the database not ready yet, wrong CPU architecture, volume permissions, memory limits and a main process that goes to the background. "Unhealthy" doesn't restart anything in plain Docker or Compose. Only Swarm replaces unhealthy containers. Step 1: Read the logs and the exit code docker ps -a --filter name=api docker logs --tail 100 api docker inspect -f 'exit={{.State.ExitCode}} oom={{.State.OOMKilled}} restarts={{.RestartCount}} err={{.State.Error}}' api With Compose, docker compose logs --tail 100 api and docker compose ps -a do the same. Watching the loop live also helps: docker events --filter container=api prints every die and start with the exit code. Step 2: Match the exit code Exit 1 (or another small number): the application threw an error and quit. The last log lines say what. Most often a missing or wrong environment variable, a database or Redis it can't reach, or a migration that failed. Exit 0: nothing crashed, the main process just finished. Common when the command starts a daemon in the background ( service nginx start , npm start & ) or runs a one-off script. A container lives only as long as its main process, so that process must stay in the foreground. Exit 126 / 127: the command isn't executable or doesn't exist. Check the CMD / ENTRYPOINT path, the script's execute bit, and Windows line endings in shell scripts ( \r breaks the shebang). Exit 137: the process got SIGKILL. If OOMKilled is true , it hit a memory limit. If not, something else killed it: a manual docker kill , or the kernel when the whole host ran out of memory. Exit 139: segmentation fault, usually a native module built for a different libc or CPU (Alpine musl vs glibc images are a classic). Exit 143: SIGTERM. Something stopped it on purpose, such as a deploy script or an orchestrator. Docker restarts a crashed container after a delay that doubles each time, starting at 100 ms. The loop itself is working as designed; the answer is in the exit code and the logs it leaves behind. Step 3: Stop the loop and get a shell Debugging is easier when the container isn't restarting under you: docker update --restart=no api # pause the policy docker compose run --rm --entrypoint sh api # same image, env and volumes, a shell instead of the app # inside: check env, files, and run the start command by hand env | sort ls -la /app node dist/main.js Running the start command by hand inside the container usually shows the error with more context than the logs. The six causes I see most 1. Missing or wrong environment variables The app reads DATABASE_URL or a secret at boot and exits when it's missing. Typical after a deploy where .env wasn't copied, a variable was renamed, or env_file points at the wrong path. Check with docker compose config , which prints the final resolved values. 2. The database isn't ready yet depends_on alone only waits for the database container to start , not to accept connections. Add a healthcheck to the database and wait for it: services: api: build: . restart: unless-stopped depends_on: db: condition: service_healthy db: image: postgres:16 healthcheck: test: ["CMD-SHELL", "pg_isready -U app"] interval: 5s retries: 10 If the API restarts later with "too many clients", that's a different problem: see fixing Postgres "too many clients already" . 3. "exec format error": wrong CPU architecture An image built on an Apple Silicon Mac (arm64) won't run on most x86 VPSs, and the log shows exec format error . Build for the server's platform: docker buildx build --platform linux/amd64 -t myapp . , or build in CI on the right architecture. 4. Volume permissions Images that run as a non-root user (Postgres, many Node images, anything with USER in the Dockerfile) crash with "permission denied" when a bind-mounted host folder belongs to another UID. Check with ls -ln on the host and id inside, then chown the folder to the container's UID. 5. Out of memory (exit 137) A mem_limit that's too small, or a host that's out of RAM. Raise the limit or fix the leak; for the host side, my guide to the Linux OOM killer shows how to confirm it and stop it. 6. Port or file conflicts at start The app can't bind its port inside the container because two processes try to use it, or it finds a stale PID or lock file in a volume and refuses to start. The log says so plainly; delete the stale file or fix the duplicate process. About healthchecks and "unhealthy" A failing HEALTHCHECK marks the container unhealthy , but Docker Engine and Compose don't restart it for that. The restart policy only reacts when the main process exits. If you need automatic recovery from a hung (not crashed) app, make the app exit on fatal errors, or run a small watcher that restarts unhealthy containers. Swarm and Kubernetes handle this natively. Stop it happening on the next deploy Health-check after every deploy and roll back automatically. My GitHub Actions deploy with auto-rollback does this, so a crash-looping release never stays live. Validate config at build time: fail the CI job if required env vars are missing instead of finding out on the server. Don't publish database ports to the internet while you're at it: Docker bypasses UFW. See why Docker ignores UFW and how to fix it . Watch restart counts: any container with a rising RestartCount deserves an alert. Frequently asked questions Why does my Docker container keep restarting? Its main process exits, and the restart policy (always, unless-stopped or on-failure) starts it again. Check docker logs and the exit code with docker inspect to see why it exits; the cause is almost always in the app's configuration or environment. What does exit code 137 mean in Docker? The process was killed with SIGKILL. If docker inspect shows OOMKilled as true, the container hit its memory limit; otherwise it was killed by the host's OOM killer or by a manual docker kill. How do I see logs of a container that keeps restarting? docker logs works on a restarting container and keeps output from previous runs, so docker logs --tail 100 name shows the last crash. With Compose, use docker compose logs for the service. How do I stop a Docker restart loop? Run docker update --restart=no with the container name to disable the policy, or docker stop it. Then start a shell with docker compose run --rm --entrypoint sh to debug with the same image and environment. Does Docker restart unhealthy containers? No. Outside Swarm, a failing healthcheck only changes the status to unhealthy. Restart policies trigger when the process exits, so the app has to exit or an external tool has to restart it. Container stuck in a restart loop right now? I debug and fix Docker and Compose setups on Linux servers, then add the health checks, rollbacks and alerts that keep them up. See my DevOps services or contact me now with the output of docker logs and the exit code.

Read article →