fork retry Resource Temporarily Unavailable With Eight Gigabytes Free

Share
fork retry Resource Temporarily Unavailable With Eight Gigabytes Free. Abstract error autopsy illustration in orange and dark grey on debugly.dev

The worker crashed with fork: retry: Resource temporarily unavailable, and the monitoring showed eight gigabytes of memory free, which made the error read as a lie. Memory was not the resource in question. The error is EAGAIN from fork, and fork fails with EAGAIN when the caller is not allowed to create another task, which is a count limit, not a memory limit.

This is one of the most misleading error messages in Linux, because it says "resource" and everyone's first resource is RAM, and the second is file descriptors, and the actual resource is the number of threads and processes, which has its own ceilings that memory monitoring never shows.

This was a containerised worker on Docker 27.5 with cgroup v2, Linux 6.8, and the ceilings were pids.max plus the process rlimit.

The short answer

fork creates a new task, and the kernel enforces several independent ceilings on task creation: the per cgroup pids.max, the per user process rlimit (nproc), and system wide limits. When the relevant count is at its ceiling, fork returns EAGAIN, which surfaces as "resource temporarily unavailable", and retrying does nothing because the count does not drop until tasks exit.

Eight gigabytes free is consistent with all of these, because none of them is memory. The machine can be mostly idle and still refuse to fork.

Why the count, not the memory

A fork copies the process's task structure, not its memory, and the copy on write pages are not allocated up front. So the cost of a fork is a task slot, and the limits exist to stop a runaway from creating unbounded tasks, which is a much faster way to kill a host than memory exhaustion. The kernel is defending the task table, and the cgroup is defending the host from one container's thread storm.

The container flavour is the one that bites in production. A container with pids.max set to a few hundred, running a runtime that spawns a thread per request or a child per job, hits the ceiling under load while using a fraction of its memory. The error is the cgroup doing its job, exactly as the OOM killer chooses a victim is the memory controller doing its job. Different resource, same pattern.

Finding which ceiling

The error does not say which limit fired, so check them in order.

The cgroup pids current and max:

cat /sys/fs/cgroup/pids.current
cat /sys/fs/cgroup/pids.max

If current equals max, that is your ceiling, and the fix is raising pids.max or, better, finding the thread leak.

The process rlimit:

ulimit -u

And the current count of tasks for the user, from ps or the proc filesystem. If the user is at nproc, that is the ceiling.

The thread count of the suspect process, which is often the real story:

ps -o nlwp= -p PID

A process with thousands of threads is not being limited unfairly. It is leaking threads, and the limit is the symptom that saved the host.

The thread leak shapes

A few shapes produce a runaway task count, and they are worth recognising.

A thread per unit of work with no pool. A job runner that spawns a child or thread per task and joins lazily, or not at all, grows with the queue depth. Under a backlog it creates tasks faster than they exit and hits the ceiling, at which point even the cleanup code cannot fork to run, which is a nasty self locking failure.

Unbounded concurrency in an async runtime. A task spawner that launches a task per request without a semaphore is the same shape in a runtime that maps tasks to threads, and the ceiling arrives as fork or thread creation failures at the runtime's edge.

Zombie children. A parent that spawns children and never reaps them accumulates zombie tasks that count against the limits. The children are dead but their task slots are not freed until reaped, so the count stays high with no live work, which is a particularly confusing variant because nothing looks busy.

The fixes

The immediate relief is raising the ceiling, and it is legitimate, because ceilings are often set conservatively. But raise with a number derived from observed peak, not by removing the limit, because the limit is the only thing between a leak and a host level failure.

The real fix is bounding concurrency. A worker pool with a fixed size, a semaphore on the task spawner, and reaping of children turn the task count from a function of the backlog into a constant, and constants do not hit ceilings.

And add the metric that memory monitoring lacks: the task count per container and per process, next to pids.max. The gap between current and max is your headroom, and alerting on the ratio catches the leak while there is still headroom, instead of discovering it as a fork failure at the peak.

The rule

"Resource temporarily unavailable" from fork names a count, not a quantity of memory. Check pids.current against pids.max, nproc against the user's task count, and the thread count of the suspect process, and treat a process at the ceiling as a leak until proven otherwise.

The host is refusing to create a task because some count you are not graphing has hit a ceiling you are not watching. Graph the count, bound the concurrency, and the eight free gigabytes stop being a contradiction.

The CPU sibling of this, where a limit bites as a stall rather than a refusal, is your container has two CPUs and is using one point three. And the case where the ceiling is hit by design, because the workload genuinely needs more tasks than the default allows, is also real: batch and fan out workloads legitimately create many short lived children, and for them the fix is a pids.max sized to the measured peak concurrency plus margin, declared in the deployment rather than discovered at the peak, so the ceiling is a stated capacity decision instead of an incident.