Kubernetes troubleshooting

OOMKilled in Kubernetes: Fix Exit Code 137

OOMKilled means the container used more memory than its limit, and the kernel killed its process. The kubelet records the reason and exit code 137. Raising the limit is sometimes right, but first confirm the workload really needs that memory.

By , Founder at SeaGit

·

Key takeaways

  • OOMKilled with exit code 137 means the memory limit was exceeded. It is not a node problem unless the node itself ran out of memory.
  • CPU limits throttle a container. They never cause OOMKilled.
  • A JVM or Node.js process can exceed its container limit when its heap settings are not capped below the limit.
  • Size the memory limit from observed peak usage, not from a guess. Keep the request near typical use so the scheduler packs nodes sensibly.

Confirm it is OOMKilled

Exit code 137 also appears when a container is killed for other reasons, so check the reason field, not only the code.

bash
kubectl describe pod my-api-6d8f9b7c5-x2k4q -n prod | grep -A6 'Last State'

# The same fields as JSON, for every container in the namespace
kubectl get pods -n prod -o custom-columns=\
NAME:.metadata.name,\
REASON:.status.containerStatuses[*].lastState.terminated.reason,\
EXIT:.status.containerStatuses[*].lastState.terminated.exitCode,\
RESTARTS:.status.containerStatuses[*].restartCount
Last State: Terminated, Reason: OOMKilled, Exit Code: 137 is the signature. Restart count tells you how often it happened.

To see whether the node ran out of memory rather than the container, look for node conditions and recent events:

bash
kubectl describe node ip-10-0-12-34.eu-west-1.compute.internal | grep -A8 'Conditions:'

kubectl get events -A --field-selector reason=OOMKilling --sort-by=.lastTimestamp | tail -20
A kubelet OOMKilling event names the container. System OOM events on the node are a different problem, usually from too many pods without memory requests.

Compare the limit with real usage

Read the configured request and limit, then compare them with the memory the container actually uses. kubectl top needs metrics-server, which most managed clusters install by default.

bash
# Configured requests and limits
kubectl get pod my-api-6d8f9b7c5-x2k4q -n prod \
  -o jsonpath='{range .spec.containers[*]}{.name}{" request="}{.resources.requests.memory}{" limit="}{.resources.limits.memory}{"\n"}{end}'

# Current usage per container
kubectl top pod my-api-6d8f9b7c5-x2k4q -n prod --containers
A limit that is close to the steady-state usage will be exceeded during a spike, so OOMKilled shows up under load rather than at idle.

A single reading is not enough. Watch the working set over a day, or use the Prometheus metric container_memory_working_set_bytes for that container, and take the peak. Set the limit above the peak with headroom, then re-check after the next release.

Size the request and the limit

The request is what the scheduler reserves on a node. The limit is the ceiling the kernel enforces. A common starting point is a request near the typical working set and a limit near the observed peak plus headroom. Adjust it to your own numbers.

yaml
resources:
  requests:
    cpu: 250m
    memory: 512Mi
  limits:
    memory: 768Mi   # above the observed peak, not the average
    # no cpu limit: CPU limits throttle, they do not kill
An example for one API container. The numbers are illustrative, so measure yours first.

Do not raise the limit without limit. A limit that keeps climbing usually hides a leak. If usage grows steadily after each restart, fix the leak in the code and keep the limit.

Cap the runtime heap below the limit

A runtime that sizes its heap from the host, not the container, will grow past the limit. Cap the heap so the process stays under the container limit, and leave room for stacks, native memory and buffers.

Java

Modern JVMs read the container limit. Set a percentage instead of a fixed size, so the heap follows the limit:

yaml
env:
  - name: JAVA_TOOL_OPTIONS
    value: "-XX:MaxRAMPercentage=70.0"
With a 768Mi limit, a 70 percent heap leaves about 230Mi for metaspace, threads and native memory.

Node.js

V8 does not read the cgroup limit for its old-generation heap. Set the heap size explicitly, below the container limit:

yaml
env:
  - name: NODE_OPTIONS
    value: "--max-old-space-size=512"
Use a value well under the limit. Buffers and the code cache sit outside the heap.

Worker processes

A web server with several workers multiplies its memory. Four gunicorn or Puma workers that each use 200Mi need about 800Mi, which is above a 768Mi limit. Lower the worker count, or raise the limit to match the workers you run.

Verify the fix

  • Restart count stays flat for a full traffic cycle, not just a few minutes after the rollout.
  • Memory working set stays below the limit at peak, with a visible margin.
  • No new OOMKilling events in the namespace after the rollout.
bash
kubectl rollout status deployment/my-api -n prod
kubectl get pods -n prod -w
Watch the restart column during the first traffic peak after the change.

How SeaGit handles this

SeaGit sets the memory request and limit for an app from its resource settings, and writes them into the pod spec of that deployment. When an app is OOMKilled, SeaGit's pod diagnosis reports the container as killed for running out of memory, and it includes the limit that was in force. That tells you which number to change.

SeaGit does not pick a memory size for you, and it does not tune your runtime heap. The request, the limit and the heap cap are decisions for the app, and the steps above apply on any cluster.

Frequently asked questions

OOMKilled in Kubernetes: Fix Exit Code 137: FAQ

What causes OOMKilled in Kubernetes?

The container used more memory than its limit. The kernel killed the process, and the kubelet recorded the reason as OOMKilled with exit code 137. Common causes are a limit set below real peak usage, a memory leak, a runtime heap that is not capped, or several workers in one container.

Does a CPU limit cause OOMKilled?

No. CPU limits throttle a container, so it runs slower. Only the memory limit causes OOMKilled.

How do I know the right memory limit?

Measure the working set over a full traffic cycle and take the peak. Set the limit above that peak with headroom, then confirm the restart count stays flat after the rollout.

Why is my pod OOMKilled when the node has free memory?

The limit is per container, and the kernel enforces it no matter how much memory the node has free. A container can be killed while the node is idle.

Sources

Checked 10 October 2026.

  1. Resource management for pods and containers (Kubernetes documentation) — requests, limits, CPU throttling and memory limits that cause OOM kills
  2. Assign Memory Resources to Containers and Pods (Kubernetes documentation) — OOMKilled reason and exit code when a container exceeds its memory limit

Read next