Profile
Back to NewsBack
Dev.to 7 min
Reader Mode
Why Kubernetes Says ‘Insufficient CPU’ When the Node Looks Idle

Why Kubernetes Says ‘Insufficient CPU’ When the Node Looks Idle

1 day ago

Your deployment is stuck. The Pod has been Pending for several minutes, and the event looks conclusive:

0/3 nodes are available: 3 Insufficient cpu.

You open the dashboard. The nodes are using 30–40% CPU.

So where did the rest of the CPU go?

Nowhere. Kubernetes and your dashboard are simply answering different questions:

  • Your dashboard: How much CPU is being used right now?
  • The scheduler: How much CPU has already been accounted for, and can this Pod's request fit on one eligible node?

That difference—usage versus requests—is the entire mystery. But it leads to a few traps that are easy to miss in production.

The short answer

The Kubernetes scheduler does not place a Pod based on live CPU usage. It primarily uses the Pod's CPU request and each node's allocatable CPU.

In simplified form, a Pod fits when:

new Pod request
    <=
node allocatable CPU - CPU already requested on that node

Notice what is missing from that calculation: current CPU usage.

This is deliberate. If the scheduler packed nodes according to a quiet moment in a monitoring graph, several workloads could spike together and overwhelm the node.

A node can be idle and still be “full”

Imagine a node with 8 CPU cores allocatable to Pods:

Workload CPU request Current usage
API 2 CPU 0.7 CPU
Worker 2 CPU 0.5 CPU
Search 2 CPU 0.6 CPU
Total 6 CPU 1.8 CPU

Your monitoring system reports roughly 22.5% CPU usage. The node looks quiet.

Now you deploy another Pod requesting 3 CPU:

6 CPU already requested + 3 CPU requested = 9 CPU

The node has only 8 CPU allocatable, so the Pod does not fit. The scheduler reports Insufficient cpu even though actual usage is below 2 CPU.

There is no contradiction. The dashboard is showing activity; the scheduler is checking commitments.

One nuance matters here: a CPU request is not necessarily a dedicated core locked away for that container. Think of it as an entry in the scheduler's capacity ledger. At runtime, it also influences the container's CPU weight when the node is contended.

Request, limit, and usage are three different numbers

Consider this container:

resources:
  requests:
    cpu: "500m"
  limits:
    cpu: "2"

In Kubernetes CPU units, 1000m is 1 CPU, so 500m is half a CPU.

The three numbers mean:

Number What it answers
Request How much CPU should the scheduler account for when placing the Pod?
Limit How much CPU time may the container consume before it is throttled?
Usage How much CPU is the container consuming at this moment?

The scheduler uses the request for placement. It does not use kubectl top as a live packing signal, and a higher CPU limit does not normally require that whole limit to be available before scheduling.

This also explains two opposite failure modes:

  • Requests set too high: Pods are hard to schedule and the cluster looks underused.
  • Requests set too low: Pods schedule easily, but too many may land together and fight for CPU during a traffic spike.

The goal is not to make requests as small as possible. The goal is to make them honest.

Capacity is not allocatable CPU

A 16-core machine does not always offer all 16 cores to normal Pods.

Kubernetes distinguishes between:

  • Capacity: the resources present on the node.
  • Allocatable: the portion available to Pods after reservations for the operating system, kubelet, container runtime, and eviction thresholds.

Check both with:

kubectl describe node <node-name>

Look for output similar to:

Capacity:
  cpu: 16
Allocatable:
  cpu: 14

For scheduling, 14 CPU is the useful number.

Further down, kubectl describe node includes an Allocated resources section. Its requests column is much more relevant to this incident than the node's live utilization graph:

Allocated resources:
  Resource  Requests       Limits
  cpu       12500m (89%)   22000m (157%)

Seeing CPU limits above 100% can be valid because CPU limits may be overcommitted. The requests column is the important scheduling signal.

Debug it in five minutes

1. Read the complete scheduling event

kubectl describe pod <pod-name> -n <namespace>

Do not stop after reading Insufficient cpu. The complete event may say:

0/6 nodes are available:
2 Insufficient cpu,
2 node(s) had untolerated taint,
2 node(s) didn't match Pod's node affinity.

In that case, CPU is only part of the story. Your Pod effectively has just two CPU-eligible nodes, not six.

2. Inspect the Pod's actual request

kubectl get pod <pod-name> -n <namespace> -o yaml

Check resources.requests.cpu for every container, including sidecars. A Pod's normal-container CPU request is generally the sum of its containers' requests.

Also inspect init containers. Because init containers run sequentially, Kubernetes uses the largest relevant init-container request when calculating the Pod's effective request, rather than simply adding all init containers together. A heavyweight migration or setup container can therefore make a seemingly small application Pod difficult to schedule.

3. Check each eligible node's allocatable and requested CPU

kubectl describe node <node-name>

Compare:

allocatable CPU
- existing CPU requests
= schedulable headroom

Then compare that headroom with the new Pod's effective request.

4. Use live usage for diagnosis—not for placement math

kubectl top nodes
kubectl top pods -A --sort-by=cpu

These commands help answer a different but valuable question: are requests aligned with reality?

If a service requests 2 CPU per replica but stays near 100m even during representative peaks, its request may be oversized. Do not resize it from one quiet snapshot, though. Look at a meaningful time window, peak traffic, startup behavior, batch jobs, and latency under CPU contention.

Four traps that make this more confusing

1. Free CPU is fragmented across nodes

Suppose three nodes each have 2 CPU of schedulable headroom. The cluster has 6 CPU free in total, but a single Pod requesting 3 CPU still cannot run.

A Pod must fit on one node. Cluster-wide free CPU cannot be pooled across nodes for one Pod.

2. A namespace policy may add requests for you

A LimitRange can inject default requests and limits when a workload omits them. Also, when a container specifies a CPU limit but no request, Kubernetes can assign a request equal to that limit.

Check for defaults with:

kubectl get limitrange -n <namespace> -o yaml

Always inspect the created Pod, not only the Helm values or Deployment template you expected to be applied.

3. DaemonSets consume space on every node

Logging agents, security agents, CNI components, and monitoring exporters often run as DaemonSets. Their requests count too. A small per-node request becomes significant when node sizes are small or the list of agents grows.

4. “Insufficient CPU” may appear beside other hard constraints

Node selectors, required affinity, pod anti-affinity, taints, topology spread rules, and volume topology can shrink the set of valid nodes before CPU is considered.

The cluster may have plenty of CPU—just not on a node your Pod is allowed to use.

What should you fix?

The event tells you the symptom, not the correct remedy. Choose the fix that matches the evidence:

  • Requests are unrealistic: Right-size them using historical usage, representative peaks, load tests, and latency or throughput goals. A Vertical Pod Autoscaler in recommendation mode can provide another data point.
  • The request is legitimate: Add or scale suitable nodes, or use larger nodes that can fit the Pod.
  • Capacity is fragmented: Rebalance workloads, reconsider placement rules, or provide a node shape large enough for the biggest Pod.
  • Only a narrow node pool is eligible: Revisit affinity, selectors, taints, and topology rules—or scale that specific pool.
  • A rollout temporarily needs extra room: Account for maxSurge and other deployment behavior in your capacity plan.

What you should not do is blindly lower the request until the Pod schedules. That can turn a visible scheduling problem into a less obvious production performance problem.

The mental model to keep

When a node looks idle but Kubernetes says Insufficient cpu, remember:

usage       = what is happening now
request     = what the scheduler plans around
allocatable = what the node offers to Pods

Then ask four questions:

  1. What is this Pod's effective CPU request?
  2. How much CPU is allocatable on each eligible node?
  3. How much CPU has already been requested there?
  4. What other rules make nodes ineligible?

Once you separate live usage from scheduler accounting, Insufficient cpu stops looking like a Kubernetes lie. The node can be quiet and still have no room for the promise your new Pod is asking Kubernetes to make.

Further reading

Chat with me