Skip to main content

[Server] One machine-provisioner queue can block every machine class

Machine classes on one provider share one create queue. A class that keeps failing capacity blocks the others until it has its own provider.

There is no per-resource-class queue in .circleci/config.yml. All machine classes on one machine-provisioner provider share one create queue. Pending VM requests are taken oldest first, then EC2 RunInstances is called.

A request that fails with a retryable capacity error stays on that queue. It is not dropped. Newer machine jobs wait behind it, including classes that have capacity and would otherwise start. The jobs do not share VMs. They share the queue.

Docker and Nomad jobs are already on a separate path. A stall on the machine queue does not by itself stop Docker jobs.

A second provider does not add EC2 capacity. It gives the class that cannot place its own queue, so the rest of the machine fleet can move.

Example: a GPU class in front of general-purpose machines

One case is a GPU resource class whose instance type returns InsufficientInstanceCapacity in the subnets you use. Those GPU requests sit at the front of the single queue. A later linux.xlarge (or any other machine class on the same provider) waits behind them even though it would use a different instance type and a different VM. The same pattern applies to any class that keeps failing capacity: the instance family is the example, the queue is the cause.

Isolate the scarce class

Each provider name has its own create queue. The stock Helm machine_provisioner.providers.ec2 block produces a single provider. A second provider is a full document in machine_provisioner.custom_config (YAML in values, not a filesystem path). Keep the existing machine_provisioner.providers.ec2 settings for region, subnets, security group, and IRSA or keys. Those still authenticate the pods.

Reference:

Dump the live file first so AMIs, aliases, preboot, and display stay as they are:

kubectl -n  get configmap machine-provisioner-configmap -o jsonpath='{.data.config\.yml}'

Put that document under machine_provisioner.custom_config and make two edits:

  1. On the existing provider, remove the scarce resource classes from offerings.linux_amd64.resource_classes.

  2. Add a second provider with the same region, subnet IDs, and security group, offering only those classes.

Do not list the same class name on both providers. Overlapping names are selected by weight, which mixes the two queues again.

While capacity is short, the other ways to unblock the fleet are to ask AWS for that instance type in the subnets you use, try an instance type EC2 can place there, or pause jobs on the scarce class so they stop occupying the front of the single queue.

Did this answer your question?