Jobs stay on Preparing environment when Nomad has servers and no clients. A common cause is DNS: Nomad clients dial the advertised RPC address, and that name is the public UI hostname, which does not accept port 4647.
CircleCI Server is designed so one load balancer — the Kubernetes service circleci-proxy (with ACM, circleci-proxy-acm) — is where both the UI and Nomad RPC land. nginx on that service sends HTTP(S) to Kong and the frontend, and TCP 4647 to the Nomad servers.
What you see
nomad node statusreports no nodes, or nodes appear and then go down after missed heartbeats.Nomad servers are healthy and have a leader.
From a machine-provisioner VM,
ncto the public address of the UI hostname on port 4647 times out.The same check to the internal load balancer addresses on port 4647 connects.
The security group for 4647 is open in that second case. The client is dialing the wrong address.
Why
Nomad RPC advertise and retry_join use the Server hostname, typically circleci.example.com:4647. If DNS for the machine subnet returns the public UI address for that name, clients never reach the internal listener. Five preboot VMs can exist while nomad node status stays empty.
What to change
From a VM in the machine subnet:
Resolve the UI hostname and the internal NLB hostname.
Check TCP 4647 against the public address and against the internal NLB addresses.
When only the internal addresses accept 4647, point that hostname at the internal NLB from the machine VPC. Split-horizon DNS, or a Nomad-specific name that retry_join uses, both work. The UI can keep using the public name from outside the VPC.
Reboot a client only after the name resolves to the internal addresses. A VM that registers and then misses heartbeats will show initializing instead of ready until RPC stays up.
/etc/nomad on the VM is not the file to edit. On machine instances the Nomad client is a systemd service. Server-side Nomad config stays on the Nomad server pods.
