Skip to content

hardware: give up on an instance type without capacity after 10 minutes - #2478

Merged
alexey-milovidov merged 1 commit into
mainfrom
hardware-capacity-wait
Oct 9, 2026
Merged

alexey-milovidov merged 1 commit into
mainfrom
hardware-capacity-wait

Conversation

@alexey-milovidov

Copy link
Copy Markdown
Member

In the compute/memory/storage sweep (run 37873181224), c8id.96xlarge returned InsufficientInstanceCapacity for ~50 minutes. hardware/run-benchmark.sh retried it indefinitely, and since the workflow launches machines one after another, the job hit its 55-minute limit with the remaining 50 machines unlaunched.

Now InsufficientInstanceCapacity is retried for capacity_wait seconds (default 600) and then the instance type is reported as failed to launch, so the loop moves on. Quota (VcpuLimitExceeded etc.) and throttling errors are still retried without a limit, as they clear when other machines finish.

🤖 Generated with Claude Code

In the compute/memory/storage sweep, c8id.96xlarge got
InsufficientInstanceCapacity for 50 minutes, the launcher retried it forever,
and the job hit its time limit with 50 machines after it unlaunched.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@alexey-milovidov alexey-milovidov self-assigned this Oct 9, 2026
@alexey-milovidov
alexey-milovidov merged commit 7f96fbf into main Oct 9, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant