Use the Nx Cloud Public API to identify flaky tasks and collect evidence for a weekly review. An artificial intelligence (AI) agent can inspect the evidence and propose a pull request. The API itself doesn't modify code or open pull requests.
Select a settled period
Section titled “Select a settled period”Choose seven completed days in Coordinated Universal Time (UTC). Set windowStartAfter to the first day and windowEndBefore to the last day. Do not include the current day, because its metric window can still change.
Use the authentication and path discovery from the request guide. Discover the metric path from the downloaded specification:
FLAKY_TASKS_PATH=$(jq -r '.paths | keys[] | select(endswith("/flaky-tasks"))' openapi.json)
curl --fail-with-body --silent --show-error --get \ "$NX_CLOUD_HOST$FLAKY_TASKS_PATH" \ --header "Authorization: Bearer $NX_CLOUD_ACCESS_TOKEN" \ --data-urlencode "windowStartAfter=2026-09-21" \ --data-urlencode "windowEndBefore=2026-09-27" \ --data-urlencode "limit=100" \ --output flaky-tasks.jsonReplace the example dates with your review period. Follow nextCursor until you have every page in that period. Keep the dates and other filters unchanged. Add a project filter if you only need one project.
Prioritize five tasks
Section titled “Prioritize five tasks”Group rows by project, target, and configuration. Each row represents one daily window, not one task's aggregate for the whole week.
For each group, calculate a sample-weighted rate:
weekly rate = sum(flakinessRate * sampleSizeFlakinessRate) / sum(sampleSizeFlakinessRate)Report the denominator alongside the rate. A task with one sample should not outrank a frequently retried task without additional evidence. Select a minimum sample size suitable for your workspace, then rank eligible groups by the weighted rate. Use totalReruns to show the retry workload and deflakedAutomatically to show automatic recovery.
Do not assume the first five API rows are the five flakiest tasks. The API pages by metric window, not by flakiness rank. Record your period, sample threshold, page count, and ranking method in the report.
Inspect contributing executions
Section titled “Inspect contributing executions”For each selected task:
- Follow each metric row's
links.executionsand fetch all its pages. - Deduplicate execution IDs across the selected daily windows.
- Follow
links.runfor run metadata andlinks.tasksfor the matching task rows. - Inspect the task's final result and
priorAttempts. - Download the final and earlier failed logs through their asset URLs.
Earlier task attempts use an attempt index that starts at zero, oldest-first. Preserve batchId when the task ID repeats within a run. Follow the asset download rules. Do not send tokens to the signed storage URL.
A retry alone doesn't prove a task is flaky. Compare failures and successes for the same computation, including the recorded hash and task configuration. Inspect the logs for a repeatable cause, such as shared ports, timing assumptions, or a resource limit.
Some contributing execution links can return 404. Report those missing records instead of treating them as successful attempts.
Ask an agent for a focused fix
Section titled “Ask an agent for a focused fix”Give the agent the selected task, metric period, sample count, relevant attempt logs, and repository revision. Ask it to reproduce the failure and propose the smallest change supported by that evidence.
Require these outputs before a pull request:
- The failure signature and its supporting execution IDs.
- A reproduction or an explicit statement that reproduction failed.
- A focused code change and relevant test results.
- A clear distinction between verified findings and hypotheses.
Do not let the agent change cache settings because a metric suggests flakiness. API records don't establish the current Nx target configuration. Inspect the matching repository before a configuration change.
Review each proposed pull request before merge. Compare the next settled period with the same filters to check whether the task improved.