Skip to content

UI: Info tooltip when a Dag's active runs exceed max_active_runs - #73693

Open
seanmuth wants to merge 19 commits into
apache:mainfrom
seanmuth:seanmuth/warn-icon-max-active-runs-header
Open

seanmuth wants to merge 19 commits into
apache:mainfrom
seanmuth:seanmuth/warn-icon-max-active-runs-header

Conversation

@seanmuth

@seanmuth seanmuth commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

The Dag header's "Active Runs" stat shows X of Y once a Dag has active runs, but nothing explains when there's more going on than that number shows, or what happens to a run that's stuck waiting. On main, X (active_runs_count) is RUNNING-only and capped by the scheduler's own promotion gate, so it can never actually exceed Y — meaning a Dag can have runs genuinely queued up behind the limit with zero indication of it anywhere in the UI.

This adds:

  • A new queued_runs_count field (#73692's sibling addition to DAGDetailsResponse, computed live in the same route, same pattern as the existing active_runs_count).
  • An info icon (FiInfo) with a tooltip next to the "Active Runs" label, shown whenever queued_runs_count > 0, explaining that runs beyond the limit won't start until an existing active run completes.
  • The queued count displayed directly in the stat itself — 1 of 1 (2 queued) — so the information doesn't require a hover at all.

Why queued_runs_count instead of exceeds_max_active_runs (#73692): that flag answers "is this Dag at or over capacity" (RUNNING + QUEUED >= max_active_runs), which is the right question for its original internal use (should the scheduler bother re-evaluating this Dag) but the wrong one for the UI — it's true even when nothing is actually queued (exactly at capacity, zero waiting), and it's a periodically-recomputed flag rather than a live value. queued_runs_count answers "are runs actually waiting right now," computed fresh on every request, so the UI's own correctness doesn't depend on the flag's refresh cadence at all.

Why an info icon rather than a warning: originally went with a warning triangle on the theory that X of Y showing X > Y is an inherently unusual, "illogical-looking" state. That's not accurate on current main — the number itself can never exceed the limit (see above), so the confusing state is invisible, not alarming-looking. An info icon fits the "here's some extra context" framing better than a warning would.

Verified against a live repro (schedule=None Dag, max_active_runs=1, three manual triggers) — screenshot in the PR conversation.

Depends on #73692 for the queued_runs_count and exceeds_max_non_backfill fix; built directly on top of it.

Part of #73686.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Sonnet 5

Generated-by: Claude Sonnet 5 following the guidelines

DagModel.exceeds_max_non_backfill has been persisted since 3.2.0
(migration 0099) and is already used internally by the scheduler to
avoid re-evaluating Dags that are already at their concurrency limit,
but it was never surfaced anywhere in the public API. There is
currently no way for a client to tell "is this Dag currently blocked
on max_active_runs" without independently computing active-run counts
against the limit.

This exposes it on GET /dags/{dag_id}/details as
exceeds_max_active_runs, aliased to the underlying model column via
the existing DAG_ALIAS_MAPPING mechanism.

Part of apache#73686.
The dag-processor's active-run-count calculation short-circuited to
0 for any Dag whose timetable can't be scheduled (e.g. schedule=None,
triggered manually or via the API), regardless of how many runs were
actually active. Since exceeds_max_non_backfill is recomputed on
every parse cycle, this made it permanently unreliable for exactly
the kind of Dag that most needs it: one triggered manually more
often than its max_active_runs allows never has a schedule to be
"scheduled" against, but still has a real concurrency limit.

Only the latest-run lookup (used for scheduling the next run) is
skippable for such Dags; the active-run count is not.
@seanmuth

Copy link
Copy Markdown
Contributor Author

there's behavior drift between this PR and the in-the-wild investigation that produced this (Airflow 3.2.2) (duh) which corrects a fair bit of the confusing parts, namely queued DRs don't count toward "active" DR anymore, so it should be functionally impossible to render "3 of 1" anymore. Reworking this improvement to drop the warning icon in favor of a less-scary info icon, then rendering "Active Runs: 1 of 1 (2 queued) (i)" (i) == info tooltip

screenshots to follow

@seanmuth
seanmuth force-pushed the seanmuth/warn-icon-max-active-runs-header branch from ec65457 to 2f7ed78 Compare September 24, 2026 21:30
@seanmuth

Copy link
Copy Markdown
Contributor Author

screens:
image

image

@seanmuth seanmuth changed the title UI: Warn when a Dag's active runs exceed max_active_runs UI: Info tooltip when a Dag's active runs exceed max_active_runs Sep 24, 2026
Comment thread airflow-core/src/airflow/ui/openapi-gen/requests/schemas.gen.ts
@ashb

ashb commented Sep 25, 2026

Copy link
Copy Markdown
Member

screens: image
image

Can you show that in a larger context? I can't quite work out where that is being shown.

Review feedback from apache#73692:

- "exceeds" was inaccurate: the flag is true once a Dag is at or
  above its limit, not only when strictly over it. Renamed the
  exposed field via the existing DAG_ALIAS_MAPPING mechanism; the
  underlying exceeds_max_non_backfill column is unchanged, since
  renaming a persisted column is a separate, more deliberate change.
- The dag-processor's active-run-count calculation was already an
  N+1 (one query per Dag per parse cycle) before this PR's fix to
  the non-schedulable-Dag case; that fix just extended the same
  pattern to more Dags. DagRun.active_runs_of_dags already accepts
  a batch of dag_ids, so hoist the call out of the per-Dag loop in
  update_dags into a single call across every Dag in the update.
@seanmuth
seanmuth force-pushed the seanmuth/warn-icon-max-active-runs-header branch from 2f7ed78 to be8af5d Compare September 25, 2026 16:13
… Dags

test_bulk_write_to_db_interval_save_runtime encoded the old, buggy
short-circuit this PR removes: active_runs_of_dags is now always
batched once per update, regardless of whether any Dag in the batch
can be scheduled, so exceeds_max_non_backfill stays accurate for
schedule=None Dags too.

Co-Authored-By: Claude <noreply@anthropic.com>
@seanmuth
seanmuth force-pushed the seanmuth/warn-icon-max-active-runs-header branch from be8af5d to d9c962d Compare September 25, 2026 17:02
@seanmuth

Copy link
Copy Markdown
Contributor Author

screens: image
image

Can you show that in a larger context? I can't quite work out where that is being shown.

Attached larger screen, main's UI def looks a bit different now!

image

@seanmuth

Copy link
Copy Markdown
Contributor Author

FWIW here's light mode with the Dark Reader extension turned off:

image

Comment thread airflow-core/src/airflow/api_fastapi/core_api/routes/public/dags.py Outdated
Comment thread airflow-core/src/airflow/ui/src/pages/Dag/Header.tsx Outdated
Comment thread airflow-core/src/airflow/ui/src/pages/Dag/Header.tsx Outdated
Comment thread airflow-core/src/airflow/ui/src/components/HeaderCard.tsx Outdated
Comment thread airflow-core/src/airflow/ui/src/pages/Dag/Header.test.tsx Outdated
Comment thread airflow-core/tests/unit/api_fastapi/core_api/routes/public/test_dags.py Outdated
seanmuth and others added 2 commits September 29, 2026 12:12
The N+1 fix moved DagRun.active_runs_of_dags from once per Dag to
once per persistence call, so it now counts toward FIXED_PER_CALL
instead of UNCHANGED_PER_DAG/REWRITE_PER_DAG.

Co-Authored-By: Claude <noreply@anthropic.com>
…tale scheduler cache

exceeds_max_non_backfill is a scheduler-side cache written on parse
and by a handful of scheduler events (scheduled-run creation, dagrun
timeout, run finish) -- but never by a manual/API/operator trigger.
For a schedule=None Dag, triggering a run left is_at_max_active_runs
false until the next parse cycle (up to min_file_process_interval,
30s by default), or for the run's entire lifetime if it finished
before that parse -- exactly the case this field exists to cover.
Marking a run failed through the API had the reverse problem.

get_dag_details now computes it fresh, the same way it already does
for active_runs_count/queued_runs_count, via
DagRun.active_runs_of_dags(exclude_backfill=True) compared against
max_active_runs. Removed the now-unused alias to
exceeds_max_non_backfill from DAGDetailsResponse; that column stays
as the scheduler's own internal optimization, just no longer exposed
through this field.

Documented via Field(description=...) that this counts differently
than active_runs_count: RUNNING+QUEUED with backfill runs excluded
(matching the scheduler's own promotion check), vs active_runs_count's
RUNNING-only that includes backfill runs.

Also simplified _RunInfo/_RunInfo.calculate in the dag-processor:
num_active_runs was only ever passed through unchanged since the N+1
fix moved its computation to the caller, so update_dags now reads it
directly from the batched query result instead of round-tripping it
through calculate's parameter and return value.

Co-Authored-By: Claude <noreply@anthropic.com>
@seanmuth
seanmuth force-pushed the seanmuth/warn-icon-max-active-runs-header branch from e16f487 to 8825a6e Compare September 29, 2026 16:36
"Whether this Dag currently has as many active runs as its max_active_runs allows. "
"Counted differently from active_runs_count above: this counts RUNNING and QUEUED "
"runs (excluding backfill runs), matching the scheduler's own promotion check, while "
"active_runs_count counts RUNNING runs only and includes backfill runs. A Dag with "

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now that active_runs_count excludes backfill runs, this description is wrong: it still says active_runs_count "includes backfill runs", and the one-running-backfill-run example now returns active_runs_count: 0. It also calls RUNNING + QUEUED the scheduler's promotion check, but get_queued_dag_runs_to_set_running gates on RUNNING only; RUNNING + QUEUED is the run-creation gate (exceeds_max_non_backfill in dags_needing_dagruns). Since active_runs_count has now meant three things (RUNNING + QUEUED in 3.3.0, RUNNING with backfill in 3.3.2, RUNNING without backfill here), could it and queued_runs_count get a Field(description=...) of their own, with the spec, UI client and airflowctl models regenerated?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. is_at_max_active_runs now has the short description from #73692, and active_runs_count and queued_runs_count each have their own (running / queued, backfill runs excluded). The spec, UI client and airflowctl are regenerated.


Drafted-by: Claude Code (Opus 5.5); reviewed by @seanmuth before posting

non_backfill_active_runs_count = DagRun.active_runs_of_dags(
dag_ids=[dag_id], exclude_backfill=True, session=session
).get(dag_id, 0)
is_at_max_active_runs = non_backfill_active_runs_count >= (dag_model.max_active_runs or 0)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With both counts now excluding backfill, active_runs_of_dags(exclude_backfill=True) is just active_runs_count + queued_runs_count, so this third SELECT can go: one select(DagRun.state, func.count())...group_by(DagRun.state) gives both counts, and is_at_max_active_runs is their sum against the limit. That also stops the three values disagreeing when a run is triggered between the separate queries, and leaves one backfill filter instead of two (run_type != BACKFILL_JOB here vs backfill_id IS NULL above).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. One query grouped by state gives both counts, and is_at_max_active_runs is their sum against the limit. There's a single backfill filter, run_type != BACKFILL_JOB, matching active_runs_of_dags.


Drafted-by: Claude Code (Opus 5.5); reviewed by @seanmuth before posting

maxActiveRuns: dag.max_active_runs,
queuedRuns: dag.queued_runs_count,
})
: `${dag.active_runs_count ?? 0} of ${dag.max_active_runs}`,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only the queued branch goes through i18n now; this one still hard-codes of in English, so in a translated locale the stat switches language as the queue drains. An activeRunsOfMax key next to activeRunsWithQueued would keep the two consistent.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added activeRunsOfMax next to activeRunsWithQueued, so both branches go through i18n.


Drafted-by: Claude Code (Opus 5.5); reviewed by @seanmuth before posting

it("shows an info icon and the queued count when runs are queued behind the maximum", () => {
render(
<Wrapper>
<Header dag={{ ...mockDag, active_runs_count: 1, max_active_runs: 1, queued_runs_count: 2 }} />

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Every icon-shown case is exactly at capacity, so changing the >= in isBlockedByMaxActiveRuns to === still passes all of these. Over capacity is reachable (lower max_active_runs while runs are active), so an active_runs_count: 3, max_active_runs: 1, queued_runs_count: 2 case asserting the icon and "3 of 1 (2 queued)" would pin it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added the 3 of 1 (2 queued) over-capacity case.


Drafted-by: Claude Code (Opus 5.5); reviewed by @seanmuth before posting

self, session, test_client
):
"""Backfill runs don't count against the Dag's own max_active_runs, so they're excluded."""
from airflow.models.backfill import Backfill

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This import can go at the top of the module, as test_backfills.py and models/test_backfill.py already do; there's no cycle to avoid here. The Backfill row committed here also outlives the test: _clear_db doesn't call clear_db_backfills(), and Backfill.dag_id has no FK, so clear_db_dags won't sweep it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dropped the Backfill rows and the import instead. With the filter on run_type, a DagRun with run_type=BACKFILL_JOB covers it and nothing outlives the test.


Drafted-by: Claude Code (Opus 5.5); reviewed by @seanmuth before posting

seanmuth and others added 13 commits September 30, 2026 10:23
The previous description wrongly called RUNNING + QUEUED the
scheduler's promotion check (that gate counts RUNNING only). State
what the field is instead: whether the Dag is currently at its limit,
counting running and queued non-backfill runs.

Test docstrings still described the old design, where the API aliased
exceeds_max_non_backfill and so depended on the dag-processor keeping
it accurate for schedule=None Dags. The backfill test no longer needs
a Backfill row: active_runs_of_dags filters on run_type alone.

Co-Authored-By: Claude <noreply@anthropic.com>
The Dag header's "Active Runs" stat shows "X of Y" once a Dag is at
or over its max_active_runs limit, but nothing explains what that
means or what happens next. A newly created run beyond the limit
simply sits queued with no explanation visible in the UI, and "3 of
1" reads as a plain oddity rather than a Dag waiting on capacity.

This adds a warning-triangle icon with a tooltip next to the stat
label, shown only when the Dag has exceeded max_active_runs,
explaining that additional runs will not start until an existing
active run completes.

Builds on apache#73692, which adds the exceeds_max_active_runs field this
consumes on DAGDetailsResponse.

Part of apache#73686.
Two refinements based on testing this against a live reproduction:

- The number displayed can never actually show more active runs than
  the limit allows (RUNNING is capped by the scheduler's promotion
  gate), so a warning-severity icon overstated the situation. Switch
  to a plain info icon.
- Key the tooltip off a new, live-computed queued_runs_count field
  instead of exceeds_max_active_runs. The flag answers "is this Dag
  at or over capacity" (useful on its own, via apache#73692), which is a
  slightly different question from "are there runs actually waiting
  right now" -- the latter is what the UI needs, and computing it
  fresh on every request sidesteps any staleness in the persisted
  flag entirely. Also display the queued count directly ("1 of 1 (2
  queued)"), so the information doesn't require a hover at all.
Drop the opening sentence -- now that the tooltip is keyed off
queued_runs_count rather than an "active vs max" comparison, "more
active runs than its limit allows" no longer accurately describes
the condition, and the second sentence already says what matters.
The rebase conflict resolution took a placeholder version of these
generated files; regenerate them fresh so they reflect the
is_at_max_active_runs rename and queued_runs_count addition.

Co-Authored-By: Claude <noreply@anthropic.com>
Same placeholder-conflict-resolution issue as the previous commit --
this file still had exceeds_max_active_runs from before the rename.

Co-Authored-By: Claude <noreply@anthropic.com>
This test suite's i18n setup doesn't resolve real translations --
every translate() call renders its raw key, as every other assertion
in this file already accounts for by comparing against i18n.t(...)
rather than a hardcoded final string. The new queued-count assertion
missed that pattern and compared against the literal English output,
which never matches in CI.

Co-Authored-By: Claude <noreply@anthropic.com>
Both counts were counting every run regardless of backfill_id, but a
backfill's own runs are gated by Backfill.max_active_runs, not the
Dag's. A 100-date backfill on a max_active_runs=16 Dag would render
"10 of 16 (90 queued)" pointing at the wrong limit. Filters both
queries on DagRun.backfill_id.is_(None) to match how the promotion
queries (get_queued_dag_runs_to_set_running, _start_queued_dagruns)
already scope concurrency per (dag_id, backfill_id).

Also strengthened the existing count test to seed two queued runs
instead of one, so it can't pass by coincidence if queued_runs_count
accidentally queried RUNNING instead of QUEUED.

Co-Authored-By: Claude <noreply@anthropic.com>
The info icon showed whenever queued_runs_count > 0, but the tooltip
claims a run is "waiting for an active run to complete" -- not true
for a paused Dag (queued runs never promote at all while paused) or
for the brief window where the scheduler just hasn't picked up a
queued run yet on a busy deployment. Now only shows the icon when
active_runs_count is actually at or over max_active_runs and the Dag
isn't paused; the "(N queued)" text still always shows whenever
queued_runs_count > 0, regardless of the icon.

Added `portalled` to the Tooltip: without it, the tooltip's content
rendered inline inside HeaderCard's stat-label Box, which sets
textTransform="uppercase" -- so the tooltip text appeared in caps,
overlapping the stat's value.

Reverted HeaderCard's key={stat.key ?? index} back to
key={stat.key ?? stat.label}: the array-index fallback slipped past
react/no-array-index-key (which doesn't inspect ?? expressions) and
silently switched every other HeaderCard caller from a label-keyed to
an index-keyed list. The active-runs stat now passes key: "activeRuns"
explicitly instead, and the stats prop is typed as a discriminated
union so a non-string label always requires an explicit key going
forward. The key expression itself narrows label with typeof, since
label's type still includes non-Key ReactNode values like false.

Rewrote Header.test.tsx's i18n-dependent assertions to load the real
en/common locale bundle (matching RenderedJsonField.test.tsx) instead
of comparing against i18n.t(...) with no bundle loaded, which just
compared the raw translation key to itself and would pass regardless
of the actual interpolated values. Also fixed two pre-existing
assertions that were querying the "dag" namespace for a key
(dagDetails.nextRun) that actually lives in "common" -- previously
invisible because both sides always fell back to the same raw key.

Added coverage for the corrected icon condition: exactly at capacity
with nothing queued, below capacity with something queued, and a
paused Dag at capacity with something queued -- all cases where the
icon must stay hidden even though it previously would have shown (or,
for the first case, was already covered).

Co-Authored-By: Claude <noreply@anthropic.com>
The rebase conflict resolution took a placeholder version of these
generated files; regenerate them fresh so they reflect the
is_at_max_active_runs description added on apache#73692.

Co-Authored-By: Claude <noreply@anthropic.com>
Same placeholder-conflict-resolution issue as the previous commit --
this file was still missing is_at_max_active_runs' description field.

Co-Authored-By: Claude <noreply@anthropic.com>
The test inherited from apache#73692's branch expected active_runs_count to
still include the backfill run (true there, since that branch doesn't
have this branch's own backfill-exclusion fix for active_runs_count).
On this branch both fields exclude it, so rewrote the test to add a
manual running run alongside the backfill one and assert the backfill
run doesn't push either count from 1 to 2 -- falsifiable regardless of
which of the two fields' exclusion logic might regress.

Co-Authored-By: Claude <noreply@anthropic.com>
active_runs_count, queued_runs_count and is_at_max_active_runs came
from three separate SELECTs with two different backfill filters, so
they could disagree if a run changed state between them. One query
grouped by state now gives both counts, and is_at_max_active_runs is
their sum against the limit.

active_runs_count has meant something different in each of the last
few releases, so both counts now carry a description of what they
count.

The non-queued "X of Y" stat now goes through i18n like the queued
one, so a translated UI doesn't switch language as the queue drains.
Added an over-capacity test so the icon condition's >= is pinned, and
dropped the Backfill row from the backfill test (the filter is on
run_type, so the row isn't needed and it outlived the test).

Co-Authored-By: Claude <noreply@anthropic.com>
@seanmuth
seanmuth force-pushed the seanmuth/warn-icon-max-active-runs-header branch from 8825a6e to f49dbc9 Compare September 30, 2026 14:35
@bbovenzi bbovenzi added this to the Airflow 3.4.0 milestone Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants