Conversation
DagModel.exceeds_max_non_backfill has been persisted since 3.2.0
(migration 0099) and is already used internally by the scheduler to
avoid re-evaluating Dags that are already at their concurrency limit,
but it was never surfaced anywhere in the public API. There is
currently no way for a client to tell "is this Dag currently blocked
on max_active_runs" without independently computing active-run counts
against the limit.
This exposes it on GET /dags/{dag_id}/details as
exceeds_max_active_runs, aliased to the underlying model column via
the existing DAG_ALIAS_MAPPING mechanism.
Part of apache#73686.
The dag-processor's active-run-count calculation short-circuited to 0 for any Dag whose timetable can't be scheduled (e.g. schedule=None, triggered manually or via the API), regardless of how many runs were actually active. Since exceeds_max_non_backfill is recomputed on every parse cycle, this made it permanently unreliable for exactly the kind of Dag that most needs it: one triggered manually more often than its max_active_runs allows never has a schedule to be "scheduled" against, but still has a real concurrency limit. Only the latest-run lookup (used for scheduling the next run) is skippable for such Dags; the active-run count is not.
The Dag header's "Active Runs" stat shows "X of Y" once a Dag is at or over its max_active_runs limit, but nothing explains what that means or what happens next. A newly created run beyond the limit simply sits queued with no explanation visible in the UI, and "3 of 1" reads as a plain oddity rather than a Dag waiting on capacity. This adds a warning-triangle icon with a tooltip next to the stat label, shown only when the Dag has exceeded max_active_runs, explaining that additional runs will not start until an existing active run completes. Builds on apache#73692, which adds the exceeds_max_active_runs field this consumes on DAGDetailsResponse. Part of apache#73686.
Two refinements based on testing this against a live reproduction: - The number displayed can never actually show more active runs than the limit allows (RUNNING is capped by the scheduler's promotion gate), so a warning-severity icon overstated the situation. Switch to a plain info icon. - Key the tooltip off a new, live-computed queued_runs_count field instead of exceeds_max_active_runs. The flag answers "is this Dag at or over capacity" (useful on its own, via apache#73692), which is a slightly different question from "are there runs actually waiting right now" -- the latter is what the UI needs, and computing it fresh on every request sidesteps any staleness in the persisted flag entirely. Also display the queued count directly ("1 of 1 (2 queued)"), so the information doesn't require a hover at all.
| - timezone | ||
| - last_parsed | ||
| - default_args | ||
| - exceeds_max_active_runs |
There was a problem hiding this comment.
Should this be required, or only sent when it's true? (Not present being implicitly false)
There was a problem hiding this comment.
I'd argue the column should persist (it's NOT NULL in the db) and the expected use case pattern is for checking whether a given Dag is at run capacity before triggering another run — requiring client-side logic to parse an if exists(exceeded_max_active_runs) feels annoying compared to the client-side if exceeded_max_active_runs then ... (agreed on the API field naming change, more in the response comment below).
Drafted-by: Claude Sonnet 5; reviewed by @seanmuth before posting
| :param dags: dict of dags to query | ||
| """ | ||
| # Skip these queries entirely if no Dags can be scheduled to save time. | ||
| active_run_counts = DagRun.active_runs_of_dags( |
There was a problem hiding this comment.
At the very least, we should only do this query when the dag defines max_active_runs.
Further though, I'm worried about this beocming an n+1 query and tanking performance.
There was a problem hiding this comment.
If a Dag doesn't define max_active_runs it will still inherit the deployment's config for [core] max_active_runs_per_dag, so that value will never be NULL.
On the N+1: you're right, and it turns out this was already an N+1 for schedulable Dags before this PR — update_dags() calls _RunInfo.calculate() once per Dag inside a loop over every Dag in the current parse-result update, and that function was calling DagRun.active_runs_of_dags() with a single-element dag_ids list each time, even though that function already accepts a list and batches internally. My change to the non-schedulable path was just extending that existing per-Dag-query pattern to more Dags, not introducing a new shape of the problem. Fixed by hoisting the call out of the loop into one batched call across every Dag in the update, and passing each Dag's precomputed count into _RunInfo.calculate() — pushed in the latest commit.
Drafted-by: Claude Sonnet 5; reviewed by @seanmuth before posting
| }, | ||
| exceeds_max_active_runs: { | ||
| type: 'boolean', | ||
| title: 'Exceeds Max Active Runs' |
There was a problem hiding this comment.
Isn't "exceeds" wrong? Since this will be true when it's at or above max active runs? and "at" isn't exceeding, but it still won't schedule new dag runs.
There was a problem hiding this comment.
On naming: agreed "exceeds" is imprecise — it'll be True when a Dag is sitting exactly at capacity too, not just over it, and that's actually the source-level naming, not something introduced by exposing it. Renamed the exposed field to is_at_max_active_runs, using the existing DAG_ALIAS_MAPPING mechanism so the API name doesn't have to match the internal one. Didn't rename exceeds_max_non_backfill itself here — that's a persisted column, and while a rename migration would be pretty trivial, it felt like a separate, deliberate change rather than something to fold into this PR.
Also — while chasing the "at vs exceeds" semantics, found something worth flagging separately: max_active_runs <= 0 behaves inconsistently across the codebase (the SQL promotion gate treats both 0 and -1 as "never satisfiable" → Dag Runs queue up forever with no error, while exceeds_max_non_backfill's own comparison evaluates permanently True for both, even with zero active runs). Not something this PR touches — filed as a finding on #73686 for visibility.
Drafted-by: Claude Sonnet 5; reviewed by @seanmuth before posting
Review feedback from apache#73692: - "exceeds" was inaccurate: the flag is true once a Dag is at or above its limit, not only when strictly over it. Renamed the exposed field via the existing DAG_ALIAS_MAPPING mechanism; the underlying exceeds_max_non_backfill column is unchanged, since renaming a persisted column is a separate, more deliberate change. - The dag-processor's active-run-count calculation was already an N+1 (one query per Dag per parse cycle) before this PR's fix to the non-schedulable-Dag case; that fix just extended the same pattern to more Dags. DagRun.active_runs_of_dags already accepts a batch of dag_ids, so hoist the call out of the per-Dag loop in update_dags into a single call across every Dag in the update.
The Dag header's "Active Runs" stat shows "X of Y" once a Dag is at or over its max_active_runs limit, but nothing explains what that means or what happens next. A newly created run beyond the limit simply sits queued with no explanation visible in the UI, and "3 of 1" reads as a plain oddity rather than a Dag waiting on capacity. This adds a warning-triangle icon with a tooltip next to the stat label, shown only when the Dag has exceeded max_active_runs, explaining that additional runs will not start until an existing active run completes. Builds on apache#73692, which adds the exceeds_max_active_runs field this consumes on DAGDetailsResponse. Part of apache#73686.
Two refinements based on testing this against a live reproduction: - The number displayed can never actually show more active runs than the limit allows (RUNNING is capped by the scheduler's promotion gate), so a warning-severity icon overstated the situation. Switch to a plain info icon. - Key the tooltip off a new, live-computed queued_runs_count field instead of exceeds_max_active_runs. The flag answers "is this Dag at or over capacity" (useful on its own, via apache#73692), which is a slightly different question from "are there runs actually waiting right now" -- the latter is what the UI needs, and computing it fresh on every request sidesteps any staleness in the persisted flag entirely. Also display the queued count directly ("1 of 1 (2 queued)"), so the information doesn't require a hover at all.
… Dags test_bulk_write_to_db_interval_save_runtime encoded the old, buggy short-circuit this PR removes: active_runs_of_dags is now always batched once per update, regardless of whether any Dag in the batch can be scheduled, so exceeds_max_non_backfill stays accurate for schedule=None Dags too. Co-Authored-By: Claude <noreply@anthropic.com>
The Dag header's "Active Runs" stat shows "X of Y" once a Dag is at or over its max_active_runs limit, but nothing explains what that means or what happens next. A newly created run beyond the limit simply sits queued with no explanation visible in the UI, and "3 of 1" reads as a plain oddity rather than a Dag waiting on capacity. This adds a warning-triangle icon with a tooltip next to the stat label, shown only when the Dag has exceeded max_active_runs, explaining that additional runs will not start until an existing active run completes. Builds on apache#73692, which adds the exceeds_max_active_runs field this consumes on DAGDetailsResponse. Part of apache#73686.
Two refinements based on testing this against a live reproduction: - The number displayed can never actually show more active runs than the limit allows (RUNNING is capped by the scheduler's promotion gate), so a warning-severity icon overstated the situation. Switch to a plain info icon. - Key the tooltip off a new, live-computed queued_runs_count field instead of exceeds_max_active_runs. The flag answers "is this Dag at or over capacity" (useful on its own, via apache#73692), which is a slightly different question from "are there runs actually waiting right now" -- the latter is what the UI needs, and computing it fresh on every request sidesteps any staleness in the persisted flag entirely. Also display the queued count directly ("1 of 1 (2 queued)"), so the information doesn't require a hover at all.
| "dag_run_timeout": "dagrun_timeout", | ||
| "last_parsed": "last_loaded", | ||
| "template_search_path": "template_searchpath", | ||
| "is_at_max_active_runs": "exceeds_max_non_backfill", |
There was a problem hiding this comment.
This aliases a scheduler cache, and nothing refreshes that cache when a run is triggered. exceeds_max_non_backfill is only written by the dag-processor on parse (collection.py:704) and by _set_exceeds_max_active_runs, which the scheduler calls on scheduled-run creation, dagrun timeout and run finish. trigger_dag_run never touches it. So for a schedule=None Dag with max_active_runs=1, triggering a run leaves this false until the next parse (up to min_file_process_interval, 30s by default), and for the whole run if it finishes before that parse. Marking a run failed through the API has the reverse problem and leaves it true. That's the "check capacity before triggering another run" case from the thread above.
get_dag_details already runs a count query for active_runs_count, so computing this at request time from DagRun.active_runs_of_dags(dag_ids=[dag_id], exclude_backfill=True, session=session) against dag_model.max_active_runs would keep it accurate without depending on the parse cycle. A test that triggers a run after parsing and reads details without reparsing would cover it; the current one sets the column directly.
There was a problem hiding this comment.
Confirmed — you're right, trigger_dag_run never touches exceeds_max_non_backfill, so for a schedule=None Dag it stayed stale exactly as you described. Switched is_at_max_active_runs to compute fresh at request time via DagRun.active_runs_of_dags(exclude_backfill=True) against max_active_runs, the same way active_runs_count/queued_runs_count already do — no more dependency on the parse cycle or scheduler events. Removed the now-unused alias to exceeds_max_non_backfill; that column stays as the scheduler's own internal optimization, just no longer exposed here. Added a test that triggers a run and reads details without reparsing — pushed.
Drafted-by: Claude Sonnet 5; reviewed by @seanmuth before posting
| owner_links: dict[str, str] | None = None | ||
| is_favorite: bool = False | ||
| active_runs_count: int = 0 | ||
| is_at_max_active_runs: bool |
There was a problem hiding this comment.
Could this get a Field(description=...) like is_backfillable has? It counts differently from active_runs_count right above it. That one is RUNNING only and includes backfill runs, while this is RUNNING + QUEUED with backfill excluded. So a Dag with max_active_runs=1 and one running backfill returns active_runs_count: 1 next to is_at_max_active_runs: false, and a Dag with one queued run returns 0 next to true. Saying what's counted would help the #73693 tooltip explain that.
There was a problem hiding this comment.
Added, using your own exact example (the backfill and queued-run cases) — pushed.
Drafted-by: Claude Sonnet 5; reviewed by @seanmuth before posting
|
|
||
| @classmethod | ||
| def calculate(cls, dag: LazyDeserializedDAG, *, session: Session) -> Self: | ||
| def calculate(cls, dag: LazyDeserializedDAG, *, num_active_runs: int, session: Session) -> Self: |
There was a problem hiding this comment.
calculate now takes num_active_runs only to hand it back in the tuple, and the comment below says it "must always be computed" even though this method no longer computes it. update_dags is the only caller, so reading active_run_counts.get(dag_id, 0) there at line 704 and leaving calculate to the latest-run lookup would be simpler.
There was a problem hiding this comment.
Done — _RunInfo and calculate now only carry latest_run; update_dags reads active_run_counts.get(dag_id, 0) directly instead of round-tripping it through calculate's parameter and return value — pushed.
Drafted-by: Claude Sonnet 5; reviewed by @seanmuth before posting
The N+1 fix moved DagRun.active_runs_of_dags from once per Dag to once per persistence call, so it now counts toward FIXED_PER_CALL instead of UNCHANGED_PER_DAG/REWRITE_PER_DAG. Co-Authored-By: Claude <noreply@anthropic.com>
…tale scheduler cache exceeds_max_non_backfill is a scheduler-side cache written on parse and by a handful of scheduler events (scheduled-run creation, dagrun timeout, run finish) -- but never by a manual/API/operator trigger. For a schedule=None Dag, triggering a run left is_at_max_active_runs false until the next parse cycle (up to min_file_process_interval, 30s by default), or for the run's entire lifetime if it finished before that parse -- exactly the case this field exists to cover. Marking a run failed through the API had the reverse problem. get_dag_details now computes it fresh, the same way it already does for active_runs_count/queued_runs_count, via DagRun.active_runs_of_dags(exclude_backfill=True) compared against max_active_runs. Removed the now-unused alias to exceeds_max_non_backfill from DAGDetailsResponse; that column stays as the scheduler's own internal optimization, just no longer exposed through this field. Documented via Field(description=...) that this counts differently than active_runs_count: RUNNING+QUEUED with backfill runs excluded (matching the scheduler's own promotion check), vs active_runs_count's RUNNING-only that includes backfill runs. Also simplified _RunInfo/_RunInfo.calculate in the dag-processor: num_active_runs was only ever passed through unchanged since the N+1 fix moved its computation to the caller, so update_dags now reads it directly from the batched query result instead of round-tripping it through calculate's parameter and return value. Co-Authored-By: Claude <noreply@anthropic.com>
The Dag header's "Active Runs" stat shows "X of Y" once a Dag is at or over its max_active_runs limit, but nothing explains what that means or what happens next. A newly created run beyond the limit simply sits queued with no explanation visible in the UI, and "3 of 1" reads as a plain oddity rather than a Dag waiting on capacity. This adds a warning-triangle icon with a tooltip next to the stat label, shown only when the Dag has exceeded max_active_runs, explaining that additional runs will not start until an existing active run completes. Builds on apache#73692, which adds the exceeds_max_active_runs field this consumes on DAGDetailsResponse. Part of apache#73686.
Two refinements based on testing this against a live reproduction: - The number displayed can never actually show more active runs than the limit allows (RUNNING is capped by the scheduler's promotion gate), so a warning-severity icon overstated the situation. Switch to a plain info icon. - Key the tooltip off a new, live-computed queued_runs_count field instead of exceeds_max_active_runs. The flag answers "is this Dag at or over capacity" (useful on its own, via apache#73692), which is a slightly different question from "are there runs actually waiting right now" -- the latter is what the UI needs, and computing it fresh on every request sidesteps any staleness in the persisted flag entirely. Also display the queued count directly ("1 of 1 (2 queued)"), so the information doesn't require a hover at all.
The rebase conflict resolution took a placeholder version of these generated files; regenerate them fresh so they reflect the is_at_max_active_runs description added on apache#73692. Co-Authored-By: Claude <noreply@anthropic.com>
The test inherited from apache#73692's branch expected active_runs_count to still include the backfill run (true there, since that branch doesn't have this branch's own backfill-exclusion fix for active_runs_count). On this branch both fields exclude it, so rewrote the test to add a manual running run alongside the backfill one and assert the backfill run doesn't push either count from 1 to 2 -- falsifiable regardless of which of the two fields' exclusion logic might regress. Co-Authored-By: Claude <noreply@anthropic.com>
| description=( | ||
| "Whether this Dag currently has as many active runs as its max_active_runs allows. " | ||
| "Counted differently from active_runs_count above: this counts RUNNING and QUEUED " | ||
| "runs (excluding backfill runs), matching the scheduler's own promotion check, while " |
There was a problem hiding this comment.
"Matching the scheduler's own promotion check" isn't right. The QUEUED to RUNNING gate (get_queued_dag_runs_to_set_running in dagrun.py, and _start_queued_dagruns) counts RUNNING runs only. RUNNING plus QUEUED minus backfill is the check the scheduler makes before creating another scheduled run (_set_exceeds_max_active_runs, read by dags_needing_dagruns), and the second example here shows the gap: one queued run with max_active_runs=1 reads true, yet the scheduler promotes that run on its next loop. Could you describe it as the run-creation check? This text is copied into the YAML, both TS files and airflowctl, so those need regenerating too.
Separately, #73693 doesn't read this field. Its tooltip in Header.tsx works out its own answer from active_runs_count, max_active_runs and queued_runs_count, which is closer to the question #73686 asks (are queued runs held back by the limit). So this adds a required public field that nothing in-tree reads, and it answers a different question. Does it need to ship in this PR, or could it wait until there's a caller that wants the run-creation answer?
There was a problem hiding this comment.
Fixed the wording. It now says what the field is: whether the Dag is currently at its max_active_runs limit, counting running and queued runs (backfill runs excluded). The spec, TS clients and airflowctl are regenerated.
On whether it needs to ship: it's meant for external clients, not in-tree readers. The main case is schedule=None Dags that are only triggered manually or through the API, where the caller wants to know before triggering whether another run would start or have to wait. RUNNING + QUEUED is the right count for that, because queued runs already waiting get promoted first. In your example (max_active_runs=1, one queued, none running), a newly triggered run would wait, so true is the answer that caller needs.
The scheduler already tracks this exact state (exceeds_max_non_backfill, persisted since 3.2.0). This just exposes it, computed fresh because the cached copy goes stale on manual triggers. That costs one COUNT per details request here, and nothing extra once #73693 folds it into its single count query.
The #73693 tooltip answers a different question, whether an existing queued run is held back, which is why it doesn't read this field. Together the three PRs cover each observability layer for #73686: logs (#73689), API (this PR), UI (#73693).
Drafted-by: Claude Code (Opus 5.5); reviewed by @seanmuth before posting
|
|
||
| reference_run: DagRun | None = run_info.latest_run | ||
| dm.exceeds_max_non_backfill = run_info.num_active_runs >= dm.max_active_runs | ||
| dm.exceeds_max_non_backfill = active_run_counts.get(dag_id, 0) >= dm.max_active_runs |
There was a problem hiding this comment.
Now that the route computes is_at_max_active_runs itself, nothing reads exceeds_max_non_backfill for a Dag with can_be_scheduled=False. Its only readers are dags_needing_dagruns and _create_dag_runs, and a schedule=None Dag never gets next_dagrun_create_after set or an asset trigger. The batching is still a good cut in per-Dag queries, but the PR body (the DAG_ALIAS_MAPPING alias, "permanently stuck False", #73693 depending on it) and the docstrings in test_collection.py and test_dag.py (the removed num_active_runs, "the flag is read regardless") still describe the old design. Could you update them to say this is a batching change with no scheduling effect for schedule=None Dags?
There was a problem hiding this comment.
Agreed. I updated the docstrings in test_collection.py and test_dag.py. The PR body now describes the dag-processor side as a batching change with no scheduling effect for schedule=None Dags.
Drafted-by: Claude Code (Opus 5.5); reviewed by @seanmuth before posting
|
|
||
| def test_dag_details_is_at_max_active_runs_excludes_backfill_runs(self, session, test_client): | ||
| """A running backfill run doesn't count toward the Dag's own is_at_max_active_runs.""" | ||
| from airflow.models.backfill import Backfill |
There was a problem hiding this comment.
This import can go at the top of the file. The Backfill row also outlives the test, because _clear_db doesn't call clear_db_backfills(). You don't need the row at all: active_runs_of_dags filters on run_type alone and backfill_id is nullable, so a DagRun with run_type=DagRunType.BACKFILL_JOB covers the same case.
There was a problem hiding this comment.
Dropped the Backfill row and its import. A DagRun with run_type=BACKFILL_JOB covers it.
Drafted-by: Claude Code (Opus 5.5); reviewed by @seanmuth before posting
The previous description wrongly called RUNNING + QUEUED the scheduler's promotion check (that gate counts RUNNING only). State what the field is instead: whether the Dag is currently at its limit, counting running and queued non-backfill runs. Test docstrings still described the old design, where the API aliased exceeds_max_non_backfill and so depended on the dag-processor keeping it accurate for schedule=None Dags. The backfill test no longer needs a Backfill row: active_runs_of_dags filters on run_type alone. Co-Authored-By: Claude <noreply@anthropic.com>
The Dag header's "Active Runs" stat shows "X of Y" once a Dag is at or over its max_active_runs limit, but nothing explains what that means or what happens next. A newly created run beyond the limit simply sits queued with no explanation visible in the UI, and "3 of 1" reads as a plain oddity rather than a Dag waiting on capacity. This adds a warning-triangle icon with a tooltip next to the stat label, shown only when the Dag has exceeded max_active_runs, explaining that additional runs will not start until an existing active run completes. Builds on apache#73692, which adds the exceeds_max_active_runs field this consumes on DAGDetailsResponse. Part of apache#73686.
Two refinements based on testing this against a live reproduction: - The number displayed can never actually show more active runs than the limit allows (RUNNING is capped by the scheduler's promotion gate), so a warning-severity icon overstated the situation. Switch to a plain info icon. - Key the tooltip off a new, live-computed queued_runs_count field instead of exceeds_max_active_runs. The flag answers "is this Dag at or over capacity" (useful on its own, via apache#73692), which is a slightly different question from "are there runs actually waiting right now" -- the latter is what the UI needs, and computing it fresh on every request sidesteps any staleness in the persisted flag entirely. Also display the queued count directly ("1 of 1 (2 queued)"), so the information doesn't require a hover at all.
The rebase conflict resolution took a placeholder version of these generated files; regenerate them fresh so they reflect the is_at_max_active_runs description added on apache#73692. Co-Authored-By: Claude <noreply@anthropic.com>
The test inherited from apache#73692's branch expected active_runs_count to still include the backfill run (true there, since that branch doesn't have this branch's own backfill-exclusion fix for active_runs_count). On this branch both fields exclude it, so rewrote the test to add a manual running run alongside the backfill one and assert the backfill run doesn't push either count from 1 to 2 -- falsifiable regardless of which of the two fields' exclusion logic might regress. Co-Authored-By: Claude <noreply@anthropic.com>
Clients that trigger Dags manually or through the API have no direct way to ask whether a run triggered right now would start or have to wait behind
max_active_runs. This matters most forschedule=NoneDags, which are only ever triggered that way. Today a client has to count running and queued runs itself and compare them to the limit.This adds
is_at_max_active_runstoGET /dags/{dag_id}/details: whether the Dag is currently at itsmax_active_runslimit, counting running and queued runs (backfill runs excluded). The scheduler already tracks this exact state inDagModel.exceeds_max_non_backfill(persisted since 3.2.0), but that cached copy isn't refreshed when a run is triggered manually or through the API, so the endpoint computes it fresh. That costs one COUNT per details request here, and nothing extra once #73693 folds it into its single count query. It's scoped to the details endpoint, next toactive_runs_count.It also batches the dag-processor's active-run count into one query per parse update instead of one per Dag (an existing N+1 for scheduled Dags), and computes it for
schedule=NoneDags too, soexceeds_max_non_backfillstays accurate for them. That part has no scheduling effect: nothing reads the flag for a Dag that can't be scheduled.Part of #73686, which covers each observability layer: logs (#73689), API (this PR), UI (#73693).
Was generative AI tooling used to co-author this PR?
Generated-by: Claude Code (Sonnet 5, Opus 5.5) following the guidelines