Skip to content

JIT: compile kernels with -fno-math-errno by default (#834) - #835

Merged
lmoresi merged 2 commits into
developmentfrom
feature/jit-fno-math-errno
Oct 10, 2026
Merged

lmoresi merged 2 commits into
developmentfrom
feature/jit-fno-math-errno

Conversation

@lmoresi

@lmoresi lmoresi commented Oct 8, 2026

Copy link
Copy Markdown
Member

Closes #834.

What changes

The JIT's default compile flags for the pointwise kernels become -O3 -g0 -fno-math-errno, after the fixed -std=c99. UW3_JIT_CFLAGS still replaces the defaults, so a compiler that rejects a flag can drop it: nvc rejects both -g0 and -fno-math-errno. docs/developer/subsystems/jit-cache.md now documents UW3_JIT_CFLAGS.

Why

gcc, and clang on Linux, default to -fmath-errno. Under it, sqrt, pow and exp may set errno, so the compiler treats each call as a side effect and cannot merge a repeated call with the same argument. Our kernels never read errno. clang targeting macOS already defaults to -fno-math-errno, so this makes Linux kernels behave as macOS ones do.

Measured on Linux (conda-forge gcc 14.3, a loaded 2× Xeon Gold 6240R), as the cost of the law's callbacks per quadrature point in one Jacobian assembly:

fixture default -fno-math-errno
viscoplastic box 8.2 µs 2.8 µs (repeat: 8.3 / 2.9)
visco-elasto-plastic box 2.7 µs 3.1 µs (no effect within the noise)

With and without the flag, these are bit-identical on both fixtures:

  • the assembled residual and Jacobian at a fixed state;
  • the solutions and residual histories;
  • the iteration counts.

The generated C is unchanged. The flags are part of the generated setup.py, which is in the cache key, so every cached module recompiles once.

Tests

test_0025_jit_compile_flags.py pins the compile arguments, with the defaults and with an override. The default case fails on development, where -fno-math-errno is missing. The bundle capture that test_0021 already used is now a conftest fixture, jit_bundles, which both tests share.

The new test is marked tier C, as TESTING-RELIABILITY-SYSTEM.md asks of a new test. Recent new tests (test_0021–test_0023) went straight to tier A, so either that practice or the document is out of date. A tier C test does not gate CI. For a pin like this one, promoting it to A is a one-line change once reviewed.

Review

The adversarial review is posted below as a comment.

Underworld development team with AI support from Claude Code

gcc's default -fmath-errno makes each sqrt, pow and exp call a side effect, so a
repeated call with the same argument is computed again. The kernels never read errno.
With the flag, a viscoplastic Jacobian's callbacks cost about a third as much on gcc 14
(8.2 to 2.8 microseconds per quadrature point), and the assembled residual and
Jacobian, the solutions and the iteration counts are bit-identical (measured on Linux
for #823). Apple clang never sets errno, so this is the macOS behaviour already.

The flag is prepended with -std=c99, so a UW3_JIT_CFLAGS override, which replaces the
optimisation flags, keeps it. The flags are in the generated setup.py, which is part of
the cache key: each cached module recompiles once.

test_0025 pins the compile arguments, with and without an override.

Underworld development team with AI support from Claude Code
- nvc rejects -fno-math-errno (and -g0). Always prepending the flag left a JIT on such
  a compiler with no way to compile; it is now one of the default flags, which
  UW3_JIT_CFLAGS replaces. -std=c99 stays fixed.
- The comment names clang on Linux as well as gcc and states the bit-identity for the
  C SymPy prints.
- jit-cache.md documents UW3_JIT_CFLAGS.
- The bundle capture shared with test_0021 is a conftest fixture, jit_bundles.
- test_0025 starts at tier C, as TESTING-RELIABILITY-SYSTEM.md asks of a new test.

Underworld development team with AI support from Claude Code
Copilot AI balanced review requested due to automatic review settings October 8, 2026 02:58
@lmoresi

lmoresi commented Oct 8, 2026

Copy link
Copy Markdown
Member Author

Adversarial review (before opening), and what we did

Findings

  1. should-fix, fixed. Always prepending -fno-math-errno left a compiler that rejects it (nvc: nvc-Error-Unknown switch: -fno-math-errno, nvc 24.9 and 26.9) with no way to compile, where UW3_JIT_CFLAGS="-O3" used to work around nvc's rejection of -g0. The flag is now one of the defaults that UW3_JIT_CFLAGS replaces.
  2. nit, fixed. The comment named gcc only. clang on Linux also defaults to -fmath-errno, so the comment now names both. The bit-identity claim is stated for the C that SymPy prints, because two clang-on-Linux effects fall outside it:
    • a literal pow(x, 0.5) becomes fabs(sqrt(x)), which differs from glibc's pow by 1 ulp on 2,691 of 3.0M inputs. SymPy prints that exponent as sqrt, so the kernels never contain it.
    • a guarded sqrt is computed speculatively, which raises FE_INVALID under PETSc -fp_trap. Nothing in the repository uses -fp_trap.
  3. nit, fixed. The bundle capture duplicated test_0021's; it is now the conftest fixture jit_bundles.
  4. nit, fixed. New test marked tier C per TESTING-RELIABILITY-SYSTEM.md; the drift in recent practice is noted in the PR body.
  5. nit, fixed. jit-cache.md did not document UW3_JIT_CFLAGS; it now has a row. The same file's MPI section, which still says a cross-rank mismatch raises and that every rank compiles, is stale and outside this change.
  6. nit, fixed. The flag list is built in two plain steps.

Attacks that found nothing

  • Numerical safety, gcc. A real viscoplastic JIT header (17 kernels) and a synthetic kernel with 40+ libm calls were run on 12,000 samples, NaNs included. The output hash was identical with and without the flag on:
    • x86-64 gcc 14.3 and 15.2;
    • gcc 14.3 with -march=x86-64-v3;
    • ARM64 gcc 14.3;
    • icc 2021.10 and icx.
  • Value-changing rewrites (pow to sqrt, libmvec) still need -funsafe-math-optimizations or fast-math, which we don't pass.
  • Nothing reads errno: not the Heaviside helpers, not the Cython wrapper, not the analytic-solution calls.
  • Toolchains that accept the flag: gcc 4.8 to 15, clang 3 to 20, AOCC, icx, icc, ARM64 and POWER gcc, Cray CCE (clang-based).
  • Cache and parallel:
    • the flags are in the hashed source, so modules recompile once and no rank-dependent input is added;
    • the disk cache's manifest check is unaffected.
  • Override: a later -fmath-errno restores the old code generation on gcc, clang and icc.
  • The test: it fails when the flag is stripped and asserts the exact list. It is reachable from scripts/test.sh (332 of 332 files), and monkeypatch restores the environment.

Issues: #834 is closed by this PR. #752 is untouched: the flags add no rank-dependent input.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The regression test compiles unsupported defaults on nvc and is incorrectly excluded from gating as Tier C.

2 open findings
What changed in this PR

Adds -fno-math-errno to default JIT kernel flags to improve Linux performance while preserving overrides.

Changes:

  • Updates default JIT compile flags.
  • Documents UW3_JIT_CFLAGS.
  • Adds shared bundle capture and regression tests.
File Description
src/​underworld3/​utilities/​_jitextension.py Adds the new default flag.
docs/​developer/​subsystems/​jit-cache.md Documents compile-flag overrides.
tests/​conftest.py Adds shared JIT bundle capture.
tests/​test_0021_jit_finds_petsc_headers.py Reuses the shared fixture.
tests/​test_0025_jit_compile_flags.py Tests default and overridden flags.

🧠 Review effort: Balanced


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

Comment on lines +20 to +30
def _compile_args(bundles):
"""The ``extra_compile_args`` of the bundle a small integral generates."""
mesh = uw.meshing.UnstructuredSimplexBox(
minCoords=(0.0, 0.0), maxCoords=(1.0, 1.0), cellSize=0.5)
# PETSc integrates only on a mesh that carries a field
uw.discretisation.MeshVariable("U0025", mesh, 1, degree=1)
uw.maths.Integral(mesh, 1.0 + mesh.X[0]).evaluate()
assert bundles, "no JIT bundle was generated"
match = re.search(r"extra_compile_args=(\[.*?\])", bundles[-1]["setup.py"])
assert match, bundles[-1]["setup.py"]
return ast.literal_eval(match.group(1))

import underworld3 as uw

pytestmark = [pytest.mark.level_1, pytest.mark.tier_c]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants