Skip to content

Graph-based Language Models - #222

Open
hunhoffe wants to merge 707 commits into
develfrom
iron-next
Open

hunhoffe wants to merge 707 commits into
develfrom
iron-next

Conversation

@hunhoffe

@hunhoffe hunhoffe commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Note, this PR relies on: Xilinx/mlir-aie#3884

PR Merge Checklist

  1. The PR is rebased on the latest devel commit and pointing to devel.
  2. Your PR has been reviewed and approved.
  3. All checks are passing.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

CI performance trends

fe5111b (2026-10-10T00:43:50Z)

Phoenix - Applications

Operators added: llama_3.2_1b

⚠️ No results from: Krackan - Operators (cancelled), Krackan - Applications (failure), Phoenix - Operators (failure). This comment covers the remaining suites only.

@hunhoffe hunhoffe left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Early notes

Comment thread .github/workflows/ci-lint.yml Outdated
Comment thread .github/workflows/krackan-examples.yml Outdated
Comment thread ci/scripts/pretty_common.py
Comment thread iron/common/declare/bound.py Outdated
Comment thread iron/common/declare/bound.py Outdated
Comment thread iron/common/design/build.py
Comment thread iron/common/design/external.py Outdated
Comment thread iron/common/design/runtime.py Outdated
Comment thread iron/common/design/runtime.py Outdated
Comment thread iron/common/design/target.py
@hunhoffe
hunhoffe marked this pull request as ready for review October 6, 2026 18:55
@hunhoffe hunhoffe changed the title [Proof of Concept] Graph-based Llama 3.2 1B Graph-based Llama 3.2 1B Oct 6, 2026
@hunhoffe hunhoffe changed the title Graph-based Llama 3.2 1B Graph-based Language Models Oct 6, 2026
hunhoffe and others added 21 commits October 6, 2026 18:35
A traced call replaces a view operand with the buffer it views
(_take_views), and a flat-declared output then took the shape of the
operand it was the size of. For a Copy of a view that operand is the
parent. A same-size parent gave the parent's shape: Copy(x.reshape(N, G,
D).transpose(1, 0, 2)) traced as (N, G, D), not (G, N, D). A larger
parent, as with Copy(cache[:, 0:N]), left the output flat. The graph
reference did the same with the whole host tensor.

The tracer now keeps the operands as the call gave them and takes the
output's shape from the view, and its bounds from the buffer behind it.
The reference matches on the view's host pattern. A (16, 512) input and a
(8, 128, 64) cache traced [(16, 8, 64), (8192,)] and referenced
[(8, 16, 64), (8192,)]; both now give (8, 16, 64) twice.

A device-free test in copy_taps checks both shapes in the trace and the
reference, and the transposed values. iron/tests/common: 219 passed; the
one failure, fused_identity's cross-process key, fails the same on HEAD.

Co-Authored-By: Claude <noreply@anthropic.com>
Now that the operators suite runs iron/tests/, these tests reach every
runner. Each one dispatches on, or compiles for, the npu2 the test binds,
so supported_devices("npu2") lets the root conftest skip them, with its
reason, on a host without an NPU or with another device. The tests are
graph_versions (4), graph_dispatch's numeric and packed tests,
sequence_output_sync, and jit_compile_path's bound-device key test.

The four modules: 34 passed on npu2 (Strix Halo).

Co-Authored-By: Claude <noreply@anthropic.com>
A probe writes every per-call value its steps read. One not given is now
derived, as a call derives it, from the bounded extents given (GQAScores'
`calls` from `valid`). The values are read from the resolved operator,
since one derived through a tunable is only seen once the tunable is set:
Softmax's `blocks` follows `streamed`, so the unresolved operator listed
two of its three values and the probe's dispatch missed the third.

Verified on Phoenix (npu1): the Llama decode step's Softmax and GQAScores
widths now measure, every width exact.

Co-Authored-By: Claude <noreply@anthropic.com>
… tuning

GEMM's A shims and L2 tiles, MHA's shim split and lanes, flm GEMM's grid
and B storage, and Dequant's packed tile are auto() fields resolve()
computes from the other fields whatever it is given. Declaring them
auto(derived=True) keeps them out of Operator.widths, with_tunables and
profile entries, which take the settable tunables (_tunable_fields) only.

Before, a derived field in a per= (GEMM's n_shim_mem_a, MHA's q_shims and
kv_lanes) became a width: variants() built one variant per value, resolve()
overwrote it, and the probe measured the same design once per value.

Co-Authored-By: Claude <noreply@anthropic.com>
…raph

A cost table is one graph's working set. The probe now also writes each
design it measures to NPU_CACHE_HOME/iron/costs/<platform>/<power mode>/,
one JSON entry per design, and measure_graph reads from there before
running anything. A design shared by two graphs, or by two versions of
one, is measured once. Configure calibrations are cached the same way,
keyed by their pair.

An entry is keyed on what its time depends on: mlir-aie's recipe and
artifact hashes for the design's build (generator, Python, kernel
sources, tools, device), the per-call values it ran at, and its given
inputs. Editing a design's generator misses the cache rather than reusing
a stale time. The directory is keyed by xrt-smi's platform name and power
mode, because ~/.npu can be NFS-shared across hosts with different NPUs.

The cache holds each output's sha256 rather than a verdict, so exactness
against whichever width a table takes as its default needs no rerun.
measure_graph now returns what it ran on the device.

On Strix Halo, measuring a graph whose add and silu designs were cached
by another graph ran nothing.

Co-Authored-By: Claude <noreply@anthropic.com>
GEMM resolves n_shim_mem_a to at most its column count, so of the
sixteen settings variants() tried for GEMM(2560, 768, 3072), twelve
resolved to a design already in the list: four designs, each measured
four times over and loaded into several probe contexts at once. A
design the measurement already holds is now skipped, so that GEMM
measures its four widths once (it took 246 s in the vision tune with
the repeats).

Co-Authored-By: Claude <noreply@anthropic.com>
token_states and __call__ take soft, each placeholder id's soft tokens,
which replace the placeholders' embeddings in order and unscaled, as
Hugging Face's masked_scatter does. A count that is not the soft tokens'
is refused.

Co-Authored-By: Claude <noreply@anthropic.com>
GEMM picks tile_m in resolve() when the call leaves it open, so
validate() saw no tile_m at construction and an M that no 64, 32, 16
or 8 row tile splits over the device's rows (SwiGLU's 300 rows) was
only refused at build time. The choice is now one method,
_auto_tile_m(dev), that resolve() and validate() share, so the trace
names the rule again (swiglu_no_padding.py).

Co-Authored-By: Claude <noreply@anthropic.com>
GEMM(A[:, :n], B[:n]) binds a new Extent, valid_k, and each core
reduces over the ceil(n / tile_k) K tiles the bound covers, acquiring
and releasing the rest without multiplying them, as the bound on M
already does for its output tiles. The DMAs still move all of A and
B. The last covered tile is multiplied whole, so A's columns past n
in it must be zero, which a bounded Softmax's output is, and B's rows
there finite.

The count is a scratchpad word, so the bound needs a full ELF; an
xclbin's cores would reduce over every tile, and compile refuses it.

gemm_bounded_k.py checks it on npu2, in accurate and default modes,
at ten lengths across tile edges, with A and B NaN past the last
covered tile. It also checks Softmax(s[:, :n]) into GEMM run long
before short, where stale blocks from a longer call would be summed.
Making the cores multiply every tile fails both checks: NaN in every
element, and a 22x error at n=300.

Co-Authored-By: Claude <noreply@anthropic.com>
Standalone wrote a word for every per-call value of its operators and
needed each one named. A bounded operator has two kinds that the graph
already handles: an extent read only through its derivations has no
word in the image (Softmax's length), and a derived word follows from
the extents (its valid_cols). The probe now writes only the words its
image reads, as CompiledVersion does, and derives the rest with
derived_at, so measure_graph's Call.op_values, which gives only bound
members, measures a bounded Softmax or GEMM.

narrowing.py's Masked graph bounds its Softmax as Softmax(x[:, :n])
and gives the probe the length.

Co-Authored-By: Claude <noreply@anthropic.com>
The encoder's Softmax took the prompt length as vector_size, which
Softmax(x[:, :n]) replaced. Its attention is now
GEMM(Softmax(scores[:, :n]), v[g, :n]), so the weights' consumer
reduces over the valid keys alone, and the mean pooling a Softmax of
zeros bounded to n, 32 rows of it per version, into the same bounded
GEMM over h[:n]. It needs no Transpose of h.

On npu2 with random weights, q and k normed small so attention is not
one-hot, the embedding's cosine to EmbeddingGemmaOracle is at least
0.9997 at 500, 300, 250, 130, 64, 60, 33 and 1 tokens, run long
before short in each version, against at most 0.9985 to the oracle of
one token fewer.

Co-Authored-By: Claude <noreply@anthropic.com>
A table indexed by an int32 input of the graph is now a gather the device
runs per call. The traced Copy becomes a Gather (Copy.per_call_gather):
shim (1, 0) streams control packets to shim (0, 0)'s TileControl, which
writes (0, 0)'s buffer descriptors with each row's address and pushes
them onto MM2S 1 in batches of 8 over two alternating sets; MM2S 1
streams the rows to (1, 0)'s S2MM 0, which drains them into the output.
The host encodes each call's ids into those words (Gather.control_words,
the trace's encoder for that input) from the table's device address,
read once from its buffer and cached on the view. The address is
physical, so a gather per call needs a full ELF and refuses an xclbin.
Ids follow numpy: negative ones count from the end, and ones past the
table raise IndexError on the device path and in the reference alike.

The DDR aperture is one named constant, APERTURE, pointing at
kDDRAIEAddrOffset (TxnEncoding.h). ShimChannel names the far end of a
shim-to-shim route, which aie.iron allocates one end of.

The TileControl flow sets keep_pkt_header=True: aie.iron's PacketFlow
writes an explicit false by default, which sets the drop-header bit of
(0, 0)'s switch master (CDO 0x88 rather than 0x08) and hangs the run
(16 rows: ERT_CMD_STATE_TIMEOUT with false, bit-exact with true or the
attribute absent).

Measured on npu2, the gather graph alone, median of 10 calls, bit-exact
against np.take: 0.26 / 0.30 / 0.38 ms at 15 / 256 / 512 rows.

Co-Authored-By: Claude <noreply@anthropic.com>
Since GEMM picks an open tile_m in validate() too, an M no tile splits
on the bound device is refused when the operator is constructed, and
one that only the target device cannot split when it is resolved for
it. test_an_m_no_tile_splits_is_rejected accepts either point.

Co-Authored-By: Claude <noreply@anthropic.com>
The small and extensive suites ran iron/operators/ and the catalog;
they now run all of iron/tests/, so the library, toolchain and
operator tests beside the catalog are checked on every PR. Some of
those import a model package, so the shared prereqs install
requirements_examples.txt, and the Krackan and Phoenix example
workflows drop the step that installed it themselves.
sequence_output_sync.py names the devices it runs on rather than
binding npu2, so it runs on Phoenix too. AGENTS.md and the README
give the same paths.

Measured on Strix (xcoradaie211): the non-extensive suite is 979
passed, 13 skipped, in 6:40 with a warm cache, and the toolchain tests
take 12:51 cold.

Co-Authored-By: Claude <noreply@anthropic.com>
The vision tower and embed_vision as one IRON graph, a version each at
1280 and 2560 patch rows, with its float32 oracle, a device test, a
cost-table tune and a GEMM profile. Patches arrive in pooling-window
order, so pooling is one GEMM by a 0/1 matrix; each head's channels are
reordered so the axial RoPE is one rotation of halves. The position
embeddings and RoPE angles are gathered on the host (Vision.gathered).

Measured on Strix Halo, default widths, 280-token image: min cosine
0.9705, mean 0.99924, relative error 3.9e-2 against the oracle, which
matches Hugging Face's float32 tower to 2.4e-6 (HF's own bf16 tower is
2.3e-2 off it: the gap is bf16 growth with depth). 1082 ms per image.
Transposing the PV product (N = T spans every column) took it from
1417 ms; the profile's narrow GEMM tiles speed the N = 768 GEMMs by
25-27% and a frame's PV by 31%.

The cost table is partly measured, and the accuracy gates are set from
tuned measurements in a later commit.

Co-Authored-By: Claude <noreply@anthropic.com>
VisionTower holds the tower's weights and its body, as AudioTower does
the audio tower's, so the multimodal graph composes it and names its
weights under vision.*; Vision is that tower as a graph of its own and
keeps its profile, shapes, load and embed. The cost table keys designs,
not paths, so it holds as it was: a device-free trace of both versions
finds the same 20 designs.

Co-Authored-By: Claude <noreply@anthropic.com>
A probe of an operator bounded by a per-call extent (`Softmax(x[:, :n])`)
holds values its sequence reads only through their derivations; those have
no scratchpad word, so writing one failed with "unknown parameter". The
probe now writes the words the built image declares. The narrowing device
test bounds its Softmax by a slice, as Softmax declares it today.

Co-Authored-By: Claude <noreply@anthropic.com>
Designs whose devices are equal once their names and runtime sequences are
set aside (`array_text`, after the fusion's shared words) are one array:
the fusion now gives them one device carrying a runtime sequence per
design, so consecutive steps over it share one configure. `merge_devices`
keeps one copy of an equal array and moves each member's sequence into it;
`Packing.sharing` unions those classes with the co-residence groups. A pair
with one `array_key` whose devices still differ is logged and kept apart.

JointNarrowing takes the designs of one default array as one unit at one
width, and `model_us` loads an array once per device entry, so the model
counts what the fusion builds.

Llama 3.2 1B on npu2 (Strix Halo, random weights, 1024-token prompt),
8 interleaved rounds against 15c3be5:
  decode: 19 -> 16 devices, 246 -> 214 configures,
          14.16 -> 14.71 tok/s median (70.6 -> 68.0 ms a token)
  logits over 8 greedy steps bit-identical
The prompt does not share yet: its four GEMMs and two RoPEs of one array
read per-design scratchpad words, which one core buffer cannot take from
two sequences without an mlir-aie lowering change.

Co-Authored-By: Claude <noreply@anthropic.com>
On NPU1 a graph is an xclbin whose steps are dispatched one at a time, so
the full-ELF cost model's pack, base and reset configure do not exist
there. A design run again back to back skips its reconfigure, so its entry
cost is measured by alternation against two reference designs instead:
[v, r] * 4 against [v] * 4 + [r] * 4 gives E(v) + E(r), and with the same
of (v, q) and (r, q), E(v). Graph.compile plans the packaging first and
asks the tuner for packs only on a fused dispatch.

The npu1 narrowing test measures add/silu/mul this way, checks the tuned
chain bit-identical to the plain one and against the numpy reference.
On Phoenix a switch between two designs measures 117-130 us.

Co-Authored-By: Claude <noreply@anthropic.com>
StepCallable pulled every non-input buffer back to the host after each
call, weights and caches included, so that a later read would see what
the steps wrote. The pull claimed those buffers for the host, and XRT
pushes every kernel argument before it dispatches, so the next call
flushed all of them back to the device as well. The callable now only
marks them device-resident, as FullELFCallable's wait already does: a
read through numpy() or to("cpu") pulls what the steps wrote, and a
dispatch pushes nothing it did not need to.

The three reads that went through numpy_view(), the write path that
skips the pull, now use numpy(): the probe's output comparison and
three output checks in tests/infrastructure/sequence.py.

Llama 3.2 1B decode on Phoenix (each step its own dispatch, 491 buffers,
2.4 GiB), interleaved over 16 rounds: 864.3 ms a token before (a 32.1 ms
pull, then 832.2 ms of call), 716.5 ms after, 17% less. Accuracy against
the float32 oracle and determinism (0/8 differing runs) are unchanged,
and sequence_output_sync's every-dispatch-returns-its-own-output check
passes on the NPU1 device.

Co-Authored-By: Claude <noreply@anthropic.com>
The text path looked its embeddings up on the host and uploaded them as
the graph's input. The scaled table is now a weight of the graph, one
zero row past it, and a call takes the token ids: body() is
Copy(self.embedding[ids]), the per-call gather, fed to encoder(x, n),
which holds the encoder from the per-layer embeddings down so the
mixed-modal graph can reuse it. Config also names the placeholder and
marker ids an image's or a clip's soft tokens take. inputs() pads the
ids with vocab_size, that zero row, so the rows past n stay zero as
before; it refuses tokens outside [-vocab_size, vocab_size) with
IndexError and folds negative ones as numpy does. The table, 256 MB, now
lives on the device beside the other weights. tune.py keys its calls by
the ids' length.

Encode latency on npu2, median of 10, before / after:
   15 tokens: 139.10 / 140.64 ms
  256 tokens: 186.69 / 187.68 ms
  512 tokens: 253.32 / 255.55 ms
The before and after runs were taken at different times on a shared NPU.
The gather alone measures 0.26 / 0.30 / 0.38 ms at those lengths.

Co-Authored-By: Claude <noreply@anthropic.com>
hunhoffe and others added 30 commits October 9, 2026 17:18
`Scratchpad(np.float32)` raises, since the scratchpad encoding zeroes a
value's top two bits. A graph's `*, scale: Scratchpad[np.float32]`
went through `__class_getitem__`, which had no such check, so the float
reached the scratchpad and came back corrupted. The check now lives in
`__class_getitem__`, and the constructor goes through it.

Co-Authored-By: Claude <noreply@anthropic.com>
`compatible()` asked `mac_dims` for aie2p's micro-tiles on every
device, on the ground that NPU2's hold NPU1's. It now passes the bound
device, as GEMM's `compatible()` does, so the check states the tiles
the kernel is built with. For the block sizes MHA accepts today the
verdicts are unchanged (rejected_shapes passes as before).

Co-Authored-By: Claude <noreply@anthropic.com>
The transfer-block comment said 6 rows of tiles; the code takes 4, or 2
with a column-major C. The default-case comment said a graph's
projections run GEMM's defaults; Llama's run flm's GEMM and EG2's pass
the accurate settings, so the comment now names the defaults alone.

Co-Authored-By: Claude <noreply@anthropic.com>
`iron` and `iron.operators` load their names on first access, which
keeps `import iron` cheap; this keeps that and makes the names visible:
- `iron.__dir__` lists the lazy names beside the loaded ones, so
  completion and `dir(iron)` see `Graph`, `state` and the rest.
- `iron.operators` imports every operator under `TYPE_CHECKING`, so a
  checker or an editor resolves `iron.operators.GEMM` to its class.
- `iron.common.__all__` is sorted.
- `Linear.is_integral` had no callers and goes.

Co-Authored-By: Claude <noreply@anthropic.com>
The runner printed each token as it was drawn, decoded on its own. A
character that spans two tokens (much non-ASCII text, and emoji) came
out as two replacement characters. The runner now holds tokens back
while their text ends in U+FFFD, prints them once the next token
completes the character, and prints what is still held when generation
ends.

Co-Authored-By: Claude <noreply@anthropic.com>
`flm.dataflow.grid` restated `dev.core_rows`; `flm.Dequant.packed_size`
and `flm.GEMM.packed_B_size` sized buffers the declared operands now
size themselves; dequant's `BFP16_GROUP_BYTES` served only the first.
`packing.packed_b_size` stays: the build tests check the declared shape
against it.

Co-Authored-By: Claude <noreply@anthropic.com>
…ale note

`pytest_configure` makes the session's one `CSVReporter`, and
`pytest_sessionfinish` writes it. The `csv_reporter` fixture beside them
made a second reporter on the same path; no test requested it, and one
that did would have written an empty CSV there at teardown.

The toolchain conftest's docstring described keeping the first of each
test's `--iterations` repeats. Only tests that take `npu_runtime` are
repeated now, so it says what the module holds.

Co-Authored-By: Claude <noreply@anthropic.com>
flm's GEMM design stated v8bfp16ebs8's block as literals, 8 values in 9
bytes, because mlir-aie exposed no width query when it was written.
`aie.utils.bfp.BLOCK` and `BLOCK_BYTES` state it now, so the flm GEMM
and dequant designs, their operators and their tests read those, and
`BFP16_GROUP`/`BFP16_GROUP_BYTES` go.

Co-Authored-By: Claude <noreply@anthropic.com>
The build test reached `packing.packed_b_size` through the flm GEMM
module's import of it, which went with `packed_B_size` in "Drop four flm
helpers no caller reaches". It imports it from `flm.packing` itself.

Co-Authored-By: Claude <noreply@anthropic.com>
58203c4 refused a field that hides an attribute the library reads, but
looked for the name in every class above the operator. flm's
PrefillSlidingAttention turns its base's own `window = None` into a
`window` field, which is the base's to give, and the module stopped
importing: probe_identity, which imports every flm operator, failed.
declare() now takes the class whose attributes the library reads,
Operator, and checks a new field against its classes alone.

Co-Authored-By: Claude <noreply@anthropic.com>
GEMV's L1 account, added with the prologue fold, read `self.dev` in
compatible(), the current device, not the one `resolved(dev)` was given.
Where none is bound (a device-free run, or an NPU busy elsewhere), the
check raised AttributeError, and rejected_shapes' finish-input test
failed; where another device is bound, GEMV was held to its memory.
The account now runs in resolve(dev), on the resolved tiles, with the
same refusal.

Co-Authored-By: Claude <noreply@anthropic.com>
…h lower=

tests/common/cases.py kept a second list of operator shapes for the
device-free lowering gate, beside each operator's own Testing; the two
had to be kept in step by hand, and it chose the shapes that exercise
a layout or dtype decision well. This moves that choice onto the cases
themselves: `Case(lower=True)` (and `Sweep(lower=True)` for a sweep's
first default case) adds a case to the gate, which lowers each
operator's first default case on every device in DEVICES plus every
case it flags. The flagged ones carry over cases.py's coverage: Copy's
float32, slot and two-channel by-head forms, Repeat's int32, weighted
RMSNorm, Softmax at block 1024, GQA's three groups a column, MHA's
chunk over a cache and interleaved decode, GEMV's first batched case,
and GEMM's column-major layouts. GEMM's f32-out case joins its Testing,
so it is also checked on a device.

lowering.py lowers 98 operator cases across npu1 and npu2; the device-free
set with it: 521 passed, 3 skipped.

Co-Authored-By: Claude <noreply@anthropic.com>
A chain (Prepare, Finish) is gated by the interval its steps' tolerances
admit around each reference value, composed step by step. `_interval`
widened that interval past what `aie.utils.verify.compare` passes:

- `_rounded` stepped `nextafter` away from the side it rounded toward,
  so each edge landed one dtype step outside the admitted range;
- the relative radius was the same on both sides of the centre, where
  `compare`'s `rtol * (|a| + |b|)` reaches further away from zero than
  toward it;
- the strict `<` of `nearly_equal` and of an ulps `atol` was read as `<=`;
- `range_frac` widened the bound and exact kinds, which `compare`
  applies it to neither of;
- ulps steps were counted from the float value rather than from its
  bfloat16 rounding, as `bf16_ulp_distance` counts them.

Each kind is now the set of tests `compare` runs, each an interval
around the centre with its own radius below and above and its own
strictness, and an edge is the dtype value nearest it on the inside.
`test_a_chains_interval_is_what_compare_admits` holds every kind to
`compare` itself: both edges pass, and one step past either fails, over
4096 random values with zeros and subnormals. Before this, each kind
admitted thousands of values `compare` rejects.

The chain gates tighten by those steps.

Co-Authored-By: Claude <noreply@anthropic.com>
The emulated AIE2P float multiply splits each operand into three
round-to-nearest bfloat16 limbs, which is what
`aie.iron.kernels.datamovement.limbs_f32_split` computes for the Limbs
kernel's reference. fmul now takes its limbs from there.

Bit-identical to the inline split over 200k products at each of five
scales (1e-30 to 1e30, overflow included) and the edge values.

Co-Authored-By: Claude <noreply@anthropic.com>
`NPUKernel.__call__` is the runtime's load_and_run with the call's
dispatch scalars split from its keywords, which is what
`OperatorImage.__call__` spelled out by hand; the image now calls its
kernel. The scalars are checked as before, by the kernel's dispatch
bridge.

Co-Authored-By: Claude <noreply@anthropic.com>
mlir-aie's ExternalFunction takes `symbol_prefix=` and has no
`digest_prefix=`; the factories prefix their symbols with the digest of
their source and flags inside `_make_extern`, and flm's fused_mm_tile.cc
is a factory now. An operator's own kernel names its configuration in
its prefix, as merge.py does.

Co-Authored-By: Claude <noreply@anthropic.com>
compilable_design_contract.py pins where CompilableDesign's cache key
draws its line. Three of its tests restate the key's own behaviour (a
closure's plain values, compile_kwargs and a stable text each change or
keep the key), which is mlir-aie's contract to hold. The two that stay
are the ones IRON's design rests on: an object closed over alone does
not change the key, which is why IRON's generators carry their identity
in compile_kwargs, and full_elf is part of it, which separates a fused
dispatch's build from a separate one's.

Co-Authored-By: Claude <noreply@anthropic.com>
lowering.py's lower() ran aiecc by hand, with its own flag list and
Peano path. It now calls compile_mlir_module, the path a compiled
design takes, so the gate lowers with the flags production uses
(--unified among them) and follows them as mlir-aie changes them.

The kernels are still built beforehand with compile_external_kernels:
compile_mlir_module(device=) builds only the kernels with a
source_file, so the inline kernel of inline_kernel.py would be left
unlinked. An issue draft for that filter is in issues-to-file.

lowering.py, inline_kernel.py and lowering_graph.py: 130 passed,
1 skipped.

Co-Authored-By: Claude <noreply@anthropic.com>
The graph's tracer and its version lookup named a tensor's dtype
through a table of six names. np.dtype(t.dtype).type gives the same
scalar type for those six, and for the rest the scalar type rather
than a dtype object, which Handle and bfp.dtype_name both take; it
also reads a runtime tensor, whose dtype is the bare class.

Device-free set (iron/tests/common/ and the device-free operator
tests): 415 passed, 2 skipped.

Co-Authored-By: Claude <noreply@anthropic.com>
compile() and __call__ named a version by the same (name, shape,
dtype) tuple, built twice, once from the traced handles and once from
the given tensors. _signature now takes the shapes compile() is called
with, so __call__ builds them once, looks the version up by them, and
passes them on when it compiles. load() reaches the runtime through the
callable property, as a first call does.

Device-free set (iron/tests/common/, the device-free operator tests and
iron/tests/toolchain/narrowing.py): 501 passed, 2 skipped.

Co-Authored-By: Claude <noreply@anthropic.com>
CostCache.put wrote its record to a per-process file and renamed it
into place, as narrowing's _write does for a fit verdict. put now calls
_write, whose parameter is named for any text.

Device-free set (iron/tests/common/, the device-free operator tests and
iron/tests/toolchain/narrowing.py): 501 passed, 2 skipped.

Co-Authored-By: Claude <noreply@anthropic.com>
The dequant layout test and the dequant-feeds-GEMM test each rounded
float32 to bfloat16 toward negative infinity by hand, as the cores do;
aie2p_math_emulation's rb is that rounding, so both call it.

Device-free set (including iron/tests/operators/flm_dequant_layout.py):
501 passed, 2 skipped. flm/dequant/test.py runs on npu2 with the flm queue.

Co-Authored-By: Claude <noreply@anthropic.com>
mlir-aie is a requirement of IRON, so the toolchain conftest imports
it at the top rather than through importorskip.

Device-free set (including iron/tests/toolchain/narrowing.py, which
the conftest serves): 501 passed, 2 skipped.

Co-Authored-By: Claude <noreply@anthropic.com>
The root conftest's npu1 and npu2 fixtures, the toolchain conftest's
device fixture and npu2 override, and two npu1 builds in the xclbin
gate each bound a device and restored the previous one by hand. The
root conftest now makes its npu1 and npu2 from one factory; the
toolchain's device fixture is the root fixture its parameter names,
and the npu1 builds take the npu1 fixture. tools.DEVICES stays for
collection, where lowering.py resolves each catalog case's shapes, and
for compile(dev=), which binds a device of its own.

inline_kernel.py, xclbin.py's npu1 GEMV build, and the device-free tests
that take npu1 or npu2 (common/build.py, elementwise.py, graph.py,
mha_buffer_shapes.py, gemm_tile_divisibility.py): 138 passed.
inline_kernel.py, xclbin.py and lowering.py collect 126 tests over both
devices.

Co-Authored-By: Claude <noreply@anthropic.com>
The flm GEMM keyword-construction test passed `9 / 8`, the bytes a
bfp16ebs8 element takes, where the operator computes it as
`bfp.BLOCK_BYTES / bfp.BLOCK`. It takes the same expression.

Co-Authored-By: Claude <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant