Repository navigation
Conversation
This was referenced Sep 29, 2026
Merged
Closed
Contributor
CI performance trends
Phoenix - ApplicationsOperators added: |
hunhoffe
commented
Oct 5, 2026
hunhoffe
marked this pull request as ready for review
October 6, 2026 18:55
A traced call replaces a view operand with the buffer it views (_take_views), and a flat-declared output then took the shape of the operand it was the size of. For a Copy of a view that operand is the parent. A same-size parent gave the parent's shape: Copy(x.reshape(N, G, D).transpose(1, 0, 2)) traced as (N, G, D), not (G, N, D). A larger parent, as with Copy(cache[:, 0:N]), left the output flat. The graph reference did the same with the whole host tensor. The tracer now keeps the operands as the call gave them and takes the output's shape from the view, and its bounds from the buffer behind it. The reference matches on the view's host pattern. A (16, 512) input and a (8, 128, 64) cache traced [(16, 8, 64), (8192,)] and referenced [(8, 16, 64), (8192,)]; both now give (8, 16, 64) twice. A device-free test in copy_taps checks both shapes in the trace and the reference, and the transposed values. iron/tests/common: 219 passed; the one failure, fused_identity's cross-process key, fails the same on HEAD. Co-Authored-By: Claude <noreply@anthropic.com>
Now that the operators suite runs iron/tests/, these tests reach every
runner. Each one dispatches on, or compiles for, the npu2 the test binds,
so supported_devices("npu2") lets the root conftest skip them, with its
reason, on a host without an NPU or with another device. The tests are
graph_versions (4), graph_dispatch's numeric and packed tests,
sequence_output_sync, and jit_compile_path's bound-device key test.
The four modules: 34 passed on npu2 (Strix Halo).
Co-Authored-By: Claude <noreply@anthropic.com>
A probe writes every per-call value its steps read. One not given is now derived, as a call derives it, from the bounded extents given (GQAScores' `calls` from `valid`). The values are read from the resolved operator, since one derived through a tunable is only seen once the tunable is set: Softmax's `blocks` follows `streamed`, so the unresolved operator listed two of its three values and the probe's dispatch missed the third. Verified on Phoenix (npu1): the Llama decode step's Softmax and GQAScores widths now measure, every width exact. Co-Authored-By: Claude <noreply@anthropic.com>
… tuning GEMM's A shims and L2 tiles, MHA's shim split and lanes, flm GEMM's grid and B storage, and Dequant's packed tile are auto() fields resolve() computes from the other fields whatever it is given. Declaring them auto(derived=True) keeps them out of Operator.widths, with_tunables and profile entries, which take the settable tunables (_tunable_fields) only. Before, a derived field in a per= (GEMM's n_shim_mem_a, MHA's q_shims and kv_lanes) became a width: variants() built one variant per value, resolve() overwrote it, and the probe measured the same design once per value. Co-Authored-By: Claude <noreply@anthropic.com>
…raph A cost table is one graph's working set. The probe now also writes each design it measures to NPU_CACHE_HOME/iron/costs/<platform>/<power mode>/, one JSON entry per design, and measure_graph reads from there before running anything. A design shared by two graphs, or by two versions of one, is measured once. Configure calibrations are cached the same way, keyed by their pair. An entry is keyed on what its time depends on: mlir-aie's recipe and artifact hashes for the design's build (generator, Python, kernel sources, tools, device), the per-call values it ran at, and its given inputs. Editing a design's generator misses the cache rather than reusing a stale time. The directory is keyed by xrt-smi's platform name and power mode, because ~/.npu can be NFS-shared across hosts with different NPUs. The cache holds each output's sha256 rather than a verdict, so exactness against whichever width a table takes as its default needs no rerun. measure_graph now returns what it ran on the device. On Strix Halo, measuring a graph whose add and silu designs were cached by another graph ran nothing. Co-Authored-By: Claude <noreply@anthropic.com>
GEMM resolves n_shim_mem_a to at most its column count, so of the sixteen settings variants() tried for GEMM(2560, 768, 3072), twelve resolved to a design already in the list: four designs, each measured four times over and loaded into several probe contexts at once. A design the measurement already holds is now skipped, so that GEMM measures its four widths once (it took 246 s in the vision tune with the repeats). Co-Authored-By: Claude <noreply@anthropic.com>
token_states and __call__ take soft, each placeholder id's soft tokens, which replace the placeholders' embeddings in order and unscaled, as Hugging Face's masked_scatter does. A count that is not the soft tokens' is refused. Co-Authored-By: Claude <noreply@anthropic.com>
GEMM picks tile_m in resolve() when the call leaves it open, so validate() saw no tile_m at construction and an M that no 64, 32, 16 or 8 row tile splits over the device's rows (SwiGLU's 300 rows) was only refused at build time. The choice is now one method, _auto_tile_m(dev), that resolve() and validate() share, so the trace names the rule again (swiglu_no_padding.py). Co-Authored-By: Claude <noreply@anthropic.com>
GEMM(A[:, :n], B[:n]) binds a new Extent, valid_k, and each core reduces over the ceil(n / tile_k) K tiles the bound covers, acquiring and releasing the rest without multiplying them, as the bound on M already does for its output tiles. The DMAs still move all of A and B. The last covered tile is multiplied whole, so A's columns past n in it must be zero, which a bounded Softmax's output is, and B's rows there finite. The count is a scratchpad word, so the bound needs a full ELF; an xclbin's cores would reduce over every tile, and compile refuses it. gemm_bounded_k.py checks it on npu2, in accurate and default modes, at ten lengths across tile edges, with A and B NaN past the last covered tile. It also checks Softmax(s[:, :n]) into GEMM run long before short, where stale blocks from a longer call would be summed. Making the cores multiply every tile fails both checks: NaN in every element, and a 22x error at n=300. Co-Authored-By: Claude <noreply@anthropic.com>
Standalone wrote a word for every per-call value of its operators and needed each one named. A bounded operator has two kinds that the graph already handles: an extent read only through its derivations has no word in the image (Softmax's length), and a derived word follows from the extents (its valid_cols). The probe now writes only the words its image reads, as CompiledVersion does, and derives the rest with derived_at, so measure_graph's Call.op_values, which gives only bound members, measures a bounded Softmax or GEMM. narrowing.py's Masked graph bounds its Softmax as Softmax(x[:, :n]) and gives the probe the length. Co-Authored-By: Claude <noreply@anthropic.com>
The encoder's Softmax took the prompt length as vector_size, which Softmax(x[:, :n]) replaced. Its attention is now GEMM(Softmax(scores[:, :n]), v[g, :n]), so the weights' consumer reduces over the valid keys alone, and the mean pooling a Softmax of zeros bounded to n, 32 rows of it per version, into the same bounded GEMM over h[:n]. It needs no Transpose of h. On npu2 with random weights, q and k normed small so attention is not one-hot, the embedding's cosine to EmbeddingGemmaOracle is at least 0.9997 at 500, 300, 250, 130, 64, 60, 33 and 1 tokens, run long before short in each version, against at most 0.9985 to the oracle of one token fewer. Co-Authored-By: Claude <noreply@anthropic.com>
A table indexed by an int32 input of the graph is now a gather the device runs per call. The traced Copy becomes a Gather (Copy.per_call_gather): shim (1, 0) streams control packets to shim (0, 0)'s TileControl, which writes (0, 0)'s buffer descriptors with each row's address and pushes them onto MM2S 1 in batches of 8 over two alternating sets; MM2S 1 streams the rows to (1, 0)'s S2MM 0, which drains them into the output. The host encodes each call's ids into those words (Gather.control_words, the trace's encoder for that input) from the table's device address, read once from its buffer and cached on the view. The address is physical, so a gather per call needs a full ELF and refuses an xclbin. Ids follow numpy: negative ones count from the end, and ones past the table raise IndexError on the device path and in the reference alike. The DDR aperture is one named constant, APERTURE, pointing at kDDRAIEAddrOffset (TxnEncoding.h). ShimChannel names the far end of a shim-to-shim route, which aie.iron allocates one end of. The TileControl flow sets keep_pkt_header=True: aie.iron's PacketFlow writes an explicit false by default, which sets the drop-header bit of (0, 0)'s switch master (CDO 0x88 rather than 0x08) and hangs the run (16 rows: ERT_CMD_STATE_TIMEOUT with false, bit-exact with true or the attribute absent). Measured on npu2, the gather graph alone, median of 10 calls, bit-exact against np.take: 0.26 / 0.30 / 0.38 ms at 15 / 256 / 512 rows. Co-Authored-By: Claude <noreply@anthropic.com>
Since GEMM picks an open tile_m in validate() too, an M no tile splits on the bound device is refused when the operator is constructed, and one that only the target device cannot split when it is resolved for it. test_an_m_no_tile_splits_is_rejected accepts either point. Co-Authored-By: Claude <noreply@anthropic.com>
The small and extensive suites ran iron/operators/ and the catalog; they now run all of iron/tests/, so the library, toolchain and operator tests beside the catalog are checked on every PR. Some of those import a model package, so the shared prereqs install requirements_examples.txt, and the Krackan and Phoenix example workflows drop the step that installed it themselves. sequence_output_sync.py names the devices it runs on rather than binding npu2, so it runs on Phoenix too. AGENTS.md and the README give the same paths. Measured on Strix (xcoradaie211): the non-extensive suite is 979 passed, 13 skipped, in 6:40 with a warm cache, and the toolchain tests take 12:51 cold. Co-Authored-By: Claude <noreply@anthropic.com>
The vision tower and embed_vision as one IRON graph, a version each at 1280 and 2560 patch rows, with its float32 oracle, a device test, a cost-table tune and a GEMM profile. Patches arrive in pooling-window order, so pooling is one GEMM by a 0/1 matrix; each head's channels are reordered so the axial RoPE is one rotation of halves. The position embeddings and RoPE angles are gathered on the host (Vision.gathered). Measured on Strix Halo, default widths, 280-token image: min cosine 0.9705, mean 0.99924, relative error 3.9e-2 against the oracle, which matches Hugging Face's float32 tower to 2.4e-6 (HF's own bf16 tower is 2.3e-2 off it: the gap is bf16 growth with depth). 1082 ms per image. Transposing the PV product (N = T spans every column) took it from 1417 ms; the profile's narrow GEMM tiles speed the N = 768 GEMMs by 25-27% and a frame's PV by 31%. The cost table is partly measured, and the accuracy gates are set from tuned measurements in a later commit. Co-Authored-By: Claude <noreply@anthropic.com>
VisionTower holds the tower's weights and its body, as AudioTower does the audio tower's, so the multimodal graph composes it and names its weights under vision.*; Vision is that tower as a graph of its own and keeps its profile, shapes, load and embed. The cost table keys designs, not paths, so it holds as it was: a device-free trace of both versions finds the same 20 designs. Co-Authored-By: Claude <noreply@anthropic.com>
A probe of an operator bounded by a per-call extent (`Softmax(x[:, :n])`) holds values its sequence reads only through their derivations; those have no scratchpad word, so writing one failed with "unknown parameter". The probe now writes the words the built image declares. The narrowing device test bounds its Softmax by a slice, as Softmax declares it today. Co-Authored-By: Claude <noreply@anthropic.com>
Designs whose devices are equal once their names and runtime sequences are set aside (`array_text`, after the fusion's shared words) are one array: the fusion now gives them one device carrying a runtime sequence per design, so consecutive steps over it share one configure. `merge_devices` keeps one copy of an equal array and moves each member's sequence into it; `Packing.sharing` unions those classes with the co-residence groups. A pair with one `array_key` whose devices still differ is logged and kept apart. JointNarrowing takes the designs of one default array as one unit at one width, and `model_us` loads an array once per device entry, so the model counts what the fusion builds. Llama 3.2 1B on npu2 (Strix Halo, random weights, 1024-token prompt), 8 interleaved rounds against 15c3be5: decode: 19 -> 16 devices, 246 -> 214 configures, 14.16 -> 14.71 tok/s median (70.6 -> 68.0 ms a token) logits over 8 greedy steps bit-identical The prompt does not share yet: its four GEMMs and two RoPEs of one array read per-design scratchpad words, which one core buffer cannot take from two sequences without an mlir-aie lowering change. Co-Authored-By: Claude <noreply@anthropic.com>
On NPU1 a graph is an xclbin whose steps are dispatched one at a time, so the full-ELF cost model's pack, base and reset configure do not exist there. A design run again back to back skips its reconfigure, so its entry cost is measured by alternation against two reference designs instead: [v, r] * 4 against [v] * 4 + [r] * 4 gives E(v) + E(r), and with the same of (v, q) and (r, q), E(v). Graph.compile plans the packaging first and asks the tuner for packs only on a fused dispatch. The npu1 narrowing test measures add/silu/mul this way, checks the tuned chain bit-identical to the plain one and against the numpy reference. On Phoenix a switch between two designs measures 117-130 us. Co-Authored-By: Claude <noreply@anthropic.com>
StepCallable pulled every non-input buffer back to the host after each
call, weights and caches included, so that a later read would see what
the steps wrote. The pull claimed those buffers for the host, and XRT
pushes every kernel argument before it dispatches, so the next call
flushed all of them back to the device as well. The callable now only
marks them device-resident, as FullELFCallable's wait already does: a
read through numpy() or to("cpu") pulls what the steps wrote, and a
dispatch pushes nothing it did not need to.
The three reads that went through numpy_view(), the write path that
skips the pull, now use numpy(): the probe's output comparison and
three output checks in tests/infrastructure/sequence.py.
Llama 3.2 1B decode on Phoenix (each step its own dispatch, 491 buffers,
2.4 GiB), interleaved over 16 rounds: 864.3 ms a token before (a 32.1 ms
pull, then 832.2 ms of call), 716.5 ms after, 17% less. Accuracy against
the float32 oracle and determinism (0/8 differing runs) are unchanged,
and sequence_output_sync's every-dispatch-returns-its-own-output check
passes on the NPU1 device.
Co-Authored-By: Claude <noreply@anthropic.com>
The text path looked its embeddings up on the host and uploaded them as the graph's input. The scaled table is now a weight of the graph, one zero row past it, and a call takes the token ids: body() is Copy(self.embedding[ids]), the per-call gather, fed to encoder(x, n), which holds the encoder from the per-layer embeddings down so the mixed-modal graph can reuse it. Config also names the placeholder and marker ids an image's or a clip's soft tokens take. inputs() pads the ids with vocab_size, that zero row, so the rows past n stay zero as before; it refuses tokens outside [-vocab_size, vocab_size) with IndexError and folds negative ones as numpy does. The table, 256 MB, now lives on the device beside the other weights. tune.py keys its calls by the ids' length. Encode latency on npu2, median of 10, before / after: 15 tokens: 139.10 / 140.64 ms 256 tokens: 186.69 / 187.68 ms 512 tokens: 253.32 / 255.55 ms The before and after runs were taken at different times on a shared NPU. The gather alone measures 0.26 / 0.30 / 0.38 ms at those lengths. Co-Authored-By: Claude <noreply@anthropic.com>
`Scratchpad(np.float32)` raises, since the scratchpad encoding zeroes a value's top two bits. A graph's `*, scale: Scratchpad[np.float32]` went through `__class_getitem__`, which had no such check, so the float reached the scratchpad and came back corrupted. The check now lives in `__class_getitem__`, and the constructor goes through it. Co-Authored-By: Claude <noreply@anthropic.com>
`compatible()` asked `mac_dims` for aie2p's micro-tiles on every device, on the ground that NPU2's hold NPU1's. It now passes the bound device, as GEMM's `compatible()` does, so the check states the tiles the kernel is built with. For the block sizes MHA accepts today the verdicts are unchanged (rejected_shapes passes as before). Co-Authored-By: Claude <noreply@anthropic.com>
The transfer-block comment said 6 rows of tiles; the code takes 4, or 2 with a column-major C. The default-case comment said a graph's projections run GEMM's defaults; Llama's run flm's GEMM and EG2's pass the accurate settings, so the comment now names the defaults alone. Co-Authored-By: Claude <noreply@anthropic.com>
`iron` and `iron.operators` load their names on first access, which keeps `import iron` cheap; this keeps that and makes the names visible: - `iron.__dir__` lists the lazy names beside the loaded ones, so completion and `dir(iron)` see `Graph`, `state` and the rest. - `iron.operators` imports every operator under `TYPE_CHECKING`, so a checker or an editor resolves `iron.operators.GEMM` to its class. - `iron.common.__all__` is sorted. - `Linear.is_integral` had no callers and goes. Co-Authored-By: Claude <noreply@anthropic.com>
The runner printed each token as it was drawn, decoded on its own. A character that spans two tokens (much non-ASCII text, and emoji) came out as two replacement characters. The runner now holds tokens back while their text ends in U+FFFD, prints them once the next token completes the character, and prints what is still held when generation ends. Co-Authored-By: Claude <noreply@anthropic.com>
`flm.dataflow.grid` restated `dev.core_rows`; `flm.Dequant.packed_size` and `flm.GEMM.packed_B_size` sized buffers the declared operands now size themselves; dequant's `BFP16_GROUP_BYTES` served only the first. `packing.packed_b_size` stays: the build tests check the declared shape against it. Co-Authored-By: Claude <noreply@anthropic.com>
…ale note `pytest_configure` makes the session's one `CSVReporter`, and `pytest_sessionfinish` writes it. The `csv_reporter` fixture beside them made a second reporter on the same path; no test requested it, and one that did would have written an empty CSV there at teardown. The toolchain conftest's docstring described keeping the first of each test's `--iterations` repeats. Only tests that take `npu_runtime` are repeated now, so it says what the module holds. Co-Authored-By: Claude <noreply@anthropic.com>
flm's GEMM design stated v8bfp16ebs8's block as literals, 8 values in 9 bytes, because mlir-aie exposed no width query when it was written. `aie.utils.bfp.BLOCK` and `BLOCK_BYTES` state it now, so the flm GEMM and dequant designs, their operators and their tests read those, and `BFP16_GROUP`/`BFP16_GROUP_BYTES` go. Co-Authored-By: Claude <noreply@anthropic.com>
The build test reached `packing.packed_b_size` through the flm GEMM module's import of it, which went with `packed_B_size` in "Drop four flm helpers no caller reaches". It imports it from `flm.packing` itself. Co-Authored-By: Claude <noreply@anthropic.com>
58203c4 refused a field that hides an attribute the library reads, but looked for the name in every class above the operator. flm's PrefillSlidingAttention turns its base's own `window = None` into a `window` field, which is the base's to give, and the module stopped importing: probe_identity, which imports every flm operator, failed. declare() now takes the class whose attributes the library reads, Operator, and checks a new field against its classes alone. Co-Authored-By: Claude <noreply@anthropic.com>
GEMV's L1 account, added with the prologue fold, read `self.dev` in compatible(), the current device, not the one `resolved(dev)` was given. Where none is bound (a device-free run, or an NPU busy elsewhere), the check raised AttributeError, and rejected_shapes' finish-input test failed; where another device is bound, GEMV was held to its memory. The account now runs in resolve(dev), on the resolved tiles, with the same refusal. Co-Authored-By: Claude <noreply@anthropic.com>
…h lower= tests/common/cases.py kept a second list of operator shapes for the device-free lowering gate, beside each operator's own Testing; the two had to be kept in step by hand, and it chose the shapes that exercise a layout or dtype decision well. This moves that choice onto the cases themselves: `Case(lower=True)` (and `Sweep(lower=True)` for a sweep's first default case) adds a case to the gate, which lowers each operator's first default case on every device in DEVICES plus every case it flags. The flagged ones carry over cases.py's coverage: Copy's float32, slot and two-channel by-head forms, Repeat's int32, weighted RMSNorm, Softmax at block 1024, GQA's three groups a column, MHA's chunk over a cache and interleaved decode, GEMV's first batched case, and GEMM's column-major layouts. GEMM's f32-out case joins its Testing, so it is also checked on a device. lowering.py lowers 98 operator cases across npu1 and npu2; the device-free set with it: 521 passed, 3 skipped. Co-Authored-By: Claude <noreply@anthropic.com>
A chain (Prepare, Finish) is gated by the interval its steps' tolerances admit around each reference value, composed step by step. `_interval` widened that interval past what `aie.utils.verify.compare` passes: - `_rounded` stepped `nextafter` away from the side it rounded toward, so each edge landed one dtype step outside the admitted range; - the relative radius was the same on both sides of the centre, where `compare`'s `rtol * (|a| + |b|)` reaches further away from zero than toward it; - the strict `<` of `nearly_equal` and of an ulps `atol` was read as `<=`; - `range_frac` widened the bound and exact kinds, which `compare` applies it to neither of; - ulps steps were counted from the float value rather than from its bfloat16 rounding, as `bf16_ulp_distance` counts them. Each kind is now the set of tests `compare` runs, each an interval around the centre with its own radius below and above and its own strictness, and an edge is the dtype value nearest it on the inside. `test_a_chains_interval_is_what_compare_admits` holds every kind to `compare` itself: both edges pass, and one step past either fails, over 4096 random values with zeros and subnormals. Before this, each kind admitted thousands of values `compare` rejects. The chain gates tighten by those steps. Co-Authored-By: Claude <noreply@anthropic.com>
The emulated AIE2P float multiply splits each operand into three round-to-nearest bfloat16 limbs, which is what `aie.iron.kernels.datamovement.limbs_f32_split` computes for the Limbs kernel's reference. fmul now takes its limbs from there. Bit-identical to the inline split over 200k products at each of five scales (1e-30 to 1e30, overflow included) and the edge values. Co-Authored-By: Claude <noreply@anthropic.com>
`NPUKernel.__call__` is the runtime's load_and_run with the call's dispatch scalars split from its keywords, which is what `OperatorImage.__call__` spelled out by hand; the image now calls its kernel. The scalars are checked as before, by the kernel's dispatch bridge. Co-Authored-By: Claude <noreply@anthropic.com>
mlir-aie's ExternalFunction takes `symbol_prefix=` and has no `digest_prefix=`; the factories prefix their symbols with the digest of their source and flags inside `_make_extern`, and flm's fused_mm_tile.cc is a factory now. An operator's own kernel names its configuration in its prefix, as merge.py does. Co-Authored-By: Claude <noreply@anthropic.com>
compilable_design_contract.py pins where CompilableDesign's cache key draws its line. Three of its tests restate the key's own behaviour (a closure's plain values, compile_kwargs and a stable text each change or keep the key), which is mlir-aie's contract to hold. The two that stay are the ones IRON's design rests on: an object closed over alone does not change the key, which is why IRON's generators carry their identity in compile_kwargs, and full_elf is part of it, which separates a fused dispatch's build from a separate one's. Co-Authored-By: Claude <noreply@anthropic.com>
lowering.py's lower() ran aiecc by hand, with its own flag list and Peano path. It now calls compile_mlir_module, the path a compiled design takes, so the gate lowers with the flags production uses (--unified among them) and follows them as mlir-aie changes them. The kernels are still built beforehand with compile_external_kernels: compile_mlir_module(device=) builds only the kernels with a source_file, so the inline kernel of inline_kernel.py would be left unlinked. An issue draft for that filter is in issues-to-file. lowering.py, inline_kernel.py and lowering_graph.py: 130 passed, 1 skipped. Co-Authored-By: Claude <noreply@anthropic.com>
The graph's tracer and its version lookup named a tensor's dtype through a table of six names. np.dtype(t.dtype).type gives the same scalar type for those six, and for the rest the scalar type rather than a dtype object, which Handle and bfp.dtype_name both take; it also reads a runtime tensor, whose dtype is the bare class. Device-free set (iron/tests/common/ and the device-free operator tests): 415 passed, 2 skipped. Co-Authored-By: Claude <noreply@anthropic.com>
compile() and __call__ named a version by the same (name, shape, dtype) tuple, built twice, once from the traced handles and once from the given tensors. _signature now takes the shapes compile() is called with, so __call__ builds them once, looks the version up by them, and passes them on when it compiles. load() reaches the runtime through the callable property, as a first call does. Device-free set (iron/tests/common/, the device-free operator tests and iron/tests/toolchain/narrowing.py): 501 passed, 2 skipped. Co-Authored-By: Claude <noreply@anthropic.com>
CostCache.put wrote its record to a per-process file and renamed it into place, as narrowing's _write does for a fit verdict. put now calls _write, whose parameter is named for any text. Device-free set (iron/tests/common/, the device-free operator tests and iron/tests/toolchain/narrowing.py): 501 passed, 2 skipped. Co-Authored-By: Claude <noreply@anthropic.com>
The dequant layout test and the dequant-feeds-GEMM test each rounded float32 to bfloat16 toward negative infinity by hand, as the cores do; aie2p_math_emulation's rb is that rounding, so both call it. Device-free set (including iron/tests/operators/flm_dequant_layout.py): 501 passed, 2 skipped. flm/dequant/test.py runs on npu2 with the flm queue. Co-Authored-By: Claude <noreply@anthropic.com>
mlir-aie is a requirement of IRON, so the toolchain conftest imports it at the top rather than through importorskip. Device-free set (including iron/tests/toolchain/narrowing.py, which the conftest serves): 501 passed, 2 skipped. Co-Authored-By: Claude <noreply@anthropic.com>
The root conftest's npu1 and npu2 fixtures, the toolchain conftest's device fixture and npu2 override, and two npu1 builds in the xclbin gate each bound a device and restored the previous one by hand. The root conftest now makes its npu1 and npu2 from one factory; the toolchain's device fixture is the root fixture its parameter names, and the npu1 builds take the npu1 fixture. tools.DEVICES stays for collection, where lowering.py resolves each catalog case's shapes, and for compile(dev=), which binds a device of its own. inline_kernel.py, xclbin.py's npu1 GEMV build, and the device-free tests that take npu1 or npu2 (common/build.py, elementwise.py, graph.py, mha_buffer_shapes.py, gemm_tile_divisibility.py): 138 passed. inline_kernel.py, xclbin.py and lowering.py collect 126 tests over both devices. Co-Authored-By: Claude <noreply@anthropic.com>
The flm GEMM keyword-construction test passed `9 / 8`, the bytes a bfp16ebs8 element takes, where the operator computes it as `bfp.BLOCK_BYTES / bfp.BLOCK`. It takes the same expression. Co-Authored-By: Claude <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Note, this PR relies on: Xilinx/mlir-aie#3884
PR Merge Checklist
develcommit and pointing todevel.