Sparsity on the ET-SoC-1: what the silicon skips, and where it can beat an A100

18 September 2026 · first measured on one card, aifoundry3, in the AI Foundry lab; every probe and every power figure re-measured on three cards (aifoundry2, aifoundry3 and aifoundry1 card 1) on 25–26 September · 1,024 compute minions, minion clock 600 MHz in every sample · code: workloads/sparsity; raw data: docs/reports/data/2026-09-18-sparsity-aifoundry3, and for the three cards docs/reports/data/2026-09-25-claims-v3 · A100 figures are published measurements, or estimates marked as such (sources at the end) · part of the ET-SoC-1 measurement reports

The question was whether sparse or irregular compute favours this chip over an A100, and where. The silicon turns zeros into saved power, never into saved time, so any speed-up has to come from software that issues less work, and even then one batch-1 layer only reaches an A100's latency.

Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan on aifoundry3, aifoundry2 and aifoundry1 card 1 (three passes of every probe per card, three strict-start runs of every energy configuration). Of 33 claims tested here, this page counts 23 held, 6 corrected and 4 differ by card; the hub’s scoreboard, 10 “proven on the cards”, 19 “a test behind it failed” and 4 “differs by card”. The cycle counts of the TensorFMA, the TensorLoad sweeps, the layer and the divergence loop repeat on every card, most to 0.1%. The corrections: a lone minion's masked TensorLoads bottom out at 42–43 cycles, not 45–47, and its 16-line DRAM load takes 749–751 cycles, not 730; zeros save 86–93% of the loop's power, not 86%; the layer's energies are now per card; and a launch from the host takes up to 0.6 ms. What differs by card: the power left once A is all zeros, whether masking rows matches zeros on aifoundry1 card 1, and how often the power meter refreshes. The check's above-idle watts, given here as registered, carry a launch-temperature offset (the check's note C2, revised): its leakage correction references fixed launch temperatures (80.9 °C on aifoundry2 and aifoundry1 card 1, 55.8 °C on this card), while the runs launched at a whole-degree reading of the die sensor (about 80, 80 and 57 °C), and whether the die sat at that reading or up to 0.96 °C above it cannot be settled. So the registered values read from 0.5 W high to 0.3 W low on aifoundry2, 0.8–1.3 W low on this card and from 0.7 W high to 0.05 W low on aifoundry1 card 1. At the same die temperature this card switches about 0.96–0.97 of aifoundry2's power in the ablations (the energy manual's catalogue, measured separately, gives 0.972), so most of this card's lower readings below come from the reduction, not the card. Differences between runs on one card (the tensor unit on zeros against the spin loop) cancel the offset; ratios such as the saving do not. Every power and energy figure on this page is the check's; the TensorLoad chart and table also offer this page's first run, and the first run's power figures, measured without temperature control, are one note in the method. Record: docs/reports/data/2026-09-25-claims-v3.

Time saved by zero-skip
0 cycles
546 cycles per fp32 op at 0–100% zeros; int8 318 at any sparsity; on all three cards
Power saved by zero-skip
86–93%
+17.5 → +2.5 W above idle on aifoundry2, +17.3 → +1.4 W on aifoundry1 card 1, +15.5 → +1.0 W on this card (registered values, each with a launch-temperature offset of up to about a watt; with busy and idle power compared at the same die temperature, over note C2's two readings of the launch temperature, the saving is 86–90% on aifoundry2, 92–96% on aifoundry1 card 1 and 86–88% on this card), in a TensorFMA-only loop, operands −3…3 (3.4–3.8 → 0.2–0.5 pJ per multiply slot); 82–97% from all ones or random normal data to all zeros, A and B alike
Batch-1 layer, 99% zeros
2.1 µs
7.5 µs dense (7.3 with plain loads); 1024×4096 fp32 in SRAM, 1,024 minions, on-chip reduction; on all three cards
Board energy per layer
60–114 µJ
at 99% zeros, including idle: 60 on this card (idle 25.8 W), 85 on aifoundry2 (36.4 W), 114 on aifoundry1 card 1 (50.2 W); dense 246–446 µJ (host-reduced kernel, 7.4 µs per layer)
Terms used on this page

Kernels run on 1,024 small RISC-V cores called minions, 32 to a shire; each has two hardware threads (harts), an 8-lane vector unit and a tensor unit, and only hart 0 issues tensor operations. A shire's 4 MB of SRAM holds a 512 KB L2 cache, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB scratchpad that any shire can address (the "L2 scratchpad" below), and 8 memory shires on the mesh network-on-chip drive the DRAM. TensorFMA is the tensor unit's matrix multiply-accumulate (16×16×16 in fp32, 16×16×32 in fp16, 16×16×64 in int8), TensorLoad fills the minion's L1 scratchpad (part of its L1 data cache) and TenB streams B from L2 instead; a multiply slot is one of an fp32 op's 4,096 multiply-adds, zero or not. More in the hub's glossary.

Zeros save power, never cycles

Method, and the full manual-vs-RTL-vs-measured table for all five hooks

The Programmer's Reference Manual describes three sparsity hooks: zero-skip in TensorFMA, and tensor_mask on TensorFMA and on TensorLoad. The table sets what the manual promises against what the open RTL implies and what the card measured, and adds rows for fp16 and int8, which behave differently. The RTL is core-et's Erbium branch, read for this report: the same Minion core lineage in a later MCU-class configuration, not the taped-out ET-SoC-1. Each probe ran on one minion; the fp32 TensorFMA and the TensorLoad probes also ran on all 1,024 at once. Inputs were nonzero integers from −3 to 3, plus the zeros under test, so every result could be checked exactly on the host. All of them were exact, and so were all of them again in three passes on each of three cards in the three-card check, whose TensorFMA cycle counts matched this table's.

HookProgrammer's Reference ManualRTLMeasured on the card
TensorFMA zero-skipfp32: the multiply-add is skipped when an A or B element is zero (PRM 9.4 pseudo-code).A zero 32-bit word gates that lane's multiplier. The sequencer still issues one 8-lane micro-op per cycle. The vector unit's specification (the Minion VPU Specification) says the gating is there "to save power".No cycles saved. 546 cycles per 16×16×16 op at 0–100% zeros, whether zeros are scattered, whole columns or whole rows, in A or in B. Power falls in step with the zeros (chart below).
fp16 zero-skipThe pseudo-code updates only when all four of a₁, a₂, b₁, b₂ are nonzero. Taken literally, a valid product would be dropped whenever its partner element is zero.Skips only when a whole 32-bit pair is zero, as the simulator does.Correct arithmetic. A pair with one zero still adds its other product. The manual's rule is a documentation bug. 546 cycles at any sparsity.
int8 (TensorIMA8A32)No skip in the pseudo-code.Gates a lane when its 4-byte A group is zero, or when both of its B words (columns j and j+8) are zero.318 cycles per 16×16×64 op at any sparsity.
tensor_mask on TensorFMAMasked rows of C "generate no calculations".Masked rows issue no multiply-adds, but their micro-op slots are still spent.No cycles saved: 544 cycles with 8, 4, 1 or 0 rows enabled (546 with all 16). Masking 8 of 16 rows cuts power like zeros do: +9.2 instead of +17.5 W on aifoundry2 and +7.8 instead of +15.5 W on this card, each as much as half zeros within its 99% range; on aifoundry1 card 1, +8.8 instead of +17.3 W, the match could not be confirmed.
tensor_mask on TensorLoadA masked row generates no memory access.No request is sent for a masked row.Time scales with the rows loaded: 160 → 80 → 42–43 cycles for 16, 8 and 4 of 16 rows from L2 or the scratchpad, and 749–751 → 320 → 144 from DRAM on one minion, on all three cards (160 → 81 → 45–47 and 730 → 320 → 144 in the first run). Not when the whole chip streams from DRAM (next section).
Cycles per TensorFMA against the fraction of A that is zero
fp32 and fp16 (identical)int8fp32 if zeros saved their cycles

One minion, A and B held in the L1 scratchpad. With B streamed from L2 through TenB, as in the matmul report, an fp32 op takes 529 cycles, also flat (that report's int8 loop, which streams both tiles, takes 280 on all three cards, against 318 here; with A held here and only B through TenB an int8 op takes 270). All 1,024 minions at once: 546.1. The same counts on all three cards.

Board power above idle against the fraction of multiplies that are not zero, on three cards (26 September)

The three-card check: each point the mean of three runs of 5 s, each launched at a set die temperature (about 57 °C on this card, 80 °C on the other two) and measured above the idle just before it; each card has its own least-squares line. The points are the registered values, which carry each card's launch-temperature offset (the note: from 0.5 W high to 0.3 W low on aifoundry2, 0.8–1.3 W low on aifoundry3, from 0.7 W high to 0.05 W low on aifoundry1 card 1). Every point runs 1.10–1.12 × 10⁹ ops/s, launch gaps included. Operands are nonzero integers from −3 to 3 apart from the zeros. The row-mask point has a dense A with 8 of its 16 rows masked off. The same loop on other values is in the chart below.

So the gating works, and it is worth a lot of energy. On these operands a dense op costs 3.4–3.8 pJ per multiply slot above idle on the three cards (5.4–6.1 pJ on random normal data, below; registered values, see the note). An all-zero op still costs 0.2–0.5 pJ per slot, and part of that is the harts simply running: an integer loop on the same harts draws 1.6 of 2.5 W on aifoundry2, 0.7 of 1.4 W on aifoundry1 card 1 and 0.3 of 1.0 W on this card. These are V3-ABL-B's own spin runs, three blocks on each card; V3-ABL-A's integer loop, which the other pages quote, drew 1.5, 0.7 and 0.4 W; the launch-temperature offset is as large as these, so which card's harts draw more is not settled. With all zeros the tensor unit itself adds about 0.7 W on this card and on aifoundry1 card 1 (99% ranges 0.1–1.3 and 0.04–1.3 W), against 15–17 W dense; on aifoundry2 its 0.9 W is not separated from the spin baseline (99% range −0.9 to 2.6 W). In this loop 87–94% of the power above idle follows the nonzero count (87% on aifoundry2, 91% on aifoundry1 card 1, 94% on this card), 94–96% of the tensor unit's own 15–17 W. But no zero is ever faster than a nonzero.

The values matter as much as the zeros. The three-card check also ran this loop (all 1,024 minions, A and B in the L1 scratchpad, 546 cycles per op) with every operand zero, every operand one and random normal data, four runs of 7 s per card launched as above (V3-ABL-A). Above idle it drew +1.9, +10.4 and +27.2 W on aifoundry2, +0.8, +8.8 and +24.7 W on this card and +1.2, +10.1 and +27.8 W on aifoundry1 card 1: 0.40, 2.26 and 5.92 pJ per multiply-add on aifoundry2, 0.16, 1.93 and 5.37 on this card and 0.25, 2.21 and 6.06 on aifoundry1 card 1, the per-card values of the energy manual, §3.2. So zero-skip saves 82–92% of the loop's power above idle against all ones and 93–97% against random normal data (78–95% and 92–98% with busy and idle power compared at the same die temperature, over note C2's two readings; these registered values read from 0.5 W high to 0.3 W low on aifoundry2, 0.9–1.5 W low on this card and from 0.7 W high to 0.05 W low on aifoundry1 card 1). The −3…3 operands of the chart above sit between the two. The Horace experiment explains why by simulating the open RTL of the multiply-add unit: power follows the register bits clocked and the nets toggled, not the count of nonzero operands, and its energies per event, fitted on aifoundry2, carry over to the other two cards up to one scale factor per card (its §10).

Board power above idle of the same loop by operand values, per card (three-card check, 26 September)

Each stem is one card's mean power above the idle just before its runs, as registered (the note); the column on the right repeats it with the energy per multiply slot, all 4,096 of an op's slots counted, zero or not. The rows with A and B alike are V3-ABL-A's runs (four of 7 s per card); the two rows of this page's loop, whose nonzero operands are −3…3, are the end points of the chart above (V3-ABL-B, three runs of 5 s). Hover or focus a stem for its runs' range, its launch temperature and its idle.

A kernel that wants time back from sparsity has to issue fewer or smaller TensorFMAs: skip all-zero tiles, or shrink the op's row and column counts (its AROWS and ACOLS fields, each the count minus one). The tensor unit will not do it by itself. (On A0 silicon, erratum 1.29 type D rules out some small fp32 and fp16 ops whose ACOLS field is nonzero: 4-row ops with B in the L1 scratchpad, and 1- to 4-row ops with B streamed through TenB.)

Masked loads save time, except when the chip saturates DRAM

A TensorLoad moves up to 16 lines of 64 B into the L1 scratchpad. With tensor_mask set, a masked line should never be requested. On one minion the time falls with the lines requested. From L2 and the shire's scratchpad it falls until a floor, the load's latency: 42 and 43 cycles on all three cards in the three-card check (45 and 47 in the first run). From DRAM it falls almost in proportion, from 749–751 cycles for 16 lines to 73 for one on all three cards (730 in the first run): a lone minion's requests go out at about 47 cycles per line, far below what DRAM can deliver. In the first of the check's three passes the 16-line DRAM load took 1,132 cycles on every card, while 15 lines and fewer took the same time in every pass; the cause was not found. (The DRAM buffer is 64 MB, larger than the 32 MB L3, so every load misses the caches.)

Cycles per 16-line TensorLoad against the lines the mask lets through, one minion (log scale)
DRAML2 (8 KB per minion)local L2 scratchpad (dashed; it lies on the L2 line except at 4 lines)

With all 1,024 minions loading at once, on-chip memory behaves the same way. DRAM does not. At 16 lines per load this short probe (2,000 loads per minion) streams 72 GB/s from DRAM counting the slowest minion (71–73 GB/s on all three cards). The same probe with 20,000 loads per minion reads 75.6–75.9 GB/s on all three cards, the 76 GB/s measured in the memory-hierarchy report. Asking for fewer lines hardly shortens a load, so the useful bandwidth falls with the mask instead of the time. Each minion streams 1 KB blocks, and every mask tested keeps the same low rows of every block, so all the kept lines are homed in a few shires' L3 slices (see Caveats). Masks that vary from block to block were not tested. What matters for kernels is that weights to be skipped should sit in SRAM.

Why: the home-shire hypothesis, charted, and the full bandwidth table

This chart, the table below and the TensorLoad chart above show: first run, 18 Sep, aifoundry3 (the Data buttons above the TensorLoad chart choose).

The page's hypothesis: fewer home shires answering, not less DRAM bandwidth from each
(a) 32 home shires (L3 slices, address bits 10:6)
(b) 8 memory shires (DRAM controllers)

Shire positions here are illustrative, not the physical layout; only which count is lit is measured, from the address mapping in Caveats. GB/s per shire divides the chip's measured DRAM bandwidth at that mask (from the table below) by the shires it should reach. If the hypothesis holds, (a)'s per-shire rate stays flat across masks while (b)'s does not, since the 8-line mask still reaches all 8 memory shires.

All 1,024 minions16 lines8 lines4 lines1 line

Cycles per load, mean over the 1,024 minions. In brackets, the bytes the mask let through per second over the whole chip, using the slowest minion's total time at 600 MHz (so they are a few percent below what the mean cycle count implies). The three cards give the median of their three passes: they repeat the first run's L2 and scratchpad counts to within 1.5% (the L2 at 16 lines varies from 286 to 306 cycles between passes), and read DRAM 3–4% faster at 16 lines, within 2% at 8 and 4, and 4–5% slower at 1 line. At 16 lines L2 reads 2.03 TB/s here (1.96–2.13 TB/s over the check's nine passes), below the 2.45 TB/s the memory-hierarchy probe reads on both cards, including this one on 23 September at a pinned 600 MHz (Ridge points). This probe is shorter (2,000 loads per minion). Run with 20,000 or 200,000 loads per minion it reads 2.45 and 2.46 TB/s on all three cards, so the gap belongs to the short probe; the three-card check could not tie it to a cold start, since a separate 2,000-load run varied from 1.63 to 2.37 TB/s between passes.

A batch-1 layer that skips zeros

Method: how the layer kernel works and was timed

This is the case most relevant to ML: one output of a ReLU layer, y = W x with a batch of one, where most of x is zero. The layer is 1024 outputs × 4096 inputs in fp32 (16.8 MB of weights). It is spread over all 32 shires, and each shire keeps its 512 KB share of W in its own L2 scratchpad. Minion j of shire s owns 16 outputs and 256 inputs, in 16 slices of 16. For each slice it loads only the rows of Wᵀ whose x is nonzero (a masked TensorLoad), then runs one single-row TensorFMA. Where all 16 x values of a slice are zero, it skips the slice. Sixteen minions then add their partial rows with a TensorReduce tree that stays inside the shire.

Latency runs from a shire barrier to the moment the block's result is in registers, averaged over 2,000 repeats per sparsity level. Each level uses one random x, the same for every variant. Every output was checked exactly against the host. The host supplies a bitmask of which x are nonzero. In a real pipeline the layer that produced x would build it, which takes two vector instructions per 8 values (fltm.ps, then read the mask).

What the layer skips at each sparsity, and its latency against an A100
(a) One minion's 256 inputs: 16 slices (columns) of 16 rows
(b) Latency of one layer (µs) against the zeros in x (uneven steps)

The slider and buttons choose a zero fraction and a kernel; press Enter on a point of the curve to choose it. (a) draws a representative minion with the measured means: how many rows of Wᵀ it loaded and how many of its 16 slices it computed; which rows and which slices are illustrative. The same random x served every kernel at each level. The A100 marks are references, not bounds (see below).

All layer timings (µs) and energies
Zeros in x0%50%75%90%95%99%100%

µs at 600 MHz: the time of each 16-output block's slowest minion, mean over the 64 blocks. The slowest of the 64 blocks averaged 7.62 µs dense and 2.22 µs at 99% zeros, and a layer is not finished until its slowest block is. Results stay in each block's root registers; writing y out is not included. "Compute only" leaves out the reduction: every minion writes its partial row and the host adds them.

The energy rows come from separate runs of the masked + skip compute-only kernel with the power logger on, which drew their own x. In the three-card check each of the three runs per card drew a different x (seeds 1 to 3), so at 99% zeros its layer took 2.7 µs in one run and 1.9 µs in the other two; the rows give the mean of the three, and each card's board energy includes its own idle (36.4 W on aifoundry2, 25.8 W on aifoundry3, 50.2 W on aifoundry1 card 1). Each card's above-idle rows carry its launch-temperature offset times the layer time (the note): up to about 5, 4 and 2 µJ at 0, 90 and 99% zeros on aifoundry3 (read low), and less on the other two cards. Layer times come from the layer rate, which also includes the shire barrier between layers.

Three effects add up. Masked loads cut the bytes, so at 50% zeros the layer drops from 7.5 to 5.7 µs. After that the per-slice fixed costs take over. Each nonempty slice waits for a load (at least 42 cycles) and an FMA, and the curve flattens at 5.4 µs. These times repeat to within 0.01 µs on all three cards. The big gains come only from skipping whole slices, which random sparsity allows only near the top: for independent zeros 18.5% of slices are empty at 90% zeros and 85% at 99% (0.9¹⁶ and 0.99¹⁶; the x drawn here had 14.9% and 87.5%). Energy per layer above idle falls from 67 to 5.9 µJ on aifoundry2 (19 at 90% zeros), from 75 to 4.5 µJ on aifoundry1 card 1 (16) and from 56 to 3.3 µJ on this card (12), mostly because skipped slices and masked loads do less work in less time. The tensor unit's own gating adds no speed and little energy here: with every row loaded, 90% zeros in x cut the layer's power above idle only from 9.9 to 8.2 W on aifoundry2, 11.0 to 9.3 W on aifoundry1 card 1 and 8.4 to 6.8 W on this card, 16–19% (the same every-row kernel at both points).

What the whole board spends per layer is another matter: once most slices are skipped, the power the card draws doing nothing, over the layer's few microseconds, is nearly all of it.

Energy per layer on the board: the card's idle over the layer's time, and what the layer added above it

How the A100 comparison numbers were built up

Against the A100. No one has published a timing for this exact layer on an A100, so the reference lines are rough references, not bounds. Launching even an empty kernel costs 2.26 µs, plus 1.44 µs for the kernel itself (measured on an A100-SXM4-80GB). Kernels queued back to back in a stream cost less, about 1.5–2.3 µs each on an A100 (on-chip communication report). cuBLAS half-precision GEMM has a latency floor of about 4.5 µs. Streaming 16.8 MB of weights from L2 at 2.5–4.5 TB/s takes 3.7–6.7 µs. Added to the 2.26 µs launch, or to cuBLAS's 4.5 µs floor, a dense fp32 layer probably takes 6–11 µs, against 7.3–7.5 µs here. CUDA graphs cut launch cost to about 1.3 µs per kernel (measured on H100). At batch 1 a GPU's sparse GEMV is floor-bound too: TEAL's kernel gets 3× at 99–100% sparsity on a 4096×14336 fp16 layer, with 14× this one's weights and 7× its bytes (102 → 33 µs), and Deja Vu reported 1.8–2× end to end at about 75% sparsity (OPT-175B on 8 A100s, against FasterTransformer).

So at batch 1, on these runs, ET-SoC-1 reaches A100-class latency on one layer. It beats a single isolated A100 launch (3.7 µs) only at 99% zeros or more, and there (2.1 µs) it only matches a queued A100 kernel (1.5–2.3 µs). That comparison favours ET: its 2.1 µs is a layer inside a kernel that is already running (a launch from the host took 0.2–0.6 ms on the three cards, median 0.36 ms on this one), and a GPU can avoid launches too, with a persistent kernel whose grid-wide sync costs 1.1–1.2 µs (on-chip communication report).

Board energy, and what a stack of layers would still need

The whole board spends 60–246 µJ per layer on this card, idle included (85–337 on aifoundry2 and 114–446 on aifoundry1 card 1, whose idles are higher). No A100 energy for this layer has been published; the benchmark in scenario 2 would measure it. The larger prize is a stack of such layers. A pipeline of tens of small layers could stay on chip with no launch between layers, but each layer would still have to pass its outputs across shires (a chip-wide allreduce, §5), and a GPU could run the same stack as one persistent kernel. None of this is measured here.

Divergence: narrow lanes help, but not enough

Method: the workload and the three lane-scheduling variants

The other half of the question is irregular work, here divergence. Each work item needs k iterations of an FMA loop (8 independent chains), and k is drawn from a heavy-tailed Pareto distribution with mean 64. The case for ET-SoC-1 is that its SIMD is only 8 lanes wide and each core runs its own code, so a long item wastes less. All 2,048 harts took 32-item chunks from a shared work queue (a global atomic counter), 524,288 items in all, and every result was checked. Three variants ran, once each, and again in three passes on each of three cards in the three-card check, which repeated every figure in this section. Static runs 8 items per 8 lanes until the longest finishes, as a GPU warp would. Refill gives a lane its next item as soon as it finishes. Scalar runs one item at a time on one lane.

Lane efficiency against the heaviness of the tail
ET refillET static, 8 lanesA100 32-lane warp, model
Useful FMAs per second, whole chip (T/s, log scale)
ET refillET staticA100 naive kernel, model

Static 8-lane groups keep busy exactly the share of lanes that grouping the same items in eights predicts: 0.55 at α = 3 and 0.21 at α = 1.2, against 0.35 and 0.08 for a 32-lane warp on the same items. (The continuous order-statistics formula gives 0.55 and 0.19; the cap at 16,384 iterations trims the heaviest tail.) Refilling lanes raises ET's efficiency to 0.96 and 0.46 (0.962–0.963 and 0.453–0.459 over the three cards' passes; which lane gets which item depends on timing, so the refill figures move in the third decimal), but its bookkeeping costs cycles: at α = 3 it does fewer useful FMAs per second than static (0.59 against 0.73 T/s, on every card), and it pulls ahead only from α = 2 (by 0.016 T/s there, a 99% range within 0.012–0.020 T/s on each card over five runs).

It is not enough to beat a modelled A100. This loop peaks at about 1.5 useful FMAs per cycle per minion, or 0.9 T/s for the chip, under a fifth of the 8-lane peak (4.9 T/s at 600 MHz). An A100 has 9.75 T FP32 FMA/s. A naive kernel that keeps only 8% of its lanes busy at α = 1.2 would still do 0.74 T/s if it issued FMAs at the A100's peak (a model, not a measurement), against ET's 0.39 T/s measured on each of three cards. That 1.9× gap is within what tuning might recover, but at α = 1.5–3 the modelled A100 stays 2.3–4.6× ahead of ET's faster variant, and GPUs have compaction and refill tricks of their own.

Investigating the work-queue hypothesis, charted

What limits the loop was not measured, and the shared work queue may be part of it. A contended global atomic retires one operation every 10 cycles (One hot line stops a shire). The 16,384 chunk grabs plus each hart's final empty grab need about 184,000 cycles of it, more than the average hart's whole run without divergence; the chart below sets that floor against every run. An 8-lane iteration also takes 84 cycles against 33 for a one-lane scalar iteration, so the loop's mask test and branch cost too. The scalar variant, which uses the queue at a third of the rate, held 0.30 T/s at every α. Larger chunks or one queue per shire would separate the two.

Could the one global work queue be what limits the loop? Mean hart cycles against the queue's floor

Bars: the mean hart's cycles in each run; the slowest hart is in each bar's tooltip. Dashed line: the cycles the queue alone needs if every grab waits its turn, (items ÷ chunk + harts) × cycles per contended atomic, here (524,288 ÷ 32 + 2,048) × the slider's value. 10 is the ceiling measured in One hot line stops a shire; the slider shows how the floor moves if the true cost differs. Only 32-item chunks were run.

On this evidence (three cards, against a modelled A100), divergence in how long items run is not an ET-SoC-1 advantage. What the chip has that a GPU lacks is one instruction stream per core. That matters when items run different code, as in interpreters, simulators and heterogeneous state machines, which this sweep does not test.

Where ET-SoC-1 should beat an A100, and how to prove it

These measurements and the published literature give one recipe for a win. The work must be latency-bound or event-driven, so that a GPU sits at its per-kernel or per-sync floor. Its state must fit in the 80 MB of on-chip scratchpad, or be regenerated rather than fetched. And each step's surviving work must be dense enough for 8-lane SIMD or the tensor unit.

The sync-cost numbers behind the estimates below

The chip's sync costs come from the on-chip communication report, re-measured on three cards on 26 September (aifoundry2, aifoundry3 and aifoundry1 card 1, three passes each, within 0.3% of each other). A 1,024-minion allreduce takes 2.3 µs for 32 B and 5.4 µs for 1 KB, a chip-wide barrier from atomics and credits 8.3 µs, and a shire allreduce 0.74 µs (444 cycles). A message round trip takes 113–190 ns inside a shire and 250 ns + 20 ns per mesh hop across shires.

Each scenario below gives why a GPU struggles, the best published baseline (an A100 result where one exists; otherwise the hardware is named), an estimate for this chip from the measurements above, and the benchmark that would settle it. Only scenario 2 is measured here.

Scenario 1: spiking cortical microcircuit (unmeasured estimate)

1. Spiking cortical microcircuit

The Potjans–Diesmann model (PD14): 77,169 neurons, 3×10⁸ synapses, 0.1 ms steps. RTF is the real-time factor, wall-clock time over model time.

Why a GPU struggles
About 25 spikes per step. Each step is several tiny kernels, and launch latency alone caps GeNN at RTF ≈ 0.1.
Published baseline
A100 (Golosio et al. 2023): GeNN RTF 0.539 (54 µs of wall time per step); NEST GPU 0.73–0.85. No A100 energy per synaptic event has been published. RTX 4090: GeNN 0.272; about 74 nJ per synaptic event, estimated by Senk et al. Fastest full-scale run: neuroAIx (35 FPGAs) 0.049 at 48 nJ, the lowest energy measured at the outlet.
ET-SoC-1 estimate
About 75 neurons per minion, resident in L1/L2, with synapses regenerated from a counter-based RNG. Per step: an exchange of the ~25 spike IDs through the hardware tree (2.3–5.4 µs for 32 B to 1 KB, above; an all-gather of a variable-length list needs fixed slots or a prefix sum), plus the compute: about 95 synaptic events and 75 neuron updates per minion per step, each event's target, weight and delay regenerated from the RNG, with no hardware divide or square root. That compute was not measured. If it fits in about a microsecond, a step takes about 4–10 µs, an RTF of 0.04–0.1, 5–13× the A100. At about 30 W this would be a few nJ per synaptic event: an estimate, not a measurement.
Benchmark that would settle it
PD14 under the community recipe (Senk et al. 2026): RTF, and energy per synaptic event measured at the outlet, over ≥ 10 s of model time and 10 seeds. Validate with rate, CV and correlation distributions, and declare procedural connectivity.

2. Batch-1 activation-sparse inference, SRAM-resident

Why a GPU struggles
Batch 1 is floor-bound: launch plus sync per layer. Gains from sparsity saturate on fixed costs.
Published baseline
A100: one isolated launch plus a null kernel 3.7 µs; cuBLAS fp16 GEMM floor ~4.5 µs. TEAL 1.33–1.45× end to end at 50% sparsity; Deja Vu 1.8–2× at ~75% (8×A100).
ET-SoC-1 estimate
Measured in §3, dense against 99% zeros: 7.5 → 2.1 µs per layer with the on-chip reduction, and 246 → 60 µJ of board energy per layer on aifoundry3 (337 → 85 on aifoundry2, 446 → 114 on aifoundry1 card 1) from the energy runs, which used the compute-only kernel whose partial sums the host adds. A model of ≤ 40 M fp16 parameters fits on chip, and every layer after the first avoids a kernel launch, though each still has to pass its outputs across shires (not measured).
Benchmark that would settle it
The same layer on an A100 with CUDA graphs, cuBLAS against a gather GEMV, 0–99% sparsity, time and NVML power. Then a whole small ReLU model at batch 1: µs and J per token.
Scenario 3: sparse triangular solve (unmeasured estimate)

3. Sparse triangular solve (ILU/IC preconditioner, deep DAG)

Why a GPU struggles
Thousands of thin dependency levels, each needing a sync.
Published baseline
A100 (FP64): cuSPARSE and a sync-free solver mostly ≤ 2 GFLOPS on DCSolver's 38 ILU(0)/IC(0) factors (7 to 916,791 levels). AG-SpTRSV is 3.99× faster than cuSPARSE (geometric mean over 2,219 matrices, FP64). tmt_sym takes about 0.57 µs per level with cuSPARSE on a Titan RTX (FP64, derived from its published 0.014 GFLOPS).
ET-SoC-1 estimate
Inside a shire, a level handoff by credit or message takes 113–400 ns. A chip-wide barrier per level (2.3–8.3 µs) would lose, so it only wins with per-shire subdomains and point-to-point credits. FP32 only: a mixed-precision preconditioner.
Benchmark that would settle it
A thin-level SuiteSparse subset: cusparseSpSV and AG-SpTRSV on the A100 against an ET version partitioned into per-shire blocks. Report solve time and time per level.
Scenario 4: RTL and gate-level simulation (unmeasured estimate)

4. RTL and gate-level simulation

Why a GPU struggles
Tiny tasks, and a sync every simulated cycle. The GPU needs thousands of stimuli to fill up.
Published baseline
GEM on an A100: 7.3–65 kHz for one stimulus. Verilator: 1–1000 kHz. Manticore (225-core FPGA, bulk-synchronous) is 2.1–4.2× over Verilator (geometric means over three x86 hosts, serial and multithreaded). Parendi, from Manticore's first and last authors, runs RTL simulation on 5,888 Graphcore IPU cores and reports up to 4× over the fastest x64 multicore hosts.
ET-SoC-1 estimate
Same bulk-synchronous model at 4.5× Manticore's core count. Synchronisation alone caps a shire at about 2.5 MHz (a 395 ns barrier every cycle) and the whole chip at about 430 kHz (the 2.3 µs tree). Gate evaluation and exchanging signal values between shires come on top, so real rates would be well below these ceilings. Several million gates fit in scratchpad.
Benchmark that would settle it
Manticore's designs: simulated kHz against Verilator at 1 and N threads, and GEM on the A100 with 1, 8 and 1K stimuli. Needs a partitioning compiler.
Scenario 5: interpreters over per-item programs (unmeasured estimate)

5. Interpreters over per-item programs (symbolic regression, GP)

GP is genetic programming. GPop/s counts GP operations per second, one expression node evaluated on one data row; it is not 10⁹ of anything.

Why a GPU struggles
Every item runs different code, so SIMT serialises the branches.
Published baseline
EvoGP 10¹⁰–10¹¹ GPop/s (RTX 3090/4090); Operon 6.8 × 10⁹ on one CPU thread, 94 × 10⁹ on 24 threads.
ET-SoC-1 estimate
One expression per minion, one data row per lane (8 rows per vector instruction), so the lanes never diverge. Guessing from this issue rate: about 10¹¹ GPop/s; 10¹² would need the interpreter to match the plain FMA loop's 0.9 × 10¹² lane-ops/s with no overhead. The sweep above says to measure it.
Benchmark that would settle it
SRBench-style evaluation: GPop/s against EvoGP on the A100 and Operon on x86.
Scenario 6: tree ensembles and motion planning (unmeasured estimate)

6. Tree ensembles and motion planning at batch 1

Why a GPU struggles
Small batches don't fill a GPU; per-query paths are data-dependent.
Published baseline
No published A100 batch-1 forest latency. Lettich et al. (IEEE TPDS 2019) parallelised QuickScorer (Lucchese et al., SIGIR 2015), a fast traversal of tree ensembles, with CPU SIMD, multi-core CPUs and GPUs, and got their best results on the GPU. Motion planning (Panda set, RTX 4090): pRRTC 0.48 ms median, cuRobo 30–62 ms (median–mean); VAMP 40 µs median on one x86 core.
ET-SoC-1 estimate
Throughput plays (1,024 independent queries), with x86 as the rival more than the A100. A single query on one minion will probably be slower than on one x86 core.
Benchmark that would settle it
MSLR-WEB30K with a LightGBM model: µs/doc at batch 1–100 against RAPIDS FIL. MotionBenchMaker: plans/s and tail latency against VAMP and cuRobo.

Where the A100 wins

Answers to the questions this work started from

The work began from two research briefs (not published). The sections above answer most of their questions; these two answers are easy to miss:

How it was measured

Method for every probe: kernel, timing, and each probe's design
Version history

Versions: published 18 September 2026; revised 24 September 2026 (operand values, later measurements, the masked-DRAM cause, the work queue, the A100 launch comparison) and 25 September 2026 (the layer explorer, the later runs on the power chart and the queue-floor chart; this card's meter rate; shorter answers); 25 September (version 3): which card and how many runs each figure rests on, aifoundry2's values beside this card's, and the tensor unit's own all-zero power no longer quoted as a number (record: docs/reports/data/2026-09-25-claims-v3); 26 September: the three-card check (aifoundry2, aifoundry3, aifoundry1 card 1) in the text, the KPIs, the power chart, the TensorLoad chart and table and the layer table; 27 September: the masked-load hypothesis charted (per card), scenario 2 names the kernel behind each figure and takes the three-card energies, and aifoundry3's 600 MHz traced to the boot service that sets its TDP to 0 W; later that day, the layer's energy per card charted, idle against what the layer adds; 28 September: a repeat cut; the note's counts given by both rules; later that day, every power figure the three-card check's: the operand values (all zeros, all ones, random normal) from its V3-ABL-A runs, in the text and a new chart, in place of the 22 September session on this card and the "Later measurements" box; the power chart, the layer table and the explorer without the first run, which is one note in the method; aifoundry1 card 1's saving at the same die temperature 92–96% on V3-ABL-B's registered metric (it said 91–96%, what the runs' dropout-rule metric gives).

Caveats

Caveats in full

Reproduce

Reproduce this
scripts/deploy-lab.sh aifoundry3 workloads/sparsity
ssh aifoundry3 'cd ~/nekko && bash workloads/sparsity/run_lab.sh build/sparsity/host/sparsity_host build/sparsity-data'
ssh aifoundry3 'cd ~/nekko && python3 workloads/sparsity/run_energy.py --host-bin build/sparsity/host/sparsity_host --out build/sparsity-data/energy-a'
ssh aifoundry3 'cd ~/nekko && python3 workloads/sparsity/run_energy.py --host-bin build/sparsity/host/sparsity_host --out build/sparsity-data/energy-b --only gemv-skip-99,gemv-skip-90,gemv-skip-0,gemv-dense-90,fma-rowmask,fma-col50,fma-zero,fma-875,fma-50,fma-dense,spin'
scp -r aifoundry3:nekko/build/sparsity-data/. docs/reports/data/DATE-sparsity-aifoundry3/
python3 workloads/sparsity/analyze.py docs/reports/data/2026-09-18-sparsity-aifoundry3 \
    --embed docs/reports/2026-09-18-et-soc1-sparsity.html

scripts/deploy-lab.sh copies the sources to ~/nekko on the lab machine and builds them there. Every card command runs in its own timeout 10 process with --budget 8, and waits until no other process holds the card. Raw data: docs/reports/data/2026-09-18-sparsity-aifoundry3/ (its README lists what each file is). analyze.py also reads the three-card check (docs/reports/data/2026-09-25-claims-v3: V3-LAT's raw TensorLoad sweeps and the per-run tables of V3-ABL-B and V3-ABL-A, results/ablb.runs.json and results/abla.runs.json) into the three-card views of the charts and tables, and prints the values the text quotes; the check's runs are tools/claims-v3/lat/, tools/claims-v3/ablb/ and tools/claims-v3/abla/. Without a card, add --sysemu to any sparsity_host command to check a kernel (the simulator's cycle counts mean nothing), e.g. build/sparsity/host/sparsity_host --sysemu --test fma --type fp16 --pattern pair --sweep 0,0.5,1 --shires 0x1 --per-shire 1.

Sources

Full source list
  1. ET-SoC-1 Programmer's Reference Manual (aifoundry-org/et-man), §9.2.1 tensor_mask, §9.4 TensorFMA32 / TensorFMA16A32 / TensorIMA8A32 / TensorLoad pseudo-code; ET-SoC errata 1.29. core-et (Erbium branch) RTL vpu_ctrl.v, vpu_tensorfma.v and the Minion VPU Specification, for zero gating and issue.
  2. Vellaisamy et al., Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures, arXiv 2504.11750, Table V (A100 launch 2.26 µs + null kernel 1.44 µs). Hazy Research, "Look Ma, No Bubbles!" (2025), H100 launch 2.1 µs, 1.3 µs with CUDA graphs. Erdil & Schneider-Joseph, arXiv 2411.01137 §5.1.1 (cuBLAS fp16 GEMM floor ~4.5 µs). Luo et al., arXiv 2402.13499, Chips and Cheese, and Zhang et al., GPGPU 2025, Fig. 6a (A100 L2 bandwidth).
  3. Liu et al., Deja Vu, ICML 2023 (arXiv 2310.17157). Liu et al., TEAL, ICLR 2025 (arXiv 2408.14690), Fig. 3 and Table 3. Mishra et al., arXiv 2104.08378, and Frantar & Alistarh, SparseGPT, ICML 2023 (arXiv 2301.00774), Table 8 (2:4 sparsity). Tsai, Cojean, Anzt, arXiv 2008.08478, and Zhang et al., GPGPU 2025 (A100 SpMV).
  4. Potjans & Diesmann, Cerebral Cortex 2014. Golosio et al., Appl. Sci. 13:9598 (2023), A100 GeNN and NEST GPU. Senk et al., Neuromorph. Comput. Eng. 6:012001 (2026), benchmark recipe and survey. Kauth et al., Front. Comput. Neurosci. 2023 (neuroAIx). Knight & Nowotny, Nat. Comput. Sci. 2021 (procedural connectivity).
  5. Hu et al., AG-SpTRSV, ACM TACO 21(4) 2024. Qiu et al., DCSolver, ACM TACO 22(3) 2025. Lu, Niu, Liu, ICPP 2020 (tmt_sym).
  6. Emami et al., Manticore, arXiv 2301.09413v4 (ASPLOS 2024). Guo et al., GEM, DAC 2025. Emami, Bourgeat, Larus, Parendi: Thousand-Way Parallel RTL Simulation, ASPLOS 2025 (arXiv 2403.04714). Thomason, Kingston, Kavraki, VAMP, ICRA 2024. Huang et al., pRRTC, arXiv 2503.06757. Sundaralingam et al., cuRobo, arXiv 2310.17274. Lucchese et al., QuickScorer, SIGIR 2015. Lettich et al., Parallel Traversal of Large Ensembles of Decision Trees, IEEE TPDS 30(9) 2019. Wu et al., EvoGP, arXiv 2501.17168. Burlacu et al., Operon, GECCO 2020.
  7. Pareto order statistics: E[max of n] = Γ(n+1) Γ(1−1/α) / Γ(n+1−1/α). Laine, Karras, Aila, HPG 2013, and Aila & Laine, HPG 2009, for GPU wavefront and refill techniques.

← All ET-SoC-1 measurement reports