Sparsity on the ET-SoC-1: what the silicon skips, and where it can beat an A100
The question was whether sparse or irregular compute favours this chip over an A100, and where. The silicon turns zeros into saved power, never into saved time, so any speed-up has to come from software that issues less work, and even then one batch-1 layer only reaches an A100's latency.
- Zero-skip saves power, not cycles. A 16×16×16 fp32 TensorFMA takes 546 cycles whether A is dense or all zeros, and the fp16 (546), int8 (318) and row-masked (544) ops are just as flat, on each of three cards. Board power above idle, though, falls by 86–93% as A goes from dense to all zeros in a loop of TensorFMAs on small-integer operands: from 17.5 W to 2.5 W on aifoundry2, 17.3 to 1.4 W on aifoundry1 card 1 and 15.5 to 1.0 W on this card in the three-card check (each card's values carry a launch-temperature offset of up to about a watt, see the note). The values matter as much as the zeros: with every operand nonzero the same loop draws 8.8–10.4 W above idle on all ones and 24.7–27.8 W on random normal data, and with every operand zero 0.8–1.9 W (§1). In a real layer, where loads dominate, zeros alone save far less.
- A layer gets faster only when the kernel skips work. Masked loads fetch only the rows a nonzero activation needs, and all-zero tiles are skipped outright. With both, a batch-1 1024×4096 fp32 layer held in the on-chip scratchpads goes from 7.5 µs dense (7.3 µs with plain loads) to 2.1 µs at 99% zero activations, including the on-chip reduction, to within 0.01 µs on all three cards (two other random draws of x at 99% zeros took 2.40 and 2.60 µs). That is about an A100's latency, not a multiple of it.
- Where it should win, and where it loses. The chip should win where a GPU is stuck at its per-kernel floor and the state fits in SRAM: event-driven simulation, batch-1 pipelines of many small layers, fine-grained wavefronts. It loses wherever the sparse data has to stream from DRAM, and, against a modelled A100, on divergence alone.
Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan on aifoundry3, aifoundry2 and aifoundry1 card 1 (three passes of every probe per card, three strict-start runs of every energy configuration). Of 33 claims tested here, this page counts 23 held, 6 corrected and 4 differ by card; the hub’s scoreboard, 10 “proven on the cards”, 19 “a test behind it failed” and 4 “differs by card”. The cycle counts of the TensorFMA, the TensorLoad sweeps, the layer and the divergence loop repeat on every card, most to 0.1%. The corrections: a lone minion's masked TensorLoads bottom out at 42–43 cycles, not 45–47, and its 16-line DRAM load takes 749–751 cycles, not 730; zeros save 86–93% of the loop's power, not 86%; the layer's energies are now per card; and a launch from the host takes up to 0.6 ms. What differs by card: the power left once A is all zeros, whether masking rows matches zeros on aifoundry1 card 1, and how often the power meter refreshes. The check's above-idle watts, given here as registered, carry a launch-temperature offset (the check's note C2, revised): its leakage correction references fixed launch temperatures (80.9 °C on aifoundry2 and aifoundry1 card 1, 55.8 °C on this card), while the runs launched at a whole-degree reading of the die sensor (about 80, 80 and 57 °C), and whether the die sat at that reading or up to 0.96 °C above it cannot be settled. So the registered values read from 0.5 W high to 0.3 W low on aifoundry2, 0.8–1.3 W low on this card and from 0.7 W high to 0.05 W low on aifoundry1 card 1. At the same die temperature this card switches about 0.96–0.97 of aifoundry2's power in the ablations (the energy manual's catalogue, measured separately, gives 0.972), so most of this card's lower readings below come from the reduction, not the card. Differences between runs on one card (the tensor unit on zeros against the spin loop) cancel the offset; ratios such as the saving do not. Every power and energy figure on this page is the check's; the TensorLoad chart and table also offer this page's first run, and the first run's power figures, measured without temperature control, are one note in the method. Record: docs/reports/data/2026-09-25-claims-v3.
Terms used on this page
Kernels run on 1,024 small RISC-V cores called minions, 32 to a shire; each has two hardware threads (harts), an 8-lane vector unit and a tensor unit, and only hart 0 issues tensor operations. A shire's 4 MB of SRAM holds a 512 KB L2 cache, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB scratchpad that any shire can address (the "L2 scratchpad" below), and 8 memory shires on the mesh network-on-chip drive the DRAM. TensorFMA is the tensor unit's matrix multiply-accumulate (16×16×16 in fp32, 16×16×32 in fp16, 16×16×64 in int8), TensorLoad fills the minion's L1 scratchpad (part of its L1 data cache) and TenB streams B from L2 instead; a multiply slot is one of an fp32 op's 4,096 multiply-adds, zero or not. More in the hub's glossary.
Zeros save power, never cycles
Method, and the full manual-vs-RTL-vs-measured table for all five hooks
The Programmer's Reference Manual describes three sparsity hooks: zero-skip in TensorFMA, and tensor_mask on TensorFMA and on TensorLoad. The table sets what the manual promises against what the open RTL implies and what the card measured, and adds rows for fp16 and int8, which behave differently. The RTL is core-et's Erbium branch, read for this report: the same Minion core lineage in a later MCU-class configuration, not the taped-out ET-SoC-1. Each probe ran on one minion; the fp32 TensorFMA and the TensorLoad probes also ran on all 1,024 at once. Inputs were nonzero integers from −3 to 3, plus the zeros under test, so every result could be checked exactly on the host. All of them were exact, and so were all of them again in three passes on each of three cards in the three-card check, whose TensorFMA cycle counts matched this table's.
| Hook | Programmer's Reference Manual | RTL | Measured on the card |
|---|---|---|---|
| TensorFMA zero-skip | fp32: the multiply-add is skipped when an A or B element is zero (PRM 9.4 pseudo-code). | A zero 32-bit word gates that lane's multiplier. The sequencer still issues one 8-lane micro-op per cycle. The vector unit's specification (the Minion VPU Specification) says the gating is there "to save power". | No cycles saved. 546 cycles per 16×16×16 op at 0–100% zeros, whether zeros are scattered, whole columns or whole rows, in A or in B. Power falls in step with the zeros (chart below). |
| fp16 zero-skip | The pseudo-code updates only when all four of a₁, a₂, b₁, b₂ are nonzero. Taken literally, a valid product would be dropped whenever its partner element is zero. | Skips only when a whole 32-bit pair is zero, as the simulator does. | Correct arithmetic. A pair with one zero still adds its other product. The manual's rule is a documentation bug. 546 cycles at any sparsity. |
| int8 (TensorIMA8A32) | No skip in the pseudo-code. | Gates a lane when its 4-byte A group is zero, or when both of its B words (columns j and j+8) are zero. | 318 cycles per 16×16×64 op at any sparsity. |
tensor_mask on TensorFMA | Masked rows of C "generate no calculations". | Masked rows issue no multiply-adds, but their micro-op slots are still spent. | No cycles saved: 544 cycles with 8, 4, 1 or 0 rows enabled (546 with all 16). Masking 8 of 16 rows cuts power like zeros do: +9.2 instead of +17.5 W on aifoundry2 and +7.8 instead of +15.5 W on this card, each as much as half zeros within its 99% range; on aifoundry1 card 1, +8.8 instead of +17.3 W, the match could not be confirmed. |
tensor_mask on TensorLoad | A masked row generates no memory access. | No request is sent for a masked row. | Time scales with the rows loaded: 160 → 80 → 42–43 cycles for 16, 8 and 4 of 16 rows from L2 or the scratchpad, and 749–751 → 320 → 144 from DRAM on one minion, on all three cards (160 → 81 → 45–47 and 730 → 320 → 144 in the first run). Not when the whole chip streams from DRAM (next section). |
One minion, A and B held in the L1 scratchpad. With B streamed from L2 through TenB, as in the matmul report, an fp32 op takes 529 cycles, also flat (that report's int8 loop, which streams both tiles, takes 280 on all three cards, against 318 here; with A held here and only B through TenB an int8 op takes 270). All 1,024 minions at once: 546.1. The same counts on all three cards.
The three-card check: each point the mean of three runs of 5 s, each launched at a set die temperature (about 57 °C on this card, 80 °C on the other two) and measured above the idle just before it; each card has its own least-squares line. The points are the registered values, which carry each card's launch-temperature offset (the note: from 0.5 W high to 0.3 W low on aifoundry2, 0.8–1.3 W low on aifoundry3, from 0.7 W high to 0.05 W low on aifoundry1 card 1). Every point runs 1.10–1.12 × 10⁹ ops/s, launch gaps included. Operands are nonzero integers from −3 to 3 apart from the zeros. The row-mask point has a dense A with 8 of its 16 rows masked off. The same loop on other values is in the chart below.
So the gating works, and it is worth a lot of energy. On these operands a dense op costs 3.4–3.8 pJ per multiply slot above idle on the three cards (5.4–6.1 pJ on random normal data, below; registered values, see the note). An all-zero op still costs 0.2–0.5 pJ per slot, and part of that is the harts simply running: an integer loop on the same harts draws 1.6 of 2.5 W on aifoundry2, 0.7 of 1.4 W on aifoundry1 card 1 and 0.3 of 1.0 W on this card. These are V3-ABL-B's own spin runs, three blocks on each card; V3-ABL-A's integer loop, which the other pages quote, drew 1.5, 0.7 and 0.4 W; the launch-temperature offset is as large as these, so which card's harts draw more is not settled. With all zeros the tensor unit itself adds about 0.7 W on this card and on aifoundry1 card 1 (99% ranges 0.1–1.3 and 0.04–1.3 W), against 15–17 W dense; on aifoundry2 its 0.9 W is not separated from the spin baseline (99% range −0.9 to 2.6 W). In this loop 87–94% of the power above idle follows the nonzero count (87% on aifoundry2, 91% on aifoundry1 card 1, 94% on this card), 94–96% of the tensor unit's own 15–17 W. But no zero is ever faster than a nonzero.
The values matter as much as the zeros. The three-card check also ran this loop (all 1,024 minions, A and B in the L1 scratchpad, 546 cycles per op) with every operand zero, every operand one and random normal data, four runs of 7 s per card launched as above (V3-ABL-A). Above idle it drew +1.9, +10.4 and +27.2 W on aifoundry2, +0.8, +8.8 and +24.7 W on this card and +1.2, +10.1 and +27.8 W on aifoundry1 card 1: 0.40, 2.26 and 5.92 pJ per multiply-add on aifoundry2, 0.16, 1.93 and 5.37 on this card and 0.25, 2.21 and 6.06 on aifoundry1 card 1, the per-card values of the energy manual, §3.2. So zero-skip saves 82–92% of the loop's power above idle against all ones and 93–97% against random normal data (78–95% and 92–98% with busy and idle power compared at the same die temperature, over note C2's two readings; these registered values read from 0.5 W high to 0.3 W low on aifoundry2, 0.9–1.5 W low on this card and from 0.7 W high to 0.05 W low on aifoundry1 card 1). The −3…3 operands of the chart above sit between the two. The Horace experiment explains why by simulating the open RTL of the multiply-add unit: power follows the register bits clocked and the nets toggled, not the count of nonzero operands, and its energies per event, fitted on aifoundry2, carry over to the other two cards up to one scale factor per card (its §10).
Each stem is one card's mean power above the idle just before its runs, as registered (the note); the column on the right repeats it with the energy per multiply slot, all 4,096 of an op's slots counted, zero or not. The rows with A and B alike are V3-ABL-A's runs (four of 7 s per card); the two rows of this page's loop, whose nonzero operands are −3…3, are the end points of the chart above (V3-ABL-B, three runs of 5 s). Hover or focus a stem for its runs' range, its launch temperature and its idle.
A kernel that wants time back from sparsity has to issue fewer or smaller TensorFMAs: skip all-zero tiles, or shrink the op's row and column counts (its AROWS and ACOLS fields, each the count minus one). The tensor unit will not do it by itself. (On A0 silicon, erratum 1.29 type D rules out some small fp32 and fp16 ops whose ACOLS field is nonzero: 4-row ops with B in the L1 scratchpad, and 1- to 4-row ops with B streamed through TenB.)
Masked loads save time, except when the chip saturates DRAM
A TensorLoad moves up to 16 lines of 64 B into the L1 scratchpad. With tensor_mask set, a masked line should never be requested. On one minion the time falls with the lines requested. From L2 and the shire's scratchpad it falls until a floor, the load's latency: 42 and 43 cycles on all three cards in the three-card check (45 and 47 in the first run). From DRAM it falls almost in proportion, from 749–751 cycles for 16 lines to 73 for one on all three cards (730 in the first run): a lone minion's requests go out at about 47 cycles per line, far below what DRAM can deliver. In the first of the check's three passes the 16-line DRAM load took 1,132 cycles on every card, while 15 lines and fewer took the same time in every pass; the cause was not found. (The DRAM buffer is 64 MB, larger than the 32 MB L3, so every load misses the caches.)
With all 1,024 minions loading at once, on-chip memory behaves the same way. DRAM does not. At 16 lines per load this short probe (2,000 loads per minion) streams 72 GB/s from DRAM counting the slowest minion (71–73 GB/s on all three cards). The same probe with 20,000 loads per minion reads 75.6–75.9 GB/s on all three cards, the 76 GB/s measured in the memory-hierarchy report. Asking for fewer lines hardly shortens a load, so the useful bandwidth falls with the mask instead of the time. Each minion streams 1 KB blocks, and every mask tested keeps the same low rows of every block, so all the kept lines are homed in a few shires' L3 slices (see Caveats). Masks that vary from block to block were not tested. What matters for kernels is that weights to be skipped should sit in SRAM.
Why: the home-shire hypothesis, charted, and the full bandwidth table
This chart, the table below and the TensorLoad chart above show: first run, 18 Sep, aifoundry3 (the Data buttons above the TensorLoad chart choose).
Shire positions here are illustrative, not the physical layout; only which count is lit is measured, from the address mapping in Caveats. GB/s per shire divides the chip's measured DRAM bandwidth at that mask (from the table below) by the shires it should reach. If the hypothesis holds, (a)'s per-shire rate stays flat across masks while (b)'s does not, since the 8-line mask still reaches all 8 memory shires.
| All 1,024 minions | 16 lines | 8 lines | 4 lines | 1 line |
|---|
Cycles per load, mean over the 1,024 minions. In brackets, the bytes the mask let through per second over the whole chip, using the slowest minion's total time at 600 MHz (so they are a few percent below what the mean cycle count implies). The three cards give the median of their three passes: they repeat the first run's L2 and scratchpad counts to within 1.5% (the L2 at 16 lines varies from 286 to 306 cycles between passes), and read DRAM 3–4% faster at 16 lines, within 2% at 8 and 4, and 4–5% slower at 1 line. At 16 lines L2 reads 2.03 TB/s here (1.96–2.13 TB/s over the check's nine passes), below the 2.45 TB/s the memory-hierarchy probe reads on both cards, including this one on 23 September at a pinned 600 MHz (Ridge points). This probe is shorter (2,000 loads per minion). Run with 20,000 or 200,000 loads per minion it reads 2.45 and 2.46 TB/s on all three cards, so the gap belongs to the short probe; the three-card check could not tie it to a cold start, since a separate 2,000-load run varied from 1.63 to 2.37 TB/s between passes.
A batch-1 layer that skips zeros
Method: how the layer kernel works and was timed
This is the case most relevant to ML: one output of a ReLU layer, y = W x with a batch of one, where most of x is zero. The layer is 1024 outputs × 4096 inputs in fp32 (16.8 MB of weights). It is spread over all 32 shires, and each shire keeps its 512 KB share of W in its own L2 scratchpad. Minion j of shire s owns 16 outputs and 256 inputs, in 16 slices of 16. For each slice it loads only the rows of Wᵀ whose x is nonzero (a masked TensorLoad), then runs one single-row TensorFMA. Where all 16 x values of a slice are zero, it skips the slice. Sixteen minions then add their partial rows with a TensorReduce tree that stays inside the shire.
Latency runs from a shire barrier to the moment the block's result is in registers, averaged over 2,000 repeats per sparsity level. Each level uses one random x, the same for every variant. Every output was checked exactly against the host. The host supplies a bitmask of which x are nonzero. In a real pipeline the layer that produced x would build it, which takes two vector instructions per 8 values (fltm.ps, then read the mask).
The slider and buttons choose a zero fraction and a kernel; press Enter on a point of the curve to choose it. (a) draws a representative minion with the measured means: how many rows of Wᵀ it loaded and how many of its 16 slices it computed; which rows and which slices are illustrative. The same random x served every kernel at each level. The A100 marks are references, not bounds (see below).
All layer timings (µs) and energies
| Zeros in x | 0% | 50% | 75% | 90% | 95% | 99% | 100% |
|---|
µs at 600 MHz: the time of each 16-output block's slowest minion, mean over the 64 blocks. The slowest of the 64 blocks averaged 7.62 µs dense and 2.22 µs at 99% zeros, and a layer is not finished until its slowest block is. Results stay in each block's root registers; writing y out is not included. "Compute only" leaves out the reduction: every minion writes its partial row and the host adds them.
The energy rows come from separate runs of the masked + skip compute-only kernel with the power logger on, which drew their own x. In the three-card check each of the three runs per card drew a different x (seeds 1 to 3), so at 99% zeros its layer took 2.7 µs in one run and 1.9 µs in the other two; the rows give the mean of the three, and each card's board energy includes its own idle (36.4 W on aifoundry2, 25.8 W on aifoundry3, 50.2 W on aifoundry1 card 1). Each card's above-idle rows carry its launch-temperature offset times the layer time (the note): up to about 5, 4 and 2 µJ at 0, 90 and 99% zeros on aifoundry3 (read low), and less on the other two cards. Layer times come from the layer rate, which also includes the shire barrier between layers.
Three effects add up. Masked loads cut the bytes, so at 50% zeros the layer drops from 7.5 to 5.7 µs. After that the per-slice fixed costs take over. Each nonempty slice waits for a load (at least 42 cycles) and an FMA, and the curve flattens at 5.4 µs. These times repeat to within 0.01 µs on all three cards. The big gains come only from skipping whole slices, which random sparsity allows only near the top: for independent zeros 18.5% of slices are empty at 90% zeros and 85% at 99% (0.9¹⁶ and 0.99¹⁶; the x drawn here had 14.9% and 87.5%). Energy per layer above idle falls from 67 to 5.9 µJ on aifoundry2 (19 at 90% zeros), from 75 to 4.5 µJ on aifoundry1 card 1 (16) and from 56 to 3.3 µJ on this card (12), mostly because skipped slices and masked loads do less work in less time. The tensor unit's own gating adds no speed and little energy here: with every row loaded, 90% zeros in x cut the layer's power above idle only from 9.9 to 8.2 W on aifoundry2, 11.0 to 9.3 W on aifoundry1 card 1 and 8.4 to 6.8 W on this card, 16–19% (the same every-row kernel at both points).
What the whole board spends per layer is another matter: once most slices are skipped, the power the card draws doing nothing, over the layer's few microseconds, is nearly all of it.
How the A100 comparison numbers were built up
Against the A100. No one has published a timing for this exact layer on an A100, so the reference lines are rough references, not bounds. Launching even an empty kernel costs 2.26 µs, plus 1.44 µs for the kernel itself (measured on an A100-SXM4-80GB). Kernels queued back to back in a stream cost less, about 1.5–2.3 µs each on an A100 (on-chip communication report). cuBLAS half-precision GEMM has a latency floor of about 4.5 µs. Streaming 16.8 MB of weights from L2 at 2.5–4.5 TB/s takes 3.7–6.7 µs. Added to the 2.26 µs launch, or to cuBLAS's 4.5 µs floor, a dense fp32 layer probably takes 6–11 µs, against 7.3–7.5 µs here. CUDA graphs cut launch cost to about 1.3 µs per kernel (measured on H100). At batch 1 a GPU's sparse GEMV is floor-bound too: TEAL's kernel gets 3× at 99–100% sparsity on a 4096×14336 fp16 layer, with 14× this one's weights and 7× its bytes (102 → 33 µs), and Deja Vu reported 1.8–2× end to end at about 75% sparsity (OPT-175B on 8 A100s, against FasterTransformer).
So at batch 1, on these runs, ET-SoC-1 reaches A100-class latency on one layer. It beats a single isolated A100 launch (3.7 µs) only at 99% zeros or more, and there (2.1 µs) it only matches a queued A100 kernel (1.5–2.3 µs). That comparison favours ET: its 2.1 µs is a layer inside a kernel that is already running (a launch from the host took 0.2–0.6 ms on the three cards, median 0.36 ms on this one), and a GPU can avoid launches too, with a persistent kernel whose grid-wide sync costs 1.1–1.2 µs (on-chip communication report).
Board energy, and what a stack of layers would still need
The whole board spends 60–246 µJ per layer on this card, idle included (85–337 on aifoundry2 and 114–446 on aifoundry1 card 1, whose idles are higher). No A100 energy for this layer has been published; the benchmark in scenario 2 would measure it. The larger prize is a stack of such layers. A pipeline of tens of small layers could stay on chip with no launch between layers, but each layer would still have to pass its outputs across shires (a chip-wide allreduce, §5), and a GPU could run the same stack as one persistent kernel. None of this is measured here.
Divergence: narrow lanes help, but not enough
Method: the workload and the three lane-scheduling variants
The other half of the question is irregular work, here divergence. Each work item needs k iterations of an FMA loop (8 independent chains), and k is drawn from a heavy-tailed Pareto distribution with mean 64. The case for ET-SoC-1 is that its SIMD is only 8 lanes wide and each core runs its own code, so a long item wastes less. All 2,048 harts took 32-item chunks from a shared work queue (a global atomic counter), 524,288 items in all, and every result was checked. Three variants ran, once each, and again in three passes on each of three cards in the three-card check, which repeated every figure in this section. Static runs 8 items per 8 lanes until the longest finishes, as a GPU warp would. Refill gives a lane its next item as soon as it finishes. Scalar runs one item at a time on one lane.
Static 8-lane groups keep busy exactly the share of lanes that grouping the same items in eights predicts: 0.55 at α = 3 and 0.21 at α = 1.2, against 0.35 and 0.08 for a 32-lane warp on the same items. (The continuous order-statistics formula gives 0.55 and 0.19; the cap at 16,384 iterations trims the heaviest tail.) Refilling lanes raises ET's efficiency to 0.96 and 0.46 (0.962–0.963 and 0.453–0.459 over the three cards' passes; which lane gets which item depends on timing, so the refill figures move in the third decimal), but its bookkeeping costs cycles: at α = 3 it does fewer useful FMAs per second than static (0.59 against 0.73 T/s, on every card), and it pulls ahead only from α = 2 (by 0.016 T/s there, a 99% range within 0.012–0.020 T/s on each card over five runs).
It is not enough to beat a modelled A100. This loop peaks at about 1.5 useful FMAs per cycle per minion, or 0.9 T/s for the chip, under a fifth of the 8-lane peak (4.9 T/s at 600 MHz). An A100 has 9.75 T FP32 FMA/s. A naive kernel that keeps only 8% of its lanes busy at α = 1.2 would still do 0.74 T/s if it issued FMAs at the A100's peak (a model, not a measurement), against ET's 0.39 T/s measured on each of three cards. That 1.9× gap is within what tuning might recover, but at α = 1.5–3 the modelled A100 stays 2.3–4.6× ahead of ET's faster variant, and GPUs have compaction and refill tricks of their own.
Investigating the work-queue hypothesis, charted
What limits the loop was not measured, and the shared work queue may be part of it. A contended global atomic retires one operation every 10 cycles (One hot line stops a shire). The 16,384 chunk grabs plus each hart's final empty grab need about 184,000 cycles of it, more than the average hart's whole run without divergence; the chart below sets that floor against every run. An 8-lane iteration also takes 84 cycles against 33 for a one-lane scalar iteration, so the loop's mask test and branch cost too. The scalar variant, which uses the queue at a third of the rate, held 0.30 T/s at every α. Larger chunks or one queue per shire would separate the two.
Bars: the mean hart's cycles in each run; the slowest hart is in each bar's tooltip. Dashed line: the cycles the queue alone needs if every grab waits its turn, (items ÷ chunk + harts) × cycles per contended atomic, here (524,288 ÷ 32 + 2,048) × the slider's value. 10 is the ceiling measured in One hot line stops a shire; the slider shows how the floor moves if the true cost differs. Only 32-item chunks were run.
On this evidence (three cards, against a modelled A100), divergence in how long items run is not an ET-SoC-1 advantage. What the chip has that a GPU lacks is one instruction stream per core. That matters when items run different code, as in interpreters, simulators and heterogeneous state machines, which this sweep does not test.
Where ET-SoC-1 should beat an A100, and how to prove it
These measurements and the published literature give one recipe for a win. The work must be latency-bound or event-driven, so that a GPU sits at its per-kernel or per-sync floor. Its state must fit in the 80 MB of on-chip scratchpad, or be regenerated rather than fetched. And each step's surviving work must be dense enough for 8-lane SIMD or the tensor unit.
The sync-cost numbers behind the estimates below
The chip's sync costs come from the on-chip communication report, re-measured on three cards on 26 September (aifoundry2, aifoundry3 and aifoundry1 card 1, three passes each, within 0.3% of each other). A 1,024-minion allreduce takes 2.3 µs for 32 B and 5.4 µs for 1 KB, a chip-wide barrier from atomics and credits 8.3 µs, and a shire allreduce 0.74 µs (444 cycles). A message round trip takes 113–190 ns inside a shire and 250 ns + 20 ns per mesh hop across shires.
Each scenario below gives why a GPU struggles, the best published baseline (an A100 result where one exists; otherwise the hardware is named), an estimate for this chip from the measurements above, and the benchmark that would settle it. Only scenario 2 is measured here.
Scenario 1: spiking cortical microcircuit (unmeasured estimate)
1. Spiking cortical microcircuit
The Potjans–Diesmann model (PD14): 77,169 neurons, 3×10⁸ synapses, 0.1 ms steps. RTF is the real-time factor, wall-clock time over model time.
- Why a GPU struggles
- About 25 spikes per step. Each step is several tiny kernels, and launch latency alone caps GeNN at RTF ≈ 0.1.
- Published baseline
- A100 (Golosio et al. 2023): GeNN RTF 0.539 (54 µs of wall time per step); NEST GPU 0.73–0.85. No A100 energy per synaptic event has been published. RTX 4090: GeNN 0.272; about 74 nJ per synaptic event, estimated by Senk et al. Fastest full-scale run: neuroAIx (35 FPGAs) 0.049 at 48 nJ, the lowest energy measured at the outlet.
- ET-SoC-1 estimate
- About 75 neurons per minion, resident in L1/L2, with synapses regenerated from a counter-based RNG. Per step: an exchange of the ~25 spike IDs through the hardware tree (2.3–5.4 µs for 32 B to 1 KB, above; an all-gather of a variable-length list needs fixed slots or a prefix sum), plus the compute: about 95 synaptic events and 75 neuron updates per minion per step, each event's target, weight and delay regenerated from the RNG, with no hardware divide or square root. That compute was not measured. If it fits in about a microsecond, a step takes about 4–10 µs, an RTF of 0.04–0.1, 5–13× the A100. At about 30 W this would be a few nJ per synaptic event: an estimate, not a measurement.
- Benchmark that would settle it
- PD14 under the community recipe (Senk et al. 2026): RTF, and energy per synaptic event measured at the outlet, over ≥ 10 s of model time and 10 seeds. Validate with rate, CV and correlation distributions, and declare procedural connectivity.
2. Batch-1 activation-sparse inference, SRAM-resident
- Why a GPU struggles
- Batch 1 is floor-bound: launch plus sync per layer. Gains from sparsity saturate on fixed costs.
- Published baseline
- A100: one isolated launch plus a null kernel 3.7 µs; cuBLAS fp16 GEMM floor ~4.5 µs. TEAL 1.33–1.45× end to end at 50% sparsity; Deja Vu 1.8–2× at ~75% (8×A100).
- ET-SoC-1 estimate
- Measured in §3, dense against 99% zeros: 7.5 → 2.1 µs per layer with the on-chip reduction, and 246 → 60 µJ of board energy per layer on aifoundry3 (337 → 85 on aifoundry2, 446 → 114 on aifoundry1 card 1) from the energy runs, which used the compute-only kernel whose partial sums the host adds. A model of ≤ 40 M fp16 parameters fits on chip, and every layer after the first avoids a kernel launch, though each still has to pass its outputs across shires (not measured).
- Benchmark that would settle it
- The same layer on an A100 with CUDA graphs, cuBLAS against a gather GEMV, 0–99% sparsity, time and NVML power. Then a whole small ReLU model at batch 1: µs and J per token.
Scenario 3: sparse triangular solve (unmeasured estimate)
3. Sparse triangular solve (ILU/IC preconditioner, deep DAG)
- Why a GPU struggles
- Thousands of thin dependency levels, each needing a sync.
- Published baseline
- A100 (FP64): cuSPARSE and a sync-free solver mostly ≤ 2 GFLOPS on DCSolver's 38 ILU(0)/IC(0) factors (7 to 916,791 levels). AG-SpTRSV is 3.99× faster than cuSPARSE (geometric mean over 2,219 matrices, FP64). tmt_sym takes about 0.57 µs per level with cuSPARSE on a Titan RTX (FP64, derived from its published 0.014 GFLOPS).
- ET-SoC-1 estimate
- Inside a shire, a level handoff by credit or message takes 113–400 ns. A chip-wide barrier per level (2.3–8.3 µs) would lose, so it only wins with per-shire subdomains and point-to-point credits. FP32 only: a mixed-precision preconditioner.
- Benchmark that would settle it
- A thin-level SuiteSparse subset:
cusparseSpSVand AG-SpTRSV on the A100 against an ET version partitioned into per-shire blocks. Report solve time and time per level.
Scenario 4: RTL and gate-level simulation (unmeasured estimate)
4. RTL and gate-level simulation
- Why a GPU struggles
- Tiny tasks, and a sync every simulated cycle. The GPU needs thousands of stimuli to fill up.
- Published baseline
- GEM on an A100: 7.3–65 kHz for one stimulus. Verilator: 1–1000 kHz. Manticore (225-core FPGA, bulk-synchronous) is 2.1–4.2× over Verilator (geometric means over three x86 hosts, serial and multithreaded). Parendi, from Manticore's first and last authors, runs RTL simulation on 5,888 Graphcore IPU cores and reports up to 4× over the fastest x64 multicore hosts.
- ET-SoC-1 estimate
- Same bulk-synchronous model at 4.5× Manticore's core count. Synchronisation alone caps a shire at about 2.5 MHz (a 395 ns barrier every cycle) and the whole chip at about 430 kHz (the 2.3 µs tree). Gate evaluation and exchanging signal values between shires come on top, so real rates would be well below these ceilings. Several million gates fit in scratchpad.
- Benchmark that would settle it
- Manticore's designs: simulated kHz against Verilator at 1 and N threads, and GEM on the A100 with 1, 8 and 1K stimuli. Needs a partitioning compiler.
Scenario 5: interpreters over per-item programs (unmeasured estimate)
5. Interpreters over per-item programs (symbolic regression, GP)
GP is genetic programming. GPop/s counts GP operations per second, one expression node evaluated on one data row; it is not 10⁹ of anything.
- Why a GPU struggles
- Every item runs different code, so SIMT serialises the branches.
- Published baseline
- EvoGP 10¹⁰–10¹¹ GPop/s (RTX 3090/4090); Operon 6.8 × 10⁹ on one CPU thread, 94 × 10⁹ on 24 threads.
- ET-SoC-1 estimate
- One expression per minion, one data row per lane (8 rows per vector instruction), so the lanes never diverge. Guessing from this issue rate: about 10¹¹ GPop/s; 10¹² would need the interpreter to match the plain FMA loop's 0.9 × 10¹² lane-ops/s with no overhead. The sweep above says to measure it.
- Benchmark that would settle it
- SRBench-style evaluation: GPop/s against EvoGP on the A100 and Operon on x86.
Scenario 6: tree ensembles and motion planning (unmeasured estimate)
6. Tree ensembles and motion planning at batch 1
- Why a GPU struggles
- Small batches don't fill a GPU; per-query paths are data-dependent.
- Published baseline
- No published A100 batch-1 forest latency. Lettich et al. (IEEE TPDS 2019) parallelised QuickScorer (Lucchese et al., SIGIR 2015), a fast traversal of tree ensembles, with CPU SIMD, multi-core CPUs and GPUs, and got their best results on the GPU. Motion planning (Panda set, RTX 4090): pRRTC 0.48 ms median, cuRobo 30–62 ms (median–mean); VAMP 40 µs median on one x86 core.
- ET-SoC-1 estimate
- Throughput plays (1,024 independent queries), with x86 as the rival more than the A100. A single query on one minion will probably be slower than on one x86 core.
- Benchmark that would settle it
- MSLR-WEB30K with a LightGBM model: µs/doc at batch 1–100 against RAPIDS FIL. MotionBenchMaker: plans/s and tail latency against VAMP and cuRobo.
Where the A100 wins
- Sparse data that streams from DRAM: SpMV, embedding tables, MoE experts. The A100 does FP64 CSR SpMV at 50–245 GFLOP/s, reaching 1.2–1.5 TB/s of effective bandwidth on the best large matrices, while this card streams 76 GB/s. Masks that keep the same rows of every block did not help once the whole chip streamed from DRAM, and even a perfect mask cannot lift 76 GB/s.
- Anything that batches: big GEMMs, batched inference, batched sequence search. 2:4 structured sparsity also favours the A100: it doubles tensor-core math (1.5–1.9× measured on large GEMMs), but it needs pre-pruned static weights and does nothing for dynamic activation zeros at batch 1.
- Divergence in run length alone (§4): narrow lanes help, but not enough.
Answers to the questions this work started from
The work began from two research briefs (not published). The sections above answer most of their questions; these two answers are easy to miss:
- Is the manual's fp16 skip rule right? No; the silicon computes the correct sum (§1).
- Does point-to-point messaging work across shires? Yes on aifoundry2, with one ready flag per minion: change partners only after a barrier (on-chip communication, A trap: one ready flag per minion).
How it was measured
Method for every probe: kernel, timing, and each probe's design
- Kernel and host.
workloads/sparsityis a standalone kernel plus host built on the runtime API, as in the other workloads. It was built on the lab machine against its/opt/et, and every probe was validated first insys_emu(the platform's functional simulator) with its vector-unit hazard checker (-vpurf_warn). The checker found two register-file hazards (errata 1.29 types B and C), which the kernel now avoids. - Timing is the minion cycle counter (
hpmcounter3) on the timed section of each hart. On aifoundry3 the service processor (the on-die management core) reported 600 MHz in every sample of this session, logged about 4 times a second during the sweeps, so cycles convert to µs at 1.667 ns. - TensorFMA probe. One A tile and one B tile are loaded into the L1 scratchpad once, then 20,000 ops (4,000 on all 1,024 minions) run back to back, accumulating in f0–f31. Operands are nonzero integers from −3 to 3. Zeros are placed by the host: scattered, whole columns, whole rows, or one element of each fp16 pair. Results are compared exactly while fp32 sums stay below 2²⁴, and modulo 2³² for int8.
- TensorLoad probe. 20,000 16-line loads per minion on one minion (2,000 on all 1,024), two in flight, over private 1 KB blocks. For DRAM the buffer is 64 MB for a lone minion and 256 KB per minion for the whole chip, more than the 32 MB L3 either way. For L2 it is 8 KB per minion, and for the scratchpad 64 KB of the shire's. GB/s are the bytes the mask let through over the slowest minion's cycle count.
- Layer probe. W is staged into the scratchpads with TensorLoadL2Scp. x is 4,096 small integers, a random fraction of them zero, with a precomputed bitmask. A block's time is its slowest minion's. Each layer starts with a fast-local-barrier plus credit barrier in every shire. The barrier polls its credit counter with a timeout, so a missing minion fails the run rather than hanging the shared card.
- Divergence probe. k = ceil(Pareto(α)) with mean 64, capped at 16,384. Chunks of 32 items come off a global atomic counter. Lane state lives in f1–f9, the running lanes are the mask m0, and every item's result must equal its k. Lane efficiency is useful lane-iterations over 8 × vector iterations. Throughput is taken over the mean hart time, which is the steady state; the slowest hart is set by the single longest item. The A100 lines apply the warp-of-32 order statistics to the same draws, times 9.75 T FMA/s.
- Energy. Every power figure on this page is the three-card check's (26 September), from two of its experiments on each card: V3-ABL-B (
tools/claims-v3/ablb/) ran this page's configurations for 5 s each in three shuffled blocks, and V3-ABL-A (tools/claims-v3/abla/) the same loop on all zeros, all ones and random normal data, four runs of 7 s. Each run launched at a set die temperature (about 57 °C on this card, 80 °C on the other two) and was measured above the idle just before it, corrected for leakage to that temperature, with a 10 Hz sampler reading board power, clock and voltage from the service processor; every sample of a kept run was at 600 MHz. The board-power reading itself changes less often, and how often depends on the card and the poller: in the three-card check, on this card about every 223–224 ms under a poll of about 8 per second and 262–264 ms under a 10 Hz sampler, on aifoundry2 about every 126–135 and 155–157 ms, and on aifoundry1 card 1 133–139 and 156–159 ms. First measured on 18 September on this card alone, with no temperature control or correction: a thread polled the service processor about 9 times a second while each configuration ran 4 s of launches in its owntimeout 10process, with 6 s of idle at the start and end, twice (the second time in reverse order; the two agree within 0.5 W above idle, the integer loop within 0.6 W), against a 24.6–24.7 W idle. It read +17.3 W dense and +2.4 W with A all zeros (86% saved; 3.8 pJ per multiply slot dense), +9.2 W with 8 rows masked, +1.7 W for the integer loop, and 70, 19 and 7.1 µJ per layer above idle at 0, 90 and 99% zeros; its 0% layer ran the masked kernel, which with no zeros also loads every row, so its gating alone read 9.4 to 8.3 W, about 12%.
Version history
Versions: published 18 September 2026; revised 24 September 2026 (operand values, later measurements, the masked-DRAM cause, the work queue, the A100 launch comparison) and 25 September 2026 (the layer explorer, the later runs on the power chart and the queue-floor chart; this card's meter rate; shorter answers); 25 September (version 3): which card and how many runs each figure rests on, aifoundry2's values beside this card's, and the tensor unit's own all-zero power no longer quoted as a number (record: docs/reports/data/2026-09-25-claims-v3); 26 September: the three-card check (aifoundry2, aifoundry3, aifoundry1 card 1) in the text, the KPIs, the power chart, the TensorLoad chart and table and the layer table; 27 September: the masked-load hypothesis charted (per card), scenario 2 names the kernel behind each figure and takes the three-card energies, and aifoundry3's 600 MHz traced to the boot service that sets its TDP to 0 W; later that day, the layer's energy per card charted, idle against what the layer adds; 28 September: a repeat cut; the note's counts given by both rules; later that day, every power figure the three-card check's: the operand values (all zeros, all ones, random normal) from its V3-ABL-A runs, in the text and a new chart, in place of the 22 September session on this card and the "Later measurements" box; the power chart, the layer table and the explorer without the first run, which is one note in the method; aifoundry1 card 1's saving at the same die temperature 92–96% on V3-ABL-B's registered metric (it said 91–96%, what the runs' dropout-rule metric gives).
Caveats
Caveats in full
- One card, one day, then three.
aifoundry3held 600 MHz throughout; later work found why: a boot service sets its TDP to 0 W at every boot, so its governor can never raise the clock (The DVFS loop and its leakage, §3). Its sister cardaifoundry2runs at 600, 700 or 800 MHz depending on die temperature and board power. Cycle counts carry over between the cards: the three-card check read the same tensor-op counts (546 cycles per fp32 or fp16 op, 318 per int8 op and 529 per fp32 op with B streamed through TenB), the same masked and full-line TensorLoads, the same layer times and the same work-queue results on aifoundry2, aifoundry3 and aifoundry1 card 1 (though a lone minion's load floor and its 16-line DRAM load read 42–43 and 749–751 cycles on all three, against 45–47 and 730 in this page's first run). Switching power on aifoundry3 ran 10% below aifoundry2's for the same work in the three-card check as registered, with its die 23 °C cooler, and 3–4% below at the same die temperature (the Horace experiment, §10; post-data note C2), and idle power differs with die temperature. - The flat DRAM time under a mask is probably caused by which lines the masks kept. A line's L3 home shire is physical-address bits 10:6 (Anatomy of a memory access), and these masks keep the same low lines of every 1 KB block, so at 8, 4 and 1 lines all 1,024 minions' misses land on 16, 8 and 2 of the 32 home shires (and at 4 and 1 lines on 4 and 1 of the 8 memory shires). Per home shire the chip streams 2.3–2.6 GB/s at every mask in this one run. The 8-line mask still reaches all 8 memory shires yet halves the bandwidth, which points at the home shires rather than the memory controllers. A mask that rotates from block to block would test this.
- The layer's bitmask of nonzero x comes from the host. The tree reduction assumes that all 16 minions of a block take part in every layer.
- The divergence loop is not hand-optimised. The performance-counter events that would count its micro-ops (such as TXFMA ops) can only be programmed in machine mode (M-mode), where the firmware runs, not from a user kernel, so micro-op counts were not read.
Reproduce
Reproduce this
scripts/deploy-lab.sh aifoundry3 workloads/sparsity
ssh aifoundry3 'cd ~/nekko && bash workloads/sparsity/run_lab.sh build/sparsity/host/sparsity_host build/sparsity-data'
ssh aifoundry3 'cd ~/nekko && python3 workloads/sparsity/run_energy.py --host-bin build/sparsity/host/sparsity_host --out build/sparsity-data/energy-a'
ssh aifoundry3 'cd ~/nekko && python3 workloads/sparsity/run_energy.py --host-bin build/sparsity/host/sparsity_host --out build/sparsity-data/energy-b --only gemv-skip-99,gemv-skip-90,gemv-skip-0,gemv-dense-90,fma-rowmask,fma-col50,fma-zero,fma-875,fma-50,fma-dense,spin'
scp -r aifoundry3:nekko/build/sparsity-data/. docs/reports/data/DATE-sparsity-aifoundry3/
python3 workloads/sparsity/analyze.py docs/reports/data/2026-09-18-sparsity-aifoundry3 \
--embed docs/reports/2026-09-18-et-soc1-sparsity.html
scripts/deploy-lab.sh copies the sources to ~/nekko on the lab machine and builds them there. Every card command runs in its own timeout 10 process with --budget 8, and waits until no other process holds the card. Raw data: docs/reports/data/2026-09-18-sparsity-aifoundry3/ (its README lists what each file is). analyze.py also reads the three-card check (docs/reports/data/2026-09-25-claims-v3: V3-LAT's raw TensorLoad sweeps and the per-run tables of V3-ABL-B and V3-ABL-A, results/ablb.runs.json and results/abla.runs.json) into the three-card views of the charts and tables, and prints the values the text quotes; the check's runs are tools/claims-v3/lat/, tools/claims-v3/ablb/ and tools/claims-v3/abla/. Without a card, add --sysemu to any sparsity_host command to check a kernel (the simulator's cycle counts mean nothing), e.g. build/sparsity/host/sparsity_host --sysemu --test fma --type fp16 --pattern pair --sweep 0,0.5,1 --shires 0x1 --per-shire 1.
Sources
Full source list
- ET-SoC-1 Programmer's Reference Manual (aifoundry-org/et-man), §9.2.1 tensor_mask, §9.4 TensorFMA32 / TensorFMA16A32 / TensorIMA8A32 / TensorLoad pseudo-code; ET-SoC errata 1.29. core-et (Erbium branch) RTL
vpu_ctrl.v,vpu_tensorfma.vand the Minion VPU Specification, for zero gating and issue. - Vellaisamy et al., Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures, arXiv 2504.11750, Table V (A100 launch 2.26 µs + null kernel 1.44 µs). Hazy Research, "Look Ma, No Bubbles!" (2025), H100 launch 2.1 µs, 1.3 µs with CUDA graphs. Erdil & Schneider-Joseph, arXiv 2411.01137 §5.1.1 (cuBLAS fp16 GEMM floor ~4.5 µs). Luo et al., arXiv 2402.13499, Chips and Cheese, and Zhang et al., GPGPU 2025, Fig. 6a (A100 L2 bandwidth).
- Liu et al., Deja Vu, ICML 2023 (arXiv 2310.17157). Liu et al., TEAL, ICLR 2025 (arXiv 2408.14690), Fig. 3 and Table 3. Mishra et al., arXiv 2104.08378, and Frantar & Alistarh, SparseGPT, ICML 2023 (arXiv 2301.00774), Table 8 (2:4 sparsity). Tsai, Cojean, Anzt, arXiv 2008.08478, and Zhang et al., GPGPU 2025 (A100 SpMV).
- Potjans & Diesmann, Cerebral Cortex 2014. Golosio et al., Appl. Sci. 13:9598 (2023), A100 GeNN and NEST GPU. Senk et al., Neuromorph. Comput. Eng. 6:012001 (2026), benchmark recipe and survey. Kauth et al., Front. Comput. Neurosci. 2023 (neuroAIx). Knight & Nowotny, Nat. Comput. Sci. 2021 (procedural connectivity).
- Hu et al., AG-SpTRSV, ACM TACO 21(4) 2024. Qiu et al., DCSolver, ACM TACO 22(3) 2025. Lu, Niu, Liu, ICPP 2020 (tmt_sym).
- Emami et al., Manticore, arXiv 2301.09413v4 (ASPLOS 2024). Guo et al., GEM, DAC 2025. Emami, Bourgeat, Larus, Parendi: Thousand-Way Parallel RTL Simulation, ASPLOS 2025 (arXiv 2403.04714). Thomason, Kingston, Kavraki, VAMP, ICRA 2024. Huang et al., pRRTC, arXiv 2503.06757. Sundaralingam et al., cuRobo, arXiv 2310.17274. Lucchese et al., QuickScorer, SIGIR 2015. Lettich et al., Parallel Traversal of Large Ensembles of Decision Trees, IEEE TPDS 30(9) 2019. Wu et al., EvoGP, arXiv 2501.17168. Burlacu et al., Operon, GECCO 2020.
- Pareto order statistics: E[max of n] = Γ(n+1) Γ(1−1/α) / Γ(n+1−1/α). Laine, Karras, Aila, HPG 2013, and Aila & Laine, HPG 2009, for GPU wavefront and refill techniques.
Related reports
- The Horace experiment — the follow-up: why the tensor unit's power depends on the operand values, not only on how many are zero, measured on aifoundry2 and repeated on three cards.
- The energy manual, §3.2 — temperature-corrected energy per multiply-add on zeros, ones and random data, on three cards.
- On-chip communication — the sync, allreduce and messaging costs used in the scenarios.
- Anatomy of a memory access — the address map behind the masked-DRAM result.
- Memory hierarchy — 76 GB/s from DRAM, and the sizes and bandwidths of L2, L3 and the scratchpads.
- Matmul efficiency — tensor ops with B streamed through TenB: 9.5 TFLOP/s fp32 on all 1,024 minions.
- Ridge points — where this page's batch-1 layer sits against the chip's bandwidth ceilings.
- Why is it low power? — the tensor unit's energy per FLOP against an A100's.