The ET-SoC-1 memory hierarchy, measured on silicon, next to the A100
Pointer-chase and streaming probes give the card's memory hierarchy as a spec sheet. ET-SoC-1's on-chip memories are fast and cheap per byte: a load hits L1 in 8.8 ns and the shire's L2 in 78 ns, against 23 ns and 148 ns on an A100, at 12–24× less energy per byte (the energy manual's 600 MHz values, three cards). They are also tiny per core: the firmware leaves each hart only 512 B of L1 data cache. Off chip the picture flips. DRAM latency is about 490 ns at the 600 MHz base clock, a fifth more than the A100's HBM (405 ns), and the card streams 76 GB/s, about an eighteenth of an A100's 1.40 TB/s, at roughly 1.1× the energy per byte (1.0–1.2× card by card).
Terms used on this page
The ET-SoC-1's compute cores are minions, small in-order RISC-V cores that each run two hardware threads (harts); eight minions form a neighbourhood, 32 a shire, and 32 shires (1,024 minions) run kernels. Each shire's 4 MB of SRAM is split by these cards' firmware into a 512 KB L2 cache, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB scratchpad, software-managed memory that any shire can address. The shires and eight memory shires, which hold the LPDDR4X controllers, sit on a 2D mesh network-on-chip clocked at 400 MHz, where a mesh hop is one step between neighbouring stops; the minions run at 600 MHz unless the card's power governor raises the clock. More in the hub's glossary.
Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan, three passes of every latency chase on each of aifoundry2, aifoundry3 and aifoundry1's card 1 at 600 MHz, and six passes of the energy probes on each card (the record). Of 20 claims tested here, this page counts 16 held, 2 corrected and 2 differ by card; the hub’s scoreboard, 16 “proven on the cards”, 3 “a test behind it failed” and 1 “differs by card”. Every latency in the spec sheet agreed on each card to within a cycle, and the latency chart shows each card's chases beside the 18 September session. The energy per byte is now the energy manual's three-card values of 26 September, with the scratchpads read full of zeros and full of random data. The same three cards then measured gathers and scatters, random elements level by level (E48, below), each tested item decided card by card.
Load latency against working-set size
How to read this chart
What does a load at each level cost at 600, 700 and 800 MHz, and which part of it speeds up with the minion clock? Only aifoundry2's data can answer this: its governor moved the clock during this page's one session of chases (18 September), while a boot service holds aifoundry3 at 600 MHz (it sets that card's TDP to 0 W at every boot). One hart follows a random pointer chain, so every load waits for the one before it; each step in the curve is a level of the hierarchy. On-chip levels are fixed in cycles, but L3 and DRAM loads also spend fixed time in the 400 MHz mesh and the DDR memory, so the lines are redrawn for the selected clock from two fits (What the numbers say, item 5). Circles are the chases as measured, each at its own clock; filled ones ran at the selected clock (none ran at 700 MHz). From 16 to 64 MB a chain mixes L3 and DRAM hits, and a dashed line joins the measured points. Grey marks are the A100's published pointer-chase plateaus. The card marks are the three-card check of 26 September, every chase at 600 MHz (each card's median of three passes, from shires 0 and 24), shown when 600 MHz is selected; there the chains up to 32 MB stayed at L3 latency on every card.
The latency split by clock
Where the time goes at the selected clock: the part of each load that scales with the minion clock, and the fixed part that does not (L3 as chased from shire 0).
Scratchpad latency across the mesh
How to read these charts
How much does mesh distance cost a remote scratchpad load, and why does the L3 latency depend on the requesting shire? Chases ran from shires 0, 7, 24 and 31 to all 32 scratchpads, over a 64 KB chain so that every load misses L1. The scatter puts each load against the mesh hops between the two shires, with the model fitted to rows 0 and 31 (66 minion cycles + 56 ns + 20 ns per hop) at 600, 700 and 800 MHz; each point takes the colour of the nearest line. The map shades each shire by the load's time in nanoseconds, with the requesting shire outlined. L3 average trip shows instead how far each of the 32 L3 slices is from the requesting shire. The three-card check re-ran the four rows at 600 MHz on each card (26 September): every remote load fell within 0.2 cycles of 99.84 + 12.00 cycles per hop, the model's 600 MHz line, and each pair of shires read the same in both directions to within 0.04 cycles.
The hierarchy as a spec sheet
ET-SoC-1 (measured)
Every ET figure is at the 600 MHz base clock.
| Level | Latency | Bandwidth, whole chip | Energy per byte | Random 4 B element | Capacity | Notes |
|---|---|---|---|---|---|---|
| L1 data cache | 5.25 cyc 8.8 ns | 6.2 TB/s | 0.75 pJ | 452 G/s 12.8 pJ | 512 B per hart (2 sets × 4 ways × 64 B) | The firmware puts the L1 in scratchpad mode before every launch: sets 0–11 become a 3 KB tensor scratchpad. Hart 1 is the same size but pays +3 cycles on misses. Bandwidth is about 5 B per hart-cycle. |
| L2 read buffer | 36 cyc 60 ns | – | – | – | 8 lines per bank × 4 banks = 2 KB | A small clean-line cache in front of each L2 bank; spec: 10 vs 21 shire cycles. |
| L2 cache | 47 cyc 78 ns | 2.45 TB/s | 3.1 pJ | 27.4 G/s 354 pJ | 512 KB per shire · 16 MB chip, private to each shire | 4.0 B per minion-cycle, or 128 B per shire-cycle by the cycle counters: half the four banks' 256 B/cycle. |
| L2 scratchpad, local | 47 cyc 78 ns | 2.46 TB/s | 2.3 / 4.4 pJ (zeros / random) | 27.4 G/s 368 pJ | 2.5 MB per shire · 80 MB chip | Same SRAM, latency and bandwidth as L2; also cached in L1 and the read buffer. |
| L2 scratchpad, remote | 112–220 cyc 186–366 ns | 0.96 TB/s | 5.1 / 11.8 pJ (zeros / random) | 10.2 G/s 903 pJ | any other shire's 2.5 MB | 20 ns per mesh hop (12 cycles at 600 MHz, 16 at 800 MHz), the same in both directions. Bandwidth and energy are for the scratchpad 16 shire IDs away, 2.1 mesh hops on average. |
| L3 | 159–169 cyc 265–282 ns | 0.98 TB/s | 14.7 pJ | 6.46 G/s 1.53 nJ | 1 MB slice per shire · 32 MB chip, shared | Latency depends on the requesting shire (item 4). |
| DRAM | 287–297 cyc 479–495 ns | 76 GB/s | 115 pJ | 1.19 G/s 9.8 nJ | 32 GB LPDDR4X, 16 × 16-bit channels, 933 MHz DDR clock (3,733 MT/s) | About 64% of the 119 GB/s that 16 × 16-bit channels carry at the 3,733 MT/s implied by the 933 MHz DDR clock. The 128,000 MB/s the runtime reports is a placeholder constant in the service-processor firmware. At 800 MHz, which only aifoundry2's governor reaches, a load took 352–368 cycles, 440–460 ns. |
Method notes for this table
Every latency in this table held at 600 MHz on each of three cards on 26 September (three passes each), to within a cycle. Bandwidth is bytes over wall time for the launches here that ran at 600 MHz. These runs have few launches at 600 MHz (L1 2 of 21, DRAM 3 of 14); the same probes re-run at a steady 600 MHz on 23 September (six passes, three on each of two cards, 9–11 launches per level per pass; Ridge points) agree within 1%, and so do the three cards' energy re-runs of 26 September, which stream the same probes (below). Energy per byte is from the energy manual, §4: the same probes re-run at a steady 600 MHz on 26 September, six passes on each of three cards, with the range over passes there; each row gives the mean over every pass. The check filled the scratchpads with zeros (odd passes) or random data (even passes) before reading them, so those rows give both; it did not set the contents of the L1, L2, L3 and DRAM buffers. No two cards differ beyond the scatter between their passes at any level (99%, each pair of cards), except the local scratchpad on aifoundry1's card 1, which reads 2.5 / 5.0 pJ/B against 2.1 / 4.0–4.1 on aifoundry2 and aifoundry3, above the bands the check registered for it (1.7–2.3 and 3.7–4.7). The L2 and L3 rows move a lot from pass to pass (1.4–5.0 and 7.1–20.5 pJ/B). The L1 row is this probe's own loop; the A100 comparison below uses the energy manual's unrolled loop (§4.2). Energy is card power above idle, so it includes the busy cores. This page's two-hart spin loop, re-run at a steady 600 MHz on 26 September, added 2.37, 2.41 and 2.38 W over idle on aifoundry2, aifoundry3 and aifoundry1's card 1 (mean of six passes each; aifoundry2's passes spread from 2.17 to 2.61 W): about half of what an L1 read adds, and a sixth to a third of an L2, L3 or DRAM stream (2.4 of the 8.7 W a DRAM stream adds). The energy manual's addi loop adds 1.9–2.0 W with one hart per minion and 3.0–3.2 W with both, on each of the three cards (energy manual §2). The random 4 B element column is the word gather fgw.ps of the next section (for the L3 its global form fgwg.ps, which skips the L1 and L2; the remote scratchpad 2 mesh hops away): both harts of all 1,024 minions, each hart on a table sized to the level (512 B, 4 KB, 16 KB of scratchpad, 256 KB of DRAM per hart), elements per second over the whole chip and energy per element above idle, pooled over three passes on each of the three cards (26 September).
Three cards: the same time, and energy apart only in one scratchpad
The same design gives the same latency and bandwidth on every card. Energy per byte moves much more from pass to pass: aifoundry1's card 1 reads above the others at most levels, but only at its own scratchpad by more than the passes' scatter (99%, each pair of cards; the method note above). Each card's value at each level, against the three cards' mean. Hover, tap or focus a mark for its numbers.
Energy per byte, by level (chart)
The energy column drawn, one row per level: each card's mean energy per byte with its standard error over that card's passes, and in grey the range over every pass of every card, with a tick at their mean (the energy manual's re-runs at 600 MHz; the scratchpads once per contents; logarithmic scale). Hover, tap or focus a mark for its numbers.
Irregular access: gathers and scatters
The streams above read memory in order. A gather reads eight elements from eight addresses in one vector instruction (fgw.ps: a vector of byte offsets added to a base register); a scatter (fscw.ps) writes them. On 26 September the three cards measured them level by level (E48): both harts of all 1,024 minions, each hart on a table of its own sized to one level, eight random words on eight lines of a 4 KB tile per instruction, three passes on each card. "Random" here means random lines within 4 KB tiles, the tiles in a scrambled order, not uniform addresses over the table. From the L1 a gather is as fast and as cheap as scalar loads; past the L1 every element fetches its own 64 B line, two at a time per minion, so a random word costs 29× its streamed share from the L2 and 19× from DRAM (the energy manual's contiguous reads, per 4 B). Every tested item was decided card by card on the three cards, and their rates agree to 0.02%.
How this chart was measured, and the table
Both harts of all 1,024 minions issue the instruction back to back; the table size per hart is in each row's tooltip. Dots are the mean over every pass of the three cards (or one card's mean, with a card picked), bars the range of those passes (or ± that card's standard error); a hollow dot rests on fewer cards. The grey tick is the same level read contiguously, per 4 B, as the energy manual's §4.4 sets it: the L2, L3 and scratchpads as in the spec sheet (the scratchpads full of random data), the L1 as a 32 B vector load (flw.ps, 0.54 pJ/B) and DRAM as a tensor load (133 pJ/B), both of random data, where the spec sheet's own probes read 0.75 and 115 pJ/B. The data are the energy manual's §4.4.
Minion-cycles per instruction (eight elements) at 1,024 minions, both harts, random lines within 4 KB tiles; the dashed lines mark each level's capacity per hart when every hart of the chip holds its own table: 512 B of L1, 8 KB of L2 (512 KB over a shire's 64 harts), 16 KB of L3 (32 MB over 2,048 harts). Up to 1 KB per hart the eight offsets fold into fewer lines, so those points touch half a line, one or two lines per instruction instead of eight.
Rules for kernels
- Offsets are bytes, signed, from the base register, not element indices: a table of floats needs the index times 4. One vector register of offsets serves eight lanes; a masked lane neither reads nor faults, but in the L1 it still takes its issue slot.
- From the L1, gather freely. A word gather issues its eight elements in 10.9 minion-cycles (452 G elements/s over the chip, 12.8 pJ each), what eight scalar loads on the same addresses cost. When the eight lanes fall in one 32 B block, the block form
fg32w.psis 3.9× as fast at 0.30× the energy. - Past the L1 the cost is per line, not per element. A minion's two miss handlers fetch two lines at a time, so a gather whose lanes touch one or two lines costs one L2 latency and one touching eight lines four; the second hart adds nothing. Sort or bucket indices so that lanes share lines.
- Keep random-access tables on the chip. A random word from the L2 or the own scratchpad costs 354–368 pJ at 27.4 G/s; from DRAM 9.8 nJ at 1.19 G/s, the DRAM line rate (76 GB/s of lines for 4.8 GB/s of data).
- A scatter costs about twice a gather past the L1 (each element allocates its line and writes it back), and scatters into lines other minions touch are unsafe: the L1 is not coherent. Scatter into private lines, or use atomics.
- For scatter-add on private tables, gather, add and scatter (139 G updates/s at 0.036 nJ in the L1, 18.2 G/s at 0.80–0.85 nJ in the L2 or the own scratchpad). On shared tables the packed atomic
famoaddl.piadds eight lanes one after another, no faster than scalaramoaddl.w(12.7 against 12.8 G updates/s) and 1.45× its energy. No atomic may target a scratchpad address: that is a bus error. - The L1-bypassing forms do not help:
fgwl.psreads the L2 at 0.72× the gather through the L1, and the two harts of a minion share one limit (the second adds 3%). - They are exact. Every verify launch passed on every card (every gathered element and scattered word checked): the known erratum 1.3, elements skipped after a gather resumed from a trap, did not occur. When every lane scatters to one word, lane 7's value remains.
ET-SoC-1 against the A100, level by level
Each row puts the ET-SoC-1 value (blue) next to the A100's (grey) on a logarithmic scale. ET bandwidth and energy per byte are at 600 MHz. The ET scratchpad row is compared with the A100's shared memory, and ET's L3 with the far half of the A100's L2, the closest equivalents. The comparison is for streams: the A100's gather and scatter rates are not measured here, so the irregular-access section above has no GPU counterpart.
A100 (published) values and sources
| Level | Latency | Bandwidth, measured | Energy per byte | Capacity | Source |
|---|---|---|---|---|---|
| Registers | 4 cyc 2.8 ns | – | – | 256 KB per SM · 27 MB | Whitepaper; Abdelkhalik et al. 2022 |
| L1 / shared memory | 33 cyc 23 ns shared: 29 cyc, 21 ns | 14.8 TB/s 19.5 peak | 12.7 pJ | 192 KB per SM · 20.25 MB (≤164 KB shared) | Luo et al. 2025; SC'25 Table 3 |
| L2, near half | 208 cyc 148 ns | 4.4 TB/s 7.2 peak | 37.7 pJ | 40 MB total, 2 × 20 MB | Luo et al. 2025; Chips and Cheese; SC'25 |
| L2, far half | 357 cyc 253–313 ns | – | – | data past ~20 MB | Luo et al. 2025; Chips and Cheese |
| HBM2 / HBM2e | 570 cyc 405 ns | 1.40 / 1.77 TB/s | 105 pJ | 40 GB / 80 GB | Luo et al.; Chips and Cheese; BabelStream; SC'25 |
A100 cycles are at the 1,410 MHz boost clock. Its energy per byte is the SC'25 whole-GPU incremental energy (1.6, 4.7 and 13.1 pJ/bit), which includes instruction and memory-controller overhead. That makes it the counterpart to ET's card power above idle, not a circuit-level figure: the HBM2e devices themselves are about 4.3 pJ/bit (34 pJ/B).
What the numbers say
- Kernels get only 512 B of L1 per hart. Before every launch the firmware puts the L1 in scratchpad mode (Programmer's Reference Manual, table 8.4), turning 3 KB of the 4 KB into the tensor unit's scratchpad. The chase's first knee sits exactly at 512 B, on each of the three cards checked on 26 September. A user-mode kernel can switch to split mode, where hart 0 gets 2 KB, but the switch invalidates L1 lines without writing them back, so it was not tried on a shared card.
- There is a 2 KB level between L1 and L2: the L2's read buffer, an 8-entry fully associative cache per bank for clean lines. Working sets up to 2 KB hit it at 36 cycles, 11 fewer than the L2 array (36.0 and 47.0 on each of the three cards). That is the same 11-cycle difference the shire cache spec gives (10 vs 21 shire cycles).
- The L2 scratchpad is the L2 without the tags. It has the same 47-cycle latency, goes through the read buffer and L1, streams at the same 4.0 B per minion-cycle, and at a steady 600 MHz costs about the same energy per byte: pass by pass, the L2 minus the scratchpad included zero on each of the three cards checked on 26 September (energy manual §4). That check filled the scratchpad with zeros or random data before reading it, which it read at 2.1 and 4.1 pJ/B on aifoundry2, 2.1 and 4.0 on aifoundry3, and 2.5 and 5.0 on aifoundry1's card 1, while the L2, whose contents were not set, read 2.6, 2.8 and 3.9 pJ/B, between the two. The 4.3 against 2.8 pJ/B of this page's first runs (aifoundry2, two runs on 18 September) mostly came from the governor running the L2 runs at a higher clock.
- Distance on the mesh is visible. A remote scratchpad load costs 66 minion cycles plus 56 ns plus 20 ns per mesh hop: 186–366 ns at 600 MHz over 1–10 hops (on-chip communication maps the hops). Row 31 → 0 reads 271 cycles against 220 for 0 → 31 only because aifoundry2's governor ran that row at 800 MHz. The L3 hides distance by spreading lines over all 32 shires, so 31 of every 32 L3 hits cross the mesh and pay the average trip from the requesting shire: Anatomy of a memory access's 110 cycles + 12 per hop, averaged over the slices (4.2–5.0 hops), predicts 160–170 cycles, and the chases measured 159–169 at 600 MHz on each of three cards.
- aifoundry2 changed its clock during these runs. Its governor lifted the cool, busy card to 700 or 800 MHz (The DVFS loop and its leakage), so what follows is one card and one session (18 September). The chases carried no clock telemetry, so each chase's clock is inferred from its cycle count and wall time. On-chip latencies were the same in cycles at both clocks. L3 and DRAM loads also spend time in the 400 MHz mesh and the 933 MHz memory domain, neither of which speeds up with the minion clock, so at 800 MHz they took more cycles but fewer nanoseconds. Shire 24's L3 load is 159 cycles (265 ns) at 600 MHz (one chase) and 188 cycles (235 ns) at 800 MHz (five chases), two points that fix it at 72 cycles plus 145 ns. Of the 145 ns, 20 ns per hop is the mesh trip to the average slice (4.2 hops from shire 24, 5.0 from shire 0); 72 cycles + 61 ns + 20 ns per hop gives every requester's L3 latency here within a cycle at 600 MHz; so it did on each of three cards on 26 September (from shires 0, 7, 24 and 31; shire 0 read 9.7 cycles slower than shire 24 on every card). The DRAM chases fit 86 cycles plus 345 ns to within 10 cycles, the chases from shires 7 and 24 reading 8–10 cycles faster than those from 0 and 31 at 600 MHz (8.2–9.6 within each pass on the three cards; 10–11 at 800 MHz, shire 24 against shire 0), as their L3 hits do. Anatomy of a memory access takes a 600 MHz DRAM load on aifoundry2 apart.
How it was measured
How it was measured, in detail
- Latency. One hart chases a random cyclic chain with one pointer per 64 B line, which no prefetcher can predict. After warm-up (two passes, at least 16,384 and at most 219 loads), 40,000 dependent loads are timed (20,000 in the placement and offset repeats) with the minion cycle counter. Each chase's final pointer is checked against a host-side walk of the same chain, and every one of the 227 chases passed, as did every chase of the three-card check. Chains up to 256 MB live in DRAM. Scratchpad chains are written into the target shire's scratchpad by the chasing hart.
- Bandwidth. Hart 0 of all 1,024 minions streams 1 KB TensorLoads (tensor-unit loads into the L1 scratchpad, which bypass the L1 cache but go through the L2), two in flight, over private slices sized to hit L2 (8 KB per minion, 256 KB per shire), L3 (24 MB in total) or DRAM (256 MB), or over the local or a remote scratchpad. L1 bandwidth uses 32 B vector loads by both harts of every minion over a private 256 B buffer. Throughput is bytes over wall time; the tables use the launches that ran at 600 MHz.
- Energy. The energy per byte in the tables and charts is the energy manual's §4: the same probes re-run on 26 September at 600 MHz in the three-card check (pinned there on aifoundry3, held there on aifoundry2 and aifoundry1's card 1 by a warm die), each pass against idle readings taken before and after and corrected for the die's warming, six passes on each of three cards, with the scratchpads filled with zeros or random data before they were read (
docs/reports/data/2026-09-23-energy-manual/reruns.json, whichworkloads/memhier/analyze.pyembeds). First measured here on 18 September, on aifoundry2 alone, as board power against the median idle before and after each run (about 27.0 W) while the governor raised the clock: item 3 above gives what the clock did to the L2 and scratchpad rows, and the method is inworkloads/memhier/run_energy.py. - Lab etiquette. Every card run was its own process under
timeout 10, on an otherwise idle card. - Versions. 18 September: first published from one session on aifoundry2, with bandwidth averaged over launches at 600–800 MHz (L1 7.55, L2 2.80, local scratchpad 2.57, L3 1.03, remote scratchpad 0.98 TB/s), DRAM latency of ~440 ns (421–486 ns) over chases at both clocks, and the energy per byte measured here, the mean of two runs (L1 1.8, L2 4.3, local scratchpad 2.8, remote scratchpad 6.3, L3 10.8, DRAM 148 pJ/B). 24 September: bandwidth at 600 MHz, energy per byte from the energy manual, DRAM and L3 latency given per clock, and the apparent asymmetry of the scratchpad rows traced to the clock. 25 September (version 3): the energy per byte of L1, L2 and the local scratchpad given per card; the clock results marked as aifoundry2's, from one session; the scratchpad model's test limited to the rows it was not fitted to; the claim that the L3 model also fits shire 0 at 800 MHz, where no chase ran, removed; 600 MHz on 23 September described as held, not pinned, on aifoundry2; the cycle-counter caveat corrected. 26 September (version 4): the latencies re-measured at 600 MHz on three cards under a pre-registered plan, and drawn in the latency chart; the spin loop's power at 600 MHz given per card; the energy per byte taken from the energy manual's three-card re-runs of 26 September, the scratchpads by their contents (record:
docs/reports/data/2026-09-25-claims-v3). 27 September: the A100 comparison takes the 40 GB card's HBM2 bandwidth (1.40 TB/s, was the 80 GB card's 1.77) and shared memory's sustained 14.8 TB/s (was the 19.5 peak), as its latency and energy rows do; the item on the busy-core floor and the unpinned-clock caveat cut (their three-card figures are in the table note). 27 September (version 5): gathers and scatters on the three cards (E48, 26 September): the irregular-access section, its two charts and rules for kernels, and the spec sheet's random-element column (record:tools/claims-v3/gs,docs/reports/data/2026-09-25-claims-v3/results/gs.json). 27 September, later: the three cards' latency, bandwidth and energy per byte charted against their mean (the spec sheet; the bandwidths from the three-card check's energy re-runs, which also re-measured this page's bandwidth column on every card); the DRAM requester difference at 600 MHz is 8–10 cycles (was 9–11, which mixed in the 800 MHz chases); the lede's A100 bandwidth ratio follows the 40 GB card's 1.40 TB/s (an eighteenth; was a twentieth, from the 80 GB card's 1.77). 28 September: the review's cuts (the clock item and the repeated chase figures); the note's counts given by both rules. 28 September, later: the energies first measured here kept only as a note on how they were measured (the "Energy" item above), the energy manual's values stated as this page's own.
Caveats
Caveats in full
- The A100 column is not our measurement. Latency and bandwidth come from several published microbenchmark studies, which agree within a few percent. The energy numbers come from a single study (SC'25). Their method is close to ours but not identical.
- ET energy per byte is whole-card power above idle. It includes the busy chip's floor, DRAM chips, regulators and PCIe, so it overstates the cost of the memory arrays themselves, most of all for DRAM. The probes read L2, L3 and DRAM buffers whose contents they never set (the scratchpads were filled with zeros or random data); the energy manual gives tensor loads on zeros and on random data separately.
- Peak-rate microbenchmarks. Real kernels mix levels, and bank conflicts, the mesh, and the DRAM controller's page policy will change the picture.
Reproduce
Reproduce this
Run from a checkout of this repository. scripts/deploy-lab.sh copies the workload into ~/nekko on the lab machine and builds it there.
# build on the lab machine
scripts/deploy-lab.sh aifoundry2 workloads/memhier
# every latency chase, the scratchpad map, the clock poll and both energy runs (each probe is its own timeout-10 process)
ssh aifoundry2 'cd ~/nekko && bash workloads/memhier/run_lab.sh build/memhier/host/memhier_host build/memhier-data'
# latency curves, clock models and scratchpad map for this page
python3 workloads/memhier/analyze.py docs/reports/data/2026-09-18-memhier-aifoundry2 --v3 docs/reports/data/2026-09-25-claims-v3/raw \
--embed docs/reports/2026-09-18-et-soc1-memory-hierarchy.html # --v3: the three-card check's chases (26 September)
# the bandwidth column (launches at 600 MHz) and bytes per cycle
python3 scripts/ridge-points.py
# energy per byte: the energy manual's re-runs, docs/reports/data/2026-09-23-energy-manual/reruns.json, levels_pj_per_byte
Related reports
- On-chip communication — the mesh map and its 20 ns per hop, on aifoundry2, and the model that fits this page's scratchpad rows at 600, 700 and 800 MHz.
- Anatomy of a memory access — on aifoundry2, L3 at 110 cycles plus 12 per hop, and a 600 MHz DRAM load of 299 cycles taken apart.
- The energy manual, §4 — the source of this page's energy per byte: these levels at 600 MHz on three cards, with zeros and random data, and the writes and the fine grain (wires, lines, rows) beside them.
- Ridge points — the bandwidths at 600 MHz per cycle, and how much reuse each level needs.
- The DVFS loop and its leakage — the governor that moved the clock during these runs.
- Hand it to the next shire — the scratchpads used as a pipeline between shires, against a round trip through DRAM.
Sources
Sources and citations
- ET-SoC-1: ET Programmer's Reference Manual (aifoundry-org/et-man): table 8.4 (L1 modes) and §15.3 (scratchpad addressing); CORE-ET Shire Cache Specification (openhwfoundation/core-et, Erbium branch, which documents a later configuration of the same design): latency table and read buffer; et-platform firmware source read at
353f20e(init_l1,system/layout.h; the cards' own trace strings match an older build, from before et-platform commit60b40c10fof 24 September 2024). - NVIDIA A100 Tensor Core GPU architecture whitepaper and datasheet.
- Abdelkhalik et al., Demystifying the Nvidia Ampere Architecture through Microbenchmarking, 2022.
- Luo et al., Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis, 2025 (includes A100 near and far L2).
- Chips and Cheese, A100 SXM4 and PCIe latency data, and Nvidia's H100: Funny L2, and Tons of Bandwidth.
- Antepara et al., Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing, SC'25, table 3 (A100 per-level energy).
- BabelStream A100-SXM4-80GB results; O'Connor et al., Fine-Grained DRAM, MICRO 2017 (HBM2 ≈ 3.9 pJ/bit); SK hynix, IEDM 2023 (HBM2e ≈ 4.3 pJ/bit).