The ET-SoC-1 memory hierarchy, measured on silicon, next to the A100

18 September 2026 · aifoundry2 in the AI Foundry lab: one ET-SoC-1 PCIe card, 32 GB LPDDR4X, 1,024 compute minions; the latencies are one session on this card, re-measured at 600 MHz on three cards (aifoundry2, aifoundry3 and aifoundry1's card 1) on 26 September, the bandwidths re-measured on two cards (aifoundry2 and aifoundry3) on 23 September and on the three on 26 September, the energies per byte on the three cards on 26 September, and gathers and scatters on the same three cards the same day · A100 values from published microbenchmarks and energy studies (sources at the end) · code: workloads/memhier; raw data: docs/reports/data/2026-09-18-memhier-aifoundry2 · part of the ET-SoC-1 measurement reports

Pointer-chase and streaming probes give the card's memory hierarchy as a spec sheet. ET-SoC-1's on-chip memories are fast and cheap per byte: a load hits L1 in 8.8 ns and the shire's L2 in 78 ns, against 23 ns and 148 ns on an A100, at 12–24× less energy per byte (the energy manual's 600 MHz values, three cards). They are also tiny per core: the firmware leaves each hart only 512 B of L1 data cache. Off chip the picture flips. DRAM latency is about 490 ns at the 600 MHz base clock, a fifth more than the A100's HBM (405 ns), and the card streams 76 GB/s, about an eighteenth of an A100's 1.40 TB/s, at roughly 1.1× the energy per byte (1.0–1.2× card by card).

Terms used on this page

The ET-SoC-1's compute cores are minions, small in-order RISC-V cores that each run two hardware threads (harts); eight minions form a neighbourhood, 32 a shire, and 32 shires (1,024 minions) run kernels. Each shire's 4 MB of SRAM is split by these cards' firmware into a 512 KB L2 cache, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB scratchpad, software-managed memory that any shire can address. The shires and eight memory shires, which hold the LPDDR4X controllers, sit on a 2D mesh network-on-chip clocked at 400 MHz, where a mesh hop is one step between neighbouring stops; the minions run at 600 MHz unless the card's power governor raises the clock. More in the hub's glossary.

Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan, three passes of every latency chase on each of aifoundry2, aifoundry3 and aifoundry1's card 1 at 600 MHz, and six passes of the energy probes on each card (the record). Of 20 claims tested here, this page counts 16 held, 2 corrected and 2 differ by card; the hub’s scoreboard, 16 “proven on the cards”, 3 “a test behind it failed” and 1 “differs by card”. Every latency in the spec sheet agreed on each card to within a cycle, and the latency chart shows each card's chases beside the 18 September session. The energy per byte is now the energy manual's three-card values of 26 September, with the scratchpads read full of zeros and full of random data. The same three cards then measured gathers and scatters, random elements level by level (E48, below), each tested item decided card by card.

L1 data cache hit
8.8 ns
5.25 cycles; 512 B per hart; the same on three cards (26 Sep)
L2 / local scratchpad hit
78 ns
47 cycles; 512 KB L2 + 2.5 MB scratchpad per shire; the same on three cards (26 Sep)
L3 hit (32 MB, over the mesh)
~280 ns
159–169 cycles at 600 MHz (265–282 ns), depending on the requesting shire; 159.1–168.9 on three cards (26 Sep)
DRAM
~490 ns
287–297 cycles at 600 MHz (287.2–297.1 on three cards, 26 Sep); 440–460 ns when aifoundry2's governor ran it at 800 MHz; 76 GB/s streaming

Load latency against working-set size

How to read this chart

What does a load at each level cost at 600, 700 and 800 MHz, and which part of it speeds up with the minion clock? Only aifoundry2's data can answer this: its governor moved the clock during this page's one session of chases (18 September), while a boot service holds aifoundry3 at 600 MHz (it sets that card's TDP to 0 W at every boot). One hart follows a random pointer chain, so every load waits for the one before it; each step in the curve is a level of the hierarchy. On-chip levels are fixed in cycles, but L3 and DRAM loads also spend fixed time in the 400 MHz mesh and the DDR memory, so the lines are redrawn for the selected clock from two fits (What the numbers say, item 5). Circles are the chases as measured, each at its own clock; filled ones ran at the selected clock (none ran at 700 MHz). From 16 to 64 MB a chain mixes L3 and DRAM hits, and a dashed line joins the measured points. Grey marks are the A100's published pointer-chase plateaus. The card marks are the three-card check of 26 September, every chase at 600 MHz (each card's median of three passes, from shires 0 and 24), shown when 600 MHz is selected; there the chains up to 32 MB stayed at L3 latency on every card.

The latency split by clock

Where the time goes at the selected clock: the part of each load that scales with the minion clock, and the fixed part that does not (L3 as chased from shire 0).

Scratchpad latency across the mesh

How to read these charts

How much does mesh distance cost a remote scratchpad load, and why does the L3 latency depend on the requesting shire? Chases ran from shires 0, 7, 24 and 31 to all 32 scratchpads, over a 64 KB chain so that every load misses L1. The scatter puts each load against the mesh hops between the two shires, with the model fitted to rows 0 and 31 (66 minion cycles + 56 ns + 20 ns per hop) at 600, 700 and 800 MHz; each point takes the colour of the nearest line. The map shades each shire by the load's time in nanoseconds, with the requesting shire outlined. L3 average trip shows instead how far each of the 32 L3 slices is from the requesting shire. The three-card check re-ran the four rows at 600 MHz on each card (26 September): every remote load fell within 0.2 cycles of 99.84 + 12.00 cycles per hop, the model's 600 MHz line, and each pair of shires read the same in both directions to within 0.04 cycles.

The hierarchy as a spec sheet

ET-SoC-1 (measured)

Every ET figure is at the 600 MHz base clock.

LevelLatencyBandwidth, whole chipEnergy per byteRandom 4 B elementCapacityNotes
L1 data cache5.25 cyc
8.8 ns
6.2 TB/s0.75 pJ452 G/s
12.8 pJ
512 B per hart (2 sets × 4 ways × 64 B)The firmware puts the L1 in scratchpad mode before every launch: sets 0–11 become a 3 KB tensor scratchpad. Hart 1 is the same size but pays +3 cycles on misses. Bandwidth is about 5 B per hart-cycle.
L2 read buffer36 cyc
60 ns
–––8 lines per bank × 4 banks = 2 KBA small clean-line cache in front of each L2 bank; spec: 10 vs 21 shire cycles.
L2 cache47 cyc
78 ns
2.45 TB/s3.1 pJ27.4 G/s
354 pJ
512 KB per shire · 16 MB chip, private to each shire4.0 B per minion-cycle, or 128 B per shire-cycle by the cycle counters: half the four banks' 256 B/cycle.
L2 scratchpad, local47 cyc
78 ns
2.46 TB/s2.3 / 4.4 pJ
(zeros / random)
27.4 G/s
368 pJ
2.5 MB per shire · 80 MB chipSame SRAM, latency and bandwidth as L2; also cached in L1 and the read buffer.
L2 scratchpad, remote112–220 cyc
186–366 ns
0.96 TB/s5.1 / 11.8 pJ
(zeros / random)
10.2 G/s
903 pJ
any other shire's 2.5 MB20 ns per mesh hop (12 cycles at 600 MHz, 16 at 800 MHz), the same in both directions. Bandwidth and energy are for the scratchpad 16 shire IDs away, 2.1 mesh hops on average.
L3159–169 cyc
265–282 ns
0.98 TB/s14.7 pJ6.46 G/s
1.53 nJ
1 MB slice per shire · 32 MB chip, sharedLatency depends on the requesting shire (item 4).
DRAM287–297 cyc
479–495 ns
76 GB/s115 pJ1.19 G/s
9.8 nJ
32 GB LPDDR4X, 16 × 16-bit channels, 933 MHz DDR clock (3,733 MT/s)About 64% of the 119 GB/s that 16 × 16-bit channels carry at the 3,733 MT/s implied by the 933 MHz DDR clock. The 128,000 MB/s the runtime reports is a placeholder constant in the service-processor firmware. At 800 MHz, which only aifoundry2's governor reaches, a load took 352–368 cycles, 440–460 ns.
Method notes for this table

Every latency in this table held at 600 MHz on each of three cards on 26 September (three passes each), to within a cycle. Bandwidth is bytes over wall time for the launches here that ran at 600 MHz. These runs have few launches at 600 MHz (L1 2 of 21, DRAM 3 of 14); the same probes re-run at a steady 600 MHz on 23 September (six passes, three on each of two cards, 9–11 launches per level per pass; Ridge points) agree within 1%, and so do the three cards' energy re-runs of 26 September, which stream the same probes (below). Energy per byte is from the energy manual, §4: the same probes re-run at a steady 600 MHz on 26 September, six passes on each of three cards, with the range over passes there; each row gives the mean over every pass. The check filled the scratchpads with zeros (odd passes) or random data (even passes) before reading them, so those rows give both; it did not set the contents of the L1, L2, L3 and DRAM buffers. No two cards differ beyond the scatter between their passes at any level (99%, each pair of cards), except the local scratchpad on aifoundry1's card 1, which reads 2.5 / 5.0 pJ/B against 2.1 / 4.0–4.1 on aifoundry2 and aifoundry3, above the bands the check registered for it (1.7–2.3 and 3.7–4.7). The L2 and L3 rows move a lot from pass to pass (1.4–5.0 and 7.1–20.5 pJ/B). The L1 row is this probe's own loop; the A100 comparison below uses the energy manual's unrolled loop (§4.2). Energy is card power above idle, so it includes the busy cores. This page's two-hart spin loop, re-run at a steady 600 MHz on 26 September, added 2.37, 2.41 and 2.38 W over idle on aifoundry2, aifoundry3 and aifoundry1's card 1 (mean of six passes each; aifoundry2's passes spread from 2.17 to 2.61 W): about half of what an L1 read adds, and a sixth to a third of an L2, L3 or DRAM stream (2.4 of the 8.7 W a DRAM stream adds). The energy manual's addi loop adds 1.9–2.0 W with one hart per minion and 3.0–3.2 W with both, on each of the three cards (energy manual §2). The random 4 B element column is the word gather fgw.ps of the next section (for the L3 its global form fgwg.ps, which skips the L1 and L2; the remote scratchpad 2 mesh hops away): both harts of all 1,024 minions, each hart on a table sized to the level (512 B, 4 KB, 16 KB of scratchpad, 256 KB of DRAM per hart), elements per second over the whole chip and energy per element above idle, pooled over three passes on each of the three cards (26 September).

Three cards: the same time, and energy apart only in one scratchpad

The same design gives the same latency and bandwidth on every card. Energy per byte moves much more from pass to pass: aifoundry1's card 1 reads above the others at most levels, but only at its own scratchpad by more than the passes' scatter (99%, each pair of cards; the method note above). Each card's value at each level, against the three cards' mean. Hover, tap or focus a mark for its numbers.

Energy per byte, by level (chart)

The energy column drawn, one row per level: each card's mean energy per byte with its standard error over that card's passes, and in grey the range over every pass of every card, with a tick at their mean (the energy manual's re-runs at 600 MHz; the scratchpads once per contents; logarithmic scale). Hover, tap or focus a mark for its numbers.

Irregular access: gathers and scatters

The streams above read memory in order. A gather reads eight elements from eight addresses in one vector instruction (fgw.ps: a vector of byte offsets added to a base register); a scatter (fscw.ps) writes them. On 26 September the three cards measured them level by level (E48): both harts of all 1,024 minions, each hart on a table of its own sized to one level, eight random words on eight lines of a 4 KB tile per instruction, three passes on each card. "Random" here means random lines within 4 KB tiles, the tiles in a scrambled order, not uniform addresses over the table. From the L1 a gather is as fast and as cheap as scalar loads; past the L1 every element fetches its own 64 B line, two at a time per minion, so a random word costs 29× its streamed share from the L2 and 19× from DRAM (the energy manual's contiguous reads, per 4 B). Every tested item was decided card by card on the three cards, and their rates agree to 0.02%.

A random 4-byte element by level: gathers, scatters and scalar loads on the same addresses

How this chart was measured, and the table

Both harts of all 1,024 minions issue the instruction back to back; the table size per hart is in each row's tooltip. Dots are the mean over every pass of the three cards (or one card's mean, with a card picked), bars the range of those passes (or ± that card's standard error); a hollow dot rests on fewer cards. The grey tick is the same level read contiguously, per 4 B, as the energy manual's §4.4 sets it: the L2, L3 and scratchpads as in the spec sheet (the scratchpads full of random data), the L1 as a 32 B vector load (flw.ps, 0.54 pJ/B) and DRAM as a tensor load (133 pJ/B), both of random data, where the spec sheet's own probes read 0.75 and 115 pJ/B. The data are the energy manual's §4.4.

Cycles per gather and per scatter against the table's size per hart: the irregular counterpart of the latency curve

Minion-cycles per instruction (eight elements) at 1,024 minions, both harts, random lines within 4 KB tiles; the dashed lines mark each level's capacity per hart when every hart of the chip holds its own table: 512 B of L1, 8 KB of L2 (512 KB over a shire's 64 harts), 16 KB of L3 (32 MB over 2,048 harts). Up to 1 KB per hart the eight offsets fold into fewer lines, so those points touch half a line, one or two lines per instruction instead of eight.

Rules for kernels

  1. Offsets are bytes, signed, from the base register, not element indices: a table of floats needs the index times 4. One vector register of offsets serves eight lanes; a masked lane neither reads nor faults, but in the L1 it still takes its issue slot.
  2. From the L1, gather freely. A word gather issues its eight elements in 10.9 minion-cycles (452 G elements/s over the chip, 12.8 pJ each), what eight scalar loads on the same addresses cost. When the eight lanes fall in one 32 B block, the block form fg32w.ps is 3.9× as fast at 0.30× the energy.
  3. Past the L1 the cost is per line, not per element. A minion's two miss handlers fetch two lines at a time, so a gather whose lanes touch one or two lines costs one L2 latency and one touching eight lines four; the second hart adds nothing. Sort or bucket indices so that lanes share lines.
  4. Keep random-access tables on the chip. A random word from the L2 or the own scratchpad costs 354–368 pJ at 27.4 G/s; from DRAM 9.8 nJ at 1.19 G/s, the DRAM line rate (76 GB/s of lines for 4.8 GB/s of data).
  5. A scatter costs about twice a gather past the L1 (each element allocates its line and writes it back), and scatters into lines other minions touch are unsafe: the L1 is not coherent. Scatter into private lines, or use atomics.
  6. For scatter-add on private tables, gather, add and scatter (139 G updates/s at 0.036 nJ in the L1, 18.2 G/s at 0.80–0.85 nJ in the L2 or the own scratchpad). On shared tables the packed atomic famoaddl.pi adds eight lanes one after another, no faster than scalar amoaddl.w (12.7 against 12.8 G updates/s) and 1.45× its energy. No atomic may target a scratchpad address: that is a bus error.
  7. The L1-bypassing forms do not help: fgwl.ps reads the L2 at 0.72× the gather through the L1, and the two harts of a minion share one limit (the second adds 3%).
  8. They are exact. Every verify launch passed on every card (every gathered element and scattered word checked): the known erratum 1.3, elements skipped after a gather resumed from a trap, did not occur. When every lane scatters to one word, lane 7's value remains.

ET-SoC-1 against the A100, level by level

Each row puts the ET-SoC-1 value (blue) next to the A100's (grey) on a logarithmic scale. ET bandwidth and energy per byte are at 600 MHz. The ET scratchpad row is compared with the A100's shared memory, and ET's L3 with the far half of the A100's L2, the closest equivalents. The comparison is for streams: the A100's gather and scatter rates are not measured here, so the irregular-access section above has no GPU counterpart.

A100 (published) values and sources
LevelLatencyBandwidth, measuredEnergy per byteCapacitySource
Registers4 cyc
2.8 ns
––256 KB per SM · 27 MBWhitepaper; Abdelkhalik et al. 2022
L1 / shared memory33 cyc
23 ns
shared: 29 cyc, 21 ns
14.8 TB/s
19.5 peak
12.7 pJ192 KB per SM · 20.25 MB (≤164 KB shared)Luo et al. 2025; SC'25 Table 3
L2, near half208 cyc
148 ns
4.4 TB/s
7.2 peak
37.7 pJ40 MB total, 2 × 20 MBLuo et al. 2025; Chips and Cheese; SC'25
L2, far half357 cyc
253–313 ns
––data past ~20 MBLuo et al. 2025; Chips and Cheese
HBM2 / HBM2e570 cyc
405 ns
1.40 / 1.77 TB/s105 pJ40 GB / 80 GBLuo et al.; Chips and Cheese; BabelStream; SC'25

A100 cycles are at the 1,410 MHz boost clock. Its energy per byte is the SC'25 whole-GPU incremental energy (1.6, 4.7 and 13.1 pJ/bit), which includes instruction and memory-controller overhead. That makes it the counterpart to ET's card power above idle, not a circuit-level figure: the HBM2e devices themselves are about 4.3 pJ/bit (34 pJ/B).

What the numbers say

  1. Kernels get only 512 B of L1 per hart. Before every launch the firmware puts the L1 in scratchpad mode (Programmer's Reference Manual, table 8.4), turning 3 KB of the 4 KB into the tensor unit's scratchpad. The chase's first knee sits exactly at 512 B, on each of the three cards checked on 26 September. A user-mode kernel can switch to split mode, where hart 0 gets 2 KB, but the switch invalidates L1 lines without writing them back, so it was not tried on a shared card.
  1. There is a 2 KB level between L1 and L2: the L2's read buffer, an 8-entry fully associative cache per bank for clean lines. Working sets up to 2 KB hit it at 36 cycles, 11 fewer than the L2 array (36.0 and 47.0 on each of the three cards). That is the same 11-cycle difference the shire cache spec gives (10 vs 21 shire cycles).
  2. The L2 scratchpad is the L2 without the tags. It has the same 47-cycle latency, goes through the read buffer and L1, streams at the same 4.0 B per minion-cycle, and at a steady 600 MHz costs about the same energy per byte: pass by pass, the L2 minus the scratchpad included zero on each of the three cards checked on 26 September (energy manual §4). That check filled the scratchpad with zeros or random data before reading it, which it read at 2.1 and 4.1 pJ/B on aifoundry2, 2.1 and 4.0 on aifoundry3, and 2.5 and 5.0 on aifoundry1's card 1, while the L2, whose contents were not set, read 2.6, 2.8 and 3.9 pJ/B, between the two. The 4.3 against 2.8 pJ/B of this page's first runs (aifoundry2, two runs on 18 September) mostly came from the governor running the L2 runs at a higher clock.
  3. Distance on the mesh is visible. A remote scratchpad load costs 66 minion cycles plus 56 ns plus 20 ns per mesh hop: 186–366 ns at 600 MHz over 1–10 hops (on-chip communication maps the hops). Row 31 → 0 reads 271 cycles against 220 for 0 → 31 only because aifoundry2's governor ran that row at 800 MHz. The L3 hides distance by spreading lines over all 32 shires, so 31 of every 32 L3 hits cross the mesh and pay the average trip from the requesting shire: Anatomy of a memory access's 110 cycles + 12 per hop, averaged over the slices (4.2–5.0 hops), predicts 160–170 cycles, and the chases measured 159–169 at 600 MHz on each of three cards.
  4. aifoundry2 changed its clock during these runs. Its governor lifted the cool, busy card to 700 or 800 MHz (The DVFS loop and its leakage), so what follows is one card and one session (18 September). The chases carried no clock telemetry, so each chase's clock is inferred from its cycle count and wall time. On-chip latencies were the same in cycles at both clocks. L3 and DRAM loads also spend time in the 400 MHz mesh and the 933 MHz memory domain, neither of which speeds up with the minion clock, so at 800 MHz they took more cycles but fewer nanoseconds. Shire 24's L3 load is 159 cycles (265 ns) at 600 MHz (one chase) and 188 cycles (235 ns) at 800 MHz (five chases), two points that fix it at 72 cycles plus 145 ns. Of the 145 ns, 20 ns per hop is the mesh trip to the average slice (4.2 hops from shire 24, 5.0 from shire 0); 72 cycles + 61 ns + 20 ns per hop gives every requester's L3 latency here within a cycle at 600 MHz; so it did on each of three cards on 26 September (from shires 0, 7, 24 and 31; shire 0 read 9.7 cycles slower than shire 24 on every card). The DRAM chases fit 86 cycles plus 345 ns to within 10 cycles, the chases from shires 7 and 24 reading 8–10 cycles faster than those from 0 and 31 at 600 MHz (8.2–9.6 within each pass on the three cards; 10–11 at 800 MHz, shire 24 against shire 0), as their L3 hits do. Anatomy of a memory access takes a 600 MHz DRAM load on aifoundry2 apart.

How it was measured

How it was measured, in detail

Caveats

Caveats in full

Reproduce

Reproduce this

Run from a checkout of this repository. scripts/deploy-lab.sh copies the workload into ~/nekko on the lab machine and builds it there.

# build on the lab machine
scripts/deploy-lab.sh aifoundry2 workloads/memhier

# every latency chase, the scratchpad map, the clock poll and both energy runs (each probe is its own timeout-10 process)
ssh aifoundry2 'cd ~/nekko && bash workloads/memhier/run_lab.sh build/memhier/host/memhier_host build/memhier-data'

# latency curves, clock models and scratchpad map for this page
python3 workloads/memhier/analyze.py docs/reports/data/2026-09-18-memhier-aifoundry2 --v3 docs/reports/data/2026-09-25-claims-v3/raw \
    --embed docs/reports/2026-09-18-et-soc1-memory-hierarchy.html   # --v3: the three-card check's chases (26 September)

# the bandwidth column (launches at 600 MHz) and bytes per cycle
python3 scripts/ridge-points.py

# energy per byte: the energy manual's re-runs, docs/reports/data/2026-09-23-energy-manual/reruns.json, levels_pj_per_byte

Sources

Sources and citations

← All ET-SoC-1 measurement reports