// HACKER NEWS — CYBERSECURITY
What happens when a GPU writes memory
In our previous post we followed
one LDG.E down through all the hardware units on an RTX 4090 — through its
L1, translation, through the crossbar to its L2 slice, and thence to DRAM. The
request retrieved its result, and then came back up through its waystations,
returning its result to its warp, which then continued to execute its part in
the kernel.
The part it was playing was in a kernel that performed a vector add. The same
kernel, once it has loaded elements of both vectors, adds them together, and
then stores the result.
STG.E is the instruction that’s responsible for writing the calculated sum
back to global memory. In this post we’re going to follow STG.E through the
same waystations, figuring out what happens at each step. As before, this
information is not all publicly available; where it isn’t, we’ll run new
experiments.
Setting the scene: the LDG.E has returned to the warp, the FADD has
added together the contents of R4 and R3 into R9, and now, the contents
of that register must be stored. The warp has become eligible within its
subpartition, and its lanes start to execute STG.E.
STG.E [R6.64], R9 is a global store of the 32 bits in register R9 to the
64-bit address in R6 and R7. Where in LDG.E, we read two rows of the
register file, in STG.E, we must read three: the two making up the address to
which we’re going to store the data, and the data itself.
The instruction then issues to the load/store unit. The LSU sends on the
opcode (‘store to these addresses’), the 32-bit mask of active lanes, and the
32 computed addresses.
One warp can push a new STG.E instruction through register/LSU/coalescer/L1
about every 6.1 cycles (2.3 ns at 2.6 GHz), regardless of how many lanes it
issues for.
The exit from the SM can sustain 32 bytes stored (or loaded) per cycle, so if
all the warps are issuing, they’ll bottleneck here1.
The next stop is the coalescer. Its job is to take
32 four-byte accesses and turn them into the smallest achievable number of
32-byte sectors. Our kernel writes 128 contiguous bytes, so that’s four
sectors, or one line2.
Loads always pull in all 32 bytes per sector, and then in the LSU the results
are filtered to write to the registers what the SASS actually asked for. For
stores, each sector request issues with a byte mask, indicating which bytes
of the sector this instruction is writing.