ML-KEM and ML-DSA fill matrices/vectors with output from SHAKE, each element derived from slightly different inputs.
This is perfect for SIMD instructions, so now we have
func ReadMulti(s []*SHAKE, out [][]byte)
backed by AVX2 on amd64, and by the existing two-lane asm on arm64.
go.dev/cl/818720