mixture of experts
An architecture where only a fraction of the network runs for any given token. Lets a model be very large in total while staying affordable to run.
ai
also: MoE
Nobody says this one in the transcripts we have yet.