RLHF
Training a model on human preferences between its own outputs, which is how a raw text predictor turns into something that answers questions helpfully.
ai
also: reinforcement learning from human feedback
Nobody says this one in the transcripts we have yet.