DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459
13:10reinforcement learning from human feedback. We'll get into some of these words. And this is what they did to create the DeepSeek-V3 model.16:49Preference fine-tuning is a generalized term for what came out of reinforcement learning from human feedback, which is RLHF.2:38:50it doesn't affect math abilities as much, but something like a, if you're trying, it's just the quality of a human judgment












