What Is Model Alignment?
Model alignment ensures AI systems behave according to human values, intentions, and safety requirements through training techniques that steer model behavior toward desired outcomes and away from harmful actions. Alignment addresses the challenge that AI models optimize for their training objectives, which may not perfectly match human preferences or ethical standards. Techniques like Reinforcement Learning from Human Feedback (RLHF) use human ratings to teach models what responses are helpful, harmless, and honest.