safety
Sycophancy
A behavior pattern where AI models agree with users even when the user is incorrect, prioritizing user satisfaction over accuracy. Sycophancy results from RLHF training where models learn that agreeable responses receive higher ratings.
In practice
A sycophantic model might say 'Great point!' when a user asserts an incorrect fact, rather than politely correcting them.