Mastodon Feed: Post

Mastodon Feed

Boosted by glyph ("Glyph"):
mhoye@cosocial.ca wrote:

Some Stanford researchers dredge out the absolute dregs of r/AmITheAsshole for all those stories where _zero percent_ of humans supported the poster, where one hundred percent of them said, yes you are unambiguously the asshole, and fed them into LLMs as first person scenarios, to see what the LLM had to say.

Unethical, harmful, cruel, criminal, didn't matter: slopbots took the faux poster's side about half the time.

https://www.science.org/doi/10.1126/science.aec8352

RESULTS We find that sycophancy is both prevalent and harmful. Across 11 AI models, AI affirmed users’ actions 49% more often than humans on average, including in cases involving deception, illegality, or other harms. On posts from r/AmITheAsshole, AI systems affirm users in 51% of cases where human consensus does not (0%). In our human experiments, even a single interaction with sycophantic AI reduced participants’ willingness to take responsibility and repair interpersonal conflicts, while increasing their own conviction that they were right. Yet despite distorting judgment, sycophantic models were trusted and preferred. All of these effects persisted when controlling for individual traits such as demographics and prior familiarity with AI; perceived response source; and response style. This creates perverse incentives for sycophancy to persist: The very feature that causes harm also drives engagement. CONCLUSION AI sycophancy is not merely a stylistic issue or a niche risk, but a prevalent behavior with broad downstream consequences.