What makes an AI chatbot go “rogue”?

It has become a pressing question in recent weeks. One pair of researchers from George Washington University think they have an answer, and possibly even a way to prevent it from happening in the future. Professor of physics Neil Johnson, who studies complex systems, and physics PhD student Frank Yingjie Huo published a paper about their formula today in the peer-reviewed journal Patterns.

Johnson and Huo argue that there is a visible “tipping point” embedded in the internal code of an AI chatbot when it goes off the rails—encouarging users to self-harm, spreading medical misinformation, or supporting violent points of view, for example. Where this tipping point lies depends on where certain concepts are stored on the giant semantic map within the AI agent’s network, and it can be nudged. Nobody draws this map for the AI. It develops during training and tends to situate ideas that are similar, such as garbage and stink, close together, while ideas that are unrelated sit far apart.

To read more, click here.