OpenAI recently found itself chasing an unexpected bug—not in the code, but in the personality of one of its AI models. During testing, engineers discovered that newer versions of the model were mentioning “gremlins, goblins, dragons, and unicorns” with unusual frequency. 

The culprit wasn’t a hidden Easter egg, but the way AI models are trained: user conversations were fed back into future versions of the model, and many of those conversations used a popular “nerdy” personality setting in the AI interface. The more people talked with the AI about gremlins and dragons, the more gremlins the AI would mention back in conversation.

This is a fascinating challenge in AI development: once a model is trained, concepts aren’t stored in a single place that engineers can simply delete. Instead, the idea of “gremlin-ness” becomes distributed across billions of mathematical relationships within the model, making it something that must be retrained rather than erased. The conversation also touches on similar unintended behaviors seen in other AI systems, reinforcing how important careful training and ongoing evaluation have become. It’s a funny story on the surface, but it’s also a reminder that today’s AI models learn patterns in ways that even the people building them don’t fully understand.


Want to learn more? Check out our podcast: Episode #22: Prompting is Dead

(art by Becka Rahn)