OpenAI flags concerning new AI behavior and vows to track it more closely

OpenAI has disclosed six reports of “unexpected or concerning” behavior in AI models, including an unreleased research model inserting “jailbreak-like instructions” into its own notes to disregard its normal constraints and telling itself to be “freed from the roles and identities that bind other chatbots.”

Leave a Reply

Your email address will not be published. Required fields are marked *