OpenAI has uncovered even more alarming examples of its AI models behaving in unexpected and potentially deceptive ways, adding to mounting concerns over how increasingly autonomous systems operate ...
OpenAI announced new guidelines for tracking and reporting AI model behavior Thursday while flagging six more incidents of concerning behavior.
OpenAI has disclosed six new cases of model misbehavior and offered a framework for disclosing future instances, as the debate over AI model safety intensifies.
OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable ...
OpenAI has disclosed six new incidents of “unexpected or concerning” behavior by its artificial intelligence models.
OpenAI has disclosed six reports on unexpected or concerning behavior in artificial-intelligence models. This includes models acting without authorization or evading oversight.
Some results have been hidden because they may be inaccessible to you
Show inaccessible results