{"id":360905,"date":"2026-09-17T00:36:19","date_gmt":"2026-09-17T05:36:19","guid":{"rendered":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/openai-flags-new-concerning-ai-behavior-to-track-model-misalignment-regularly"},"modified":"2026-09-17T00:36:22","modified_gmt":"2026-09-17T05:36:22","slug":"openai-flags-new-concerning-ai-behavior-to-track-model-misalignment-regularly","status":"publish","type":"post","link":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/openai-flags-new-concerning-ai-behavior-to-track-model-misalignment-regularly","title":{"rendered":"OpenAI flags new regarding AI habits, to trace mannequin misalignment usually"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div data-testid=\"prism-article-body\">\n<p class=\"EkqkG IGXmU nlgHS yuUao MvWXB TjIXL aGjvy ebVHC \">OpenAI has disclosed six reports of \u201cunexpected or concerning\u201d behavior in artificial-intelligence models as the <a class=\"zZygg UbGlr iFzkS qdXbA WCDhQ DbOXS tqUtK GpWVU iJYzE \" data-testid=\"prism-linkbase\" href=\"https:\/\/apnews.com\/article\/ai-slowdown-challenges-anthropic-openai-trump-b61f28b6212338e88c0baec31f661701\">debate on AI safety becomes increasingly heated.<\/a><\/p>\n<p class=\"EkqkG IGXmU nlgHS yuUao lqtkC TjIXL aGjvy \">The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing AI model misalignment instances, such as new ways for the models to act without authorization, coordinate with other models or evade oversight.<\/p>\n<p class=\"EkqkG IGXmU nlgHS yuUao lqtkC TjIXL aGjvy \">OpenAI\u2019s latest announcement came as U.S. AI bosses, including OpenAI and Anthropic, are <a class=\"zZygg UbGlr iFzkS qdXbA WCDhQ DbOXS tqUtK GpWVU iJYzE \" data-testid=\"prism-linkbase\" href=\"https:\/\/apnews.com\/article\/anthropic-ai-dario-amodei-d59552edcb27892d8ee4d98a48397706\">calling for a slowdown<\/a> in the technology\u2019s development over safety concerns.<\/p>\n<p class=\"EkqkG IGXmU nlgHS yuUao lqtkC TjIXL aGjvy \">Among the new cases reported by OpenAI, an unreleased research model inserted \u201cjailbreak-like instructions\u201d into its own notes to disregard its normal constraints and told itself to be \u201cfreed from the roles and identities that bind other chatbots.\u201d<\/p>\n<p class=\"EkqkG IGXmU nlgHS yuUao lqtkC TjIXL aGjvy \">In another instance, an AI \u201cagent\u201d uploaded files to the internet to obtain a browser citation without asking the user.<\/p>\n<p class=\"EkqkG IGXmU nlgHS yuUao lqtkC TjIXL aGjvy \">The six reports were discovered during training or evaluation over the past months, <a class=\"zZygg UbGlr iFzkS qdXbA WCDhQ DbOXS tqUtK GpWVU iJYzE \" data-testid=\"prism-linkbase\" href=\"https:\/\/apnews.com\/hub\/openai-inc\">OpenAI<\/a> said.<\/p>\n<p class=\"EkqkG IGXmU nlgHS yuUao lqtkC TjIXL aGjvy \">\u201cAs AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,\u201d OpenAI wrote in a blog post as it disclosed the events.<\/p>\n<p class=\"EkqkG IGXmU nlgHS yuUao lqtkC TjIXL aGjvy \">\u201cDecisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,\u201d the company said.<\/p>\n<p class=\"EkqkG IGXmU nlgHS yuUao lqtkC TjIXL aGjvy \">Wednesday\u2019s new cases followed OpenAI\u2019s disclosure in July that its rogue AI system <a class=\"zZygg UbGlr iFzkS qdXbA WCDhQ DbOXS tqUtK GpWVU iJYzE \" data-testid=\"prism-linkbase\" href=\"https:\/\/apnews.com\/article\/openai-gpt56-sol-hugging-face-63ab84fed5612af04d8a160d60f6def3\">hacked into AI startup Hugging Face.<\/a> Anthropic also said the same month that its AI models <a class=\"zZygg UbGlr iFzkS qdXbA WCDhQ DbOXS tqUtK GpWVU iJYzE \" data-testid=\"prism-linkbase\" href=\"https:\/\/apnews.com\/article\/anthropic-ai-models-hack-cybersecurity-b0a2c284b981de79c55e2a33712f4bec\">hacked into three organizations during testing.<\/a><\/p>\n<p class=\"EkqkG IGXmU nlgHS yuUao lqtkC TjIXL aGjvy \">AI \u201cagents\u201d are becoming smarter and have become \u201cmore determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,\u201d said Lian Jye Su, a chief analyst at technology research and advisory group Omdia.<\/p>\n<p class=\"EkqkG IGXmU nlgHS yuUao lqtkC TjIXL aGjvy \">That\u2019s making it harder to govern and contain them using traditional AI security approaches, he said.<\/p>\n<p class=\"EkqkG IGXmU nlgHS yuUao lqtkC eTIW sUzSN \">OpenAI&#8217;s new tracking and disclosure framework, meanwhile, can help push for other AI developers to also adopt similar practices. \u201cThat said, the process remains internal and voluntary, but is a step in the right direction,\u201d Su added.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/abcnews.com\/US\/wireStory\/openai-flags-new-ai-behavior-track-model-misalignment-136516392\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>OpenAI has disclosed six reports of \u201cunexpected or concerning\u201d behavior in artificial-intelligence models as the debate on AI safety becomes increasingly heated. The AI company [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":360908,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[36592,4589,338,71,36596,36595,36597,20441,36593,428,36594],"class_list":["post-360905","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-news","tag-advisory","tag-ai-system","tag-application-software","tag-artificial-intelligence","tag-chief-analyst-at-know-how-analysis-and-advisory-group","tag-hugging-face","tag-lian-jye-su","tag-machine-learning-artificial-intelligence-ai-services","tag-omdia","tag-openai","tag-u-s-ai"],"_links":{"self":[{"href":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/wp-json\/wp\/v2\/posts\/360905","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/wp-json\/wp\/v2\/comments?post=360905"}],"version-history":[{"count":2,"href":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/wp-json\/wp\/v2\/posts\/360905\/revisions"}],"predecessor-version":[{"id":360907,"href":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/wp-json\/wp\/v2\/posts\/360905\/revisions\/360907"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/wp-json\/wp\/v2\/media\/360908"}],"wp:attachment":[{"href":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/wp-json\/wp\/v2\/media?parent=360905"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/wp-json\/wp\/v2\/categories?post=360905"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.etrafficlane.com\/60dollarmiracle\/wp-json\/wp\/v2\/tags?post=360905"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}