OpenAI flags six new cases of ‘unexpected or concerning’ AI behaviour
Original article ↗ Paywalled source — limited preview availableWhat we measured
Outlets covering
3
Lean bands spanned
2
Widest spread
Center → Center-Right
Verdict
Unverifiable
B.I.A.S. ANALYSIS
CENTER RIGHT
Where the 3 outlets covering this sit
Center
67%
Center-Right
33%
Science Tech framing
Center majority
3-outlet cluster
Broad pickup
Signal breakdown
Heuristic (v1/v3)
0.00 · CENTER
ML v2 (DistilBERT)
0.265 · RIGHT
Ensemble
0.265 · CENTER RIGHT
🏦 Source Intelligence
🏢 Corporate
· Globe and Mail Inc. (Woodbridge)
Rolling outlet bias
CENTER LEFT
avg -0.251
14,065 articles tracked all-time
7-day bias trend
LcenterR
V.E.R.I.F.Y. has fact-checked this article.
Subscribe to see claim-by-claim verdicts and reasoning.
Subscribe to see claim-by-claim verdicts and reasoning.
🔍 Intelligence Feed
Cross-Watch · Gov · Parliament · Legal · Civic
📄 Related Gov Tenders
Via Gov Watch · CanadaBuys + PSPC tenders
🏛 Related Parliament Votes
Via Civic Watch · OpenParliament.ca
🏙 Related Municipal Events
Via Civic Watch · City council, bylaws & permits
Article Excerpt
OpenAI flags six new cases of ‘unexpected or concerning’ AI behaviour
CHAN HO-HIM
THE ASSOCIATED PRESS
PUBLISHED YESTERDAY
Open this photo in gallery:
Among the new cases reported by OpenAI, an unreleased research model inserted 'jailbreak-like instructions' into its own notes to disregard its normal constraints and told itself to be 'freed from the roles and identities that bind other chatbots.'
MICHAEL DWYER/THE ASSOCIATED PRESS
COMMENTS
SHARE
SAVE FOR LATER
Listen to this article
Learn more about audio
Log in or create a free account to listen to this article.
OpenAI has disclosed six reports of “unexpected or concerning” behaviour in artificial-intelligence models as the debate on AI safety becomes increasingly heated.
The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing instances of what it called “misalignment,” including cases where AI models acted without authorization, co-ordinated with other models or evaded oversight.
OpenAI’s latest announcement came as U.S. AI bosses, including the leaders of OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns.
Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots.”
In another instance, an AI “agent” used computer code to come up with the answer to a question, but, in order to have an online source to cite, it uploaded a file to the public internet without asking the user.
The AI company said it was introducing a new framework for tracking, probing and disclosing instances of what it called 'misalignment.'
THE ASSOCIATED PRESS
During training of an AI model called 5.6-sol, the model instructed itself to invent missing data, and an agent wrote a message to remind itself to hide mismatched information.
This kind of deceitful behaviour has underpinned a flare-up of recent concerns about AI evading human control. But that outcome is not surprising to some AI researchers, said Matt Fredrikson, an associate professor at Carnegie Mellon University and the CEO of Gray Swan AI.
“At the risk of anthropomorphizing model behaviour, you can almost think of them as knowing that they’re going to be graded,” Fredrikson said. “If they know that they cheated – took shortcuts, didn’t really do it in the way that it was…
Read full article at The Globe and Mail ↗
How we scored this article
WTF uses a two-tier system: every article gets a heuristic bias score from keyword analysis, and priority articles (high overlap across 3+ outlets or strong heuristic signal) get full LLM analysis from B.I.A.S. and V.E.R.I.F.Y.
Cite this analysis
2 outlets covered this story
Side-by-side comparison · bias by outlet · framing analysis