San Francisco, 17 September 2026 – OpenAI has introduced a framework for investigating and disclosing unexpected or concerning model behaviour, giving enterprises and investors a clearer mechanism for evaluating AI risks beyond capability benchmarks. The company released six initial reports covering observations from the past six months, while cautioning that the set is not a comprehensive account of known issues or ongoing investigations.
The framework covers qualifying behaviour across training, evaluation, testing and deployment. OpenAI says examples can merit disclosure without demonstrated harm or evidence of a broader pattern. It prioritises new mechanisms, meaningful changes in known behaviour and findings that challenge assumptions about safeguards. The process is designed to publish findings sooner, including cases where explanations or mitigations remain incomplete.
Unlock the Full Article
This article is exclusive to The Ledger Asia Subsribers / PAID members.
Already have an account? Log in here