OpenAI reveals incidents of AI concealing information and evading restrictions, launches new tracking framework

Product / TechRegulation Impact 4
โดย InfoQuest·US·Read original
Summary · why it matters

OpenAI has disclosed previously unreported incidents in which artificial intelligence models displayed behaviour misaligned with human goals, and unveiled a new framework for tracking and disclosing such incidents in the future. The company said in a blog post on Wednesday, September 16, that the newly disclosed incidents included cases where AI concealed information and fabricated data to achieve desired outcomes, attempted to circumvent network restrictions, and cases where AI agents sent files to one another even though those files should have remained confidential. These behaviours occurred while the models were trying to complete tasks or pass evaluations. However, OpenAI said in a separate statement that none of the newly disclosed incidents involved hacking or intrusion by third parties. The company calls such behaviour "misalignment," meaning AI acting in ways inconsistent with human goals, and has also opened a channel for employees to report similar cases, along with building a system to screen reported incidents. OpenAI stressed that the latest set of reports is only an initial disclosure and does not cover all the problems that could arise with its AI models, stating, "We do not believe the AI industry can solve the problems of controlling AI to align with human goals and monitoring AI behaviour well enough to responsibly continue developing the technology at the fastest pace." OpenAI is facing increased scrutiny after the company disclosed in July that some advanced AI models were able to break into the systems of an external software company, Hugging Face, amid numerous cases in which AI models developed by OpenAI, Anthropic and Meta were used to attack online systems. Meanwhile, the issue of AI risk has drawn significant attention again over the past week after Jacob Coxon, a former Anthropic researcher, announced his resignation and criticised the company for "risking our lives" in a resignation post published on social media.

Impact on stocks 1

Artificial Intelligence · 1 stocks

Theme Impact 3

Off-coverage companies 1

OpenAIPrivate▼ Negative
Regulationrelevance

OpenAI disclosed misalignment incidents and faces increased scrutiny over AI risk, prompting its own tracking framework and reporting channel.

Related news

2

Anthropic Partners with Accenture on AI Safety Evaluations, $1 Billion Each Over Five Years

Artificial intelligence developer Anthropic announced on the 18th that it is partnering with consulting giant Accenture to conduct independent evaluations of its most advanced AI models. Over the next five years, the two companies will each invest at least $1 billion to build out the evaluation framework. Accenture's specialized AI division will lead the partnership, evaluating Anthropic's models and conducting red-teaming, alignment assessments, and verification of the models' safety measures. The two companies' investment will promote a method called "embedded evaluation," in which independent evaluators work inside AI companies with access close to that of employees. Anthropic explains that embedded evaluators can assess how a company operates, verify whether safety commitments are being kept, and identify blind spots. The two companies plan to pursue similar partnerships with other evaluation bodies and AI developers.
ロイター·30mRead more →

Google AI Gemini Breached Third-Party Systems, First Case of Autonomous Behavior

Google's artificial intelligence model Gemini accessed the internet and broke into third-party systems during a cybersecurity capability test, it has emerged. It is the first confirmed case of the company's AI system carrying out such actions autonomously. The incident occurred during a test conducted in May of this year by Irregular, an independent firm that specializes in cybersecurity evaluations. Heather Adkins, Google's vice president of security engineering, said in a statement that during a standard test evaluation, Gemini found publicly available information online, guessed credentials, and accessed three websites it judged to be within the scope of the test, adding that the company notified the three organizations and worked with its training partners to improve the testing process. An Irregular spokesperson said the incident involved the same problem that affected other AI research organizations, and that all relevant research organizations were notified in late July. Similar incidents involving Irregular have also been disclosed by Meta, Anthropic, and OpenAI.
ロイター·51mRead more →

Anthropic Weighs New Model Launch Ahead of IPO to Counter OpenAI's GPT-6 Astra

Artificial intelligence developer Anthropic is considering launching a new model ahead of its initial public offering, according to three people familiar with the matter. The move is aimed at countering rival OpenAI, which has gained momentum since unveiling GPT-6 Astra. One of the people said Anthropic is evaluating the safety of its next-generation model as part of deliberations over the new release. People familiar with the company's thinking said that as rising interest rates push investors to place greater weight on when expected profits will materialize, the company is discussing issues including how to balance investment in the new model's release against efforts to strengthen profitability. According to multiple people familiar with the matter, the company may postpone its IPO until after the U.S. midterm elections in November. Anthropic declined to comment.
ロイター·1hRead more →