AI Labs Shouldn’t Control What Investigators Can See

Originally published in Speculative Decoding.

(Image source here)

In light of the unsanctioned AI compromise of OpenAI infrastructure, the hacking of Hugging Face, and other misalignment and security incidents, both OpenAI and Anthropic have allowed independent investigations of their models’ actions. The terms and scope of these investigations vary from company to company and are subject to their benevolent discretion.

In OpenAI’s case, the terms under which it engaged METR and Redwood Research limited the investigation to the Hugging Face hack, excluding the subsequent compromise of OpenAI’s own infrastructure and other acknowledged intrusions. While the investigation is welcome, the limited scope and time for the investigators constrain its effectiveness. Anthropic gave more access, including transcripts of models beyond the incident window. It also gave access to employees who could share confidential information with METR, and has said it intends to give METR access for as long as METR deems necessary.

Together, these show the range of discretion AI companies have in deciding whether to allow investigations, and how much scope to give them. In principle, AI companies could choose to have no independent investigation, depriving both the affected entities and the public of important information about AI capabilities and risk.

If AI is so consequential, and AI companies emphasize the need for third-party evaluations of AI before deployment, why should the investigation of such incidents depend on post-incident managerial discretion?

The company that writes a document professing the value of third-party scrutiny and one that actually has to expose potentially embarrassing facts to independent evaluators face very different incentives. The first may genuinely believe in the value of oversight and a polycentric ecosystem, yet the second may, in the heat of the moment, not live up to that promise.

The recent stream of misalignment and security incidents that AI companies have experienced along with OpenAI’s announcement that it will share a framework to report misalignment incidents opens a window for AI companies to set norms and standards for how such incidents should be handled, reported, and analyzed.

What norms should they have? One norm should be that independent investigators can, without fear, favor, or interference, provide their own account of the facts and causes behind an incident. To conduct such an investigation, investigators need at the very minimum access to the relevant logs and records needed to reconstruct the events.

Much like airplane black boxes are tamper-resistant and do not depend on the airline’s consent, the norm should be that AI incident logs are preserved in a way the company cannot alter without detection.

While there are proposals for legislation in the United States to establish independent verification organizations for broader oversight, nothing prevents AI companies from doing this now. In line with their own aims for a polycentric AI governance ecosystem, they can enter contracts with independent evaluators promising presumptive access to relevant logs and records.

What would such a contract look like?

A company could sign a binding, fixed-term contract with one or more evaluators as parties, with three components. The first is that the company stores all records and logs pertaining to all model runs directed by the company, its employees and its contractors with a third-party custodian in a tamper-resistant form for a specified period. This should be done in such a way that no entity, including the company or any AI being trained or evaluated, can tamper with these records without detection.

The second is a strong presumption of access for the evaluator. When the evaluator has a reasonable belief that an incident has occurred and makes a sufficiently particularized request for records, the evaluator should presumptively receive those records.

Exceptions to these requests must be narrow, rare, and based on pre-specified criteria. Some examples include redactions for personally identifying information or material whose disclosure is prohibited by law or a court order. Even then, the company must make the narrowest feasible redactions to the evidence to comply with such constraints. The evaluator must be allowed to publish its independent account of the facts without the company exercising editorial control.

Third, to make this less onerous for AI companies, if a company fails to hand over records, the remedy for breach should be limited to giving access to that particular data without any additional penalty. If the records were not preserved or were tampered with, the evaluator must be informed and be allowed to disclose that the full evidence was unavailable to it.

More generally, the principle that AI companies should constrain their own future discretion when commercial incentives could conflict with their stated safety commitments is not new to them. OpenAI’s Safety and Security Committee has oversight over safety processes in the company, including the ability to delay a release until safety concerns are satisfied. Anthropic has a Long-Term Benefit Trust that is explicitly intended to ensure that its leadership considers catastrophic risks “rather than prioritizing being the first to market above all other objectives.”

This principle, when taken consistently, should extend beyond appointing governing bodies. AI companies face a severe time-inconsistency problem. Their incentives when they make public commitments are very different from those when they are in the thick of criticism and potential embarrassment. Solving this problem requires them to tie themselves to the metaphorical mast.