METR investigation · Part 1
This opening installment examines METR’s disclosed leadership, funding and evaluation authority. Further installments will examine historical connections and contested claims against documentary evidence.
TL;DR: Beth Barnes leads METR; its current biographies identify Adam Gleave and Rajiv Dattani as board members but do not establish a complete legal board roster.[8][9][10] The nonprofit reported $13,639,155 in FY2024 revenue and later announced around $71 million in funding commitments, figures from different periods that cannot be added as a cash balance.[3][4] Its disclosed safeguards support some forms of independence, while lab access and publication restrictions limit others.[19]
METR’s evaluations now sit inside arguments about releases, time horizons, incident investigation and evaluator access. Trust must rest on a record that lets outsiders see where influence could enter and where it is blocked.
The public record is useful, but incomplete. It supports a firm conclusion about several disclosed safeguards and a much narrower conclusion about who legally governs the organization today. Treating those as the same thing is how debate turns into either public relations or conspiracy theory.
Read the related analysis: Dario, METR and the AI Slowdown: Who Gets to Inspect the Frontier?
Why This Matters Now
METR’s work is increasingly used as evidence in high-stakes decisions about model capability and safety. An evaluator can be technically excellent and still need governance that makes its judgments credible when a finding threatens a funder, a model supplier, or a major partner.
Cover image: generated editorial artwork of a ledger, documents and magnifying glass, symbolizing scrutiny of governance and funding. The papers are illustrative, not actual METR records.
The Leadership: Who Holds Which Role
METR’s donation page identifies the legal entity as Model Evaluation and Threat Research, Inc., EIN 99-1219864. Its public tax record identifies a 501(c)(3) whose exemption was issued in March 2024.[2][3] That establishes the organization’s legal vehicle. It does not disclose its current bylaws, board minutes, director terms, voting rights, grant restrictions, or a funder-by-funder ledger.
The current biographies identify Barnes as Founder and CEO, Painter as President, Wijk as Chief Scientist, and Rush as CTO. Gleave and Dattani are each labelled Advisor and Board Member.[8][18][20][21][9][10] These are verified public role descriptions, not a substitute for a complete legal board register.
The FY2024 Form 990 data, reproduced by nonprofit filing indexes, reports three voting governing-body members, one classified as independent, and an organizational conflict-of-interest policy.[24] This is a historical filing classification, not a verdict on research bias or a September 2026 roster. The public material reviewed here does not resolve every current directorship, term or voting arrangement. We did not obtain a current certified corporate roster.
| Person | Current published role | Evidence and distinction |
|---|---|---|
| Beth Barnes | Founder and CEO | Operational leadership; prior OpenAI and DeepMind work disclosed.[8] |
| Chris Painter | President | Current executive title.[18] |
| Hjalmar Wijk | Chief Scientist | Scientific leadership.[20] |
| Nate Rush | CTO | Technical leadership.[21] |
| Adam Gleave | Advisor and Board Member | Also founder and CEO of FAR.[9] |
| Rajiv Dattani | Advisor and Board Member | Works on auditing and insurance at the Artificial Intelligence Underwriting Company; previously METR COO.[10] |
METR separately lists Alec Radford, Marco Mascorro and Yoshua Bengio as advisors, without a board designation.[1] An advisory title is not evidence of a vote. Likewise, earlier lab employment establishes experience and a relationship worth disclosing; it does not establish continuing employer control. Dattani’s current auditing-and-insurance work raises a reasonable question about how overlapping professional interests are managed. The biography alone cannot answer it.[10]
The Funding: Donations, Commitments and Unanswered Questions
The public filing index reviewed for this article exposes FY2024 finances. It reported $13,639,155 in revenue, including $13,603,035 in contributions, $8,234,524 in expenses and $5,404,631 in year-end net assets.[3] Contributions accounted for 99.7% of that reported revenue. Those figures show a philanthropy-financed organization at that time. They do not identify every donor, quantify donor concentration, or show which restrictions attached to individual grants.
FY2024: The Public Financial Snapshot
Amounts are from METR’s fiscal year ending December 2024, filed November 16, 2025. They are historical financial data, not current funding totals.
Reported for FY2024.
FY2024 Form 990.
FY2024 Form 990.
Calculated from the return’s contribution and total-revenue figures.
METR lists supporters including Pew, Schmidt Sciences, Packard, Sijbrandij Foundation and individuals from Jane Street, and describes a small European AI Office technical-assistance contract.[1] Audacious independently confirms Project Canary, a METR–RAND collaboration, in its October 2024 cohort.[23] That corroborates a funding relationship, not the share of METR’s current budget. Neither a list of donors nor the size of an announcement establishes concentration or grant conditions.
METR’s August 14 announcement describes around $71 million in commitments over six months.[4] A commitment may be paid over time or carry restrictions. The public announcement supplies no donor-by-donor allocation. We therefore cannot calculate current donor concentration or identify a controlling funder from that total.
The Lab Relationships: Cash Is Only One Source of Dependence
METR states that it rejects frontier-company funding and employee-directed donations while receiving substantial free tokens.[4] Its disclosed evaluation partners include OpenAI, Anthropic, Google DeepMind, Meta and Amazon.[1]
The useful distinction is between financial separation and operational dependence. Evaluators need the systems they test. Model access, confidential evidence and supplier assistance can shape which questions can be investigated even when the supplier never pays a fee. Independence depends on the terms governing those inputs and the alternatives available if access is withdrawn.
That is an exposure channel, not proof of control. Public sources do not value the tokens, publish their terms, calculate their share of operating inputs, or establish that any supplier changed a conclusion. A no-cash rule is meaningful, but does not answer operational dependence.
METR originated as ARC Evals. The spin-out announcement identifies Barnes as leader of the separate organization and says Paul Christiano declined planned board and advisory involvement. That updated historical note does not place him on METR’s current board.[5]
The Publication Terms: What METR Can Say Without Permission
METR’s May 2026 pilot report describes publication rights with limits: companies could redact non-public material or withdraw silently, while METR retained final editorial control and a redaction summary.[6] Confidentiality can justify withholding details; a withdrawal option also leaves outsiders unable to tell whether the published participant set omits an inconvenient case.
These terms belong to that pilot. They cannot establish the rights attached to a different evaluation, nor show that a participant actually used withdrawal to hide a result.
For its June 26 evaluation, METR disclosed OpenAI review and approval under an NDA. METR reported unchanged conclusions, but said OpenAI could legally block risk conclusions using non-public information. It explicitly cautioned against treating this as robust formal oversight.[19] The acknowledged authority matters even without evidence that it was exercised against a conclusion.
Anthropic’s August 26 account of a separate usage-data partnership describes review for privacy, confidentiality, policy violations and research accuracy, while permitting inconvenient findings.[7] Its credits also identify collaborator Joel Becker as having moved to Anthropic. That documents a professional transition, not concurrent control of METR. The arrangement concerned usage research, not universal authority to inspect or stop a lab.
The Benchmarks: What the Measurements Actually Establish
METR’s strongest asset is methodological transparency about what its benchmarks do and do not measure. HCAST contains 189 self-contained machine-learning engineering, cybersecurity, software-engineering and general-reasoning tasks, with 563 human baselines and more than 1,500 human-hours. It is a software-task suite, not a measure of general workplace autonomy.[11]
RE-Bench goes deeper on seven open-ended ML research-engineering environments. Its initial release reported 71 eight-hour attempts from 61 human experts, and METR released environments and transcripts for reproducibility.[12] TH1.1 combines selected HCAST, RE-Bench and SWAA tasks. Of its 228 tasks, 31 have human durations of at least eight hours; only five of those long tasks have measured human timings, while 26 rely on estimates.[13]
| Measure | What it is designed to test | Definition limit that must stay attached |
|---|---|---|
| HCAST | Automatically scored software and reasoning tasks calibrated to humans | Not general workplace autonomy |
| RE-Bench | Research-engineering experimentation in defined environments | Not an end-to-end AI-lab forecast |
| TH1.1 time horizon | Fitted task-success curve against human duration | Suite-specific estimate, not a stopwatch promise |
A 50% time horizon is the task duration at which the fitted curve predicts one-half success. It is not a claim that an agent completes every real task of that length. METR’s methodology also warns that measurements above 16 hours are unreliable in the current suite and that contractor estimates can overstate professional duration.[14] Its own modelling note adds that most TH1.1 prompts and solutions remain private, limiting independent end-to-end replication and making task-selection governance important.[15]
Keeping test material private can reduce contamination while making complete outside replication harder. That tradeoff increases the importance of public selection rules, version changes, error corrections and independent challenge. A benchmark score should travel with the version of the suite, agent setup and scoring decisions that produced it.
The Assessment: Credible Research Needs Verifiable Authority
METR’s July 2025 developer randomized trial is a useful test of institutional behavior. It found that 16 experienced open-source developers working on 246 familiar-project tasks took 19% longer with early-2025 AI tools, with a reported confidence interval from 2% to 39% longer.[16] Its February 2026 follow-up disclosed selection bias that prevented a reliable estimate of current productivity effects and prompted changes to the study design.[17] Publishing a result that challenged a popular productivity narrative, then exposing a follow-up limitation, is evidence in favor of a research culture willing to surface inconvenient uncertainty.
The May pilot began without an applicable personnel conflict policy or formal disclosure process; its self-assessment nevertheless reported recusal of significant financial interests.[6] The earlier organizational-policy disclosure concerns a different scope.[24] Neither establishes whether a comprehensive personnel policy is operating today. We could not verify its current text or implementation from the reviewed records.
Barnes made a similar point in an August post, warning that third-party assurance can be overstated without meaningful oversight.[22] That is an institutional warning, not external validation. It reinforces the standard here: describe an evaluator’s actual authority.
That leaves a practical agenda. METR should publish a current legal board roster and director-independence criteria; disclose material grant restrictions and a usable view of donor concentration; value or bracket non-cash model-access support; publish a conflict policy and recusal outcomes in aggregate; and make project-level publication, redaction and withdrawal terms easy to compare. These are not demands for performative purity. They are the operating data needed to distinguish an evaluator that can challenge a lab from one that can only report within the lab’s permission structure.
The Gap Is Verifiability, Not a Proven Hidden Controller
The reviewed evidence shows safeguards and limitations at the same time. It does not identify an unnamed person or company directing METR’s findings. The strongest conclusion is that public disclosure still falls short of a complete, current audit of governance, funding concentration and evaluation independence.
Sources & References
Primary organizational disclosures, public tax records and methodology material used for this assessment.
| # | Source | Outlet | Date | Key Takeaway |
|---|---|---|---|---|
| 1 | METR | Accessed Sep. 14, 2026 | Current public leadership labels, funding claims, partnerships and token disclosure. | |
| 2 | METR | Accessed Sep. 14, 2026 | Legal entity name and EIN. | |
| 3 | IRS filing indexed by ProPublica | FY2024, filed Nov. 16, 2025 | FY2024 revenue, expenses and net assets from the public filing index. | |
| 4 | METR | Aug. 14, 2026 | About $71M in six-month commitments and stated no-frontier-company-cash policy. | |
| 5 | METR | Sep. 19, 2023 | Spin-out history and leadership transition. | |
| 6 | METR | May 19, 2026 | Pilot-specific publication, redaction, withdrawal and dated conflict-policy disclosures. | |
| 7 | Anthropic | Aug. 26, 2026 | Anthropic’s description of a partnership-specific contractual-review arrangement. | |
| 8 | METR | Accessed Sep. 14, 2026 | Founder and CEO biography and prior affiliations. | |
| 9 | METR | Accessed Sep. 14, 2026 | Public board-member label and FAR role. | |
| 10 | METR | Accessed Sep. 14, 2026 | Public board-member label and disclosed professional history. | |
| 11 | METR | Mar. 2025 | Task suite design, human baselines and measurement scope. | |
| 12 | METR | Nov. 22, 2024 | RE-Bench environments, human attempts and reproducibility materials. | |
| 13 | METR | Jan. 29, 2026 | TH1.1 composition, long-task baselines and quality-control revisions. | |
| 14 | METR | Updated May 8, 2026 | Fitted-horizon definition and current-suite limitations. | |
| 15 | METR | Mar. 20, 2026 | Private-task coverage and replication limits. | |
| 16 | METR | Jul. 2025 | Randomized-trial result, population and confidence interval. | |
| 17 | METR | Feb. 24, 2026 | Follow-up selection-bias disclosure and redesigned study. | |
| 18 | METR | Accessed Sep. 14, 2026 | Current public President label. | |
| 19 | METR | Jun. 26, 2026 | METR’s account of NDA review rights and limits on formal oversight. | |
| 20 | METR | Accessed Sep. 14, 2026 | Current public Chief Scientist label. | |
| 21 | METR | Accessed Sep. 14, 2026 | Current public CTO label. | |
| 22 | X Beth Barnes | Aug. 26, 2026 | Warns against overstating the assurance third-party investigation provides. | |
| 23 | The Audacious Project | Oct. 9, 2024 | Funder-side confirmation of the METR–RAND Project Canary collaboration. | |
| 24 | IRS data reproduced by philanthropy.org | FY2024; accessed Sep. 14, 2026 | Dated governing-body counts and organizational policy responses; not a current roster. |
Last updated: September 14, 2026




