METR review · Part 2
TL;DR: METR’s people span frontier labs, academia, government and AI-risk philanthropy; those professional histories identify expertise and possible conflicts to examine, not a hidden chain of command. Its FY2024 filing reports a $4,501,424 gift, grant or capital contribution from ARC, while ARC separately disclosed a $1.25 million FTX Foundation grant received in 2022 and later returned less costs.[10][8] The e/acc authors around Guillaume Verdon advocate a distinct philosophy of acceleration; the reviewed evidence establishes no e/acc governance link to METR.[11][15]
The word “independence” becomes vague the moment a research organization is reduced to a list of former employers. It is a poor substitute for asking who decides what counts as a danger, who selects the tasks, who has access to the model, who reads the result before publication and who can explain a conflict when it matters.
That is the question Part 2 investigates. METR is influential because an evaluation result can affect a frontier lab’s public safety case, an investor’s narrative and a policymaker’s sense of urgency. Its people therefore matter. But professional history is evidence of experience and network proximity, not evidence of command. The distinction is the entire story.
Read Part 1: Leadership, Funding and Independence · Read the related analysis of Dario Amodei’s oversight proposal
Why This Matters Now
The useful audit trail runs from a risk hypothesis to a published conclusion: threat modelling, task design, evaluation infrastructure, model access, analysis, partner terms and disclosure. Public biographies reveal some of those functions. They do not reveal a complete approval chain, a current legal board roster, compensation, personal conflicts, or every contractual restriction.
The people behind the evaluations
METR
Who frames the risks, builds the tests and runs the organization?
Six people. Three areas of work.
Direction & operations
Risk & threat models
Tests & external use
Selected profiles, grouped by work. Click a name for its source. These are not reporting lines.
Origins & money: separate records
ARC
2022
ARC Evals
Incubated at ARC
METR
Spin-out announced
09 / 2023
FTX Foundation → ARC · 2022
$1.25 million
Grant. ARC later reported returning it less costs.
ARC → METR · FY2024
$4,501,424
Filed as gifts, grants or capital contributions. Donor origins untraced.
These records do not establish an FTX → METR transfer.
e/acc: the separate authorship record
Beff Jezos + Bayeslord
Beff Jezos = Guillaume Verdon
e/acc principles
Repost · October 31, 2022
No verified e/acc governance link to METR in the reviewed sources.
Read the connections and their sources
ARC / FTX: ARC reports receiving $1.25 million in 2022 and, in a 2024 update, having returned it less legal and administrative expenses. The net amount and transfer date are not disclosed.
ARC / ARC Evals / METR: METR describes ARC incubating ARC Evals after hiring Barnes in 2022; the spin-out announcement is dated September 19, 2023.
FY2024 Schedule R: METR reports $4,501,424 in the category for gifts, grants or capital contributions from ARC. This record does not identify those funds as FTX money.
Staff biographies document selected prior work and current responsibilities: Beth Barnes; Hjalmar Wijk; Ajeya Cotra; Megan Kinniment; Kit Harris; Stephanie Palla. FHI means Future of Humanity Institute; CLR means Centre on Long-Term Risk. No continuing payment or reporting relationship is implied.
e/acc: The October 31, 2022 repost credits Beff Jezos and Bayeslord. Verdon publicly discusses being Beff in his 2023 interview. No verified e/acc governance connection to METR was found in the sources reviewed. Interview transcript.
Cover: generated editorial artwork. The flywheel and measuring instrument symbolize acceleration and evaluation; they do not depict METR equipment or a measured result.
The Question Is Authority: A Team Page Is Not a Control Map
METR’s public About page separates leadership, technical staff, policy staff, operations, advisors, research collaborators and specialist partners. It identifies Beth Barnes as Founder and CEO, Chris Painter as President, Hjalmar Wijk as Chief Scientist and Nate Rush as CTO.[1] Those are the organization’s current public titles. They establish an executive layer, not every legal or practical power attached to it.
The page labels Adam Gleave and Rajiv Dattani “Advisor and Board Member.” It lists Alec Radford, Marco Mascorro and Yoshua Bengio as advisors, without that board label.[1] It lists Khalid Mahamud as a research contractor and Alexander Barry of Epoch AI as a research collaborator. Those categories matter. An advisor is not automatically a director, a collaborator is not staff, and a former executive is not automatically a current decision-maker.
The public record has a second, narrower governance signal. METR’s FY2024 Form 990 reports three voting governing-body members and one independent member for that filing period.[3] A dated tax return is not a September 2026 legal roster. Nor does a website title disclose voting rights, director terms, committee mandates or the difference between board oversight and day-to-day research control. The honest conclusion is modest: the public material shows some named roles and historical governance classifications, but it does not make METR’s complete present authority structure auditable.
That gap matters because “who runs METR” contains at least four different questions:
| Question | Publicly visible evidence | What remains unverified |
|---|---|---|
| Who directs operations? | Executive titles and individual biographies | Internal reporting lines and delegated authority |
| Who governs legally? | Two current board labels and a dated Form 990 count | Complete current roster, terms and votes |
| Who shapes research? | Bios, bylines and research roles | Project-specific authorship, review and sign-off |
| Who can constrain publication? | Selected project terms and public disclosures | Terms for every project and every unpublished case |
The real story isn’t a claim that all four questions have the same answer. It is that the public discussion often treats them as interchangeable. They are not. A board member may not choose benchmark tasks. A technical lead may not approve a grant. A lab can supply access without holding a vote. A funder can have restrictions without directing an individual result. Good accountability begins by separating those channels.
The People Who Shape the Research: From Threat Model to Published Claim
METR’s staff biographies reveal a functional division of labor. That is more meaningful than a social graph because it locates where research judgment enters the process.
METR says Barnes oversees the technical team that designs and conducts evaluations of generative models. Her biography also discloses prior professional work at OpenAI and work with DeepMind’s chief scientist on scaling-law forecasting. Barnes’s bio establishes the published responsibility and background. It does not establish current employment, equity, continuing compensation or a relationship that gives either lab control over METR.
Painter’s public remit is outward-facing. His biography says he leads engagement with governments and AI labs on frontier safety, helped shape Frontier Safety Policies and works on scaling third-party risk assessment. Before METR, it lists a Technology and National Security Fellowship at the Department of Defense Joint AI Center and earlier medical and biotechnology machine-learning work. That combination makes Painter central to the institutional translation of evaluation results. It does not tell readers which external counterpart, if any, can change a result.
Wijk’s biography places threat modelling at the center of his role: identifying capabilities that might create catastrophic risk and should be evaluated. His prior work includes AI risk and strategy at the Future of Humanity Institute and Centre on Long-Term Risk. Cotra’s biography similarly identifies threat modelling and loss-of-control risk assessment, while disclosing roles at Coefficient Giving, including leading technical AI safety in 2024 and contributing to AI-giving strategy in 2025. Threat modelling is not a clerical step. It decides which hypothetical failure modes deserve scarce evaluation time. That is why methodological assumptions and conflicts disclosure matter here most.
Kinniment leads METR’s benchmark-creation effort. Her published history includes the Centre on Long-Term Risk, the Future of Humanity Institute and 1DaySooner. Rein’s biography records prior research at NYU, where he created GPQA, and a prior Cohere research-engineering role on semantic-search embeddings. Whitfill works on AI capabilities and is on leave from an MIT economics PhD. These are the people closest to a deceptively consequential layer: turning a broad danger claim into a task, a baseline, a scoring rule and an interpretation.
Rush’s bio identifies him as CTO and says he previously founded and led Mito and worked at the Ethereum Foundation on consensus algorithms. Spiegelmock builds software for researchers conducting model evaluations. Evaluation infrastructure is not neutral plumbing. It determines which agent trajectories are observable, what data survive, how a score is reproduced and whether an outside reviewer can inspect a claim. That is a strong reason to ask for version history and methods, not a reason to imply that a former crypto role controls METR.
The policy and operations functions close the loop. Foster works on evaluation-based AI risk management after work at Finetune, now Prometric, and contributions to EleutherAI. Harris helps external partners understand and use evaluation-based risk management; the biography describes previous AI and biosecurity grant investigations and new work at Longview Philanthropy, plus earlier J.P. Morgan work. Palla oversees finance, people operations and business operations, while Chaturvedi’s bio describes startup-operations and climate-technology experience. Those roles may determine how a grant restriction, a staff conflict or a partner agreement is surfaced. Titles alone do not reveal who has final authority when those functions disagree.
Where influence can enter an evaluation
These are functional exposure points identified from public role descriptions. They are not a finding that any person used influence improperly.
Threat framing
Which capabilities count as consequential enough to test? Publicly associated with threat-modelling work.
Benchmark construction
Which tasks, human baselines, scoring rules and exclusions make up the measured object?
Evaluation infrastructure
Which logs, agent actions, versions and reproducibility records exist?
Partner engagement
Who negotiates access, explains results and receives external feedback?
Operations and conflicts
Who records funding restrictions, employment interests and recusal processes?
Publication
Who can describe, redact, delay or withdraw a result under a specific agreement?
What’s often overlooked is that affiliations are not equally relevant to each step. A past research-lab role matters most when access terms or unexpired financial interests are at issue. A prior philanthropic role matters most when a grant or recusal is at issue. A former academic affiliation matters most when methodology is being designed. No biography, by itself, answers any of those questions. The work of oversight is to request the record that connects a named person to a named decision.
Research Decisions: How People Turn Assumptions Into Results
METR’s public materials document useful pieces of this path. The March 2025 HCAST paper describes 189 self-contained tasks, 563 human baselines and more than 1,500 human-hours; it is expressly a software-task suite rather than a measure of general workplace autonomy.[4] The November 2024 RE-Bench release reports seven research-engineering environments and 71 eight-hour attempts by 61 human experts.[5] The January 29, 2026 TH1.1 release combined suites and reported 228 tasks, including 31 with human durations of at least eight hours. Only five of those longer tasks had measured human timings, while 26 relied on estimates.[6]
Those figures are not a detour from the people investigation. They show why personnel choices matter. A benchmark has a theory of relevance built into its selection rules. A 50% time horizon is the human-expert task duration at which a fitted curve predicts 50% model success, not a promise that a model can autonomously complete every real-world job of a stated length.[7] When much of a suite remains private to resist contamination, external observers cannot independently replay every choice. That makes it more important to know who authored the task policy, maintained versions, approved changes and disclosed uncertainty.
A research claim has six potential decision points
1. Frame the threat
Define the harmful capability or operational risk worth measuring.
2. Build the task
Choose the environment, baseline, scoring rule and exclusion criteria.
3. Obtain access
Set the model, tools, data, confidentiality and partner conditions.
4. Run and interpret
Control configuration, analyze failures and state uncertainty.
5. Review disclosures
Apply conflict, privacy, security and publication provisions.
6. Publish and respond
Release the finding, explain redactions and record what changed.
This is where Part 1’s funding and access findings become operational. METR says it does not accept frontier-company funding or employee-directed donations, while receiving substantial free tokens and working with leading labs on evaluations.[1] The cash rule is a meaningful safeguard. Yet model access, confidential incident data and contractual review can still affect which questions a team can investigate. The existence of this channel is not proof that a lab has altered a conclusion. It is the reason disclosure needs to cover more than cash.
The May 2026 Frontier Risk Report pilot gives a concrete, bounded example. It said companies could redact or anonymize non-public information and could leave silently before final-material approval, while METR retained stated final editorial control over the industry report and could publish certain results based on public model access and aggregate findings.[12] In a separate June evaluation note, METR said OpenAI’s NDA allowed review and approval of protected information, and that OpenAI could legally block risk conclusions based on non-public information. METR reported that conclusions had not changed and cautioned that this was not robust formal oversight.[13]
Both statements can be true. An evaluator can have real editorial independence in some defined space and still have a limited ability to publish claims dependent on a partner’s confidential evidence. The public cannot infer universal powers from a pilot, or infer silent suppression from an NDA. It can demand a clearer project-by-project record: access granted, model configuration, authors, funder disclosure, conflicts, review rights, redactions, withdrawals and the response to a material finding.
ARC, FTX and EA: The Historical Record Is Narrower Than the Online Narrative
The historical bridge is straightforward when the dates are kept intact. METR says ARC hired Barnes in 2022, incubated ARC Evals and announced in September 2023 that ARC Evals would spin out as a separate evaluation-focused organization with Barnes leading it.[9] The later METR legal entity received its federal tax-exempt ruling in March 2024.[3] These facts establish lineage. They do not turn ARC and METR into one undifferentiated funding ledger.
The FY2024 Schedule R does document related-organization transactions involving ARC and METR, with ARC named as the counterparty: $4,501,424 in the category for gifts, grants or capital contributions from ARC; $490,182 in the category for loans or loan guarantees by ARC; and $325,431 in the category for loans or loan guarantees to or for ARC.[10] The filing gives fair market value as the valuation method. It does not supply the agreements, transaction dates, year-end loan balances or donor restrictions. The category labels distinguish contributions from loans or guarantees; adding all three amounts as “funding received” would erase those differences. They should be reported as filed categories, not recast as evidence that FTX funded METR.
ARC’s own public statement says it received a $1.25 million grant from the FTX Foundation in 2022. It said it set the money aside after FTX’s collapse and, in a 2024 update, had returned the grant less legal and administrative expenses to the FTX bankruptcy estate.[8] The statement does not supply a transfer date or net returned amount.
Here is the evidence boundary: reviewed sources document an ARC grant and ARC’s stated return decision. They do not document an FTX Foundation grant to Model Evaluation and Threat Research, Inc.; an FTX grant to ARC Evals specifically; Sam Bankman-Fried holding a board, employment, equity or editorial role at METR; or a direct link from that 2022 grant to a particular METR report. A 2022 ARC relationship is a reason to investigate the separation record. It is not a license to assign guilt by association to people who later appear on METR’s team page.
The related EA context needs the same discipline. Longview’s August 2023 grants report says it recommended $220,000 to ARC Evals, described it as part of ARC, and referred to prior OpenAI and Anthropic evaluation partnerships.[14] This is evidence of an EA-adjacent philanthropic pathway and a recommendation. It is not proof of a payment date, a later grant to METR, donor control or an individual staff member’s ideological membership.
EA’s institutional history is older than this dispute. Giving What We Can’s history records Toby Ord starting the project in 2009 and launching it with Will MacAskill that November. That is a concrete effective-giving institution, not the founding date of every EA organization. The movement’s own introduction describes evidence-based choices across causes, including global health, animal welfare and risks to the future. AI safety is one field in that larger landscape.
The relevant distinction is between an intellectual agenda and a legal instruction. Prioritizing low-probability, very large harms can lead a researcher to build different tests from someone focused on near-term product reliability. That is a legitimate subject for criticism: which risks are included, whose costs count, and how sensitive is the conclusion to those choices? It becomes a claim of improper influence only when evidence connects an outside interest to a compromised decision.
“Rationalist,” “longtermist,” “EA” and “AI safety” likewise describe different things. For this article, a biography naming the Future of Humanity Institute establishes prior professional work there. It does not identify a person’s complete philosophical commitments. Treating those terms as a common employer would make the analysis less precise at the exact point it should become more demanding.
The history that can be sourced
Original events, reporting periods and later updates are distinguished.
| Date | Milestone | Significance |
|---|---|---|
| 2022 | ARC receives FTX Foundation grant | ARC says it received $1.25M. This is an ARC record, not a METR legal-entity grant. |
| 2022 | ARC incubates ARC Evals | METR’s spin-out notice says Barnes was hired by ARC and ARC Evals was incubated there. |
| Sep. 2023 | ARC Evals spin-out announced | Barnes is named to lead the separate evaluation organization. |
| Mar. 2024 | METR tax exemption issued | A separate legal-entity milestone, not a retroactive consolidation of ARC finances. |
| 2024 update | ARC reports return of the grant | The update says the grant had been returned less costs. It does not identify the transfer date. |
| FY2024 | ARC/METR transactions reported | Schedule R lists contributions and separate loan-or-guarantee categories. The return was filed November 16, 2025. |
Kevin Bass’s funding argument: Which links survive scrutiny?
In a September 14, 2026 X post, published at 22:10 UTC, Kevin Bass calls for a congressional investigation. His argument is that appreciation in donated Anthropic shares could enrich a philanthropic network supporting evaluators and journalism, creating incentives to protect Anthropic while promoting regulation. He goes further, asserting that METR is effectively on Anthropic’s payroll and that this system cannot stop amplifying fear of AI. Those are Bass’s allegations and causal interpretation, not findings established by the documents below.
The useful part is the question about indirect financial exposure. A rule against taking lab money does not automatically answer whether a donor’s wealth depends on a lab. But tracing that exposure requires identifying the asset holder, the grant decision-maker, the recipient and the conditions of each transfer. A diagram cannot replace those records.
| Claim under examination | What the evidence supports | What it does not establish |
|---|---|---|
| Moskovitz donated Anthropic shares into Good Ventures Foundation, where they dominate its portfolio. | Coefficient CEO Alexander Berger’s December 2025 statement says Moskovitz donated his stake and Coefficient was not the recipient. Bass’s research notes leave the legal recipient unresolved. | The specific Good Ventures Foundation holding, its present size or its share of that entity’s assets. An unidentified charitable vehicle is not an identified foundation balance sheet. |
| The stake rose from $500 million to more than $7.7 billion. | Bass’s methodological notes explain that the first number is a press estimate and the second is a ceiling derived from a reported ownership percentage and a later financing valuation. | A realized gain, a current appraised holding or money available for grants. Even accepting the inputs, 0.008 × $965 billion = $7.72 billion; an ownership estimate below 0.8% produces an upper bound, not proof of a value above $7.7 billion. |
| Coefficient directly funds METR. | Its December 19, 2025 donor recommendations recommend METR while explicitly saying it is not a Coefficient grantee. Historical ARC funding and partner relationships require separate accounting. | A direct grant or a continuous, earmarked cash path to METR. A recommendation to other donors is not the recommending organization’s payment. |
| Grants through intermediaries all trace back to the same Anthropic stake. | Bass’s repository identifies donor-advised accounts whose principals remain unattributed. His Vanguard follow-up acknowledges a tracing gap. | The originating donor, asset sold, restrictions or onward recipient for each grant. A grant to RAND, FAR AI or Longview cannot simply be counted as money received by METR. |
| Shared funders finance journalism that serves Anthropic. | Tarbell lists Coefficient, Longview and SFF in its $1 million-plus lifetime-support tier. It supports fellows and grantees and publishes Transformer; it says donors have no editorial control. | That Anthropic commissions the reporting, controls the other newsrooms where fellows write, or determines an article’s conclusions. Tarbell’s independence statement is its stated policy, not an independent audit. |
| Fundraising caused METR to produce alarming findings. | METR’s August 14 update reports commitments of around $71 million over six months. That is METR’s rounded figure for commitments, not cash received. | That any result was purchased or manipulated. Timing alone cannot establish causation; methods, negative results, project selection and publication decisions must be examined. |
Bass also raises staff movement, Redwood connections, media disclosure and competing evaluator candidates. Each is a legitimate reporting lead. None makes every organization in a shared professional network a subsidiary of Anthropic. The thread’s broader predictions about China and American political instability are political claims, not conclusions derivable from a grant ledger. Its funding argument also does not establish e/acc involvement.
The most consequential unanswered question is narrower and testable: how much of METR’s future funding could disappear if it published a result that materially harmed an important donor’s investments? Answering it needs donor concentration, renewal conditions, relevant investment exposure and project-level safeguards. These records could establish a serious conflict without proving a fabricated result. Conversely, an untraced payment cannot be assigned to a convenient donor merely because it would complete the story.
The Accelerationists: Authors, Businesses and Competing Visions
Effective accelerationism, written e/acc, has a public authorship record rather than a verified membership register. A June 1, 2022 introduction embedded in its newsletter credits @BasedBeff, @bayeslord, @zestular and @creatine_cycle. The October 31, 2022 repost of its principles names @BasedBeffJezos and @bayeslord as authors. These dates establish the records reviewed here, not the first private discussion or the earliest possible use of the label.[11]
Guillaume Verdon: The public person behind Beff Jezos
In his December 29, 2023 interview with Lex Fridman, Verdon publicly discusses being Beff Jezos, co-writing with Bayeslord and trying to spread technological optimism. He describes e/acc as a loose cultural framework with forks, rather than a formal organization. That is his account of its structure, not an audited finding. His public advocacy and his business role are both relevant to understanding the argument.
Extropic’s own launch account identifies Verdon as its founder and CEO, dates its founding to 2022 and describes his prior role as quantum technology lead in Alphabet X’s Physics & AI team. It names Trevor McCourt as CTO. The company’s current site describes thermodynamic computing for probabilistic AI workloads. McCourt’s corporate role is not evidence that he co-founded e/acc or adopted its philosophy.
There is a straightforward commercial interest to disclose: an AI-computing company can benefit when demand for AI infrastructure grows. That observation does not establish that Verdon’s philosophy is insincere, that Extropic finances e/acc activity, or that an investor dictated his views. It does establish why a reader should evaluate a hardware founder’s arguments about acceleration with the same attention to incentives applied to an evaluator’s arguments about risk.
Bayeslord and the other credited accounts: What the record actually identifies
Verdon describes Bayeslord as a separate, anonymous co-founder in the 2023 interview. The primary repost independently credits the account as a coauthor. Zestular and creatine_cycle appear in the embedded introduction’s credits. The reviewed material does not authenticate their civil identities, employers, financial interests or current responsibilities. It supports naming the public handles and their credited writing, not constructing biographies from guesses.[11][15]
This leaves an accountability difference. A nonprofit can be asked for directors, filings and research approvals. A decentralized current needs to be evaluated through specific authors, organizations and actions. An account using a slogan, an investor backing a company, and a person co-writing a manifesto are three different forms of involvement. A list that merges them would inflate the apparent coherence of the movement.
The philosophical wager: Adaptation instead of centralized restraint
The 2022 principles frame technological growth, market competition and intelligence as adaptive processes, invoke thermodynamics, and favor experimentation over top-down restriction. They explicitly contrast that position with anti-AGI factions of EA.[11] Those are the authors’ philosophical claims. Thermodynamics alone does not tell policymakers which model release is safe, how losses should be distributed or what level of irreversible harm is acceptable.
The strongest version of the accelerationist argument deserves a fair hearing: centralized gatekeepers can make errors too. Restrictions can entrench incumbents, delay useful products and remove opportunities to learn from deployment. An evaluator with privileged access should not acquire unchallengeable policy authority merely by measuring risk. A useful test is whether its evidence can be contested and whether proposed restrictions are proportionate to the capability actually measured.
The weakness appears when a general faith in adaptation substitutes for a model-specific safety case. Markets can reward a useful product while costs fall on people outside the transaction. Some consequences arrive too quickly for learning-by-failure to be adequate. These are analytical reasons to demand evidence on both sides: which restriction causes which opportunity cost, and which deployment creates which risk under which conditions? Neither an optimistic slogan nor a catastrophic scenario answers those questions by itself.
Marc Andreessen: An influential argument with a commercial setting
Andreessen’s October 16, 2023 Techno-Optimist Manifesto makes a related case for markets, technological development, intelligence and energy. It credits Nick Land in its discussion of the techno-capital machine, endorses acceleration, and attacks several ideas associated with caution, including existential-risk and risk-management rhetoric. That is evidence of Andreessen’s published position and an intellectual reference. It does not make Land an e/acc founder or establish an organizational command chain.
The location of that argument matters: it is published by a venture-capital firm whose business involves financing companies. That is an identifiable commercial context. It is not proof that every a16z partner agrees with every sentence, or that an investment relationship gives Andreessen authority over a nonprofit evaluator.
There is a particularly instructive professional overlap. METR lists Marco Mascorro as an advisor and an Andreessen Horowitz partner. This is a real advisory connection worth disclosing. It does not establish a board vote, an a16z grant to METR, Mascorro’s adherence to e/acc, or Andreessen’s control over METR. The appropriate follow-up is the advisor’s scope, information access and conflict arrangements. A person’s advisory role cannot be replaced by a theory about their employer’s ideological unity.
Vitalik Buterin: Why two camps are not enough
Buterin’s November 27, 2023 essay My techno-optimism offers d/acc, emphasizing defensive, decentralized, democratic and differential acceleration. He recognizes costs from delaying beneficial technologies while arguing for special caution about advanced AI and concentrated power. This is an authored alternative to treating every technological advance as the same policy choice, not evidence of a METR affiliation.
That distinction helps clarify the dispute. “Accelerate what?” is a better question than “Are you for progress?” Capabilities, independent evaluation, defensive tools, distributed infrastructure and public oversight can advance at different rates. A policy can accelerate one while constraining another. Calling all such choices “deceleration” hides the tradeoff instead of resolving it.
METR should therefore be scrutinized as a particular institution with particular methods. The e/acc writers should be scrutinized through their documented statements and interests. The record reviewed here does not establish an e/acc role in METR’s governance. It does show why a benchmark’s public interpretation can become a contest over who gets to define progress, acceptable risk and sufficient evidence.
What Independence Would Look Like in Public: Make the Decision Trail Auditable
METR has published evidence that cuts against a simple capture story. Its 2025 randomized developer study reported that 16 experienced open-source developers working across 246 familiar-project tasks took 19% longer with early-2025 AI tools, with a 2% to 39% confidence interval. A 2026 update then disclosed selection bias that prevented a reliable estimate of current productivity effects and changed the design. Publishing an inconvenient result and then disclosing a limitation is evidence of a research culture prepared to surface uncertainty. It is not a substitute for institutional controls.
The May 2026 pilot disclosed that it began without an applicable personnel conflict policy and reported close personal ties to lab employees among at least six participating staff and collaborators.[12] That historical gap must be distinguished from METR’s published policy, version 1.0, dated August 28, 2026. It covers company-identifying risk assessments and specifies conflict review and disclosure. It permits grants from public grantmakers previously funded by lab employees if those employees do not participate in grant decisions. For unclear donor direction, confirmation is required at 0.1% of annual income. This corrects the earlier draft’s statement that the existence of a current policy was unresolved. A published policy still requires evidence of implementation.
What the public record establishes, and what it does not
Counts are public-role or study figures, not measures of organizational independence.
Founder/CEO, President, Chief Scientist and CTO on METR’s current directory.
Gleave and Dattani are each labelled Advisor and Board Member.
A 2022 ARC grant, not established METR funding.
FY2024 Schedule R: gift, grant or capital contribution from ARC.
Let’s be clear about the standard. The strongest conclusion available is neither that METR is secretly run by a movement nor that every safeguard is complete. It is that public disclosures show a real research organization with identifiable functional roles and some meaningful safeguards, alongside incomplete visibility into the exact points where authority can affect an evaluation.
The repair is practical. For each consequential report, publish the author and contributor roles, the benchmark version, task-selection and exclusion criteria, model-access agreement, relevant funding disclosure, conflict process, external review rights, redactions, withdrawals and management response. Publish a current legal board roster, director-independence criteria and aggregate recusal information. Value or bracket non-cash model access. Give a reviewer enough evidence to distinguish a former affiliation from an active conflict and a formal independence claim from actual publication power.
A Network Is Not a Chain of Command
The evidence supports historical links, current titles and intellectual disagreements. It does not identify an unnamed person, donor or ideology directing METR’s results. The real test is whether the organization can show, report by report, how it protects a bad finding from the interests it may inconvenience.
The uncomfortable truth is that frontier governance will increasingly depend on evaluators whose work makes somebody unhappy. That is precisely why the relevant question is not who shares a discourse, a former employer or a donor ecosystem. It is who had authority at each decision point, what they were allowed to see, what they were allowed to publish and whether outsiders can verify the answer. Until that trail is public, trust remains a request rather than a demonstrated property.
Sources & References
Primary organizational disclosures, public biographies, benchmark documentation and historical records. Individual profile links appear in the text where the role is discussed.
| # | Source | Outlet | Date | Key Takeaway |
|---|---|---|---|---|
| 1 | METR | Accessed Sep. 14, 2026 | Current public team categories, leadership titles, advisor labels, collaborator labels and access/funding disclosures. | |
| 2 | Extropic | Accessed Sep. 14, 2026 | Company account of its founding, leadership and AI-computing focus. | |
| 3 | IRS filing indexed by ProPublica | FY2024, filed Nov. 16, 2025 | Historical governing-body count, independent-member classification and tax-exempt entity context. | |
| 4 | METR | Mar. 2025 | HCAST task count, human baselines, hours and scope. | |
| 5 | METR | Nov. 22, 2024 | RE-Bench environments and human attempts. | |
| 6 | METR | Jan. 29, 2026 | TH1.1 composition and long-task timing qualifications. | |
| 7 | METR | Updated May 8, 2026 | Definition of the fitted 50% time-horizon measure and current limits. | |
| 8 | Alignment Research Center | 2022 statement; 2024 update | ARC’s stated $1.25M receipt, set-aside and return to the estate less legal and administrative costs. | |
| 9 | METR | Sep. 19, 2023 | ARC Evals incubation and spin-out history. | |
| 10 | IRS filing rendering | FY2024 | Historical related-organization transaction categories involving ARC and METR; not a complete explanation of terms or purpose. | |
| 11 | Effective Accelerationism | Oct. 31, 2022 | Primary repost with authorship labels, embedded June 2022 introduction and explicit contrast with anti-AGI EA factions. | |
| 12 | METR | May 19, 2026 | Pilot-specific publication terms, redactions, withdrawals and personnel-conflict disclosure. | |
| 13 | METR | Jun. 26, 2026 | METR’s account of NDA review rights and limits on formal oversight. | |
| 14 | Longtermism Fund | Aug. 2023 | Public $220,000 recommendation to ARC Evals, not proof of a later METR payment. | |
| 15 | Lex Fridman Podcast | Dec. 29, 2023 | Verdon’s public identification with Based Beff Jezos and description of Bayeslord as a distinct anonymous co-founder. |
Last updated: September 15, 2026. Added the Bass claim review and corrected the current conflict-policy status.




