← Evidence & methodology

US AISI Anthropic

1.0.0-rc.1 · Evidence cutoff 12 September 2026

Directly documented interface

Evaluation

Named evaluator performed a model-specific exercise.

evaluator → provider. Evaluator → provider; domains follow the reported exercise.

Displayed claim → analysis (if present) → coded attributes → interface / instrument → atomic evidence → source version

Instrument and timeline

Instrument
Claude 3.5 Sonnet evaluation
Bindingness
No regulatory power inferred
Announced
Unknown
Observed by
2024-11-19 · upper-bound
Effective period
UnknownUnknown
Event
Conducted
Status at cutoff
historical event only
Last confirmation
2024-11-19

Model-release dates conflict within the full report; report publication supplies conservative upper bound. No exact testing date asserted.

Status investigation

Checked 2026-09-12; assessed as of 2026-09-12. Last confirmation does not establish uninterrupted continuity.

amendments
unknown
expiry
unknown
termination
investigated within listed sources
restrictions
investigated within listed sources
succession
see succession.json
conflicts
Model-release dates conflict within the full report; report publication supplies conservative upper bound. No exact testing date asserted.
Search queries
  • "US AISI" "anthropic" evaluation 2026 amendment termination expiry status before:2026-09-13

Analytical assessment

No reviewed analytical conclusion is attached. This does not establish independence or absence of an institutional constraint.

Coded attributes

evaluation-capacity: substantial-access-constrained

Ability to design and execute a substantive assessment independently, conditional on access.

Evaluator × evaluation exercise or programme × date. Claude 3.5 Sonnet 2024 pre-deployment exercise; assessed 2024-11-19.

Completed independent assessment plus explicit access/time limitation satisfies anchor. Not a statement of current CAISI resources.

Supporting evidence: independent-tests access-limits . Counterevidence: independent-tests

  • Coding applies only to stated context and date, not a permanent institutional rank.
Published rubric anchors
absent
Affirmative evidence of no ability to perform this assessment.
nascent
Initial methods/team with no completed substantive assessment.
substantial-access-constrained
Completed independent assessment with documented access or time constraints materially limiting conclusions.
substantial-durable-access
Completed assessment and enforceable/reproducible access sufficient for the task over its defined period.

Unresolved when necessary evidence is missing or adjacent anchors cannot be distinguished. No numeric conversion. Absence requires affirmative remit or operational evidence, not missing sources.

technical-control: shared-constrained

Practical ability to alter the model, its availability, or safeguards in the specified deployment. Contract clauses alone are insufficient.

Actor × exact model/deployment × period. Anthropic provision and replacement of Claude models for DoW, contrasted with deployed operational control; 2025–August 2026; assessed 2026-08-27.

Anthropic provides model artifacts and responds to feedback; DoW and cloud providers independently gate operational deployment. Existing deployed models are outside Anthropic technical alteration or shutdown. Consequential functions are divided across institutions, satisfying shared/constrained control without attributing a real-time veto.

Supporting evidence: court-model-provision court-update-gate . Counterevidence: contractual-control court-static-deployment

  • Coding applies only to stated context and date, not a permanent institutional rank.
Published rubric anchors
little-control
No demonstrated practical control over the specified consequential operation, with affirmative contrary evidence.
shared-constrained
Consequential controls divided among actor, host, customer or enforceable technical constraints.
substantial
Actor demonstrably retains major development, update or access controls in this environment.
dominant-lifecycle
Actor demonstrably controls development, deployment, updates and access across the specified lifecycle without a consequential independent controller.

Unresolved when necessary evidence is missing or adjacent anchors cannot be distinguished. No numeric conversion. Absence requires affirmative remit or operational evidence, not missing sources.

Atomic evidence and exact sources

US AISI and UK AISI conducted a joint evaluation before the upgraded Claude 3.5 Sonnet release.

Source
Pre-deployment evaluation of upgraded Claude 3.5 Sonnet · NIST / US AISI and UK AISI
Version
2025-03-20; retrieved 2026-09-12
Exact locator
Introduction; Overview of the Joint Safety Research & Testing Exercise, paragraphs 1–2
Limit
A completed exercise does not establish continuing access or regulatory authority.

Retrieved bytes SHA-256: eb91f397ed29e9e2684cd9dd518dea0b16fba6fbc5634d37cfc25297c78050b5

The exercise included biological and cyber capabilities; UK biological findings were not published in this summary.

Source
Pre-deployment evaluation of upgraded Claude 3.5 Sonnet · NIST / US AISI and UK AISI
Version
2025-03-20; retrieved 2026-09-12
Exact locator
Overview, paragraph 2; Biological Capabilities, paragraph 2; Cyber Capabilities
Limit
Domain tags describe the exercise scope, not a finding about its effectiveness.

Retrieved bytes SHA-256: eb91f397ed29e9e2684cd9dd518dea0b16fba6fbc5634d37cfc25297c78050b5

The evaluation had a limited access and testing period.

Source
Claude 3.5 Sonnet joint testing report · US AISI / UK AISI
Version
2024-11-19; retrieved 2026-09-12
Exact locator
Section 2.1, printed pages 1–2
Limit
Applies to this exercise; does not identify who caused the limits.

Retrieved bytes SHA-256: cee68f384975897e9ae8fd556baa9ebae9919fe08fe1e1abb363cbead7802903

The institutes each ran independent tests and worked together on methodology and interpretation of findings.

Source
Claude 3.5 Sonnet joint testing report · US AISI / UK AISI
Version
2024-11-19; retrieved 2026-09-12
Exact locator
Section 1, printed page 1
Limit
Independence of tests does not establish durable model access.

Retrieved bytes SHA-256: cee68f384975897e9ae8fd556baa9ebae9919fe08fe1e1abb363cbead7802903

Section 2.2 says October 10; section 1 says October 22 for the model release.

Source
Claude 3.5 Sonnet joint testing report · US AISI / UK AISI
Version
2024-11-19; retrieved 2026-09-12
Exact locator
Sections 1 and 2.2, printed pages 1–2
Limit
Retain discrepancy; use report publication as upper bound, not an exact evaluation date.

Retrieved bytes SHA-256: cee68f384975897e9ae8fd556baa9ebae9919fe08fe1e1abb363cbead7802903

Anthropic could not technologically enforce contractual usage restrictions and lacked direct visibility into DoW use.

Source
Anthropic v DoW: summary judgment, document 250 · US District Court, Northern District of California
Version
2026-08-27; retrieved 2026-09-12
Exact locator
Printed page 6
Limit
Specific deployed environment; does not imply no control over future development, contracts or provision.

Byte capture unavailable. HTTP Error 403: Forbidden Published document identity and locator remain recorded.

The court describes deployed DoW Claude models as static and beyond Anthropic technological access, alteration or shutdown.

Source
Anthropic v DoW: summary judgment, document 250 · US District Court, Northern District of California
Version
2026-08-27; retrieved 2026-09-12
Exact locator
Printed page 17, Section II.F, first paragraph
Limit
Applies to deployed models described in this litigation, not future model provision or hosted consumer services.

Byte capture unavailable. HTTP Error 403: Forbidden Published document identity and locator remain recorded.

The court records that replacement Claude models undergo third-party cloud-provider and DoW security testing and evaluation before operational approval.

Source
Anthropic v DoW: summary judgment, document 250 · US District Court, Northern District of California
Version
2026-08-27; retrieved 2026-09-12
Exact locator
Printed page 17, Section II.F, first paragraph
Limit
Identifies a deployment approval gate, not complete interpretability, universal evaluation sufficiency or provider substitutability.

Byte capture unavailable. HTTP Error 403: Forbidden Published document identity and locator remain recorded.

The court records Anthropic providing models to DoW or a primary defense contractor and responding to DoW evaluation feedback and requests.

Source
Anthropic v DoW: summary judgment, document 250 · US District Court, Northern District of California
Version
2026-08-27; retrieved 2026-09-12
Exact locator
Printed page 17, Section II.F, first paragraph
Limit
Provision is distinct from control over deployed inference; future supply terms and update alternatives are unspecified.

Byte capture unavailable. HTTP Error 403: Forbidden Published document identity and locator remain recorded.

Supporting and conflicting evidence

Interface support: joint-evaluation, evaluation-domains, independent-tests. Counterevidence: release-date-conflict

Substantive review

evidence/joint-evaluation: approved · agent reviewer codex-research-checker · 2026-09-12

NIST explicitly describes the joint US/UK exercise on the upgraded Sonnet. The full report has inconsistent model-release dates; this proposition is not used as an exact testing date. Checked locator: Introduction; Overview of the Joint Safety Research & Testing Exercise, paragraphs 1–2. Source version: 2025-03-20.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

evidence/evaluation-domains: approved · agent reviewer codex-research-checker · 2026-09-12

NIST biological/cyber sections establish the exercise topics and nonpublication of the UK biological findings; topic coverage is not effectiveness. Checked locator: Overview, paragraph 2; Biological Capabilities, paragraph 2; Cyber Capabilities. Source version: 2025-03-20.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

evidence/access-limits: approved · agent reviewer codex-research-checker · 2026-09-12

Full joint report section 2.1 explicitly limits access and the testing window. It does not identify the institution responsible for either constraint. Checked locator: Section 2.1, printed pages 1–2. Source version: 2024-11-19.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

evidence/release-date-conflict: approved · agent reviewer codex-research-checker · 2026-09-12

Compared report sections 1 and 2.2: October 22 versus October 10 is a real internal discrepancy, preserved rather than reconciled by inference. Checked locator: Sections 1 and 2.2, printed pages 1–2. Source version: 2024-11-19.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

evidence/contractual-control: approved · agent reviewer codex-research-checker · 2026-09-12

Judgment page6 distinguishes contractual conditions from technical enforcement and use visibility in the specified environment; later provision and lifecycle control are expressly not excluded. Checked locator: Printed page 6. Source version: 2026-08-27.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

instruments/sonnet-evaluation: approved · agent reviewer codex-research-checker · 2026-09-12

NIST/full report identify a shared model-specific evaluation event; no regulatory power is implied.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

relationships/us-anthropic-evaluation: approved · agent reviewer codex-research-checker · 2026-09-12

NIST/full report support actual US testing of Sonnet and evaluator-to-provider direction. Historical event and conservative publication upper bound preserve internal date conflict.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

institution-attributes/us-sonnet-capacity: approved · agent reviewer codex-research-checker · 2026-09-12

Report methods and constraints support the context-bound evaluation-capacity anchor for the US AISI 2024 exercise. Independent testing counters provider-directed-results inference; no present CAISI resource coding follows.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

evidence/independent-tests: approved · agent reviewer codex-primary · 2026-09-12

Independently inspected the primary source at the recorded locator. Narrow proposition, attribution, date/version and stated limit supported; conflicting access/technical/status accounts remain separate.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

evidence/court-static-deployment: approved · agent reviewer codex-primary · 2026-09-12

Independently inspected the primary source at the recorded locator. Narrow proposition, attribution, date/version and stated limit supported; conflicting access/technical/status accounts remain separate.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

evidence/court-update-gate: approved · agent reviewer codex-primary · 2026-09-12

Independently inspected the primary source at the recorded locator. Narrow proposition, attribution, date/version and stated limit supported; conflicting access/technical/status accounts remain separate.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

evidence/court-model-provision: approved · agent reviewer codex-primary · 2026-09-12

Independently inspected the primary source at the recorded locator. Narrow proposition, attribution, date/version and stated limit supported; conflicting access/technical/status accounts remain separate.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

institution-attributes/dod-claude-control: approved · agent reviewer codex-primary · 2026-09-12

Checked rubric and contextual evidence, including Augustcourt p17, MarchCIO pp27–29 and MaySenate pp59–68. Divided control is not remote veto; meaningful reliance is not structural indispensability; substitution ordinal remains unresolved.

Approval fingerprint: 1e7930e6d70cddee066b09a959f5046c89228cfdbf91ecb37db5b43d6212271a

Release-pinned methodology →