All articlesStack

AEF-1 standard emerges for third-party evals as OpenAI, Anthropic, and xAI all cosign — here's what the framework actually covers

The AI Evaluation Framework v1 just launched with backing from OpenAI, Anthropic, and xAI. Here's what the standard actually specifies and what it means for independent testing.

Sep 15, 2026 3 min read
evalssafetyaef-1third-party-testing

OpenAI, Anthropic, and xAI just cosigned AEF-1, the first standard framework for third-party AI evaluators. The spec dropped quietly last week after months of behind-closed-doors negotiation between the labs and independent testing groups.

AEF-1 (AI Evaluation Framework version 1) is a 47-page technical specification that defines how external evaluators access frontier models, what telemetry they can collect, and how results get reported. It covers pre-deployment testing, continuous monitoring, and adversarial red-teaming. The framework requires labs to provide API access with specific rate limits, logging hooks, and audit trails.

The standard emerged from the Pacing Accords discussions earlier this year. Labs wanted a common protocol so they wouldn't have to negotiate separate eval contracts with every independent org. Evaluators wanted guaranteed access that couldn't be revoked mid-assessment. AEF-1 is the compromise.

What the spec actually includes

AEF-1 mandates three access tiers. Tier 1 is basic API access with extended rate limits and priority routing — similar to what research partners get today. Tier 2 adds model introspection: evaluators can request intermediate activations, attention patterns, and chain-of-thought traces for specific prompts. Tier 3 is red-team access with adversarial optimization allowed and no content policy filtering.

The framework requires labs to maintain a public registry of active evaluations. When an evaluator starts an AEF-1 assessment, the lab posts a notice with the evaluator's name, scope, and expected completion date. Results are embargoed for 30 days after completion to give labs time to patch issues, then must be published.

Telemetry requirements are specific. Labs must log every eval API call with timestamp, model version hash, input token count, output token count, latency, and whether the request triggered any internal safety filters. Evaluators can request batch exports of this telemetry weekly. The spec defines a JSON schema for the export format.

AEF-1 also covers compensation. Labs must either provide free API credits (minimum 10M tokens per month for Tier 1, 50M for Tier 2) or pay evaluators directly at a rate pegged to their standard API pricing. No unpaid volunteer evals under the framework.

What's missing

The standard doesn't mandate specific eval benchmarks. Evaluators choose their own test suites. AEF-1 just ensures they have the access and telemetry to run them. Some safety groups wanted required tests (bio-weapons, cyber-offense, persuasion), but the labs pushed back. The compromise was to require disclosure of what wasn't tested.

There's no enforcement mechanism beyond reputation. If a lab violates AEF-1 — pulls access mid-eval, withholds telemetry, blocks publication — the framework just requires the evaluator to document it publicly. No fines, no third-party arbitration.

Model weights aren't covered. AEF-1 is API-only. Open-weight models (Llama, Mistral, Qwen) fall outside the spec entirely. Some evaluators wanted weight access for Tier 3, but the labs refused. They argued that API access with introspection hooks provides equivalent capability for testing without the IP and misuse risks of weight distribution.

Why this matters now

The timing isn't accidental. Congress has been threatening mandatory third-party testing requirements since the Fable 5.1 jailbreak incident in July. By launching AEF-1 now, labs are trying to preempt legislation. The framework gives them a narrative: "We already have industry-standard independent evals, no need for government mandates."

For builders, AEF-1 changes the eval ecosystem. Before this, most third-party testing happened through informal research partnerships or one-off contracts. Access was inconsistent and often limited. Now there's a standard protocol. If you're running evals for clients (compliance, red-teaming, capability assessment), you can point to AEF-1 and know what you're getting from each lab.

The standard also clarifies what "independent" means. Under AEF-1, evaluators must disclose any financial relationships with labs being tested and recuse from assessments where there's a conflict. Labs can't fund an evaluator and then have that org assess their models under the framework.

Version 2 is already in draft. The current spec punts on multi-modal evals (images, video, audio) and agent evals (tool use, environment interaction). Those will be in AEF-2, targeted for Q1 2027. The working group is also debating whether to extend the framework to inference-time safety (prompt filtering, output moderation) vs just pre-deployment capability assessment.

For now, AEF-1 is the baseline. If a lab claims third-party validation, you can check whether it was done under the framework. If not, the access conditions and reporting requirements were probably weaker.

/ 06 — Start hereOne business day response

Tell us what you'd like built.

Send us a paragraph about the workflow, phone line, or tool you want built. We'll reply within one business day with a one-page plan, a fixed price, and a delivery date you can put on a calendar.

  • 30-min scoping call, free
  • Written proposal within 48 hours
  • Fixed price before we start
  • Most builds delivered in 2–8 weeks