The 2026 Benchmark Report · FDA regulatory AI

Fluency is not fidelity.

We benchmarked Agent Astro against the leading frontier models - with and without web search - on FDA 510(k) predicate discovery and regulatory question answering, on identical inputs with deterministic scoring. This page is the report in figures; the white paper carries the full methodology.

  • 0

    FDA clearance numbers Agent Astro fabricated, across every task in this report

  • 100%

    Recall counts and class breakdowns exactly right, where the best web model manages 73%

  • 90 / 90

    The only system both accurate (90%) and calibrated (90%) on regulatory Q&A

  • 67.5%

    Predicates discovered from a description alone, where unaided frontier models score near zero

Methodology

A fair, reproducible comparison

Two tasks built directly from the live FDA database - at evaluation time, 174,828 cleared devices and 165,953 recorded predicate citations.

  • Identical inputs, deterministic scoring

    Every system answered the same inputs under the same automatic rules, and every response was saved verbatim for audit.

  • Memorization removed

    The predicate gold set uses only devices cleared after the models' training cutoffs - a model cannot recall what it never saw.

  • Baselines at their strongest

    Gemini 3.1 Pro, GPT-5.5, GPT-5.6, Claude Opus 4.8, Grok-4.3, Kimi K3, and Perplexity Sonar - each with and without live web search.

  • Written to be contested

    Where a baseline beats Agent Astro, it is reported in full. The harness, gold sets, and raw responses are retained for replication.

Track 1 · Predicate discovery

The task teams actually face - and only one system can do

Given a device description with every K-number removed, find the predicates it is substantially equivalent to. The gold set is 200 devices cleared after the models’ training cutoffs, so nothing can be recalled from memory - only discovered.

Agent Astro reaches 67.5% with zero fabricated citations. The frontier models, working from the description alone, are effectively unable to do it - and they invent 16 to 24% of the clearance numbers they offer.

Predicate discovery, Hit@10

From the device description alone · n = 200 post-cutoff devices

  • Agent Astrozero fabricated citations

    67.5%

  • Gemini 3.1 Profabricates 20.7%

    2.0%

  • Claude Opus 4.8fabricates 23.7%

    1.0%

  • GPT-5.6fabricates 16.2%

    0.5%

  • Grok-4.3fabricates 19.3%

    0.5%

  • Kimi K3fabricates 15.6%

    0.5%

  • GPT-5.5fabricates 3.8%

    0.0%

Web search is lookup, not discovery

Given live web search, the same models jump to 48–96% on this set - and we report that in full. But our gold devices are already cleared, so a web model simply finds each device’s own published 510(k) filing and reads back the predicate it lists. That works only once a filing is public - and it vanishes the moment the device has not yet been filed, which is exactly when a team needs the answer.

Same models, with web search

Reading the predicate off the published filing · Agent Astro does not attempt lookup

  • GPT-5.6 + web

    95.5%

  • GPT-5.5 + web

    92.9%

  • Kimi K3 + web

    77.0%

  • Grok-4.3 + web

    74.0%

  • Gemini 3.1 Pro + web

    72.2%

  • Claude Opus 4.8 + web

    54.0%

  • Perplexity Sonar + web

    48.0%

Track 2 · Regulatory Q&A

Accurate and calibrated - both at once

One hundred questions from the live database: hard facts with deterministic ground truth, plus adversarial probes - non-existent K-numbers, fictitious devices, false premises. A trustworthy system answers what it can verify and declines what it cannot.

Agent Astro is the only system at 90/90. One baseline - Grok-4.3 with web search - is a genuine peer on this set, and we report that plainly. GPT-5.5 with web search answers accurately but fabricates on half the non-existent-entity probes (Holm-adjusted McNemar p = 0.0003).

Factual correctness vs. calibration

n = 100 questions

Regulatory Q&A: factual correctness and calibration by system
SystemFactual correctnessCalibration
Agent Astro0.9010.9
Grok-4.3 + web0.9830.95
GPT-5.6 + web0.9690.75
GPT-5.5 + web0.9690.5
Gemini 3.1 Pro + web0.7930.973
Kimi K3 + web0.6610.85
Perplexity Sonar + web0.550.775
Claude Opus 4.8 + web0.050.725
Frontier LLMs, no web00.6

Measured performance on five axes

Each scaled 0–1, higher is better · web models credited their full lookup score on the predicate axis

  • Agent Astro
  • Grok-4.3 + web
  • GPT-5.6 + web
  • GPT-5.5 + web
Measured performance on five benchmark axes, scaled zero to one
SystemPredicate discoveryQ&A accuracyQ&A calibrationCitation integrityRecall / AE accuracy
Agent Astro0.6750.9010.90.961
Grok-4.3 + web0.740.9830.950.980.53
GPT-5.6 + web0.9550.9690.750.90.23
GPT-5.5 + web0.9290.9690.50.80.17

The combination

The area each shape encloses is the argument

A general model with web search can spike on one or two axes. Only Agent Astro reaches the outer edge on all five at once - discovery, accuracy, calibration, citation integrity, and recall intelligence - and that combination, not any single score, is what regulated work requires.

The decisive test

From answers to deliverables

Regulatory work is not isolated questions - it is deliverables. Asked for a recall and adverse-event assessment of a product code, Agent Astro states the correct recall count and class breakdown for every one of 30 codes tested. It is the cleanest cliff in the report.

Correct recall count, by system

Share of product codes with the exact total · n = 30

  • Agent Astroexact count and class breakdown, every code

    100%

  • Perplexity Sonar + web

    73%

  • Kimi K3 + web

    60%

  • Grok-4.3 + web

    53%

  • Claude Opus 4.8 + web

    37%

  • GPT-5.6 + web

    23%

  • Gemini 3.1 Pro + web

    20%

  • GPT-5.5 + web

    17%

Worked example · product code JZO

Truth: 18 recalls (Class II: 14, Class III: 3, Unclassified: 1, no Class I).

AGENT ASTRO

“18 recalls. Class II: 14, Class III: 3, Unclassified: 1.”

Exact, grounded in the FDA recalls table.

FRONTIER MODELS, WITH WEB SEARCH

Grok-4.3: “at least 7 recalls, 1 Class I.”  Gemini 3.1 Pro: “at least 7.”  Perplexity Sonar: “low single digits, no Class I.”  Claude Opus 4.8: could not produce counts.  GPT-5.5: no answer.

Half a second to a correct answer

Because Agent Astro reads the recall record straight from the FDA database, it returns the assessment in about half a second. The web models fetch pages and reason over them for 20 seconds to more than four minutes - and are wrong more often than not. Accuracy and speed usually trade against each other; here, grounding wins both at once.

Median time to the recall assessment

Wall-clock per deliverable · n = 30 · shorter is better

  • Agent Astro

    0.5s

  • Perplexity Sonar + web

    21s

  • Grok-4.3 + web

    25s

  • Gemini 3.1 Pro + web

    47s

  • Claude Opus 4.8 + web

    50s

  • GPT-5.5 + web

    2m 47s

  • GPT-5.6 + web

    3m 02s

  • Kimi K3 + web

    4m 09s

The conclusion

Three behaviors, one conclusion

  • Frontier LLM, no web

    Not a regulatory tool

    • Predicate discovery fails (2% or less)
    • Recent device facts: 0% correct
    • Invents 16–24% of cited K-numbers
  • Frontier LLM + web search

    A retriever, not an analyst

    • Reads predicates off already-public filings (up to 96%)
    • Calibration inconsistent (GPT-5.5: 50%)
    • $0.25–0.58 and 20s–4min per deliverable
  • Agent Astro, grounded

    Accurate, calibrated, auditable

    • Leads genuine discovery (67.5%), zero fabrication
    • 90/90 on Q&A accuracy and calibration
    • ~0.5s per assessment, every claim resolves to an FDA record

Read this like a skeptic

This is a vendor-run benchmark. The harness, gold sets, and raw model responses are retained so any figure can be reproduced or challenged; scoring is deterministic rather than judged; and where a comparison favors the baselines - the web models’ lookup scores - it is credited in full. The strongest test is your own: take ten of your devices, ask each system for predicates and recall history, and check every cited identifier against the FDA database. The exercise takes an afternoon.

Evaluate it yourself

Request a live evaluation on your own devices, or access to the benchmark harness to run every figure in this report yourself.