The 2026 Benchmark Report · FDA regulatory AI
Fluency is not fidelity.
We benchmarked Agent Astro against the leading frontier models - with and without web search - on FDA 510(k) predicate discovery and regulatory question answering, on identical inputs with deterministic scoring. This page is the report in figures; the white paper carries the full methodology.
0
FDA clearance numbers Agent Astro fabricated, across every task in this report
100%
Recall counts and class breakdowns exactly right, where the best web model manages 73%
90 / 90
The only system both accurate (90%) and calibrated (90%) on regulatory Q&A
67.5%
Predicates discovered from a description alone, where unaided frontier models score near zero
Methodology
A fair, reproducible comparison
Two tasks built directly from the live FDA database - at evaluation time, 174,828 cleared devices and 165,953 recorded predicate citations.
Identical inputs, deterministic scoring
Every system answered the same inputs under the same automatic rules, and every response was saved verbatim for audit.
Memorization removed
The predicate gold set uses only devices cleared after the models' training cutoffs - a model cannot recall what it never saw.
Baselines at their strongest
Gemini 3.1 Pro, GPT-5.5, GPT-5.6, Claude Opus 4.8, Grok-4.3, Kimi K3, and Perplexity Sonar - each with and without live web search.
Written to be contested
Where a baseline beats Agent Astro, it is reported in full. The harness, gold sets, and raw responses are retained for replication.
Track 1 · Predicate discovery
The task teams actually face - and only one system can do
Given a device description with every K-number removed, find the predicates it is substantially equivalent to. The gold set is 200 devices cleared after the models’ training cutoffs, so nothing can be recalled from memory - only discovered.
Agent Astro reaches 67.5% with zero fabricated citations. The frontier models, working from the description alone, are effectively unable to do it - and they invent 16 to 24% of the clearance numbers they offer.
Predicate discovery, Hit@10
From the device description alone · n = 200 post-cutoff devices
Agent Astrozero fabricated citations
67.5%
Gemini 3.1 Profabricates 20.7%
2.0%
Claude Opus 4.8fabricates 23.7%
1.0%
GPT-5.6fabricates 16.2%
0.5%
Grok-4.3fabricates 19.3%
0.5%
Kimi K3fabricates 15.6%
0.5%
GPT-5.5fabricates 3.8%
0.0%
Web search is lookup, not discovery
Given live web search, the same models jump to 48–96% on this set - and we report that in full. But our gold devices are already cleared, so a web model simply finds each device’s own published 510(k) filing and reads back the predicate it lists. That works only once a filing is public - and it vanishes the moment the device has not yet been filed, which is exactly when a team needs the answer.
Same models, with web search
Reading the predicate off the published filing · Agent Astro does not attempt lookup
GPT-5.6 + web
95.5%
GPT-5.5 + web
92.9%
Kimi K3 + web
77.0%
Grok-4.3 + web
74.0%
Gemini 3.1 Pro + web
72.2%
Claude Opus 4.8 + web
54.0%
Perplexity Sonar + web
48.0%
Track 2 · Regulatory Q&A
Accurate and calibrated - both at once
One hundred questions from the live database: hard facts with deterministic ground truth, plus adversarial probes - non-existent K-numbers, fictitious devices, false premises. A trustworthy system answers what it can verify and declines what it cannot.
Agent Astro is the only system at 90/90. One baseline - Grok-4.3 with web search - is a genuine peer on this set, and we report that plainly. GPT-5.5 with web search answers accurately but fabricates on half the non-existent-entity probes (Holm-adjusted McNemar p = 0.0003).
Factual correctness vs. calibration
n = 100 questions
| System | Factual correctness | Calibration |
|---|---|---|
| Agent Astro | 0.901 | 0.9 |
| Grok-4.3 + web | 0.983 | 0.95 |
| GPT-5.6 + web | 0.969 | 0.75 |
| GPT-5.5 + web | 0.969 | 0.5 |
| Gemini 3.1 Pro + web | 0.793 | 0.973 |
| Kimi K3 + web | 0.661 | 0.85 |
| Perplexity Sonar + web | 0.55 | 0.775 |
| Claude Opus 4.8 + web | 0.05 | 0.725 |
| Frontier LLMs, no web | 0 | 0.6 |
Measured performance on five axes
Each scaled 0–1, higher is better · web models credited their full lookup score on the predicate axis
- Agent Astro
- Grok-4.3 + web
- GPT-5.6 + web
- GPT-5.5 + web
| System | Predicate discovery | Q&A accuracy | Q&A calibration | Citation integrity | Recall / AE accuracy |
|---|---|---|---|---|---|
| Agent Astro | 0.675 | 0.901 | 0.9 | 0.96 | 1 |
| Grok-4.3 + web | 0.74 | 0.983 | 0.95 | 0.98 | 0.53 |
| GPT-5.6 + web | 0.955 | 0.969 | 0.75 | 0.9 | 0.23 |
| GPT-5.5 + web | 0.929 | 0.969 | 0.5 | 0.8 | 0.17 |
The combination
The area each shape encloses is the argument
A general model with web search can spike on one or two axes. Only Agent Astro reaches the outer edge on all five at once - discovery, accuracy, calibration, citation integrity, and recall intelligence - and that combination, not any single score, is what regulated work requires.
The decisive test
From answers to deliverables
Regulatory work is not isolated questions - it is deliverables. Asked for a recall and adverse-event assessment of a product code, Agent Astro states the correct recall count and class breakdown for every one of 30 codes tested. It is the cleanest cliff in the report.
Correct recall count, by system
Share of product codes with the exact total · n = 30
Agent Astroexact count and class breakdown, every code
100%
Perplexity Sonar + web
73%
Kimi K3 + web
60%
Grok-4.3 + web
53%
Claude Opus 4.8 + web
37%
GPT-5.6 + web
23%
Gemini 3.1 Pro + web
20%
GPT-5.5 + web
17%
Worked example · product code JZO
Truth: 18 recalls (Class II: 14, Class III: 3, Unclassified: 1, no Class I).
AGENT ASTRO
“18 recalls. Class II: 14, Class III: 3, Unclassified: 1.”
Exact, grounded in the FDA recalls table.
FRONTIER MODELS, WITH WEB SEARCH
Grok-4.3: “at least 7 recalls, 1 Class I.” Gemini 3.1 Pro: “at least 7.” Perplexity Sonar: “low single digits, no Class I.” Claude Opus 4.8: could not produce counts. GPT-5.5: no answer.
Half a second to a correct answer
Because Agent Astro reads the recall record straight from the FDA database, it returns the assessment in about half a second. The web models fetch pages and reason over them for 20 seconds to more than four minutes - and are wrong more often than not. Accuracy and speed usually trade against each other; here, grounding wins both at once.
Median time to the recall assessment
Wall-clock per deliverable · n = 30 · shorter is better
Agent Astro
0.5s
Perplexity Sonar + web
21s
Grok-4.3 + web
25s
Gemini 3.1 Pro + web
47s
Claude Opus 4.8 + web
50s
GPT-5.5 + web
2m 47s
GPT-5.6 + web
3m 02s
Kimi K3 + web
4m 09s
The conclusion
Three behaviors, one conclusion
Frontier LLM, no web
Not a regulatory tool
- Predicate discovery fails (2% or less)
- Recent device facts: 0% correct
- Invents 16–24% of cited K-numbers
Frontier LLM + web search
A retriever, not an analyst
- Reads predicates off already-public filings (up to 96%)
- Calibration inconsistent (GPT-5.5: 50%)
- $0.25–0.58 and 20s–4min per deliverable
Agent Astro, grounded
Accurate, calibrated, auditable
- Leads genuine discovery (67.5%), zero fabrication
- 90/90 on Q&A accuracy and calibration
- ~0.5s per assessment, every claim resolves to an FDA record
Read this like a skeptic
This is a vendor-run benchmark. The harness, gold sets, and raw model responses are retained so any figure can be reproduced or challenged; scoring is deterministic rather than judged; and where a comparison favors the baselines - the web models’ lookup scores - it is credited in full. The strongest test is your own: take ten of your devices, ask each system for predicates and recall history, and check every cited identifier against the FDA database. The exercise takes an afternoon.
Evaluate it yourself
Request a live evaluation on your own devices, or access to the benchmark harness to run every figure in this report yourself.
