Evaluation

模型评测

Offline suites, model scorecards and hallucination triage.

19 suites7 failing
Suites

19

3.7K cases
Avg pass rate

86.4%

Avg score

83.0

Failing suites

7

Regressions (30d)

116

60 runs
Hallucinations

24

6 critical · $547.28 cost
Model scorecard
Dimension scores across top models

Eval suites

Pass rate by target

SuiteStatusPassScore
Safety & Refusal Model Regression
Support Desk AI · model: all candidate models · 187 cases
failing73.3%70.3
Support Triage Agent Trajectory Suite
Support Desk AI · agent: Support Triage Agent · 48 cases
failing75.4%71.0
Docs Answer Agent Trajectory Suite
Docs Chat · agent: Docs Answer Agent · 120 cases
failing75.9%72.6
Long-context Retrieval Model Regression
Docs Chat · model: all candidate models · 318 cases
failing76.1%72.7
Structured Output Conformance Model Regression
Invoice AI · model: all candidate models · 429 cases
degraded78.3%78.6
Refund Policy Agent Trajectory Suite
Support Desk AI · agent: Refund Policy Agent · 78 cases
failing79.4%76.2
Meeting Agent Trajectory Suite
Recorder AI · agent: Meeting Agent · 48 cases
degraded83.1%78.4
Index Watcher Agent Trajectory Suite
Docs Chat · agent: Index Watcher Agent · 47 cases
degraded83.9%82.5
RAG Answer Golden Set
Docs Chat · prompt: RAG Answer · 70 cases
failing84.1%79.8
CRM Agent Trajectory Suite
CRM Copilot · agent: CRM Agent · 108 cases
degraded85.4%83.9
Interview Question Generator Golden Set
Resume AI · prompt: Interview Question Generator · 245 cases
failing85.9%86.0
Resume Rewrite Golden Set
Resume AI · prompt: Resume Rewrite · 219 cases
degraded87%88.0
Invoice Reconcile Agent Trajectory Suite
Invoice AI · agent: Invoice Reconcile Agent · 118 cases
passing89.9%89.5
Invoice Extraction Golden Set
Invoice AI · prompt: Invoice Extraction · 304 cases
passing90.6%84.6
Ticket Triage Golden Set
Support Desk AI · prompt: Ticket Triage · 99 cases
passing91.3%87.9
Answer Composer Golden Set
Support Desk AI · prompt: Answer Composer · 302 cases
passing91.9%86.9
Summarisation Benchmark Model Regression
Recorder AI · model: all candidate models · 634 cases
passing93.2%94.6
String Localisation Golden Set
Translate Pro · prompt: String Localisation · 296 cases
passing97.5%95.3
Clause Risk Review Golden Set
Legal AI · prompt: Clause Risk Review · 46 cases
passing99.4%97.5

Hallucinations

Recent groundedness failures

ClaimSeverityWhen
Claimed `validateSession` already checks the tenant id.
Code Review Agent · claude-4.5-opus · truth: `validateSession` only verifies the sig…
Critical14h ago
There is no auto-renewal provision in this contract.
Clause Risk Review · claude-4.5-opus · truth: Clause 4.2 auto-renews for successive 1…
High2d ago
Used `users.last_active_at` to define active users.
Natural Language to SQL · gpt-5 · truth: The canonical definition in metric_defi…
High2d ago
Suggested a test for `parseInvoiceTotals`, saying it is currently untested.
Test Gap Agent · deepseek-v4 · truth: The function has 14 existing cases in `…
Low2d ago
Decision: the billing migration ships on 14 March, approved by the CFO.
Meeting Summary · gpt-5 · truth: The CFO was not on the call. Priya prop…
High8d ago
Stated that annual plans are refundable pro-rata at any time.
Refund Policy Agent · gpt-5-mini · truth: Annual plans are refundable within 14 d…
Critical8d ago
The staging environment is redeployed automatically on every merge to develop.
RAG Answer · gpt-5-mini · truth: The runbook says staging deploys are ma…
Medium10d ago
Ad copy claimed "trusted by over 10,000 teams".
Marketing Copy Agent · claude-4.5-sonnet · truth: The brand guide caps the claim at 3,000…
High13d ago
Promised the export bug would be fixed in the next release.
Answer Composer · claude-4.5-sonnet · truth: No fix was scheduled; the linked issue …
Critical17d ago
Extracted a VAT rate of 20% for a Dutch supplier invoice.
Invoice Extraction · gpt-5-mini · truth: The invoice states 21% BTW; 20% was inf…
High18d ago