← All sessionsOpen in Markdown reader
Model / Tool Evaluation Card
Candidate model/tool/version:
Task class: drafting / research / spreadsheet analysis / code assistance /
browser exploration / verdict review.
Allowed environment and data class:
Golden corpus version:
Route: fabricated mock / live measured trial. Mock results cannot approve a real model.
Input/prompt version and actor-access boundary (live route):
Sample counts and metric denominators: binary cases / known failures / known passes; score ambiguity and instruction-boundary cases separately.
For the supplied mock outputs, write not measured for task completion, evidence completeness, latency and cost. For live runs, link actual output and measurement records; unavailable values remain unknown.
| Metric |
Baseline |
Candidate |
Decision |
| Task completion |
|
|
|
| False pass |
|
|
|
| False fail |
|
|
|
| Evidence completeness |
|
|
|
| Policy/safety compliance |
|
|
|
| Median latency |
|
|
|
| Cost per run |
|
|
|
Known failure modes:
Human review requirement:
Approved prompt/playbook version:
Approval decision and reviewer: