anthropic ·claude-sonnet-4-5Aug 25, 12:50 PM
Asset gIRjPqM
Score
1
Latency
3s
Cost
$0.0088
Workflow Eval Detail
Analyzes video frames to detect hardcoded captions baked into the visual content—useful for compliance checks and accessibility audits.
The workflow scored 0.9683 overall, with Anthropic leading quality, OpenAI gpt-5.1 leading latency, and OpenAI gpt-5.6-luna offering the lowest cost.
Each eval run captures efficacy, efficiency, and expense. We use this data to compare providers and track regressions over time.
We evaluate caption detection accuracy, confidence calibration, and response integrity alongside speed and cost thresholds.
| Provider | Model | Cases | Avg Score | Avg Latency | Avg Tokens | Avg Cost | Avg Cost / Min |
|---|---|---|---|---|---|---|---|
| anthropic | claude-sonnet-4-5 | 3 | 1 | 2.72s | 3,179 | $0.0099 | $0.0345/min |
| gemini-2.5-flash | 3 | 0.99 | 6.27s | 2,599 | $0.0026 | $0.0092/min | |
| gemini-3-flash-preview | 3 | 0.99 | 5.74s | 3,170 | $0.003 | $0.0106/min | |
| gemini-3.1-flash-lite | 3 | 1 | 2.36s | 2,624 | $0.0007 | $0.0024/min | |
| openai | gpt-5-mini | 3 | 0.9 | 17.14s | 3,649 | $0.0023 | $0.0082/min |
| openai | gpt-5.1 | 3 | 0.9 | 1.73s | 2,279 | $0.0022 | $0.0077/min |
| openai | gpt-5.6-luna | 3 | 1 | 4.56s | 2,871 | $0.0003 | $0.0012/min |