anthropic ·claude-sonnet-4-5Aug 25, 12:50 PM
Asset 88Lb01q
Score
0.89
Latency
5.28s
Cost
$0.0154
Workflow Eval Detail
Answers natural-language questions about a video by retrieving relevant context and answering with a concise response.
The workflow performed strongly overall, with Google leading quality and latency while GPT-5.6 Luna minimized cost.
Each eval run captures efficacy, efficiency, and expense. We use this data to compare providers and track regressions over time.
We score answer accuracy and response integrity while tracking latency, token usage, and cost.
| Provider | Model | Cases | Avg Score | Avg Latency | Avg Tokens | Avg Cost | Avg Cost / Min |
|---|---|---|---|---|---|---|---|
| anthropic | claude-sonnet-4-5 | 1 | 0.89 | 5.28s | 4,562 | $0.0154 | $0.0268/min |
| gemini-2.5-flash | 1 | 1 | 6.32s | 3,008 | $0.0017 | $0.003/min | |
| gemini-3-flash-preview | 1 | 0.95 | 5.81s | 3,972 | $0.0033 | $0.0058/min | |
| gemini-3.1-flash-lite | 1 | 0.97 | 3.09s | 3,598 | $0.0011 | $0.0019/min | |
| openai | gpt-5-mini | 1 | 0.85 | 12.52s | 4,764 | $0.0013 | $0.0022/min |
| openai | gpt-5.1 | 1 | 0.99 | 5.76s | 3,123 | $0.0052 | $0.0091/min |
| openai | gpt-5.6-luna | 1 | 0.94 | 3.48s | 4,377 | $0.0003 | $0.0005/min |