anthropic ·claude-sonnet-4-5Aug 25, 12:50 PM
Asset 88Lb01q
Score
0.97
Latency
6.44s
Cost
$0.0129
Workflow Eval Detail
Generates concise summaries and smart tags from your content—perfect for search, discovery, and quick recaps.
Anthropic led on accuracy, Google led on latency, and OpenAI gpt-5.6-luna was the lowest-cost option, but comparisons are limited by the small sample.
Each eval run captures efficacy, efficiency, and expense. We use this data to compare providers and track regressions over time.
We score summary quality, tag relevance, and semantic similarity while tracking latency, token usage, and cost.
| Provider | Model | Cases | Avg Score | Avg Latency | Avg Tokens | Avg Cost | Avg Cost / Min |
|---|---|---|---|---|---|---|---|
| anthropic | claude-sonnet-4-5 | 4 | 0.99 | 6.4s | 3,840 | $0.0132 | $0.0226/min |
| gemini-2.5-flash | 4 | 0.99 | 8.19s | 3,151 | $0.0033 | $0.0059/min | |
| gemini-3-flash-preview | 4 | 0.98 | 8.21s | 3,427 | $0.0031 | $0.0053/min | |
| gemini-3.1-flash-lite | 4 | 0.99 | 2.88s | 2,980 | $0.0009 | $0.0016/min | |
| openai | gpt-5-mini | 4 | 0.94 | 12.99s | 4,436 | $0.0018 | $0.0036/min |
| openai | gpt-5.1 | 4 | 0.99 | 3.25s | 2,509 | $0.0026 | $0.0071/min |
| openai | gpt-5.6-luna | 3 | 0.99 | 4.4s | 3,827 | $0.0003 | $0.0005/min |