anthropic ·claude-sonnet-4-5Aug 25, 12:50 PM
Asset 1XIUcA9
Score
0.97
Latency
4.47s
Cost
$0.013
Workflow Eval Detail
Automatically segments long-form video content into navigable chapters with timestamps and titles—enabling viewers to jump to key moments instantly.
Chapters performed strongly overall, with gpt-5.1 best for quality, gemini-3.1-flash-lite best for latency, and gpt-5.6-luna best for cost.
Each eval run captures efficacy, efficiency, and expense. We use this data to compare providers and track regressions over time.
We evaluate chapter segmentation quality, timestamp accuracy, and title relevance alongside latency and cost metrics.
| Provider | Model | Cases | Avg Score | Avg Latency | Avg Tokens | Avg Cost | Avg Cost / Min |
|---|---|---|---|---|---|---|---|
| anthropic | claude-sonnet-4-5 | 5 | 0.99 | 4.31s | 3,831 | $0.013 | $0.0015/min |
| gemini-2.5-flash | 5 | 0.98 | 6.77s | 4,923 | $0.004 | $0.0004/min | |
| gemini-3-flash-preview | 5 | 0.96 | 10.01s | 5,395 | $0.0071 | $0.0011/min | |
| gemini-3.1-flash-lite | 5 | 0.99 | 1.61s | 3,794 | $0.0012 | $0.0001/min | |
| openai | gpt-5-mini | 5 | 0.93 | 18.05s | 4,502 | $0.0027 | $0.0003/min |
| openai | gpt-5.1 | 4 | 0.99 | 2.8s | 3,362 | $0.0037 | $0.0006/min |
| openai | gpt-5.6-luna | 3 | 0.98 | 4.17s | 3,418 | $0.0003 | $0.00/min |