Workflow Eval Detail

Summarization

Generates concise summaries and smart tags from your content—perfect for search, discovery, and quick recaps.

Latest Runcompleted
muxinc/ai
maind35ae5a·@mux/ai v0.35.0
Cases
27
Avg Score
0.98
Avg Latency
6.7s
Avg Cost
$0.0037
Avg Cost / Min
$0.0066/min
Avg Tokens
3,439
TL;DR

Anthropic led on accuracy, Google led on latency, and OpenAI gpt-5.6-luna was the lowest-cost option, but comparisons are limited by the small sample.

Best Quality
anthropic
claude-sonnet-4-5
Fastest
google
gemini-3.1-flash-lite
Most Economical
openai
gpt-5.6-luna

What we measure

Each eval run captures efficacy, efficiency, and expense. We use this data to compare providers and track regressions over time.

Efficacy
Quality + correctness
Efficiency
Latency + token usage
Expense
Cost per request

Workflow snapshot

Suite statussuccess
Suite average score0.95
Suite duration52.4s
Last suite runAug 25, 12:48 PM

Evaluation criteria

From eval tests

We score summary quality, tag relevance, and semantic similarity while tracking latency, token usage, and cost.

...and so when we look at the numbers...
...the growth trajectory has been...
...our teams have worked incredibly hard...
...which brings me to the next point...
Analyzing
AI Summary8:00 duration
6 tags extracted
Complete
Semantic Extraction
Efficacy checks
  • Title is non-empty, <=100 chars, and avoids filler starters.
  • Description is non-empty, <=1000 chars, and avoids meta phrases.
  • Tags are non-empty strings, unique, and <=10 items.
  • Title, description, and tags are semantically similar to references.
  • Response includes asset ID and HTTPS storyboard URL.
Efficiency targets
  • Latency: scores are normalized between 0 and 1. Under 8s earns 1.0; past 20s trends toward 0.
  • Token usage: scores are normalized between 0 and 1. Under 4,000 tokens earns 1.0; higher usage reduces the score.
  • Usage data must include input and output tokens > 0.
Expense guardrails
  • Estimated cost under $0.015 per request for full score.

Provider breakdown

Run d35ae5a
Efficacy scoreHigher is better
LatencyLower is better
Token UsageLower is better
CostLower is better
ProviderModelCasesAvg ScoreAvg LatencyAvg TokensAvg CostAvg Cost / Min
anthropicclaude-sonnet-4-540.996.4s3,840$0.0132$0.0226/min
googlegemini-2.5-flash40.998.19s3,151$0.0033$0.0059/min
googlegemini-3-flash-preview40.988.21s3,427$0.0031$0.0053/min
googlegemini-3.1-flash-lite40.992.88s2,980$0.0009$0.0016/min
openaigpt-5-mini40.9412.99s4,436$0.0018$0.0036/min
openaigpt-5.140.993.25s2,509$0.0026$0.0071/min
openaigpt-5.6-luna30.994.4s3,827$0.0003$0.0005/min

Recent cases

Latest run · 27 cases
anthropic ·claude-sonnet-4-5Aug 25, 12:50 PM
Asset 88Lb01q
Score
0.97
Latency
6.44s
Cost
$0.0129
anthropic ·claude-sonnet-4-5Aug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
6.73s
Cost
$0.0131
anthropic ·claude-sonnet-4-5Aug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
5.76s
Cost
$0.0129
anthropic ·claude-sonnet-4-5Aug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
6.66s
Cost
$0.0138
google ·gemini-2.5-flashAug 25, 12:50 PM
Asset 88Lb01q
Score
0.97
Latency
8.46s
Cost
$0.0034
google ·gemini-2.5-flashAug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
8.3s
Cost
$0.0036
google ·gemini-2.5-flashAug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
8.1s
Cost
$0.0028
google ·gemini-2.5-flashAug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
7.9s
Cost
$0.0033
google ·gemini-3-flash-previewAug 25, 12:50 PM
Asset 88Lb01q
Score
0.96
Latency
7.01s
Cost
$0.0031
google ·gemini-3-flash-previewAug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
7.28s
Cost
$0.0026
google ·gemini-3-flash-previewAug 25, 12:50 PM
Asset 88Lb01q
Score
0.98
Latency
10.96s
Cost
$0.0033
google ·gemini-3-flash-previewAug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
7.61s
Cost
$0.0036
google ·gemini-3.1-flash-liteAug 25, 12:50 PM
Asset 88Lb01q
Score
0.96
Latency
3.08s
Cost
$0.0009
google ·gemini-3.1-flash-liteAug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
2.59s
Cost
$0.0009
google ·gemini-3.1-flash-liteAug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
2.94s
Cost
$0.0009
google ·gemini-3.1-flash-liteAug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
2.91s
Cost
$0.0009
openai ·gpt-5-miniAug 25, 12:50 PM
Asset 88Lb01q
Score
0.91
Latency
14.6s
Cost
$0.002
openai ·gpt-5-miniAug 25, 12:50 PM
Asset 88Lb01q
Score
0.95
Latency
12.47s
Cost
$0.0016
openai ·gpt-5-miniAug 25, 12:50 PM
Asset 88Lb01q
Score
0.95
Latency
12.54s
Cost
$0.0018
openai ·gpt-5-miniAug 25, 12:50 PM
Asset 88Lb01q
Score
0.95
Latency
12.37s
Cost
$0.0017
openai ·gpt-5.1Aug 25, 12:50 PM
Asset 88Lb01q
Score
0.95
Latency
3.64s
Cost
$0.0041
openai ·gpt-5.1Aug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
3.07s
Cost
$0.0017
openai ·gpt-5.1Aug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
3.15s
Cost
$0.0017
openai ·gpt-5.1Aug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
3.16s
Cost
$0.0027
openai ·gpt-5.6-lunaAug 25, 12:50 PM
Asset 88Lb01q
Score
0.96
Latency
5.42s
Cost
$0.0003
openai ·gpt-5.6-lunaAug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
3.4s
Cost
$0.0003
openai ·gpt-5.6-lunaAug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
4.39s
Cost
$0.0003