Workflow Eval Detail

Ask Questions

Answers natural-language questions about a video by retrieving relevant context and answering with a concise response.

Latest Runcompleted
muxinc/ai
maind35ae5a·@mux/ai v0.35.0
Cases
7
Avg Score
0.94
Avg Latency
6.04s
Avg Cost
$0.004
Avg Cost / Min
$0.007/min
Avg Tokens
3,915
TL;DR

The workflow performed strongly overall, with Google leading quality and latency while GPT-5.6 Luna minimized cost.

Best Quality
google
gemini-2.5-flash
Fastest
google
gemini-3.1-flash-lite
Most Economical
openai
gpt-5.6-luna

What we measure

Each eval run captures efficacy, efficiency, and expense. We use this data to compare providers and track regressions over time.

Efficacy
Quality + correctness
Efficiency
Latency + token usage
Expense
Cost per request

Workflow snapshot

Suite statussuccess
Suite average score0.94
Suite duration42.28s
Last suite runAug 25, 12:49 PM

Evaluation criteria

From eval tests

We score answer accuracy and response integrity while tracking latency, token usage, and cost.

Question set3 checks
Question
Has on-screen text?
Answer: YesConf0.94
Question
Is the speaker visible?
Answer: NoConf0.88
Question
Scene indoors?
Answer: YesConf0.91
Accuracy
Match expected yes/no outputs.
Format
Required fields + response shape.
Integrity
Confidence and reasoning present.
Throughput
Latency and cost stay within targets.
Response validation
Efficacy checks
  • Answers match expected yes/no outputs.
  • Response includes required fields and answer structure.
  • Confidence is 0-1 and reasoning strings are non-empty.
  • Response preserves asset ID and storyboard URL.
Efficiency targets
  • Latency: scores are normalized between 0 and 1. Under 8s earns 1.0; past 12s trends toward 0.
  • Token usage: scores are normalized between 0 and 1. Under 2,900 tokens earns 1.0; higher usage reduces the score.
  • Usage data must include total tokens for cost analysis.
Expense guardrails
  • Estimated cost under $0.012 per request for full score.

Provider breakdown

Run d35ae5a
Efficacy scoreHigher is better
LatencyLower is better
Token UsageLower is better
CostLower is better
ProviderModelCasesAvg ScoreAvg LatencyAvg TokensAvg CostAvg Cost / Min
anthropicclaude-sonnet-4-510.895.28s4,562$0.0154$0.0268/min
googlegemini-2.5-flash116.32s3,008$0.0017$0.003/min
googlegemini-3-flash-preview10.955.81s3,972$0.0033$0.0058/min
googlegemini-3.1-flash-lite10.973.09s3,598$0.0011$0.0019/min
openaigpt-5-mini10.8512.52s4,764$0.0013$0.0022/min
openaigpt-5.110.995.76s3,123$0.0052$0.0091/min
openaigpt-5.6-luna10.943.48s4,377$0.0003$0.0005/min

Recent cases

Latest run · 7 cases
anthropic ·claude-sonnet-4-5Aug 25, 12:50 PM
Asset 88Lb01q
Score
0.89
Latency
5.28s
Cost
$0.0154
google ·gemini-2.5-flashAug 25, 12:50 PM
Asset 88Lb01q
Score
1
Latency
6.32s
Cost
$0.0017
google ·gemini-3-flash-previewAug 25, 12:50 PM
Asset 88Lb01q
Score
0.95
Latency
5.81s
Cost
$0.0033
google ·gemini-3.1-flash-liteAug 25, 12:50 PM
Asset 88Lb01q
Score
0.97
Latency
3.09s
Cost
$0.0011
openai ·gpt-5-miniAug 25, 12:50 PM
Asset 88Lb01q
Score
0.85
Latency
12.52s
Cost
$0.0013
openai ·gpt-5.1Aug 25, 12:50 PM
Asset 88Lb01q
Score
0.99
Latency
5.76s
Cost
$0.0052
openai ·gpt-5.6-lunaAug 25, 12:50 PM
Asset 88Lb01q
Score
0.94
Latency
3.48s
Cost
$0.0003