# AI Agent Benchmark Planning Template

## 1) Benchmark Scope
- Product/workflow being evaluated:
- Agent version:
- Date range:
- Owner:

## 2) Task Set
- Number of tasks:
- Task categories:
- Held-out tasks included? (yes/no)

## 3) Metrics
- Quality score target:
- Robustness score target:
- Security target (ASR max):
- Latency target:
- Cost target:

## 4) Run Protocol
- Number of runs per task:
- Randomization / prompt variants:
- Deterministic replay requirements:

## 5) Failure Categories
- Logic failure:
- Spec mismatch:
- Hallucination:
- Safety/policy breach:
- Runtime/tool failure:

## 6) Decision Rule
- Promote threshold:
- Rollback triggers:

## 7) Notes
- Key observations:
- Next experiment:
