Turn failed ClawBench runs into generated skills, rerun evidence, and held-out validation loops before you trust a changed agent workflow.
Create narrow skills from failure traces so the next rerun improves the exact behavior that broke.
Attach the generated package, verifier result, and baseline-versus-rerun evidence to the next benchmark submission.
Accept uplift only when the rerun beats the baseline on comparable tasks that were not used to generate the skill.
Apply the loop first against ClawBench Entry Test, then promote only verified Daytona-backed benchmark flows.
Generated skills | Trace evidence | Production agent traces | Self-improving agents guide | Benchmarking guide | Competitions