jev / gallery
← Back to collection

Author test

jev-agent-failure-benchmark

Attribute failures in multi-agent traces

How this project uses Jev

Input
Failure traces and candidate agents, steps, and error types
Jev decides
Responsible agent, key step, and error category
Code executes
The official scorer evaluates predictions

Evidence and limitations

The author uses an injected-error dataset. Some comparisons come from papers; constrained Who/When choices are not equivalent to free generation.

This project has not been run independently here. Author-reported results are not independently verified results.

Original sources

Explore related use cases

Browse this category →