| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| Agent success vs baseline Does it follow best practices? Average score across 7 eval scenarios Reviewed: Version: 1.0.0 | 1.64x | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |