| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| Agent success vs baseline Does it follow best practices? Average score across 10 eval scenarios Reviewed: Version: 46.0.0 | 1.18x | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |
| No change in agent success vs baseline Does it follow best practices? Average score across 10 eval scenarios Reviewed: Version: 2.1.0 | 1.00x | — | |
| | Pending | — | |
| | Pending | — | |
| | Pending | — | |