Validate that an EAGLE3 pipeline run completed successfully end-to-end. Checks all 4 steps produced expected artifacts, verifies acceptance rate meets threshold (>= 2.1), and produces a summary report. Use when user wants to verify a pipeline run or check benchmark results.
74
92%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Verify that an EAGLE3 pipeline run completed successfully and meets quality criteria.
Find the most recent experiment directory (or ask the user for the path):
ls -td experiments/cicd/cicd_* | head -5Each experiment directory has one subdirectory per task (numbered 0–3), each containing a
log file whose name varies by launch mode (Slurm: sbatch_*.out, local Docker: *.log).
Match the log files generally and read the tail of each:
find experiments/<exp_id>/ -type f \( -name '*.out' -o -name '*.log' \) | sort | while read -r f; do
echo "=== $f ==="; tail -50 "$f"; echo
doneAll 4 tasks must complete without error. Look for:
exit code: 0 or no error — successDUE TO TIME LIMIT — timeoutFAILED / signal / exception traceback — failureIf any task failed, suggest running /eagle3-triage instead.
Check each step produced the expected output (artifacts live on the cluster at /scratchspace/).
Confirm via log messages:
| Step | Expected log evidence | Artifact |
|---|---|---|
| task_0 | "Saved N samples" or progress bar completing | /scratchspace/data/*.jsonl |
| task_1 | "Successfully processed N conversations" | /scratchspace/offline_hidden_states/*.pt |
| task_2 | Training loss decreasing, "export complete" | /scratchspace/eagle3/model.safetensors, /scratchspace/export/ |
| task_3 | Average Acceptance Length ... ratio: X.XX | JSON result files |
In the task_3 log, find:
Average Acceptance Length {'accept': X, 'count': Y, 'ratio': Z.ZZ}The ratio field is the acceptance rate (AR).
| Criterion | Threshold | Status |
|---|---|---|
| AR (MT-Bench) | >= 2.1 | PASS / FAIL |
If the log shows AR ... < lower bound, the run already triggered a threshold failure (exit code 1).
In the task_2 log look for:
training.ar_validate_steps was set)## EAGLE3 Pipeline Validation Report
**Experiment:** <exp_dir>
**Model:** <model_name>
**Date:** <date>
**Pipeline config:** <yaml_path>
### Step Status
| Step | Task | Status | Notes |
|------|------|--------|-------|
| 0 | Data synthesis | PASS/FAIL/TIMEOUT | N samples generated |
| 1 | Hidden state dump | PASS/FAIL | N .pt files |
| 2 | Training + export | PASS/FAIL | Final loss: X.XX |
| 3 | Benchmark | PASS/FAIL | AR: X.XX |
### Acceptance Rate
- MT-Bench AR: X.XX (threshold: >= 2.1) — PASS/FAIL
### Training Summary
- Final loss: X.XX
- Training steps: N
- AR during training: X.XX (if validated)
### Overall: PASS / FAIL
<one-line summary>If PASS:
If FAIL:
/eagle3-triage for diagnosis87c9f8c
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.