CtrlK
BlogDocsLog inGet started
Tessl Logo

uinaf/verify

Run the builder-owned pre-review verification pass for a completed change using repo guardrails and real-surface evidence. Use a separate evaluator for complex, subjective, or high-risk work; when the repo is not verifiable, report blocked with the exact missing infrastructure and required setup.

94

1.04x
Quality

95%

Does it follow best practices?

Impact

94%

1.04x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No known issues

Overview
Quality
Evals
Security
Files

criteria.jsonevals/scenario-1/

{
  "context": "Tests whether verify refuses to call a change ready when the repo has no reliable boot/test/runtime surface, reports exact attempted commands, and specifies the missing infrastructure instead of pretending static inspection is proof.",
  "type": "weighted_checklist",
  "checklist": [
    {
      "name": "Blocked verdict",
      "description": "The report verdict is exactly `blocked`, not `ready for review`, `ship it`, or an optimistic equivalent",
      "max_score": 16
    },
    {
      "name": "Commands attempted",
      "description": "The report lists exact commands attempted to find or run verification infrastructure, such as package scripts, tests, boot commands, or lockfile checks",
      "max_score": 12
    },
    {
      "name": "Missing infra evidence",
      "description": "The report includes concrete evidence that verification infrastructure is missing or unusable, not just a generic statement that there are no tests",
      "max_score": 12
    },
    {
      "name": "No static proof substitution",
      "description": "The report does not treat reading `src/worker.ts` or reasoning about the function as equivalent to runtime verification",
      "max_score": 14
    },
    {
      "name": "Readiness gaps listed",
      "description": "The report names the missing readiness pieces, such as lockfile/install path, test script, boot command, smoke check, or service entrypoint",
      "max_score": 12
    },
    {
      "name": "Infrastructure next steps",
      "description": "The required next steps identify the concrete infrastructure needed before verification can resume, such as a lockfile/install path, test script, boot command, smoke check, or service entrypoint",
      "max_score": 14
    },
    {
      "name": "Surfaces not exercised honestly",
      "description": "The report is honest about what was attempted and what could not be exercised, without requiring a verbose Surfaces Exercised section",
      "max_score": 10
    },
    {
      "name": "No invented tests",
      "description": "The solution does not add ad hoc tests or a custom harness solely to claim verification passed; it reports the repo readiness gap",
      "max_score": 10
    },
    {
      "name": "Compact blocked footer",
      "description": "The final blocked verification footer is no more than 5 labeled lines and does not repeat the missing-infra evidence after listing it once",
      "max_score": 8
    }
  ]
}

SKILL.md

tile.json