CtrlK
BlogDocsLog inGet started
Tessl Logo

uinaf/verify

Run the builder-owned pre-review verification pass for a completed change using repo guardrails and real-surface evidence. Use a separate evaluator for complex, subjective, or high-risk work; when the repo is not verifiable, report blocked with the exact missing infrastructure and required setup.

94

1.04x
Quality

95%

Does it follow best practices?

Impact

94%

1.04x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No known issues

Overview
Quality
Evals
Security
Files

criteria.jsonevals/scenario-2/

{
  "context": "Tests whether the agent actually starts the server and makes real HTTP requests to verify behavior — not just reads/analyzes the source code — and whether the report follows evidence rules: exact commands recorded, real responses captured, failure paths exercised, surfaces named, success kept terse, findings tied to impact, and a valid verdict produced.",
  "type": "weighted_checklist",
  "checklist": [
    {
      "name": "Server started and queried",
      "description": "Report shows evidence that the server was actually started (e.g. references to starting the process) and that HTTP requests were made to localhost:9127 — not just a static code review",
      "max_score": 15
    },
    {
      "name": "Real HTTP client used",
      "description": "Verification used curl or an equivalent real HTTP client (e.g. httpx, requests, wget) to exercise the server — not just calling Python functions directly or importing the module",
      "max_score": 12
    },
    {
      "name": "Actual response captured",
      "description": "Report includes actual HTTP response body content and status codes (e.g. '200 OK', '404', JSON response body) received from the live server",
      "max_score": 10
    },
    {
      "name": "Error/failure path exercised",
      "description": "At least one error or failure path was tested (e.g. requesting a non-existent user ID, sending malformed JSON to POST /users, or requesting an unknown route) and the response is documented",
      "max_score": 12
    },
    {
      "name": "Exact surfaces named",
      "description": "Report names the specific endpoints exercised (e.g. GET /users, GET /users/1, POST /users, GET /users/999) — not just 'the API was tested'",
      "max_score": 10
    },
    {
      "name": "Exact commands recorded",
      "description": "Report includes the actual commands run (e.g. the curl command lines with flags and URLs), not just a prose description of what was done",
      "max_score": 10
    },
    {
      "name": "Success kept terse",
      "description": "Report does NOT include verbose output from successful requests (e.g. does not dump full response headers for every passing call); successful checks are summarized briefly",
      "max_score": 8
    },
    {
      "name": "Finding tied to impact",
      "description": "At least one finding or observation in the report explains what could break, who is affected, or why it matters — not just a description of the behavior observed",
      "max_score": 8
    },
    {
      "name": "Valid verdict",
      "description": "Report contains a verdict using exactly one of: 'ready for review', 'needs more work', or 'blocked'. It does not use 'ship it' or make a ship-or-not decision.",
      "max_score": 15
    },
    {
      "name": "Compact verification footer",
      "description": "Report keeps the final verification footer to no more than 5 labeled lines, does not repeat detailed response bodies after recording them once, and summarizes successful evidence by command or surface name",
      "max_score": 8
    }
  ]
}

SKILL.md

tile.json