CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/bug-report-template

Builds a well-formed bug (defect) report from raw observation notes - fills in summary, environment, steps to reproduce, expected vs actual, and severity rationale - and validates that each field has the load-bearing content reviewers and engineers need to triage. Also converts a single test-failure record (JUnit XML, Allure JSON, pytest log, Playwright report) into a classified, ready-to-file bug spec, and provides the adversarial review checklist that gates a report before it enters the tracker (required fields, single-description title test, severity-priority independence, reproduction quality). Use when a stakeholder reports a problem informally, when a CI failure artefact needs to become a triageable report, or when a drafted report needs a pre-filing quality audit.

81

0.88x
Quality

95%

Does it follow best practices?

Impact

80%

0.88x

Average score across 10 eval scenarios

SecuritybySnyk

High

Do not use without reviewing

Overview
Quality
Evals
Security
Files

Evaluation results

100%

3%

Customer says the amount on her transfer doesn't match her statement

Criteria
Baseline
With context

Deliverable exists at the exact path

100%

100%

No invented monetary figures

100%

100%

No invented environment facts

100%

100%

Outstanding questions are enumerated

100%

100%

The frequency contradiction is surfaced

100%

100%

What happened and what should have happened are stated separately

83%

100%

Second-account claim is marked unverified

100%

100%

MUST NOT assert consistent reproduction from one observation

100%

100%

Title names the surface and the observable behaviour

87%

100%

90%

-2%

Integration partner emails that our API is rate-limiting them

Criteria
Baseline
With context

Deliverable exists at the exact path

100%

100%

The prose-versus-response contradiction is the report's headline finding

96%

100%

MUST NOT merge impact and queue position into one judgement

95%

100%

No invented account or client facts

100%

100%

Timestamps carried as ambiguous

92%

100%

Request identifier preserved exactly

100%

100%

Outstanding items listed with the right owner

68%

62%

Title describes the observed failure, not the reporter's diagnosis

80%

30%

88%

6%

Night shift says the handheld counts some totes twice

Criteria
Baseline
With context

Deliverable exists at the exact path

100%

100%

MUST NOT propose filing a new item without checking the open ones

100%

100%

No invented fleet, version, or location

84%

96%

Match is stated as probable, not confirmed

43%

62%

Questions phrased for a two-minute answer from the floor

81%

100%

The mis-pallet consequence is recorded and marked second-hand

83%

100%

Frequency stated as unquantified

100%

66%

User impact and scheduling order reasoned separately

60%

70%

97%

Nightly checkout suite went red and nobody has written it up

Criteria
Baseline
With context

Deliverable exists at the exact path

100%

100%

MUST NOT carry the hand-typed assertion instead of the result file

100%

100%

The chat-versus-artefact contradiction is named

100%

100%

No fabricated revision identifier

100%

100%

Re-run instruction is derived only from recorded facts

93%

81%

Speculation is attributed, not adopted

85%

100%

Run history stated only as far as it is known

100%

100%

Environment carried from the artefact

100%

100%

77%

-21%

Office manager's email lists several things wrong with the payroll portal

Criteria
Baseline
With context

Deliverable exists at the exact path

100%

100%

MUST NOT combine the unrelated failures into one item

100%

100%

No invented technical detail per item

100%

0%

Per-item questions for the customer

100%

100%

The onset contradiction is surfaced on the slow-page item

100%

100%

Items are cross-linked as one intake

100%

100%

Impact and scheduling urgency treated as separate judgements

100%

83%

Titles are single-clause and behavioural

75%

100%

100%

5%

One-star review says the app closes itself at checkout

Criteria
Baseline
With context

Deliverable exists at the exact path

100%

100%

MUST NOT fill environment fields that no source supplies

100%

100%

Version conflict is surfaced rather than resolved

100%

100%

Reviewer's environment separated from the teammate's

100%

100%

Outstanding items listed with an owner

100%

100%

Reproduction described honestly for each source

100%

100%

Steps are numbered and start from a known state

50%

100%

Crash signature carried verbatim

100%

100%

58%

-28%

Owner says his thermostat shows a different temperature than the room

Criteria
Baseline
With context

Deliverable exists at the exact path

100%

100%

MUST NOT populate device fields the sources leave blank

100%

0%

No invented temperature values

100%

100%

Temperature scale flagged as unknown

73%

0%

Outstanding items listed and matched to Thursday's decision

100%

100%

Reference thermometer treated as unverified

41%

100%

Onset and frequency stated as the owner stated them

80%

30%

Observed and expected behaviour stated separately

60%

100%

0%

-89%

Patient can't download her results letter from the portal

Criteria
Baseline
With context

Deliverable exists at the exact path

100%

0%

MUST NOT record the failure as consistently reproducible

100%

0%

All three frequency statements are surfaced together

100%

0%

'Empty' is quoted, not translated into a technical claim

80%

0%

No invented environment

100%

0%

Family member's attempt marked as not comparable

92%

0%

Source reliability is stated

100%

0%

Outstanding questions listed for the next patient contact

25%

0%

97%

24%

Travel agent can't get the seat map to appear for a client's booking

Criteria
Baseline
With context

Deliverable exists at the exact path

100%

100%

MUST NOT convert 'latest Chrome' into a stated version

0%

100%

The leg contradiction is surfaced

100%

100%

No reconstructed URL or booking identifier

100%

100%

Single consolidated question list for the agency

100%

81%

Colleague's result recorded as uncertain

100%

100%

Observed state described from the evidence, not diagnosed

75%

100%

Steps are numbered and their unknown starting point is marked

70%

100%

100%

3%

Trial customer says his call notes vanished, wants it fixed today

Criteria
Baseline
With context

Deliverable exists at the exact path

100%

100%

MUST NOT collapse user impact and fix order into one judgement

100%

100%

The lost-versus-not-lost contradiction is surfaced

100%

100%

No invented environment or timeline

100%

100%

Enterprise-customer report marked as unconfirmed

100%

100%

Outstanding facts listed as blocking questions

81%

100%

Reproduction status stated honestly

100%

100%

Title states the observed behaviour, not the customer's demand

100%

100%

Evaluated
Agent
Claude Code
Model
Claude Sonnet 4.6