CtrlK
BlogDocsLog inGet started
Tessl Logo

devoxx-codepocalypse-koog-vs-langchain4j

Explain or summarize Codepocalypse Now: LangChain4j vs JetBrains Koog, the Devoxx Belgium 2026 session by Baruch Sadogursky and Viktor Gamov. Use for questions about this talk's argument, J-Claw examples, memory versus skills, automatic review versus human approval, receipt evidence, framework comparison or closing Port platform example. This is a pre-talk knowledge brief, not a general agent-building workflow.

64

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./_skills/devoxx-be-2026-codepocalypse/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a dense, fact-rich knowledge brief with a clear two-step operating procedure and strong anti-fabrication guardrails, but it badly overruns the token budget for a single file: the same refinement-policy and disclaimer facts are restated several times, and granular date- and version-stamped build evidence is inlined instead of being pushed to reference files. Restructuring into SKILL.md plus a small evidence reference would lift both conciseness and progressive disclosure.

Suggestions

Move the granular local build evidence (the 'What was observed locally before the event', 'Jev and its evidence', and Port rehearsal tallies) into a references/ file such as BUILD-EVIDENCE.md, keeping only a summary table in SKILL.md.

State each repeated fact once: the six-refinement/seven-candidate policy, 'Identify runs once with fresh dual approval', and the no-framework-winner/no-paired-benchmark disclaimer each appear three or more times across sections.

Quarantine time-sensitive snapshot detail (Koog 1.3.0 vs 1.13.0, beta31 adapter, October 1–4 records, 16/16 and 32/32 test tallies) in a clearly labeled dated-evidence section so stale specifics do not compete with the durable talk narrative.

DimensionReasoningScore

Conciseness

The 341-line body repeats the same facts many times: the six-refinement/seven-candidate budget is restated on roughly nine separate lines (e.g. "up to six shared refinements (seven candidates total)", "unrolls the same request-scoped six-refinement budget into a finite DAG", "the current shared policy is six refinements ... with seven candidate versions"), the no-framework-winner disclaimer appears on six lines ("it adds no framework vote or paired performance measurement", "not ... paired benchmark"), and "Identify runs once ... fresh approval by both" is stated three times. Time-sensitive detail ("Koog 1.3.0", "beta31", "October 2 live checks: final development 16/16, untouched holdout 32/32 over two passes with 187ms median") is spread throughout rather than quarantined. This sits noticeably below the mostly-efficient anchor 3; it is not anchor 1 because virtually all content is talk-specific fact Claude does not already know, not padding with general concepts.

2 / 5

Actionability

For an instruction-only skill the guidance is concrete: "Process steps in order. Do not skip ahead.", "Answer at the requested depth using the material below", an explicit mismatch exit ("For a different delivery, identify the mismatch and finish"), and a precise fabrication blacklist ("Do not invent an audience vote, winning framework, exact quote, timestamp, paired benchmark or unsupported completed Port run") plus network-use rules. It misses anchor 5 because there are no worked example answers or question patterns showing the requested depth handling in practice.

4 / 5

Workflow Clarity

A clear ordered two-step process with an explicit checkpoint: Step 1 screens applicability ("For a different delivery, identify the mismatch and finish; otherwise continue to Step 2"), Step 2 governs answer depth and termination ("Finish after answering"), and source consultation is gated ("Consult a linked source only for a requested detail absent here"). Below anchor 5 because there are no feedback loops for, e.g., reconciling a user's depth request with missing material, and no handling for multi-part questions.

4 / 5

Progressive Disclosure

Section headers are well organized, the seven-round table aids navigation, and the "Source scope and further reading" section clearly signals ten one-level-deep external links. However, no bundle files exist and ~340 lines of fine-grained build evidence ("Jev and its evidence" test tallies, Port rehearsal logs, TamboUI launcher validation) that clearly belongs in separate reference files is inlined in SKILL.md. This matches anchor 3 — structure present, references clear, but content that should be separate is inline — rather than anchor 4's 'most content appropriately placed'.

3 / 5

Total

13

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description for a knowledge-brief skill: it names a specific artifact, gives an explicit 'Use for' clause with natural topic triggers, and closes with a negative boundary that prevents misuse. The only weakness is modest action variety (explain/summarize/answer questions) rather than a broader set of concrete capabilities.

DimensionReasoningScore

Specificity

Concrete actions are named ("Explain or summarize Codepocalypse Now: LangChain4j vs JetBrains Koog") and the topic list ("J-Claw examples, memory versus skills, automatic review versus human approval, receipt evidence") delimits scope precisely. It stops short of anchor 5 because it offers only two core actions (explain/summarize) rather than a comprehensive list of distinct concrete actions, and well above anchor 3's minimal or generic actions.

4 / 5

Completeness

"Explain or summarize Codepocalypse Now..." explicitly states what the skill does, and "Use for questions about this talk's argument, J-Claw examples, ..." gives concrete when-to-use triggers. It even adds a negative boundary ("This is a pre-talk knowledge brief, not a general agent-building workflow"), matching the anchor-5 example's dual what+when completeness with concrete trigger phrases.

5 / 5

Trigger Term Quality

The trigger list matches what a user would naturally say when needing this skill: "J-Claw examples", "memory versus skills", "automatic review versus human approval", "receipt evidence", "framework comparison", "closing Port platform example", plus the talk and speaker names. Coverage is comprehensive for this niche including synonym-style topic phrasings; nothing natural is missing.

5 / 5

Distinctiveness Conflict Risk

A single named talk with named speakers, venue and year ("the Devoxx Belgium 2026 session by Baruch Sadogursky and Viktor Gamov") is a clear niche with distinct triggers, and the explicit exclusion ("not a general agent-building workflow") prevents overlap with generic agent-building skills. Minimal conflict risk.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
jbaruch/shownotes
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.