CtrlK
BlogDocsLog inGet started
Tessl Logo

agui-dotnet-integration-tests

Write integration tests for the AG-UI .NET SDK. USE FOR: adding a new AG-UI event type and covering it end-to-end, testing SSE or protobuf streaming through the hosting pipeline, verifying AGUIChatClient maps events to ChatResponseUpdate, multi-turn conversation tests, parameterizing a test over Json/Protobuf transports, adding or updating a GettingStarted sample replay/Verify snapshot. Covers: WebApplicationFactory + DelegatingStreamingChatClient setup, IChatClient-based assertions with Assert.Collection, TransportFormat [Theory] (Json/Protobuf), recording/replay capture infrastructure, and the 8-capture-point Verify baselines. DO NOT USE FOR: unit tests of event serialization (use tests/AGUI.Abstractions.UnitTests) or stream-conversion unit tests (use tests/AGUI.Server.UnitTests).

74

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, execution-ready reference: complete code templates, a sequenced authoring procedure, an explicit baseline-review workflow, and sharp project-specific gotchas (namespace/folder mismatch, .NET-only events that force Json-only tests). The main improvement room is consolidation of a few repeated rules and a slightly leaner SKILL.md by moving sample-replay and baseline-naming detail into reference files.

Suggestions

Consolidate the repeated 'emit helpers yield ChatResponseUpdate, not BaseEvent' and 'assert on u.RawRepresentation' rules into one place (e.g., the Key rules section only) to tighten conciseness.

Add an explicit final step to the authoring procedure — run `dotnet test tests/AGUI.Hosting.AspNetCore.IntegrationTests/` and confirm both transport variants pass — so the workflow has a built-in validation checkpoint.

Consider moving the sample recording/replay mechanics and the 8-capture-point baseline naming table (~60 lines) into a references/ file linked from a short overview section, keeping SKILL.md itself leaner.

DimensionReasoningScore

Conciseness

The body is dense and project-specific with no filler or explanation of concepts Claude already knows (no "what is SSE" padding), and the layout tree and capture-point table are irreducible operational detail. It falls just short of the score-5 anchor because a few rules are repeated verbatim across sections — "yield ChatResponseUpdate, not BaseEvent" appears in Emit helpers, Procedure step 6, and Key rule 3, and "Assert on u.RawRepresentation" similarly recurs — which could be consolidated. It is clearly above score 3, which requires unnecessary explanation rather than deliberate cross-reference.

4 / 5

Actionability

Guidance is fully executable: complete copy-paste C# for class declaration, CreateClient with an inline handler, a full [Theory] test with Assert.Collection assertions, host registration snippets ("builder.Services.AddAGUI()", "ProtobufEventStreamFormatter"), the run command ("dotnet test tests/AGUI.Hosting.AspNetCore.IntegrationTests/"), and the baseline-accept command ("dotnet test --environment VERIFY_ACCEPT=true"). The worked tool-call example covers the common case end-to-end, matching the score-5 anchor exactly; score 4 would require missing key details that are not missing.

5 / 5

Workflow Clarity

The 'Procedure: writing a new integration test' section gives a clearly sequenced 6-step process, and the 'Updating baselines' section provides a review-then-accept feedback loop ("review the .received.json diff, then accept") plus "review them before committing" for new fixtures. It stops short of the score-5 anchor because the procedure itself lacks an explicit run-and-verify step (the dotnet test command lives only in the closing Build & run section) and error-recovery guidance is limited to the baseline flow. It exceeds score 3, which would require missing or only implicit checkpoints.

4 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are all absent), so this is a single-file skill; it opens with clear one-level-deep external pointers ("Read those first: sdks/dotnet/AGENTS.md ... docs/architecture.md") and its sections (Orientation, Project Layout, Architecture, Procedure, Sample replay, Key rules, Build & run) are well-organized for navigation. It falls short of the score-5 anchor because at ~245 lines, self-contained sub-topics like the sample recording/replay mechanics and the 8-capture-point baseline naming scheme are inlined in SKILL.md where a reference file would keep the overview leaner. It is above score 3, since what is inline is operational detail the skill needs in context and all pointers that do exist are clearly signaled, not buried.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a model example: a concrete 'what', an explicit and scenario-rich 'USE FOR' trigger list, concrete infrastructure coverage, and a 'DO NOT USE FOR' clause that disambiguates against sibling unit-test projects. All four dimensions sit at the top anchor.

DimensionReasoningScore

Specificity

Quotes like "Write integration tests for the AG-UI .NET SDK", "testing SSE or protobuf streaming through the hosting pipeline", "parameterizing a test over Json/Protobuf transports", and "adding or updating a GettingStarted sample replay/Verify snapshot" name multiple concrete actions, plus concrete infrastructure coverage ("WebApplicationFactory + DelegatingStreamingChatClient setup", "Assert.Collection", "the 8-capture-point Verify baselines"). This matches the score-5 anchor (multiple specific concrete actions, comprehensive) and is clearly above score 4, which would require minor coverage gaps that are not present.

5 / 5

Completeness

It explicitly answers what ("Write integration tests for the AG-UI .NET SDK" plus a "Covers:" list), when (an explicit "USE FOR:" clause with six concrete trigger scenarios), and even when-not ("DO NOT USE FOR: unit tests of event serialization..."). This is the score-5 anchor — both what and when stated with concrete trigger phrases — and exceeds score 4, where the 'when' would be less explicit; the 'USE FOR' clause also satisfies the rubric's requirement that a missing trigger clause would cap this dimension at 3.

5 / 5

Trigger Term Quality

Natural phrases a user would say are densely present: "integration tests", "SSE", "protobuf", "streaming", "multi-turn conversation tests", "Json/Protobuf transports", "replay", "Verify snapshot", "AG-UI event type", plus concrete redirect targets ("tests/AGUI.Abstractions.UnitTests", "tests/AGUI.Server.UnitTests"). Coverage is comprehensive for this domain, matching the score-5 anchor rather than score 4 ('a few natural terms missing'), since both the event-level and transport-level vocabularies users would invoke are present.

5 / 5

Distinctiveness Conflict Risk

The niche is sharply defined ("integration tests for the AG-UI .NET SDK") and the "DO NOT USE FOR" clause actively redirects the two nearest overlap cases — unit tests of event serialization and stream-conversion unit tests — to specific sibling test projects. That explicit disambiguation puts it at the score-5 anchor (clear niche, distinct triggers, minimal conflict risk) rather than score 4, which allows minor overlap risk with closely related skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ag-ui-protocol/ag-ui
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.