CtrlK
BlogDocsLog inGet started
Tessl Logo

data-scraper-agent

Build a fully automated AI-powered data collection agent for any public source — job boards, prices, news, GitHub, sports, anything. Runs on a schedule, enriches data with a free LLM (Gemini Flash), stores results in Notion/Sheets/Supabase, and learns from user feedback. Runs 100% free on GitHub Actions. Use when the user wants to monitor, collect, or track any public data automatically.

60

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/data-scraper-agent/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable — near-complete, executable code for every layer of the architecture — but it pays for that with significant verbosity (duplicated pattern sections, time-sensitive quota/version tables) and a monolithic structure with no reference files. The batch-oriented workflow also lacks runtime validation and retry loops, capping workflow clarity.

Suggestions

Deduplicate the "Common Scraping Patterns" section against Step 3 (the REST, HTML, and RSS patterns appear twice) and move the time-sensitive "Free Tier Limits Reference" and pinned requirements versions into a separate, clearly dated reference file so they can be updated without touching the core instructions.

Split the full code listings (ai/client.py, ai/pipeline.py, storage/notion_sync.py, the GitHub Actions workflow) into references/ files such as references/gemini-client.md, references/storage-sync.md, and references/workflow.md, keeping SKILL.md as an overview of the 10-step build with links to each.

Add explicit validation checkpoints to the workflow — verify AI-enriched rows against a sample before batch writes, retry or surface failed Notion pushes per item, and confirm feedback.json is valid JSON before the commit step — so the batch pipeline has a validate-and-retry loop rather than print-and-continue.

DimensionReasoningScore

Conciseness

The ~770-line body has several padded/duplicated sections: the REST/HTML/RSS patterns appear verbatim in both Step 3 and "Common Scraping Patterns", the "Free Tier Limits Reference" embeds time-sensitive RPM/RPD numbers and pinned library versions outside any deprecated/old-patterns section, and "Real-World Examples" plus the closing "Reference Implementation" paragraph restate content already covered. This is noticeably verbose rather than just occasionally loose — not severe enough to be pure padding since the bulk is genuinely useful code.

2 / 5

Actionability

Mostly copy-paste-ready: complete implementations for the Gemini client, batch pipeline, feedback memory, Notion sync, orchestrator, GitHub Actions workflow, config.yaml, and requirements. Minor gaps keep it below a 5 — `scraper/filters.py`'s `is_relevant` is referenced but never shown, `setup.py`/`enrich_existing.py` are promised in the tree and checklist without implementation, and the Sheets/Supabase storage alternatives are named but not provided.

4 / 5

Workflow Clarity

The 10-step sequence is clearly ordered with a quality checklist, but this is a batch skill (batched AI calls, batched database writes) and the runtime workflow lacks validation checkpoints: failed Notion pushes and failed sources are only printed, there is no validate-and-retry loop over stored rows, and no step verifies output before committing feedback history. Per the batch-operations cap, a batch skill without validation cannot score above 3.

3 / 5

Progressive Disclosure

Section headers and the step structure give reasonable in-file navigation, but the skill is a monolithic ~770-line SKILL.md with no references/ files at all — full per-file implementations, the scraping-pattern cookbook, and the free-tier limits table are all inlined content that clearly belongs in one-level-deep reference files. Not a 2 because headers and a consistent step layout keep it navigable, but the split is absent rather than merely imperfect.

3 / 5

Total

12

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete capabilities with named technologies, an explicit "Use when" trigger clause, and good natural-language trigger coverage. Its only notable gap is omission of the domain's most common user verb ("scrape"), which slightly narrows trigger-term coverage.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "Runs on a schedule, enriches data with a free LLM (Gemini Flash), stores results in Notion/Sheets/Supabase, and learns from user feedback" — with named technologies, giving comprehensive coverage of what the skill does. No vague filler; every clause names a concrete capability.

5 / 5

Completeness

Both parts are explicit: the "what" is a fully enumerated capability list (collect, schedule, enrich, store, learn), and the "when" is a dedicated clause — "Use when the user wants to monitor, collect, or track any public data automatically" — with concrete trigger phrasing.

5 / 5

Trigger Term Quality

Strong natural-phrase coverage: "monitor, collect, or track" plus concrete domains ("job boards, prices, news, GitHub, sports") users would actually say. However the most common user verb for this domain — "scrape" / "scraping" — is absent, and no near-synonyms ("crawler", "watcher", "alert on") appear, so a few natural terms are missing.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche — scheduled, free-stack, AI-enriched data collection agents — with triggers ("monitor/collect/track any public data automatically") that don't collide with one-off scraping or analysis skills. Over-claims like "any public source... anything" are broad within the niche but the trigger conditions remain distinct.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (777 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

Total

14

/

16

Passed

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.