CtrlK
BlogDocsLog inGet started
Tessl Logo

spark-optimization

Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.

79

1.28x
Quality

75%

Does it follow best practices?

Impact

77%

1.28x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./tests/ext_conformance/artifacts/agents-wshobson/data-engineering/skills/spark-optimization/SKILL.md

The canonical home for this skill is spark-optimization in wshobson/agents

SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Rich, concrete, and largely executable Spark optimization reference material, weakened by the absence of any diagnostic workflow with validation checkpoints, no progressive disclosure into reference files, and some re-explanation of concepts Claude already knows.

Suggestions

Add a sequenced diagnostic workflow (inspect Spark UI / stage metrics → identify bottleneck: skew, spill, shuffle → apply the matching pattern → re-measure and verify improvement) so the patterns are applied in a validated order rather than presented as a catalog.

Split the configuration cheat sheet and the monitoring/debugging helpers into reference files (e.g., references/config.md, references/monitoring.py) and link them one level deep, keeping SKILL.md as a lean overview.

Fix the invalid `spark.sql.shuffle.partitions = "auto"` setting (it is an integer config; use AQE's advisoryPartitionSizeInBytes instead) and trim the execution-model diagram and factors table, which restate knowledge Claude already has.

DimensionReasoningScore

Conciseness

Largely code-forward and efficient, but spends tokens on concepts Claude already knows: the Driver→Job→Stage→Task execution-model diagram, a Key Performance Factors table that restates the later patterns, and commented storage-level glosses.

3 / 5

Actionability

Nearly all code is copy-paste-ready with concrete config values, but `spark.conf.set("spark.sql.shuffle.partitions", "auto")` is an invalid value for an integer config and would throw at runtime, and the memory monitor uses private `_jsc` APIs — minor executable gaps.

4 / 5

Workflow Clarity

The body is a pattern catalog with no sequenced tuning workflow — no 'diagnose via Spark UI → identify skew/spill → apply pattern → verify improvement' ordering, and no validation checkpoints despite batch/overwrite write operations.

3 / 5

Progressive Disclosure

A single ~420-line SKILL.md with no bundle files and no offloading; the configuration cheat sheet and monitoring/debug scripts clearly belong in separate reference files, though internal section structure is well organized.

3 / 5

Total

13

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that pairs concrete capabilities with an explicit, natural 'Use when' trigger clause. Trigger coverage is good but could add common synonyms such as PySpark or Databricks.

DimensionReasoningScore

Specificity

Lists four concrete, specific capabilities — "partitioning, caching, shuffle optimization, and memory tuning" — which comprehensively covers the Spark optimization domain rather than exhibiting minor gaps.

5 / 5

Completeness

Explicitly answers both 'what' (optimize Spark jobs with partitioning, caching, shuffle optimization, memory tuning) and 'when' ("Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines") with concrete trigger phrases.

5 / 5

Trigger Term Quality

Includes natural phrases users would say ("debugging slow jobs", "improving Spark performance", "scaling data processing pipelines") but misses common variations like "PySpark", "Databricks", or "executor" issues.

4 / 5

Distinctiveness Conflict Risk

"Apache Spark" establishes a clear niche with distinct triggers; minimal risk of conflicting with other data-processing skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Dicklesworthstone/pi_agent_rust
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.