CtrlK
BlogDocsLog inGet started
Tessl Logo

sn-da-large-file-analysis

万行以上 Excel 数据集的高性能分析引擎。提供 openpyxl read_only 流式读取(iter_rows 支持 10 万行以上)、Parquet 转换加速、内存优化、分块处理和大文件写入模式。**遇到以下任一情况就主动使用本 skill**:①数据行数 ≥ 10k(由 sn-da-excel-workflow 的行数评估步骤触发);②用户出现触发词:大文件 / 大数据量 / 性能优化 / 内存不足 / OOM / 百万行 / 十万行 / 流式读取 / Parquet / 分块处理 / large file / big data / streaming read / chunked processing;③直接使用 pd.read_excel() 导致超时或内存溢出;④用户明确要求对大规模数据集进行高性能处理。仅不用于:小于 10k 行的常规 Excel 分析(使用 sn-da-excel-workflow 即可)。

72

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with complete executable code and a clear scale-based workflow, but it is a large monolithic file with no progressive disclosure and minor duplication of guidance. Strengthening it would mean splitting reference material into bundle files and adding explicit validation feedback loops.

Suggestions

Externalize the Vectorized Operations Cheat Sheet, Memory Estimation table, and the two long worked examples into reference files under references/ to apply genuine progressive disclosure and slim the main SKILL.md.

Add an explicit fix→retry feedback loop for the streaming-to-Parquet path (e.g., verify row count written == inspected row count; re-run or fall back on mismatch) to strengthen validation for these memory-intensive batch operations.

De-duplicate the data-scale strategy and the prohibited-operations list, which appear in both the Mandatory Rules section and later in Best Practices / the cheat sheet.

DimensionReasoningScore

Conciseness

The body is dense, code-first, and largely assumes Claude's competence, but the CJK font-setup block is tangential to the core purpose and the data-scale table / prohibited list are duplicated in Best Practices and the cheat sheet—minor trimmable instances per the score-4 anchor rather than the lean score-5 anchor.

4 / 5

Actionability

Provides fully executable, copy-paste-ready functions (inspect_excel, stream_excel_to_parquet, optimize_dtypes, write_large_excel) plus two complete worked examples for 100k and 1M rows, matching the score-5 anchor.

5 / 5

Workflow Clarity

A clear row-count-keyed strategy table and numbered example steps with pre-flight validation (inspect-before-load, memory estimation) are present, but there are no explicit fix→retry feedback loops for the memory-intensive batch operations, leaving minor validation gaps per the score-4 anchor.

4 / 5

Progressive Disclosure

No bundle files exist and the entire ~365-line skill is inlined under good section headers, but reference-based disclosure is absent and content like the cheat sheet and long examples could be externalized, fitting the score-3 anchor better than score-4 (which expects clearly signaled references).

3 / 5

Total

16

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: it states concrete capabilities, enumerates natural trigger terms in both languages, gives explicit use-when guidance with four conditions, and cleanly delineates scope from a sibling skill. It hits the top anchor on every dimension.

DimensionReasoningScore

Specificity

Lists multiple concrete actions—"openpyxl read_only 流式读取", "Parquet 转换加速", "内存优化", "分块处理", "大文件写入模式"—giving comprehensive coverage with no gaps, matching the score-5 anchor.

5 / 5

Completeness

Explicitly answers both what (streaming read/Parquet/memory/chunking/writing) and when ("遇到以下任一情况就主动使用本 skill" with four numbered trigger conditions) plus a negative scope clause, matching the score-5 anchor.

5 / 5

Trigger Term Quality

Comprehensive natural-term coverage across both Chinese and English—"大文件 / 大数据量 / 内存不足 / OOM / 百万行 / large file / big data / streaming read / chunked processing"—including synonyms, matching the score-5 anchor.

5 / 5

Distinctiveness Conflict Risk

Clear niche (≥10k-row Excel analysis) with an explicit boundary against sibling skill sn-da-excel-workflow for <10k rows and distinct scale-based triggers, giving minimal conflict risk per the score-5 anchor.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
OpenSenseNova/SenseNova-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.