CtrlK
BlogDocsLog inGet started
Tessl Logo

tts-voice-synthesis

影片与视频编辑、内容创作者在制作视频配音或有声书时,当需要克隆音色、生成情感化配音或流式实时语音合成请用此技能。支持1.7B高质量与0.6B快速双模型,一键实现多语言方言配音,让语音创作更高效更自然。

67

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized skill body: clear mode-based workflows, executable examples, and exemplary progressive disclosure with a clean one-level reference structure. Weaknesses are minor — a redundant goals section, one incorrect flag in the streaming example, and missing validation in the clone/streaming workflows.

Suggestions

Fix the streaming example to match the script's boolean flag: use '--streaming' instead of '--streaming true'.

Add an explicit validation checkpoint in mode 2 (generate a short test utterance with the cloned voice and verify similarity before reuse) and in mode 3 (verify merged output duration/quality).

Trim the '任务目标' section, which duplicates the frontmatter description, and remove filler steps like '确认待合成的文本内容'.

DimensionReasoningScore

Conciseness

The body is well-sectioned and assumes Claude's competence (no explanations of TTS concepts), but '任务目标' largely restates the frontmatter description and some workflow steps are filler ('确认待合成的文本内容', '确保分段自然,不会截断语义'). Below 5 due to this trimmable redundancy.

4 / 5

Actionability

Provides concrete, copy-paste bash commands for all four modes, with flags (verified against scripts) like --emotion, --speed, --pitch, --streaming. Not 5 because the streaming example passes '--streaming true' while tts_generate.py defines --streaming as a store_true flag, so that example fails as written.

4 / 5

Workflow Clarity

Each mode has a clear numbered sequence, and modes 1 and 4 include explicit validation steps ('验证输出', '如有需要,调整参数重新生成'). Modes 2 and 3 lack explicit quality checkpoints (e.g. verifying the cloned voice before reuse), which keeps it below 5.

4 / 5

Progressive Disclosure

The body is a true overview: a '资源索引' section lists all bundle files with purposes and links, and detailed material (model config, emotion guide, usage examples) lives in references/. All referenced paths exist, and references are one level deep — the .md files point only to scripts, never to further .md files.

5 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit use-when trigger clause, a clear target audience, and a concrete capability list. Its main weaknesses are marketing-style filler and missing common synonyms (TTS, 文字转语音) that would broaden natural trigger coverage.

Suggestions

Remove the marketing filler '一键实现…让语音创作更高效更自然' and replace it with concrete capability wording, e.g. '将文本转换为语音'.

Add common synonyms users naturally say — 'TTS', '文字转语音', '语音朗读', 'text to speech' — to broaden trigger coverage.

DimensionReasoningScore

Specificity

Names the domain (视频配音、有声书) and several concrete actions (克隆音色、生成情感化配音、流式实时语音合成、多语言方言配音、双模型选择). Falls short of 5 because of buzzword padding ('一键实现', '让语音创作更高效更自然') that adds no concrete capability.

4 / 5

Completeness

Explicitly answers both what (克隆音色、情感化配音、流式合成、多语言方言、双模型) and when ('当需要克隆音色、生成情感化配音或流式实时语音合成请用此技能'), with an additional audience/context clause ('在制作视频配音或有声书时').

5 / 5

Trigger Term Quality

Contains natural user-facing phrases like '视频配音', '有声书', '克隆音色', '流式实时语音合成', '方言配音'. Missing common synonyms such as '文字转语音', 'TTS', '朗读' that a user might naturally say.

4 / 5

Distinctiveness Conflict Risk

Voice cloning, emotional dubbing, and streaming TTS are distinct triggers unlikely to fire for unrelated skills. Minor overlap risk with generic audio-editing or speech skills keeps it below 5.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
anbeime/skill
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.