Generates Chinese/Japanese speech with StepFun's Contextual TTS — default stepaudio-2.5-tts, stepaudio-3-tts for whisper/inline-() prosody. Replaces step-tts-2's voice_label with natural-language instruction. Use for emotional/prosody-controlled synthesis, batch voice lines, migrating from step-tts-2, or cloned voices (2.5/step-tts-2, not v3). Triggers on 阶跃 TTS, 语音合成, 配音. Not for transcription (use stepfun-asr).
This skill hasn't been reviewed yet
bb6ad55
Table of Contents
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.