Loading...
Loading...
Found 37 Skills
Interacting with Mediabunny
Multimedia handling with the Mediabunny library
Generate AI music with Producer via AceDataCloud API. Use when creating songs, generating lyrics, extending tracks, creating covers, swapping vocals/instrumentals, replacing song sections, or uploading reference audio. Supports custom lyrics, instrumental-only mode, and multiple creative actions.
Music generation queueing, retrieval, and completion endpoints via Venice.ai. Suited for jingles, background loops, and prototype scoring.
MediaKit is a professional toolset for audio-visual and image processing, covering workflows such as video editing and synthesis, audio processing, video understanding and enhancement, image processing and content understanding. When users clearly propose goals like editing, splicing, cropping, transitions, filters, camera movements, audio mixing, subtitle extraction, speech-to-subtitle, audio-visual processing, image processing, video analysis or image quality enhancement, load this Skill first, then select audio, editing, image or video according to the target object and goal. It does not provide explanations of specific capability parameters.
Master the essential audio post-production techniques—normalization, compression, EQ, and noise reduction—using the correct processing order to achieve professional-quality audio. Use when: Editing podcast episodes or video soundtracks; Cleaning up recorded voiceovers; Improving audio quality for marketing content; Preparing audio files for distribution; Troubleshooting common audio issues
Use Chanjing TTS API to synthesize speech from text, using user-provided voice
Transcribe audio files to text using OpenAI Whisper
Video and audio processing with FFmpeg. Use for format conversion, resizing, compression, audio extraction, and preparing assets for Remotion. Triggers include converting GIF to MP4, resizing video, extracting audio, compressing files, or any media transformation task.
Local speech-to-text with the Whisper CLI (no API key).
Text-to-Speech Tool - Supports script parsing, emotion tagging, and post-processing, based on Edge TTS
Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.