Loading...
Loading...
Found 35 Skills
Interacting with Mediabunny
Multimedia handling with the Mediabunny library
Turn video moments into AI-generated hand-drawn storyboards with Gemini-powered frame analysis and social media content generation
Download and analyze social videos using frames + transcript for AI agent understanding at 50× lower cost than multimodal APIs
Watch any video (URL, stream, or local path) via Watch Skill. Downloads, extracts scene-aware deduped frames, OCRs them, transcribes (captions first, then local Whisper — offline by default), indexes everything, and hands the result to the agent. Follow-up questions are answered from the persistent index without re-processing.
Turn a plain video clip into a cinematic AI-VFX shot at 1080p by default (or 4K / 720p on request). Give Claude a video and the change you want; it reads EVERY frame via local contact sheets and understands the audio, writes a Seedance-faithful prompt that locks your face, gestures, and camera move, then re-renders the same shot with the VFX baked in via Seedance reference-to-video.
Targeted at the editing and production of finished content composed of audio, video or image materials, it applies to operations such as timeline cropping and splicing, speed and volume adjustment, video filters, camera motion effects, transitions, frame cropping, rotation and flipping, frame overlay, subtitle embedding, animated image extraction, fade in/out, audio-video extraction and muxing, audio mixing, text scrolling video production, image-to-video conversion, and multi-frame spatial composition. If users want to add filter effects to videos, or if the target clearly involves editing, synthesizing, overlaying, mixing or arranging existing materials but the specific method is uncertain, this Skill can be loaded first for exploration.
AI Face Swap - Swap face in video, deepfake face replacement, face swap for portraits. Use from command line. Supports local video files, YouTube, Bilibili URLs, auto-download, real-time progress tracking.
Watch and understand footage with Diffusion Studio via the `dapi` CLI: answer questions about a video or audio file, summarize it, find scenes and moments, pull quotes, and describe what happens and when. Use whenever the user asks what's in a piece of footage, wants a summary or recap, wants to locate a moment ("where does X happen", "find the scene where..."), or needs a claim about a video or audio file checked.
Video and audio processing with FFmpeg. Use for format conversion, resizing, compression, audio extraction, and preparing assets for Remotion. Triggers include converting GIF to MP4, resizing video, extracting audio, compressing files, or any media transformation task.
Expert guidance for computer vision development using OpenCV, PyTorch, and modern deep learning techniques for image and video processing.
Extracts first and/or last frames of every shot from a video using adaptive scene detection. Use this skill when the user says "extract frames", "get shot frames", "pull frames", "shot breakdown", "scene detect", "first frame of each shot", "last frame of each shot", "extract shots from video", or wants to extract key frames at shot cut points from a video file.