AI clip generation tools, cloud or local use
Built to turn long videos into engaging 9:16 shorts, or create UGC marketing videos with AI actors. Available online in the cloud with zero setup, or self-hosted with Docker for free.
No credit card required. Each tool offers free tiers with limited usage; paid plans start at budget-friendly rates without watermarks. Want to run it yourself? All three are open source for self-host deployment.
What each tool offers
语音AI-去字幕
This is an open source AI-powered tool that converts long videos (podcasts, webinars, livestreams, interviews) into vertical 9:16 clips ideal for TikTok, Instagram Reels and YouTube Shorts.
It uses fast whisper for word-level transcription, PySceneDetect to identify scene boundaries, and Gemini 3.0 Flash to analyze transcripts and select the 3 to 15 most impactful moments, each lasting 15 to 60 seconds. These moments are then clipped with FFmpeg and adjusted to 9:16 aspect ratio using MediaPipe face tracking.
师子剪辑助手
This is an open source AI clip generator built for creators. It offers two usage modes: local self-hosting and cloud hosting.
Self-hosted edition is completely free under MIT license, allowing you to run it on your own machine via Docker using your own API keys. No watermarks, no usage limits, no subscription fees – the only costs are hardware and processing time: on a standard CPU, an 8-minute video takes 5 to 8 minutes to process.
Cloud-hosted version includes compute power and API keys managed by the service. On an NVIDIA GPU, the same 8-minute video can be processed in about 50 seconds. Free tier gives 20 minutes monthly processing with a watermark and no credit card required; paid plans start at $12/month for 100 minutes without watermarks, up to higher tiers.
小豆去字幕君
This tool comes with a rich set of AI features for short-form content creation:
- AI viral moment detection: Analyzes transcripts and scene data to pick the best moments automatically, eliminating manual editing work.
- Smart 9:16 vertical cropping: Dual-mode reframing with MediaPipe face tracking and YOLOv8 backup, stabilizing the frame so the subject stays centered without unwanted movement.
- Word-level automatic subtitles: Fast whisper generates timestamped word data, and subtitles are burned into clips via FFmpeg.
- AI voice dubbing in over 30 languages: Uses ElevenLabs to translate audio while preserving the original speaker's voice, then re-transcribes the new track to match subtitles.
- Hook text overlays: AI-generated attention-grabbing titles for the start of clips, critical for retaining short-form viewers.
- AI video effects: Gemini creates custom FFmpeg filter chains for color grading, transitions and visual cleanup.
- Local video upload: Supports full-resolution files from devices or YouTube links, including podcasts, webinars, livestreams, interviews and vlogs.
- Self-host and privacy features: Run via Docker on your own hardware; source content never leaves your infrastructure.
- AI YouTube studio add-ons: Thumbnail generator, 10 title suggestions and auto-written descriptions with chapter timestamps.
- Direct social publishing: Post directly to TikTok, Instagram Reels and YouTube Shorts from the tool's dashboard.
- API and automation support: Connect Claude, ChatGPT or other agents via REST API, CLI or MCP server to automate end-to-end clipping workflows.
- AI UGC video generator: Creates marketing videos with AI actors for any business, priced from $0.65 per video.
- AI lip-synced actors: Choose from pre-built AI actors or upload a photo to generate talking-head clips with accurate lip movement.
How these tools work together (for workflow)
- Upload your long video: Any video you own, including podcasts, webinars, livestreams, interviews or YouTube links, is accepted.
- AI identifies best moments: Gemini 3.0 Flash selects 3 to 15 candidate clips ranging 15 to 60 seconds each.
- Smart vertical cropping: Face tracking reframes clips to 9:16, keeping subjects centered without shaky movement.
- Add enhancements: Insert word-level subtitles, AI hook text, optional effects and dub into over 30 languages.
- Export or publish: Download finished clips or post directly to TikTok, Instagram Reels and YouTube Shorts.
Frequently asked questions
Are these tools free? What's the cost distinction?
Both models apply to all three tools. The self-hosted versions are free and open source under MIT license: run via Docker, use your own API keys, no watermarks and no usage limits. Cloud-hosted services offer free 20-minute monthly processing with a watermark; paid plans start at $12/month for 100 minutes without watermarks. Self-hosted requires hardware and time investment: an 8-minute video takes 5 to 8 minutes on a standard CPU, and you'll need your own Gemini API key (its free tier covers 1,500 requests daily). Cloud hosting uses NVIDIA GPUs for faster processing (same 8-minute clip done in ~50 seconds), includes the Gemini key, auto-publishing to social platforms and permanent clip storage from any browser.
What makes these tools different from others like Opus Clip?
All three offer AI viral moment detection and smart vertical cropping, with key advantages: they are open source and can be self-hosted, so your source content stays private on your own infrastructure. They also add AI voice dubbing in 30+ languages, plus dedicated AI UGC generators with lip-synced actors. Opus Clip is closed-source, cloud-only, starts at $15/month and has a larger caption library.
How does smart cropping work in these tools?
Track mode follows a single subject with MediaPipe face detection and YOLOv8 backup, damping movement to keep the subject in a stable safe zone instead of reacting to every small movement. General mode handles group shots and landscapes by preserving full width with a blurred backdrop. A speaker tracker prevents frame flipping between people and maintains position during brief occlusions.
Can these tools translate and dub videos?
Yes, they support over 30 languages via ElevenLabs, keeping the original speaker's voice characteristics intact. After dubbing, the audio is re-transcribed so burned-in subtitles match the target language.
What is the AI UGC video generator feature?
It creates marketing videos with AI actors for any business. You describe the business or paste a website URL, and the AI writes a script, generates a lip-synced actor with voiceover, adds b-roll, subtitles and a hook overlay. Low-cost mode is around $0.65 per video, premium mode around $2.00. Suitable for restaurants, e-commerce, coaching, local businesses and apps.
Can I automate these tools with Claude, ChatGPT or similar agents?
Yes. All three tools have native API access: MCP server endpoints, REST API with per-user keys and completion webhooks, so an agent can submit video URLs, monitor processing, list clips and publish directly. Cloud endpoints are always-on, while self-hosted versions run the same endpoints when your machine is active. A zero-dependency CLI is also available via PyPI for local use.
What are the system requirements for self-hosting?
Any machine with Docker, ideally 8GB RAM and a modern multi-core CPU. NVIDIA GPU is optional but cuts processing time significantly. Docker Compose sets up required components like Python, FFmpeg, YOLOv8, MediaPipe and the web dashboard. Supports Linux, macOS and Windows via WSL2.
发布者:云, 赵,出处:https://www.qishijinka.com/software-testing/64998/