FROM LONG VIDEO TO NEXT POST.
Turn long videos into short clips.
Make each moment count.
Turn a podcast, interview, webinar or stream into captioned short clips for YouTube Shorts, Instagram Reels or TikTok. Here is the workflow, from the original recording to an exported MP4.
Before you start: install Cutawan on Windows, macOS or Linux and have a source video you can use. For AI clip finding, choose ChatGPT sign-in via Codex with local Whisper or an OpenAI-compatible API. Codex uses your plan's allowance; API usage is billed separately. Local-only captions can edit a whole video but do not find clips. See the first-run and installation guide.
1. Bring in your original video
Import a local file, such as an MP4, MOV, MKV or WEBM. You can also paste a supported video URL. Use material you own or have permission to repurpose, and start with a single recording while you get familiar with the editor.
For spoken content, clear audio gives transcription a better starting point. Set the transcription language in Settings when you know it.
2. Choose clips or a whole-video edit
Choose the clip-finding mode to let AI suggest highlights. Cutawan transcribes the recording, identifies moments and scores their potential. You can steer discovery with a prompt, such as asking for practical advice or funny exchanges.
If you already know you want the whole recording, choose Caption the whole video. This skips highlight discovery and gives you an end-to-end captioned edit that you can trim, style and reframe.
3. Review the moment, not just the score
A score is an editing aid, not a prediction of views. Play each candidate and check whether it makes sense without the original video around it.
- The opening: does the first sentence give someone a reason to keep watching?
- The context: can a new viewer understand the names, references and point?
- The ending: does it finish a thought, or cut off before the payoff?
Open a promising clip in the editor. Adjust the trim and use the transcript to navigate to the part you want. Preview after enabling pause and filler-word removal so you can judge the pacing.
4. Frame the shot and style the captions
Choose 9:16 for a vertical edit, or use 1:1 or 16:9 when another layout suits the content. Speaker-aware reframing can follow the active speaker. Always preview cuts between people and check that important visuals remain visible.
Pick a caption style, then check names, specialist terms and line breaks. You can use your own fonts and add your branding. Keep captions readable and away from important faces or visuals.
5. Export, review and upload
Export an H.264/AAC MP4 with captions burned in. Rendering runs on your computer. Play the exported file once to check the framing, caption timing, audio and ending before publishing.
Cutawan can help write a social post caption, but you upload the final video yourself. Check the destination’s current upload requirements, then choose a title or caption that accurately sets up the moment.
What is processed locally?
Editing, face tracking, rendering and export happen locally. On the API route, extracted audio may be sent for transcription, and transcript text and sampled frames are sent for analysis. ChatGPT/Codex keeps speech local but sends analysis data. Local-only whole-video captions need no AI service. Your full source video is not uploaded to a hosted editor.
Build a repeatable habit
Try a few different moments from your recording rather than choosing only the highest score. Keep the ones that stand on their own and sound like you. The aim is to make more useful work from material you already made.