One idea in, a finished reel out.
Social Video Pipeline is a pipeline, not a single endpoint. Give it a topic and it runs the whole chain — writes the script, renders a cover image, generates the video, adds a voice track with lip-sync, and burns in word-timed subtitles — then hands back a post-ready clip. Each step is its own pay-per-call API. Billed per run. No subscription, no credit packs.
A pipeline, not an endpoint. Every step is its own paid call.
The clip is assembled from several independent generation APIs, each billed per call. You pay for one run; the pipeline pays each model as it goes, then stitches the result locally.
Write the script
Turn your topic into a hook, body, and caption via the model-access endpoint. Tuned for short-form pacing, not blog prose.
Render the cover
Generate a thumbnail / first frame via the image endpoint — the still that conditions the video and stops the scroll.
Generate the video
Image-to-video via the Kling 2.1 Pro video endpoint: cinematic 1080p motion from the cover frame and script direction.
Voice + lip-sync
Synthesize a voice track and align mouth movement via the audio endpoint, so the talent actually speaks the script.
Caption + assemble
Transcribe with word-level timing (Deepgram), burn in styled subtitles, and mux audio + video into one post-ready MP4.
For anyone who needs clips faster than an edit suite allows.
One pipeline, many formats. You bring the idea; it brings the script, the footage, the voice, and the captions.
A posting cadence without an editor.
Turn a one-line idea into a captioned, voiced clip you can post today. No timeline scrubbing, no separate caption tool, no monthly seats for software you use in bursts.
Feature clips and release teasers on demand.
Spin a changelog or feature into a short, captioned video for social — without booking design time. Pay per clip, bill it to the launch instead of a yearly tool.
Render variations to test, cheaply.
Produce dozens of captioned hooks and angles to A/B, without standing up generation seats for every occasional contributor. The per-run model fits a one-week campaign.
Video production as one tool call.
Expose the whole pipeline to your agent: it writes, renders, voices, captions, and returns a finished MP4 — no human juggling five generation accounts and an editor in the loop.
Give it an idea. Get a finished clip.
Run the pipeline once and pay for that clip. No subscription, no credit packs, no seats. Idea in, captioned and voiced MP4 out — ready to post.
- Pipeline of pay-per-call APIs
- Script → video → captions
- Voice + lip-sync built in
- Agent-ready
The honest answers.
Straight answers about what the pipeline does. No SDR funnel.
How is this different from the Video Generation API?
+
The Video Generation endpoint renders one clip from a prompt or image — it's a single atomic call. This is the whole pipeline around it: it writes the script, makes the cover, calls that video endpoint, adds a voice track with lip-sync, and burns in subtitles — returning a post-ready clip instead of raw footage. Want just a clip? Use the video endpoint. Want a finished social video? Use this.
What does each step use?
+
Script from the model-access endpoint, cover image from the image endpoint, video from Kling 2.1 Pro, voice + lip-sync from the audio endpoint, and word-timed captions via Deepgram transcription. Each is an independent paid call; the pipeline orchestrates them and muxes the result.
What does one clip cost?
+
A few dollars per finished clip, with the video generation step being the bulk of it. You pay per run — no subscription, no credit packs. Short clips cost less; longer or multi-shot sequences cost more.
Can I customize voice, captions, or style?
+
Yes — voice, caption styling, aspect ratio, and clip length are parameters on the run. You can also start from your own cover image to keep a product or character consistent.
Can my AI agent run the whole pipeline?
+
Yes. It's built for AI agents and pays automatically per step, so an agent can run script-to-captioned-clip end to end and return the finished MP4 URL without a human managing five generation accounts.
Who owns the output?
+
Generated output follows the underlying models' commercial-use terms. Don't generate content that infringes third-party rights or violates a model's safety policy; standard content filters apply.