A pipeline · Several paid steps · Pay per run

One idea in, a finished reel out.

Social Video Pipeline is a pipeline, not a single endpoint. Give it a topic and it runs the whole chain — writes the script, renders a cover image, generates the video, adds a voice track with lip-sync, and burns in word-timed subtitles — then hands back a post-ready clip. Each step is its own pay-per-call API. Billed per run. No subscription, no credit packs.

  • Idea → script → video → captioned clip
  • Composes 5+ generation APIs in one run
  • Voice with lip-sync + word-timed subtitles
  • Agent-ready · pay per run
Your AI agent
Any agent
Social Video
Video pipeline
What runs under the hood

A pipeline, not an endpoint. Every step is its own paid call.

The clip is assembled from several independent generation APIs, each billed per call. You pay for one run; the pipeline pays each model as it goes, then stitches the result locally.

01

Write the script

Turn your topic into a hook, body, and caption via the model-access endpoint. Tuned for short-form pacing, not blog prose.

02

Render the cover

Generate a thumbnail / first frame via the image endpoint — the still that conditions the video and stops the scroll.

03

Generate the video

Image-to-video via the Kling 2.1 Pro video endpoint: cinematic 1080p motion from the cover frame and script direction.

04

Voice + lip-sync

Synthesize a voice track and align mouth movement via the audio endpoint, so the talent actually speaks the script.

05

Caption + assemble

Transcribe with word-level timing (Deepgram), burn in styled subtitles, and mux audio + video into one post-ready MP4.

Typical cost ~a few dollars per finished clip — video generation is the bulk of it
Output A single captioned, voiced MP4, ready to post — no editor required
Steps in one run 5 Script, cover, video, voice + lip-sync, and subtitles — five generation APIs orchestrated into one finished clip instead of five tools and an editor.
Output quality 1080p Native 1080p video from Kling 2.1 Pro, with a voice track, lip-sync, and word-timed captions burned in.
Pricing Per run Pay for the clips you actually render. No monthly plan, no credit packs, no seats for occasional creators.
Who it's for

For anyone who needs clips faster than an edit suite allows.

One pipeline, many formats. You bring the idea; it brings the script, the footage, the voice, and the captions.

For creators & solo marketers

A posting cadence without an editor.

Turn a one-line idea into a captioned, voiced clip you can post today. No timeline scrubbing, no separate caption tool, no monthly seats for software you use in bursts.

Make a 15-second explainer reel about Lightning payments with captions
For SaaS & product teams

Feature clips and release teasers on demand.

Spin a changelog or feature into a short, captioned video for social — without booking design time. Pay per clip, bill it to the launch instead of a yearly tool.

A 10-second teaser for our new card top-up feature, voiced and captioned
For agencies & ad studios

Render variations to test, cheaply.

Produce dozens of captioned hooks and angles to A/B, without standing up generation seats for every occasional contributor. The per-run model fits a one-week campaign.

Five captioned variations of this product hook for paid social
For AI agent developers

Video production as one tool call.

Expose the whole pipeline to your agent: it writes, renders, voices, captions, and returns a finished MP4 — no human juggling five generation accounts and an editor in the loop.

Produce a captioned short from this prompt and return the MP4 URL
Start in one line

Give it an idea. Get a finished clip.

Run the pipeline once and pay for that clip. No subscription, no credit packs, no seats. Idea in, captioned and voiced MP4 out — ready to post.

  • Pipeline of pay-per-call APIs
  • Script → video → captions
  • Voice + lip-sync built in
  • Agent-ready
FAQ

The honest answers.

Straight answers about what the pipeline does. No SDR funnel.

How is this different from the Video Generation API?

+

The Video Generation endpoint renders one clip from a prompt or image — it's a single atomic call. This is the whole pipeline around it: it writes the script, makes the cover, calls that video endpoint, adds a voice track with lip-sync, and burns in subtitles — returning a post-ready clip instead of raw footage. Want just a clip? Use the video endpoint. Want a finished social video? Use this.

What does each step use?

+

Script from the model-access endpoint, cover image from the image endpoint, video from Kling 2.1 Pro, voice + lip-sync from the audio endpoint, and word-timed captions via Deepgram transcription. Each is an independent paid call; the pipeline orchestrates them and muxes the result.

What does one clip cost?

+

A few dollars per finished clip, with the video generation step being the bulk of it. You pay per run — no subscription, no credit packs. Short clips cost less; longer or multi-shot sequences cost more.

Can I customize voice, captions, or style?

+

Yes — voice, caption styling, aspect ratio, and clip length are parameters on the run. You can also start from your own cover image to keep a product or character consistent.

Can my AI agent run the whole pipeline?

+

Yes. It's built for AI agents and pays automatically per step, so an agent can run script-to-captioned-clip end to end and return the finished MP4 URL without a human managing five generation accounts.

Who owns the output?

+

Generated output follows the underlying models' commercial-use terms. Don't generate content that infringes third-party rights or violates a model's safety policy; standard content filters apply.