Back to the lab

The comment-to-video pipeline

How we turn a typed answer into a finished whiteboard-style animated explainer video — AI-generated script, illustrations, narration, and captions for about a dollar, with human-in-the-loop approval at every point that matters. The automated video production pipeline we run for our own studio and build for clients.

Content pipelines · July 2026

Last Sunday we published a post about building your own software. Someone asked a good question in the comments: does small SaaS still have value? Daniel typed an answer on his phone — a few rough sentences, the way you'd answer anyone.

By the next morning that comment had become this:

82 seconds, drawn, narrated, and captioned from a comment reply. Total cost: about seventy-five cents.

Nobody opened an editor. Nobody recorded anything. The script, the drawings, the narration, the word-timed captions, the music mix — all generated from the comment while Daniel did something else. It went into the schedule as a draft and waited for one tap of approval.

This note is about how that pipeline works, because the shape of it — not the animation style — is the thing we keep building for clients.

The loop

A question or a trend comes in. A human writes the answer they would have written anyway. From there the pipeline takes over:

  1. Script. An agent rewrites the answer into the video's format — a hook, five to seven beats, a closing question — following written style rules (register, pacing, what a hook is allowed to look like).
  2. Stills. Each beat becomes one hand-drawn-style illustration from an image model. Cents per image, so rejects are cheap.
  3. The gate. Every drawing is laid out on a contact sheet and reviewed before any money is spent animating. More on this below — it's the important part.
  4. Motion. Approved stills get animated — a mix of paid image-to-video models and free programmatic effects (the pen draw-on you see is code, not a model).
  5. Assembly. Local ffmpeg stitches beats, narration, word-timed captions, and a music bed into the final vertical video. This step costs nothing and is deterministic.
  6. Draft. The finished video lands in the content calendar as a draft. Nothing publishes itself.

About twenty-five minutes of machine time end to end. Under a dollar.

The human stays in the loop — as much as you want

The pipeline has two built-in approval points, placed where judgment actually matters.

The first is the contact sheet: every drawing, next to the script line it illustrates, reviewed before the animation spend. This gate isn't a yes/no button — it's where you art-direct. Reject a drawing with a reason. Tweak the prompt and regenerate. Ask for the same scene from a different angle, swap the metaphor entirely, or polish a beat until it's exactly right. Because illustrations cost cents, iterating here is essentially free — bad art caught after animation costs the whole beat.

The second is the draft: the finished video sits in the calendar until a human taps approve. The pipeline can make content around the clock; it cannot post anything on its own.

How much you stand between those two points is a dial, not a setting. Some clients want hands on every frame; some want to review the first ten videos and then only spot-check. Both are the same pipeline — the only thing that changes is how often it waits for you.

Teaching the gate your taste

For the first videos, Daniel reviewed every contact sheet himself. Each rejection came with a reason: no floating heads, no hands without a body, if you can't tell what a drawing shows in one second the viewer never will.

Those reasons became written rules. Now the agent holds its own drawings against them before a human sees anything. On a recent video it rejected three of its own eight drawings — one had a man with two laptops, one was a pair of disembodied hands — redrew them, and moved on. The review still happens. It just happens without anyone watching it happen.

The video about the gate, made by the pipeline it describes — which rejected three of its own drawings on the way.

That's the general pattern, and it applies well beyond videos: a human review step doesn't have to disappear to stop costing you time. Write down why you say no, and an agent can apply your taste at machine speed — with spot checks instead of full-time supervision.

Why response time is the point

Content pipelines usually optimize for volume. This one optimizes for latency — the time between something happening and your answer existing as finished media.

A customer asks a question: the answer can be a video the same day, in your brand's style, at a cost that rounds to zero. Your FAQ stops being a page nobody reads and becomes a video library, one question at a time. Something shifts in your industry: your take can be out the same afternoon, not next quarter after a production cycle. The knowledge was never the bottleneck. Production was.

The business case, as the pipeline tells it: every question your customers ask is a video you already know the answer to.

What one video costs

Illustrations run about ten cents, animation fifty to ninety cents, narration a few cents, and assembly is free and local. Call it a dollar a video, plus twenty-five minutes of machine time nobody has to sit through. The style — black marker on white, one red accent, a recurring character — is locked in written specs, so video forty looks like video four.

The same shape, your operation

The animation style is ours. The shape is reusable, and it's what we actually sell: an unattended pipeline with human approval gates exactly where judgment matters, tuned until the gates can run on your written taste. We run this one for our own studio every week — this page is the evidence.

If your team answers the same questions over and over, or needs to respond to things faster than a production cycle allows, that's a good conversation to have.

Common questions

What is a comment-to-video pipeline?

An automated video production workflow that turns a short written answer — a comment reply, an email, a support ticket — into a finished animated explainer video: script, illustrations, narration, and word-timed captions, with human approval before anything is published.

How much does an AI-generated explainer video cost to produce?

Our generation costs run about a dollar per video: roughly ten cents of illustrations, fifty to ninety cents of animation, and a few cents of narration. Building a pipeline in your brand's style is a fixed-scope project; running it afterward costs almost nothing per video.

Can a human stay in the loop with AI-generated video?

As much as you want. The pipeline has two built-in gates — art approval before any animation spend, and final approval before anything publishes. You can art-direct every drawing, tweak prompts and regenerate for cents, or write down your review rules and let the pipeline apply them while you spot-check.

How fast can a business answer a question with a video?

Same day. The pipeline needs about twenty-five minutes of machine time, so a customer question that arrives in the morning can be an approved, on-brand animated video by the afternoon.

Will the videos match my brand?

The look is locked in written style specs — palette, recurring characters, narrator, pacing rules — so video forty looks like video four. We build the spec around your brand, not ours.

Can the videos use my voice?

Only if you want them to. Videos use professional stock narrators by default; a cloned voice is an opt-in upgrade that only ever happens with your explicit consent.