Skip to main content

Deliverable · Explainer video

Explainer video production for software

The short answer

An explainer video makes a product understood in one sitting. This studio builds them as code, not as a timeline file. On the narrated one every scene boundary derives from the narration track, so the cut cannot drift from the voice. 2 explainers here, 0 of them made for a real client, and no model draws a frame of either.

An explainer earns its runtime by being understood, not just watched. Blue Snow Developers makes the two here as motion graphics rather than live action — every frame designed and rendered, none of it filmed — and both are built as numbered walkthroughs rather than moods: the problem first, then one step at a time, with each chapter showing the single thing the sentence over it is claiming.

Laneflow (Concept) — a product walkthrough in three numbered steps, each on its own ground, each carrying one drawn interface. Read how it was built, including where its qualifier stops.

Most explainer videos fail by padding, not by being ugly

A software product is abstract, so there is nothing to point a camera at, and the runtime has to be filled with something. That is the whole problem. The film drifts into atmosphere — swooping gradients, abstract nodes, a voice saying the word “seamless” — and a viewer comes out the other side having watched a mood rather than understood a tool. Every one of those seconds was paid for and none of them explained anything.

The alternative is unglamorous and it works: name the problem before selling anything, advance exactly one idea per beat, and let the picture show that idea and nothing beside it. The test is not whether the film looked expensive, it is whether someone can describe what the product does after it stops. That argument is why an explainer sits differently from a brand film, which is bought for how a company feels, and from a logo animation, which has one asset to get exactly right, and from a promo, which is bought to be posted and judged where it lands, and from ad creative, which is bought as a set rather than as a film. What the studio makes, and how a project runs sets them against each other.

2 explainers — and neither was made for a client

Both explainers on this page are for products that do not exist. The studio invented the software, drew every screen in them and cut both films — and on the narrated one, wrote the script it is spoken from. Each piece is labelled (Concept) wherever it appears. The consequence is sharper for this format than for the others the studio sells: an explainer is a claim that a real product survived being explained, and a product invented in-house cannot argue back. There was no support team to say “that is not how customers actually use step three”, because there are no customers and there is no support team. That is a real difference, not a technicality.

What concept work does prove is the part a brief cannot describe — structure, pacing, how legible an interface stays at a phone size, whether the film has the discipline to stop. The full portfolio is filterable by industry, deliverable and format, and the one engagement a client did pay for is film work rather than an explainer: a brand promo and a market-update format for a mortgage lender.

Cadence (Concept) — B2B-SaaS operations-room data UI

Cadence (Concept)

B2B-SaaS operations-room data UI. Delivered 16:9 and 1:1. This one is tagged twice — it is an explainer and a brand film, so it also sits with the brand films. It is the heavier, later-stage register: an operations room rather than a walkthrough, for a buyer who already knows what the category is and wants to see the product under load.

The cut is locked to the voice, not the voice to the cut

In most explainer pipelines the audio is laid against a finished animation and somebody nudges keyframes until the two roughly agree. It works until the script changes. Move a sentence, cut a clause, re-record a line, and every downstream cut is a few frames wrong in a way that reads as cheap without a viewer being able to say why.

The film above is assembled the other way round. The narration is rendered first and timestamped, and every narrated scene’s boundary is computed from those timestamps rather than typed in — so the step-two chapter starts on the sentence that begins “Two”, and there is no state in which the cut and the voice disagree. The one hand-set number in the whole timeline is where the closing card lets go, which was shortened on purpose to take dead air off the end. Re-cut the script and the film re-times itself on the next render. The craft note on that piece goes through the frames, including what it gets wrong.

Be precise about what that does and does not buy, because the neighbouring claim is one it is easy to state a notch too strongly. The build holds the CUT to the voice; the caption file shipped beside the film is cued by narration segment, not by word. The same machinery drives the narrated work on the data-driven films, where the voice has to land on a figure at the moment the figure resolves — a world-electricity film (Concept) is the hardest version of that problem in the library.

The interfaces are drawn — and for your product they may not need to be

Every screen in the film above is drawn for the frame: a shared board with its columns labelled, a task card joined by a lit line to the one person marked free, a sync node fanning out to the plates it just updated. None of it is a screen recording, and the reason is honesty rather than preference — the product is fictional, so there is nothing to record, and a shot of a real screen would be the one element in the piece that could not be stood behind.

For a product that exists, a capture of the actual software is often the better answer, and a studio that will not say so is selling its own preference. What drawing buys, when it is the right call, is legibility: type set for the frame instead of for a browser at whatever zoom the recording happened at, one element lit while the rest stays quiet, and a flow shown in the order the script needs rather than the order the app enforces. The usual answer for a real product is both — the interface recorded, and the moment that matters rebuilt so it can be read at a phone size.

Drawn does not mean simplified into meaninglessness. A six-city multi-market template (Concept) and a narrated data film that prints its own source codes (Concept) are drawn the same way, at real density, from real datasets.

The figure a viewer repeats is the one that has to be labelled

When an explainer has a number in it, that frame is the most dangerous one in the film. A big figure counting up under a confident line is the thing a viewer remembers, screenshots and repeats — and if it was invented for the edit, the film has put a claim into the world that nobody in the company ever signed off. Laneflow (Concept) handles that in the frame: the payoff chapter sets the figure counting and prints a chip directly beneath it naming what the figure is a sample of.

Now the part a portfolio page normally leaves out. That chip is small type under a figure many times its height, so it is a footnote — what makes it defensible is that it is a footnote UNDER the number rather than an end card nobody screenshots. And the other film on this page does it less well: Cadence (Concept) labels that figure too, and more thoroughly than the film above does: the line directly beneath its headline number reads “illustrative figure · concept commission”, naming the provenance as well as the status. Where it falls short is one plate. The square cut is a real square build — its chrome re-lays and it prints its own 1080×1080 — but the 24-week ridgeline is laid out at the wide film’s coordinates and clips at both edges, so its title reads “RATIONS LANDSCAPE”, its legend loses the word it starts with, and the annotation marking that change illustrative is pushed past the right edge. That plate is ours to re-render.

Nobody speaks the narration on the film above

One of the two explainers here carries narration, and that voice is a text-to-speech render — an open-weights speech model, run locally, from the same script printed verbatim on the piece’s own page. No voice is cloned from a real person and no performance is imitated. It is said here, in a heading, rather than in a footnote, because on this format the voice is not a garnish: an explainer is a voice with pictures under it, and a buyer choosing this studio is choosing that voice more directly than on any other page of this site.

The picture is a different matter and the distinction is the whole point. No generative model draws a frame of this work — not the interfaces, not the type, not the transitions; every frame is a component rendered from code, which is why the same build returns the same film. The synthetic voice is a production choice on concept pieces made with no budget and nobody to book. If a project wants a human read, the build composites the voice track it is handed and nothing else about the film changes — and a script written to be spoken by a person is a better script either way. One person writes the script, draws the screens and renders the film, which is why there is somebody to ask.

How an explainer project runs

It starts with the script, because on this format the script is the product and everything visual is a consequence of it. You send what the software does, who it is for and the one thing they must understand; what comes back is a written walkthrough with the steps numbered, and that is the thing to argue about — arguing about it later, over rendered frames, is what makes explainers expensive. Once the steps are agreed the interfaces are drawn against them, the voice is laid, and the film assembles itself around the timings. Tell us what your product does — the message reaches the founder, not a queue.

Afterwards you own the source, not just the export, and on this format that matters more than usual: a product changes, and a film you cannot open is a film you have to buy again. The field notes are where the engineering behind it gets written down, and the studio’s own identity system is the reference for how the token discipline works when a film, a mark and a document all have to agree.

How long should explainer videos be?

There is no correct number, and a studio that hands you one is quoting a habit rather than a measurement. Runtime is an output here, not an input: it falls out of how many ideas the product genuinely needs, one to a beat, and it can only be settled honestly once that count is known. Choosing the length first is what manufactures the padding above, because the film is committed to filling a slot before anybody has decided what belongs in it.

The explainer at the top of this page, Laneflow (Concept), came to 23 seconds. That figure is measured off the encoded file this page serves rather than taken from a brief, and the wide and the vertical are separate builds rather than crops of one another, yet both land on it. Read it as evidence of a method, not as a length to aim at: it is what one product's ideas amounted to once nothing was added to round the figure up.

What to ask before you commission an explainer

Has a real client ever asked you for an explainer video?

No. Both explainers on this page are for products that do not exist — software invented by the studio so the format could be built end to end and shown, and each is labelled “(Concept)” wherever it appears. The studio does have one real client, a mortgage lender, but the work delivered for them was a brand promo and a market-update format, not an explainer. That gap is worth being blunt about, because an explainer lives or dies on whether a real product survived being explained, and a product we invented could not argue back.

What actually makes an explainer video work?

Structure, and the discipline to keep it. The failure mode is padding: the product is abstract, the runtime has to be filled, and the film drifts into atmosphere until the viewer has watched a mood instead of understanding a tool. The alternative is to state the problem first, then take one step at a time, and put on screen the single thing the sentence over it is claiming and nothing else. Laneflow (Concept) is built that way — a numbered walkthrough where each of the three steps carries its own ground colour, its own step numeral and one drawn interface that shows the thing just claimed.

Do you screen-record our product, or draw the interface?

For a real product, both are on the table, and a recording of your actual software is often the more honest of the two. On this page the interfaces are drawn, and that is a consequence of the brands being invented rather than a house style: there is no Laneflow to record, so a shot of a real screen would be the one thing in the piece we could not stand behind. Drawing them also buys legibility a capture cannot — type set for the frame instead of for a browser, and the one element the narration is talking about lit while the rest stays quiet.

Who speaks the narration?

Nobody does, on the film shown here, and that is worth knowing before you commission one rather than after. One of the two explainers on this page carries narration, and that voice is a text-to-speech render — Kokoro-82M, an open-weights speech model, run locally. No voice is cloned from a real person and no performance is imitated. The picture is a different matter: no generative model draws a frame of this work, which is the claim this studio actually stakes itself on. If a project wants a human read, the build composites the voice track it is given, and a recorded voice changes nothing else about how the film is made.

Our product changes every few weeks. Does the film go stale?

Less than a timeline file does, which is the practical reason these are built in code. An explainer ages fast, because it is a portrait of a product on the day it was filmed — a new screen, a renamed feature or a changed flow dates it immediately. When the film is code, a changed screen is an edit and a re-render rather than a rebuild of the piece, the chapters you did not touch come back identical because nothing in the frame path is random, and the version you replace is still there to fall back to. Every source file is yours, so this does not depend on us being the ones who open it.

What ratios does an explainer ship in, and are they real builds?

The ones the film is actually going to be watched in — a wide cut for a site or a sales call, a vertical for the feed. Laneflow (Concept) ships both and they are two builds rather than one build and a crop: the wide cut sets the step numeral and heading against the interface plate, and the vertical stacks the same elements and re-sets the plate to the narrower measure, so nothing runs off an edge in either. The library is not uniformly that good and this page would rather point at it than let you find it — Cadence (Concept) also delivers a square cut, and while that one is a real square build, a single chart plate inside it is still laid out at the wide film’s coordinates, so that chapter’s title and legend are clipped at the frame edge.

Can you explain something complicated without dumbing it down?

That is the whole job, and the test of it is what a film does with its numbers. The riskiest frame in any explainer is the figure a viewer remembers and repeats, and the honest move is to mark it for what it is. Laneflow (Concept) does: its payoff chapter sets a large figure counting up and prints a chip directly beneath it naming what the figure is a sample of. Be exact about how far that goes — the chip is small type under a figure many times its height, so it is a footnote, but it is a footnote under the number rather than parked on an end card nobody screenshots.

Watch your own product get explained

Two invented products show how this studio structures an explanation. The test that counts is software somebody actually uses: send the product and a line about who it is for, and for the products we take on that round a watermarked sample gets built against your real interface. Nothing to sign, and no call to sit through first.

Written by Baljeet Aulakh, who builds every piece in this portfolio. Last updated . No AI-generated frames.