How to Make a Training Video: 6 Steps, Per-Type Production and Costs

Key takeaways

  1. Making a training video takes six steps: define the audience and goal, pick the type, build the deck and script, record or generate, review against the source, then publish with a revision rule.
  2. The type decides the route: demonstration and instructor-led video need a camera, while slide-based, screen-recording, role-play and animated training come from a deck, a script or a screen capture.
  3. Studio rate guides published in 2026 quote $500 to $15,000 per finished minute with a 4 to 8 week lead time.
SHARE

Making a training video takes six steps, and five of them are the same whichever type you are building. What changes is step four: some trainings have to be filmed, some can be recorded out of PowerPoint, and some can be generated from a deck you already own.

This article is written for the person doing the building: the six types and how each is produced, what studios charge, the free PowerPoint-only route, the AI route from a PDF or PPTX file, and the checks to run before an AI-narrated video reaches your staff. If you are still deciding which trainings are worth turning into video, start there and come back.

Try NoLang for Free

What a training video is, and the six types

A training video is a video program assigned to employees so they learn a defined skill or rule — onboarding, compliance, product knowledge, safety — and finish it with the same understanding wherever they work.

Training videos sort into six types, and the type decides the equipment, who can build it, and the price.

The six types of training video
TypeUse it forWhat makes it work
① Slide-basedCompliance, product knowledge, company rules, AI literacyThe deck exists, so the content is written before production starts
② Instructor-ledMessages from leadership, sessions where credibility mattersA named person on camera saying it out loud
③ DemonstrationMachine operation, assembly, inspection, safetyThe real motion, hands and workspace
④ Screen recordingInternal systems, SaaS tools, back-office proceduresThe actual screen, in the actual order
⑤ Role-playCustomer handling, harassment, negotiation, interviewingA bad example next to a good one
⑥ AnimatedAbstract subjects: bias, ethics, organizational conceptsDiagrams carrying an idea that has no footage

Most of a company’s training library falls into types ①, ④ and ⑤, where the content is already written down somewhere.

Maverick-kun

Put your own training in one of those six rows before you read on. Everything after this depends on the row you landed in.

What the build commits you to

Video pays off where the same content is taught repeatedly to different people; where the session happens once, live is still cheaper. Two consequences bind the build rather than the decision.

  • The first version is front-loaded: the return arrives on the second and third run, so plan for one build and many re-issues.
  • The revision trigger goes in before publishing: a video that still states the rule that changed in April is worse than no video.

The case for video, the sessions to keep in a room and the blended split are argued in the companion article on choosing which trainings to convert. If a whole program becomes self-paced, delivery and hosting turn into their own project — that ground belongs to building e-learning video.

How to make a training video in six steps

Making a training video is six steps: define the audience and learning goal, pick the type, build the deck and script, record or generate, review against the source, then publish with a revision rule. Every later section of this article is one of these steps in detail.

Step three and step four are where the schedule goes: a 2010 Chapman Alliance survey of 249 organizations, still the benchmark the L&D field quotes, put 43 hours of work behind one hour of instructor-led training. Writing the content, not pointing a camera, consumes the weeks.

Step one is where most builds go wrong. “Raise awareness of information security” is not a learning goal; “a new joiner can tell which files may leave the company network” is. Step six is the one teams skip: name the destination, the owner and the event that forces a revision before the video ships.

How each of the six types is produced

Same six steps, six versions of step four. Each subsection gives the three production stages with the equipment named, and whether a team without video experience can build it.

① Slide-based training videos

Slide-based training is the type where the content already exists, because the deck from last year’s session is the script. It suits compliance, product knowledge, internal rules and literacy training.

  • Prepare: cut the existing deck to one message per slide and write the narration as full spoken sentences.
  • Produce: record narration over the slides, or generate the narrated video from the file. No camera, no studio.
  • Finish: add captions, split it into sections, and check numbers and dates against the source deck.

In-house without filming: yes, and this is the type to start with. There is a slide-based sample further down.

② Instructor-led training videos

Instructor-led video works when who is speaking matters as much as what is said.

  • Prepare: script the talk, choose the location, and decide the frame — seated or standing, with slides beside the speaker or not.
  • Produce: camera on a tripod, a lavalier microphone (never the camera’s built-in mic), two lights, and a slot long enough to absorb retakes.
  • Finish: cut the retakes, drop in slide inserts, add captions and lower thirds.

An avatar presenter — an AI-generated presenter shown on screen speaking the narration — reproduces the lecture format with no shoot. Be straight about the trade: in the TechSmith study, 87% of viewers said they prefer a real person over an animated character or an AI avatar. So film the message from your CEO, and take the avatar route for the annual security refresher, where the job is accurate delivery — there is an instructor-style sample below.

③ Demonstration training videos

Demonstration video is the one type where filming is genuinely unavoidable, because the information is the physical motion itself. Machine operation, assembly, inspection routines, safety procedures.

  • Prepare: get site access, book the equipment and the operator, and agree the safety rules for filming.
  • Produce: shoot three angles — wide for context, hand-level for the technique, close-up on the part that matters.
  • Finish: slow or freeze the critical moment, label the parts on screen, and narrate why each step is done that way.

The real cost here is rarely the crew; it is stopping a line or fitting the shoot around a shift.

④ Screen-recording training videos

Screen-recording training shows an internal system in the exact order a person will use it, which no screenshot deck ever quite manages. A screen recorder and a basic editor are the only tools involved.

  • Prepare: set up a test account with dummy data — never record live customer or employee data — and rehearse the flow once.
  • Produce: capture at the resolution people actually work at, moving deliberately and pausing at each decision point.
  • Finish: trim the waiting, highlight the click targets, and add narration and captions.

A single interface change invalidates the recording, which makes this the type worst served by outsourcing: a studio revision cycle costs more than the original recording did. Where the goal is a reference opened mid-task, you are making a video manual instead. A screen-recording sample is below.

⑤ Role-play training videos

Role-play training teaches judgement rather than facts, by showing the same situation handled badly and then handled well. Customer complaints, harassment scenarios, interviewing, negotiation.

  • Prepare: write the exchange as a two-speaker dialogue, poor version first and corrected version second, with a specific difference between them.
  • Produce: cast two people and film the exchange, or have generated speakers perform the dialogue.
  • Finish: add a short commentary after each version naming what went wrong and what changed.

Casting colleagues is what kills this type: people are busy and self-conscious on camera. Generating the dialogue from the script removes that blocker, as the role-play sample below shows.

⑥ Animated training videos

Animation earns its cost when the subject has no footage: bias, ethics, organizational concepts, anything that lives in a diagram.

  • Prepare: storyboard the sequence and decide what each visual has to make clear.
  • Produce: build the diagrams and motion, then time them to the narration rather than the other way around.
  • Finish: caption everything, and watch it once at the screen size your staff will use.

Text-and-diagram animation is within reach of generation from a deck; character performance needs an animator’s time. The unconscious bias sample shows the generated end of that range.

What each type costs, and how long it takes

Studios publish rate guides rather than survey data, so read these as asking prices for planning. Two 2026 guides cover the six types between them — Vidico and DMAK Productions — and each row names its source. Both put a standard single video at four to eight weeks.

Published rates per finished minute by type, and what it takes to build in-house
TypePublished rate per finished minuteBuilding it in-house
① Slide-basedNeither guide prices a narrated deck on its ownNo filming: deck edit plus script, hours to two days
② Instructor-ledTalking head $2,000–$5,000 (Vidico); $500–$1,500 basic (DMAK)Half to one shoot day plus edit, or an avatar presenter
③ DemonstrationHigh-production live action $5,000–$15,000 (Vidico)Filming unavoidable: a shoot day, site downtime, an edit
④ Screen recordingScreen-capture tutorial $1,000–$3,000 (Vidico)No filming: test account plus half a day per flow
⑤ Role-playProfessional live action $1,500–$5,000 (DMAK)From a script; writing the two versions is the work
⑥ Animated2D explainer $2,000–$5,000 (Vidico); $2,000–$8,000 (DMAK)Text and diagrams only: storyboard plus build, days
  • Where the guides disagree: DMAK starts live action at $1,500, Vidico at $5,000. Budget against the wider figure.
  • What the do-it-yourself end costs: Vidico alone prices it, at $26 to $200 in total for a screen recording and $25 to $75 per finished minute for AI-generated avatar video.
  • Why review rounds move the invoice: Vidico splits a budget into pre-production 25–30%, production 35–40% and post-production 30–40%, and DMAK says outright that instructional design and review cycles cost more than runtime.

The in-house column is stated in effort rather than dollars, and those hours are planning estimates rather than survey data; the constraint is the time of the person who knows the subject. For rates outside training specifically, what video production costs has the fuller breakdown.

Making a training video from PowerPoint alone

With a deck, a microphone and no budget, PowerPoint will produce a finished file on its own. PowerPoint records narration, animations, pointer movement and slide timings into the presentation itself, and exports the result as a video through File > Export > Create a Video, saved as .mp4 or .wmv. In newer builds the ribbon tab is called Record (Source: Microsoft Support).

What makes the route usable is that narration is stored per slide: if slide 14 has a wrong figure, you re-record slide 14 and export again.

The limit is time. Recording runs in real time: a 20-slide deck with 30 seconds of talking per slide is a 10-minute take, longer once you restart a few slides. Multiply that by twenty titles and an annual revision cycle, and the route that cost nothing to start becomes the reason the refresh never happens.

Real trainings mix the types

Almost no useful training is a single type end to end. A SaaS rollout runs slide-based for the rules, screen recording for the workflow, then slide-based again for the mistakes to avoid. Build the pieces as separate videos with one goal each: when one rule changes you re-do one video, not the program.

Maverick-kun

Notice that the pieces you would rebuild most often are the slide-based ones — which happen to be the pieces you can generate from a file.

How to make a training video with NoLang

NoLang, from Mavericks, Inc., is an AI video generation service that produces narrated, subtitled video from material you already have. For L&D the case is narrow and common: the compliance deck or the policy PDF on your drive becomes a narrated video without a camera or an editing suite.

  • Documents: PDF and PPTX files — the route for types ① and ⑥, and the one most training teams need.
  • Text or a script: paste the content, or supply a two-speaker dialogue for role-play.

Footage you have filmed and web pages are accepted too; the formats and the Chrome-extension requirement are in the FAQ. Note the division of labor: NoLang produces the video, while assignment and completion records come from wherever it is hosted.

  1. Choose From document and upload your deck

    Generate from documents — the From document tab — takes the PDF or PPTX you already deliver the training from. A deck ordered the way you teach beats a tidied-up summary.

  2. Set how much of the deck to cover

    Choose Explain all pages when every slide carries a rule people must hear, or Explain key points only for a refresher. Detail level is separate: Briefly, Normal or Detailed.

  3. Choose the output language and format

    Pick from 34 output languages, choose Desktop or Mobile, and turn Subtitles on: training is often watched on a shared floor with the sound down.

  4. Generate the video

    NoLang writes the narration, builds the visuals and paces the sections from the deck.

  5. Review, correct and share

    Rewrite any mispronounced term in the deck the way it should be said, adjust the intervals between sections, then download the file or share the URL.

A deck becomes a narrated training video without a shoot, a voice-over session or editing software, which is what makes an annual revision cycle survivable: next April’s version is a re-generation, not a re-production.

Turn Your Deck into a Training Video
Maverick-kun

Start with the deck that has to change for a policy update this quarter. You find out what a re-generation costs at the moment it matters.

Quality assurance when you generate training videos with AI

Caution about AI-made video is not irrational. In the TechSmith study, 75% of people were receptive to AI-generated video, but 90% had concerns, the top one being accuracy at 45%. A training video states your rules to the whole company, so the review has to be a written process, not a feeling that it looked fine.

The quality of a generated training video is decided by two things: the deck you upload, and the check you run before you publish.

How to write the narration script from your deck

The script is the deck rewritten to be heard, and the learning goal decides what survives the rewrite.

  • One message per slide: a slide carrying three rules produces narration that races. Split it.
  • Turn fragments into sentences: “Report within 24h” is a bullet; “Report the incident to your manager within 24 hours” is what someone should hear.
  • Say the reason before the instruction: a rule people understand is one they follow in the situation the slide did not cover.
  • Keep the spoken version in the notes: the speaker notes already hold your explanation; tidy them before you upload.
  • Write awkward terms the way they should be said: names and abbreviations are read as written, so fixing one in the deck fixes it in every version you generate afterwards.
  • Read the script aloud once: anything you stumble over on the page sounds worse in narration.

What to check before you publish

Pre-publish checks for a generated training video
CheckWhat to look atHow to fix it
PronunciationCompany, product and people’s names; abbreviations read as wordsRewrite the term in the deck the way it should be said, then regenerate
Figures and referencesAmounts, dates, deadlines, clause and policy numbersCorrect them in the deck, then generate again
Slide-to-narration matchWhether the narration matches what is on screenAdjust the interval, or split the slide
Statements not in the sourceAnything asserted that your deck does not sayRemove it, or add the point to the deck
Length and section breaksWhether it runs as one undifferentiated blockAdd intervals, or split into several videos

The rule that keeps this simple: fix facts in the source deck and generate again; fix delivery in the editor. Patching a wrong figure in the narration text leaves it wrong in the deck, and the deck is what next year’s version is built from.

Which trainings not to generate

Two cases. Hands-on demonstration, where the physical motion is the content: film it, then bring the footage in to have subtitles added over audio that stays as recorded. And any session where the speaker’s authority is the point — a message from your CEO about a serious incident should have your CEO in it. Everything else, and that is the bulk of what companies deliver, is where generation earns its place.

Six training videos made with NoLang

Six training videos produced with NoLang, one per production route. Narration and captions are in Japanese because the samples were made for a Japanese audience; the same deck can be generated in any of the 34 output languages. More examples are in the Video Gallery.

Introduction to generative AI literacy

  • How it was made: an existing introductory deck generated as slides with narration over them.
  • Where it fits: a first AI session for non-technical departments, or onboarding.

Personal data compliance in recruiting

  • How it was made: a two-speaker script generated as a conversation between a senior manager and a junior recruiter, with nobody cast.
  • Where it fits: compliance training for HR and recruiting.

Unconscious bias training

  • How it was made: built from a deck, with an on-screen instructor character working through examples of how assumptions shift an evaluation.
  • Where it fits: appraiser training and diversity education.

Generative AI for business

  • How it was made: a slide deck turned into a short narrated video, with no filming and no editing.
  • Where it fits: raising AI literacy company-wide as part of a digital program.

Information security training

  • How it was made: an avatar presenter delivering the session in lecture format, produced without a shoot.
  • Where it fits: security onboarding and the annual refresher.

Confidentiality, integrity and availability, opened the way an instructor would — same format as type ②, no studio day.

Tool operation walkthrough

  • How it was made: a screen recording of the real interface, with narration and captions added over it.
  • Where it fits: training staff on an internal tool, or turning a written manual into video.
Make a Training Video Like These

In-house, a studio, or AI generation?

Step two and step four come down to one routing question, answered per training, not per company.

Three routes for building a training video, compared
What you are comparingIn-house filming and editingA production studioAI generation
Time to a first versionDays, once the shoot is scheduled4–8 weeks for a standard projectAs soon as the deck is ready
Cost of a revisionRe-shoot or re-edit the affected partBack into the pipeline and billed againUpdate the deck and generate again
Other languagesA new recording per languageA new session and edit per languageChoose another output language
SuitsDemonstrations on your own sitesBrand-facing programs where polish decides itCompliance, product and process training at volume

Choosing per training, not per company

Route by revision frequency: content you rewrite every year belongs on the generation route, filming is for material only you can capture, and a studio is for the few programs where polish decides the outcome. Running all three at once is normal rather than a compromise. Moving the routine bulk in-house is what frees the budget to shoot the two videos that genuinely need a crew.

Maverick-kun

Do not decide this for your whole library at once. Decide it for the next video, and let the revision count make the argument.

Frequently asked questions

The questions below come up most often once a team starts producing its own training videos. Each answer points back to the section of this guide that covers it in full.

FAQ

How long should a training video be?
It depends on what the viewer came for. TechSmith's viewer study found half of respondents prefer 30 seconds or less for a general topic, while 67% would watch more than an hour to learn a job skill, with stated preferences clustering at 10 to 19 minutes. Give each video one learning goal and let that set the runtime; if a session runs long, split it into parts.
How much does it cost to produce a training video?
Two 2026 rate guides differ on the talking-head figure — $2,000 to $5,000 per finished minute at Vidico, $500 to $1,500 at DMAK Productions — and price screen-capture tutorials at $1,000 to $3,000 and 2D animation from $2,000. These are asking prices, not survey averages. Vidico puts the do-it-yourself end at $26 to $200 in total for a screen recording and $25 to $75 per finished minute for AI-generated avatar video.
Do I need special equipment to make a training video?
Only for two of the six types. Demonstration needs a camera, and instructor-led needs a camera, a lavalier microphone and lighting unless you use an avatar presenter. Slide-based, screen-recording, role-play and animated training can be built from a deck, a script or a screen capture on the computer you already have.
How do I write a training video script?
Work from the learning goal written as one sentence, then go slide by slide, turning each bullet into a full spoken sentence and saying the reason before the instruction. If your deck has speaker notes, most of the script already exists. The full checklist is in the section above on writing the script.
Which file formats can NoLang turn into a training video?
Document input accepts PDF and PPTX. You can also start from plain text or a two-speaker script, or bring in footage as .mp4, .mov or .webm, in which case the original audio of that footage is used as it is.
Can I control how long and how detailed the generated video is?
Yes, through two separate settings. Detail level is Briefly, Normal or Detailed, and coverage is either Explain all pages or Explain key points only. A full compliance run wants every page at normal detail; a refresher for people who sat the session last year wants key points only.
Can I produce training videos in other languages?
Yes. NoLang offers 34 output languages, and the language is chosen before generating rather than commissioned as a separate production. For a group with sites in several countries that is the biggest single difference from the traditional route, which prices a fresh recording and edit per market.
Can I edit a generated training video before publishing it?
Yes. You can adjust the intervals between sections when the pacing is wrong, and the editor keeps a reading list for correcting pronunciation, with readings entered in hiragana or katakana. For English narration, the reliable fix for a mispronounced name is to write it in the source deck the way it should be said and generate again. Delivery problems belong in the editor; factual problems belong in the deck.
SHARE
NoLang Service Overview
Download Materials

Considering NoLang
for your business?

Already using
NoLang?