NoLang Adds a Feature That Automatically Turns Audio Files Into Captioned Video

Logo of Mavericks, Inc., the company behind the AI video generation service NoLang

Key takeaways

  1. NoLang, the AI video generation service from Mavericks, Inc., can now generate a high-quality video automatically from nothing more than an uploaded audio file.
  2. Companies can reuse audio assets that sit unused inside the organization, such as webinar and interview recordings, as video that carries far more content value.
  3. AI transcribes the audio and adds captions, so a finished video is produced in a short time without scriptwriting, filming or manual editing.
  4. The feature raises productivity for audio-based content production across internal training, recruiting sessions, product introductions and PR video.
  5. It is available now inside NoLang, adding an audio-based production flow alongside the existing option of starting from text.

Mavericks, Inc., the Japanese company behind the AI video generation service NoLang, has released a new feature that generates a video with automatic captions from an uploaded audio file such as an mp3. The feature uses AI to fully automate work that is expensive to outsource, including transcription and the insertion of sound effects. It lets companies reuse the audio data sitting unused inside their organization as “video knowledge,” and makes it easier to distribute audio content on other platforms such as YouTube, maximizing the value of their information assets.

Audio data that is only ever recorded, never used

As digitalization advances, audio data is rapidly gaining importance in business and communications.

Online meetings have become standard, so companies now hold recordings of sales meetings, call center conversations recorded in full for compliance reasons, and recordings of online webinars. Enormous volumes of audio data accumulate inside organizations. In practice, however, these valuable audio assets are hard to use in their raw form. Taking in information through hearing alone is monotonous, making concentration difficult and placing a heavy cognitive load on the listener. As a result, even when audio is used for training or knowledge sharing, deep understanding and retention suffer, and the value of the audio asset is never fully realized. At call centers, close to 100% of calls are recorded, yet conventional quality management monitors only about 1% to 3% of them1, leaving most of the audio unused.

The problem is not limited to business settings. Consumer audio content such as podcasts, which have grown quickly in popularity, faces the same issue. This content reaches only people who listen, and it is hard to publish as is on video platforms such as YouTube and social media, so its potential audience stays narrow. In Japan, only 17.2% of people listen to a podcast at least once a month2, while YouTube has more than 73 million monthly users in Japan alone3 and is used by 79% of people aged 45 to 64, making it part of everyday life. Because conventional audio content is difficult to publish on video platforms in its original file form, it reaches only the limited audience that chooses to listen.

Corporate audio assets and consumer audio content alike become genuinely valuable assets that can be searched, analyzed and used for education only once they are transcribed and made visible with on-screen captions. Until now, the transcription and other outsourced work required to turn audio into video cost tens of thousands of yen per piece, which made it hard to justify the budget.

A new feature that automatically turns audio files into captioned video

To solve the problem shared by recorded-but-unused audio assets and audio content with limited reach, Mavericks, Inc. has released a new feature in NoLang, the Japanese AI video generation service, that automatically produces a video with captions from an uploaded audio file such as an mp3.

The feature uses AI to fully automate the task of turning audio into video, which previously required substantial cost and editing skill. Companies can reuse the audio data sitting unused inside their organization as “video knowledge,” and podcasts and other audio content can be distributed on large platforms such as YouTube, maximizing the value of their information assets.

Key capabilities and benefits of the mp3 feature

The feature delivers three innovations: AI-driven conversion of audio files into video, reuse of corporate audio assets, and a dramatic reduction in production cost.

Capability 1: AI transcription and automatic caption generation

AI analyzes the uploaded audio (mp3) with high accuracy and transcribes it automatically, then instantly generates a video with captions. The generated captions are easy to correct and customize. Users can also freely combine the more than 100 AI avatars available in NoLang with royalty-free background footage and background music, so anyone can create visually appealing, rich content that an audio file alone could never deliver. Podcasts and other audio content that people could only listen to become video content people watch, dramatically expanding the channels available for distribution.

Capability 2: Reusing existing audio assets as video

The feature lets companies put the audio assets they already hold to work as video: sales meeting recordings accumulated through online meetings, call center conversations recorded in full for compliance purposes, and recordings of company introductions from online webinars and new-graduate recruiting events. Video is said to be roughly nine times more effective than text for retention, so turning seminar recordings into learning content as video substantially increases both their educational impact and their value as an asset. Building a library of video knowledge from strong sales conversations and customer support calls also strengthens training for sales staff and operators and helps standardize service quality.

Capability 3: A dramatic cut in production time and cost

Adding captions to a one-hour seminar recording and editing it into a video used to start with transcription at just under 20,000 yen, and the editing needed to finish the video brought the total outsourcing cost to tens of thousands of yen or more. NoLang has AI take over the captioning and editing work, so even people with no editing experience can complete the conversion without any specialist skills. The Business plan allows unlimited video generation for a fixed monthly fee, cutting the production cost per video to roughly a few thousand yen, one-tenth of the previous level, and making high-volume content production practical.

How companies can use audio-to-video in NoLang

The feature delivers immediate value across a wide range of industries and business situations.

Content distribution (podcast producers, influencers)

Convert podcast mp3 files instantly into video optimized for YouTube or TikTok, in horizontal or vertical format. Distributing across multiple platforms maximizes both reach and revenue.

Education and training (HR and learning teams)

Turn recordings of past seminars and lectures into captioned review videos that become training assets. Companies improve learning outcomes while reducing training costs.

Call centers and customer support (voice-of-customer analysis, quality management)

AI automatically converts the customer conversations recorded every day at call centers and support desks into captioned video with a transcript. Making these audio assets visible and searchable as text, instead of something staff can only listen back to, makes it far easier to analyze what customers actually say and to extract frequently asked questions. Strong examples of customer handling can also be reused as training videos, supporting operator training and consistent service quality.

Knowledge sharing (sales, digital transformation teams)

Teams can use audio notes from important meetings and sales calls as shared video that stands in for minutes. Captions are added automatically through speech recognition, so nuances that text alone conveys poorly come across accurately, and information is absorbed more efficiently.

What comes next

Mavericks, Inc. will continue to improve the accuracy of its AI speech recognition, expand the supported file formats (wav, m4a and others), and strengthen the integration with AI summarization, including automatic generation of highlight videos from long recordings. Through the recently launched NoLang API, the company will also embed video generation into the workflows of any organization, contributing to digital transformation and better communication as a solution that maximizes the value of information assets with AI.

Contact us

Footnotes

  1. Quality-assurance commentary from Enthu.ai and others: “traditionally only 1-3% of calls were sampled for evaluation” enthu.ai

  2. Otonal and The Asahi Shimbun, “5th Podcast Usage Survey in Japan” (in Japanese) audio-marketing.jp

  3. Google, “YouTube user numbers in Japan” (Think with Google, in Japanese) Google Business

FAQ

Where can the new audio-to-video feature be used?
It is available in the AI video generation service NoLang, where uploading an audio file generates a video automatically.
What kinds of audio assets can be used with the feature?
Audio accumulated inside a company, such as webinar recordings and interview audio, can be reused as video.

Mavericks AI News

Latest Case Studies

Download Materials

Considering NoLang
for your business?

Already using
NoLang?

Create a Video