Mavericks, Inc., the Japanese company behind the AI video generation service NoLang, has added more than 300 high-quality voices to NoLang and, at the same time, released a voice search function that pulls up the right voice instantly. NoLang is the first and largest AI video generation service in Japan1. With this update, users can pick anything from an energetic voice that keeps viewers watching a short video to a deep, weighty voice that builds trust in IR and branding content. By optimizing the audio side of a video, which shapes its quality as much as the visuals do, the update supports marketing, PR, and internal education programs.
Changing business results through the voice in your videos
As companies use video more and more, audio has become as important as the visuals in deciding how good a piece of video content is.
In marketing and PR, for example, short videos on TikTok or YouTube Shorts live or die by the impact of the first few seconds of audio, which directly affects how many people keep watching. Product videos aimed at decision-makers in a sales process, or IR videos for investors, need something different: a low, calm, weighty voice that conveys trust and authority.
Preparing that kind of voice in-house is not easy. The result depends on the voices of the employees available, so the ideal voice is rarely at hand. High-quality recording also requires buying dedicated equipment and having the skills to use it. Hiring a professional narrator instead costs hundreds of thousands of yen, which is a serious hurdle for companies with limited budgets. And because a recording is hard to change once it is made, projects tend to move cautiously, which stretches the lead time from planning to delivery.
More than 300 new voices added to NoLang
With this update, NoLang has added more than 300 varied voices. Anyone can now use professional-grade narration without expensive outsourcing or equipment. A new search function makes that library usable, so a voice that fits any business situation can always be found. Changing a line of narration, which in outsourced production meant extra fees and a delayed delivery, now takes nothing more than editing the text in the editor. The AI regenerates the audio immediately, so content can be refined as many times as needed at no extra cost and with no added lead time.
Watch a product video with an anime-style voiceNew feature 1: A wide range of new voices for every audience
The library of more than 300 new voices was selected based on the impression and the effect each voice creates, and it covers the situations that come up in business.
A young, energetic voice, for instance, works as a strong hook for Gen Z and social media users. It grabs attention in the opening moment even in a fast-scrolling feed, which makes it a good fit for short video ads, campaign announcements, and other content where completion rate matters.
A low, calm, weighty voice that recalls a movie trailer or a documentary shows its value when the goal is to convey trust and authority to decision-makers and investors. That depth suits formal settings that have to carry the dignity and credibility of the brand, such as corporate philosophy messages, promotions for premium products, and IR videos.
A casual, friendly voice is right for content that aims to resonate with a broad audience. Its storytelling quality works for videos that dramatize how a product is used, or internal newsletters that convey what employees are like, closing the psychological distance with viewers and raising engagement.
New feature 2: Voice search that cuts production time
Along with the more than 300 new voices, the search function was rebuilt so that users reach the right voice in the shortest possible path through a large library. Combining six criteria, Voice Engine, Gender, Age, Voice quality, Purpose, and Tone, pinpoints the intended voice instantly and cuts the time spent on selection.

Filter by audience attributes
Selecting Gender (Male, Female, Unknown) or Age (Kids, young, adult, Elderly) immediately surfaces candidates that match the target persona of the video. Choosing a voice that feels familiar to the audience prevents drop-off and helps improve viewer retention.
Select by tone of voice
Voices can be filtered by emotion parameters such as Fun, Glad, and Anger, and by pitch information. Whether the intent is to sound approachable or to build confidence, the tone that sets the mood of the video can be chosen freely.
Work backward from the use case
Selecting a concrete business situation such as Advertisement・PR, Training, Presentation, or Customer Support calls up voices suited to that purpose. Even when it is unclear which voice is appropriate, the right one for the goal can be settled on immediately.
Practical use cases
With this feature, the people responsible for video at a company can run concrete programs like the following and address their business challenges.
Marketing and social media: prevent drop-off in the first three seconds and maximize the conversion rate of short video ads
The biggest problem with short video ads on TikTok or YouTube Shorts is that viewers skip within the first few seconds and leave before the product has made its case. This feature makes it possible to search for and apply a voice optimized for the ad’s target persona by combining Gender, Age, and Tone. Well-paced narration strengthens the opening hook by appealing to hearing as well as sight. Viewers stop scrolling, retention improves sharply, and click-through and conversion rates rise.
Watch a vertical product video with an anime-style voiceIR, company introductions, and PR: establish a sonic brand and build trust in the company
When the quality and tone of the narrator differ from video to video in company introductions or IR materials, the brand image never settles and the impression is inconsistent. Defining a weighty, composed voice from the NoLang library and using it consistently unifies the company’s voice. An authoritative low voice for earnings briefings, an intelligent and sincere voice for recruiting videos: a deliberate audio identity like this builds unwavering trust among stakeholders and spreads the value of the brand.
HR and internal training: raise e-learning completion and retention with an instructor’s delivery
The flat, mechanical narration common in e-learning materials reduces learners’ concentration and weakens the effect of the training. Selecting the Training or Teacher tag in the new purpose search instantly calls up an instructor-style voice with the intonation and pauses of a professional teacher. A natural, human voice does not wear learners down over a long session. Content is watched to the end, which improves retention of the material and strengthens compliance awareness across the organization.
Watch a training video with an announcer-style voiceWhat comes next
Beyond the greatly expanded range of voices, Mavericks plans to deepen the expressive power of non-verbal communication as well, including subtle changes in avatar expressions and gestures. The aim is a video experience that moves viewers more deeply and fundamentally improves the quality of business communication in marketing, IR, and internal education.
Contact usFootnotes
-
Based on Mavericks’ own research into registered user counts for AI video generation services ↩