Skip to content

Captions & text ​

"Captions & text" in LmBox handles the whole caption flow on the desktop: fetch, recognize, proofread and write back. You can read existing captions from the timeline, or recognize dialogue online or on your device, fix the text in a list, split or merge lines, then export or import into your editing app. The sidebar icon is .

Video tutorial ​

Video coming soon

What you will learn

  • Get captions from the timeline, import from a file, and recognize online or on your device
  • Edit the list and use the common shortcuts
  • Tasks and versions, find and replace, merge and split, script align, AI correction
  • Create & export to the timeline / media pool / SRT / text
  • Convert captions to templates: batch-convert to Essential Graphics (MOGRT) or Fusion titles

Guide

This is the main Lab guide for the LmBox desktop app. Screenshots on this page show the Chinese interface. You can switch the app to English in Settings. If you want to remove repeats, mark highlights or use AI on a word-level transcript, use Talking rough cut.

What it solves ​

Pain pointWhat this page does
Captions already on the timeline need editingGet from Timeline, edit in the list, then write back
You have sound but no captionsOnline or On device recognition makes sentence-level captions
Recognized text is too choppy or too longMerge and Split, and list shortcuts to split or join lines
You already have a final script and want captions from the audioScript align: drop audio or take it from the timeline, then generate from the script
You want template stylesConvert to template: batch-write Premiere Essential Graphics (MOGRT) or DaVinci Resolve Fusion titles

Before you start ​

  1. Pick the target editing app in the sidebar ( / / ). Its extension icon should be lit.
  2. Online recognition: sign in first, and follow the on-screen steps to set up an available recognition service.
  3. On-device recognition: in Settings, on the Large models page, download and apply the runtime and the caption model pack. Use an English path for the model folder.
  4. Before recognizing, make sure the timeline has audio that can be exported (or have a local audio/video file ready).

The desktop screen ​

Open from the sidebar. From top to bottom: top bar actions → (when recognizing) recognition panel and waveform → caption list.

Captions & text in the LmBox desktop app

Figure: the top bar (get / import / recognize and "Create & Export Subtitles") and, below it, the caption list with editing hints. The screenshot shows the Chinese interface.

Screen areas ​

AreaWhat it does
Get from TimelineReads captions / titles / graphics from the connected extension (choose the type)
Import from LocalImports a caption file from your computer into the current list
OnlineOpens the online speech recognition panel and takes audio from the timeline or the selected layers
On deviceRecognizes on your computer. Take audio from the timeline, or drop audio/video to preview before you start
Tool iconsTask and Version History, Script align, AI Correction, Find and Replace, Merge and Split, Convert to template
Create & Export SubtitlesImport to the timeline / media pool, or export SRT or text
Caption listTick lines, edit text, see character counts and in/out points; use shortcuts to tidy up
  1. Connect the extension and open this page.
  2. If you already have captions, click Get from Timeline or Import from Local. If not, use Online or On device.
  3. Edit text and split or join lines in the list. When needed, use Find and Replace, Merge and Split, Script align or AI Correction.
  4. Click Create & Export Subtitles and choose the timeline, the media pool, or a file.
  5. At important points, save a version in Task and Version History so you can go back.

Get and import ​

Get from Timeline ​

Choose how to get captions and the caption type (captions, titles, Essential Graphics, Fusion, source text and so on, depending on what the target extension supports), then start. The result goes into the current task list.

Import from Local ​

Pick a caption file on your computer. It becomes a new task or replaces what you are editing (follow the confirmation on screen).

Online and on-device recognition ​

Each mode opens its own panel and keeps its own options. Starting a new recognition usually clears the previous captions, so scripts do not mix. Click Start to take audio from the timeline, or drop audio/video into the dashed area first. The waveform below lets you listen to the prepared audio.

OnlineOn device
Audio sourceTimeline audio; in After Effects, export by the selected layersTimeline audio, or drop local audio/video to preview before you start
Audio range / tracksChoose a range (such as the timeline in/out points) and audio tracks; you can click "Refresh"Same
Common optionsDigits, soften filler words, add punctuation, semantic sentence breaks, timestamp align, speakers, sensitive words, max characters per line, and moreRecognition model (such as Qwen), speakers, show punctuation, hot words, and more
NoteNeeds sign-in and available capsSet up the runtime and model pack first; your computer's performance affects speed and quality

Online ​

Click Online in the top bar to open the panel. Set the range and options, then click Start.

Online speech recognition panel

Figure: online recognition with the drop area, audio range and tracks, option switches and waveform preview. The screenshot shows the Chinese interface.

On device ​

Click On device in the top bar to open the panel. Before recognizing, apply the runtime and the caption model in Settings. The hot words box takes names, brands and other proper nouns.

On-device speech recognition panel

Figure: on-device recognition with model choice, speakers / show punctuation, hot words and waveform preview. The screenshot shows the Chinese interface.

While recognizing, the list can highlight along with playback. Next to the waveform you can use Scroll to current subtitle to bring the playing line into view.

Caption list and shortcuts ​

The list shows the number, track, text, character-count hint and in/out points. When it is empty, follow the top bar hint to get, import or recognize.

Common shortcuts while editing (also shown at the top of the screen):

ShortcutWhat it does
Shift + EnterSplit caption
Shift + ↑ / ↓Merge captions
Tab / Shift + TabMove between lines
EscLeave editing
SpacePlay / pause the recognition audio

Tick several lines to use the tools in bulk. Clearing the list asks you to confirm first; after that you need to get or import again.

Tool drawers (top bar icons) ​

The row of icons in the middle of the top bar opens a drawer on the right. Task and version history comes first, then the others in the table.

Task list and version history ​

Click to open the drawer on the right. It has two tabs:

TabWhat it does
TasksPast recognition / import tasks, with status, line count, version count, the Online / On device tag and "Done"
VersionsSaved versions of the current task. You can switch back to a version, and the edit area supports undo / redo

Click a task to switch the current caption draft. You can pin or delete a card. At important points, save a version before you start a new recognition or rewrite a lot, so nothing is lost when the list is cleared.

Task list and version history

Figure: the "Tasks" tab in the right drawer, with recognition task cards, line and version counts and the Online tag. The screenshot shows the Chinese interface.

Script align ​

Click Generate subtitles in the top bar and choose Script align. Drop audio/video, or export audio from the timeline, then paste your final script one sentence per line (a new line makes a new caption). Format script breaks lines at periods, question marks and commas, and Remove punctuation strips the punctuation at the end of each line. Then click Generate from script. You do not need to run recognition first. If a pack is missing, the app jumps to the Large models page in Settings. The result is a new task, with caption blocks fitted to the speech and starting slightly early to cover the first sound.

Script align

Figure: "Script align" inside Generate subtitles, with the audio and the final script, and Generate from script. The screenshot shows the Chinese interface.

AI correction ​

Click to open the drawer on the right. You can choose a route such as Official cloud / Golden key / Keychain / your own model (follow what the screen shows), tick the jobs to run, then click Submit repair job. Long videos may be written into the list in several rounds.

Job (pick several)What it does
ReferencePaste the original script or proper nouns to help check the text
Fix typosFixes wrong words and awkward lines
Split and merge linesSplits long lines to read better and merges very short ones
Remove fillersRemoves um, ah and similar filler words
Fine-tune timingReduces overlaps and follows the speaking rhythm
TranslateTranslates the caption text
Unify names and spellingMakes names, brands and similar terms consistent

An optional "Include word-level text" uses more caps and is off by default. When split and merge is ticked, caption lines are really added or removed.

AI correction

Figure: AI correction with the Official cloud route, the job checkboxes and "Submit repair job". The screenshot shows the Chinese interface.

Find and replace ​

Click to open the drawer on the right. Matches in the list are colored to tell the original from the replacement.

SectionWhat it does
Block Words SettingsModal particles, punctuation, sensitive words, and custom sensitive words (press Enter to add)
Replace SettingsA list of "Replace A with B" rules; add or delete rules

You can also detect blocked words from the timeline audio and insert or overwrite censor beeps in one click, without changing the caption text. Rules apply on import / export and when saving.

Find and replace

Figure: find and replace with block word tags, replace rules, and the red / green highlight in the list. The screenshot shows the Chinese interface.

Merge and split ​

Click to open the drawer on the right. Merge and Split have separate switches. We suggest clicking Preview to see how the line count changes, then Apply to current subtitles (this is saved as a new version).

SectionCommon options
MergeMerge overlapping lines, short neighbors, and lines that do not end a sentence; same track only; how to join text (space / new line / no separator); max duration, characters and gap for short lines
SplitSplit at punctuation (period, exclamation and so on); split very long lines evenly as a fallback; split reference duration and characters

Merge and split

Figure: merge and split with the merge options, split at punctuation, and Preview / Apply to current subtitles. The screenshot shows the Chinese interface.

Convert to template ​

Click to open the drawer on the right. It batch-writes the captions in the list into styled template clips:

Target appHow to use it
DaVinci ResolveDrag your finished Fusion style into the media pool → click Load Fusion media-pool templates → choose the template and track → Convert
Premiere ProSelect an Essential Graphics (MOGRT) clip with text on the timeline → Load template from selection → choose the text control → Convert (the matching extension must be connected; what it can do depends on the version)

You can also set the time reference for creation (such as the playhead), the target track, and the caption type / source filters. The screen shows "N line(s) will convert", and you can Stop while it runs.

Convert to template

Figure: convert to template with Load Fusion media-pool templates, time reference, track and Convert (DaVinci Resolve side shown). The screenshot shows the Chinese interface.

Create & Export Subtitles ​

Click Create & Export Subtitles at the right of the top bar to open the dialog. Choose the Create Type, then Time Alignment (when writing to the timeline), and finally click Start Create/Export.

Create TypeDescription
Import to TimelineWrites to the timeline of the connected editing app
Import to Media PoolWrites to the media pool (DaVinci Resolve and similar)
Export SRTExports a caption file to your computer
Export TextExports plain text
Time AlignmentDescription
Original TimePlaces captions by their own in/out points
Timeline StartAligns to the start of the timeline
Current PlayheadAligns to the playhead
Clip StartAligns to the start of the clip

You cannot create when the list is empty. Get or recognize captions first.

Create and export captions

Figure: Create & Export Subtitles with the create type and time alignment, then Start Create/Export. The screenshot shows the Chinese interface.

FAQ ​

On-device recognition says the model is not set up?
Go to the Large models page in Settings, download or upgrade the runtime, apply the caption model pack, then try again.

The result is mixed with the previous round?
A new recognition clears the previous one. If you want to keep it, save a version in Task and Version History first.

Blocked words / replacements do not apply after import?
Check that the rules are set in Find and Replace. A newer version fixed export still writing the original text, so keep LmBox up to date.

I only want a word-level rough cut of talking video?
This page is a sentence-level caption workflow. For word-level cuts and highlights, use Talking rough cut.

Release notes ​

  • 2.9.0 With older Premiere Pro versions that cannot put imported captions on the track, the app now opens the project caption folder so you can drag them onto the timeline by hand.
  • 2.8.0 Script align formatting now also breaks lines at commas, and you can remove the punctuation at the end of each line in one click.
  • 2.8.0 Fixed taking audio from the timeline still using the old file after you changed in/out points or tracks, and side tracks leaking in. It now exports again by the chosen range and tracks.
  • 2.7.0 Script align now generates straight from your final script after you drop audio or take it from the timeline. Caption blocks fit the speech and start slightly early to cover the first sound. No need to recognize first.
  • 2.1.1 Fixed a problem where audio could not be taken from a track after a transition was added in DaVinci Resolve. Recognition now works.
  • 2.0.0 Fixed on-device recognition failing to prepare audio when a DaVinci Resolve timeline used proxy media. It now works.
  • 1.6.0 List editing supports Tab / Shift+Tab to move between lines, Esc to leave editing, and Space to play / pause the recognition audio. Next to the recognition waveform you can click "Scroll to current subtitle" to bring the line into view.
  • 1.6.0 Caption AI correction can use "Official cloud / Golden key", so you do not need your own model key.
  • 1.6.0 Fixed the list jumping back to the top or jittering when merging or splitting with shortcuts. The editing position now stays steady.
  • 1.6.0 Caption AI correction options now use clearer job names (fix typos, split and merge lines, remove fillers and more). With split and merge ticked, lines are really added or removed, and long videos are repaired in rounds, written into the list round by round.
  • 0.3.0 Online and on-device recognition can choose the audio range and track. DaVinci Resolve audio follows the timeline in/out points, and opening the panel refreshes the track list.
  • 0.2.9 Improved the stability of on-device recognition. Once the runtime and model are ready, recognition runs more smoothly.
  • 0.2.8 Fixed a false "Set up the on-device model in Settings first" message with the installed app. Recognition works once setup is ready.
  • 0.2.8 Fixed recognition failing when the runtime from the offline pack was old. Download the runtime again on the Large models page in Settings, or apply the pack again.
  • 0.2.8 Fixed on-device recognition saying the model was not set up or could not start on Windows when the model path had Chinese characters.
  • 0.2.6 Fixed "could not prepare" when running on-device recognition from the editing timeline. Captions are now recognized normally.
  • 0.2.5 In Find and Replace you can detect blocked words from timeline audio and insert or overwrite censor beeps in one click, without changing the captions.
  • 0.2.4 Added "Script align": paste a final script (one sentence per line) and align the time of each sentence.
  • 0.2.4 Fixed speaker labels still showing after turning off "Speakers" in online recognition.
  • 0.2.3 You can choose an on-device model for recognition and get sentence-level and word-level captions on your computer.
  • 0.2.3 Reworked the online and on-device recognition flow: options are saved separately, and starting a new recognition clears the previous captions so they do not mix.
  • 0.2.3 On-device recognition lets you drop audio/video to preview the waveform and start by hand. The list and the waveform progress highlight together while recognizing.
  • 0.2.3 On-device recognition can separate speakers and splits caption lines by punctuation first. The list hides punctuation by default (sentences are still split by punctuation).
  • 0.2.3 Better task history and version management: browse many recognition tasks, save and switch versions, and undo / redo in the edit area.
  • 0.2.3 Better "Merge and Split": merge and split have separate switches, with split at punctuation and merging of continued lines to tidy recognized text.
  • 0.2.3 The Premiere Pro LmBox extension now works with this page for speech-to-caption and reading captions from the timeline.
  • 0.2.3 Task list and version history are clearer, and the list and current content are less likely to disagree when you switch tasks.
  • 0.2.3 Fixed blocked words or global replacements not applying when importing to the editing app, exporting or saving; the original text was still written.
  • 0.2.3 Fixed the list not highlighting with progress when playing the recognition preview, and list checkboxes only sliding but not ticking on a click.
  • 0.2.2 With After Effects connected, "Online" can export audio from the selected layers and generate captions.