@remotion/whisper-webgpuv4.0.518
Transcribe audio locally in the browser or Node.js using timestamped Whisper models and WebGPU through Transformers.js, and convert the result to @remotion/captions.
Installation
- Remotion CLI
- npm
- bun
- pnpm
- yarn
npx remotion add @remotion/whisper-webgpu @huggingface/transformers
This assumes you are currently using v4.0.534 of Remotion.npm i --save-exact @remotion/[email protected] @huggingface/[email protected]
Also update
remotion and all `@remotion/*` packages to the same version.Remove all
^ character in front of the version numbers of it as it can lead to a version conflict.This assumes you are currently using v4.0.534 of Remotion.pnpm i @remotion/[email protected] @huggingface/[email protected]
Also update
remotion and all `@remotion/*` packages to the same version.Remove all
^ character in front of the version numbers of it as it can lead to a version conflict.This assumes you are currently using v4.0.534 of Remotion.bun i @remotion/[email protected] @huggingface/[email protected]
Also update
remotion and all `@remotion/*` packages to the same version.Remove all
^ character in front of the version numbers of it as it can lead to a version conflict.This assumes you are currently using v4.0.534 of Remotion.yarn --exact add @remotion/[email protected] @huggingface/[email protected]
Also update
remotion and all `@remotion/*` packages to the same version.Remove all
^ character in front of the version numbers of it as it can lead to a version conflict.Example
transcribe.tsimport {clearStaleModels ,downloadWhisperModel ,resampleTo16Khz ,toCaptions ,transcribe } from '@remotion/whisper-webgpu'; export consttranscribeFile = async (file :File ) => { awaitclearStaleModels (); awaitdownloadWhisperModel ({model : 'small.en'}); constchannelWaveform = awaitresampleTo16Khz ({file }); consttranscription = awaittranscribe ({channelWaveform ,model : 'small.en',language : 'en', }); const {captions } =toCaptions ({whisperWebGpuOutput :transcription }); returncaptions ; };
Model hosting
We noticed that hosting the models on R2 leads to faster downloading than through Hugging Face directly.
We mirrored the models to remotion.media, keeping them byte-identical.
We also notified Hugging Face and they are investigating the issue. We may migrate back to the official hosting in the future.
On the server
See @remotion/whisper-webgpu in Node.js for a complete Node.js example.
APIs
canUseWhisperWebGpu()
Check whether transcription is possible
getAvailableModels()
List models and their download sizes
clearStaleModels()
Remove models discontinued by newer versions
isWhisperModelCached()
Check whether a model is downloaded
downloadWhisperModel()
Download a model
loadWhisperModel()
Initialize a downloaded model
disposeWhisperModel()
Release model memory
removeWhisperModel()
Remove a model from the persistent cache
transcribe()
Transcribe a waveform with word-level timestamps
toCaptions()
Convert a transcription to
@remotion/captionsresampleTo16Khz()
Decode and resample browser audio
Requirements
You obviously need a GPU.
Use canUseWhisperWebGpu() to determine if it is supported.
License
MIT