openAiWhisperApiToCaptions()v4.0.217
Turns the output from openai.audio.transcriptions.create from the openai package into an array of Caption objects.
This package performs processing on the captions in order to retain the punctuation in the words, which is not by default included in the OpenAI response.
This function can be used in any JavaScript environment, but you should not use the OpenAI API in the browser because your API key will be exposed to the browser.
Example usageimportfs from 'fs'; import {OpenAI } from 'openai'; import {openAiWhisperApiToCaptions } from '@remotion/openai-whisper'; constopenai = newOpenAI (); consttranscription = awaitopenai .audio .transcriptions .create ({file :fs .createReadStream ('audio.mp3'),model : 'whisper-1',response_format : 'verbose_json',prompt : 'Hello, welcome to my lecture.',timestamp_granularities : ['word'], }); const {captions } =openAiWhisperApiToCaptions ({transcription });
Inputv4.0.530
Accepts a verbose JSON transcription with text and timed words (word, start, end in seconds). Word-level conversion reconstructs punctuation from the full text. If words are absent or empty, timed segments (text, start, end in seconds) produce one caption per segment, without invented word timings. Invalid supplied word timings throw an error rather than falling back to segments. A transcription without timed entries throws an actionable error. Conversion is local and never calls OpenAI.
Compatibility
| Browsers | Servers | Environments | |||||||
|---|---|---|---|---|---|---|---|---|---|
Chrome | Firefox | Safari | Node.js | Bun | Serverless Functions | ||||