Turn an audio recording into SRT subtitles with Scribe and the Ofox API

Transcribe audio with Scribe through Ofox, inspect word timestamps, create SRT subtitles and add a selectable MP4 track using a real recording and scripts.

Typewriter drawn in black ink on a pale card, blue-gray background, mustard circle and the title Scribe: Audio to Subtitles.

Use Ofox’s /v1/audio/transcriptions endpoint to upload a WAV or MP3, request verbose_json, and inspect the returned word timestamps. Then generate SRT locally and check the subtitles against the recording. On the gateway route verified for this tutorial, requesting SRT directly is not the same supported workflow: the adapter accepts JSON output, while the downloadable converter builds the subtitle file from the actual response.

Our October 10, 2026 test transcribed a real 10.00-second ElevenLabs audio file. Scribe returned the script’s words, with punctuation differences, and word timings from 0.14 to 9.78 seconds. We converted those timings into three cues and placed them in an MP4 as a selectable English subtitle track. This demonstrates one complete path; it is not a speech-recognition accuracy benchmark or a test of noisy meetings.

The output bundle and what each file proves

The downloadable kit contains the generated source audio, original API JSON, converter, SRT and MP4 with selectable subtitles. Keep the files together so a subtitle editor can reconstruct where a cue’s timing came from.

ElevenLabs sample audio

FileRoleWhat it does not prove
elevenlabs.mp3The actual input recordingPerformance on human or noisy speech
narration-transcript.jsonUnmodified Scribe responseA manually corrected final transcript
narration.srtLocally grouped, timed captionsThat every target player renders SRT identically
subtitled-selectable.mp4Video, audio and a subtitle streamBurned-in captions visible without player support

The sample voice is synthetic and the text is original tutorial material. It contains no private customer recording. Use material you are authorized to upload when replacing it with a real interview, lecture or call; permission to hear a recording does not automatically establish permission to send it to a cloud transcription service.

1. Prepare the recording without destroying evidence

The tested gateway accepts WAV and MP3. If your source is an MP4 or another container, extract an audio copy first. Preserve the original file and record the conversion command. You should be able to return to the original timeline if you discover that silence, edits or a sample-rate conversion changed an offset.

ffmpeg -i original-video.mp4 -vn -c:a pcm_s16le recording.wav
ffprobe -v error -show_entries format=duration \
  -of default=noprint_wrappers=1:nokey=1 recording.wav

This extraction creates uncompressed audio, so it may be much larger than a compressed source. Do not interpret a larger WAV as better underlying speech. Choose an appropriate supported format for your workflow and check your account’s current upload limits rather than guessing that every long recording will fit in a single request.

Avoid unnecessary edits before the first transcription. Removing a two-second introduction changes every later subtitle’s position relative to the untouched video. If you deliberately transcribe a clip, store its source start time and add that offset when returning the captions to the full video.

For a recording with multiple tracks, select the intended audio stream explicitly. Otherwise an extraction may choose commentary, a translation or an empty track. Probe the input’s streams and listen to the exported file before paying to transcribe the wrong material.

2. Send a multipart request to the correct model

Set your Ofox API key in OFOX_API_KEY, install Python requests, and run the supplied client:

python3 audio_api.py transcribe \
  --input elevenlabs.mp3 \
  --output my-transcript.json

The client sends model=elevenlabs/scribe_v2 and response_format=verbose_json. It lets the HTTP library construct the multipart boundary, opens the audio as a binary file, checks the response, and preserves metadata including the input duration. It does not automatically repeat a failed or timed-out POST.

An equivalent cURL request is:

curl --fail-with-body --silent --show-error \
  https://api.ofox.run/v1/audio/transcriptions \
  -H "Authorization: Bearer $OFOX_API_KEY" \
  -F 'model=elevenlabs/scribe_v2' \
  -F 'response_format=verbose_json' \
  -F 'file=@elevenlabs.mp3;type=audio/mpeg' \
  --output my-transcript.json

Use one client, not both, unless you intend to make another billable call. Do not manually add Content-Type: multipart/form-data without its generated boundary. Authentication headers and multipart field declarations have different jobs; the key goes in Authorization, while the model and file are form fields.

Check the current Ofox Scribe model page for the route exposed to your account. Native ElevenLabs transcription documentation describes the provider interface, which can expose fields that an adapter does not forward. A native provider’s capabilities do not establish that the same parameter works through Ofox.

3. Inspect the real JSON before writing a converter

Our response includes text, language, duration, usage, logprobs and words. Each item in words contains word, start and end. Code written for a different schema might expect text inside each word or a segments list; neither assumption should replace inspecting the file you actually received.

For the example, the returned text begins “A clear product video starts with a clear brief.” The final word ends at 9.78 seconds, while the audio container lasts 10.00 seconds. The end of speech and the end of the file need not be identical: there may be trailing silence or encoding padding.

The following inspection reads the saved response without another API call:

import json
from pathlib import Path
result = json.loads(Path('my-transcript.json').read_text())
print(result['text'])
print(result.get('language'))
print(result.get('usage'))
for item in result.get('words', [])[:5]:
    print(item['start'], item['end'], item['word'])

Preserve the original response before changing spelling or punctuation. If words is absent, a plain transcript does not supply enough evidence to recreate exact subtitle timing. Use a supported timestamp response or a separate alignment step. Dividing the total duration equally among words invents timings and is not an equivalent replacement.

Do not infer speaker identity from a name spoken in the text. This example’s response has word timings, not verified speaker identities. The related meeting-to-action-items workflow keeps that distinction when extracting owners and decisions.

4. Group words into readable subtitle cues

A subtitle cue needs a sequential number, a start and end time, and readable text. The supplied make_subtitles.py groups actual word timings into cues with a simple character and duration limit. It preserves the first word’s start and last word’s end for each group; it does not create a new alignment.

python3 make_subtitles.py my-transcript.json my-subtitles.srt

For the supplied sample, the first cue is:

1
00:00:00,140 --> 00:00:02,980
A clear product video starts with a clear brief.

The next two cues cover 3.56–7.36 and 7.42–9.78 seconds. The gap after the first sentence is preserved rather than filled with a caption merely to make the timeline continuous.

The converter’s 58-character and five-second grouping limits are implementation choices for this short English example, not universal accessibility or broadcast standards. Japanese, Korean and languages with different word spacing need language-appropriate segmentation. Reading speed, line breaks and the amount of text on a phone all matter more than mechanically reusing the English threshold.

Review every boundary. The example’s second cue ends with “and,” which is technically valid but may be a poor editorial break. An editor can move a word between neighboring cues while retaining that word’s actual timestamp. Keep a corrected subtitle version and describe the edit; do not overwrite the raw transcription to conceal the difference.

The converter rejects missing, negative, non-finite and non-monotonic start times. That catches structural problems, not every visual defect. It is still possible to have captions that are too short to read or badly phrased even when the timestamps pass a schema check.

5. Correct the transcript without losing its provenance

Compare the transcript with the recording and your approved source material. Our synthetic example returned the intended words but omitted final punctuation in the short narration. Punctuation normalization is not the same error as a changed product name or missing negation, and should be logged separately.

For real recordings, prioritize names, numbers, dates, units and “not” statements. These are often the details that change what a viewer should do. Do not silently replace an uncertain word with what a presentation slide suggests if the speaker said something else; mark uncertainty or consult an authorized reviewer.

A useful correction table has the original phrase, proposed correction, audio time range, reason and review status. Keep the audio, raw JSON and edited SRT as separate artifacts. When someone asks why a subtitle changed, you can point to a specific source interval instead of relying on a model’s summary.

6. Put the captions in a video and verify the stream

To keep the video and audio unchanged while adding a selectable MP4 subtitle track:

ffmpeg -i narration-video.mp4 -i my-subtitles.srt \
  -map 0:v:0 -map 0:a:0 -map 1:0 \
  -c:v copy -c:a copy -c:s mov_text \
  -metadata:s:s:0 language=eng \
  -disposition:s:0 default subtitled-selectable.mp4

Use the correct language code for your captions. This command adds a text track; it does not burn text into the picture. Some web players ignore that track even when a desktop player can display it. A default-track flag expresses a preference, not a guarantee that every platform will show captions.

Our supplied ten-second MP4 contains H.264 video, AAC audio and a mov_text subtitle stream. The picture is an excerpt from the earlier tutorial’s demo video, paired with the real narration. It is a packaging demonstration, not evidence that the API generated the video itself.

Check the streams and extract the embedded captions for comparison:

ffprobe -v error -show_entries stream=codec_type,codec_name \
  -of json subtitled-selectable.mp4
ffmpeg -i subtitled-selectable.mp4 -map 0:s:0 \
  recovered-subtitles.srt

For burned-in captions, use a subtitle-capable renderer and re-encode the video. The FFmpeg build used for this tutorial did not support the attempted subtitle-rendering filter, so the delivered example uses selectable subtitles. We do not label it as a burned-in or visually player-verified export. Before posting to a specific platform, test the actual uploaded result on its intended player.

7. Diagnose failures without repeating the wrong request

An unsupported-format error calls for checking the gateway’s accepted input and response formats. It does not mean Scribe lacks every corresponding feature in its native API. Convert a copy to WAV or MP3 and retry only after fixing the mismatch.

A quota-related 401 requires the full response body and request ID. Earlier calls in this project failed at the upstream route, and fresh calls worked after restoration. A positive wallet balance alone cannot tell you which provider account handled a request. Preserve evidence rather than repeatedly sending the same recording.

If captions drift steadily, compare the audio actually transcribed with the video audio. A constant offset often suggests a trimmed introduction or clip start. A changing offset can indicate different edits or playback speed. Do not “fix” every cue by guessing until you have identified whether the two timelines are the same.

If the file is long, design a chunking plan with source offsets and overlap review before splitting it. Duplicate words at chunk boundaries and missing context can alter captions. The short-file converter here does not claim to solve long-form alignment automatically.

8. Count the right unit and keep a review gate

Transcription usage relates to audio duration, not the number of returned words. Our file has a 10.00-second probed container duration, while the raw API response reports usage.seconds=10.083265306122449. Preserve both values rather than rounding away the difference. Do not substitute the final spoken word’s end time for the input’s billable duration, and do not infer a settled charge from this usage field alone.

Before treating subtitles as complete, require five checks: the correct authorized recording was uploaded; the saved JSON is a successful response; timing comes from actual words; edited text has been checked; and the target player displays the intended subtitle version. The downloadable kit establishes the API, conversion and container checks. Audience-facing playback remains a separate publication check for your environment.

Frequently Asked Questions

Can I request SRT directly from this Ofox endpoint?
The route verified here accepts JSON or verbose JSON. This guide converts the returned word timestamps locally instead of assuming a native provider's SRT option is forwarded.
What if the response has text but no words array?
Keep the transcript, but do not invent word timing. Obtain a supported timestamp response or run a separate alignment process before making timed subtitles.
Are these burned-in captions?
No. The supplied MP4 has a selectable mov_text subtitle stream. Player support varies; burned-in captions require rendering text into the video frames.
Does this test establish Scribe's accuracy on meetings?
No. It is a short, clean synthetic narration. Real meetings add overlap, noise, names and speaker ambiguity that require their own evaluation and correction process.