ElevenLabs through Ofox: generate, save and verify your first API voiceover
Generate your first ElevenLabs voiceover through Ofox with Python or cURL. Verify real MP3 output, model and voice settings, errors and downloadable examples.
To generate speech through Ofox, send text to /v1/audio/speech with an Ofox API key, the model ID elevenlabs/eleven_v4, a compatible voice ID and an accepted audio format. Save a successful binary response as audio, then check that it decodes. A file named speech.mp3 is not proof that the request worked: an authentication error saved under that name is still JSON.
This tutorial follows a real request completed on October 10, 2026. The supplied original English script produced a 10.00-second MP3, and a separate Scribe request transcribed that audio. The downloadable files and client let you inspect those results. This is one successful example through the Ofox gateway, not a reliability benchmark or a claim that every ElevenLabs feature is exposed by this route.
What you will finish
The deliverable is a playable voiceover, its source text, a request configuration without secrets, and a small verification record. You can use it in a product video or pass it to a transcription workflow. You do not need to start with a video editor, a cloned voice or a paid ElevenLabs account of your own to reproduce the Ofox request; you need an Ofox key with access and sufficient available balance, and the configured upstream route must work.
The example uses an existing voice ID. It does not enroll a new voice or clone a person. If you later use a voice associated with a real person, obtain the relevant permission and check the applicable rights separately. A successful synthesis call does not establish permission to impersonate someone or publish every use of their voice.
Download the complete working kit, including audio_api.py, comparison.txt, elevenlabs.mp3 and the saved response metadata. The kit uses environment variables for authentication and stops on unsuccessful requests rather than automatically repeating a potentially billable POST.
ElevenLabs sample audio
1. Prepare the key, tools and an exact script
Install Python, the requests package, and FFmpeg with ffprobe. The HTTP request itself needs only an HTTP client; FFmpeg is included here because decoding and duration checks are part of the acceptance criteria. Confirm the commands resolve in the same terminal where you will run the tutorial:
python3 -m pip install requests
ffmpeg -version
ffprobe -version
Set OFOX_API_KEY using your usual secret manager or a private environment configuration. Do not put the key in a Markdown article, source file, screenshot or committed .env. If you have already exported it, the provided client reads it directly. The script deliberately does not print the Authorization header.
Use the text in comparison.txt for the first test:
A clear product video starts with a clear brief. Show the real interface, explain one useful task, and check the exported video before sharing it.
Save the file as UTF-8. Keep the exact wording while troubleshooting: changing the text, voice, model and format together makes failures harder to isolate. For a multilingual project, prepare and review each language’s script before generating it; translation and synthesis are different operations.
For your own script, replace unexplained abbreviations and resolve ambiguous dates before submitting. “October 15, 2026” is less ambiguous than “10/15.” A pronunciation requirement should be tested in the final audio rather than assumed from a text preview. The example avoids promising a particular emotion or pronunciation control that has not been verified through this gateway.
2. Use the gateway’s model and voice fields
The tested request has four essential content choices:
| Field | Tested value | Why it matters |
|---|---|---|
model | elevenlabs/eleven_v4 | Ofox’s provider-prefixed model ID |
voice | JBFqnCBsd6RMkjVDRZzb | The voice identifier accepted in this test |
response_format | mp3_22050_32 | The requested MP3 preset, not a filename |
speed | 1.0 | The submitted speed value; not a guaranteed duration |
Check the current ElevenLabs model page before using a different model or format. Do not substitute a display label for the exact ID, and do not assume that a voice name from another provider is a valid ElevenLabs identifier.
The Ofox endpoint and the native ElevenLabs API are different interfaces. Native examples may put a voice ID in the path or accept provider-specific fields. For this tutorial, keep the voice in the JSON field shown here and use the Ofox base URL. Passing every option from a native provider reference into an adapter is not a compatibility test.
The current example is deliberately small. Add a language field, alternative voice or advanced setting only after checking that it is accepted by your current Ofox route. If you are building a reusable wrapper, keep provider-specific options in a separate configuration object rather than presenting every speech model as interchangeable.
3. Generate the first MP3 with the verified Python client
From the unpacked kit directory, run:
python3 audio_api.py speech \
--engine elevenlabs \
--text comparison.txt \
--output my-first-voiceover.mp3
Choose a new output name. The client rejects an existing output path so that a second experiment cannot silently overwrite evidence from the first. It saves the submitted content configuration alongside the response metadata, while obtaining the key only from the environment.
On success, the client checks the response type, saves bytes, probes the audio duration and attempts decoding. The expected result is an MP3 plus metadata, not a JSON object containing a link to an audio file. On failure, inspect the separate error record; do not feed that response to a video editor as if it were speech.
Our example returned HTTP 200 with audio/mpeg, 40,456 bytes and a probed duration of 10.00 seconds. The request took 2.861 seconds in the client measurement. That elapsed time includes this request’s network and processing conditions; it is neither time to first audio nor a service latency guarantee.
Listen to or download the actual ElevenLabs sample. Keep the original audio unchanged when documenting the request. If you normalize loudness or trim silence for a video, save that as a separate editorial derivative so another developer can still examine the raw result.
4. Reproduce the HTTP request with cURL
The following alternative writes a temporary response and moves it into place only after a successful HTTP response:
python3 - <<'PY'
import json
from pathlib import Path
payload = {
'model': 'elevenlabs/eleven_v4',
'voice': 'JBFqnCBsd6RMkjVDRZzb',
'input': Path('comparison.txt').read_text().strip(),
'response_format': 'mp3_22050_32',
'speed': 1.0,
}
Path('speech-request.json').write_text(json.dumps(payload))
PY
curl --fail-with-body --silent --show-error \
https://api.ofox.run/v1/audio/speech \
-H "Authorization: Bearer $OFOX_API_KEY" \
-H 'Content-Type: application/json' \
--data-binary @speech-request.json \
--output speech-response.tmp \
&& mv speech-response.tmp curl-voiceover.mp3
This is an alternative invocation, not a required second paid request. Running both clients generates speech twice. The article’s successful result was produced with the supplied Python client; the cURL snippet shows the equivalent request construction.
With an older cURL that lacks --fail-with-body, use the Python client or explicitly check the HTTP status before accepting the output. The temporary file may contain a useful error body after failure. Delete or archive it only after you have read the error; never rename a failed body to .mp3 simply to make a player open it.
5. Verify the file before calling the task complete
Use ffprobe to inspect the actual container and stream, then decode the whole file:
ffprobe -v error -show_entries \
stream=codec_name,sample_rate,channels:format=duration \
-of json my-first-voiceover.mp3
ffmpeg -v error -i my-first-voiceover.mp3 -f null -
A zero-error decode is a technical check. It does not tell you whether the brand name is pronounced correctly, whether a pause matches your visual cut, or whether the voice suits the audience. For publication, listen through the complete export at normal volume and compare it with the approved script.
For our sample, a follow-up Scribe transcription returned the same words with punctuation differences. That is a useful secondary check, not a replacement for listening: speech recognition can normalize mistakes or miss an unwanted sound. Do not use a second model’s agreement as proof of perfect audio quality.
A practical acceptance record should name the input version, model, voice, requested format, actual duration, whether decoding passed, and who checked pronunciation. Record unanswered items explicitly. A successful API request with unchecked content is ready for editorial review, not automatically ready for a customer campaign.
6. Fit the voiceover to a video without hiding a timing problem
Plan the video around the measured duration. The same script may take a different amount of time with another voice or a later generation. Setting speed to 1.0 does not make a ten-second timeline fit every result.
If narration is longer than the picture, extend or restructure the relevant scene before exporting. If it is shorter, a deliberate visual hold can be preferable to stretching speech. Do not use a shortest-stream option without understanding whether it will truncate the end of the video or voiceover.
The related Ofox video voiceover tutorial covers measuring audio and assembling an MP4. For subtitles, use the Scribe recording-to-subtitles guide to preserve actual timing rather than guessing cue boundaries from text length.
Keep the development example and the marketing claim separate. This audio demonstrates that the request produced a usable file under the recorded configuration. It does not establish that the voice will handle every language, every technical term or every commercial delivery format equally well.
7. Understand cost without confusing counters
Speech costs need a unit. A text-character price, an audio-token price and a per-second transcription price cannot be compared by placing their raw numbers side by side. Consult the model page and account’s request-level billing record for the applicable configuration.
We do not label the file size or the request elapsed time as billable usage. A 40-kilobyte download does not imply a particular fee. The tutorial preserves the model and request so you can reconcile the relevant entry in your own usage history. It does not present a catalog estimate as the settled bill for this sample.
For a production workflow, separate accepted outputs from retries and rejected takes. Ten requested takes that produce one approved voiceover have a different effective cost from one successful take. Track all billable calls, then calculate cost per accepted deliverable. A lower nominal rate alone does not establish the less expensive workflow.
8. Troubleshoot the exact failing stage
| Symptom | What to inspect | Safe next action |
|---|---|---|
| HTTP 401 mentioning quota | Response body and request ID | Check the account or upstream route actually used; do not assume the Ofox wallet is empty |
| HTTP 401 mentioning credentials | Key environment and authentication | Correct the credential source without printing it |
| HTTP 400 or unsupported format | Model, voice and format combination | Return to the tested configuration, then change one field |
| A tiny “MP3” will not play | Content type and first response bytes | Read the error as text; regenerate only after fixing the cause |
| Timeout with unknown outcome | Provider/request logs | Determine whether the call completed before repeating it |
| Words are missing or unsuitable | Original audio and script | Revise the script or voice configuration and save a new take |
Earlier attempts for this tutorial returned an upstream quota error even while the Ofox wallet was reported positive. After the route was restored, fresh calls succeeded. This is evidence about those requests, not a rule that every 401 comes from a provider balance problem.
Frequently Asked Questions
- Do I put an ElevenLabs key in this example?
- No. This example authenticates to the Ofox gateway with an Ofox API key. Native ElevenLabs authentication is a different integration.
- Can I use any voice name?
- Do not assume so. The successful request used the exact identifier shown above. Check compatibility before replacing it with a display name or another provider's voice.
- Why does my saved MP3 contain JSON?
- The server likely returned an error that your client saved without checking the HTTP status and response type. Read that error first; changing the extension does not fix it.
- Does HTTP 200 mean the narration is ready to publish?
- It confirms a successful response, not editorial approval. Decode the file, check its full duration and listen for omitted words, pronunciation and timing before publishing.


