Convert text to speech with VisionStory smart TTS. Pick a voice from GET /api/v1/voices (public or your cloned voices). The service automatically selects the best synthesis strategy for the voice and the text language. Billing: 2 credits per 1000 characters (rounded up), only charged on success.
Synchronous — the response body is the MP3 audio itself (audio/mpeg). Duration and billing are returned in the X-Audio-Duration-Sec, X-Usage-Characters and X-Cost-Credit headers. The audio is not stored on our side, so save the response body; a repeat call regenerates and bills again. Per-key concurrency is limited during beta; requests beyond the limit are rejected, not queued.
Headers
X-API-KeystringrequiredYour VisionStory API key (sk-vs-...), kept server-side. Create or manage keys in API keys (Pro plan and up).
Request body
application/jsonrequiredtextstringrequiredText to synthesize, up to 3000 characters
voice_idstringrequiredPublic or cloned voice id, check GET /api/v1/voices
Responses
200MP3 audio (44.1kHz). Usage metadata is returned in response headers.
audio/mpeg
defaultError response. All failures share one envelope: an error object with a numeric code, a human-readable message, an optional details string, and an optional hint giving an actionable next step (useful for AI agents).
application/json
errorErrorDetailrequiredcodeintegerrequiredMachine-readable error code. Mirrors the HTTP status for transport-level failures (e.g. 401, 404, 422, 500) and may carry a business-specific code otherwise.
detailsstring | nullnullableOptional structured detail about the failure, e.g. a JSON string of per-field validation errors on a 422. Absent when there is nothing extra to report.
hintstring | nullnullableActionable next step for resolving the error, written for both humans and AI agents (e.g. how to fix the request, or where to obtain an API key). May be absent.
messagestringrequiredHuman-readable explanation of what went wrong. Safe to log or surface to end users; not localized.