AI agents can generate VisionStory videos, images, speech, avatars, and voices through MCP, the Python SDK, CLI, or machine-readable documentation. The Agent Skill covers talking-avatar generation, voice discovery, localized speech, transcription, and word alignment. Every channel uses the same REST API and VISIONSTORY_API_KEY; choose the integration that matches where your agent runs.
Choose an AI agent integration
| Channel | Install | Best when |
|---|---|---|
| MCP server | uvx visionstory-mcp — your client runs it, nothing to host | Claude Desktop, coding agents, and other local MCP clients |
| CLI | curl -fsSL https://developers.visionstory.ai/cli | bash (latest release) | People and agents that want one persistent terminal command |
| Python SDK | pip install visionstory | Agents that write and run Python |
| Agent Skill | npx skills add visionstory-ai/skills --skill visionstory-api | Skill-compatible coding agents working with avatars, speech, or audio inside a repo |
llms.txt + .md pages | nothing — just fetch the URLs | Agents that browse the web and read docs on demand |
Whichever channel you use, configure the key the same way — the agent reads it from an environment variable. If the variable is missing, the agent must explain how to set it locally and must never ask you to paste the key into chat:
export VISIONSTORY_API_KEY="sk-vs-xxxxxxxxxxxxxxxxxxx"
Before asking an agent to submit a paid generation, run the read-only visionstory credits command or call the MCP get_credits tool. A 401 response means the local key is missing, expired, or invalid and must be replaced on the API keys page.
Connect the VisionStory MCP server
For most users, the recommended local setup below is enough. Your MCP client starts VisionStory on demand and receives the same 27 public operations available through the SDK and CLI, including audio transcription and alignment.
Recommended setup — Run with uvx 👍
Install uv once. Skip this command if uvx --version already works:
curl -LsSf https://astral.sh/uv/install.sh | sh
Confirm VisionStory MCP can start:
uvx --refresh visionstory-mcp --check
When it prints ✓ VisionStory MCP started successfully, add this configuration to your MCP client. In Claude Desktop, open Settings → Developer → Edit Config:
{
"mcpServers": {
"visionstory": {
"command": "uvx",
"args": ["visionstory-mcp"],
"env": { "VISIONSTORY_API_KEY": "YOUR_LOCAL_VISIONSTORY_API_KEY" }
}
}
}
Replace the API-key placeholder only in your local client configuration, then restart the client. Ask it to call get_credits and list_models; the connection is ready when both succeed.
Alternative — Install with pip
Use this only if you prefer a permanently installed visionstory-mcp command:
python3 -m pip install --upgrade visionstory-mcp
In the configuration above, replace the command and arguments with:
"command": "visionstory-mcp",
"args": []
Advanced — Connect to a remote MCP endpoint
The local uvx visionstory-mcp setup above remains the simplest option and continues to support files on your computer. Teams that want one centrally managed MCP service can additionally deploy the repository's stateless Streamable HTTP service and connect to its /mcp endpoint.
{
"mcpServers": {
"visionstory-remote": {
"url": "https://YOUR_MCP_HOST/mcp",
"headers": {
"Authorization": "Bearer YOUR_LOCAL_VISIONSTORY_API_KEY"
}
}
}
}
This option is for teams operating a shared MCP service. Local files must be uploaded or available through HTTP(S); choose uvx when the agent needs direct access to files on your computer.
For every setup, keep the API key local and never paste it into chat or commit it to Git. Generation tools wait for results by default, and deletion tools require explicit confirmation.
The local MCP server checks for a newer package when it starts and exposes visionstory://release with the installed version, latest version, feature highlights, release notes, and upgrade command. Update checks are cached and fail silently when offline. Set VISIONSTORY_UPDATE_CHECK=0 to disable them. The remote MCP service is managed centrally and exposes the same resource, but it updates automatically; reconnect the MCP client to refresh its available tools.
Use the Python SDK from an agent
If your agent writes and runs code, pip install visionstory gives it a typed, zero-dependency client:
from visionstory import VisionStoryClient, build_video_payload
client = VisionStoryClient.from_env() # reads VISIONSTORY_API_KEY
video = client.generate_video(build_video_payload(
avatar_id="YOUR_AVATAR_ID", text="Hello from VisionStory.", voice_id="YOUR_VOICE_ID"))
Select both IDs from client.list_avatars() and client.list_voices() immediately before building a request. Values shown in examples are placeholders, not guaranteed resources.
Prefer the terminal? pip install visionstory-cli gives a visionstory command-line tool, so an agent can run visionstory create-video ... without writing code. See the Python SDK guide for the full client.
Install the VisionStory Agent Skill
For skill-compatible coding agents (Claude Code, Codex, …), install the skill straight from GitHub. The skills CLI auto-detects your agent and drops the skill into .agents/skills/:
npx skills add visionstory-ai/skills --skill visionstory-api
The installer may display Snyk warning W011 because the text you ask the avatar to speak is sent to POST /api/v1/video. That network transfer is the purpose of the skill, not an undeclared dependency; review the generated security-audit link and the public skill source before accepting it.
No Node? Clone the repo and copy the skill folder in yourself:
git clone https://github.com/visionstory-ai/skills.git
mkdir -p .agents/skills
cp -r skills/visionstory-api .agents/skills/
The package contains SKILL.md (workflow instructions the agent follows) and scripts/visionstory_api.py, a zero-dependency Python helper for resource discovery, talking-avatar generation, localized text to speech, audio transcription, word alignment, polling, and downloads. Use MCP, the standalone CLI, or the Python SDK when the agent needs the complete set of 27 public operations, including AI Video, image generation, and reusable assets. Skills-compatible agents pick it up automatically; you can also view the developer-site SKILL.md without downloading, or use the public distribution mirror at github.com/visionstory-ai/skills. The helper works standalone too:
python3 .agents/skills/visionstory-api/scripts/visionstory_api.py models
python3 .agents/skills/visionstory-api/scripts/visionstory_api.py create-video --avatar-id YOUR_AVATAR_ID --text "Hello from VisionStory." --voice-id YOUR_VOICE_ID --output result.mp4
Give agents llms.txt and OpenAPI
Point any web-capable agent at the docs index; it links every guide in agent-readable form:
Read https://developers.visionstory.ai/llms.txt for the VisionStory API docs index.
Append .md to any guide URL to get that guide as Markdown.
Load the one guide that matches the task, then use OpenAPI for exact fields.
Use https://developers.visionstory.ai/llms-full.txt only when the task spans several capabilities.
Use https://developers.visionstory.ai/openapi.json for the complete machine-readable API contract.
You can also hand a single .md guide URL or the OpenAPI JSON URL to an agent directly.
Generate a video with one agent instruction
With the Skill installed (or the MCP server connected), a single instruction produces a video:
Create a talking-avatar video that says "Welcome to our launch week!", pick a friendly public avatar and an energetic voice, wait for it to finish, and save it as launch.mp4.
The agent will list avatars and voices, submit POST /api/v1/video, poll GET /api/v1/video?video_id=... until created, and download the video_url — the same flow as the Quick start.
Security and reliability rules for agents
- Read the key from
VISIONSTORY_API_KEY; never print, log, or commit it. - If the key is missing, explain how the user can set it locally; never ask them to paste it into chat.
- Discover resources (
GET /api/v1/models,/api/v1/avatars,/api/v1/voices) instead of guessing IDs. - Treat every ID shown in documentation as a placeholder until it has been returned by a discovery or creation request.
- Pass
client_request_idwhen the create operation supports it so network retries cannot create a second charged task. - Poll no faster than every 5 seconds; stop after 10 minutes unless asked to keep waiting.
- Handle the
failedstatus explicitly and report the server's error message (with credentials redacted). - Never call a
DELETEoperation without an explicit user request and confirmation immediately before the call.
Related VisionStory developer guides
- Quick start — the underlying five-step flow.
- AI Video — text-to-video and image-to-video for agents using an active Pro API key.
- Audio Transcription and Alignment — speech-to-text, SRT, speaker labels, and word timing.
- API reference — every endpoint the channels are built on.
- Changelog — documentation version, API changes, and deprecation status.
Structured media understanding
Use understand_media(prompt, inputs, schema) for structured extraction from 1–8 images, audio clips, or videos. The prompt must contain 1–5000 characters and the JSON Schema must describe a top-level object. The synchronous call allows 180 seconds and returns output, token usage, and cost_credit; there is no model selector or free-text mode. Successful calls are charged, so do not automatically repeat timeouts.
Remote MCP accepts media URLs or asset IDs; local MCP additionally accepts inline_data. The Agent Skill helper exposes understand-media --prompt ... --inputs ... --schema .... Both MCP variants also accept optional speech_rate (slow, normal, fast) on create_speech; omission or null uses normal speed.
See Media Understanding and Text to Speech for complete contracts.