# Cartesia integration

> Low-latency text-to-speech (Sonic), voice cloning, voice changing, and speech-to-text (Ink) REST API.

- Authentication: apiKey

## Tools (38)

- **Clone Voice**: Create a new voice by cloning from an audio sample.
- **Create Access Token**: Mint a short-lived scoped token for client-side use.
- **Create Dataset**: Create a dataset to hold training audio for fine-tuning.
- **Create Fine-Tune**: Train a custom model from a dataset.
- **Create Pronunciation Dictionary**: Create a dictionary that rewrites text before synthesis.
- **Delete Dataset**: Delete a dataset.
- **Delete Dataset File**: Remove a file from a dataset.
- **Delete Fine-Tune**: Delete a fine-tune.
- **Delete Pronunciation Dictionary**: Delete a pronunciation dictionary.
- **Delete Voice**: Delete a voice you own.
- **Get Agent Usage**: Report usage for the voice-agent platform.
- **Get API Key**: Retrieve one API key record by id (metadata only).
- **Get API Status**: Public health and latest-version check; no auth required.
- **Get Credit Usage**: Report credit consumption and remaining balance.
- **Get Dataset**: Retrieve a single dataset by id.
- **Get Fine-Tune**: Retrieve a fine-tune, including its training status.
- **Get Pronunciation Dictionary**: Retrieve a pronunciation dictionary by id.
- **Get Voice**: Retrieve a single voice by id.
- **Infill**: Generate audio that fills between a left and right clip so the join sounds natural.
- **List API Keys**: List the API keys on your account (metadata only).
- **List Dataset Files**: List the files in a dataset.
- **List Datasets**: List datasets in your organization.
- **List Fine-Tunes**: List fine-tunes in your organization.
- **List Fine-Tune Voices**: List the voices produced by a completed fine-tune.
- **List Pronunciation Dictionaries**: List pronunciation dictionaries in your organization.
- **List Voices**: List voices in the library; also the clean credential validator.
- **Localize Voice**: Create a new voice that speaks an existing voice in another language.
- **Speech to Text (WebSocket)**: Real-time streaming transcription over a WebSocket connection.
- **Text to Speech (Bytes)**: Synthesize speech and return the full audio file as bytes.
- **Text to Speech (SSE)**: Stream synthesized speech as Server-Sent Events.
- **Text to Speech (WebSocket)**: Real-time bidirectional TTS over a WebSocket connection.
- **Transcribe Audio (STT)**: Batch speech-to-text with the Ink Whisper model.
- **Update Dataset**: Update a dataset's name or description.
- **Update Pronunciation Dictionary**: Update a pronunciation dictionary's fields or entries.
- **Update Voice**: Update an existing voice's editable fields.
- **Upload Dataset File**: Add an audio file to a dataset.
- **Voice Changer (Bytes)**: Convert the speaker in an input clip to a target voice, returning full audio.
- **Voice Changer (SSE)**: Streaming voice changer returning SSE audio chunks.

## Related prompts

- [Weekly wins voice recap in Slack from HubSpot deals](https://www.generalinput.com/prompts/weekly-wins-voice-recap-in-slack-from-hubspot-deals.md)
- [Turn Notion drafts into narrated blog audio and podcast episodes](https://www.generalinput.com/prompts/turn-notion-drafts-into-narrated-blog-audio-and-podcast-episodes.md)
- [Nightly voice memo transcription in a Google Drive folder](https://www.generalinput.com/prompts/nightly-voice-memo-transcription-in-a-google-drive-folder.md)
- [Add multilingual audio to every new Webflow blog post](https://www.generalinput.com/prompts/add-multilingual-audio-to-every-new-webflow-blog-post.md)

Connect Cartesia in General Input: https://www.generalinput.com/apps/cartesia