Speech Studio
Type Lao text and turn it into natural speech — with a built-in voice or your own cloned voice. Runs fully offline.
Output
Recent generations
This session onlyVoice Clone
Record 15–25 seconds of clear Lao speech, and the system will read any text in that voice.
http://127.0.0.1:8000 on this computer, or upload a recording file instead.
- Quiet room, no music or TV
- Phone or mic 20–30 cm from your mouth
- One speaker only, natural pace
- Speak the way you want the voice to sound
API Keys
Connect your website, app or chatbot to SiangMuan. Send the key in the Authorization header.
Create a key
Your keys
Plan & Usage
Your plan, this month's usage and your account.
Characters this month
Characters per day
Last 30 daysYour plan
Profile
Admin- Name
Users
Manage accounts, packages and passwords
Packages
Create packages with custom character (token) quotas and limits
Top-up packages (LaoQR)
Characters users buy once; they never expire and are used after the monthly quota.
Top up characters
Buy extra characters with LaoQR. They never expire and are used after your monthly quota.
Your orders
Payments
LaoQR orders and webhook calls
Webhook calls from the bank
Recorded only — not appliedPronunciation
Teach the voice how to say brand names, people's names or foreign words. Saved permanently.
Saved words 0
API Documentation
Add Lao speech to your website, app or chatbot with a simple HTTP call.
Getting started
The API is plain JSON over HTTP. Every request is authenticated with an API key from your account.
Sign up free — the Free plan includes 20,000 characters a month.
Open API Keys and create a key for your app. Copy it once and keep it secret.
POST /tts with your key in the Authorization header. The response is the audio file.
Authentication
Send your API key with every request, in either header. Requests without a valid key get 401.
Keep keys on your server. Never put an API key in public web pages or mobile apps — anyone could copy it and use your quota. Call SiangMuan from your backend and pass the audio to your users. If a key leaks, revoke it on the API Keys page.
Plans & limits
Limits depend on your plan. Characters are counted from the text you send to /tts and /tts/json; /normalize and the other endpoints are free.
| Plan | Characters / month | Per request | Requests / min | API keys | Voice cloning | Verify |
|---|---|---|---|---|---|---|
| Loading… | ||||||
Each successful response includes X-Quota-Remaining (characters left this month). The monthly quota resets on the 1st (UTC).
Convert Lao text to speech and return the audio file directly. Best for downloading or streaming to a player.
| Field | Type | Description |
|---|---|---|
| textrequired | string | Lao text, up to 5000 characters. Numbers, dates, times, prices and abbreviations are expanded automatically. |
| format | string | wav (default), mp3, ogg or flac. |
| voice | string | Voice name from GET /voices. Default voice if omitted. |
| speed | number | 0.5 – 2.0, default 1.0. Pitch is preserved. |
| seed | integer | null | Default 1234. A fixed seed returns identical audio for identical text (good for caching). null = random. |
| verify | boolean | Check every sentence with a Lao speech recogniser and keep the best take. 2–3× slower, most accurate. Recommended for pre-recorded content. |
| candidates | integer | 1 – 8. Maximum takes per sentence when verify is true (default 3). |
| cfg_value | number | 1.0 – 5.0. Advanced: how closely the model follows the text. Voice default if omitted. |
| inference_timesteps | integer | 4 – 50. Advanced: quality/speed trade-off. Voice default if omitted. |
The audio file with Content-Type audio/wav, audio/mpeg, audio/ogg or audio/flac. Extra headers: X-Audio-Duration (seconds of audio before any speed change) and X-Processing-Time (seconds).
Same request body as /tts, but returns JSON: the audio as base64, the exact text that was spoken, and per-sentence quality details. Useful for chatbots and apps that show subtitles.
With verify: true each chunk also has cer (character error rate, lower is better). A warning field appears if a sentence still sounded wrong after all retries.
Show exactly what will be spoken (numbers, dates and abbreviations written out as Lao words) and how the text is split into sentences. Returns instantly.
Manage pronunciation overrides for words the voice reads wrongly, such as brand or person names. Entries are saved to data/lexicon.json and apply to all future requests. Changing the list (PUT/DELETE) requires an administrator account. Same as the Pronunciation page.
List the available voices (built-in and cloned) and the default one. custom: true marks voices cloned in the web app or via POST /voices.
Create a new voice from a recording of one speaker. Send 15–25 s of clear speech (minimum 6 s) as multipart/form-data. Silences are removed, the level is normalised, and the voice can be used right away with "voice": "<name>". Same as the Voice Clone page.
| Field | Type | Description |
|---|---|---|
| audiorequired | file | The recording: WAV, MP3, M4A, WEBM, OGG or FLAC (max 50 MB). Only the first 30 s of speech are used. |
| namerequired | string | Voice ID used in requests: 2–32 characters, lowercase a–z, 0–9, -, _. |
| label | string | Display name, Lao allowed (e.g. ສົມພອນ). |
| overwrite | boolean | Replace an existing cloned voice with the same ID. Default false. |
Download the cleaned reference recording a voice is cloned from (WAV).
Delete a cloned voice. Built-in voices are protected (403).
Check that the server is up and the model is loaded. Public — no API key needed. Use it for monitoring or a readiness probe.
Errors
Errors return a JSON body with a detail field explaining what went wrong.
| Status | Meaning |
|---|---|
| 200 | Success. |
| 401 | Missing, invalid or revoked API key. |
| 403 | Not allowed: a feature not in your plan (e.g. verify on Free), too many API keys, a second cloned voice (one per account), deleting a built-in voice, or changing the lexicon without admin rights. |
| 404 | Unknown voice, or lexicon word not found. |
| 409 | A voice with that ID already exists (send overwrite=true to replace a cloned voice). |
| 422 | Invalid request: empty text, text longer than 5000 characters, a value out of range, text with no Lao to speak (e.g. only English), or a voice recording that is silent, unreadable or shorter than 6 s of speech. |
| 413 | Text is longer than your plan allows per request — split it into smaller parts. |
| 429 | Rate limit (requests per minute — wait and retry; see Retry-After) or monthly character quota reached. |
| 500 | Audio encoding failed (check that ffmpeg is installed for MP3/OGG/FLAC or a speed other than 1.0). |
Integration guide
Tips for using the API from other apps and devices.
By default the server only listens on this computer. To allow other devices on your network:
Browsers block requests from other sites unless the server allows them. List the allowed sites (comma-separated), or * for any:
- Timeouts: generation takes about 1–2 s of compute per second of audio (more with
verify). Set your client timeout to at least 120–300 s for long texts. - Caching: with a fixed
seed, the same text always gives the same audio. Cache results by text to answer repeat requests instantly. - One at a time: the server generates one request at a time; others wait in line. For many users, pre-generate common phrases.
- Long content: split articles into paragraphs and call the API per paragraph, so playback can start sooner.
- Wrong pronunciation? Add the word to the lexicon instead of changing your text.
More references
Generated automatically from the server code, so they are always up to date.