Turns
Turn-level speech-to-text over WebSocket. The server emits events at turn boundaries. See the Turns guide for concepts.
wss://api.reson8.dev/v1/speech-to-text/turnsRequest
Section titled “Request”Headers
Section titled “Headers”| Header | Value |
|---|---|
| Authorization | ApiKey <api_key> or Bearer <access_token> |
| Sec-WebSocket-Protocol | bearer, <access_token> |
See Authentication for which header to use in different situations.
Query Parameters
Section titled “Query Parameters”| Parameter | Type | Default | Description |
|---|---|---|---|
encoding |
string | auto |
Audio encoding: auto for detected container formats, pcm_s16le for raw PCM, or mulaw / alaw for raw G.711 telephony audio. See Audio Formats |
sample_rate |
number | 16000 |
Sample rate in Hz (only used depending on encoding) |
channels |
number | 1 |
Number of audio channels, 1-10 (only used depending on encoding) |
language |
string | Language to transcribe. Recommended for best quality. A comma-separated list (e.g. nl,en) constrains per-utterance auto-detection to those candidates. When omitted, the server auto-detects each utterance independently. See Languages for supported codes |
|
phrases |
string | Comma-separated phrases to bias transcription toward, up to 250. See Custom Models | |
custom_model_id |
string | ID of a custom model to bias transcription. Overrides the model configured on the API client | |
bias_strength |
number | 0.45 |
Strength of contextual biasing. Must be a non-negative number |
include_timestamps |
boolean | false |
Include start_ms and duration_ms on turn_end_candidate |
include_words |
boolean | false |
Include word-level detail on turn_end_candidate |
include_language |
boolean | false |
Include detected language on turn_end_candidate |
include_confidence |
boolean | false |
Include confidence on words |
eager_turn_probability |
number | 0.5 |
Emit an interim turn_end_candidate when the combined turn-end probability exceeds this threshold. Must be between 0 and 1 |
final_turn_probability |
number | 0.92 |
Emit the final turn_end_candidate and turn_end when the combined turn-end probability exceeds this threshold. Must be between 0 and 1 |
patterns |
string | Comma-separated regex-style patterns for short alphanumeric tokens (order codes, licence plates) to recover. Only set when the token is likely present; cannot be combined with phrases or a custom model - see Patterns |
|
filler_mode |
string | natural |
Controls filler words in transcripts: clean removes them, natural lets the model decide, and verbatim preserves them |
Example
Section titled “Example”import asyncioimport jsonimport websockets
async def transcribe(): url = "wss://api.reson8.dev/v1/speech-to-text/turns" headers = {"Authorization": "ApiKey <your_api_key>"}
async with websockets.connect(url, additional_headers=headers) as ws: async def send_audio(): try: with open("recording.wav", "rb") as f: while chunk := f.read(8192): await ws.send(chunk) await ws.send(json.dumps({"type": "flush_request"})) except Exception: await ws.close() raise
sender = asyncio.create_task(send_audio())
async for message in ws: event = json.loads(message) if event["type"] == "turn_start": print("-- turn start") elif event["type"] == "turn_end_candidate": print(event["text"]) elif event["type"] == "turn_end": print("-- turn end") if sender.done(): break
await sender
asyncio.run(transcribe())const token = "<your_access_token>";const url = "wss://api.reson8.dev/v1/speech-to-text/turns";
// Passes token via Sec-WebSocket-Protocol headerconst ws = new WebSocket(url, ["bearer", token]);
ws.onopen = async () => { const audio = await (await fetch("/recording.wav")).arrayBuffer(); for (let i = 0; i < audio.byteLength; i += 8192) { ws.send(audio.slice(i, i + 8192)); } ws.send(JSON.stringify({ type: "flush_request" }));};
ws.onmessage = (event) => { const message = JSON.parse(event.data); if (message.type === "turn_start") console.log("-- turn start"); else if (message.type === "turn_end_candidate") console.log(message.text); else if (message.type === "turn_end") ws.close();};Sending Messages
Section titled “Sending Messages”Binary WebSocket frame containing audio data.
Flush Request
Section titled “Flush Request”Force the current turn to finish immediately rather than waiting for turn-end detection. When a turn is active, the server emits a final turn_end_candidate followed by turn_end.
{ "type": "flush_request"}The request is a JSON text WebSocket frame and has no additional fields.
Receiving Messages
Section titled “Receiving Messages”Turn Start
Section titled “Turn Start”Sent when the server detects the beginning of a new turn.
{ "type": "turn_start"}Turn End Candidate
Section titled “Turn End Candidate”Sent when the server detects that a turn may be ending. Candidates can be emitted before the final Turn End; the final candidate contains the complete transcript.
{ "type": "turn_end_candidate", "text": "the patient presented with chest pain"}{ "type": "turn_end_candidate", "text": "the patient presented with chest pain", "language": "en", "start_ms": 1200, "duration_ms": 2400, "words": [ { "text": "the", "start_ms": 1200, "duration_ms": 200, "confidence": 0.990 }, { "text": "patient", "start_ms": 1410, "duration_ms": 450, "confidence": 0.980 }, { "text": "presented", "start_ms": 1880, "duration_ms": 500, "confidence": 0.970 }, { "text": "with", "start_ms": 2400, "duration_ms": 200, "confidence": 0.990 }, { "text": "chest", "start_ms": 2620, "duration_ms": 350, "confidence": 0.960 }, { "text": "pain", "start_ms": 3000, "duration_ms": 600, "confidence": 0.970 } ]}| Field | Type | Included | Description |
|---|---|---|---|
text |
string | Always | The recognized text |
language |
string | When include_language=true |
The detected language code. Empty string when no language has been detected yet |
start_ms |
number | When include_timestamps=true |
Start time in milliseconds |
duration_ms |
number | When include_timestamps=true |
Duration in milliseconds |
words |
array | When include_words=true |
Word-level detail |
Each word contains:
| Field | Type | Included | Description |
|---|---|---|---|
text |
string | Always | The recognized word |
start_ms |
number | When include_timestamps=true |
Start time in milliseconds |
duration_ms |
number | When include_timestamps=true |
Duration in milliseconds |
confidence |
number | When include_confidence=true |
Probability in (0, 1] |
Turn End
Section titled “Turn End”Confirms that the previous turn end candidate is the final end for the turn.
{ "type": "turn_end"}Errors
Section titled “Errors”Rejected WebSocket upgrades return no body; where available, the reason is in the X-Error-Message response header.
| Status | Description |
|---|---|
| 400 Bad Request | Invalid query parameter, unknown custom_model_id, or patterns combined with a custom model |
| 401 Unauthorized | Missing or invalid credentials |
| 402 Payment Required | Credit limit exceeded - see Limits |
| 429 Too Many Requests | Concurrent connection limit exceeded - see Limits |
Invalid Client Messages
Section titled “Invalid Client Messages”Malformed or unrecognized client messages are ignored; the session continues.
Close Codes
Section titled “Close Codes”After the connection is established, errors are signalled through the WebSocket close code and reason (in browsers: event.code and event.reason on the close event). The reason is a stable machine-readable token: branch on the close code for coarse handling and on the reason for specifics.
| Close code | Reason | Description |
|---|---|---|
| 1000 | Normal closure | |
| 1011 | internal_error |
Unexpected server failure. Reconnect and resend unconfirmed audio |
| 1011 | internal_timeout |
An internal service did not recover within its reconnect budget (up to ~5 seconds). Reconnect and resend unconfirmed audio; back off if it recurs |
| 4000 | message_backlog |
Client did not read messages for 10 seconds. Consume promptly, then reconnect |
A connection that drops without a close frame (surfaced as code 1006 in browsers) indicates a network failure or server crash. Treat it like 1011: reconnect and resend unconfirmed audio.