# Turns Turn-level speech-to-text over WebSocket. The server emits events at turn boundaries. See the [Turns guide](/speech-to-text/turns/) for concepts. ``` wss://api.reson8.dev/v1/speech-to-text/turns ``` ## Request ### Headers | Header | Value | |------------------------|-----------------------------------------------| | Authorization | `ApiKey ` or `Bearer ` | | Sec-WebSocket-Protocol | `bearer, ` | See [Authentication](/authentication/) for which header to use in different situations. ### Query Parameters | Parameter | Type | Default | Description | |--------------------------|---------|-----------|--------------------------------------------------------------| | `encoding` | string | `auto` | Audio encoding: `auto` for detected container formats, `pcm_s16le` for raw PCM, or `mulaw` / `alaw` for raw G.711 telephony audio. See [Audio Formats](/speech-to-text/features/audio-formats/) | | `sample_rate` | number | `16000` | Sample rate in Hz (only used depending on encoding) | | `channels` | number | `1` | Number of audio channels, 1-10 (only used depending on encoding) | | `language` | string | | Language to transcribe. Recommended for best quality. A comma-separated list (e.g. `nl,en`) constrains per-utterance auto-detection to those candidates. When omitted, the server auto-detects each utterance independently. See [Languages](/speech-to-text/features/languages/) for supported codes | | `phrases` | string | | Comma-separated phrases to bias transcription toward, up to 250. See [Custom Models](/speech-to-text/features/custom-models/) | | `custom_model_id` | string | | ID of a [custom model](/api/custom-model/create/) to bias transcription. Overrides the model configured on the API client | | `bias_strength` | number | `0.45` | Strength of contextual biasing. Must be a non-negative number | | `include_timestamps` | boolean | `false` | Include `start_ms` and `duration_ms` on `turn_end_candidate` | | `include_words` | boolean | `false` | Include word-level detail on `turn_end_candidate` | | `include_language` | boolean | `false` | Include detected `language` on `turn_end_candidate` | | `include_confidence` | boolean | `false` | Include `confidence` on words | | `eager_turn_probability` | number | `0.5` | Emit an interim `turn_end_candidate` when the combined turn-end probability exceeds this threshold. Must be between `0` and `1` | | `final_turn_probability` | number | `0.92` | Emit the final `turn_end_candidate` and `turn_end` when the combined turn-end probability exceeds this threshold. Must be between `0` and `1` | | `patterns` | string | | Comma-separated regex-style patterns for short alphanumeric tokens (order codes, licence plates) to recover. Only set when the token is likely present; cannot be combined with `phrases` or a custom model - see [Patterns](/speech-to-text/features/patterns/) | | `filler_mode` | string | `natural` | Controls filler words in transcripts: `clean` removes them, `natural` lets the model decide, and `verbatim` preserves them | ### Example ```python import asyncio import json import websockets async def transcribe(): url = "wss://api.reson8.dev/v1/speech-to-text/turns" headers = {"Authorization": "ApiKey "} async with websockets.connect(url, additional_headers=headers) as ws: async def send_audio(): try: with open("recording.wav", "rb") as f: while chunk := f.read(8192): await ws.send(chunk) await ws.send(json.dumps({"type": "flush_request"})) except Exception: await ws.close() raise sender = asyncio.create_task(send_audio()) async for message in ws: event = json.loads(message) if event["type"] == "turn_start": print("-- turn start") elif event["type"] == "turn_end_candidate": print(event["text"]) elif event["type"] == "turn_end": print("-- turn end") if sender.done(): break await sender asyncio.run(transcribe()) ``` ```javascript const token = ""; const url = "wss://api.reson8.dev/v1/speech-to-text/turns"; // Passes token via Sec-WebSocket-Protocol header const ws = new WebSocket(url, ["bearer", token]); ws.onopen = async () => { const audio = await (await fetch("/recording.wav")).arrayBuffer(); for (let i = 0; i < audio.byteLength; i += 8192) { ws.send(audio.slice(i, i + 8192)); } ws.send(JSON.stringify({ type: "flush_request" })); }; ws.onmessage = (event) => { const message = JSON.parse(event.data); if (message.type === "turn_start") console.log("-- turn start"); else if (message.type === "turn_end_candidate") console.log(message.text); else if (message.type === "turn_end") ws.close(); }; ``` ## Sending Messages ### Audio Binary WebSocket frame containing audio data. ### Flush Request Force the current turn to finish immediately rather than waiting for turn-end detection. When a turn is active, the server emits a final `turn_end_candidate` followed by `turn_end`. ```json { "type": "flush_request" } ``` The request is a JSON text WebSocket frame and has no additional fields. ## Receiving Messages ### Turn Start Sent when the server detects the beginning of a new turn. ```json { "type": "turn_start" } ``` ### Turn End Candidate Sent when the server detects that a turn may be ending. Candidates can be emitted before the final `Turn End`; the final candidate contains the complete transcript. ```json { "type": "turn_end_candidate", "text": "the patient presented with chest pain" } ``` ```json { "type": "turn_end_candidate", "text": "the patient presented with chest pain", "language": "en", "start_ms": 1200, "duration_ms": 2400, "words": [ { "text": "the", "start_ms": 1200, "duration_ms": 200, "confidence": 0.990 }, { "text": "patient", "start_ms": 1410, "duration_ms": 450, "confidence": 0.980 }, { "text": "presented", "start_ms": 1880, "duration_ms": 500, "confidence": 0.970 }, { "text": "with", "start_ms": 2400, "duration_ms": 200, "confidence": 0.990 }, { "text": "chest", "start_ms": 2620, "duration_ms": 350, "confidence": 0.960 }, { "text": "pain", "start_ms": 3000, "duration_ms": 600, "confidence": 0.970 } ] } ``` | Field | Type | Included | Description | |---------------|--------|--------------------------------|----------------------------| | `text` | string | Always | The recognized text | | `language` | string | When `include_language=true` | The detected language code. Empty string when no language has been detected yet | | `start_ms` | number | When `include_timestamps=true` | Start time in milliseconds | | `duration_ms` | number | When `include_timestamps=true` | Duration in milliseconds | | `words` | array | When `include_words=true` | Word-level detail | Each word contains: | Field | Type | Included | Description | |---------------|--------|--------------------------------|----------------------------| | `text` | string | Always | The recognized word | | `start_ms` | number | When `include_timestamps=true` | Start time in milliseconds | | `duration_ms` | number | When `include_timestamps=true` | Duration in milliseconds | | `confidence` | number | When `include_confidence=true` | Probability in `(0, 1]` | ### Turn End Confirms that the previous turn end candidate is the final end for the turn. ```json { "type": "turn_end" } ``` ## Errors Rejected WebSocket upgrades return no body; where available, the reason is in the `X-Error-Message` response header. | Status | Description | |-----------------------|--------------------------------------------------------------| | 400 Bad Request | Invalid query parameter, unknown `custom_model_id`, or `patterns` combined with a custom model | | 401 Unauthorized | Missing or invalid credentials | | 402 Payment Required | Credit limit exceeded - see [Limits](/limits/) | | 429 Too Many Requests | Concurrent connection limit exceeded - see [Limits](/limits/) | ### Invalid Client Messages Malformed or unrecognized client messages are ignored; the session continues. ### Close Codes After the connection is established, errors are signalled through the WebSocket close code and reason (in browsers: `event.code` and `event.reason` on the `close` event). The reason is a stable machine-readable token: branch on the close code for coarse handling and on the reason for specifics. | Close code | Reason | Description | |------------|--------------------|-------------------------------------------------------------------------| | 1000 | | Normal closure | | 1011 | `internal_error` | Unexpected server failure. Reconnect and resend unconfirmed audio | | 1011 | `internal_timeout` | An internal service did not recover within its reconnect budget (up to ~5 seconds). Reconnect and resend unconfirmed audio; back off if it recurs | | 4000 | `message_backlog` | Client did not read messages for 10 seconds. Consume promptly, then reconnect | A connection that drops without a close frame (surfaced as code 1006 in browsers) indicates a network failure or server crash. Treat it like 1011: reconnect and resend unconfirmed audio.