Skip to content

Turns

Turn-level speech-to-text over WebSocket. The server emits events at turn boundaries. See the Turns guide for concepts.

wss://api.reson8.dev/v1/speech-to-text/turns
Header Value
Authorization ApiKey <api_key> or Bearer <access_token>
Sec-WebSocket-Protocol bearer, <access_token>

See Authentication for which header to use in different situations.

Parameter Type Default Description
encoding string auto Audio encoding: auto for detected container formats, pcm_s16le for raw PCM, or mulaw / alaw for raw G.711 telephony audio. See Audio Formats
sample_rate number 16000 Sample rate in Hz (only used depending on encoding)
channels number 1 Number of audio channels, 1-10 (only used depending on encoding)
language string Language to transcribe. Recommended for best quality. A comma-separated list (e.g. nl,en) constrains per-utterance auto-detection to those candidates. When omitted, the server auto-detects each utterance independently. See Languages for supported codes
phrases string Comma-separated phrases to bias transcription toward, up to 250. See Custom Models
custom_model_id string ID of a custom model to bias transcription. Overrides the model configured on the API client
bias_strength number 0.45 Strength of contextual biasing. Must be a non-negative number
include_timestamps boolean false Include start_ms and duration_ms on turn_end_candidate
include_words boolean false Include word-level detail on turn_end_candidate
include_language boolean false Include detected language on turn_end_candidate
include_confidence boolean false Include confidence on words
eager_turn_probability number 0.5 Emit an interim turn_end_candidate when the combined turn-end probability exceeds this threshold. Must be between 0 and 1
final_turn_probability number 0.92 Emit the final turn_end_candidate and turn_end when the combined turn-end probability exceeds this threshold. Must be between 0 and 1
patterns string Comma-separated regex-style patterns for short alphanumeric tokens (order codes, licence plates) to recover. Only set when the token is likely present; cannot be combined with phrases or a custom model - see Patterns
filler_mode string natural Controls filler words in transcripts: clean removes them, natural lets the model decide, and verbatim preserves them
import asyncio
import json
import websockets
async def transcribe():
url = "wss://api.reson8.dev/v1/speech-to-text/turns"
headers = {"Authorization": "ApiKey <your_api_key>"}
async with websockets.connect(url, additional_headers=headers) as ws:
async def send_audio():
try:
with open("recording.wav", "rb") as f:
while chunk := f.read(8192):
await ws.send(chunk)
await ws.send(json.dumps({"type": "flush_request"}))
except Exception:
await ws.close()
raise
sender = asyncio.create_task(send_audio())
async for message in ws:
event = json.loads(message)
if event["type"] == "turn_start":
print("-- turn start")
elif event["type"] == "turn_end_candidate":
print(event["text"])
elif event["type"] == "turn_end":
print("-- turn end")
if sender.done():
break
await sender
asyncio.run(transcribe())

Binary WebSocket frame containing audio data.

Force the current turn to finish immediately rather than waiting for turn-end detection. When a turn is active, the server emits a final turn_end_candidate followed by turn_end.

{
"type": "flush_request"
}

The request is a JSON text WebSocket frame and has no additional fields.

Sent when the server detects the beginning of a new turn.

{
"type": "turn_start"
}

Sent when the server detects that a turn may be ending. Candidates can be emitted before the final Turn End; the final candidate contains the complete transcript.

{
"type": "turn_end_candidate",
"text": "the patient presented with chest pain"
}
Field Type Included Description
text string Always The recognized text
language string When include_language=true The detected language code. Empty string when no language has been detected yet
start_ms number When include_timestamps=true Start time in milliseconds
duration_ms number When include_timestamps=true Duration in milliseconds
words array When include_words=true Word-level detail

Each word contains:

Field Type Included Description
text string Always The recognized word
start_ms number When include_timestamps=true Start time in milliseconds
duration_ms number When include_timestamps=true Duration in milliseconds
confidence number When include_confidence=true Probability in (0, 1]

Confirms that the previous turn end candidate is the final end for the turn.

{
"type": "turn_end"
}

Rejected WebSocket upgrades return no body; where available, the reason is in the X-Error-Message response header.

Status Description
400 Bad Request Invalid query parameter, unknown custom_model_id, or patterns combined with a custom model
401 Unauthorized Missing or invalid credentials
402 Payment Required Credit limit exceeded - see Limits
429 Too Many Requests Concurrent connection limit exceeded - see Limits

Malformed or unrecognized client messages are ignored; the session continues.

After the connection is established, errors are signalled through the WebSocket close code and reason (in browsers: event.code and event.reason on the close event). The reason is a stable machine-readable token: branch on the close code for coarse handling and on the reason for specifics.

Close code Reason Description
1000 Normal closure
1011 internal_error Unexpected server failure. Reconnect and resend unconfirmed audio
1011 internal_timeout An internal service did not recover within its reconnect budget (up to ~5 seconds). Reconnect and resend unconfirmed audio; back off if it recurs
4000 message_backlog Client did not read messages for 10 seconds. Consume promptly, then reconnect

A connection that drops without a close frame (surfaced as code 1006 in browsers) indicates a network failure or server crash. Treat it like 1011: reconnect and resend unconfirmed audio.