Skip to main content

Dialog engine

Sellable product

This capability is granted by an API key scoped to any of the ai-suite, vagary-voice products (product face). See the product reference below.

Conversation orchestration as a shared capability — turn management, intent detection, tool-calling, safety-gating, barge-in/cancel, streaming assembly. It CONSUMES other capabilities (provider-gateway for the LLM, retrieval for RAG context, stt/tts for the audio legs) and owns only the dialogue control flow. D-2: retrieval is a SEPARATE capability (C9) — dialog-core calls it, never embeds it.

  • Group: Voice & AI
  • Contract: contracts/dialog-core/v1/openapi.yaml
  • Console: this capability is granted by more than one product — manage whichever one your key is scoped to:
  • Runbook: operational checklist for dialog-core — internal (Vagary Labs ops; not part of this public site): docs/runbooks/capability-operations.md#dialog-core
  • Public base: https://api.vagarylabs.com (the consolidated API gateway — one host, per-brand sibling api.<zone>)
  • Auth: a product API key (vgk_…) issued from the console — Authorization: Bearer vgk_…
  • Product face (customer-keyed):
    • https://api.vagarylabs.com/product/v1/dialog/turns
  • Capability face (internal first-party — NOT customer-keyed):
    • https://api.vagarylabs.com/v1/turn

Edge route table​

The routes this capability actually serves on the public edge — live-synced from GET https://api.vagarylabs.com/product/v1/_meta/catalog (snapshot v1). The metric is the request-granularity billing counter; the scope is the key permission required.

Product face (customer-keyed)​

MethodPathMetricScope
POST/product/v1/dialog/turnsdialog_turnsread

Capability face (internal first-party — NOT customer-keyed)​

  • /v1/turn

Endpoints​

MethodPathSummary
POST/v1/turnProcess one conversational turn (intent → optional tools + retrieval → LLM → safety → stream)
POST/v1/turn/&#123;session_id&#125;/cancelbarge-in — cancel the in-flight turn
POST/v1/greeting/prerenderProactively render + cache a generation's opening line, so the FIRST call on it is fast
GET/healthliveness
GET/context-healthwhich context legs are enabled/disabled (default-deployment posture is otherwise silent)
GET/metricsPrometheus

Product face​

The external, paying-customer surface served by the edge at /product/v1/* (contracts/dialog-core/product/v1/openapi.yaml).

MethodPathSummary
POST/turnsProcess one conversational turn (managed, metered) — wraps the capability's orchestration

Schemas​

TurnRequest​

FieldTypeDescription
organization_idstring
session_idstring
inputstring
streamboolean
modelstringLLM model hint forwarded to provider-gateway (which owns provider selection)
systemstringitem 6 — per-caller system-prompt override; default is the service preamble
historybooleanitem 1 — enable/disable multi-turn conversation history for this turn (overrides the DIALOG_HISTORY_ENABLED default)
messagesarrayitem 1 — OPTIONAL client-supplied PRIOR conversation window ({role,content} list, EXCLUDING the current input). When present it is used as the history for this turn (durable/cross-replica product-face state), taking precedence over the in-process store. Absent => backward-compatible (in-process store when history is on, else stateless).
retrievalobjectopt-in RAG params — dialog-core calls the retrieval capability (C9), never embeds it
toolsarrayfunction-calling tool definitions passed through to provider-gateway; returned tool_calls are surfaced (guarded) — or EXECUTED when execute_tools=true (T5-3)
execute_toolsbooleanT5-3 — OPT-IN: run the agentic execute-then-continue loop (execute each tool_call, feed the result back, re-call, until a final answer). Requires tools. Absent/false => tool_calls are only surfaced (byte-identical to the pre-loop turn).
tool_dispatchobjectT5-3 — how the executor DISPATCHES each tool by name to the caller's own HTTP endpoint (I5: core hosts no business tool). name -> {url, method?, headers?}. A tool with no entry yields a fail-open tool_not_dispatchable result the model can recover from.

GreetingPrerenderRequest​

FieldTypeDescription
organization_idstring
flow_idstring
versionintegerthe EXACT definition version to warm — never the live pointer
definition_hashstringthe flow content pin; also the cache key component, so changed content is automatically a miss
timeout_secondsnumberhow long to wait for the out-of-band render to land before returning pending (default 8, capped at 30). A timeout is not a failure — the render may still land.

GreetingPrerenderResponse​

FieldTypeDescription
statusstringcached = an asset already existed for this content (no render paid); rendered = a fresh render landed within the wait; pending = the capture was requested but had not landed yet; skipped = this profile carries no extractable opening line, so no asset will ever be produced for it (permanent, reported synchronously rather than as a never-resolving pending); not_permitted = the flow failed the same admission a live call runs; unavailable = no Redis configured; error = malformed request or capture publish failed.
definition_hashstring
reasonstring

TurnResponse​

FieldTypeDescription
session_idstring
contentstring
intentstring
tool_callsarray
tool_tracearrayT5-3 — present only when execute_tools ran: the ordered trace of executed tool_calls (args + result + latency), never silent.
tool_stop_reasonstringT5-3 — why the execute-then-continue loop stopped (present only when it ran)
safetyobject

Error​

FieldTypeDescription
errorstring

ContextHealthResponse​

FieldTypeDescription
retrievalobject
historyobject
memoryobject
fast_pathobject
tool_executorobject
plugin_registryobject
realtimeobject
default_posturestringliteral summary: "[system, user] only" when both history and memory are disabled, else "context-enriched"

LlmTokensEmit​

FieldTypeDescription
typestringchunk = streamed word-buffer content; clear = flush-buffer control (no text); filler/fallback retained for voice parity (dialog-core emits chunk/clear — provider-gateway owns failover); prerendered = greeting pre-render asset reference (2026-09-19, see greeting_prerender.py); greeting_capture = one-shot capture request on the synthetic session:greeting_capture:llm_tokens channel
textstringREQUIRED for chunk/filler/fallback/greeting_capture; OMITTED for clear and prerendered (control/reference messages)
timestampnumberunix float seconds at emit; present on every message
org_idstringADDITIVE tenant org_id threaded from the transcript envelope; absent on clear; nullable for API-key (no-org) sessions; also carries the capture-request org on greeting_capture
consent_refstringADDITIVE D22 consent lineage threaded from the transcript envelope; present on text-bearing kinds, absent on clear
pcm_keystringADDITIVE, prerendered only: Redis key holding the base64 pre-rendered PCM to replay instead of live synthesis
definition_hashstringADDITIVE, greeting_capture only: the published flow content-hash cache key component
voice_idstringADDITIVE, greeting_capture only: informational voice identifier

FlowTraceEmit​

FieldTypeDescription
typestring
session_store_idstringActual session-store-generated ID; distinct from the Voice session UUID.
flow_idstring
flow_versioninteger

Generated by scripts/gen-capability-docs.py from contracts/dialog-core/v1/openapi.yaml — the contract IS the source of truth; edit the contract, not this page. The Edge route table section is synced from the live edge catalog snapshot (v1) — re-sync with python3 scripts/sync-edge-catalog.py.