Skip to main content

Provider gateway

scripts/gen-capability-docs.py. Edit it here; editing the generated page directly is writing to a cache — it works, it reads correctly, and the next regeneration erases it. That is exactly what happened: 355 lines and 20 code blocks were hand-added straight to

Quick start​

Prerequisites​

  1. An API key from the console (vgk_…)
  2. The base URL: https://api.vagarylabs.com

First chat call (curl)​

curl -X POST https://api.vagarylabs.com/v1/llm/chat \
-H "Content-Type: application/json" \
-H "Authorization: Bearer vgk_YOUR_API_KEY" \
-d '{
"messages": [
{"role": "user", "content": "Hello, world!"}
],
"model": "gpt-4o"
}'

Response:

{
"content": "Hello! How can I help you today?",
"usage": {
"prompt_tokens": 10,
"completion_tokens": 8,
"provider_used": "openai",
"cost_cents": 0.00032
}
}

TypeScript (using the generated client)​

import { Client } from "@vagary/shared-sdks";

const client = new Client({
baseUrl: "https://api.vagarylabs.com",
headers: { Authorization: "Bearer vgk_YOUR_API_KEY" },
});

const response = await client.default.llmChat({
messages: [{ role: "user", content: "Hello, world!" }],
model: "gpt-4o",
});

console.log(response.content);
console.log(`Cost: $${response.usage.cost_cents / 100}`);

Python (using the generated client)​

from clients.provider_gateway_client import Client

client = Client(
base_url="https://api.vagarylabs.com",
headers={"Authorization": "Bearer vgk_YOUR_API_KEY"},
)

response = client.default.llm_chat(
messages=[{"role": "user", "content": "Hello, world!"}],
model="gpt-4o",
)

print(response.content)
print(f"Cost: ${response.usage.cost_cents / 100}")

Streaming response​

curl -X POST https://api.vagarylabs.com/v1/llm/chat \
-H "Content-Type: application/json" \
-H "Authorization: Bearer vgk_YOUR_API_KEY" \
-d '{
"messages": [
{"role": "user", "content": "Write a haiku about code"}
],
"model": "gpt-4o",
"stream": true
}'

Embeddings​

curl -X POST https://api.vagarylabs.com/v1/embeddings \
-H "Content-Type: application/json" \
-H "Authorization: Bearer vgk_YOUR_API_KEY" \
-d '{
"input": "The quick brown fox jumps over the lazy dog",
"model": "text-embedding-3-small"
}'

Error handling​

All error responses include an error_class field for programmatic branching:

{
"error": "rate_limited",
"error_class": "rate_limit",
"retryable": true,
"retry_after": 30,
"reason": "rate limit exceeded"
}
error_classretryableClient action
transient✅Retry with backoff
rate_limit✅Fail over to next provider
server_error✅Retry or fail over
auth❌Fix your API key
policy❌Fix your request
context_overflow❌Trim message history
content_policy❌Change your prompt
  • Group: Auth & gateway
  • Contract: contracts/provider-gateway/v1/openapi.yaml
  • Console: Manage the ai-suite product in the console →
  • Runbook: operational checklist for provider-gateway — internal (Vagary Labs ops; not part of this public site): docs/runbooks/capability-operations.md#provider-gateway
  • Public base: https://api.vagarylabs.com (the consolidated API gateway — one host, per-brand sibling api.<zone>)
  • Auth: a product API key (vgk_…) issued from the console — Authorization: Bearer vgk_…
  • Product face (customer-keyed):
    • https://api.vagarylabs.com/product/v1/ai/chat
  • Capability face (internal first-party — NOT customer-keyed):
    • https://api.vagarylabs.com/v1/embeddings
    • https://api.vagarylabs.com/v1/llm
Sellable product

This capability is granted by an API key scoped to any of the ai-suite, vagaris, vagary-voice products (product face). See the product reference below.

A unified egress to external AI providers (LLM + embeddings; STT/TTS have their own capabilities) — key custody (BYOK + fleet keys), provider-selection (cost/latency/quality), failover, per-provider circuit-breaker + bulkhead, and per-request usage-metering. Products stop re-implementing provider plumbing. Keys resolved server-side (Infisical/BYOK vault), NEVER in body/DB/logs. Every response carries usage (prompt/completion tokens + provider + cost_cents) — the billed feed consumed by billing-metering (C5) and reconciled into analytics (D4).

  • Group: Auth & gateway
  • Contract: contracts/provider-gateway/v1/openapi.yaml
  • Console: this capability is granted by more than one product — manage whichever one your key is scoped to:
  • Runbook: operational checklist for provider-gateway — internal (Vagary Labs ops; not part of this public site): docs/runbooks/capability-operations.md#provider-gateway
  • Public base: https://api.vagarylabs.com (the consolidated API gateway — one host, per-brand sibling api.<zone>)
  • Auth: a product API key (vgk_…) issued from the console — Authorization: Bearer vgk_…
  • Product face (customer-keyed):
    • https://api.vagarylabs.com/product/v1/ai/chat
  • Capability face (internal first-party — NOT customer-keyed):
    • https://api.vagarylabs.com/v1/embeddings
    • https://api.vagarylabs.com/v1/llm

Edge route table​

The routes this capability actually serves on the public edge — live-synced from GET https://api.vagarylabs.com/product/v1/_meta/catalog (snapshot v1). The metric is the request-granularity billing counter; the scope is the key permission required.

Product face (customer-keyed)​

MethodPathMetricScope
POST/product/v1/ai/chatllm_requestsread

Capability face (internal first-party — NOT customer-keyed)​

  • /v1/embeddings
  • /v1/llm

Endpoints​

MethodPathSummary
POST/v1/llm/chatChat/completion via the selected provider (streaming or sync), with failover
POST/v1/embeddingsEmbeddings via the selected provider, with failover + usage
POST/v1/visionVision completion (image + prompt) via the selected vision-capable provider, with failover + usage
GET/v1/providersprovider health + selection state
GET/v1/byok/configList the org's active provider configs — METADATA + masked hint only, NEVER the raw secret
POST/v1/byok/configCreate/replace (upsert) the org's config for a provider
GET/v1/byok/config/&#123;provider&#125;One config's METADATA (masked). 404 if the org has no active config for provider
DELETE/v1/byok/config/&#123;provider&#125;Soft-delete (deactivate) the org's config for provider — reversible via POST /v1/byok/config
POST/v1/byok/config/&#123;provider&#125;/rotateRotate the credential on an EXISTING active config. 404 if none exists
POST/v1/byok/purgeERASE every stored provider credential the org holds — the BYOK leg of an org-delete cascade. IDEMPOTENT: success is asserted on configs_remaining == 0, so a re-run erasing 0 is still success.
POST/v1/images/generationsImage generation through the gateway's provider routing (OpenAI-shaped request)
GET/healthliveness
GET/v1/availabilitySPOF documentation and dependency status
GET/metricsPrometheus
POST/v1/prompt-cache/purgePurge cached prompt/completion content for an org (GDPR Art-17 DSR propagation)
GET/v1/byok/config/&#123;provider&#125;/lifecycleCredential lifecycle metadata — expiry, staleness, rotation schedule, owner principal
POST/v1/byok/config/&#123;provider&#125;/transferTransfer credential ownership to a different identity principal
GET/v1/byok/staleFleet-wide list of stale active BYOK credentials
POST/v1/approvalsCreate a pending approval record for an MCP tool execution
GET/v1/approvals/&#123;ref&#125;Get the status of an approval record
POST/v1/approvals/&#123;ref&#125;/approveApprove a pending MCP tool execution
POST/v1/approvals/&#123;ref&#125;/denyDeny a pending MCP tool execution
POST/v1/mcp/trustRecord a server trust assignment
GET/v1/mcp/trust/&#123;org_id&#125;List all trust records for an org
POST/v1/mcp/trust/&#123;org_id&#125;/&#123;server_id&#125;/reviewReview a trust record (required before VERIFIED or FIRST_PARTY takes effect)
GET/v1/replay/&#123;request_id&#125;Get a captured execution's full context packet (for inspection or manual replay)
POST/v1/replay/&#123;request_id&#125;Re-issue a captured request for real, against the CALLER's own tenant credential
GET/v1/tracesList recent execution traces for the caller's own org (content included, PII-redacted)
GET/v1/support/trace/&#123;request_id&#125;Content-redacted, cross-tenant trace lookup for support triage (internal-only)
GET/v1/lineage/&#123;request_id&#125;Content-free provenance lookup for one request (model/routing/prompt-template lineage)
GET/v1/lineageList recent lineage events for the caller's own org (content-free)
GET/health/liveLiveness probe — is the process alive? (K8s/docker restart signal)
GET/health/readyReadiness probe — can the gateway accept traffic?
GET/v1/fleet/statusFleet-wide status, distinguishing gateway-down from vendor-down
GET/v1/canaryRun the golden-set canary against the configured model (CI gate; SPENDS MONEY)
POST/v1/canary/consistencyResponse-consistency sampling — send the same prompt N times and check variance (SPENDS MONEY)
GET/v1/models/livenessProbe each configured default model for liveness (SPENDS MONEY)
GET/v1/models/catalogModel cards for every configured provider (EU AI Act Art.50 + enterprise DPA disclosure; UNAUTHENTICATED)
GET/v1/models/catalog/&#123;provider&#125;Model card for one provider, canonical or alias (UNAUTHENTICATED)
GET/v1/models/registryFull model registry — owner, successor, cutoff, and status for every model (UNAUTHENTICATED)
POST/v1/models/registry/&#123;model_id&#125;/deprecateMark a model as deprecated with a successor and reason
GET/v1/models/deprecation-notificationsModels needing customer notification — deprecated, plus expiring within 90 days (UNAUTHENTICATED)
POST/v1/admin/reload-policyHot-reload the routing policy from model-routing-policy.yaml without a redeploy
GET/v1/admin/routing-configCurrent routing config — canary state, provider ranks, fallback ladder
POST/v1/admin/pinPin routing to a specific provider/model (demo-safe mode; disables failover)
POST/v1/admin/unpinClear the routing pin, restoring normal failover
GET/v1/admin/pin-statusCurrent routing pin status
POST/v1/admin/canarySet percentage-based canary routing
POST/v1/admin/canary-stopStop canary routing — all traffic goes to stable
GET/v1/admin/canary-statusCurrent canary routing status
POST/v1/admin/orgs/&#123;org_id&#125;/suspendSuspend an org — every egress route (llm/chat, embeddings, vision, images) refuses it (403)
POST/v1/admin/orgs/&#123;org_id&#125;/unsuspendUnsuspend an org — restore fleet-key access
POST/v1/admin/sessions/&#123;run_id&#125;/abortRequest a cooperative abort of an in-flight session
GET/v1/admin/sessions/&#123;run_id&#125;/spendReal-time spend for one in-flight session
GET/v1/admin/orgs/&#123;org_id&#125;/spendReal-time aggregated spend for an org across all active sessions, against its budget ceiling
POST/v1/runs/mintMint a new run authority (called by the run's initiator — e.g. dialog-core, flow-builder)
GET/v1/runs/&#123;run_id&#125;/authorityGet the current state of a run authority (debugging / operator visibility)
POST/v1/runs/&#123;run_id&#125;/revokeRevoke a run authority — the runaway-loop kill switch
GET/v1/security/blast-radiusBlast-radius query for a leaked provider key — which orgs/services are currently affected
GET/v1/pricing/driftPricing rows past their review cadence (UNAUTHENTICATED)
POST/v1/pricing/reconcileCompare metered cost (from billing-metering) against vendor invoice amounts

Product face​

The external, paying-customer surface served by the edge at /product/v1/* (contracts/provider-gateway/product/v1/openapi.yaml).

MethodPathSummary
POST/chatChat/completion (managed, metered) — wraps the capability's multi-provider LLM egress

Schemas​

Usage​

FieldTypeDescription
prompt_tokensinteger
completion_tokensinteger
provider_usedstring
cost_centsnumber
estimatedbooleanfalse => provider-reported (real) — see D4
requested_providerstringprovider the caller asked for (absent when no substitution risk)
requested_goalstringoptimization_goal the caller requested
substitutedbooleantrue when a different provider served than was requested
cache_hitbooleantrue when the response was served from cache

LlmRequest​

FieldTypeDescription
organization_idstring
modelstring
messagesarray
streamboolean
toolsarrayfunction-calling tool schemas (OpenAI shape); non-streaming OpenAI surfaces tool_calls. Additive — absent => unchanged completion path
optimization_goalstring
byokobjectbring-your-own-key ref (never a raw key in body)
allow_substitutionbooleanfalse => return 502 with original provider error instead of silent substitution
data_classificationstringdata classification for tenant routing policy
localestringISO 639-1 language code (e.g. 'en', 'es', 'fr') or BCP-47 locale (e.g. 'en-US', 'es-MX'). Controls the language the model responds in. When absent, the model infers language from input. When set, a system instruction is prepended: 'Respond in {locale}.' Two modes: - 'auto' (default): model infers from input text (current implicit behavior) - explicit locale: force response language regardless of input
language_modestringHow to interpret the locale field: - 'auto' (default): locale is used as a hint; model may still use input language - 'force': locale is a hard requirement; model MUST respond in this language

LlmResponse​

FieldTypeDescription
contentstring
tool_callsarrayfunction-calling tool calls the model emitted (OpenAI non-streaming); absent/None when the model returned none
usageobject

EmbeddingsRequest​

FieldTypeDescription
organization_idstring
modelstring
inputobject
optimization_goalstringprovider-selection goal for the embeddings failover chain (openai, gemini)
byokobjectbring-your-own-key ref (never a raw key in body); resolved server-side per (org, provider)

EmbeddingsResponse​

FieldTypeDescription
dataarray
usageobject

ImageRef​

the image to analyze — supply EXACTLY ONE of url or base64 (base64 may include mime_type; the raw bytes are NEVER logged)

FieldTypeDescription
urlstringpublicly-fetchable image url
base64stringbase64-encoded image bytes (inline)
mime_typestringe.g. image/png — used with base64; defaults image/jpeg

VisionRequest​

FieldTypeDescription
organization_idstring
modelstringoverride the vision model (default gpt-4o / gemini-2.0-flash)
promptstringthe task prompt — classification/alt-text/OCR instructions supplied by the caller
imageobject
optimization_goalstringprovider-selection goal for the vision failover chain (openai, gemini)
byokobjectbring-your-own-key ref (never a raw key in body); resolved server-side per (org, provider)

VisionResponse​

FieldTypeDescription
contentstringraw vision completion (parse JSON for classification; use directly for alt-text)
usageobject

Generated by scripts/gen-capability-docs.py from contracts/provider-gateway/v1/openapi.yaml — the contract IS the source of truth; edit the contract, not this page. The Edge route table section is synced from the live edge catalog snapshot (v1) — re-sync with python3 scripts/sync-edge-catalog.py.