Provider gateway
scripts/gen-capability-docs.py. Edit it here; editing the generated page directly is writing to a cache — it works, it reads correctly, and the next regeneration erases it. That is exactly what happened: 355 lines and 20 code blocks were hand-added straight to
Quick start
Prerequisites
- An API key from the console (
vgk_…) - The base URL:
https://api.vagarylabs.com
First chat call (curl)
curl -X POST https://api.vagarylabs.com/v1/llm/chat \
-H "Content-Type: application/json" \
-H "Authorization: Bearer vgk_YOUR_API_KEY" \
-d '{
"messages": [
{"role": "user", "content": "Hello, world!"}
],
"model": "gpt-4o"
}'
Response:
{
"content": "Hello! How can I help you today?",
"usage": {
"prompt_tokens": 10,
"completion_tokens": 8,
"provider_used": "openai",
"cost_cents": 0.00032
}
}
TypeScript (using the generated client)
import { Client } from "@vagary/shared-sdks";
const client = new Client({
baseUrl: "https://api.vagarylabs.com",
headers: { Authorization: "Bearer vgk_YOUR_API_KEY" },
});
const response = await client.default.llmChat({
messages: [{ role: "user", content: "Hello, world!" }],
model: "gpt-4o",
});
console.log(response.content);
console.log(`Cost: $${response.usage.cost_cents / 100}`);
Python (using the generated client)
from clients.provider_gateway_client import Client
client = Client(
base_url="https://api.vagarylabs.com",
headers={"Authorization": "Bearer vgk_YOUR_API_KEY"},
)
response = client.default.llm_chat(
messages=[{"role": "user", "content": "Hello, world!"}],
model="gpt-4o",
)
print(response.content)
print(f"Cost: ${response.usage.cost_cents / 100}")
Streaming response
curl -X POST https://api.vagarylabs.com/v1/llm/chat \
-H "Content-Type: application/json" \
-H "Authorization: Bearer vgk_YOUR_API_KEY" \
-d '{
"messages": [
{"role": "user", "content": "Write a haiku about code"}
],
"model": "gpt-4o",
"stream": true
}'
Embeddings
curl -X POST https://api.vagarylabs.com/v1/embeddings \
-H "Content-Type: application/json" \
-H "Authorization: Bearer vgk_YOUR_API_KEY" \
-d '{
"input": "The quick brown fox jumps over the lazy dog",
"model": "text-embedding-3-small"
}'
Error handling
All error responses include an error_class field for programmatic branching:
{
"error": "rate_limited",
"error_class": "rate_limit",
"retryable": true,
"retry_after": 30,
"reason": "rate limit exceeded"
}
error_class | retryable | Client action |
|---|---|---|
transient | ✅ | Retry with backoff |
rate_limit | ✅ | Fail over to next provider |
server_error | ✅ | Retry or fail over |
auth | ❌ | Fix your API key |
policy | ❌ | Fix your request |
context_overflow | ❌ | Trim message history |
content_policy | ❌ | Change your prompt |
- Group: Auth & gateway
- Contract:
contracts/provider-gateway/v1/openapi.yaml - Console: Manage the
ai-suiteproduct in the console → - Runbook: operational checklist for
provider-gateway— internal (Vagary Labs ops; not part of this public site):docs/runbooks/capability-operations.md#provider-gateway - Public base:
https://api.vagarylabs.com(the consolidated API gateway — one host, per-brand siblingapi.<zone>) - Auth: a product API key (
vgk_…) issued from the console —Authorization: Bearer vgk_… - Product face (customer-keyed):
https://api.vagarylabs.com/product/v1/ai/chat
- Capability face (internal first-party — NOT customer-keyed):
https://api.vagarylabs.com/v1/embeddingshttps://api.vagarylabs.com/v1/llm
This capability is granted by an API key scoped to any of the ai-suite, vagaris, vagary-voice products (product face). See the product reference below.
A unified egress to external AI providers (LLM + embeddings; STT/TTS have their own capabilities) — key custody (BYOK + fleet keys), provider-selection (cost/latency/quality), failover, per-provider circuit-breaker + bulkhead, and per-request usage-metering. Products stop re-implementing provider plumbing. Keys resolved server-side (Infisical/BYOK vault), NEVER in body/DB/logs. Every response carries usage (prompt/completion tokens + provider + cost_cents) — the billed feed consumed by billing-metering (C5) and reconciled into analytics (D4).
- Group: Auth & gateway
- Contract:
contracts/provider-gateway/v1/openapi.yaml - Console: this capability is granted by more than one product — manage whichever one your key is scoped to:
- Runbook: operational checklist for
provider-gateway— internal (Vagary Labs ops; not part of this public site):docs/runbooks/capability-operations.md#provider-gateway - Public base:
https://api.vagarylabs.com(the consolidated API gateway — one host, per-brand siblingapi.<zone>) - Auth: a product API key (
vgk_…) issued from the console —Authorization: Bearer vgk_… - Product face (customer-keyed):
https://api.vagarylabs.com/product/v1/ai/chat
- Capability face (internal first-party — NOT customer-keyed):
https://api.vagarylabs.com/v1/embeddingshttps://api.vagarylabs.com/v1/llm
Edge route table
The routes this capability actually serves on the public edge — live-synced from GET https://api.vagarylabs.com/product/v1/_meta/catalog (snapshot v1). The metric is the request-granularity billing counter; the scope is the key permission required.
Product face (customer-keyed)
| Method | Path | Metric | Scope |
|---|---|---|---|
POST | /product/v1/ai/chat | llm_requests | read |
Capability face (internal first-party — NOT customer-keyed)
/v1/embeddings/v1/llm
Endpoints
| Method | Path | Summary |
|---|---|---|
POST | /v1/llm/chat | Chat/completion via the selected provider (streaming or sync), with failover |
POST | /v1/embeddings | Embeddings via the selected provider, with failover + usage |
POST | /v1/vision | Vision completion (image + prompt) via the selected vision-capable provider, with failover + usage |
GET | /v1/providers | provider health + selection state |
GET | /v1/byok/config | List the org's active provider configs — METADATA + masked hint only, NEVER the raw secret |
POST | /v1/byok/config | Create/replace (upsert) the org's config for a provider |
GET | /v1/byok/config/{provider} | One config's METADATA (masked). 404 if the org has no active config for provider |
DELETE | /v1/byok/config/{provider} | Soft-delete (deactivate) the org's config for provider — reversible via POST /v1/byok/config |
POST | /v1/byok/config/{provider}/rotate | Rotate the credential on an EXISTING active config. 404 if none exists |
POST | /v1/byok/purge | ERASE every stored provider credential the org holds — the BYOK leg of an org-delete cascade. IDEMPOTENT: success is asserted on configs_remaining == 0, so a re-run erasing 0 is still success. |
POST | /v1/images/generations | Image generation through the gateway's provider routing (OpenAI-shaped request) |
GET | /health | liveness |
GET | /v1/availability | SPOF documentation and dependency status |
GET | /metrics | Prometheus |
POST | /v1/prompt-cache/purge | Purge cached prompt/completion content for an org (GDPR Art-17 DSR propagation) |
GET | /v1/byok/config/{provider}/lifecycle | Credential lifecycle metadata — expiry, staleness, rotation schedule, owner principal |
POST | /v1/byok/config/{provider}/transfer | Transfer credential ownership to a different identity principal |
GET | /v1/byok/stale | Fleet-wide list of stale active BYOK credentials |
POST | /v1/approvals | Create a pending approval record for an MCP tool execution |
GET | /v1/approvals/{ref} | Get the status of an approval record |
POST | /v1/approvals/{ref}/approve | Approve a pending MCP tool execution |
POST | /v1/approvals/{ref}/deny | Deny a pending MCP tool execution |
POST | /v1/mcp/trust | Record a server trust assignment |
GET | /v1/mcp/trust/{org_id} | List all trust records for an org |
POST | /v1/mcp/trust/{org_id}/{server_id}/review | Review a trust record (required before VERIFIED or FIRST_PARTY takes effect) |
GET | /v1/replay/{request_id} | Get a captured execution's full context packet (for inspection or manual replay) |
POST | /v1/replay/{request_id} | Re-issue a captured request for real, against the CALLER's own tenant credential |
GET | /v1/traces | List recent execution traces for the caller's own org (content included, PII-redacted) |
GET | /v1/support/trace/{request_id} | Content-redacted, cross-tenant trace lookup for support triage (internal-only) |
GET | /v1/lineage/{request_id} | Content-free provenance lookup for one request (model/routing/prompt-template lineage) |
GET | /v1/lineage | List recent lineage events for the caller's own org (content-free) |
GET | /health/live | Liveness probe — is the process alive? (K8s/docker restart signal) |
GET | /health/ready | Readiness probe — can the gateway accept traffic? |
GET | /v1/fleet/status | Fleet-wide status, distinguishing gateway-down from vendor-down |
GET | /v1/canary | Run the golden-set canary against the configured model (CI gate; SPENDS MONEY) |
POST | /v1/canary/consistency | Response-consistency sampling — send the same prompt N times and check variance (SPENDS MONEY) |
GET | /v1/models/liveness | Probe each configured default model for liveness (SPENDS MONEY) |
GET | /v1/models/catalog | Model cards for every configured provider (EU AI Act Art.50 + enterprise DPA disclosure; UNAUTHENTICATED) |
GET | /v1/models/catalog/{provider} | Model card for one provider, canonical or alias (UNAUTHENTICATED) |
GET | /v1/models/registry | Full model registry — owner, successor, cutoff, and status for every model (UNAUTHENTICATED) |
POST | /v1/models/registry/{model_id}/deprecate | Mark a model as deprecated with a successor and reason |
GET | /v1/models/deprecation-notifications | Models needing customer notification — deprecated, plus expiring within 90 days (UNAUTHENTICATED) |
POST | /v1/admin/reload-policy | Hot-reload the routing policy from model-routing-policy.yaml without a redeploy |
GET | /v1/admin/routing-config | Current routing config — canary state, provider ranks, fallback ladder |
POST | /v1/admin/pin | Pin routing to a specific provider/model (demo-safe mode; disables failover) |
POST | /v1/admin/unpin | Clear the routing pin, restoring normal failover |
GET | /v1/admin/pin-status | Current routing pin status |
POST | /v1/admin/canary | Set percentage-based canary routing |
POST | /v1/admin/canary-stop | Stop canary routing — all traffic goes to stable |
GET | /v1/admin/canary-status | Current canary routing status |
POST | /v1/admin/orgs/{org_id}/suspend | Suspend an org — every egress route (llm/chat, embeddings, vision, images) refuses it (403) |
POST | /v1/admin/orgs/{org_id}/unsuspend | Unsuspend an org — restore fleet-key access |
POST | /v1/admin/sessions/{run_id}/abort | Request a cooperative abort of an in-flight session |
GET | /v1/admin/sessions/{run_id}/spend | Real-time spend for one in-flight session |
GET | /v1/admin/orgs/{org_id}/spend | Real-time aggregated spend for an org across all active sessions, against its budget ceiling |
POST | /v1/runs/mint | Mint a new run authority (called by the run's initiator — e.g. dialog-core, flow-builder) |
GET | /v1/runs/{run_id}/authority | Get the current state of a run authority (debugging / operator visibility) |
POST | /v1/runs/{run_id}/revoke | Revoke a run authority — the runaway-loop kill switch |
GET | /v1/security/blast-radius | Blast-radius query for a leaked provider key — which orgs/services are currently affected |
GET | /v1/pricing/drift | Pricing rows past their review cadence (UNAUTHENTICATED) |
POST | /v1/pricing/reconcile | Compare metered cost (from billing-metering) against vendor invoice amounts |
Product face
The external, paying-customer surface served by the edge at /product/v1/* (contracts/provider-gateway/product/v1/openapi.yaml).
| Method | Path | Summary |
|---|---|---|
POST | /chat | Chat/completion (managed, metered) — wraps the capability's multi-provider LLM egress |
Schemas
Usage
| Field | Type | Description |
|---|---|---|
prompt_tokens | integer | |
completion_tokens | integer | |
provider_used | string | |
cost_cents | number | |
estimated | boolean | false => provider-reported (real) — see D4 |
requested_provider | string | provider the caller asked for (absent when no substitution risk) |
requested_goal | string | optimization_goal the caller requested |
substituted | boolean | true when a different provider served than was requested |
cache_hit | boolean | true when the response was served from cache |
LlmRequest
| Field | Type | Description |
|---|---|---|
organization_id | string | |
model | string | |
messages | array | |
stream | boolean | |
tools | array | function-calling tool schemas (OpenAI shape); non-streaming OpenAI surfaces tool_calls. Additive — absent => unchanged completion path |
optimization_goal | string | |
byok | object | bring-your-own-key ref (never a raw key in body) |
allow_substitution | boolean | false => return 502 with original provider error instead of silent substitution |
data_classification | string | data classification for tenant routing policy |
locale | string | ISO 639-1 language code (e.g. 'en', 'es', 'fr') or BCP-47 locale (e.g. 'en-US', 'es-MX'). Controls the language the model responds in. When absent, the model infers language from input. When set, a system instruction is prepended: 'Respond in {locale}.' Two modes: - 'auto' (default): model infers from input text (current implicit behavior) - explicit locale: force response language regardless of input |
language_mode | string | How to interpret the locale field: - 'auto' (default): locale is used as a hint; model may still use input language - 'force': locale is a hard requirement; model MUST respond in this language |
LlmResponse
| Field | Type | Description |
|---|---|---|
content | string | |
tool_calls | array | function-calling tool calls the model emitted (OpenAI non-streaming); absent/None when the model returned none |
usage | object |
EmbeddingsRequest
| Field | Type | Description |
|---|---|---|
organization_id | string | |
model | string | |
input | object | |
optimization_goal | string | provider-selection goal for the embeddings failover chain (openai, gemini) |
byok | object | bring-your-own-key ref (never a raw key in body); resolved server-side per (org, provider) |
EmbeddingsResponse
| Field | Type | Description |
|---|---|---|
data | array | |
usage | object |
ImageRef
the image to analyze — supply EXACTLY ONE of url or base64 (base64 may include mime_type; the raw bytes are NEVER logged)
| Field | Type | Description |
|---|---|---|
url | string | publicly-fetchable image url |
base64 | string | base64-encoded image bytes (inline) |
mime_type | string | e.g. image/png — used with base64; defaults image/jpeg |
VisionRequest
| Field | Type | Description |
|---|---|---|
organization_id | string | |
model | string | override the vision model (default gpt-4o / gemini-2.0-flash) |
prompt | string | the task prompt — classification/alt-text/OCR instructions supplied by the caller |
image | object | |
optimization_goal | string | provider-selection goal for the vision failover chain (openai, gemini) |
byok | object | bring-your-own-key ref (never a raw key in body); resolved server-side per (org, provider) |
VisionResponse
| Field | Type | Description |
|---|---|---|
content | string | raw vision completion (parse JSON for classification; use directly for alt-text) |
usage | object |
Generated by scripts/gen-capability-docs.py from contracts/provider-gateway/v1/openapi.yaml — the contract IS the source of truth; edit the contract, not this page. The Edge route table section is synced from the live edge catalog snapshot (v1) — re-sync with python3 scripts/sync-edge-catalog.py.