Documentation

Gateway API documentation

This page describes the gateway as it exists in the repository: the route table, the accepted request body, the numeric constraints, every error the gateway can return, and the metadata it stores. Where a sentence elsewhere on this site disagrees with this page, this page is the contract.

API reference Manually onboarded One configured model Non-streaming text only

Base URL

https://guntech.cloud

Gateway route

POST /v1/messages

Upstream

Anthropic Messages API

anthropic-version

2023-06-01

Transport

HTTPS, JSON request and response

Access

Operator-issued project key

01

Status, base URL and access

Reading this document

Access is granted to approved teams by a human operator, one project at a time. There is no self-serve registration and no automatically provisioned account; a project key is issued when an account is created.

A live domain is not an approval

The domain guntech.cloud is published as a website. That does not mean the gateway is open for public signup, that a request will be served for an arbitrary caller, or that a free key can be obtained. An operator issues every gt_live_ key, and revokes it when a pilot ends.

This build is not certified for production. "Implemented" means the corresponding source exists in this repository and can be verified locally; it does not mean the feature is deployed, and it does not mean a release gate has passed.

Two public endpoints report the runtime condition so a reader does not have to trust a status page:

  • GET /api/health returns status, upstream_configured, waitlist_ready and time. upstream_configured is reported separately on purpose: a reachable process is not proof that the Anthropic credential exists.
  • GET /api/public/config returns the configured model, waitlist readiness, the public Turnstile site key or null, bootstrap availability and the numeric request limits below.
02

Route table

Every route the gateway serves. Paths outside this table are not routed: /api/* returns a JSON error, and any other unknown path returns the site's HTML 404 page.

Gateway and console routes, current implementation
Method Path Authentication Result
GET/api/healthPublicBasic process health, upstream configuration flag, waitlist readiness
GET/api/public/configPublicConfigured model, waitlist readiness, public Turnstile site key, request limits
POST/api/waitlistPublic; Turnstile additionally, when configuredSave one pilot request; 503 only when the operator has closed intake
POST/api/setupOne-time bootstrap secretCreate the first administrator; 409 on a second attempt
POST/api/auth/loginEmail and passwordSet the HttpOnly session cookie
POST/api/auth/logoutAdministrator sessionRevoke the session and clear the cookie
GET/api/meAdministrator sessionCurrent user
GET, POST/api/projectsAdministrator sessionList projects with current-month usage, or create a project
PATCH/api/projects/:idAdministrator session, ownerUpdate monthly token ceiling and rate limit
GET, POST/api/projects/:id/keysAdministrator session, ownerList key prefixes, or generate a key (plaintext returned once)
DELETE/api/keys/:idAdministrator session, ownerRevoke a key
GET/api/projects/:id/usageAdministrator session, ownerCurrent month, six months of totals, latest 30 request events
GET/api/admin/waitlistAdministrator sessionLatest 100 pilot requests
POST/v1/messagesAuthorization: Bearer gt_live_…Relay a validated non-streaming text request to Claude

Requests to another owner's project return 404 not_found, not 403, so the console does not confirm that a project exists. Non-GET methods on static paths return 405 Method Not Allowed as plain text.

Sorting, filtering, pagination and export are not implemented on any route. The usage route returns a fixed window: the current month, the previous six months, and the most recent 30 request events.

03

Authentication and project keys

The gateway authenticates a caller with a project key in a bearer header. The key is scoped to one project, and the project's policy decides the rate limit and the monthly token ceiling.

Authorization header
Authorization: Bearer gt_live_YOUR_PROJECT_KEY
Content-Type: application/json

Key format

A key is the literal prefix gt_live_ followed by 24 cryptographically random bytes encoded as base64url. A key listing shows only the prefix and the first six characters of the random part, which is enough to tell two keys apart and far too short to be useful to an attacker.

Storage

Only a SHA-256 verifier of the key is stored. The plaintext value appears in exactly one response in the whole API: the 201 from POST /api/projects/:id/keys. It cannot be retrieved again, so a lost key is replaced, not recovered.

Revocation

Key lookup happens on every request, so revocation takes effect on the next call. A revoked key and an unknown key return the same 401 unauthorized body, with no hint about which one it was.

Upstream credential

The Anthropic credential is held by the deployment and sent only in the outbound request to Anthropic. It is never returned to a caller, never included in an error body, and never written to the metadata ledger.

Console sessions

The console uses an HttpOnly, SameSite=Strict cookie named gt_session with a seven-day lifetime, marked Secure whenever the public URL is HTTPS. Only a SHA-256 digest of the session token is stored. State-changing console requests must be same-origin, otherwise they are rejected with 403 origin_rejected. Login attempts are limited per client address in a fifteen-minute window and return 429 rate_limited when exhausted.

04

Request contract

The gateway accepts four top-level keys and nothing else: model, max_tokens, messages and system. It rebuilds the payload from those validated fields before forwarding, so an unvalidated field cannot reach Anthropic even if validation were skipped later.

POST /v1/messages — request only
curl https://guntech.cloud/v1/messages \
  -H 'Authorization: Bearer gt_live_YOUR_PROJECT_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "claude-sonnet-5-5",
    "max_tokens": 256,
    "messages": [{"role":"user","content":"Explain this error"}]
  }'

This example shows a request only. No response body is reproduced anywhere in this documentation, because the gateway does not fabricate one: a 200 carries the real Claude Messages payload, including the usage counters the gateway settles against.

model

Required string. It must equal the model configured for this deployment exactly. One model is configured per deployment; a different value returns 422 rather than being silently substituted, so a caller is never billed for a model they did not ask for.

max_tokens

Required integer, 16 to 4096 inclusive. It is also the amount reserved against the monthly ceiling before the request is forwarded, because output usage is only known after the model has produced it.

messages

Required array of 1 to 20 objects. Each object accepts only role (user or assistant) and content (a non-empty string). Content blocks, including image and tool-result blocks, are rejected: content must be text.

system

Optional string of at most 12,000 characters. It counts toward the combined content limit together with every message body.

What the gateway sends upstream

The forwarded request contains the configured model, max_tokens, the validated messages and, when present, system. The outbound call carries the Anthropic credential in x-api-key and anthropic-version: 2023-06-01. The response is returned to the caller unchanged.

05

Request constraints

These are the numbers the running code enforces, not targets. A value outside its range is rejected before anything is forwarded and before any quota is consumed.

Enforced limits
Rule Constraint Failure
max_tokensInteger 16–4096 inclusive422 validation_error
messagesArray of 1–20 items422 validation_error
messages[].roleuser or assistant422 validation_error
messages[].contentNon-empty string; content blocks rejected422 validation_error
messages[].*No fields other than role and content422 validation_error
systemOptional string, at most 12,000 characters422 validation_error
combined contentAll message bodies plus system: at most 40,000 characters422 validation_error
modelExactly the model configured for the deployment422 validation_error
top-level keysOnly model, max_tokens, messages, system422 validation_error, with details.unsupported for known-unsupported parameters
request bodyAt most 256 KiB, enforced while streaming422 validation_error
monthly token ceilingInteger 10,000–100,000,000 per project400 on policy update; 429 budget_exceeded at the gateway
rate limitInteger 1–300 requests per minute per key400 on policy update; 429 rate_limited at the gateway
reservation lifetime5 minutes, then a stale reservation is released by the sweepSwept every 15 minutes
upstream timeout55 seconds by default504 upstream_timeout
project name2–70 characters, no control characters400 name_invalid
administrator passwordAt least 14 characters400 password_too_short
bootstrap secretAt least 24 characters403 secret_invalid
pilot use case4–200 characters after trimming400 use_case_invalid
session lifetime7 days401 unauthorized after expiry
pilot request bodyAt most 16 KiB400 payload_too_large

Policy defaults

A new project starts at a monthly ceiling of 1,000,000 tokens and 30 requests per minute. Both values are configurable within the ranges above; the limits take effect on the next gateway request, because policy is read on every call.

How a reservation is computed

The gateway estimates input tokens conservatively at roughly 3.5 characters per token, adds four tokens of structural overhead per message and eight for the request, then reserves that estimate plus the full max_tokens. If used + reserved + estimate would exceed the ceiling, the request is rejected with 429 budget_exceeded before it is forwarded.

06

Error reference

Every error body has the same shape. Branch on error.type and the HTTP status; do not parse the human-readable message.

Error body shape
{
  "error": {
    "type": "validation_error",
    "message": "The \"max_tokens\" parameter must be between 16 and 4096.",
    "request_id": "req_0123456789abcdef01234567",
    "details": { "unsupported": ["stream"] }
  }
}

Human-readable message strings are currently written in Indonesian, matching the operator console. error.type, the HTTP status and details are the stable parts of the contract. request_id is generated per request and appears in the response body; persisting it in the usage ledger for cross-referencing is still pending work.

Gateway errors, POST /v1/messages
Status error.type Meaning What to do
401 unauthorized The Authorization header is missing or malformed, or the key is unknown or revoked. Send Authorization: Bearer gt_live_…. If the key was revoked, ask the operator for a new one; do not retry.
422 validation_error The body is not valid JSON, exceeds 256 KiB, or a field breaks a constraint above. Unknown top-level keys are named in error.details.unsupported. Fix the request. The same body will fail the same way; no quota was consumed.
429 rate_limited The key reached its project rate limit inside the current fixed one-minute window. Wait for the next window. The gateway performs no automatic retries and never queues a request.
429 budget_exceeded Used + reserved + this request's estimate would pass the monthly token ceiling. Nothing was forwarded to Anthropic. Raise the project ceiling, reduce max_tokens, or wait for the next UTC month. Retrying unchanged fails identically.
429 upstream_rate_limited Anthropic itself rate-limited the call. This is a provider limit, not a GunTech cap, and the project's own counters were released. Back off and retry later. Provider limits apply independently of the project ceiling.
503 not_configured The deployment has no upstream credential configured. The operator must set ANTHROPIC_API_KEY. The gateway will not synthesise an answer in the meantime.
502 upstream_unavailable The gateway could not reach the Anthropic service (network failure). Retry with backoff if the request is safe to repeat. The reservation was released.
502 upstream_rejected Anthropic rejected the call with a 4xx that the gateway's own validation did not catch. Check the model identifier configured for the deployment and the request shape. The upstream diagnosis is not echoed, because it can quote the prompt.
502 upstream_error Anthropic returned a 5xx, or returned a 200 whose body held no usable assistant text. Retry later. A 200 without usable text is treated as a failure so the console never shows a completed request that produced nothing.
504 upstream_timeout The upstream call passed the configured timeout, 55 seconds by default. Retry with a smaller max_tokens if appropriate. The upstream call may still complete and be billed, so do not retry blindly.
Console API error codes
Status error.type When
400invalid_json, payload_too_largeBody could not be parsed or exceeded the endpoint's size cap
400name_invalid, limit_invalid, rpm_invalid, policy_emptyProject creation or policy update failed validation
400email_invalid, use_case_invalid, consent_required, rejectedPilot intake fields failed validation. rejected is the generic answer to a failed honeypot or dwell-time check.
401unauthorized, invalid_credentialsNo live session, or the login pair did not match
403origin_rejected, secret_invalidCross-origin state change, or a wrong bootstrap secret
404not_foundUnknown /api/* path, or a project or key the caller does not own
409already_bootstrappedAn administrator already exists
429rate_limitedToo many login attempts from the same client address
503waitlist_disabledPilot intake has been closed by the operator (WAITLIST_ENABLED=false)
07

Not supported yet

These capabilities are planned, not available. A request that uses most of them is rejected with 422 validation_error and never reaches Anthropic, so a caller cannot accidentally buy a result this gateway does not understand.

Rejected with 422 today

  • Streaming — stream is an unknown top-level key.
  • Tool use — tools and tool_choice are unknown top-level keys.
  • Images and tool results — content must be a string, so content blocks are refused.
  • Model routing and multiple models — a value other than the configured model is refused.
  • Idempotency keys — an Idempotency-Key header is ignored, and there is no duplicate detection.
  • Automatic retries — the gateway never retries an upstream call.
  • Temperature and other sampling parameters — not in the accepted key list.

Not implemented at all

  • Batches — there is no /v1/messages/batches route; that path returns the site's HTML 404 page.
  • Response caching and prompt storage — prompts are not persisted and cannot be replayed.
  • Fine-tuning or model management.
  • Organizations, roles and invitations — the data model is single-operator.
  • Subscription billing, invoices and self-serve signup.
  • Usage exports, webhooks and notifications.
  • Additional providers, SSO/SAML, policy-as-code and formal compliance certifications.

If a request returns 422 and the reason is not obvious, read error.message and error.details.unsupported before changing anything else.

08

Token ceilings are not a spend guarantee

Read this before relying on a limit

Tokens are not money.

The monthly ceiling bounds the tokens the gateway observes for requests that pass through it. It is not a guaranteed monetary spending limit, it is not a billing cap, and it cannot see usage that never reaches this gateway. Anthropic's own billing is separate and must be reconciled against your provider invoice.

  • The reservation uses an estimate: roughly 3.5 characters per token, plus the full max_tokens. An estimate can be wrong in either direction.
  • Settlement uses the input_tokens and output_tokens counters Claude reports. If they exceed the reservation, the difference is still consumed, deliberately: the ceiling may be exceeded slightly by one expensive request rather than under-counted silently.
  • Failed forwarding releases the reservation in full, because nothing was inferred. Only successful inference consumes quota.
  • There is no pricing catalog yet, so the console reports tokens, not currency. Deriving a cash budget from a single blanket token rate is explicitly out of scope until a versioned pricing catalog and invoice reconciliation exist.
  • Provider-level limits (requests per minute, input and output tokens per minute) are enforced by Anthropic and can return 429 upstream_rate_limited regardless of the project ceiling. Do not read them as GunTech caps.
  • The ledger stores the two aggregate token counters. It does not decompose cache-specific token categories, so token counts here are usage telemetry, not an invoice line.
09

What we log

One metadata event is written per gateway request. The ledger is designed so that a request can be explained without storing what the caller asked or what the model answered.

Stored in the ledger

  • Key reference: the project and the API key that made the call.
  • Model identifier as configured for the deployment.
  • HTTP status and outcome: success, validation_error, rate_limited, budget_exceeded, upstream_error, not_configured or unauthorized.
  • Latency in milliseconds, gateway-side, including upstream time.
  • Input and output token counters for successful calls.
  • Timestamp, and the UTC month bucket the consumption settles into.

Never stored in the ledger

  • The prompt body, in any form.
  • The response body.
  • The plaintext project key. Only its SHA-256 verifier is written, and only the key prefix is shown in a listing.
  • The Anthropic credential.
  • The session token; only its digest is stored.
  • The Authorization header. Infrastructure logs are single-line JSON records carrying an event name, a correlation id, status, outcome and latency.

Pilot intake stores the email address, the described use case, an explicit consent flag, the source label and the submission time. Four layers protect the form, in the order a submission meets them: a per-address rate limit at the reverse proxy; a honeypot field named website, rejected when it is not empty; a minimum dwell time, checked against an elapsed_ms value the client sends, which must be at least 2000 and at most 3600000; and, when both Turnstile keys are configured, a challenge. A challenge failure returns 403 turnstile_failed. The honeypot and dwell checks return the same generic 400 rejected response on purpose, so a script cannot tell which check it failed. Dwell time is a cheap filter, not a security boundary: it stops a client that posts on load and nothing more determined. Intake is open by default; an operator can close it by setting WAITLIST_ENABLED=false, after which the endpoint answers 503 waitlist_disabled.

Retention and deletion windows for lead and usage data are not finalized. They depend on the legal review that is still open, so ask the operator for the current handling rather than assuming a schedule.

10

Console API

The console is an operator surface, not a customer dashboard. It authenticates with a session cookie and shows the persisted rows a deployment actually has, including empty states on a fresh install.

Bootstrap

POST /api/setup creates the single administrator when a bootstrap secret of at least 24 characters is configured and no user exists yet. A wrong secret returns 403 secret_invalid, a second attempt returns 409 already_bootstrapped, and the route returns 404 not_found once no bootstrap secret is configured. There is no self-registration route.

Key issuance

POST /api/projects/:id/keys returns the plaintext key once, alongside its prefix and identifier, with a notice that the value cannot be shown again. DELETE /api/keys/:id revokes a key; a revoked key retains minimal audit metadata but can never be used again.

Usage view

GET /api/projects/:id/usage returns the current UTC month's consumed and reserved totals against the ceiling, the previous six months, and the latest 30 request events. Empty results are returned as empty results; no sample or seeded figures are inserted.

Public configuration

The shape below is what GET /api/public/config returns. The value shown for model is substituted by the server with the model this deployment is configured with. The literal values shown correspond to a deployment with intake open and no Turnstile keys set: waitlist_ready is true and turnstile_site_key is null. Setting both Turnstile keys populates the site key; setting WAITLIST_ENABLED=false makes waitlist_ready false.

GET /api/public/config — response shape
{
  "model": "claude-sonnet-5-5",
  "waitlist_ready": true,
  "turnstile_site_key": null,
  "bootstrap_available": false,
  "limits": {
    "max_tokens_min": 16,
    "max_tokens_max": 4096,
    "messages_max": 20,
    "system_max_chars": 12000,
    "combined_max_chars": 40000
  }
}

Field names are exactly as returned. Treat waitlist_ready as authoritative: a client that renders an intake form must refuse to submit when it is false, and the API enforces the same rule with 503.

11

Retries, timeouts and duplicates

The gateway performs no automatic retries and implements no idempotency keys. A retry is the caller's decision, and the caller owns the risk of a duplicate inference.

  • 504 upstream_timeout: the gateway stopped waiting. Anthropic may still finish the call and bill it, so retrying immediately can produce two charges for one intended request.
  • 502: the call did not return usable output and the reservation was released. A bounded retry with backoff is reasonable.
  • 429 rate_limited and 429 upstream_rate_limited: wait. The gateway has no queue, so hammering the endpoint only produces more 429s.
  • 429 budget_exceeded: an unchanged retry repeats the same failure. Change the policy or the request size first.
  • Quote error.request_id when reporting a problem. Correlating that identifier with the stored event is planned work, not a current feature.

Health for a caller means two things, and they are reported separately: process health at /api/health, and whether an upstream credential exists in the same response.

12

Contract version and change control

The gateway forwards to the Anthropic Messages API with a fixed anthropic-version: 2023-06-01. This documentation describes the gateway's own current implementation; it is not a versioned public API contract, and there is no deprecation schedule yet.

Everything on this page is derived from the running source: the HTTP routes and their error bodies, the validation rules, the numeric limits, the upstream client and its failure mapping. Where a limitation is not implemented, this page says so rather than describing the intended end state.

Nothing here is a service guarantee. Availability, latency, recovery and security targets are measurement goals for a controlled pilot, not commitments, and no uptime figure or benchmark is published.

Still open

Before an external private pilot this deployment still needs release-gate work: hosted database and migration proof, backup and restore drills, real upstream success and failure validation against a funded Anthropic account, attack-path testing, settlement-failure reconciliation, alerting and an incident runbook, and finalized legal identity, retention and deletion policy.