SEAGIT DOCS
LLM Gateway

LLM Gateway

One OpenAI-compatible endpoint for Amazon Bedrock models, running in your own AWS account: per-team keys and budgets, an input guardrail, a log of every call in your S3 bucket, and an optional usage dashboard. Built on the open-source LiteLLM proxy.

What gets deployed

The LLM Gateway stack deploys up to three SeaGit applications into one cluster of an environment, managed together from the Stack tab:

  • litellm: the gateway. It serves /v1/chat/completions, /v1/models and the LiteLLM Admin UI over HTTPS on your environment’s domain. It calls Bedrock with its own IAM role, which may invoke only the models you chose. No AWS keys are stored.
  • db (only with “Create new Postgres”): Postgres 17 on an encrypted EBS volume, kept off spot nodes, with a nightly backup to the stack’s S3 bucket. LiteLLM keeps keys, teams, budgets and spend here.
  • llm-dashboard (optional): the usage dashboard, which reads the gateway’s logs with read-only access.

Each component appears in the Stack tab with an open link (↗) once it is running.

Requirements

  • A paid plan.
  • A started cluster on the native (sgt-k8s) executor, in an environment with a domain. The gateway and dashboard are exposed on that domain.
  • The EBS CSI add-on on the cluster, when you create a new database.
  • Model access enabled in the Amazon Bedrock console for the models you choose. Amazon Nova models need no request; Anthropic, Meta and Mistral models do.

Deploy

Open Templates, choose LLM Gateway, pick the environment (and the cluster, if it has more than one), and answer:

QuestionWhat it does
ModelsThe Amazon Bedrock chat models the gateway serves. Only models your organization approves are listed.
Default modelAlso served under the model name "default", so applications can switch models without a code change.
Database"Create new Postgres" adds a Postgres database to the stack, backed up nightly to the stack’s S3 bucket. "Use existing Postgres" takes a postgresql:// URL that must be reachable from the cluster; it is stored encrypted and never shown again.
Database storage / Keep backups forOnly for a new database: the volume size (GiB) and how many days of nightly backups to keep.
Include usage dashboardAdds the usage dashboard (cost, usage, blocked requests, traces). On by default.

A deploy usually takes 1–3 minutes. When it finishes, the Stack tab shows:

OutputMeaning
Gateway URLThe OpenAI-compatible endpoint, e.g. https://<name>.<your domain>.
Admin UI<Gateway URL>/ui. Sign in as admin with the master key.
Master keyFull admin access to the gateway. Hidden until you reveal it; every reveal is audited.
Dashboard / Dashboard keyThe usage dashboard and its sign-in key.
Database hostOnly for a new database: its in-cluster address.
Call logs / Database backupsWhere the per-call log and the database backups are written in your S3 bucket.

Call the gateway

Any OpenAI-compatible client works. Point it at the Gateway URL and use a key. GET /v1/models lists the model names: each chosen model under a short name (for example amazon-nova-micro), plus default.

curl -s https://<gateway-url>/v1/chat/completions \
  -H "Authorization: Bearer $LLM_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "default",
    "messages": [{"role": "user", "content": "Summarise our travel policy"}],
    "metadata": {"team": "support", "feature": "faq", "trace_id": "run-42"}
  }'

The optional metadata fields (team, feature, env, session, trace_id) end up in the call log and the dashboard. Use the same trace_id for every call of one agent run to see it as a timeline.

Keys, teams and budgets

Give each team or application its own virtual key rather than the master key. Create keys in the Admin UI (Virtual Keys) or through the API:

curl -s https://<gateway-url>/key/generate \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"key_alias": "team-search", "models": ["default"], "max_budget": 50}'
  • A key can call only the models listed for it. Any other model is refused with 401.
  • max_budget (USD) stops the key once its recorded spend reaches the budget. Spend is recorded per key and per team.
  • Requests without a key, or with an unknown key, are refused with 401.

Keys, teams and spend live in the stack’s database. They survive Stop/Start and redeploys.

Guardrail

Every request is checked before it reaches the model:

  • Prompt attacks (for example “ignore all previous instructions”) are blocked with 400. The model is never called and nothing is charged.
  • Personal data (email addresses, US social security numbers, card numbers) is replaced with a placeholder such as {EMAIL} before the model sees it.

The guardrail matches patterns. It catches common cases, not every possible attack or format. Each finding is logged, so you can see what it caught in the dashboard.

Call log

Every call (answered, blocked or failed) is written as one JSON line to your S3 bucket under gateway-logs/dt=YYYY-MM-DD/, within about 10 seconds. Prompts and answers are never logged, only:

FieldMeaning
ts, request_idWhen the call finished and its id.
model, upstream_modelThe model name the client asked for and the Bedrock model that answered.
status, stop_reasonsuccess, blocked or error; and how the answer ended (stop, length, guardrail_intervened).
input_tokens, output_tokens, cost_usdUsage and estimated cost.
latency_ms, guardrail_msEnd-to-end time and guardrail time.
key_aliasWhich virtual key made the call.
team, feature, env, session, trace_idCopied from the request’s metadata, when you send it.
guardrail_findingsWhat the guardrail blocked or masked.

The files are plain NDJSON, so Athena, DuckDB or any log tool can read them directly.

Usage dashboard

Sign in with the Dashboard key. Choose a window (1–90 days) and switch between tabs:

  • Usage: calls over time by outcome, and per model and per key (latency, truncated answers).
  • Cost: spend over time by model, and by key, team and model.
  • Blocked: guardrail findings over time, and the blocked requests.
  • Tracing: one trace_id as a timeline, and the latest calls.

Costs are estimates from public Bedrock prices. Your AWS bill is authoritative.

Database and backups

  • New database: backed up right after deploy and then nightly to backups/ in the stack’s bucket, kept for the number of days you chose. Deleting the stack takes a final backup first; if that backup fails, the delete stops and offers Delete anyway.
  • Existing database: SeaGit only connects to it. Its backups, upgrades and availability are yours. The gateway waits for the database at start-up and is not marked ready until it can reach it.

Stop, start and delete

  • Stop scales every component to zero. The database volume, keys and logs are kept. Start brings it back.
  • Delete removes the components, IAM roles and DNS records. Tick Also delete my backups to also delete the stack’s S3 bucket, including the call logs and database backups; otherwise the bucket and its contents are kept.

Current limits

  • Amazon Bedrock models only.
  • The Admin UI is protected by the master key. Single sign-on is not set up by the stack yet.
  • LiteLLM’s audit log of admin actions (who created a key, changed a budget) is a LiteLLM Enterprise feature and is not included. The call log records every model call.
  • One gateway replica.

Troubleshooting

  • A component stays pending: the cluster may be out of capacity. A new database needs a non-spot node with about 512 MiB free. See Troubleshooting.
  • The gateway never becomes ready: it cannot reach its database. For an existing database, check the URL and that the cluster can reach it on port 5432.
  • 401 for one model only: that key isn’t allowed to use the model.
  • A model returns an access error: enable model access for it in the Amazon Bedrock console.