Deploy a RAG Stack in Minutes, Not Weeks
A step-by-step guide to launching a complete retrieval-augmented generation (RAG) stack on your own AWS account with SeaGit: a vector database, an answering API with citations, a private chat UI and document sync — from one form, in a few minutes.
What you get
One deploy gives you a working “chat with your documents” service that runs in your cloud account. Your documents never leave it.
- A vector database (Qdrant) on its own encrypted disk, holding the index of your documents.
- The RAG API — answers questions from your documents with numbered citations, through a simple REST API and an OpenAI-compatible
/v1/chat/completionsendpoint. - A private web UI — chat with citations, drag-and-drop upload, a model picker, protected by an access key.
- Sync jobs that keep the index up to date with your S3 buckets — on a schedule, within seconds of a change, or both.
- HTTPS URLs for the API and the UI, with certificates issued automatically.
Why this usually takes weeks
A production RAG service is much more than a model call. Doing it by hand means building and wiring every row below yourself. With SeaGit each one is part of the stack template.
| Piece of work | By hand | With SeaGit |
|---|---|---|
| Kubernetes cluster, networking, ingress, DNS and TLS | Days of Terraform and Helm | Pick a cluster SeaGit already runs |
| Vector database with persistent storage | Choose, deploy, size and back a volume | Qdrant on an EBS volume, sized in the form |
| Chunking, embedding and indexing pipeline | Write and operate ingestion code | Built in; chunk size and overlap are form fields |
| Keeping the index in sync with S3 | Cron jobs, S3 events, queues, retries, deletes | Pick "on a schedule", "on change" or both |
| Least-privilege IAM for S3 and Bedrock | Hand-written roles and bucket policies | Per-stack roles scoped to your folders and approved models |
| Answering API with citations | Build, secure and host an API | REST + OpenAI-compatible API with an API key |
| Chat UI with access control | Build a frontend and sign-in | Private web UI behind an access key |
| Which models people may use | Policy docs nobody enforces | Account and organization approved-model lists, enforced |
Before you start
- An AWS account connected as a provider.
- A started cluster in an environment. The Getting Started guide walks through both.
- At least one approved chat model. An account admin chooses which Bedrock models the account allows, and each organization can narrow that list further (Approved models in the sidebar). See Approved models.
- For Anthropic, Meta and Mistral models: model access enabled for the AWS account in the Amazon Bedrock console. Amazon Nova models work without extra steps.
Step 1 — Open the RAG Stack template
In your organization, go to Applications → Templates and choose RAG Stack, then Deploy. The deploy wizard has four short steps; every field has a sensible default.
Step 2 — Where it runs
- Environment and cluster — the cluster list follows the environment you pick.
- Ingestion endpoints — single runs one stack on one cluster. Multi-cluster runs a complete stack on every started cluster of the environment, each building its own index from the same sources, for resilience and lower latency.
Step 3 — Your documents
Data sources are S3 buckets the stack reads. For each one choose a folder, the file types to index, whether to import the files already there, and how to keep it in sync.
Uploads are files people add through the web UI or the API. Choose where they are kept and how they become searchable:
- Where uploaded files are kept — a new S3 bucket created for the stack, or your existing bucket (in this or another AWS account). For an existing bucket you also choose the folder and whether to import existing files now.
- Keep uploads in sync — On change (S3 events) makes a new file searchable within seconds; On a schedule picks files up on each run; On change + safety sweep (the default) does both; Off turns uploads off for the stack.
The details — access options, cross-account buckets, what each sync mode deploys — are in RAG Stacks: documents and sync.
Step 4 — Pipeline and answering
- Chunk size and overlap — smaller chunks give more precise citations; larger ones give the model more context.
- Embedding model — Amazon Titan Text Embeddings v2. Questions and documents use the same model.
- Vector index disk size — the Qdrant volume; grow it later if needed.
- Chat models — the models people can pick in the web UI and the API. Only models your organization approves are listed. Pick a default model for questions that don’t name one.
- Passages per answer (top-k) — how many matching chunks the model sees; 5 suits most documents.
- Include web UI — adds the chat and upload UI.
Step 5 — Deploy and save your keys
Select Deploy. SeaGit shows two keys once — copy them before closing the dialog:
- API key — for applications calling the RAG API.
- Web UI access key — for people signing in to the web UI.
The stack’s Stack tab shows each component as it starts, the URLs of the API and the UI, and the models the stack uses. Components are usually running within a few minutes; the HTTPS certificates follow shortly after.
Step 6 — Ask your first question
Open the web UI URL, sign in with the web UI access key, drop a document into the upload box and ask about it. Answers cite the passages they came from.
From code, call the API with the API key:
curl -s https://<api-host>/v1/query \
-H "Authorization: Bearer $RAG_API_KEY" \
-H 'Content-Type: application/json' \
-d '{"question": "What is our refund policy?"}'Upload from code with POST /v1/ingest (multipart file). The response returns stored once the file is saved; with an on-change sync mode it is searchable a few seconds later.
Next steps
- RAG Stacks reference — architecture, sync modes, approved models, access keys and the full API.
- Troubleshooting — when a component doesn’t start.
Frequently asked questions
How long does it take to deploy a RAG stack with SeaGit?
A few minutes. On our own clusters the four components (vector database, RAG API, web UI and sync jobs) report started in about one to three minutes, and the HTTPS certificates for the stack URLs are issued a couple of minutes later. Building the same stack by hand usually takes weeks.
Which cloud and which models are supported?
Stacks run on your own AWS account, on a cluster SeaGit manages. Answers come from Amazon Bedrock models (Amazon Nova, Anthropic Claude, Meta Llama, Mistral) that your account and organization approve; documents are embedded with Amazon Titan Text Embeddings v2 and indexed in Qdrant.
Can the stack read documents from an S3 bucket I already have?
Yes. Add the bucket as a data source, or keep uploads in it. SeaGit only adds its own access rules for the folder you choose, works with buckets in another AWS account, and can import the files already there when the stack is deployed.
How quickly does a new document become searchable?
With "On change (S3 events)" a new or updated file is searchable within seconds. With "On a schedule" it is picked up on the next scheduled run. "On change + safety sweep" does both, so nothing is missed.
Who can use the web UI?
Only people with the stack’s web UI access key. It is shown once when the stack is created and can be rotated from the Stack tab; rotating signs everyone out.