Production-style Helm chart for the Vigil SOC platform. Ships the backend API (with bundled SPA), the autonomous SOC daemon, the LLM worker, and in-cluster Postgres + Redis as a single release.
TL;DR
helm install vigil ./infra/helm/vigil \
--namespace vigil --create-namespace \
--set secrets.anthropicApiKey=$ANTHROPIC_API_KEY \
--set secrets.postgresPassword=$(openssl rand -hex 24)
For a dev install with auth bypass:
helm install vigil ./infra/helm/vigil \
-f ./infra/helm/vigil/values-dev.yaml \
--namespace vigil --create-namespace \
--set secrets.anthropicApiKey=$ANTHROPIC_API_KEY
Then port-forward the backend to try it out:
kubectl port-forward -n vigil svc/vigil-backend 6987:6987
# open http://localhost:6987
What gets deployed
| Workload | Kind | Replicas | Notes |
|---|---|---|---|
vigil-backend |
Deployment | 2 (default) | FastAPI API + bundled SPA on port 6987 |
vigil-daemon |
StatefulSet | 1 (singleton) | Autonomous orchestrator; webhook=8081, metrics=9090, health=9091 |
vigil-llm-worker |
Deployment | 2 (default) | ARQ worker for Claude requests off Redis queue |
vigil-agent-worker |
Deployment | 2 (default) | Drains the BullMQ agent-runs queue; probes on 6990, no Service. Opt-out via agentWorker.enabled=false |
vigil-agent-serve |
Deployment | 2 (default) | Agent chat (SSE) and run projections on 6989; ClusterIP, reached by the backend at AGENT_URL. Opt-out via agentServe.enabled=false |
vigil-postgres |
StatefulSet | 1 | Opt-out via postgresql.enabled=false |
vigil-redis |
StatefulSet | 1 | Opt-out via redis.enabled=false |
vigil-db-init |
Job (Helm hook) | 1 per install/upgrade | Applies infra/database/init/*.sql idempotently |
Images
Three, not two. vigil-agent-worker and vigil-agent-serve run the same image with different commands — it is a Node image and cannot reuse the backend’s the way llm-worker does.
| Image | Used by |
|---|---|
ghcr.io/vigil-soc/vigil-backend |
backend, llm-worker, db-init |
ghcr.io/vigil-soc/vigil-daemon |
daemon |
ghcr.io/vigil-soc/vigil-agent |
agent-worker, agent-serve |
vigil-agent is published from the release that first ships these Deployments. If you mirror images into a private registry or run air-gapped, add the third one before upgrading — the agent Deployments are enabled by default and will otherwise sit in ImagePullBackOff. To upgrade without them, set agentWorker.enabled=false and agentServe.enabled=false.
Required inputs
At minimum you need:
secrets.anthropicApiKey— Claude API key for AI agentssecrets.postgresPassword— used for the in-chart Postgres; skip if using an external DB withpostgresql.existingSecret
When config.DEV_MODE=false (the default), also set:
secrets.jwtSecretKey— generate withpython -c "import secrets; print(secrets.token_urlsafe(64))"
When the agent Deployments are enabled (the default), also set:
secrets.agentInternalToken— generate the same way. It is the whole gate on the seam between the backend and the agent layer since ADR 0014, and it fails closed: leave it unset and every/internalcall answers 503 and every agent call answers 401, so no run gets past resolving its playbook.
External Postgres or Redis
Point the chart at existing infrastructure instead of running in-cluster:
postgresql:
enabled: false
external:
host: my-rds.example.com
port: 5432
database: vigil
username: vigil
existingSecret: my-rds-credentials
existingSecretKey: password
sslRequired: true
redis:
enabled: false
external:
url: "rediss://:password@my-elasticache.example.com:6379/0"
Pre-created Secret
If you manage secrets with ExternalSecrets Operator, SOPS, or Sealed Secrets, create the Secret yourself and point the chart at it:
secrets:
existingSecret: vigil-secrets
The Secret must define keys matching env var names:
ANTHROPIC_API_KEY, POSTGRES_PASSWORD, JWT_SECRET_KEY, plus whichever
integration creds you use (SPLUNK_PASSWORD, SLACK_BOT_TOKEN, …).
Upgrades
helm upgrade vigil ./infra/helm/vigil -n vigil --reuse-values
The chart’s default image tag resolves to Chart.AppVersion, which
release-please bumps in lockstep with the chart version on every
release — so a helm upgrade after pulling the new chart version
picks up the matching images automatically. Override only if you need
to pin to a different tag than the chart’s appVersion (for example,
to deploy a :latest build for testing):
helm upgrade vigil ./infra/helm/vigil -n vigil --reuse-values \
--set backend.image.tag=latest \
--set daemon.image.tag=latest
The db-init Job re-runs on every upgrade but is idempotent — it tracks
applied files in a _vigil_schema_versions table.
⚠️ First upgrade when
dbInit.sqlFileshas changed —helm upgrade --reuse-valuesreuses the previous release’s coalesced values, which means a longerdbInit.sqlFileslist in the new chart is silently overwritten by the previous (shorter) one. Any new SQL files in the bump won’t run, and code that touches their tables crashes at runtime. On the first upgrade after a chart bump that added init SQL, use one of:# Helm 3.14+ — reset to new defaults, then layer user overrides on top helm upgrade vigil ./infra/helm/vigil -n vigil --reset-then-reuse-values # Or pass an explicit values file so the new defaults aren't lost helm upgrade vigil ./infra/helm/vigil -n vigil -f my-values.yamlSubsequent upgrades that don’t touch
dbInit.sqlFilescan go back to plain--reuse-values.
Values reference
See values.yaml for the full schema. Non-obvious choices:
- Daemon singleton:
replicas: 1is hardcoded indaemon-statefulset.yamlbecause the orchestrator holds in-memory state. Do not template this. - LLM worker image: inherits from
backend.imageunlessllmWorker.image.repositoryis set. The only difference at runtime is the entrypoint (services.worker). - Daemon probes: target port
9091(/health), not9090. Port9090is the Prometheus/metricsport, which is only served whenconfig.VIGIL_OTEL_ENABLED=true.
Optional features (all off by default)
| Feature | Flag | Notes |
|---|---|---|
| NetworkPolicies | networkPolicies.enabled |
Default-deny + per-component allow rules |
| Prometheus ServiceMonitor | observability.serviceMonitor.enabled |
Requires Prometheus Operator CRDs |
| ExternalSecrets | secrets.externalSecret.enabled |
Pulls from AWS SM / Vault / GCP SM |
| Bitnami postgres subchart | postgresql.bitnami.enabled |
Run helm dependency update once |
| Bitnami redis subchart | redis.bitnami.enabled |
Same |
| OTEL Collector subchart | otelCollector.enabled |
In-cluster OTLP endpoint |
| KEDA queue-depth autoscaling | llmWorker.autoscaling.keda.enabled |
Requires KEDA operator |
| Splunk sidecar | splunk.enabled |
Dev/demo only |
| pgAdmin sidecar | pgadmin.enabled |
Dev/demo only |
See docs/HELM.md for end-to-end examples of each.
Development
# Lint
helm lint infra/helm/vigil
helm lint infra/helm/vigil -f infra/helm/vigil/values-dev.yaml
# Render without applying
helm template vigil infra/helm/vigil \
--set secrets.anthropicApiKey=test \
--set secrets.postgresPassword=test
# Dry-run install
helm install --dry-run --debug vigil infra/helm/vigil \
--set secrets.anthropicApiKey=test \
--set secrets.postgresPassword=test
Keeping SQL files in sync
The chart bundles copies of infra/database/init/*.sql under
infra/helm/vigil/files/database-init/. CI (.github/workflows/helm-chart.yml) will
fail on drift. To sync after adding new init SQL:
cp infra/database/init/*.sql infra/helm/vigil/files/database-init/
# then add the new filename to values.yaml -> dbInit.sqlFiles in order