Configuration
The chart is the supported surface; the environment variables are what the binary reads. Both are listed because you will meet both: the chart when you install, and the env vars when you read a pod spec at three in the morning.
The machine-checkable version of this page is
charts/bosun/values.schema.json,
which helm install validates against before it renders anything.
Credentials as files, not environment
Section titled “Credentials as files, not environment”Both forms are read once at start-up by the binary’s configuration loader, before anything is served. Each value goes to exactly one client: the git host, the model provider, ArgoCD. None of them reaches a prompt, and nothing the model returns is consulted about them. This section is about how the process is handed a credential, which is a different question from what can see it.
Every credential below reads from either its variable or a _FILE variant
holding a path: GIT_TOKEN_FILE, GITHUB_APP_PRIVATE_KEY_FILE,
LLM_API_KEY_FILE, ARGOCD_TOKEN_FILE, PROMOTION_TOKEN_FILE. A trailing
newline is trimmed, so a mounted Secret file works as it is.
credentials.mountAsFiles (default false) makes the chart do this for you: a
projected volume at /etc/bosun/credentials, mounted read-only, and the _FILE
variants set in place of every secretKeyRef. An environment variable is
visible to kubectl exec env, to /proc, and to every child process, and this
service shells out to git and helm, so the App private key was riding along
in their environment. With a path there instead, a child inherits a path.
It defaults off because _FILE is a convention rather than a Kubernetes
feature. The kubelet mounts the file and sets the variable to a path; opening
that path is code in the binary. image.tag follows appVersion, so an image
older than the release that added that code sees a path it does not act on and
starts with every credential empty. Turn it on once you are running an image
that supports it.
Setting both forms for one credential is a start-up error, not a precedence rule, and so is a path that cannot be read or a file that is empty. Two answers where the loader wants one is a question with no right answer, and guessing wrong surfaces as a rejected token, the symptom of a dozen unrelated mistakes.
Git host
Section titled “Git host”| Value | Env | Default | |
|---|---|---|---|
git.provider | GIT_PROVIDER | github | github, gitea implemented; gitlab, bitbucket are extension points |
git.owner | GIT_OWNER | n/a | REQUIRED |
git.repo | GIT_REPO | n/a | REQUIRED |
git.repoURL | GIT_REPO_URL | n/a | REQUIRED. Clone URL, reachable from the cluster. A credential embedded in it is taken out before git is run and supplied through the environment, so it reaches neither a command line nor the status page’s repository link |
git.existingSecret | n/a | n/a | REQUIRED. Existing Secret holding the credential; the chart creates none |
git.tokenKey | → GIT_TOKEN | token | Key within that Secret. Its value becomes GIT_TOKEN, which is the name the startup error uses |
git.apiBase | GIT_API_BASE | (unset) | See below; it means different things per host |
git.insecureSkipTLSVerify | GIT_INSECURE_SKIP_TLS_VERIFY | false | Scoped to the agent’s git client and its push clone, never global |
git.author.name | GIT_AUTHOR_NAME | (derived) | See below; leave empty unless you have a reason |
git.author.email | GIT_AUTHOR_EMAIL | (derived) | ” |
git.app.appId | GITHUB_APP_ID | (unset) | Set to authenticate as a GitHub App instead of a token |
git.app.installationId | GITHUB_APP_INSTALLATION_ID | (discovered) | Optional; discovered from the repository when unset |
git.app.existingSecret | n/a | git.existingSecret | Secret holding the App private key |
git.app.privateKeyKey | → GITHUB_APP_PRIVATE_KEY | private-key | Key within it. Its value becomes GITHUB_APP_PRIVATE_KEY |
When git.app.appId is set, GIT_TOKEN is not set at all: installation
tokens are minted from the key at runtime and live about an hour, so there is
nothing static to hold.
git.apiBase is host-specific
Section titled “git.apiBase is host-specific”The hosts differ, so the value does:
- github: the API root,
https://ghe.example.com/api/v3 - gitea: the instance root,
https://gitea.example.com. The client appends/api/v1itself, and also needs that root to build a push remote.
Leave git.author empty
Section titled “Leave git.author empty”Empty means the agent derives it. As a GitHub App that is its own bot identity
(<slug>[bot] <id+slug[bot]@users.noreply.github.com>), which is what makes the
pushed commits attribute to the App’s avatar rather than to a stranger.
If you do set an email, never use a users.noreply.github.com address that is
not your bot’s own: that namespace belongs to GitHub accounts. An earlier
default of bosun@users.noreply.github.com attributed every commit the first
live repair pushed, avatar and all, to an unrelated account named bosun.
| Value | Env | Default | |
|---|---|---|---|
llm.provider | LLM_PROVIDER | n/a | REQUIRED. openai or anthropic |
llm.model | LLM_MODEL | n/a | REQUIRED |
llm.baseURL | LLM_BASE_URL | (unset) | Required for openai; optional for anthropic |
llm.reasoningEffort | LLM_REASONING_EFFORT | (unset) | Passed through where supported; leave unset otherwise |
llm.timeout | LLM_TIMEOUT | 10m | |
llm.existingSecret | n/a | (unset) | Omit entirely for an unauthenticated local endpoint |
llm.apiKeyKey | → LLM_API_KEY | api-key | Key within that Secret. Its value becomes LLM_API_KEY |
There is no default provider. openai reaches OpenAI, Azure OpenAI, LM
Studio, Ollama, vLLM, llama.cpp and LiteLLM; anthropic reaches Anthropic and
gateways presenting the Messages API. See
Model providers for how to choose a model, and why
the score to optimise is unsafe actions = 0 rather than accuracy.
The gate
Section titled “The gate”| Value | Env | Default | |
|---|---|---|---|
gate.checkName | GATE_CHECK_NAME | addons-gate | Must match your branch protection rule |
gate.forkPRs | GATE_FORK_PRS | false | Render pull requests whose head is in another repository |
gate.poll | GATE_POLL | 30s | Paces the sweep for pull requests with no verdict yet. Must be positive; 0 is not a faster poll but no wait at all, and the process refuses to start with it |
gate.argocd.baseURL | ARGOCD_BASE_URL | (none) | Required. The ArgoCD API the inventory is read from; see below |
gate.argocd.podPort | (NetworkPolicy only) | 8080 | The port argocd-server’s pod listens on. Not the port in baseURL; see below |
gate.argocd.existingSecret | ARGOCD_TOKEN | (none) | Required. Secret holding the ArgoCD account token, key tokenKey |
gate.argocd.caSecret | ARGOCD_CA_FILE | (none) | Secret holding the CA that verifies argocd-server, key caKey |
gate.argocd.insecureSkipTLSVerify | ARGOCD_INSECURE_SKIP_TLS_VERIFY | false | Accept any certificate from argocd-server |
gate.registryAuth.existingSecret | HELM_REGISTRY_CONFIG | (none) | Secret holding one Docker config file, key key: the registry login helm renders with. For a chart a registry will not serve anonymously; see below |
gate.registryAuth.key | (volume only) | .dockerconfigjson | The key in that Secret, which is what a kubernetes.io/dockerconfigjson pull secret uses |
gate.registryAuth is for a chart behind a login
Section titled “gate.registryAuth is for a chart behind a login”Without it, a chart in a private registry is reported as not rendering at the new version, with the registry’s 401 for the reason, on every pull request that moves it.
The Secret holds one Docker config file, which is what helm registry login
writes and what an image pull secret already is:
{"auths": {"registry.example.com": {"auth": "<base64 of user:password>"}}}The chart mounts it read-only and names it to helm in HELM_REGISTRY_CONFIG.
That variable is helm’s own. The agent checks at start-up that the file is
there, because helm reads a missing one as no logins, and never opens it.
helm holds this login for every chart the gate renders, and a helm plugin a chart brings runs as helm. Give it a credential that can pull and nothing else.
The registry also has to be reachable: add its host to
networkPolicy.egress.fqdns.
gate.argocd is where the inventory comes from
Section titled “gate.argocd is where the inventory comes from”The gate reads four fields per cluster, name, server, labels and annotations,
from GET /api/v1/clusters on the ArgoCD API, which serves them with the
credential block redacted. Mint the token with
argocd account generate-token --account bosunand give it three read lines in argocd-rbac-cm — clusters, get,
applications, get, */* and applicationsets, get, */* — and nothing else.
The first is the cluster inventory; the other two are what the gated repository
deploys, which the gate derives rather than reading out of a file that
repository keeps in step by hand.
It is the API and not the cluster Secrets those clusters are stored in
because that read cannot be made small enough. Kubernetes RBAC has no
predicate for “the labels but not the data”: there are no deny rules,
resourceNames does not apply to list (a list request carries no name for
the authorizer to match), and a label selector in the request is a filter the
apiserver applies after authorising, so a token holding such a Role could
drop it and read every Secret in the namespace. The chart creates no Role over
Secrets at all.
What it costs, stated as plainly as the grant it replaces: a credential to rotate, a component that can be down on its own (the apiserver is up whenever the cluster is; argocd-server is not), and its own TLS story, because argocd-server serves its own certificate rather than the one the kubelet mounts into every pod. The chart adds the NetworkPolicy egress rule for the ArgoCD namespace itself, because argocd-server is a ClusterIP and forgetting that hangs with zero bytes.
gate.argocd.podPort is the pod’s port, not the URL’s
Section titled “gate.argocd.podPort is the pod’s port, not the URL’s”The rule the chart adds opens gate.argocd.podPort, and that is normally not
the port in baseURL. A NetworkPolicy matches the destination port of the
packet, and a ClusterIP is DNAT’d to the backend pod’s port before policy is
evaluated: whatever port the Service published, the packet reaching the rule is
addressed to the pod, on 8080.
Setting it to the Service port instead renders clean, passes helm lint,
passes the chart’s schema, and then drops every packet. There is no error at
either end: the connection hangs for the full HTTP timeout, and the pod dies at
start-up saying argocd-server is unreachable, which is true and points nowhere
near the values file.
8080 is argocd-server’s container port in the upstream argo-cd chart, and it
does not move with server.insecure: the Service publishes both 80 and 443
against the same container port, and argocd-server decides per connection
whether to speak TLS on it. So the two common installs differ only in the URL,
http://argocd-server.argocd.svc with nothing to verify or
https://argocd-server.argocd.svc with caSecret, and both want podPort: 8080. Confirm yours:
kubectl -n argocd get svc argocd-server -o jsonpath='{.spec.ports[*].targetPort}{"\n"}'The chart writes bosun’s egress. argocd-server’s ingress policy lives in your ArgoCD release, and until both exist the connection is dropped with nothing logged at either end; the chart README carries a copy-pasteable rule for it.
gate.forkPRs is off for a reason
Section titled “gate.forkPRs is off for a reason”The render runs helm over the pull request’s content, inside your
cluster. Whose content that is should be an operator’s decision. Off, a fork
pull request gets an error status naming this value: a refusal you can see,
rather than a required check that never reports.
Triage
Section titled “Triage”| Value | Env | Default | |
|---|---|---|---|
triage.allowPaths | ALLOW_PATHS | [] | Where the agent may ever write. Empty refuses everything, and the process refuses to start with it |
triage.denyPaths | DENY_PATHS | [] | Added to the built-in deny-list; cannot subtract from it |
triage.maxAttempts | MAX_ATTEMPTS | 2 | Attempt cap, tracked by pull-request label |
triage.explainGreen | EXPLAIN_GREEN | true | Explain green gates on held pull requests |
triage.migrateDroppedVersions | MIGRATE_DROPPED_VERSIONS | true | The deterministic apiVersion repair. No model involved |
triage.structuralMigration | STRUCTURAL_MIGRATION | true | The schema-guided repairs: reshaping a document an apiVersion swap left behind, and migrating values a new chart version refuses |
triage.migrateMaxDocs | MIGRATE_MAX_DOCS | 5 | Cap on documents reshaped in one pass |
triage.egressDeny | EGRESS_DENY | [] | Public hosts the upstream lookup must never reach |
triage.egressAllowPrivate | EGRESS_ALLOW_PRIVATE | [] | Internal networks it may reach after all. See below |
triage.upstreamNotes.enabled | UPSTREAM_NOTES | true | Fetch publisher release notes for the explain and escalate paths |
triage.upstreamNotes.maxReleases | UPSTREAM_MAX_RELEASES | 5 | |
triage.upstreamNotes.maxCommits | UPSTREAM_MAX_COMMITS | 10 | |
triage.upstreamNotes.maxBodyChars | UPSTREAM_MAX_BODY_CHARS | 4000 |
triage.allowPaths is the whole write surface
Section titled “triage.allowPaths is the whole write surface”An empty allowlist refuses everything and the process refuses to start with one. A service that can write nowhere and does not say so looks broken later, for a reason nobody will find.
Set it to the tree the agent may repair, typically [addons/**]. It is a
standing grant and deliberately coarse; the per-request bound is Scope, read
from the pull request’s own diff against its base, and both must pass. The
promotion body still carries a files list, and it reaches the prompt rather
than the applier — a caller cannot widen what it may edit by claiming to have
touched more.
triage.egressAllowPrivate is the only way past a closed network
Section titled “triage.egressAllowPrivate is the only way past a closed network”Egress is open to the public internet and closed to internal address space, and
the second half is not a default that configuration turns off. Loopback,
link-local (169.254.0.0/16), the RFC1918 blocks, CGNAT (100.64.0.0/10) and
the IPv6 equivalents are refused at the dial, where the address is finally
known, so a name that resolves into one of them is refused too.
That is right for reading public chart indexes and registries, which is what this reaches. It is wrong if yours are internal, and this is where you say so:
triage: egressAllowPrivate: - 10.42.0.0/16 # the internal chart museumAn entry is a CIDR or a single address. Naming one network opens that network and nothing else, and an entry that parses as neither opens nothing. Without the entry the pod runs and the symptom is a refusal in the log naming the network it stopped, with the brief degrading to “no evidence”.
This governs the chart-repository and registry lookups. The model, git, ArgoCD and apiserver clients do not go through it, so a model endpoint on a private address is not what this setting is for.
The deny-list cannot be shrunk
Section titled “The deny-list cannot be shrunk”triage.denyPaths adds to a built-in list that configuration cannot remove
from. Every entry is a way to make a red gate green without fixing anything:
.github/** the workflows that run the gate.gitops-gate.yaml what the gate renders, and how.bosun.yaml the same file under the name the gate is moving todelivery/** the kit itself, including this agent and its prompt.gitlab-ci.yml the GitLab and Bitbucket equivalentsbitbucket-pipelines.yml**/kargo-projects/** the merge policy and version constraints**/kargo-pipelines/** the promotion pipelines themselvesThe matcher understands ** at the start of a pattern, at the end, or at both,
not in the middle. A wildcard inside a directory name (**/kargo-*/**) is
not something the deny-list can express.
triage.upstreamNotes never feeds the write path
Section titled “triage.upstreamNotes never feeds the write path”Release notes and upstream commits are fetched only on the paths that produce prose: the green-gate explanation and an escalation. The mechanical path, the one that writes files, does not fetch them, so they are not in the evidence string the applier corroborates against.
Without that rule, a commit message containing v1.5.0 would make v1.5.0 a
corroborated value to write. See
ADR 0005.
Live cluster reads
Section titled “Live cluster reads”| Value | Env | Default | |
|---|---|---|---|
liveReads.enabled | LIVE_READS | false | Off by default: everything else the agent reads is public or already in the pull request |
liveReads.scope | (RBAC only) | groups | groups or wide |
liveReads.apiGroups | (RBAC only) | [] | The groups granted under groups scope |
liveReads.argocdNamespace | LIVE_READS_ARGOCD_NS | argocd | Also the namespace the chart opens egress to for argocd-server |
Live reads are get and list only. The chart’s ClusterRole has no create,
update, patch or delete verb anywhere.
The supervisor
Section titled “The supervisor”| Value | Env | Default | |
|---|---|---|---|
supervise.enabled | SUPERVISE_PIPELINE | true | |
supervise.interval | SUPERVISE_INTERVAL | 10m | |
metrics.serviceMonitor.enabled | n/a | false | Scrape /metrics |
metrics.serviceMonitor.namespace | n/a | n/a | Required when the ServiceMonitor is on and networkPolicy.enabled. The namespace Prometheus runs in. The port serves the whole HTTP surface, not only /metrics, so the ingress rule needs a namespace as well as a pod label — a pod label alone is chosen by whoever creates the pod |
Read-only: three LISTs and a shallow clone, using the Kargo read the ClusterRole
already grants. Both /pipeline and /metrics answer 503 before the first
sweep completes, deliberately: a scraper reading zeroes from a supervisor that
has not looked yet would record “nothing is wrong” as a measurement.
See The pipeline supervisor for what it looks for and the two alert rules worth having, including the one that fires when the supervisor itself goes quiet.
Network policy
Section titled “Network policy”| Value | Default | |
|---|---|---|
networkPolicy.enabled | true | |
networkPolicy.flavor | standard | |
networkPolicy.kargoNamespace | kargo | Which namespace may call the triage hook |
networkPolicy.kargoPodSelector | {} | Which pods in it may. Empty admits the whole namespace, which is every workload in it, not only the Kargo controller |
networkPolicy.egress.dnsNamespace | kube-system | |
networkPolicy.egress.dnsPodSelector | {} | Narrows DNS egress to the resolver pods. Empty opens 53 to the whole namespace |
networkPolicy.egress.namespaces | [] | |
networkPolicy.egress.ipBlocks | [] | Your model endpoint goes here |
networkPolicy.egress.apiServer.ipBlocks | [] | The apiserver’s real endpoints |
networkPolicy.egress.fqdns | [] | Registries the upstream lookup may reach |
networkPolicy.egress.fqdnPatterns | [] | |
networkPolicy.egress.allowPublicHTTPS | false | Your git host |
networkPolicy.egress.allowInternet | false |
Deployment shape
Section titled “Deployment shape”| Value | Env | Default |
|---|---|---|
image.repository | n/a | REQUIRED |
image.tag / image.digest | n/a | (appVersion); prefer a digest |
image.pullPolicy | n/a | IfNotPresent |
replicaCount | n/a | 1 |
service.port | AGENT_ADDR | 8080 |
branding.name | AGENT_BRAND | Bosun |
promotionAuth.existingSecret | n/a | (unset) |
promotionAuth.tokenKey | → PROMOTION_TOKEN | token |
maxConcurrentTriage | MAX_CONCURRENT_TRIAGE | 4 |
serviceAccount.create / .name | n/a | true / (fullname) |
rbac.create | n/a | true |
resources | n/a | 25m CPU / 64Mi requested, 512Mi memory limit |
nodeSelector, tolerations, affinity, podAnnotations, priorityClassName | n/a | standard |
The pod spec also carries CLONE_ROOT=/work, which is not a chart value. It is
where the agent and the gate clone a pull request’s branch, backed by an
emptyDir so the root filesystem can stay read-only. Nothing there is worth
persisting: a restart mid-triage starts clean rather than resuming.
branding.mark is deprecated and ignored since 0.17.0. Comments no longer
carry an identity header at all. Authenticating as a GitHub App puts the name
and avatar above every comment already, and a bold header under that was the
agent introducing itself twice. Still accepted so setting it does not fail an
upgrade.
Authenticating the promotion endpoint
Section titled “Authenticating the promotion endpoint”POST /v1/promotion-opened takes a pull-request number and the list of files
the agent may edit and will read into the prompt it publishes. The NetworkPolicy
restricts callers to the namespace, which is as narrow as a NetworkPolicy gets,
so every workload in that namespace qualifies.
promotionAuth.existingSecret names a Secret holding a shared token; the value
at promotionAuth.tokenKey becomes PROMOTION_TOKEN, and requests then need
Authorization: Bearer <token>. Kargo sends it via the kargo-pipelines value
triage.authorization, which is the whole header, so a Kargo expression keeps
the token out of your values file:
triage: authorization: '${{ "Bearer " + secrets.bosun.token }}'It is opt-in, so an upgrade does not silently stop answering Kargo. Unset, the endpoint is open and the pod logs a warning saying so at every start-up.
maxConcurrentTriage bounds how many pull requests are triaged at once. Each is
a clone, a helm render and a model call, so the ceiling is about the pod’s memory
and your git host’s rate limit, not throughput.
The gate’s own config file
Section titled “The gate’s own config file”Most repositories have none. The gate asks ArgoCD which Applications and ApplicationSets exist and renders their paths from the pull request’s checkout, so the pointers are read live rather than restated in a file somebody keeps in step by hand (ADR 0012).
Where a file is wanted it is .bosun.yaml, in the repository being gated,
not in this chart, and .gitops-gate.yaml is still read under its old name.
Both is an error. Configuring the gate is the whole
schema, and leads with why you probably need none of it.
Two of its keys have moved here instead: gate.concurrency and
gate.validate.* below. The renders happen in this pod, so how hard the gate
works and what it checks are decisions about your cluster rather than about the
repository under review, the same line the egress deny-list is on. Each is
unset by default, and unset leaves the gated repository’s own file alone.