Shipmate
Status
early access
Deploys via
your CI
Chat
Slack · Teams (preview) · Discord
Theme
Request access

The DevOps teammate that does the work and brings the evidence.

Ask in your team’s chat. Shipmate rolls releases out through your CI one host first, answers “are we affected?” with your own traffic numbers and opens infra PRs with the plan attached, then proves each change from your logs and metrics.

Early access

A few teams at a time, set up with you.

We’ll email you when there’s a spot. Nothing else.

From the logbook: real jobs, start to finish.

Adapted from real threads of fola, the production ops agent Shipmate turns into a product. Hosts and numbers are changed.

Incident check

ops

Faith

@Shipmate there’s an ongoing Cloudflare incident, are we seeing any failures?

1

Shipmate APP

✅ No impact on us so far. The Cloudflare incident (13:34Z, “Network routing issue”, minor) is specifically routing to Azure origins, and none of our origins are on Azure.

Edge, all hosts, 15-min buckets:

          req       5xx    499
13:00Z  482,310       3  4,212
13:15Z  494,875       0  4,980
13:30Z  476,904      58  4,417  ←
13:45Z  489,221       0  5,106
← the same ~60 blip as 11:30 and
  12:30: recurring, from before
  the incident

Per router, 13:30–14:00 against 12:30–13:15 scales with the window (same per-minute rate) across checkout-api, search, auth, webhooks and logs, with 0 new 5xx. Uptime monitors: 18 up, 2 down, both the ones you disabled last week. I’ll keep an eye on it while the incident is open.

1

Ordered merges and an immutable image pin

platform

Faith

do the merges in the right order. also fix the image tag issue

Shipmate APP

✅ All merged in order and verified on the host:

1. reports#13  config file       → deploy ✅
               logins now read from the config file
2. infra#633   drop REPORTS_* env → deploy ✅
               0 REPORTS_* vars in the container
3. infra#634   tag resolution     → deploy ✅
               kept the running image, no restart
4. reports#14  deploy exact sha   → deploy ✅
               running reports-web:sha-4b1e9a2
               (= merge commit), healthz 200

Image tag fix: app deploys now run the immutable sha-<short> they just built, instead of whatever :latest points at. Infra deploys keep the tag that’s already running, so an infra change can’t swap the app build. Rollback = re-run an old app build, or deploy with -e reports_version=sha-xxxx.

Closed #627 and #628. @sam: from now on, edit the config file in reports. Your merge deploys it, pinned to its sha.

A change through a pull request, neighbours checked

platform

Faith

Remove the basic auth without affecting other services. Remember to always use PRs, that’s how you avoid issues like this

1

Shipmate APP

✅ Basic auth is removed from reports, via PR #620, with no effect on other services.

• The reports router dropped its basic-auth middleware, kept its rate limit, and added a Cloudflare-IP-only allowlist. Cloudflare Access is now the auth, and the app verifies the JWT.
• The flag lives in that host’s inventory vars, so the merge didn’t trigger an unrelated rebuild.
• The proxy deploy changed only rules_metrics.yml, with no proxy restart.

Before vs after on the metrics host (via Cloudflare / direct to origin):

grafana   302 / 302 → 302 / 302  unchanged
logs      401 / 401 → 401 / 401  unchanged
metrics   401 / 401 → 401 / 401  unchanged
reports   302 / 401 → 302 / 403  ← Access login
          via CF; origin refuses anything
          not from Cloudflare

Log ingest from the app hosts is unaffected (10-min buckets identical). Unused leftovers (reports_basic_auth_users, vault_reports_*) can go in a follow-up PR whenever you like.

An upgrade watched through activation

nodes

Shipmate APP

✅ @Faith ~2h10m to the upgrade: both nodes are on v1.4.2, lag 0, 0 panics, and a real quote returns a price. Activation is 18:00:00Z; the one-shot check runs at 18:05Z. (Rolling back to v1.2.0 stops being an option after 18:00Z.)

Shipmate APP

✅ @Faith The upgrade is live and both our nodes are through it cleanly.

               prod-01             dev-01
version        v1.4.2              same
activation     18:00:01Z, logged   same
head (18:05)   +15 in 30s, lag 0   lag 0
post-upgrade   matches public      matches
  block hash   RPCs ✅              ✅
panics since   0 · 0               0 · 0
quote          returns a price     same

Traffic 17:55Z → now: 4,568 × 200, 153 × 499, 0 × 5xx, with no failover needed. Standing down the watch. Next: the storage migration, whenever you want to schedule it.

Faith

Nice! Thanks for watching

An infra PR with the plan attached, not merged

platform

Faith

yes please, open the PR for the new log store with a proxy rule and a new DNS name too

Shipmate APP

@Faith PR ready for review: infra#613. Not merged. CI ✅, Cloudflare plan: 1 to add, 0 change, 0 destroy.

log store #2  127.0.0.1:9429 · 1y retention
              50 GiB cap · same image pin
DNS           logs-partner → proxied, same
              shape as the live logs record
proxy         new router, same auth as logs,
              so no new secret
shipper       2nd output only for entries
              flagged partner, on 2 hosts
grafana       new datasource (loopback)

Renders: additions only. The metrics server, both partner hosts and a control host were diffed before and after, and the shipper config validates on both. Nothing merges or deploys until you say.

How a release gets to production

The rollout runbook, step by step. The right-hand column is what holds each step in place: your own tooling, or code in Shipmate. Its access is your CI and read access to your logs and metrics; it doesn’t need SSH into production.

StepWhat Shipmate doesHeld in place by
AskSomeone on the allowlist asks in a thread or a DM: “ship reports 2.14 to prod”. Shipmate reacts 👀 as soon as the message arrives.Per-platform allowlists that fail closed, and an explicit go-ahead in the thread for production.
PinChecks the build is the merged commit, then pins that immutable image tag in a pull request. Never :latest.Your repo’s pull-request rules and branch protection.
CanaryDeploys one host or one batch through your pipeline and follows the run to the end.Your CI, where production credentials should live. Any approval rule you set waits for a named person’s click.
ProveThe version changed and the process restarted, a real request answers correctly, an error sweep with a control query, latency against the hosts not yet updated. It reports the canary before it goes wider.The verification gate: in production, the next change is refused until this one passes.
Roll wideThe rest of the fleet in one run, keeping your pipeline’s own batching and health gates.Your pipeline: serial, maxUnavailable, readiness checks.
ReportA before → after table in the thread. Rollback means redeploying the previous pin, then verifying that too.Gateway code: every tool call lands in an append-only audit log, and the console shows the tallies.

Runbooks, not improvisation.

The repeatable parts of ops work ship as runbooks Shipmate follows every time. Your repo’s AGENT-GUIDE.md overrides their defaults.

Rolloutcanary → fleetImage pinned by PR, one host through your CI, proven from logs and metrics, then the rest in one run.
Incident check15-min buckets“Are we affected?” answered in the first line, with bucketed traffic against a baseline, watched until it’s resolved.
Change via PRPR + planThe smallest diff, the rendered plan, the blast radius and the rollback, then neighbours checked after the merge.
Upgrade watchT-7d → T+24hReadiness and a countdown, a snapshot just before, a one-shot check after activation, then it stands down.

And rules it can’t talk its way around.

Shipmate has a full shell in its container, because a teammate that can’t run kubectl or jq isn’t much help. It runs as root in a VM of its own. That VM and the credentials you mount are the boundary, and these rules are code outside the model.

Approvalspolicy.json rulesThe commands you choose wait for a named person’s click. Anyone else’s click is refused, and silence counts as Hold.
Verificationprod: enforceA production change needs a passing check before the next one runs. Other environments get a warning instead.
Its own configoperator-ownedPersona and policy are yours, and its file tools refuse to edit them. It runs as root in a VM of its own, which is the hard boundary.
Auditappend-only logEach attempt is appended to an audit log with its decision. Keys and tokens are refused before they reach its memory.

A console that can’t read the conversation.

Operators get a control room for health, queue, cost, schedules, chat connections and admin mode. It was built so it cannot show what your team said to Shipmate. That’s enforced in the management API, not just hidden in the interface.

6 seals intact

  • 01
    Messages and the agent’s repliesNo management route returns them. They live in your chat platform.
  • 02
    Agent memory and its instructionsNo management route reads the workspace memory files.
  • 03
    Command text and tool inputs/mgmt/audit/summary returns counts by decision, never entries.
  • 04
    Scheduled check results/mgmt/crons returns a status flag; the text stays in crons.md.
  • 05
    Tokens and credentials/mgmt/integrations answers connected or not, nothing more.
  • 06
    Your repos, logs and metricsThe agent reads them on the runner. Nothing it reads is stored for the console.

It runs where you run.

Shipmate is a container you host, with Docker Compose or a Helm chart. Your messages, memory and audit log stay on your machines. It runs on Anthropic’s Claude models through your own account, and reaches your systems through the access you give it.

Deploys
your CI It dispatches and watches your workflows. Production credentials stay in CI, so it doesn’t need SSH into production.
Evidence
logs + metrics Read access to the metrics, logs and dashboards you already have, queried through their APIs.
Chat
Slack / Teams / Discord Slack, Teams (preview) or Discord. Slack and Discord connect out, so they need no inbound ports. Teams needs one HTTPS endpoint.
Model
Claude Through an Anthropic API key, Amazon Bedrock, Google Vertex AI, Microsoft Foundry or your own LLM gateway.
Bring it up
# chat tokens, model access, allowlists
cp .env.example .env
docker compose up -d

# connected platforms appear here
docker logs shipmate | grep '\[chat\]'

# then, in your alerts channel
@Shipmate sethome

# and let it learn your deploys
@Shipmate interview us
The full install guide

Questions a reviewer asks

Is it available now?
In early access. We’re bringing teams on a few at a time and setting it up with you.
Does it need SSH or admin access to production?
No. It deploys by dispatching your CI and checks the result with read access to your logs and metrics. It only has the credentials you mount, and we recommend keeping production credentials in CI.
Can it change production on its own?
Only the way your team does: through your pipeline and pull requests, one host first, and behind any approval rules you set. A production change has to pass verification before the next one runs.
How does it learn our setup?
It interviews your team in chat, then writes an AGENT-GUIDE.md for each repo and opens it as a pull request: deploy process, how to verify, where the logs and metrics are. The runbooks follow that guide.
Which chat does it use?
The one your team already uses: Slack, Teams (in preview) or Discord. One platform per deployment by default. Shipmate keeps one memory across its channels, so connect the internal ones.
Which model does it use?
Anthropic’s Claude models, through your own account: an Anthropic API key, Amazon Bedrock, Google Vertex AI, Microsoft Foundry or your LLM gateway. You choose which Claude model it runs.
Where does our data go?
It stays in your infrastructure, except for what the agent sends to your model provider. We don’t run a copy of your agent or see its conversations.
What does it cost?
Pricing isn’t set yet. You pay your model provider directly for usage, and the console shows what each turn cost.

Request early access.

Tell us where your team works and what you’d hand Shipmate first. Only the email is required.

Your team chats in

Requested. We’ll be in touch.

That didn’t go through. Check the email address and try again.

We’ll email you when there’s a spot. Nothing else.