Runbook — repair domains rows against the SES identity that actually covers them

Three steps, two scripts, in a fixed order:

  1. platform_managed — flag the sandbox rows that already existed when the column was added. Everything below keys off that column, and the migration's DEFAULT false stamped it wrong on exactly the rows that matter. Its own script: scripts/backfill-platform-managed-domains.ts.
  2. region — rows written before the create-path fix name a region their SES identity was never in. Required on any deployment whose AWS_REGION is not us-east-1. This deployment runs AWS_REGION=ap-southeast-2, so it needs it.
  3. status / verified_at — the eleven *.sandbox.sendoka.com rows the daily recheck unverified. This half is ordered after a deploy (see below) and lives behind its own flag.

⚠️ Ordering — read this before you run anything

--restore-status must not run until the platform_managed exclusion in domain-verify-poll is live in production.

Restoring status = 'verified' while the deployed cron still reconciles platform-managed rows against SES does not fail — it re-arms the flip. The row goes back to pending on the next daily recheck and the customer receives a second domain.unverified webhook about a domain that never stopped working.

The script enforces what it can: --restore-status --apply refuses unless you pass --cron-fix-deployed and the cron route in your working tree mentions the column. The first is your attestation about production — nothing running on your laptop can check that. The second can only prove the fix is absent.

The region half has no such constraint. Run it whenever.

⚠️ Dev and prod share one Neon database. There is no staging copy to rehearse against — whatever DATABASE_URL resolves to is production. Run the dry pass and read it before --apply.

Why — part one, the region

createDomainIdentity has always created the SES identity in AWS_REGION. Until the fix, both add-domain routes then inserted a domains row with no region, so Postgres filled in the schema column default "us-east-1" (dropped in migration 0053, once every insert path supplied the column). Every row written before the fix therefore names a region its identity was never in, and every later operation keys off the row, not off the env:

  • domain-verify-poll calls checkDomainStatus(d.domain, d.region), gets NotFoundException from us-east-1, and substitutes verified: false. The "identity missing" early return is guarded on !wasVerified, so a row that is already verified skips it, falls through to justUnverified, flips to pending and fans out a domain.unverified webhook at the customer — for a domain that is verified and sending.
  • Deleting a domain (internal/domains) calls deleteDomainIdentity(record.domain, record.region), so the delete goes to us-east-1, finds nothing, and the real ap-southeast-2 identity survives and keeps sending. The row is gone by then, so nothing in the product mentions that identity again.

The fix stops new rows being written wrong. It does not touch the rows that already are — nothing self-heals.

Why — part two, the sandbox rows

Every org gets a {org-slug}.sandbox.sendoka.com row at signup, written verified by hand in the same transaction that creates the org. Neither signup path creates an SES identity for it, deliberately. The row rides the parent sendoka.com identity, which SES authorises for every subdomain — that is what lets a trial account send with no DNS at all, which the product offers in writing. The parent identity is a real operational dependency and these rows have to end up usable for live sends, not merely be kept quiet.

GetEmailIdentity is an exact-name lookup. It does not resolve up the suffix chain, so it misses on a subdomain a parent identity covers, and that miss is indistinguishable from an identity somebody deleted. Two things read it as deletion:

  • The cron, once NotFoundException was correctly reclassified from a transient error into a definitive answer (5f59cde, in prod as b5136b5). Eleven rows are already flipped to pending with verified_at cleared, each with a customer-facing domain.unverified webhook carrying a dkim_status of IDENTITY_NOT_FOUND — a value we synthesise, not one SES returned. Fixed by the platform_managed exclusion in the cohort select.
  • This script, which parked those rows in "found in no region" and told you to "re-add the domain to mint a fresh one, or delete the row". Both halves were actively harmful and both are gone: nothing re-creates a sandbox row (the insert lives only in the new-org branch of signup), so deleting it permanently removes that org's sandbox sender, and re-adding through POST /v1/domains answers 422 DOMAIN_RESERVED (no name under sendoka.com can be added by a customer).

The column, not a hostname suffix match: the sandbox domain has been renamed twice already (mrsendo.dev → chaparly.dev → sendoka.com), and a string that carries a semantic stops carrying it the next time someone renames it.

Prerequisites

Needed for Checked by the script
DATABASE_URL, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY steps 1 and 2 — step 0 touches Postgres only and needs just DATABASE_URL refuses without them
the domains.platform_managed migration applied (0054) everything refuses without the column
platform_managed backfilled on the existing sandbox rows (step 0) the sandbox half not directly — but a covered row still flagged false is reported as its own verdict rather than written or condemned
the cron exclusion deployed --restore-status --apply only --cron-fix-deployed + a grep of the cron route

ADD COLUMN … DEFAULT false stamps false on every existing row, sandbox rows included. Until step 0 runs, the eleven are still indistinguishable from customer rows by the column — they will surface as covered by a parent identity … platform_managed is false and this script will refuse to touch them, which is the intended failure. --restore-status is inert in the same window for the same reason: every write it makes is gated on platform_managed = true, so against a freshly migrated database it restores nothing.

Step 0 — flag the sandbox rows that predate the column

npx tsx --env-file=.env scripts/backfill-platform-managed-domains.ts          # dry run
npx tsx --env-file=.env scripts/backfill-platform-managed-domains.ts --apply  # writes

Postgres only — no AWS credentials, no SES call. A row is flagged when all three hold, and reported for a human when the first holds and either of the others does not:

  1. the name is {label}.{sandbox zone}, across all three zones the sandbox has ever had (sandbox.sendoka.com, sandbox.chaparly.dev, sandbox.mrsendo.dev — nothing rewrote domains.domain on a rebrand, so the oldest orgs still carry the oldest zone);
  2. label is the owning org's slug, which is the expression both signup paths actually evaluate. A customer cannot pass it: their org's signup row already holds that exact name, so both add-domain routes answered 409 DOMAIN_ALREADY_ADDED to it — and they now refuse every name under sendoka.com outright with 422 DOMAIN_RESERVED;
  3. dkim_tokens IS NULL — both add-domain routes store what createDomainIdentity returned ([], never null, when SES returns none) and nothing else has ever written that column, so NULL means no identity was ever minted for the row. That is the definition of a platform-managed one.

All three, because the wrong direction is quiet and expensive: flagging a real customer row true removes it from the verify-poll cohort permanently, so a later DKIM revocation on it is never detected and its owner goes on reading verified.

Matching on the hostname is allowed here and nowhere else. The runtime rule is a column precisely because the zone has been renamed twice and a string that carries a semantic stops carrying it at the next rename — but that argument is about consumers, which run forever against rows that do not exist yet. This is a one-shot over a frozen set of rows that already exist, read by a human before it writes, and rows written after the deploy already carry true from signup.

status and verified_at are deliberately not part of the match: the recheck has already flipped eleven of these rows, and gating on the state the bug destroyed would skip exactly the rows this exists for. Flagging a flipped row stops the next flip; it does not undo the last one. That is step 2.

No ordering constraint against the deploy — run it before and the still-old cron keeps flipping a flagged row, which is where we already are; run it after and the row is excluded from the next poll.

Step 1 — the region, dry run first

scripts/repair-domain-regions.ts is dry run by default: it prints every row it would change and writes nothing.

npx tsx --env-file=.env scripts/repair-domain-regions.ts

It does not guess the target region — it asks SES. What it asks depends on who owns the row:

  • A customer row (platform_managed = false) is asked about by exact name, in every plausible region (ALLOWED_REGIONS, plus AWS_REGION and the row's own value when those sit outside that list).
  • A platform row (platform_managed = true) has no identity of its own to find, so the suffix chain is walked instead — {slug}.sandbox.sendoka.com, then sandbox.sendoka.com, then sendoka.com — inside each candidate region, stopping at the nearest ancestor that region holds. That is SES's own resolution order: the most specific matching identity governs the send. Only a verified ancestor is allowed to name a region.

Read-only at AWS in every mode — the one call it can make is GetEmailIdentity; --apply changes what happens in Postgres and nothing else. Answers are memoised per {region, identity} for the run, so the eleven rows ask each region about the shared parent once rather than eleven times into a bucket throttled at roughly one request per second.

Read the output, then apply the region half:

npx tsx --env-file=.env scripts/repair-domain-regions.ts --apply

Idempotent and resumable: a second run finds the rows it corrected already correct and writes nothing. Exit code is 0 when every row resolved and nothing is left waiting on a human, 1 otherwise — including while flipped rows still await --restore-status.

Step 2 — then, and only then, restore the flipped rows

Once the cron's platform_managed exclusion is deployed:

npx tsx --env-file=.env scripts/repair-domain-regions.ts \
  --restore-status --cron-fix-deployed --apply

That sets status = 'verified' and verified_at = created_at on a platform-managed row that this run has just proved is covered by a verified parent identity and that is currently sitting at pending / verified_at IS NULL. Nothing else qualifies: a platform row that has simply never been verified has no prior state to put back, and is left alone.

created_at rather than now() because signup writes the sandbox row's verified_at and created_at in one statement — they were the same instant, and created_at is the one the flip could not reach. With the cron excluding these rows, nothing schedules off the column for them any more, so the value can be the faithful one.

--restore-status without --apply previews the restores and requires no attestation. The gate is on the write.

What it will not write, and what to do about it

Every verdict below is reported and left alone. A row is only written when SES gave a complete answer that names exactly one region.

Customer-owned rows

Reported What it means What to do
found in no region No identity by that name anywhere, and no parent identity covers it either — the script re-asks up the chain before saying this. Deleted at AWS, or in a region this deployment has never named. Re-add the domain (mints a fresh identity) or delete the row. Do not hand-write a region to make it look fixed — that just moves the NotFoundException somewhere harder to spot.
covered by a parent identity, platform_managed is false No identity of its own, but a verified ancestor covers it — so "the identity was deleted" is false here and the old advice would have been destructive. Decide which it is. A sandbox row that predates the column needs platform_managed = true — that is step 0, and it will have flagged the row already if the slug and dkim_tokens corroborate. A customer subdomain riding a parent identity in our account is a different conversation. The script will not flag a row on its own.

Platform-managed rows

Reported What it means What to do
covered by a parent identity in R The governing ancestor is verified in exactly one region. Nothing — this is the good outcome. If the row names another region, --apply corrects it: a live send from the sandbox address builds its SES client from this column, so until it names R the send is rejected.
no ancestor identity at all Not one link of the chain exists in any region. An alarm, not paperwork. Every org's sandbox address is dark, not just this row's, and the product is still selling "skip DNS — offer the sandbox". Re-create the parent identity at SES, then re-run. Do not delete the row.
parent verified nowhere The ancestor identity exists but has not completed verification, so it authorises nothing. Same blast radius, different fix: finish the parent's DKIM/DNS at SES, then re-run. Pointing the row at an unverified identity would leave a repaired-looking column and a send that still fails, with MessageRejected instead of a NotFoundException you would recognise. Do not delete the row.
flipped, awaiting restore Covered, but sitting at pending / verified_at IS NULL — the recheck got to it. Deploy the cron exclusion, then the --restore-status run above.

Either kind

Reported What it means What to do
found in several regions A real identity — or, for a platform row, a verified ancestor — in more than one region. Created by hand, or left behind by a region move. Decide which one is the live sender (check DNS and which is verified — the script prints verified and DKIM status per region), set region by hand. For a customer row, delete the spare at AWS; picking wrong re-creates the delete bug on purpose. For a platform row the spare is the shared parent identity, so deleting one is a decision about every org's sandbox at once — set the column and leave AWS alone.
probe incomplete At least one region did not answer — throttle, timeout, AccessDenied. Re-run. A non-answer is not an absence, so the row is refused rather than written on partial evidence. If every row skips with the same regions erroring, the IAM key is missing ses:GetEmailIdentity there (see scripts/sendoka-app-runtime-policy.json).

After

No redeploy. The domains.platform_managed migration is a prerequisite, not a follow-up, and the schema default for region stays dropped — it is no longer reachable from any of the four insert paths, which src/app/api/domain-create-region.test.ts and src/lib/auth/signup-sandbox-domain.test.ts assert between them.

Re-run the dry pass afterwards. It should exit 0.