Runbook — repair domains rows against the SES identity that actually covers them
Three steps, two scripts, in a fixed order:
platform_managed— flag the sandbox rows that already existed when the column was added. Everything below keys off that column, and the migration'sDEFAULT falsestamped it wrong on exactly the rows that matter. Its own script:scripts/backfill-platform-managed-domains.ts.region— rows written before the create-path fix name a region their SES identity was never in. Required on any deployment whoseAWS_REGIONis notus-east-1. This deployment runsAWS_REGION=ap-southeast-2, so it needs it.status/verified_at— the eleven*.sandbox.sendoka.comrows the daily recheck unverified. This half is ordered after a deploy (see below) and lives behind its own flag.
⚠️ Ordering — read this before you run anything
--restore-statusmust not run until theplatform_managedexclusion indomain-verify-pollis live in production.Restoring
status = 'verified'while the deployed cron still reconciles platform-managed rows against SES does not fail — it re-arms the flip. The row goes back topendingon the next daily recheck and the customer receives a seconddomain.unverifiedwebhook about a domain that never stopped working.The script enforces what it can:
--restore-status --applyrefuses unless you pass--cron-fix-deployedand the cron route in your working tree mentions the column. The first is your attestation about production — nothing running on your laptop can check that. The second can only prove the fix is absent.The
regionhalf has no such constraint. Run it whenever.
⚠️ Dev and prod share one Neon database. There is no staging copy to rehearse against — whatever
DATABASE_URLresolves to is production. Run the dry pass and read it before--apply.
Why — part one, the region
createDomainIdentity has always created the SES identity in AWS_REGION. Until the fix, both add-domain routes then inserted a domains row with no region, so Postgres filled in the schema column default "us-east-1" (dropped in migration 0053, once every insert path supplied the column). Every row written before the fix therefore names a region its identity was never in, and every later operation keys off the row, not off the env:
domain-verify-pollcallscheckDomainStatus(d.domain, d.region), getsNotFoundExceptionfromus-east-1, and substitutesverified: false. The "identity missing" early return is guarded on!wasVerified, so a row that is already verified skips it, falls through tojustUnverified, flips topendingand fans out adomain.unverifiedwebhook at the customer — for a domain that is verified and sending.- Deleting a domain (
internal/domains) callsdeleteDomainIdentity(record.domain, record.region), so the delete goes tous-east-1, finds nothing, and the realap-southeast-2identity survives and keeps sending. The row is gone by then, so nothing in the product mentions that identity again.
The fix stops new rows being written wrong. It does not touch the rows that already are — nothing self-heals.
Why — part two, the sandbox rows
Every org gets a {org-slug}.sandbox.sendoka.com row at signup, written verified by hand in the same transaction that creates the org. Neither signup path creates an SES identity for it, deliberately. The row rides the parent sendoka.com identity, which SES authorises for every subdomain — that is what lets a trial account send with no DNS at all, which the product offers in writing. The parent identity is a real operational dependency and these rows have to end up usable for live sends, not merely be kept quiet.
GetEmailIdentity is an exact-name lookup. It does not resolve up the suffix chain, so it misses on a subdomain a parent identity covers, and that miss is indistinguishable from an identity somebody deleted. Two things read it as deletion:
- The cron, once
NotFoundExceptionwas correctly reclassified from a transient error into a definitive answer (5f59cde, in prod asb5136b5). Eleven rows are already flipped topendingwithverified_atcleared, each with a customer-facingdomain.unverifiedwebhook carrying adkim_statusofIDENTITY_NOT_FOUND— a value we synthesise, not one SES returned. Fixed by theplatform_managedexclusion in the cohort select. - This script, which parked those rows in "found in no region" and told you to "re-add the domain to mint a fresh one, or delete the row". Both halves were actively harmful and both are gone: nothing re-creates a sandbox row (the insert lives only in the new-org branch of signup), so deleting it permanently removes that org's sandbox sender, and re-adding through
POST /v1/domainsanswers422 DOMAIN_RESERVED(no name undersendoka.comcan be added by a customer).
The column, not a hostname suffix match: the sandbox domain has been renamed twice already (mrsendo.dev → chaparly.dev → sendoka.com), and a string that carries a semantic stops carrying it the next time someone renames it.
Prerequisites
| Needed for | Checked by the script | |
|---|---|---|
DATABASE_URL, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY |
steps 1 and 2 — step 0 touches Postgres only and needs just DATABASE_URL |
refuses without them |
the domains.platform_managed migration applied (0054) |
everything | refuses without the column |
platform_managed backfilled on the existing sandbox rows (step 0) |
the sandbox half | not directly — but a covered row still flagged false is reported as its own verdict rather than written or condemned |
| the cron exclusion deployed | --restore-status --apply only |
--cron-fix-deployed + a grep of the cron route |
ADD COLUMN … DEFAULT false stamps false on every existing row, sandbox rows included. Until step 0 runs, the eleven are still indistinguishable from customer rows by the column — they will surface as covered by a parent identity … platform_managed is false and this script will refuse to touch them, which is the intended failure. --restore-status is inert in the same window for the same reason: every write it makes is gated on platform_managed = true, so against a freshly migrated database it restores nothing.
Step 0 — flag the sandbox rows that predate the column
npx tsx --env-file=.env scripts/backfill-platform-managed-domains.ts # dry run
npx tsx --env-file=.env scripts/backfill-platform-managed-domains.ts --apply # writes
Postgres only — no AWS credentials, no SES call. A row is flagged when all three hold, and reported for a human when the first holds and either of the others does not:
- the name is
{label}.{sandbox zone}, across all three zones the sandbox has ever had (sandbox.sendoka.com,sandbox.chaparly.dev,sandbox.mrsendo.dev— nothing rewrotedomains.domainon a rebrand, so the oldest orgs still carry the oldest zone); labelis the owning org's slug, which is the expression both signup paths actually evaluate. A customer cannot pass it: their org's signup row already holds that exact name, so both add-domain routes answered409 DOMAIN_ALREADY_ADDEDto it — and they now refuse every name undersendoka.comoutright with422 DOMAIN_RESERVED;dkim_tokens IS NULL— both add-domain routes store whatcreateDomainIdentityreturned ([], never null, when SES returns none) and nothing else has ever written that column, so NULL means no identity was ever minted for the row. That is the definition of a platform-managed one.
All three, because the wrong direction is quiet and expensive: flagging a real customer row true removes it from the verify-poll cohort permanently, so a later DKIM revocation on it is never detected and its owner goes on reading verified.
Matching on the hostname is allowed here and nowhere else. The runtime rule is a column precisely because the zone has been renamed twice and a string that carries a semantic stops carrying it at the next rename — but that argument is about consumers, which run forever against rows that do not exist yet. This is a one-shot over a frozen set of rows that already exist, read by a human before it writes, and rows written after the deploy already carry true from signup.
status and verified_at are deliberately not part of the match: the recheck has already flipped eleven of these rows, and gating on the state the bug destroyed would skip exactly the rows this exists for. Flagging a flipped row stops the next flip; it does not undo the last one. That is step 2.
No ordering constraint against the deploy — run it before and the still-old cron keeps flipping a flagged row, which is where we already are; run it after and the row is excluded from the next poll.
Step 1 — the region, dry run first
scripts/repair-domain-regions.ts is dry run by default: it prints every row it would change and writes nothing.
npx tsx --env-file=.env scripts/repair-domain-regions.ts
It does not guess the target region — it asks SES. What it asks depends on who owns the row:
- A customer row (
platform_managed = false) is asked about by exact name, in every plausible region (ALLOWED_REGIONS, plusAWS_REGIONand the row's own value when those sit outside that list). - A platform row (
platform_managed = true) has no identity of its own to find, so the suffix chain is walked instead —{slug}.sandbox.sendoka.com, thensandbox.sendoka.com, thensendoka.com— inside each candidate region, stopping at the nearest ancestor that region holds. That is SES's own resolution order: the most specific matching identity governs the send. Only a verified ancestor is allowed to name a region.
Read-only at AWS in every mode — the one call it can make is GetEmailIdentity; --apply changes what happens in Postgres and nothing else. Answers are memoised per {region, identity} for the run, so the eleven rows ask each region about the shared parent once rather than eleven times into a bucket throttled at roughly one request per second.
Read the output, then apply the region half:
npx tsx --env-file=.env scripts/repair-domain-regions.ts --apply
Idempotent and resumable: a second run finds the rows it corrected already correct and writes nothing. Exit code is 0 when every row resolved and nothing is left waiting on a human, 1 otherwise — including while flipped rows still await --restore-status.
Step 2 — then, and only then, restore the flipped rows
Once the cron's platform_managed exclusion is deployed:
npx tsx --env-file=.env scripts/repair-domain-regions.ts \
--restore-status --cron-fix-deployed --apply
That sets status = 'verified' and verified_at = created_at on a platform-managed row that this run has just proved is covered by a verified parent identity and that is currently sitting at pending / verified_at IS NULL. Nothing else qualifies: a platform row that has simply never been verified has no prior state to put back, and is left alone.
created_at rather than now() because signup writes the sandbox row's verified_at and created_at in one statement — they were the same instant, and created_at is the one the flip could not reach. With the cron excluding these rows, nothing schedules off the column for them any more, so the value can be the faithful one.
--restore-status without --apply previews the restores and requires no attestation. The gate is on the write.
What it will not write, and what to do about it
Every verdict below is reported and left alone. A row is only written when SES gave a complete answer that names exactly one region.
Customer-owned rows
| Reported | What it means | What to do |
|---|---|---|
| found in no region | No identity by that name anywhere, and no parent identity covers it either — the script re-asks up the chain before saying this. Deleted at AWS, or in a region this deployment has never named. | Re-add the domain (mints a fresh identity) or delete the row. Do not hand-write a region to make it look fixed — that just moves the NotFoundException somewhere harder to spot. |
covered by a parent identity, platform_managed is false |
No identity of its own, but a verified ancestor covers it — so "the identity was deleted" is false here and the old advice would have been destructive. | Decide which it is. A sandbox row that predates the column needs platform_managed = true — that is step 0, and it will have flagged the row already if the slug and dkim_tokens corroborate. A customer subdomain riding a parent identity in our account is a different conversation. The script will not flag a row on its own. |
Platform-managed rows
| Reported | What it means | What to do |
|---|---|---|
covered by a parent identity in R |
The governing ancestor is verified in exactly one region. | Nothing — this is the good outcome. If the row names another region, --apply corrects it: a live send from the sandbox address builds its SES client from this column, so until it names R the send is rejected. |
| no ancestor identity at all | Not one link of the chain exists in any region. | An alarm, not paperwork. Every org's sandbox address is dark, not just this row's, and the product is still selling "skip DNS — offer the sandbox". Re-create the parent identity at SES, then re-run. Do not delete the row. |
| parent verified nowhere | The ancestor identity exists but has not completed verification, so it authorises nothing. | Same blast radius, different fix: finish the parent's DKIM/DNS at SES, then re-run. Pointing the row at an unverified identity would leave a repaired-looking column and a send that still fails, with MessageRejected instead of a NotFoundException you would recognise. Do not delete the row. |
| flipped, awaiting restore | Covered, but sitting at pending / verified_at IS NULL — the recheck got to it. |
Deploy the cron exclusion, then the --restore-status run above. |
Either kind
| Reported | What it means | What to do |
|---|---|---|
| found in several regions | A real identity — or, for a platform row, a verified ancestor — in more than one region. Created by hand, or left behind by a region move. | Decide which one is the live sender (check DNS and which is verified — the script prints verified and DKIM status per region), set region by hand. For a customer row, delete the spare at AWS; picking wrong re-creates the delete bug on purpose. For a platform row the spare is the shared parent identity, so deleting one is a decision about every org's sandbox at once — set the column and leave AWS alone. |
| probe incomplete | At least one region did not answer — throttle, timeout, AccessDenied. |
Re-run. A non-answer is not an absence, so the row is refused rather than written on partial evidence. If every row skips with the same regions erroring, the IAM key is missing ses:GetEmailIdentity there (see scripts/sendoka-app-runtime-policy.json). |
After
No redeploy. The domains.platform_managed migration is a prerequisite, not a follow-up, and the schema default for region stays dropped — it is no longer reachable from any of the four insert paths, which src/app/api/domain-create-region.test.ts and src/lib/auth/signup-sandbox-domain.test.ts assert between them.
Re-run the dry pass afterwards. It should exit 0.