# KEEPER: the operator runbook Written 2026-10-01, rewritten 2026-10-02 for the public keeper. The keeper is `keeper/src/`, a permissionless crank runner for bundle-vault (Sheaf). It is a convenience, not a privilege: every instruction it sends is callable by any signer, its key pays fees and rent and nothing else, and the redemption path never touches it. If it stops, the vault keeps its funds and its rules; trades and fee harvests simply wait until someone (this keeper, a copy of it, anyone) runs the next cycle. Sheaf is a launchpad (D-19): any creator launches a self-funded bundle. Without a running keeper a creator's bundle never migrates, never clears its tier and never cranks. So by default the keeper is a **public keeper**: it finds every bundle of the program on its own and serves each one. Proven on a mainnet fork by `tests/e2e/keeper-discovery.sh` (discovery mode, three creators) and `tests/e2e/keeper.sh` (allowlist mode, one bundle). Results: the two "Fork proof" sections below, `tests/e2e/results/keeper-discovery.json` and `tests/e2e/results/keeper.json`. ## Two modes | Mode | When | What is served | |---|---|---| | **discovery** (default) | No `--bundle`, `BUNDLES` or `BUNDLE` | Every Bundle account of the program that is Launched, found by a scan every `DISCOVERY_INTERVAL_S` (60 s). A new bundle joins the running process; no restart | | **allowlist** | `--bundle ADDR`, `BUNDLES=a,b,...` or `BUNDLE=a` | Exactly those bundles, in any state, as before. No scan. Overrides discovery | ## Discovery `keeper/src/discovery.ts`. One `getProgramAccounts` on `9iihm312JFeK6uNoxcLRxz8DCMK7Sh1ZbtCJz18iWrX9`, filtered on the server to the Bundle size (480 bytes) and a memcmp on the 8-byte Anchor Bundle discriminator at offset 0. Each result is re-derived as `["bundle", creator, slot]` from the creator and slot its own bytes record, and dropped unless that is its address. Then, from the account bytes alone, with no further read: | Class | Served | Why | |---|---|---| | `Launched` | yes | migrate, the pre-creates, its table, clear_tier, cranks and harvests are all keeper work | | `Funding` | no | Until `execute_launch` nothing is a keeper's to do. Re-checked every scan, so it is served within one interval of its launch | | `FundingPastDeadline` | no | The only action left is the creator's own refund (`scripts/refund-bundle.ts`) | | `Failed` | no | Nothing to do | | `NotBundlePda` | never | Bundle-shaped, owned by the program, but not at its own PDA | Each class is logged once per bundle (`discovery` lines), and again only when it changes. **An RPC that refuses `getProgramAccounts`.** Many free plans do. A refusal (any answer other than a rate limit or a timeout) is logged as an `alert` line, `reason: GetProgramAccountsRefused`, with the RPC's own error text, and the keeper serves a fallback: `FALLBACK_BUNDLES` plus every bundle in its own lookup-table state file. New bundles are not found until the RPC allows the call, and the alert repeats every scan. A transient scan failure (429, timeout) keeps the served set as it is and retries next interval. Paid plans at Helius, Triton and QuickNode serve the call. ## What one cycle does Each cycle reads chain state once (the clock, bundle, strategy, pool, curve, fee accounts, Pyth account and the two venue configs, in one `getMultipleAccountsInfo` unless `RPC_BATCH_READS=false`), then runs these steps in order. Each step rebuilds its transaction from that read, simulates it exactly as it would be sent, and sends only on a passing simulation. | # | Action | When | What is sent | |---|---|---|---| | 1 | `migrate` | Bundle Launched, curve complete, no PumpSwap pool | pump's own permissionless `migrate`, its own transaction (RAILS D9). The keeper pays the 2,039,280-lamport boost-vault rent | | 2 | `precreate` | Pool exists and either account is missing | One transaction: the idempotent ATA create of the AMM creator-fee vault, then PumpSwap `init_user_volume_accumulator` for user = Bundle PDA, only if that account is absent (RAILS D6, E2E step 4 order) | | 2b | table | The bundle first needs its keeper lookup table (venue ready, and the tier not cleared or the strategy on) and none is configured | Create plus the first 20 keys in one transaction, then the other 16, then wait one slot for activation. Never a second table for a bundle ("Lookup tables" below) | | 3 | `clear_tier` | `tier_cleared` false and the venue is ready | `clear_tier` through the keeper lookup table, with the trade marker (D-13) | | 4 | `crank` | Oracle fresh and something is due (below) | The client's D-13 helper simulates without the marker; success means a non-trade (observation, reference step, halt), sent unmarked; error 3005 means a trade, re-simulated with the marker, sent marked only if that passes (otherwise `SellDeferred`); any other code is a no-op, logged and never sent. Up to `MAX_CRANK_ACTIONS_PER_CYCLE` per cycle, because an observation, a reference step and a sell can all fall due at once | | 5a | `collect_curve_fee` | pump's creator vault for `fee_auth` holds lamports above rent, and they clear the threshold | pump's permissionless `collect_creator_fee_v2(creator = fee_auth)`: moves the curve-phase creator fees onto `fee_auth` (RAILS D8) | | 5b | `harvest` | AMM creator vault + `fee_auth` lamports above rent + fee inbox clear the threshold | `harvest_fees`, through the table. If the protocol fee account (D-19, the wSOL ATA of `PROTOCOL_FEE_OWNER`) is missing, an idempotent ATA create rides in front, because a nonzero cut cannot be paid to a closed account | **When a crank is due**, from the strategy's own intervals (`strategy.rs` `next_action`): an observation when `now - obs_ts[obs_head] >= obs_interval_s`; a reference step when `consec_above_ratchet >= ratchet_confirm_obs` and `now - ref_updated_ts >= ratchet_min_interval_s`; a trade window when `now >= last_trade_ts + min_trade_interval_s` and `now >= halted_until_ts`. Before any of those, the keeper does not even simulate. Inside a trade window it simulates every cycle, because a third party's swap can move the price into the sell band at any moment. The program makes every decision; the keeper only avoids asking when the answer is known. **The oracle.** The crank passes the Pyth SOL/USD `PriceUpdateV2` account (default: the sponsored push feed `7UVimffxr9ow1uXYxsr4LHAcV58mLzhmwaeKvJ1pjLiE`; the program checks owner, discriminator and feed id, not the address). The keeper decodes it with `client/pyth.ts`, a byte-for-byte port of `oracle.rs`, and skips the crank when `|chain clock - publish_time| > oracle_max_age_s` (60 s on bundle #1), the exact rule the program would fail on with 6045. **It does not post price updates itself.** Posting means fetching a signed update from Hermes and verifying it through the Pyth receiver and Wormhole, which is neither trivial nor needed while the sponsored feed is live; when it is stale the crank waits, which is the program's own fail-closed behavior. **Harvest economics.** A harvest is sent only when the pending lamports are at least `HARVEST_MIN_LAMPORTS` AND at least `HARVEST_FEE_MULTIPLE` times the transaction's own fee (5,000 lamports plus priority fee times the CU limit). The fees never leave the vault's control while they wait: the AMM creator vault and `fee_auth` are both the bundle's. ## Scheduling across many bundles One loop (`keeper/src/index.ts`). Each served bundle has its own due time, set after every cycle by `Keeper.schedule` from that bundle's last read and the program's own intervals: | After a cycle that | Next cycle in | |---|---| | sent something | `CYCLE_INTERVAL_MS` (more may follow) | | found a trade window open | `CYCLE_INTERVAL_MS`: a third party's swap can move the price into the sell band at any moment | | found the oracle stale, or the table not ready | `CYCLE_INTERVAL_MS` | | found nothing due | the time to the next observation, reference step or trade window, clamped to [`CYCLE_INTERVAL_MS`, `MAX_IDLE_MS`] | | threw (an RPC failure) | `CYCLE_INTERVAL_MS` doubled per consecutive failure, at most `MAX_IDLE_MS`; only that bundle backs off | `MAX_IDLE_MS` (60 s) bounds the wait even when nothing is time-gated, so pending fees, clock drift and anything a third party did are still noticed. At most `MAX_CONCURRENT_BUNDLES` cycles run at once, and at most `MAX_CONCURRENT_SENDS` transactions are in flight across them. A cycle that throws ends that bundle's cycle only; every other bundle keeps its time. **Rate limits.** Every RPC request goes through one wrapper (`keeper/src/rpc.ts`). An HTTP 429 pauses every request from the process (the limit belongs to the API key, not to a bundle), for the `Retry-After` time or 0.5 s doubling per consecutive 429, at most 30 s, then retries the request; each is logged as `rpc` `RateLimited`. Every request is also aborted after `RPC_TIMEOUT_MS` (30 s), so a hung socket costs one bundle one cycle, never a concurrency slot forever. A timeout does not mean the transaction was not processed: twice on the fork a `sendTransaction` that timed out at 30 s landed afterwards. So when the broadcast call itself fails transiently, the keeper watches the signature (known from the signed bytes) for 10 s, rebroadcasting the same bytes, and reports `sent` if it lands; otherwise the line is `rpc_error` with that signature, and the retry decides again from a fresh read. It is never re-signed blindly. **A public keeper is not the only sender.** pump's own migrator, another copy of this keeper or anyone can land the same action between our simulation and our send. When a send fails and the re-read shows the action done (the pool exists, the tier is cleared, the strategy moved), it is logged as `skipped` with `raced: true` (`AlreadyMigrated`, `TierAlreadyCleared`, `RacedByAnotherCranker`, ...), never as `failed`. A crank the program refuses at preflight with 6035 `NoActionAvailable` means its own next action changed between our simulation and our send (time passed and an observation fell due ahead of a sell): nothing landed, nothing was paid, and it is logged as `skipped` `ActionChangedBeforeSend` and decided again next cycle. A crank refused at preflight with 6045 `StaleOracle` means the Pyth update aged past `oracle_max_age_s` in that same gap: logged as `skipped` `StaleOracleBeforeSend`, also decided again next cycle. (Found in the lead's discovery rerun on 2026-10-03, where it had been logged as `failed`.) ## Lookup tables, one per bundle `clear_tier` serializes at 1,249 B and a marked crank at 1,282 B without the keeper table, both over the 1,232 B limit, so no bundle is serviceable without one (`docs/SECURITY-REVIEW.md`, D-19 fixes). The keeper makes its own, one per bundle, the first time the bundle needs it (`keeper/src/tables.ts`), and pays its rent. The table holds the 36 keys of `client/alt.ts` `keeperTableAddresses`, with the bundle's address at entry 2. **The state file.** `STATE_DIR/tables-.json` maps bundle to table. It is named for the key, so a rotated key starts clean and two keepers can share a folder. Writes are atomic. **No duplicate, by construction:** 1. A table's address is a function of (authority, `recent_slot`). The keeper picks the slot, derives the address and writes it to the state file as `pending` *before* broadcasting. 2. The create transaction carries the first 20 keys too, so any table that exists already names its bundle at entry 2. A crash between the create and the second extend leaves a table that is merely short of keys; the restart extends it. 3. A pending entry is resolved only from chain: the account exists, so it is adopted; or it is absent and its `recent_slot` is more than 600 slots old. The lookup-table program only accepts a `recent_slot` still in SlotHashes (512 slots), so after that the create can never land, and only then is a new table made. 4. A send that did not confirm (expired, timed out, RPC error) is never re-signed. The entry stays pending and rule 3 decides on a later cycle. **Verified at every start, before use.** For each entry: the account exists, its authority is the keeper key, it is not deactivated, entry 2 is the bundle, and it holds every key the bundle needs (any missing key is extended in, never a new table). A table that fails any check is logged as `TableNotUsable` and dropped from the file. **A lost or corrupt state file.** A file that does not parse, or names another key or program, is moved aside (`.corrupt-`). Then the keeper rebuilds the map from chain before creating anything: every table whose authority is its key, by `getProgramAccounts` on the lookup-table program with a memcmp on the authority (offset 22); if the RPC refuses that, by walking up to `TABLE_SCAN_MAX_SIGNATURES` of the key's own signatures for table creates and table lookups. Each table is matched to its bundle by entry 2; with two for one bundle (impossible under rules 1 to 4, but possible from an older version) the fuller one wins and the other is logged. If the scan itself fails, the keeper retries five times and then refuses to run, because a new table could duplicate one that exists. ## Money: the float and the spend ledger The keeper pays every fee and every rent it causes from its own balance, the **float**: migrate's boost vault, the two pre-creates, the bundle's lookup table, and the fee of every transaction. Nothing it pays comes back to it. The amounts are measured in "Costs" below. **The minimum-float guard** (`keeper/src/float.ts`). Below `MIN_FLOAT_LAMPORTS` the keeper sends nothing new. Cycles keep running and logging (each send it would have made is a `skipped` line with `reason: FloatBelowMinimum`), the health document turns `float_ok: false` and HTTP 503, and an `alert` line, `alert: FLOAT_BELOW_MINIMUM`, names the balance, the floor and the key to fund, once a minute until the balance is back above the floor (then `float` `FloatRestored`). The guard checks a balance at most 3 s old, lowered by what this process has spent since; with sends in flight it can undershoot by at most `MAX_CONCURRENT_SENDS` transactions (the largest is migrate, about 0.0021 SOL). **Spend per bundle per day.** Every landed transaction's cost (the key's balance drop in that transaction's own metadata, so fee plus any rent) is logged as `spent_lamports` on its line and added to `STATE_DIR/spend-.json`: per UTC day, per bundle, per action, kept 62 days, written on every record. At UTC midnight and at shutdown a `spend` line per bundle gives the totals. The health document carries today's total. ## Health One JSON document (`keeper/src/health.ts`), two ways: `GET /health` on `HEALTH_HOST:HEALTH_PORT` (set `HEALTH_PORT` or pass `--health PORT`; the host defaults to 127.0.0.1), and the heartbeat file `HEARTBEAT_FILE` (default `STATE_DIR/heartbeat.json`), rewritten every 5 s. HTTP 200 when healthy, 503 when not. Healthy means: a cycle finished within max(3 x `MAX_IDLE_MS`, 120 s), the float is at or above the floor, and (allowlist mode, or the last scan succeeded, or a fallback is served). `problems` lists what is wrong. | Field | Meaning | |---|---| | `ok`, `problems` | The verdict above, and why not | | `last_cycle_at`, `last_cycle_age_s`, `cycles_total`, `in_flight` | The loop is alive | | `bundles_served`, `bundles_known`, `bundles_waiting` | Served bundles, and the others by class (`Funding`, ...) | | `float_lamports`, `min_float_lamports`, `float_ok` | The float | | `errors_last_hour`, `alerts_last_hour`, `last_error` | Counted from the log lines themselves, so the two never disagree | | `rpc` | Requests, 429s in the last hour, timeouts | | `discovery` | Last successful scan, last error, whether the RPC refuses, fallback size | | `spend_today_lamports` | Today's spend, total and per bundle | ## Running it From `keeper/` (it uses the repo root's `node_modules` and `client/`; `npm install` at the root): ```bash cp .env.example .env # then edit; .env is gitignored at the repo root node ../node_modules/ts-node/dist/bin.js src/index.ts --help node ../node_modules/ts-node/dist/bin.js src/index.ts --dry-run --once # rehearse node ../node_modules/ts-node/dist/bin.js src/index.ts --health 8935 # run (discovery mode) ``` Discovery mode needs no table setup: the keeper makes each bundle's table itself. In allowlist mode a table can still be named (`LOOKUP_TABLE`, built once with `--create-lookup-table`); a listed bundle without one gets an automatic table like any discovered bundle. `npm run start`, `npm run dry-run` and `npm run check` (type-check) wrap the same commands. Set `TS_NODE_TRANSPILE_ONLY=1` for a faster start; `npm run check` is the type gate. | Flag | Meaning | |---|---| | `--env FILE` | Config file. Default `keeper/.env`. Process environment variables override it | | `--dry-run` | Simulate every action and log `simulated`; send nothing, create no table. Each step sees the chain as it is, so after a simulated `migrate` the later steps report what blocks them (`NoPool`) | | `--once` | One cycle of every served bundle (in discovery mode: after one scan), then exit 0 | | `--bundle ADDR` | Allowlist mode with this one bundle | | `--health PORT` | Serve `GET /health` on `HEALTH_HOST:PORT` | | `--create-lookup-table` | Allowlist mode only: build the 36-key keeper table for each listed bundle, print its address, exit. Needs the bundle Launched | | `--i-mean-a-live-cluster` | Required for any `RPC_URL` that is not `127.0.0.1` or `localhost`, including a dry run. Same rule as `scripts/launch-bundle.ts` | | `--crash-after-send ACTION` | Test hook, refused off loopback: exit 86 right after broadcasting ACTION, before confirming it | `SIGINT` or `SIGTERM`: the keeper stops starting cycles, finishes every action in flight (confirmation included), writes the day's `spend` lines and a last heartbeat, logs `shutdown` with `clean exit`, and exits 0. A second signal exits at once with 130. ## Every config value | Key | Default | Meaning | |---|---|---| | `RPC_URL` | required | JSON-RPC endpoint. Non-loopback needs `--i-mean-a-live-cluster`. Discovery needs one that serves `getProgramAccounts`. Logged with any API key removed | | `KEYPAIR_PATH` | required | Fee payer keypair JSON (`KEEPER_KEYPAIR_PATH` is accepted too). A **file path only**: a value that looks like key material is refused. Relative paths resolve from the env file's folder. A file readable by other users draws a warning. Use a dedicated low-balance key, never one from `keys/` | | `PROGRAM_ID` | `9iihm312...rX9` | Must equal the id `client/` is built for; anything else is refused | | `--bundle ADDR`, `BUNDLES` or `BUNDLE` | none (discovery) | Allowlist mode. D-19: a slot alone names nothing, so `BUNDLE_SLOT` is refused. Each is checked to be `["bundle", creator, slot]` of its own bytes | | `DISCOVERY_INTERVAL_S` | 60 | Seconds between scans | | `FALLBACK_BUNDLES` | none | Served, with the state file's bundles, when the RPC refuses `getProgramAccounts` | | `STATE_DIR` | `keeper/state` | Table map, spend ledger, default heartbeat. Gitignored | | `AUTO_LOOKUP_TABLES` | true | Make and verify one table per bundle when none is named | | `LOOKUP_TABLES`, `LOOKUP_TABLE` | none | Allowlist mode only: named tables, in `BUNDLES` order. Refused in discovery mode | | `TABLE_SCAN_MAX_SIGNATURES` | 5000 | Table recovery when `getProgramAccounts` is refused: signatures of the key to walk | | `MIN_FLOAT_LAMPORTS` | 50000000 | 0.05 SOL. Below it nothing new is sent ("Money"). Size it from "Float sizing" | | `MAX_CONCURRENT_BUNDLES` | 4 | Cycles at once | | `MAX_CONCURRENT_SENDS` | `MAX_CONCURRENT_BUNDLES` | Transactions in flight at once across all bundles. 1 on a Surfpool fork | | `CYCLE_INTERVAL_MS` | 15000 | Shortest time between two cycles of one bundle, and the poll rate inside a trade window | | `MAX_IDLE_MS` | 60000 | Longest time between two cycles of one bundle | | `RPC_TIMEOUT_MS` | 30000 | Any single RPC request is aborted after this | | `HEALTH_PORT`, `HEALTH_HOST` | off, 127.0.0.1 | The health endpoint | | `HEARTBEAT_FILE` | `STATE_DIR/heartbeat.json` | `off` disables it | | `PYTH_PRICE_UPDATE` | sponsored SOL/USD | The `PriceUpdateV2` account the crank passes | | `PRIORITY_FEE_MICROLAMPORTS` | 0 | Adds `SetComputeUnitPrice`; 0 adds nothing | | `HARVEST_MIN_LAMPORTS` | 50000000 | 0.05 SOL. Applies to `collect_curve_fee` and `harvest` | | `HARVEST_FEE_MULTIPLE` | 20 | Pending must also exceed this many times the transaction fee | | `MAX_CRANK_ACTIONS_PER_CYCLE` | 4 | Crank sends per cycle | | `CU_HEADROOM_PCT` | 25 | CU limit = simulated units plus this percent, at least +10,000, at most 1,400,000 | | `SEND_ATTEMPTS` | 4 | Tries per action when a send expires or the RPC fails internally; each retry re-reads state and decides again | | `CONFIRM_TIMEOUT_MS` | 60000 | Backstop on the confirmation wait; the real bound is the blockhash's `lastValidBlockHeight` | | `STARTUP_SETTLE_S` | 0 local, 90 live | Wait before the first cycle (see "Restarts") | | `RPC_BATCH_READS` | true | One `getMultipleAccountsInfo` per cycle; `false` reads each account with `getAccountInfo` (use on a Surfpool fork) | ## Log lines stdout, one JSON object per line, one line per action. Every line has `ts` (wall clock, ISO), `action` and `outcome`. bigint values are decimal strings. | `outcome` | Meaning | |---|---| | `sent` | Landed and succeeded. Carries `signature`, `slot`, `cu` (consumed), `cu_limit`, `sim_cu`, `bytes` (serialized size) and `bytes_limit` (1,232), `trace` (top-level instructions plus CPIs, limit 64), `depth` (max stack height, limit 5), `fee_lamports`, `static_keys`, `alt_keys`, `attempt` | | `skipped` | Not sent. `reason` names why (table below); `code` is the program error code when a simulation said so | | `simulated` | `--dry-run`: what would have been sent, with `sim_cu`, `cu_limit`, `bytes` | | `expired` | Blockhash expired before it landed. The keeper backs off, re-reads, decides again | | `rpc_error` | The RPC refused it for its own reasons (timeout, 5xx, a fork's failed mainnet fetch). Not processed. Retried like `expired` | | `failed` | Landed and failed, or refused at preflight by the program. Never expected: every send follows a passing simulation. Investigate | | `error` | An RPC or client exception. The cycle moves on; the next one re-reads everything | | `info` | `start`, `lookup_table` (the table's entries and any missing keys), `cycle` (per cycle: `n`, `sent`, `idle`, `chain_ts`, `ms`), `shutdown` | | `crash_injected` | Test hook only | Per action, the extra fields: `crank` lines carry `marked`, `simulated_kind` (`state` or `trade`), `due` (`observation`, `reference`, `trade window`) and `kind`: `observe`, `reference`, `halt`, `sell` or `buy`. Trade lines add `side`, `base_amount`, `quote_amount`, `venue_fee_bps` from the `VaultTrade` event. `precreate` says which of the two it created. `harvest` and `collect_curve_fee` carry `pending_lamports` and its parts. Common `skipped` reasons, all normal: | `reason` | Meaning | |---|---| | `NoBundle`, `Funding`, `Failed` | Nothing for a keeper to do in that state | | `AlreadyMigrated`, `AlreadyExist`, `TierAlreadyCleared` | Read from chain as done; never resent | | `NoPool`, `VenueNotReady`, `CurveNotComplete` | A prerequisite step has not landed yet | | `NoLookupTable`, `TxTooLarge` | Configuration: set `LOOKUP_TABLE` | | `StaleOracle` (6045), `OracleUnreadable` | The Pyth account is older than `oracle_max_age_s`, or not a SOL/USD update | | `NotDue` | No interval has elapsed; `next_due_ts` says when one will | | `CrankTooSoon` (6034), `InsideBand` (6037 family), `BelowSellBand`, `BuyDisabled`, `Halted` (6040), `BelowMinimum` (6039), ... | The program's own no-op codes from the simulation. Logged, never sent | | `BelowThreshold`, `NothingToHarvest`, `NothingToCollect` | Fees not worth a transaction yet | | `SellDeferred` | The program chose a sell but the marked sell fails in simulation, usually PumpSwap's own 6004 min-out when the M3 median floor defers it. `venue_code` is the venue's code, not this program's. Retried next cycle | Since the public keeper: every line a bundle's keeper writes carries `bundle`, and every landed transaction's line carries `spent_lamports` (the key's balance drop: fee plus rent). New outcome `alert` means an operator must act. New actions: | `action` | Lines | |---|---| | `discovery` | `Serving` (a bundle joins, with `origin`: discovery, fallback or allowlist); `skipped` per non-served class; `info` scan summaries; `error` `ScanFailed` (transient); `alert` `GetProgramAccountsRefused` | | `lookup_table` | `StateFileMissing`, `StateFileCorrupt`, `TableRecovered`, `TablePending`, `TableGone`, `TableNotUsable`, `TableReady`, `RecoveryFailed` | | `create_lookup_table`, `extend_lookup_table` | The table's own transactions, logged like any send | | `float` | `alert` `FLOAT_BELOW_MINIMUM`, `info` `FloatRestored` | | `rpc` | `error` `RateLimited` (HTTP 429, with the backoff) | | `spend` | Per bundle per day: `txs`, `lamports`, `fees`, `by_action` | | `health` | Where the endpoint listens | New `skipped` reasons: `FloatBelowMinimum`; `ActionChangedBeforeSend` (a crank refused at preflight with 6035 because the program's next action changed); and, with `raced: true`, `AlreadyMigrated`, `AlreadyExist`, `TierAlreadyCleared`, `RacedByAnotherCranker` when another sender landed first. Repeated identical skips are logged once and again only when the reason changes, so a quiet vault writes about one line per cycle (the `cycle` line). ## Restarts and idempotency The keeper keeps no decision state between cycles. Two files persist, and neither is trusted: the table map is verified on chain at every start, and the spend ledger is a record, not an input. Every decision is a function of one chain read, so a restarted keeper decides what the dead one would have. Proven on the fork, in both modes: killed right after broadcasting `clear_tier`, a crank, and a table create; each restart read the landed action as done and sent nothing twice, and adopted the crashed table rather than making a second. On a blockhash expiry, an RPC failure or a timeout it never re-signs the old instruction. It checks whether the signature landed late, re-reads state, rebuilds and re-simulates, so an action that landed is seen as done. A table create is stricter (rule 4 under "Lookup tables"). **The one window this does not close.** A transaction broadcast just before a crash can land up to a blockhash lifetime (about 60 to 90 s) later. A restart inside that window reads state before it lands and could send the same action again. The program makes that harmless for funds (a second crank or migrate or clear_tier fails on chain: `CrankTooSoon`, the pool exists, `TierAlreadyCleared`), but the loser would land as a failed transaction and cost its fee. `STARTUP_SETTLE_S` (90 s by default on a live cluster) closes it by waiting out the lifetime before the first read. Table creates do not depend on it (rules 1 to 3). The fork proofs ran with 0, because the fork confirms in under a second. ## Failure modes | Symptom | Cause | What to do | |---|---|---| | `start` error: not a loopback fork | Live `RPC_URL` without the flag | Add `--i-mean-a-live-cluster` deliberately | | `discovery` alert `GetProgramAccountsRefused` | The RPC plan does not serve the call | Switch to one that does, or set `BUNDLES` (allowlist mode). Until then only the fallback is served and new bundles wait | | `float` alert `FLOAT_BELOW_MINIMUM` | The float ran down | Send SOL to the key named in the line. Sending resumes on its own (`FloatRestored`) | | `rpc` `RateLimited` repeating | The plan's request rate is too low for the bundle count | Raise the plan, raise `CYCLE_INTERVAL_MS`, or lower `MAX_CONCURRENT_BUNDLES` ("Host requirements" has the request rate) | | `lookup_table` `TablePending` | A create was broadcast and did not confirm | Nothing. It is adopted when it lands, or replaced after 600 slots (about 4 minutes) | | `lookup_table` `TableNotUsable` | The recorded table is not this key's, is deactivated, or names another bundle (after a key rotation, for example) | Nothing: a new one is made. Close the old one with its own authority to recover its rent | | `lookup_table` `RecoveryFailed`, then the process exits | The state file is gone and the chain scan failed five times | Fix the RPC and restart. The keeper will not create a table it cannot prove is not a duplicate | | `crank` skipped `StaleOracle` for long stretches | The sponsored Pyth feed is not updating | Nothing on the vault side: it fails closed by design. Point `PYTH_PRICE_UPDATE` at another fresh SOL/USD `PriceUpdateV2` if one exists | | `NoLookupTable` | Allowlist mode with `AUTO_LOOKUP_TABLES=false` and no `LOOKUP_TABLE`, or the table is not ready yet | Wait one cycle, or name a table | | `lookup_table` info lists `missing` keys | PumpSwap rotated its fee recipients | An automatic table is extended with them; a named one rides them as static keys (32 B each) | | `migrate` skipped `AlreadyMigrated` with `raced: true` | pump's migrator or another keeper landed it first | Nothing: done | | `clear_tier` skipped `TierTargetAlreadyMet` (6031) or `TierSpendExceeded` (6032) | Market cap already above the target, or clearing costs more than the bundle's cap | Both fail closed and spend nothing. The keeper keeps skipping; crank and harvest still run | | `rpc_error` / `expired` repeating | RPC overloaded or blockhash expiry under congestion | Raise `PRIORITY_FEE_MICROLAMPORTS`, use a better RPC. Nothing is sent twice | | `failed` | A transaction failed after a passing simulation and no race explains it | Read `code`, `reason` and `log_tail`. A failed transaction changes nothing on chain except the fee it paid | | The process dies | Anything | Restart it (systemd and Docker do). It resumes from chain state. The vault does not depend on it | ## Running it always-on Two ready files in `keeper/deploy/`: `sheaf-keeper.service` (systemd, hardened, restarts on exit) and `Dockerfile` (with `Dockerfile.dockerignore`, built from the repo root: `docker build -f keeper/deploy/Dockerfile -t sheaf-keeper .`). Both run discovery mode with `--i-mean-a-live-cluster`, take the config from an env file and the key from a mounted file, and stop with SIGTERM, which finishes the actions in flight. Install steps are in each file's header. **The key is a file, nowhere else.** `KEYPAIR_PATH` is a path; a value that looks like key material is refused. The unit reads it from `/etc/sheaf-keeper/` (mode 600), the container from a read-only mount; neither file nor image contains it, and `Dockerfile.dockerignore` keeps `*keypair*.json`, `.env` files and `keys/` out of the build context. Never commit a key or an env file. ### Host requirements | Need | Figure | |---|---| | RPC | One that serves `getProgramAccounts` (discovery and table recovery). Request rate: a cycle is about 3 to 5 requests (one batched read, then a blockhash and a simulation per decision); a bundle in an open trade window cycles every `CYCLE_INTERVAL_MS` (15 s), a quiet one about once a minute, plus about 8 requests per landed transaction. Budget about 0.3 requests per second per bundle at worst: 3 for 10 bundles, 30 for 100, plus one scan a minute | | CPU | One core is ample: the work is waiting on the RPC | | Memory | About 350 MB resident: 225 to 356 MB measured across the nine keeper processes of the passing fork proof, the most in the longest (three bundles, about 15 minutes). Not measured over days. The units cap the heap at 512 MB (`NODE_OPTIONS`) and the service at 1 GB | | Disk | Under 1 MB: the state files (`tables-*.json` about 150 B per bundle, `spend-*.json` under 1 KB per bundle per day, 62 days kept) and the heartbeat | | Software | Node 22 or later (tested on 24) and the repo with `npm ci` at its root; the keeper has no dependencies of its own | | Network | Outbound HTTPS to the RPC only. The health port listens on 127.0.0.1 unless `HEALTH_HOST` says otherwise | | Clock | NTP-synced. Due times come from the chain clock, so drift only delays a cycle, bounded by `MAX_IDLE_MS` | ### Environment for a live keeper The minimum: `RPC_URL`, `KEYPAIR_PATH`, `STATE_DIR` on persistent storage, `MIN_FLOAT_LAMPORTS` sized below, `HEALTH_PORT` (or rely on the heartbeat file) and, under congestion, `PRIORITY_FEE_MICROLAMPORTS`. Leave `RPC_BATCH_READS` and `MAX_CONCURRENT_SENDS` at their defaults (the fork values are Surfpool artifacts). Keep `STARTUP_SETTLE_S` at 90. ### Costs Measured on the fork (`tests/e2e/keeper-discovery.sh`, 2026-10-02) at a priority fee of 1,000 microlamports per CU, from each transaction's own metadata (the key's balance drop: fee plus rent). Three bundles gave the same figures to within 30 lamports. **One-time, per bundle: 15,247,282 lamports, about 0.0152 SOL**, almost all of it rent. | Item | Lamports | Of which rent | Recoverable | |---|---|---|---| | pump `migrate` (the pool's boost vault) | 2,044,730 | 2,039,280 | no (pump's account) | | pre-creates (AMM creator-fee vault ATA, PumpSwap user volume accumulator) | 3,888,721 | 3,883,680 | no (the bundle's accounts) | | lookup table (create plus 20 keys, then 16 more; 36 entries) | 9,308,614 | 9,298,560 | yes: deactivate and close it with the keeper key once the bundle is dead | | `clear_tier` | about 5,216 | 0 | | | `collect_curve_fee`, once (the curve-phase creator fees) | about 5,033 | 0 | | **Steady state, per bundle per day**, at the example bundle's 300 s observation interval (288 observations a day). Measured per transaction: observation 5,070, sell 5,210, harvest 5,066. | Day | Lamports | SOL | |---|---|---| | Flat market: 288 observations, no sell, 1 harvest | 1,465,226 | 0.0015 | | 288 observations, 24 sells, 4 harvests | 1,605,464 | 0.0016 | | Busiest: 288 observations, 72 sells (the M3 pace, about one per 20 minutes), 8 harvests | 1,875,808 | 0.0019 | **At another priority fee** each transaction costs 5,000 lamports plus price x CU limit / 10^6. The keeper's limits (simulated units plus 25%) were about 66,000 CU for an observation, 210,000 for a sell and 67,000 for a harvest. At 100,000 microlamports per CU an observation costs about 11,600 lamports and a sell about 26,000, so the busiest day is about 5,317,000 lamports (0.0053 SOL) and the typical one about 4,020,000 (0.0040 SOL). Rent dominates the one-time cost at any price (about 0.0154 SOL at 100,000). ### Float sizing ``` float = N x (one-time 0.0154 SOL + D days x daily) + MIN_FLOAT_LAMPORTS ``` Daily is 0.0019 SOL per bundle at the measured price, 0.0053 SOL in the busiest case at 100,000 microlamports per CU. With 30 days of runway at the conservative figure: | Bundles | Float | `MIN_FLOAT_LAMPORTS` | Runway at the measured price | |---|---|---|---| | 10 | **2 SOL** | 150,000,000 (0.15 SOL: two busiest days plus one new bundle) | about 90 days | | 100 | **18 SOL** | 1,200,000,000 (1.2 SOL) | about 80 days | At no priority fee 0.75 SOL (10) and 7.5 SOL (100) cover 30 days. The floor is a stop, not a warning: below it the keeper stops serving every bundle, so watch `float_lamports` on the health endpoint and top up well before it. Every new bundle draws its 0.0154 SOL the day it launches. ### Rotating the keeper key The key holds only the float and the authority over the keeper's own lookup tables; it controls nothing in any vault. A leaked key can spend the float and close the tables (the keeper then makes new ones). Rotate at once if it leaks, and on any schedule you like otherwise: 1. `solana-keygen new -o keeper-keypair.new.json`, then fund it with the float. 2. Stop the keeper (SIGTERM; it finishes what is in flight). 3. Point `KEYPAIR_PATH` at the new file and start it. The state files are named for the key, so it starts clean: its recovery scan finds no tables for the new key, and it makes one per bundle as each needs it (the one-time table cost again, per bundle). 4. With the old key, deactivate each old table listed in `STATE_DIR/tables-.json` (`solana address-lookup-table deactivate
--keypair old.json`), and about 513 slots later close it (`solana address-lookup-table close
--keypair old.json`) to recover its rent. 5. Send the old key's remaining balance to the new one, then destroy the old file. ## Surfpool: a fork artifact, and how the fork proof avoids it Measured 2026-10-01 on Surfpool 1.6.0. During the first fork attempts, transactions were refused with `Transaction verification failed ... Failed to fetch accounts from remote: error sending request`, each after about 30 s, sometimes for minutes at a time. Simulations kept passing, so the keeper never sent anything the program would refuse, but `migrate` once needed seven tries. What the probes showed: 1. **It is the fork's own HTTP client.** With the fork's datasource pointed at a logging proxy, every request that reached the proxy was answered 200, and the failing requests never arrived. The fork was reusing idle keep-alive connections the far end had already dropped. 2. **Closing every connection fixes it.** `tests/e2e/datasource-proxy.js` forwards the fork's reads to the public mainnet endpoint and answers each one with `Connection: close`. `tests/e2e/keeper.sh` starts it by default. The next full run had **zero** refusals; so did the passing run below. 3. **`getMultipleAccounts` made it worse.** Before the cause was known, isolated probes showed a `getMultipleAccounts` that made the fork fetch from mainnet (an account not yet local) was followed by about 30 s of refusals, while the same accounts read with `getAccountInfo` were not. `RPC_BATCH_READS=false` switches the keeper to `getAccountInfo` per account. The fork proof runs with both the proxy and `RPC_BATCH_READS=false`; batch reads behind the proxy were not tested. None of this applies to a real RPC. Leave `RPC_BATCH_READS` at its default (`true`, one request per cycle) there. The keeper treats a refusal of this kind as `rpc_error`, a retryable non-landing, either way. **Concurrent heavy sends froze the fork** (measured 2026-10-02, Surfpool 1.6.0, behind the proxy). The first discovery run started two bundles at once; pump `migrate` on one and a lookup-table extend on the other were in flight together, and every request to the fork, a plain `getAccountInfo` included, then timed out for about six minutes, with no datasource request in that time and nothing in Surfpool's log. A probe of four concurrent transfers, six rounds, did not reproduce it, so it is narrower than "two sends". `MAX_CONCURRENT_SENDS=1` (the send gate in `keeper/src/send.ts`) serializes broadcast and confirmation across bundles while reads and simulations stay concurrent; the passing run uses it. A real RPC does not need it. Even serialized, a `sendTransaction` occasionally took over 30 s to answer and the transaction landed after the keeper's timeout (twice in one run); that is why a failed broadcast call is followed by a 10 s watch of the signature ("Scheduling", rate limits). Two more fork artifacts the harness handles, as `lifecycle.ts` does: the fork freezes the Pyth account at first fetch, so the harness rewrites only its publish time (every 5 s once the stale-oracle path has been shown, standing in for the live feed, and right after every time travel); and for about a second after a time travel, v0 transactions fail `InvalidAddressLookupTableIndex`, which the keeper sees in simulation and skips. ## Fork proof: discovery mode (`tests/e2e/keeper-discovery.sh`) `bash tests/e2e/keeper-discovery.sh` (WSL, repo root). It waits until no other Surfpool runs and at least 2.5 GB is free, starts the datasource proxy (8932) and a fresh mainnet fork on 8930, runs `tests/e2e/keeper-discovery.ts`, and stops both. Evidence: `tests/e2e/results/keeper-discovery.json` (every assertion, the limits, the costs, each process's peak memory), `keeper-discovery-runs.ndjson` (every line every keeper process wrote) and `keeper-discovery.log`. Passing run, 2026-10-03: **KEEPER DISCOVERY E2E PASS, 103 of 103 assertions**, against the release build sha256 `b298320cebf16ac1f812ab34506cbe9bd0b6bcb56c8e806f1793458c9a9cfc9c` (645,160 bytes), pinned by the harness. Five runs before it failed and each led to a fix: two fork freezes ("Surfpool" above), four harness or logging defects (a heartbeat read before its first rewrite, a fixed clock offset that ignored real time passing, a check that counted `clear_tier` as a crank buy, an `rpc` line without an `error` field), and two keeper changes, below. The harness never names a bundle to the keeper. Nine keeper processes, all started with no `BUNDLES` and no `--bundle`: | Step | Process | What it proved | Exit | |---|---|---|---| | 1 | | Operator: deploy, then `operatorFinalize` (revoke, protocol fee account) | | | 2 | | Creators A and B launch; creator D initializes only (Funding); a decoy account owned by the program, Bundle-sized, with the Bundle discriminator and A's bytes, at a random address | | | 3 | refuse-live-rpc | A mainnet `RPC_URL` without `--i-mean-a-live-cluster` is refused before any request | 2 | | 3 | d1, crash hook on the first table create, one bundle at a time | Discovery served exactly A and B, logged D as `Funding` and the decoy as `NotBundlePda`, migrated and pre-created A, wrote A's table address to the state file as `pending`, broadcast the create and died | 86 | | 4 | d2, health on | Adopted A's crashed table (same address, extended to 36 keys, no second create), made B's, cleared both tiers once; `GET /health` 200 with the served count, last cycle, float and errors; the heartbeat file the same | | | 5 | d2 | Creator C launched while it ran: served 1.98 s after `execute_launch` (scan every 10 s), then migrated, pre-created, table, tier cleared, with no restart | | | 6 | d2 | Third-party buys on all three pools; clock to +1,500 s: observations on all three (23 in total across the run) and at least one sell each (6), 3 sells deferred and logged | 0 | | 7 | d3, crash hook on crank | Died right after broadcasting B's observation; the orphan landed | 86 | | 7 | d4 | No table created or extended, no second `clear_tier`, B's orphaned observation not re-sent; all three cranked at the next due time | 0 | | 8 | d5, state file deleted | `StateFileMissing`, then all three tables recovered from chain (getProgramAccounts on the lookup-table program by authority), each the original address; all three cranked through them; nothing created | 0 | | 8 | d6 `--once`, through a proxy answering the first 4 requests 429 and refusing getProgramAccounts | `RateLimited` logged and backed off; `alert` `GetProgramAccountsRefused` with the RPC's own text; served A, B and C from the state file; each cycled | 0 | | 8 | d7 `--once`, same proxy, state file deleted again, `FALLBACK_BUNDLES` = A | Tables recovered from the key's own signature history, all three original addresses, none created; served A from the list and B, C from the rebuilt file | 0 | | 9 | d8, health on, `MIN_FLOAT_LAMPORTS` 0.1 SOL | Float drained from 9.95 to 0.06 SOL: `alert` `FLOAT_BELOW_MINIMUM`, every due crank `skipped` `FloatBelowMinimum`, nothing sent, no keeper transaction on chain after the drain, health 503 `float_ok: false`. Topped up: `FloatRestored`, sending resumed | 0 | Assertions over everything, grouped: - **Zero failed sends.** All 54 transactions the keeper key paid for succeeded on chain; each is in the log with its signature and every logged send is on chain. The only error lines are the four injected 429s. - **One of each one-shot action per bundle,** across nine processes and two crashes: one migrate, one pre-create, one `clear_tier` each. Exactly three keeper-owned tables on chain, one per served bundle, each holding its bundle's 36 keys and none of another bundle's. - **No cross-bundle effect.** No keeper transaction touches two bundles' accounts, and none touches the Funding bundle D or the decoy. Each bundle's marker signatures are exactly its own vault swaps; `verify-bundle --strict-marker` passes for each; each is sell-only with no buy; no two observations in any ring are closer than 300 s. - **Limits:** every keeper transaction inside 1,232 B, 64 trace entries and depth 5. New ones: table create plus 20 keys 964 B, 22,459 CU; extend of 16 keys 816 B, 10,783 CU. - **Money:** the spend ledger equals the sum of every logged `spent_lamports`, and each of those equals the key's balance drop in that transaction's metadata. Costs are under "Costs" above. **The two keeper changes the runs forced.** Run 3 logged one crank as `failed`: the fork clock jumped 300 s between a sell's simulation and its send, the program's next action became an observation, and it refused the stale sell at preflight with 6035. Nothing landed. That is now `skipped` `ActionChangedBeforeSend`. Run 4 showed two `sendTransaction` calls that timed out at 30 s and landed afterwards; the retry read them as done (no double), but the log had no signature. A transient broadcast failure is now followed by a 10 s watch of the signature. ## Fork proof: allowlist mode (`tests/e2e/keeper.sh`) `bash tests/e2e/keeper.sh` (WSL, repo root). It waits until no other Surfpool runs and at least 2.5 GB is free, starts the datasource proxy and a fresh mainnet fork on 8977, runs `tests/e2e/keeper.ts`, and stops both. Evidence: `tests/e2e/results/keeper.json` (every assertion, the per-action limits), `keeper-runs.ndjson` (every line every keeper process wrote) and `keeper.log`. **Rerun on the public-keeper build, 2026-10-03: KEEPER E2E PASS, 59 of 59 assertions**, against the release build sha256 `b298320cebf16ac1f812ab34506cbe9bd0b6bcb56c8e806f1793458c9a9cfc9c` (645,160 bytes, pinned with `KEEPER_E2E_EXPECT_SHA`, the 59th assertion). Same scenario, one bundle in allowlist mode (`BUNDLE` and `LOOKUP_TABLE` set), now through the per-bundle scheduler; the harness adds `MAX_IDLE_MS=1500` and `STATE_DIR`. All 22 keeper transactions succeeded. One migrate, one pre-create, one `clear_tier` (its sender died before confirming and the orphan landed); 8 observations plus the orphaned one, 1 reference step, 2 sells, 5 sells deferred and logged. `fee_ledger_self` 266,522,562 and `fee_ledger_organic` 42,243,084 exact; the protocol cut 4,693,676, 10% of the 46,936,760 organic. Zero `rpc_error`, `expired` or `error` lines. The 2026-10-01 record follows unchanged. Passing run, 2026-10-01: **KEEPER E2E PASS, 58 of 58 assertions**, against the release build sha256 `7915e1276fff79a11efb1806513e1f65599b4e910c491a2ae2afaaad6bebb04b` (643,608 bytes, the security-fix build, the source committed afterwards as `d54f0a5`, the same build the lifecycle E2E ran 110 of 110 on), with the H3 runbook order: deploy, `initialize_bundle` signed by the upgrade authority, revoke. The harness only deploys, initializes and launches (through `scripts/launch-bundle.ts`), drives two third-party 2.5 SOL PumpSwap buys and moves the clock. Everything after `execute_launch` is the keeper process, started as an operator would, eight times: | Process | What it did | Exit | |---|---|---| | `--once` in Funding | Read `Funding`, sent nothing | 0 | | `--create-lookup-table` | Built the 36-key table (3 transactions) | 0 | | `--dry-run --once` | Simulated `migrate` and `collect_curve_fee`, sent nothing; no new keeper signature on chain | 0 | | run 1, crash hook on `clear_tier` | Sent `migrate`, then the pre-creates, then broadcast `clear_tier` and died before confirming. The orphan landed | 86 | | run 2 | Read migrate, pre-creates and the tier as done; skipped the crank on a stale oracle (6045); collected the curve fee and harvested; observations at +300, +600, +900 and the first sell at +900; second harvest; SIGINT | 0 | | run 3, crash hook on `crank` | Broadcast the +1200 observation and died; the orphan landed | 86 | | run 4 | Did not resend the +1200 observation; seven observations, two more sells, the ratchet step; SIGINT | 0 | | `--once`, 1-lamport threshold | The final harvest | 0 | Assertions, grouped: - **Zero failed sends.** All 25 transactions the keeper key paid for succeeded on chain, every one of them is in the logs and every logged signature is on chain, and no line has outcome `failed`. Zero `rpc_error`, `expired` or `error` lines. - **The marker exactly on swaps.** On every keeper transaction the marker is listed iff it emitted exactly one `VaultTrade`; the marked ones are `clear_tier` and the three crank sells only; the 12 unmarked cranks are 11 observations and 1 reference step. Successful marker signatures equal the keeper's 4 vault swaps exactly, and `verify-bundle --strict-marker` passes with no foreign entries. - **No buy.** `mode` is 1; `cum_bought_quote`, `buy_streak` and both `buy_out_*` are zero; every crank trade is side 0 (sell). - **Fee ledgers exact, the E2E pattern.** `fee_ledger_self` 269,754,266 = 254,125,198 (launch, net of the rent `fee_auth` keeps) + 5,897,695 (`clear_tier`) + 9,731,373 (three crank sells), each term measured outside the program from transaction metadata. `fee_ledger_organic` 46,936,760 = the two third-party buys' creator fees exactly. Both equal the E2E figures to the lamport (E2E.md, "E1 fixed"). Everything self is recognized; the three harvests moved exactly self + organic (316,691,026) into `vault_quote`; the creator vault and inbox end empty and `fee_auth` at its rent minimum. - **No duplicate after the restarts.** Exactly one `migrate`, one pre-create and one `clear_tier` across all processes, though the process that sent `clear_tier` died before confirming; no two ring observations closer than 300 s; no crank signature logged twice. Per action, against the limits (CU limit is what the keeper requested: simulated units plus 25%, at least +10,000): | Action | n | CU | CU limit | Bytes (of 1,232) | Trace (of 64) | Depth (of 5) | Keys static + table | |---|---|---|---|---|---|---|---| | `migrate` (pump) | 1 | 341,735 | 427,169 | 1,058 | 53 | 4 | 28 + 0 | | `precreate` | 1 | 36,591 | 46,591 | 535 | 10 | 2 | 12 + 0 | | `clear_tier`, marked (measured from chain: its sender died) | 1 | 170,790 | not logged (its sender died) | 332 | 13 | 3 | 3 + 30 | | `crank` observation | 10 + 1 orphan | 48,478 to 49,312 | 61,389 | 335 | 3 | 1 | 3 + 31 | | `crank` reference step | 1 | 47,911 | 59,638 | 335 | 3 | 1 | 3 + 31 | | `crank` sell, marked | 3 | 152,495 to 152,690 | 190,612 | 336 | 13 | 3 | 3 + 32 | | `collect_curve_fee` (pump) | 1 | 18,486 | 28,486 | 529 | 5 | 2 | 12 + 0 | | `harvest` | 3 | 49,833 to 52,924 | 66,155 | 322 | 9 | 3 | 4 + 11 | | lookup table create / extend | 3 | 9,013 to 13,128 | none | 252 to 1,052 | | | | The tightest margins are pump's own `migrate`: 1,058 of 1,232 bytes and 53 of 64 trace entries (52 in the E2E, plus the keeper's `SetComputeUnitPrice` instruction). **Deferred sells.** Twice in this run (two consecutive 300 s cycles inside one trade window) the program chose a sell (the unmarked crank failed 3005) but the marked sell failed in simulation inside PumpSwap with PumpSwap's own 6004 `ExceededSlippage`, propagated out of the crank. The program engineer attributes this to the new M3 median floor deferring a sell. The keeper logged each as `SellDeferred` with the venue code and the log tail, sent nothing, and the sell executed on a later cycle. Three crank sells executed, as in the E2E. The run before this one, on the same build, also passed 58 of 58 with four deferrals and two sells; its fee ledgers matched the same way (266,522,562 self, 46,936,760 organic). ## Not proven - Anything on devnet or mainnet: a live Pyth feed, real time, congestion, priority-fee landing, a real blockhash expiry (the retry path is built but the fork never expired a send), and a real provider's rate limits (429 is proven only against the harness's own restrictive proxy). - `getProgramAccounts` on a real paid RPC. On the fork the scan of the program and the scan of the lookup-table program by authority both answered (Surfpool forwarded them to mainnet and merged its local accounts). Whether a given provider answers the lookup-table scan, a program with millions of accounts, in time is untested; the signature-history fallback covers a refusal, and a failed scan stops the keeper rather than risk a duplicate table. - Discovery at scale: three served bundles, one Funding and one decoy on the fork, not 100. The request-rate figure in "Host requirements" is computed, not measured. - A race with pump's own migrator, or with a second keeper, in flight. The code treats a failure whose re-read shows the action done as a race (`raced: true`), but the fork never produced one. - Concurrent sends on a real RPC (the fork proof ran `MAX_CONCURRENT_SENDS=1`), and the float guard's undershoot with several sends in flight. - `RPC_BATCH_READS=true` against any RPC, and the fork without the datasource proxy. - The restart window `STARTUP_SETTLE_S` closes (a broadcast that lands after the restarted keeper's first read). The fork lands in under a second, so every orphan had landed before its restart. - The pending-table expiry path (`TableGone` after 600 slots): every crashed create in the proof landed and was adopted. - Halts, the dislocation path, `TierTargetAlreadyMet`, `TxTooLarge`, and a `failed` send: none occurred. - The systemd unit and the Docker image were written, not run: this machine has neither systemd services for it nor Docker. Both start the same command the fork proof starts.