Skip to content

2026 - 08

Goal: Rapid strategy iteration to positive margin — settle MM, and prove arb is a convergent type. | Phase: 1 (Validate)

July removed the platform as an excuse. Release time is down, non-engine deploys cost zero trading downtime, a full environment rebuild is a rehearsed 45 minutes, and the deploy tool is no longer end-of-life. August is the month that spends that: rapid iteration toward the actual Phase-1 exit bar — two convergent strategy types, positive P&L and positive margin, in one clean month at min quantity.

Two things sharpen the target. First, arb may already be halfway there: btcusdt (+9.6e-5) and ethusdt_seq (+7.3e-5) are two variants of one type showing similar metrics — which is the plan's definition of convergence, not one lucky arm. It needs a decisive sample, not a discovery. Second, MM is unmeasured, not dead — July's break coincides ~1:1 with q50 on the control's book while the regime was unchanged, so E1 decides it in 7 days for ~nothing.

With the platform no longer the constraint, the thing to compress this month is the loop from result to next experiment. July's cycles were paced by genuine measurement — MM needed an undisturbed book, MR needs fills — but where a decision is available on the evidence, August should take it in days, not weeks. That is a deliberate speed target, not a diagnosis of July.

Week 1 (Aug 3–9) — everything keys off E1's read-out

E1 is already running. eth_okx_sigcancel2_r2 has been alone on the book since ~07-31; q50, q100b and volgate_r3 are disabled. Arb runs unchanged. At ~1,400 fills/day the 10k-fill / 7-day thresholds both land Thu 06 – Fri 07 Aug.

Early read — not a verdict. Marginal P&L/fill was −3.51e-4 across the contaminated window (07-26→08-01) and −3.17e-4 in the first clean days (08-01→08-02). That is flat, not the recovery self-competition predicts. But it is ~1,400–3,000 clean fills against a 10,000 threshold, and the pre-registered metric is capture, not marginal P&L/fill. Do not call E1 early. The one thing it does justify: prepare the decay branch properly.

The organizing fact of the week: the live surface is frozen until Friday. That is not idle time — it is the one week where MM capacity is free and the thing gating whatever comes next can be fixed without contaminating anything. So the week's job is to prepare both branches so Friday's read-out converts to action the same day.

Critical path — paper POST_ONLY → DirectionUp v3. If E1 says decay (which the early read leans toward), the type-2 exit candidate becomes the reopened direction family, and it currently has no validation path: paper cancels resting orders at ~one run-interval, so v3's maker-first design cannot be tested there. That makes the paper fix the highest-leverage item of the week, and it must land before Friday to be useful.

Mon–Tue Wed–Thu Fri (E1 reads) Sat–Sun
Strategy (MJ) Track A audit — offline, from existing stamps: cancel→repost gap, 1s level-join lateness, TV freshness Audit read-out · pre-register both branch decisions Read out E1 on the capture metric · execute the pre-decided branch Begin E2 ladder or commercial-only scope
Infra (MJ) Diagnose + fix paper POST_ONLY lifetime One batched release · scope #745 (don't build) · alert-severity accumulation — (no release Friday) Build #745 only if E1 says self-competition
ML (Vicky) Re-measure q(H) on the current live config Prep DirectionUp v3 live-tiny, ready to fire v3 goes live-tiny if decay branch Read v3's first fills

Sequencing rules for the week:

  • Do not touch the MM book before Friday. Any change re-contaminates the only clean read we have.
  • Batch engine releases into one. Every release stops the book ~5 minutes and eats E1 fills. One release Wed/Thu, none Friday.
  • Pre-register both branch decisions by Thursday — the E2 ladder design and the commercial-only scope — so Friday is a lookup, not a debate. This is the "act the same week" commitment made concrete.
  • Read E1 on its own metric. Capture vs the −0.19 / −0.31 thresholds, not the marginal-P&L proxy used above.
  • Arb: no changes, but watch it. ethusdt_seq margin halved (8.2e-5 → 4.3e-5) on a 144-fill sample and usdcusdt is still at 0 fills (#269). Watch; don't act on 17 fills of movement.
  • MR: keep accumulating, set the trigger. No intervention; name the fill count and margin band this week.

1. MM — E1 decides, everything else is downstream

The full experiment program lives in its own design doc: inf-trading-strategies/docs/design/2026-08-01-mm-experiment-program.md. Not restated here — one source of truth. What the company plan commits to:

  • Run E1 first and let it gate the rest. Re-enable 2113 alone: 7 days / 10k fills, thresholds pre-registered. Capture recovers toward −0.19 → self-competition confirmed, MM alive, proceed to the time-sliced ladder (E2). Capture stays at −0.31 → genuine decay, MM goes commercial-only. Costs one operational click.
  • Do not start Track A/B/C builds before E1 reads. The audit-first rule applies: Track A begins as an offline presence/join-latency audit from stamps we already have, and only what the audit shows dominates gets built.
  • Hold the −0.90 commercial arithmetic as the fallback, now independently confirmed on ETH and BTC. If E1 says decay, MM's remaining value is commercial, not proprietary — and that is a result, not a loss.

The methodology rules from that doc are binding company-wide, not just for MM: one arm per book · marginal windows, never cumulative margin · pre-registered confound lists · offline falsification before building.

2. Arb — convert the best margin we have into a convergent type

The month's highest-expected-value trading work, because it is the only thing already positive and it is the closest to satisfying half the exit bar.

  • Drive both live arms to a decisive sample. btcusdt_okxbus_r is at 907 fills, ethusdt_okxbus_seq_r at
  • Convergence needs both arms holding similar margins on samples nobody can wave away — that is the claim to establish, and it is a waiting task as much as an engineering one.
  • Close the correctness risk before scaling attention onto it (#261). The in-flight-hedge guard — the handler re-fires before a fill reconciles — is the one defect that can corrupt the only converging cohort. It is also exactly the class the 4 Hz experiment already surfaced once as a double-hedge race.
  • Diagnose the silent arm (#269). usdcusdt_okxbus_arb has sat at 0 fills for days. Either it is mis-gated or the opportunity isn't there; both are cheap to establish and one of them is a bug.
  • Latency→capture elasticity (#228) — the BUS-gated sequence. Note this is the same underlying question as MM Track A: does queue presence, not signal, pay? Run them as one investigation with two applications, not two.

3. MR — accumulate the sample, then change how it runs

The framing that matters: this is not a keep-or-kill decision waiting on courage. All four arms are negative but fill-starved — 24 to 53 fills, at or below the decisive gate — and the real question is how to run MR differently, which needs samples to inform. Killing it now would discard the input to that decision.

  • Keep the arms running to accumulate fills. The cost is small and bounded; the information is the point. This is the same discipline as MM's freeze: an undisturbed run is what makes the next decision real.
  • Treat it as an execution problem (#245), not a signal one — and note July already narrowed it. The band-resting maker hook was killed on live episodes (−4.63 bp, t −3.9, on 55 real entries) after an offline proxy said +2.7 bp. The surviving variant is trigger-time maker-ization — POST_ONLY at current price at signal, seconds-TTL, taker fallback, capped ≈ +2.3 bp. It cannot be priced offline (the queue race at trigger is the unmeasured corner), so the only path is a live-tiny A/B. The deep sweep-catching work from MM's Track B still routes here.
  • Set the trigger in advance. Name the fill count and the margin band at which MR's execution redesign either ships or the cohort retires — pre-registered, the same rule E1 follows. Deciding the threshold now is what stops this from drifting a fourth month.

Direction (#188) is effectively resolved: every live variant is retired, and what remains runs in paper — some of it positive. The umbrella issue can be closed on that basis whenever the paper arms are read out; there is no live exposure and nothing to execute.

4. Platform — pay down the three taxes on iteration speed

August's goal is rapid iteration. These are the specific things that make it slow. Each is scoped against that test — if it doesn't buy iteration speed or protect correctness, it waits.

  • Paper POST_ONLY lifetime — do this first (week 1). See the week-1 sequence: it is the only item on this list that unblocks a strategy validation path rather than making an existing one cheaper, and the branch it unblocks is the one the early E1 read points at.
  • Warm-standby handover (engine#796). Every engine release stops the book ~5 minutes. At July's rate (86 releases) that was ~7 hours of self-inflicted downtime. In a month explicitly about iteration rate, this is the compounding tax. Built on READ_ONLY, which already exists. Deprioritized in week 1 — it compounds, but it unblocks nothing, and week 1 is release-constrained by E1.
  • Hidden/iceberg order support (engine#745). Gates the MM capacity ladder (#254) — E2 cannot answer the size question without it. Scope in week 1, build only if E1 says self-competition. If E1 says decay, this isn't needed at all.
  • Account-awareness end to end (engine#780). The structural root of July's capital-ledger corruption. Until it lands, expect more defects of that class — and every one of them lands on the ledger every convergence judgement is measured against.
  • Finish the migration (inf-trading#169) — decommission the retired Copilot stacks. Cost, and closure.
  • Fix the paper POST_ONLY lifetime defect. Resting orders die at ~one run-interval, so paper cannot validate any rest-horizon maker strategy — it silently converts every maker experiment into a fills-fast-or-never one. This blocks the validation path for the reopened direction family and the MR maker hook, which makes it the highest- leverage platform fix on this list after warm-standby.
  • Fix the ML data-refresh defect properly (ml#320). The cutover cleared the symptom by replacing tasks; the cause — nothing reliably refreshes the futures-metrics file — is untouched and will recur.
  • Make alert severity grow with duration. A three-day total outage emitted the same "2 consecutive bars" notice as a brief blip, because the counter resets instead of accumulating. This is a monitoring correctness bug, and it is cheap.

5. ML — consume #299, and get the reopened direction family a real test

The estimator August needs already exists. #299 shipped the (offset, horizon, realized-vol) fill-probability surface and the P(trade-through) × q(H) factorization in July. August's ML work is not to build it — it is to put it to work and to give the reopened direction family the live test that paper currently cannot provide.

  • Wire the fill-probability surface into the MM Track A audit. Track A asks whether queue presence pays; #299 answers the adjacent question quantitatively — trade-through is free from tape on any coin/venue, and the residual short-horizon gap is pure queue (44% @1s → 90% @5s on BTC). Use it as the gate rather than re-deriving it. Note q(H) is execution-setup-specific and must be re-measured per config — it is not a constant we can borrow.
  • Fix the paper POST_ONLY lifetime defect — it is a blocker, not a nuisance. Paper cancels resting orders at roughly one run-interval (39–71s), so a 480s strategy cancel never fires. Paper cannot validate any rest-horizon maker strategy — which is exactly what DirectionUp v3 and the MR maker hook are. Until this is fixed, the reopened direction family has no validation path short of live.
  • Give DirectionUp v3 a live-tiny test once the above clears. v1's paper read was thin and on the old design; v3 carries the validated stack (delta-below entry + maker-first exit, +5.94 bp/signal in re-score). Testing v1's numbers as if they judged v3 would repeat July's apparatus error in a new place.
  • Hold the 3-book OOS protocol (BTC + ETH + SOL, Sharpe 1.58) — no re-tuning before 2026-10-01. The pre-registration is the point; touching it early destroys the only clean read we have scheduled.
  • No new alpha-from-model probes on crypto. Unchanged. The STOP list gains: triple-barrier / up-only / anomaly-excursion labels on this tape, aggTrades size-bucket features, further hyper-parameter tuning on direction, and the band-resting MR maker hook.

6. Substrate — US equities is answered; the vol/options question isn't

US equities came back negative and was correctly deprioritized on that read. That thread is closed and should stay closed. What survived it is a different question wearing an equities costume.

  • Separate the vol/options-access thread from the dead equities one. The memo's real find is that US-listed options give the VRP / vol-selling edge an accessible home — the edge June judged real but gated on crypto derivatives access. That is a sibling to the crypto derivatives-access decision, and it should live as its own thread so a future reader looking for the vol decision can find it.
  • Pressure-test the carry leg before anything else. The barbell's engine is VRP, and VRP is the edge our own work already judged marginal and decaying (BTC net Sharpe ~0.6–1.1, ETH uninvestable), on a substrate that is more crowded, not less ($127bn vol-ETFs, 59% 0DTE). The structure inherits the weakness of its most-competed component. If that leg doesn't clear net-@-cost, the elegance is irrelevant.
  • Close out PR #174 — dangling sibling reference, title convention, and the open question: what does the model show on free kline data across assets? If we are not running HFT on equities there is no case for buying tick or order-book data, and the free-data answer is the whole decision.
  • No build either way. Research and decision only, per the year plan.

7. Company — one conversation, and the hedge it represents

Discovery has been quiet since early July. The September wedge-decision gate is now ~6 weeks out and there is still zero demand data.

  • Hold one discovery conversation. Not a sprint — one. Cody/Windhorse is warmest (already a build-partner posture); peer-notes frame, not a pitch. Two weeks of a zero-conversation "sprint" says the 12–15 framing is too big to start from.
  • Test the counter-hypothesis first: do funded pre-launch desks reflexively build rather than partner? The entire Seg-2 wedge depends on the answer and nothing else in the plan can substitute for it.
  • Recognize what this is a hedge against. The Q3-end checkpoint asks whether the model is viable if we are still in Phase 1. The segment map's argument is that a platform aggregating many edges de-risks the single-edge dependency that Phase 1 currently rests on. That argument is worth exactly as much as the demand data behind it — which today is none.
  • Tier-0 content ships only on genuinely marginal hours.

What's NOT in August

  • Same-book A/B of any kind. July's cardinal sin, and the reason a 59,394-fill sample is ambiguous. Parallelism comes from breadth of books, not concurrent arms on one — a real constraint on how "rapid" iteration can be, and the reason the venue/symbol screen in MM Track C matters more than it looks.
  • Order-quantity ramp. Min qty until Phase 1 exits. Unchanged.
  • A new edge scan. June's taxonomy stands; do not reopen killed regions.
  • Derivatives access. No, again.
  • Equities build. Decision and closure-scan only.
  • New venue integrations unless the Track C diversification screen produces a specific, justified case.
  • Backtest-fidelity investment for HF. Validate live at min qty; the July sim work already closed the gross gap.

Team Focus

MJ: E1 and the MM program · arb to a decisive sample (#261, #269) · the iteration taxes (#796, #780) · the equities decision.

Vicky: put #299 to work — the fill-probability surface as the Track A gate · the reopened direction family (v3 live-tiny once paper is fixed, SOL/TRX watch, hold the OOS protocol) · the trigger-time maker-ization A/B for MR · ml#320 root cause.

Gaddafi: execution-quality surfaces for the Track A audit — cancel→repost gap, level-join lateness, quote freshness. The month's central question needs to be visible before it can be argued about.

Phase Status

Phase 1 (Validate). The exit bar, stated exactly: two converged and profitable strategies of distinct types, demonstrated in a single clean month at min order quantity — convergence meaning multiple variants with similar metrics, profitable meaning positive P&L and positive margin, clean meaning no incident-luck asterisks.

Against that bar, August's honest position:

Requirement Status entering August
Type 1 convergent + profitable Arb — plausibly there. Two arms, both positive, similar margins. Needs sample size, not discovery.
Type 2 convergent + profitable Open, but with two candidates now. MM if E1 says self-competition — or the reopened direction family, which needs a live path (blocked on the paper POST_ONLY defect) rather than more research.
Backtest good enough for parameter selection Materially closer — the ~5–7× under-fill gap was measured and largely closed.
Risk core survives incidents Strongest it has been: exposure caps, per-strategy budgets, verify-halt gate proven in a live incident.

So August is a two-outcome month — but no longer a narrow one. If E1 says self-competition and arb holds its sample, we end with two candidate convergent types and a credible run at a clean September. If E1 says decay, MM becomes commercial-only — but unlike a month ago, the bench is not empty: the reopened direction family is the second candidate, and its blocker is a platform defect (paper POST_ONLY lifetime) rather than an unanswered research question. That is a much better position to be in, and it is the direct payoff of July's research month. Failing both, the vol/options-access question stops being a side thread and becomes the live one.

The Q3-end checkpoint is two months out, and it asks whether the model is viable or needs structural rethinking. August is where the evidence for that conversation gets made. The measurement pace is what it is — MM needs an undisturbed book, MR needs fills, arb needs sample — so the compression has to come from everything around the measurements: pre-registering the trigger before the run, reading out the moment the threshold is hit, and acting the same week rather than the next month.