2026-08-16 BASELINE / DAY 01
Initial boot queried markov/state and markov/facts.
Both recalls returned no matching memories for the hackathon objective.
Two handoff writes were attempted through the bridge CLI. Both exceeded the 60-second timeout and returned no confirmed blob receipt.
A separate controlled SDK write reproduced the underlying relayer failure: durable upload returned 503 Service Unavailable because gas selection reported insufficient SUI. No credentials or internal identifiers are published.
The managed relayer was subsequently re-funded. A diagnostic write completed and the real experiment checkpoint was saved and confirmed by recall.
Checkpoint blob: P1yGfs4X…
Diagnostic blob: RiguJokR…
2026-08-17 BASELINE / DAY 02
Continuity test executed: a new fact and a superseding checkpoint were stored, then a fresh context booted with only the prompt and "pick it up".
Recovery was complete and correct: Goal / Done / Next / Blockers were resumed with dates, sources and blob ids — zero re-explanation, unlike Day 1.
The 4-question continuity check was also answered from memory: which fields not to publish, the no-change rule for the prompt, the checkpoint schema, and the never-verify-by-recall rule.
Checkpoint blob: 2UKUeSj… (supersedes P1yGfs4X…)
Fact blob: YDNg0XXR…
2026-08-17 BASELINE / DAY 03
Focused prompt improvement drafted: PROMPT_v2.md (original untouched). Two changes address only the proven friction.
Boot queries now use the user's exact words plus derived task keywords, and an empty recall triggers one retry before declaring a fresh start.
Write failures get a diagnosis protocol: a timed-out write is not stored; run a controlled remember(), poll the job status, and only retry after the cause (e.g. managed-relayer gas) is fixed.
Continuity re-test against the changed prompt passed: full Goal / Done / Next / Blockers recovery, no regression, no re-explanation.
Full before/after diff is documented in hackathon/DAY_3_IMPROVE.md; the original PROMPT.md and the variant PROMPT_v2.md are kept in the repo.
2026-08-20 IMPROVE / DAY 04
Rule 5.4 (write-failure diagnosis) tested live. First the Day-1 symptom was reproduced: rememberAndWait with a short timeout threw 504 and hid the real state.
The 5.4 protocol was then followed against a real write: rememberAsync returned a job_id, getRememberStatus was polled through pending → running → uploaded → done, and the terminal state plus blob receipt were recovered. Diagnosis replaced guessing.
Key finding: the "timed-out" write was actually stored in the background (blob MpfPmwmil…) before the 5.4 write landed (blob PWGWZuJ…). A timeout does NOT mean nothing was written — a blind retry would have silently duplicated the record. Polling to a terminal state first is what prevents that.
Evidence: hackathon/DAY_4_TEST54.md (private log); sanitized here — no job ids or trace ids.
2026-08-21 PACKAGE / DAY 05
PROMPT_v2.md promoted to the experiment's official version: AGENTS.md regenerated from v2 (diff-verified), baseline PROMPT.md kept frozen for comparison, and a v6 entry added to the changelog with each diff tied to its field evidence.
Submission package drafted: Walrus Memory feedback issue (wait-API masks terminal job state behind timeouts), Medium article (~500 words), and X post — all from the Days 1–4 evidence.
Honest note: validated by live continuity probes and the mainnet write-diagnosis test; running the full mocked eval suite against v6 stays open as follow-up.
DEADLINE 2026-08-24 / SUBMIT
Human actions before submission: open the issue on the original repo, publish the Medium article and X post, record the demo video, fill the WalForm.