Skip to content

Latest commit

 

History

History
215 lines (161 loc) · 9.97 KB

File metadata and controls

215 lines (161 loc) · 9.97 KB

Spec gaps — interpretations made while building this reference impl

While writing this reference implementation from the architecture-v1 docs, a number of points required interpretation. Each is listed below with the decision I made and the rationale. These feed back into architecture v2.

1. Memory record id format

Gap: ARCHITECTURE §2 uses mem_xxxxxxxx in examples but doesn't specify length, character set, or entropy requirements.

Decision: mem_ + 12 lowercase hex characters (48 bits of entropy, ~2^48 collision-resistant for a single-user hub). Generated via secrets.token_hex(6).

Alt considered: ULID for lexicographic sortability. Dropped because the schema already has a created_at column and ORDER BY updated_at DESC is the dominant list query.

2. MCP transport

Gap: ARCHITECTURE §6 "narrow-waist schema" specifies MCP-over-HTTP but doesn't say whether to implement the full streamable-HTTP transport from modelcontextprotocol.io/specification, stdio, or a simpler JSON-RPC-over-HTTP subset.

Decision: Phase 0 ships JSON-RPC 2.0 over POST /mcp. This is valid MCP for any client that configures HTTP transport; it is not the full streamable variant (no SSE, no session stickiness, no resumable streams).

Alt considered: Using the official mcp Python SDK. Dropped because that SDK currently defaults to stdio transport, and wrapping it for HTTP adds a nontrivial wrapper layer that obscures what's happening — the point of a reference impl is legibility.

Future: v0.2 will add a streamable-HTTP endpoint alongside, so clients can pick. Document the transport negotiation in the architecture v2 doc.

3. Tag search semantics

Gap: ARCHITECTURE §2 says tags are a list but list_memories(tag=...) semantics aren't specified — exact match? prefix? substring?

Decision: substring match against a flattened space-separated lowercase tag string (tags_flat column). This lets callers query tag=safety and hit both safety-charter and priority:safety, which matches how tags are used semi-structured in the public Starshard memory examples. Exact-match callers can filter post-hoc.

Alt considered: exact match with a separate junction table for multi-tag AND queries. Dropped as overkill for Phase 0; FTS5 already supports text queries over tags if needed.

4. Authentication model

Gap: SAFETY-CHARTER §Agent-permission-matrix mentions bearer tokens but doesn't specify per-user vs per-agent vs single shared token.

Decision: a single HUB_API_TOKEN env var. All authenticated requests use the same bearer. This is Phase 0 — identity is expected to come from the Cloudflare Access layer in front (email OTP for humans, service tokens for agents). Per-user or per-agent tokens are a Phase 1+ extension.

Alt considered: generating per-agent tokens on first connection. Dropped because the bootstrap problem (who issues the first token?) pushes complexity into Phase 0 that can be solved cleanly with a single shared secret.

5. provenance schema

Gap: ARCHITECTURE §2 lists provenance as a field but doesn't pin a schema. The public Starshard hub uses derivation_type, confidence, source_chat_excerpt, source_session_id — unclear if those are required or recommended.

Decision: store provenance as an opaque JSON object. The server does no validation beyond "is it JSON-object-shaped". Agents are free to evolve the shape. Architecture v2 should publish a canonical provenance schema and make it a first-class validated column.

6. Archive semantics

Gap: ARCHITECTURE §2 lists archived: bool but the contract for how archived records interact with list/search isn't in the doc.

Decision: archived records are hidden by default from both list_memories and search_memories. Callers opt in via include_archived=true. Get by id still returns archived records.

7. Full-text search over Chinese content

Gap: The existing Starshard hub stores Chinese mixed with English; the architecture doc doesn't say what tokenizer the reference impl should use.

Decision: SQLite FTS5 with unicode61 tokenizer. This tokenizes Chinese characters as individual tokens, which is imperfect for phrase search but works for the "keyword soup" style queries dominant in practice. For production-grade Chinese search, swap to ICU tokenizer or a dedicated search backend (Meilisearch, OpenSearch) — that's Phase 1+.

8. Pagination

Gap: No cursor/offset convention defined.

Decision: limit parameter only, default 20, max 200. No offset — clients needing more can narrow via tag / source filters. Cursor-based pagination is a Phase 1 concern.

9. Mirror consolidation frequency and scope

Gap: ARCHITECTURE §4 describes Mirror as "offline consolidation passes (daily / weekly)" without pinning cadence or scope.

Decision: Phase 0 Mirror runs every 24h, reads memories updated in the last 24h (capped at 200), proposes up to 5 consolidations per run, writes them as a single proposal awaits-user-approval memory. Proposals are never auto-applied — surface-only.

Rationale: nightly cadence balances freshness (users want improvement visible within a day) with cost (LLM calls are not free). 5 proposals keeps the review effort bounded; if Mirror thinks there are more, it should pick the top 5 by estimated value and defer the rest.

Alt considered: continuous on-write consolidation (every new memory triggers a pass). Dropped — too expensive, too noisy, and breaks the "surface don't judge" invariant by coupling writes to inference.

10. Weekly digest format + privacy scope

Gap: no spec doc covers the user-facing digest surface, but PHILOSOPHY implies "improvement must be visible to the user".

Decision: one plain-text email per week, 400-700 words, structured as headline + five bullets + proposals + themes + stuck items. Excludes any memory tagged sensitive-* or relationship:* (hard filter at fetch time, not just prompt-level) to respect cross-tier red lines from SAFETY-CHARTER.md.

Rationale: email because it meets the user where they already are; plain text because markdown rendering varies across clients; short length because the goal is visibility, not comprehensive audit. The hard filter prevents the LLM from being asked to redact — redaction by filter is more auditable than redaction by prompting.

12. Assumption-layer conflict detection (v0.3)

Gap: ARCHITECTURE.md §consolidation mentions "contradiction detection" but doesn't specify at what layer (surface text, embedding, logical). Nor does it define a schema for tracking the assumptions a memory rests on.

Decision (v0.3): Each memory record gains three optional fields — assumptions: list[str], supersedes: mem_id?, superseded_by: mem_id?. Conflict detection runs as a separate cron (scripts/assumption_checker.py, every 12h) that asks an LLM to find conflicts at the assumption-entailment level. Writes proposals tagged proposal + awaits-user-approval + assumption-conflict. Nothing is auto-applied. Supersedes pointers are bidirectional (setting supersedes: X on a new memory automatically sets superseded_by: new_id on X in the same transaction).

Rationale: Inspired by Brian Williams' ATMS / de Kleer truth maintenance work. Dedup and embedding similarity catch textual near-duplicates; they miss conflicts where two memories textually look unrelated but rest on incompatible assumptions ("use account A" vs "use account B"). The assumption layer is the place those conflicts are legible. The original design exploration predates this repo and lives in the maintainer's private memory hub; this v0.3 implements the minimum viable subset of that design.

Out of scope for v0.3: full propositional ATMS with unit-resolution inference; explicit justification graphs. These are v0.4+ if ever needed — for the Phase 0 target audience (personal memory hub, <10k memories), an LLM-as-oracle on the assumption layer is both cheap enough and reliable enough. Formal ATMS matters at scale orders of magnitude larger.

Backward compatibility: existing v0.2 memories have empty assumptions, null supersedes/superseded_by, and the checker skips records with empty assumption lists. Upgrade is zero-friction — old memories simply don't participate in assumption-layer analysis.

11. CLAUDE.md template's behavioral contract

Gap: architecture docs describe what the hub stores but not how agents should use it turn-by-turn. The templates/CLAUDE.md file in this repo is the first attempt at a behavioral contract.

Decision: instruct the agent to (a) list_memories(tag=briefing) at session start, (b) search_memories before asking the user to re-explain past context, (c) create_memory on non-trivial decisions with at least 2-3 tags and an explicit provenance confidence.

Known weakness: this is a prompt-side contract; agents vary in how faithfully they follow CLAUDE.md over long sessions. A stronger enforcement would be harness-level hooks (e.g., auto-inject a briefing at session start) which is out of scope here and belongs in agent-specific config rather than the hub.


What this implies for architecture v2

If I were patching the architecture repo based on these gaps, I'd:

  • Add a "v0 reference wire format" section defining id format, tag semantics, pagination contract, archive behavior
  • Lock the MCP transport story (streamable HTTP) with a simpler fallback path documented as acceptable
  • Publish a canonical provenance schema
  • Add a Chinese-tokenization note in QUICKSTART so Phase 0 implementors don't each reinvent it
  • Publish a canonical Mirror cadence + proposal output schema (gap #9)
  • Define the user-facing digest contract as a first-class architectural surface, including the hard-filter rule for sensitive tags (gap #10)
  • Formalize the agent-side behavioral contract (what create_memory should and should not be called for) so it's not reinvented by every Phase 0 implementor (gap #11)

These are small patches — the architecture repo is already 90% there.