Smriti
Conversational memorial companion with zero-fabrication retrieval architecture
Technical Descriptor: "Zero-fabrication retrieval architecture" · "Structurally grounded, non-generative-fallback design"
Forbidden framing: Hallucination-prevention (implies mitigating a model flaw; Smriti makes fabrication structurally impossible)
01 · Opening Statement
"A grounded AI always has an easy way out of a hard moment: it can simply say it doesn't know. That's technically honest and emotionally hollow — the last thing someone wants from a system built to preserve a person they lost. Smriti was constrained to never invent anything outside the memories it was given, and still had to answer grief with something that felt like care, not a lookup failure. That tension — never fabricate, never feel cold — is what shaped the architecture."
02 // Context & Problem
Smriti is a conversational AI trained on a specific person's actual messages and writing, built to let someone continue a form of conversation with them after they're gone.
The Structural Risk
The obvious engineering risk in a product like this is fabrication — an AI inventing memories or words the person never said. Smriti closes that risk structurally rather than mitigating it: it only ever draws from the memory data assigned to it, and when a query falls outside that data, it doesn't guess — it names the gap and offers to add the new information as a memory going forward.
A system that's honest about the limits of what it knows can easily default to sounding like it's hitting those limits — clinical, hedging, a database returning 'no result.' For a person in grief, that reads as coldness at the exact moment they need presence. The real design problem wasn't accuracy. It was building a system with zero tolerance for invention that still had to recognize sadness and respond with warmth, not caveats.
03 // Request Execution Topology
7-Node Verified LifecycleInteractive system graph: Verified branch fork at Node 1, asymmetric tone latching at Node 4, and 600ms resilience backoff at Node 6.
04 // Grounding & Retrieval
- Memories embedded with
gemini-embedding-2-previewand persisted as vector arrays in Firestore. - User queries embedded at request time via
app/api/chat/route.ts. - Retrieval uses Application-level cosine similarity computed in-process over memory array (top 5 selected).
- Fallback mechanism: Keyword-overlap fallback for memories missing precomputed embeddings.
Decision · In-Process Cosine Similarity vs. Managed Vector Index
Verified ArchitectureChosen Implementation
In-process cosine similarity over bounded array in app/api/chat/route.ts
Rejected Alternative
Managed vector index (Pinecone, Weaviate, or Firestore native vector search)
Why: A single memorial memory vault is bounded (dozens to a few hundred entries) — not the volume a distributed vector index earns its keep on.
Tradeoff: Doesn't scale indefinitely; would require redesign if memory vault size grew by orders of magnitude.
Decision · Top-5 Retrieval vs. Full-Context Stuffing
Verified ArchitectureChosen Implementation
Top-5 vector retrieval with keyword fallback
Rejected Alternative
Stuffing entire memory vault into context window on every prompt
Why: Avoided latency/token costs and known instruction-following degradation in long contexts — exactly where anti-fabrication and tone directives live.
Tradeoff: Top-5 can under-retrieve if query wording diverges significantly from memory phrasing.
05 // Memory-Write Gating & State Isolation
Zero unconfirmed writes to Firestore. Gating writes is where anti-fabrication actually starts.
Cross-turn integrity: Model instructed never to treat its own earlier, unverified conversational warm acknowledgments as confirmed memories in subsequent turns.
Decision · Explicit UI Action Cards vs. Conversational Auto-Saving
Verified ArchitectureChosen Implementation
Explicit interactive card approval (Save / Review-Edit / Dismiss)
Rejected Alternative
Silently auto-saving details surfaced conversationally
Why: Anything saved becomes authoritative for all future retrieval. A silently saved mistake isn't just wrong once; it becomes 'true' in every later conversation.
Tradeoff: Introduces friction — sacrifices some 'it just knows me' magic for absolute data integrity.
06 // Deterministic Tone & Language Analysis
Engineered in lib/language-analysis.ts — not embedding-based.
Per-turn dynamic re-evaluation with an explicit prompt directive against dropping into English mid-Hinglish conversation.
Decision · Rule-Based Regex Classifier vs. Embedding Sentiment Similarity
Verified ArchitectureChosen Implementation
Deterministic regex script and keyword matching
Rejected Alternative
Embedding similarity against labeled emotional exemplars
Why: Interpretability and debuggability. In grief-adjacent interactions, a verifiable rule is auditable and instantly correctable. Also bypasses the lack of labeled training datasets for code-switched Hinglish grief dialogue.
Tradeoff: Brittle outside keyword coverage — lacks generalization to unanticipated phrasing.
07 // Resilience Engineering (lib/gemini.ts)
Applying distributed systems failover rigor to LLM API dependencies:
Up to 2 attempts per model with a 600ms backoff on retryable errors (503, 429, RESOURCE_EXHAUSTED). Maximum worst-case execution: 8 attempts before terminating.
Tradeoff: Fallback models aren't guaranteed to match primary model voice and pacing under identical prompts, risking subtle tonal shifts mid-session. Also introduces real latency in worst-case cascades.
08 // Trust Sequencing
Phase 1: Text-only launch. Phase 2: Real-time voice synthesis.
Risk staging. Text allows reflection and pause; synthesized voice happens in real-time with near-zero tolerance for discordant notes. Core anti-fabrication had to be proven before voice expansion.
Tradeoff: Text-only appeared incomplete to outside observers unaware of the trust-staging architecture.
09 // Isolation by Design
Rules policy: Default-deny on every unmatched path (annotated as "Zero Insecure Defaults").
Nested Ownership Chains:
/users/{userId}/memorials/{memorialId}/memories/{memoryId} (request.auth.uid == userId)
/conversations/{conversationId}/messages/{messageId} (request.auth.uid == userId)
/api/chat calls verifyAuthToken against Firebase ID token before request parsing or retrieval logic executes.
10 // Lessons & Engineering Realities
Baseline Viability
The conversational companion holds up in real testing as an experience users are willing to interact with honestly.
The Emotional Trust Tension
A concern surfaced that making the companion feel too faithful makes it harder to serve grief recovery — users relate to continuity rather than memory.
Tension: The zero-fabrication architecture protects factual trust, but leaves open the question of emotional trust — whether fidelity helps someone move through grief or suspends them within it.
Product Maturity Focus
Memory-grounding and gated save mechanics are unusually visible; surrounding UI is sparse compared to modern chat platforms. Reflects deliberate investment in retrieval integrity over superficial polish.
11 // Final Outcome
Smriti proves out its specific architectural claim: a conversational system can be architected to never fabricate, maintain warmth, degrade gracefully across LLM failure modes, and enforce least-privilege security by default. It does not claim product-market fit or resolved emotional-safety boundaries — solving the human dilemma remains the true frontier.
Forward Look: Typing-behavior mimicry (matching punctuation rhythm, message pacing, and typing cadence).
Must resolve the emotional-safety dilemma first: higher imitation precision amplifies the risk of unhealthy attachment.