Manish Prajapati_
Flagship ProjectStatus: Deployed & functional, informal validation
Live System Demo↗

Smriti

Conversational memorial companion with zero-fabrication retrieval architecture

Technical Descriptor: "Zero-fabrication retrieval architecture" · "Structurally grounded, non-generative-fallback design"

Forbidden framing: Hallucination-prevention (implies mitigating a model flaw; Smriti makes fabrication structurally impossible)

01 · Opening Statement

"A grounded AI always has an easy way out of a hard moment: it can simply say it doesn't know. That's technically honest and emotionally hollow — the last thing someone wants from a system built to preserve a person they lost. Smriti was constrained to never invent anything outside the memories it was given, and still had to answer grief with something that felt like care, not a lookup failure. That tension — never fabricate, never feel cold — is what shaped the architecture."

02 // Context & Problem

Smriti is a conversational AI trained on a specific person's actual messages and writing, built to let someone continue a form of conversation with them after they're gone.

The Structural Risk

The obvious engineering risk in a product like this is fabrication — an AI inventing memories or words the person never said. Smriti closes that risk structurally rather than mitigating it: it only ever draws from the memory data assigned to it, and when a query falls outside that data, it doesn't guess — it names the gap and offers to add the new information as a memory going forward.

A system that's honest about the limits of what it knows can easily default to sounding like it's hitting those limits — clinical, hedging, a database returning 'no result.' For a person in grief, that reads as coldness at the exact moment they need presence. The real design problem wasn't accuracy. It was building a system with zero tolerance for invention that still had to recognize sadness and respond with warmth, not caveats.

03 // Request Execution Topology

7-Node Verified Lifecycle

Interactive system graph: Verified branch fork at Node 1, asymmetric tone latching at Node 4, and 600ms resilience backoff at Node 6.

1. Message Received2. Query Embedded3. Memory Retrieval4. Tone & Script Detected5. Grounded + Toned Prompt6. Generation & Cascade7. Response Returned

04 // Grounding & Retrieval

  • Memories embedded with gemini-embedding-2-preview and persisted as vector arrays in Firestore.
  • User queries embedded at request time via app/api/chat/route.ts.
  • Retrieval uses Application-level cosine similarity computed in-process over memory array (top 5 selected).
  • Fallback mechanism: Keyword-overlap fallback for memories missing precomputed embeddings.

Decision · In-Process Cosine Similarity vs. Managed Vector Index

Verified Architecture

Chosen Implementation

In-process cosine similarity over bounded array in app/api/chat/route.ts

Rejected Alternative

Managed vector index (Pinecone, Weaviate, or Firestore native vector search)

Why: A single memorial memory vault is bounded (dozens to a few hundred entries) — not the volume a distributed vector index earns its keep on.

Tradeoff: Doesn't scale indefinitely; would require redesign if memory vault size grew by orders of magnitude.

Decision · Top-5 Retrieval vs. Full-Context Stuffing

Verified Architecture

Chosen Implementation

Top-5 vector retrieval with keyword fallback

Rejected Alternative

Stuffing entire memory vault into context window on every prompt

Why: Avoided latency/token costs and known instruction-following degradation in long contexts — exactly where anti-fabrication and tone directives live.

Tradeoff: Top-5 can under-retrieve if query wording diverges significantly from memory phrasing.

05 // Memory-Write Gating & State Isolation

Zero unconfirmed writes to Firestore. Gating writes is where anti-fabrication actually starts.

UI Action Cards:SaveReview-EditDismiss

Cross-turn integrity: Model instructed never to treat its own earlier, unverified conversational warm acknowledgments as confirmed memories in subsequent turns.

Decision · Explicit UI Action Cards vs. Conversational Auto-Saving

Verified Architecture

Chosen Implementation

Explicit interactive card approval (Save / Review-Edit / Dismiss)

Rejected Alternative

Silently auto-saving details surfaced conversationally

Why: Anything saved becomes authoritative for all future retrieval. A silently saved mistake isn't just wrong once; it becomes 'true' in every later conversation.

Tradeoff: Introduces friction — sacrifices some 'it just knows me' magic for absolute data integrity.

06 // Deterministic Tone & Language Analysis

Engineered in lib/language-analysis.ts — not embedding-based.

▸Regex Unicode-range script detection (Devanagari, Bengali, Tamil, Telugu)
▸Hardcoded Hinglish token dictionary
▸Keyword-list emotional state matcher (e.g. "yaad aati" -> Grieving & Longing; "purane din" -> Nostalgic & Reminiscing)

Per-turn dynamic re-evaluation with an explicit prompt directive against dropping into English mid-Hinglish conversation.

Decision · Rule-Based Regex Classifier vs. Embedding Sentiment Similarity

Verified Architecture

Chosen Implementation

Deterministic regex script and keyword matching

Rejected Alternative

Embedding similarity against labeled emotional exemplars

Why: Interpretability and debuggability. In grief-adjacent interactions, a verifiable rule is auditable and instantly correctable. Also bypasses the lack of labeled training datasets for code-switched Hinglish grief dialogue.

Tradeoff: Brittle outside keyword coverage — lacks generalization to unanticipated phrasing.

07 // Resilience Engineering (lib/gemini.ts)

Applying distributed systems failover rigor to LLM API dependencies:

Cascade Chain:gemini-3.1-flash-lite→gemini-3.6-flash→gemini-3.7-flash→gemini-flash-latest

Up to 2 attempts per model with a 600ms backoff on retryable errors (503, 429, RESOURCE_EXHAUSTED). Maximum worst-case execution: 8 attempts before terminating.

Tradeoff: Fallback models aren't guaranteed to match primary model voice and pacing under identical prompts, risking subtle tonal shifts mid-session. Also introduces real latency in worst-case cascades.

08 // Trust Sequencing

Phase 1: Text-only launch. Phase 2: Real-time voice synthesis.

Risk staging. Text allows reflection and pause; synthesized voice happens in real-time with near-zero tolerance for discordant notes. Core anti-fabrication had to be proven before voice expansion.

Tradeoff: Text-only appeared incomplete to outside observers unaware of the trust-staging architecture.

09 // Isolation by Design

Rules policy: Default-deny on every unmatched path (annotated as "Zero Insecure Defaults").

Nested Ownership Chains:

/users/{userId}/memorials/{memorialId}/memories/{memoryId} (request.auth.uid == userId)

/conversations/{conversationId}/messages/{messageId} (request.auth.uid == userId)

/api/chat calls verifyAuthToken against Firebase ID token before request parsing or retrieval logic executes.

10 // Lessons & Engineering Realities

Lesson 01

Baseline Viability

The conversational companion holds up in real testing as an experience users are willing to interact with honestly.

Lesson 02

The Emotional Trust Tension

A concern surfaced that making the companion feel too faithful makes it harder to serve grief recovery — users relate to continuity rather than memory.

Tension: The zero-fabrication architecture protects factual trust, but leaves open the question of emotional trust — whether fidelity helps someone move through grief or suspends them within it.

Lesson 03

Product Maturity Focus

Memory-grounding and gated save mechanics are unusually visible; surrounding UI is sparse compared to modern chat platforms. Reflects deliberate investment in retrieval integrity over superficial polish.

11 // Final Outcome

Smriti proves out its specific architectural claim: a conversational system can be architected to never fabricate, maintain warmth, degrade gracefully across LLM failure modes, and enforce least-privilege security by default. It does not claim product-market fit or resolved emotional-safety boundaries — solving the human dilemma remains the true frontier.

Forward Look: Typing-behavior mimicry (matching punctuation rhythm, message pacing, and typing cadence).

Must resolve the emotional-safety dilemma first: higher imitation precision amplifies the risk of unhealthy attachment.