Skip to content
Sagar Thakkar
← All writing
Architecture 17 min read
rag architecture banking llm-observability governance ai-architecture

Production RAG for a Tier-1 Bank: An Architecture Walkthrough

Bank RAG projects die at second-line risk review, not at the model. The retrieval architecture that survives an audit, and why it was chosen.

Sagar Thakkar
Sagar Thakkar
AI Systems Architect
Poster reading THE DEMO PASSED, RISK SAID NO: an answer card whose chain of source citations has been cut, blank tags falling into an empty document tray
TL;DR

Bank RAG projects fail at risk review, not at the model. The bank is not buying a chatbot, it is buying a defensible answer system. What clears the bar is compliance-aware orchestrated retrieval with full-trace observability, so every answer carries an audit artifact the second line can read.

For platform leads and architects standing up retrieval-augmented generation inside a regulated bank, what breaks, what the auditors ask, and the pattern that survives both.


TL;DR

  • Bank RAG projects tend to die between the notebook demo and second-line risk review. Not because the model was wrong, because there was no audit artifact.
  • The reframe that unblocks it: the bank is not buying a chatbot. It is buying a defensible answer system. The LLM is the cheapest, most replaceable piece.
  • What clears the bar: compliance-aware orchestrated RAG + full-trace observability. Beats naive RAG (no audit) and black-box managed services (opaque chunking, opaque logs).
  • Anchored on a composite Indian Tier-1 bank profile (DPDP Act + RBI), assembled from patterns common to several institutions rather than any one of them. Same pattern moves to UK · EU · Singapore by swapping the regulator + residency layer.
  • RAG is a family, not an architecture. The candidates are eliminated by one question, not by benchmark score: can the retrieval path be reconstructed by someone who was not in the room?
  • For: architects + platform leads in banking, insurance, or any regulated domain who must answer to risk and ship.

Context

Picture an Indian Tier-1 bank. It wants a natural-language Q&A assistant for branch staff, operations teams and compliance officers. Answers must be grounded in regulator master directions, recent circulars, and internal product and policy documents. The notebook pilot impressed leadership. The risk and compliance review did not move. Every answer had to cite its source. Every call had to be logged and retained. No data could leave the country. The system had to refuse rather than guess when retrieval was thin. The gap between an impressive demo and a system that passes second-line review is what this case study is about.

Composite, not a client. Nothing here describes a single institution, an engagement, or anyone’s internal system. The profile is assembled from constraints that are public (regulator directions) and patterns that are common across the sector. Anchored on India because the constraints are at their sharpest there: regulator-mandated residency, DPDP obligations on processing, and long audit retention all bind at once. The same pattern moves to the UK (PRA and FCA, UK GDPR), the EU (EBA, GDPR, EU AI Act) and Singapore (MAS). Only the regulator and the residency layer change. The architecture below does not change; the labels do.

AttributeValue
IndustryBanking, Indian Tier-1 (representative; private + PSU patterns folded together)
Scale~50k staff users · ~5 to 10k queries/day in pilot, scaling to ~50k/day · ~30k documents (RBI circulars + internal policy + product manuals)
GeoIndia. Residency is set by the stack the bank answers to: supervisory expectations on regulated data, RBI master directions on IT outsourcing, DPDP and its 2025 Rules, and a board-approved internal data policy that is stricter than all of them. Internal 7-year audit retention
Existing stackIn-region cloud tenant (Azure / AWS, in-India region) · SharePoint + file-share document stores · Oracle DB for staff identity · existing data lake (Databricks / Synapse class)
TeamPlatform engineering team + a small ML team (3 to 6 people) · limited LLM-ops maturity at start · strong SRE + risk-ops culture

The Problem

A notebook RAG demo answers questions well. A bank RAG system has to answer well and then prove itself, for every single response. Which circular or policy clause backed the answer. That no data left the country. That the call was logged and retained for seven years. That latency stayed under SLA at peak branch hours. That gap, between an impressive demo and a system that passes second-line risk and compliance review, is where this kind of project stalls. Closing it is what this case study is about.


Hard Constraints

#ConstraintWhy it mattersHard limit
1Data residencyDPDP Act + RBI master directions on IT outsourcingAll inference + storage in India region. No cross-border egress. No managed APIs that ship tokens offshore.
2AuditabilityEvery answer must be defensible to internal risk + RBI on demandEvery LLM call + retrieved chunks + final response logged. 7-year retention. Tamper-evident.
3LatencyBranch staff query between customer interactions; slow = abandonedTime to first token, p95 < 2s (redaction, retrieval, rerank, gate, prompt assembly, first token). The complete answer streams in over the following seconds; see the latency section for why the SLI is first token rather than completion.
4Cost ceilingMust scale to ~50k queries/day without per-quarter budget review< $2 per 1k queries in run cost: inference, retrieval, rerank, storage. Platform headcount sits in the programme budget, not in the per-query unit, or the unit stops being comparable to the vendor quote it is judged against
5Grounding / no hallucinationWrong policy answer → regulator + reputational hitEvery answer cites source chunks; system refuses when retrieval confidence is low rather than guessing
6PII handlingInternal policy + escalation docs sometimes reference customer casesPII detect + redact before embedding; no PII in prompts, completions, or logs

Alternatives Considered

Option A, Naive RAG (the demo that wowed leadership)

  • How: single-shot retrieve-then-generate, assembled from the framework’s stock retrieval chain, a managed embedding and completion API, and an off-the-shelf vector database.
  • Pro: Days to build. The notebook that got the budget approved.
  • Con:
    • Residency fail: managed embedding and completion APIs route via US or EU regions, so document text and staff queries leave the country. What makes that a breach here is a combination, not one clause. Supervisory expectations on where regulated data is stored and processed. The transfer restrictions attaching to a Significant Data Fiduciary. And the bank’s own board-approved data policy, usually the strictest of the three. No single clause does the work; the stack of them does.
    • Audit fail: no per-call trace; cannot tell risk why an answer was given.
    • Grounding fail: no confidence gate → the model hallucinates politely instead of refusing.
  • Verdict: ❌ Rejected. Fails constraints 1, 2, 5 and 6 on day one. Useful only for the pilot that justified the project, not what ships.

Option B, Fully managed RAG service (vendor turnkey)

  • How: the cloud vendor’s bundled RAG service. Ingestion, chunking, retrieval and generation sit behind a single API, in-region where the vendor offers it.
  • Pro:
    • In-region offering exists (South India for the frontier vendor, Mumbai for the other), so residency can be satisfied by pinning the region.
    • Lowest ops burden; vendor scales.
  • Con:
    • Audit fail: the bundle is a black box. Chunking strategy is opaque. Log format and retention are vendor-controlled, auditors will not accept “trust the vendor.”
    • Grounding fail: confidence gating + refusal logic is buried inside the service; cannot tune it to bank-specific risk tolerance.
    • Lock-in: model choice + pipeline shape locked to one vendor. Renegotiation leverage gone.
    • Model availability lags in-region. New models land in the largest regions first, so an in-country deployment usually picks from a narrower and slightly older menu. Check the region list for the specific models you need before committing, and treat portability as insurance against the gap rather than a nice-to-have.
  • Verdict: ❌ Rejected. Solves residency but breaks audit + grounding + portability. Black-box answers do not pass second-line review.

Option C, Compliance-aware orchestrated RAG + full-trace observability ✅

  • How: a composable, in-region stack:
    • Orchestrator for the pipeline, prompt assembly and gate wiring.
    • In-region model endpoint, with a self-hosted open-weight model as the portable fallback.
    • In-region vector index, on managed Postgres with pgvector, or self-hosted inside the bank VPC.
    • PII redaction before embedding and again before prompting.
    • Cross-encoder reranker, which produces the relevance score the gate acts on.
    • Confidence gate that refuses when that score is low.
    • Tracing on every step: query, retrieved chunks, prompt, completion, cost, latency.
  • Pro:
    • Residency provable per component.
    • Every answer is debuggable from the trace, the trace is the audit artifact.
    • Model-portable: swap the LLM endpoint without redesigning the system.
    • Confidence gate trades refusals for ungrounded answers, and both sides of that trade are measurable: refusal rate on one side, groundedness and faithfulness on the other.
  • Con:
    • Higher build + ops burden, you own the stack instead of buying the service.
    • Added latency from redaction + gate (worth measuring).
  • Verdict: ✅ Chosen. Only option that clears all six constraints together.

Scoring Against Hard Constraints

ConstraintA. Naive RAGB. Managed bundleC. Orchestrated ✅
1. Residency❌ tokens leave India✅ in-region✅ in-region, per-component
2. Audit (7y, tamper-evident)❌ no trace❌ vendor-format, opaque chunking, log retention outside the bank’s control✅ full trace, owned format
3. Latency, TTFT p95 < 2s✅ (no gates to pay for)✅⚠️ achievable; redaction, rerank and gate together add roughly 350 to 400 ms
4. Cost < $2 / 1k⚠️ depends on tier⚠️ bundle pricing✅ controllable per component
5. Grounding / refusal❌ no gate❌ gating logic inside the service, not tunable to the bank’s risk appetite✅ explicit confidence gate
6. PII redaction❌ none⚠️ partial / opaque✅ explicit step, in pipeline

Decision evidence: C is the only column with no red at all. A is rejected on four constraints, B on two.


Which RAG, and Why This One

RAG is a family, not an architecture. The word now covers a dozen retrieval mechanisms whose only shared property is that something gets fetched before the model writes. Inside a regulated institution they do not fail first on accuracy. They fail on whether the retrieval path can be reconstructed, months later, by someone who was not in the room.

So the filter is one question, applied before any benchmark is opened. Can this variant emit a per-answer trace that a second-line risk function can read, defend and reproduce? A mechanism that tops a leaderboard and cannot show its work is eliminated before the accuracy conversation starts.

CandidateAudit traceVerdict
Naive RAG: one query, one top-k set, one generationClean, fully reconstructibleRejected on grounding and residency, not on traceability. Simple was never the problem.
Hybrid retrieval (BM25 and dense) with index-time context and a cross-encoder rerankClean, every stage deterministic, the rerank score comes from a documented functionSelected.
GraphRAG: cross-document reasoning over entities and relationsPartial, provenance resolves to a community summaryRejected. See below, because the usual objection to it is now the wrong one.
Agentic or multi-hop RAG: the model plans its own retrieval sequenceOpaque, the path is non-deterministicRejected for this surface. Viable only when every hop is logged as its own audit event, which is a different architecture.
Long context instead of retrieval: preload the corpus, skip the fetchOpaque, there is no retrieval eventRejected at this corpus size. Nothing exists for a citation to point at.

Seven further variants were considered and eliminated on the same axis: HyDE, RAPTOR, Self-RAG, corrective RAG, late-interaction retrieval, query rewriting and summarise-then-retrieve. Late interaction is the one worth revisiting, since it loses on index size rather than on auditability. The short version for the rest. Anything that retrieves against synthetic text fails. So does anything that cites a generated summary rather than a clause, or lets the model attest to its own grounding. Each one breaks the reconstruction test by construction.

GraphRAG deserves the longer answer, because its usual objection has expired. In 2024 it lost on cost: every entity and relationship needed a model-generated summary before the first question could be asked. That is no longer true. Microsoft’s own LazyGraphRAG work defers summarisation to query time and reports indexing cost on par with plain vector RAG. The elimination survives the cost fix, which makes it an architecture decision rather than a budget one. A GraphRAG answer’s provenance resolves to a community summary, which is itself model output. A regulator asking which paragraph said this receives a paraphrase.

The selected stack wins because every stage is boring in the way that matters. BM25 is a documented function over terms that actually appear in the document. The contextual augmentation runs offline. What the system knew about a chunk at index time is therefore a written artifact under version control. A re-index becomes a reviewable change rather than a silent behaviour shift. Anthropic reports this combination cutting top-20 retrieval failure from 5.7 percent to 1.9 percent, at a one-time build cost of roughly one dollar per million document tokens. Most decisive: the rerank emits a relevance score from a documented function, computed over the query and the passage. Cross-encoders output uncalibrated logits. The threshold has to be fitted against a labelled set before the gate means anything, and that fitting step is itself an artifact someone can review. A confidence signal that is instead a token the model emitted about its own output cannot be calibrated, reviewed, or defended.

When not to pick it. If the questions require composing an answer across documents that no single chunk connects, this stack refuses far more often than it answers. The honest fix is orchestrated multi-step retrieval with every hop logged as its own audit event, not a cleverer single-shot retriever.


Decision

Chose Option C. The binding constraints are audit and residency, not build speed. Only C lets you stand in front of second-line risk with a per-response trace and an in-region guarantee.

The reframe that unblocked the project: the bank was not buying a chatbot. It was buying a defensible answer system. Once that landed in the steering committee, observability stopped being a “nice to have” line item and became the core deliverable. The traces are the audit artifact risk had been asking for. The LLM became the cheapest, most replaceable component in the stack; the grounding gate, the redaction step, and the trace store became the system. Framed this way, a black-box vendor cannot win, a black box, by definition, has no audit artifact.


Architecture

Request-time flow

Request-time flow: authenticated query, PII redaction, hybrid retrieval, confidence gate, in-region model, cited response, and an asynchronous trace write to the audit store

Ingestion-time flow (offline, daily/event-driven)

Ingestion flow: source documents normalised, PII redacted, chunked with index-time context, embedded, and written to the in-region vector index with a lineage row

Key components, each maps to the constraint it satisfies

ComponentRoleSatisfies
API gateway + AuthN/AuthZIdentity-bound queries; per-user rate limitsaudit
LangChain orchestratorControl plane, pipeline, prompt assembly, gate wiring2, 5
PII detect + redactStrips/masks customer data before embedding + before prompts6
In-region vector DBpgvector on managed Postgres (in-India region) primary, reuses existing RDBMS + simpler ops; self-hosted Qdrant in VPC as fallback if scale demands ANN-native performance1
Cross-encoder rerankerReorders retrieved candidates and emits the relevance score the gate consumes. Replicated separately, since it is the largest pre-generation cost3, 5
Confidence gateRefuses below the score threshold, so no ungrounded policy answers reach a user5
In-region model endpointFrontier vendor in South India primary; in-region self-hosted open-weight model as portable fallback1, 3
Cited responseEvery answer carries source chunk IDs + doc-version2, 5
LangfuseTraces query → chunks → prompt → completion → cost → latency. This IS the audit log.2
Lineage storeDoc-version → chunk → answer traceability for regulator2

Request-time data flow

  1. Authenticated query hits the orchestrator (identity attached for per-user logs + access policy).
  2. PII redaction runs on the inbound query.
  3. Hybrid retrieval pulls candidates from the in-region vector index: BM25 and dense, rank-fused.
  4. Cross-encoder rerank of the top 50 candidates, which produces the relevance score everything downstream depends on.
  5. Confidence gate reads that rerank score. Below threshold, the system refuses and escalates to a human reviewer, logged as a refusal event rather than a failure. Above it, continue.
  6. Prompt assembled with cited chunks and system rules, sent to the in-region model endpoint, streamed back with inline citations to document version and section.
  7. Before the answer streams, the trace record is appended to a durable in-region queue (single-digit milliseconds, the only audit work on the critical path). From there it is written asynchronously to the audit store. The record holds inputs, retrieved candidates with both retrieval and rerank scores, prompt, output, cost, latency and the refusal flag, retained seven years. A lineage row ties the response to the exact document versions used.

Does It Hold at Scale?

The constraints above are numbers, and a number nobody services is a wish. This section sums them.

Start with the least intuitive part: at this size the system is not a throughput problem. Fifty thousand queries a day across a ten hour branch window averages about 1.4 queries per second. Branch traffic is spiky, so assume a four times peak and call it 6 queries per second sustained. That is a small number. The engineering difficulty sits in the latency budget and the write path for evidence, not in serving concurrent requests.

Latency budget, in the units that make it checkable

State a latency target in milliseconds per stage and nobody can tell whether it is achievable. State it in time to first token and tokens per second and the arithmetic checks itself. That distinction matters here, because the naive version of this budget hides an impossible number.

The SLI is time to first useful token, p95. Not time to a complete answer. A streamed answer is readable while it is still being written, and a branch user reacts to the first line. The indicator has to match what the user experiences. The SLO is 2 seconds at p95, with a monthly error budget expressed as 5 percent of queries above 2 seconds and 1 percent above 3.5 seconds. Two thresholds, because a tail that is merely slow and a tail that looks broken are different failures.

StageBudget (p95)Note
PII detect and redact on inbound query60 msSmall input, pattern plus model hybrid
Hybrid retrieval (BM25 and dense, rank-fused)80 msRoughly 150k chunks, assuming around 800 tokens per chunk with overlap across 30k documents of mixed length. Well inside a single node at that size
Cross-encoder rerank of top 50250 msThe largest pre-generation cost, and the one worth a dedicated replica
Confidence gate40 msReads the rerank score; on the first-token path because a refusal short-circuits generation
Prompt assembly40 msCited chunks plus system rules
Time to first token400 msIn-region endpoint, cold-path excluded
Total to first token875 msAgainst a 2 second SLO, which leaves real headroom
Full completion, roughly 400 output tokens+3.6 to 5.0 sAssuming 80 to 110 tokens per second. That range is an assumption to be replaced with a measurement from your own endpoint on day one, because every number below it moves when it moves
Trace write to the audit store~5 msA synchronous append to a durable in-region queue, then an asynchronous write to the store. Only the append sits on the critical path, and this placement is the point

The completion row is the one people get wrong. Four hundred output tokens inside the same two second budget would have to fit in the 1,125 ms left after everything upstream. That is about 356 tokens per second sustained. Written as a single end-to-end number it looks reasonable; written in tokens per second it is obviously false. That is the whole argument for using the standard units: the wrong vocabulary lets a wrong number pass review.

So the honest shape is a fast first token and a complete answer arriving over the following few seconds. If the requirement really is a complete answer inside two seconds, the fix is a shorter answer format, not a faster model.

The trace write deserves its own emphasis. Putting the audit write on the critical path is the most common way a compliant design becomes an unusable one. Append the trace to a durable queue before the answer is released. Write it to the audit store asynchronously from there. Reconcile trace completeness as a separate job that alarms on gaps. An audit log with holes is a worse outcome than a slow answer, so the reconciliation job is not optional. Trace completeness is itself an SLI: the percentage of answers with a complete, replayable trace, with a target of 100 percent and any gap treated as an incident.

What scales, and what does not

Replicas are needed in exactly two places: the reranker and the generation endpoint. The vector database is not one of them. Thirty thousand documents is a small index by any modern measure, and pgvector on managed Postgres carries it without sharding. Choosing an expensive distributed vector store for this corpus is buying a solution to a problem the corpus does not have.

Re-indexing is a batch job, not a stream. Circulars arrive daily, not by the second, so a nightly rebuild of changed documents is sufficient. Contextual retrieval is what makes that affordable: the chunk augmentation runs offline, so the expensive step happens once per document version rather than once per query.

The cost curve, and what breaks first

Assume roughly 2,500 input and 400 output tokens per query. Assume a mid-tier model at list prices near $0.15 per million input tokens and $0.60 per million output tokens. Two caveats before the arithmetic. List pricing moves. And provisioned in-region capacity is usually bought as reserved throughput rather than per token, which changes the shape of the bill even when it leaves the conclusion intact. A thousand queries is then 2.5 million input and 0.4 million output tokens, about $0.38 and $0.24 respectively, so roughly $0.62 in model cost. Retrieval, reranking and loaded operations bring it to a little over a dollar. The stated ceiling of two dollars per thousand queries holds at 50k a day with room to spare. Publish the assumption rather than the conclusion, because the conclusion changes the moment the model tier does.

At ten times that volume two things break, and neither is the one people expect. First, provisioned throughput on the in-region model endpoint becomes the binding limit. In-country capacity is scarcer than global capacity, and routing elsewhere to relieve it would breach residency. Second, the audit store grows into a real estate. Size a trace before sizing the store. A trace holds the assembled prompt (roughly 10 KB at 2,500 tokens) and the completion (about 1.6 KB). It holds candidate identifiers with their retrieval and rerank scores, not the chunk bodies, which live once in the lineage store. Call it 15 KB. At the starting volume, 50k traces a day held for seven years is roughly 128 million records and about 2 TB. At ten times the volume it is roughly 1.28 billion records and about 19 TB before indexes. That is a different class of problem: partitioning, tiering and query performance on the audit path all become design decisions rather than defaults. Budget for tiered storage from day one, because migrating an audit log after the fact requires proving the migration did not alter it.

Availability inside the boundary

High availability has to happen inside the residency boundary, which is the genuinely hard part. Two in-country regions, active and standby, with the audit store replicated synchronously and the vector index rebuilt asynchronously from source documents.

That shape implies specific objectives, and a design that does not state them has not been reviewed. RTO of 15 minutes for the answer path, because the index is rebuildable from source and the standby model endpoint is warm. RPO of zero for the audit store and best-effort for everything else. Zero is only honest because the trace is durable in the queue before the answer is released. A trace written only after the answer is sent has a loss window, and a failover inside that window loses evidence. Losing an answer is an inconvenience. Losing its trace is a reportable gap. The failover is exercised quarterly, because an untested failover is a diagram. Cross-border failover is not a degraded mode here. It is a breach.

Of the six constraints, one did the actual work, and it is worth being precise about which. Residency eliminated the naive option, but the managed bundle satisfies residency perfectly well once the region is pinned. What eliminated the bundle was auditability: an opaque chunking strategy and a vendor-controlled log format cannot produce a trace the second line can defend. Grounding failed in both rejected options too, but grounding is fixable by adding a gate, whereas auditability is a property of who owns the pipeline. That is the test for load bearing: not which constraint failed most often, but which one could not be bought, tuned or added later. Latency and cost shaped the build; they did not decide it. Naming the load-bearing constraint matters, because a team that treats all six as equal will trade away the one that was never negotiable.


Consequences

What this gets us

OutcomeResult
Per-response audit trailAudit preparation stops being an archaeology project. Where evidence previously had to be assembled by hand across systems, it becomes a trace export. The size of that saving depends entirely on what the institution was doing before, so it is not a number this pattern can promise in advance.
Residency guaranteeProvable per-component. No token leaves India region. Auditors verify via region pinning + Langfuse storage location + lineage records.
Debuggable failuresA bad answer becomes a readable trace: which chunk misled the model, and whether the fix belongs in retrieval or in the prompt. In this shape of system the retrieval layer is where most defects are found, which is an argument for instrumenting it first rather than a law.
Grounding, measuredThe gate refuses on a low rerank score rather than guessing, which bounds retrieval quality. Generation faithfulness is a separate property and needs its own measurement: groundedness and faithfulness scored against a golden dataset, tracked as an SLI, with a regression run on every index or model change. A gate alone does not make a model faithful, and claiming otherwise is how this gets oversold.
Model portabilitySwap the managed endpoint for a self-hosted open-weight model without redesigning the pipeline. Becomes leverage in vendor renegotiation and a real Plan B.

What we accept (tradeoffs)

TradeoffCostWhy worth it
Higher ops burden~1 to 2 FTE platform engineers vs ~0 with managed bundleOwning the stack = owning the audit story. Auditors do not accept “vendor handles it.”
Added latency from gate + redaction~100 to 200 ms added to p95Still inside the 2-second budget. Refusal-on-low-confidence is the headline feature, not overhead.
Engineering time to wire Langfuse~3 to 4 weeks of platform workThe trace is the audit artifact. Non-optional.
Refusal rate as UX costDesign target of 5 to 10% of queries, not a measured result. The real figure falls out of where the threshold is fitted, and the queue has to be staffed for whatever it turns out to beBetter than a confident wrong answer landing on the regulator’s desk.

What we monitor (SLOs)

MetricTargetWhat it tells us
Retrieval precision @ 5Target ≥ 0.80, measured against a labelled golden setBelow it, chunking and embedding need work before anything else does. In this shape of system it is the largest quality lever, which is a judgement from the architecture rather than a measured ranking: everything downstream consumes what retrieval produced.
Citation coverage100% of non-refused answers carry doc-version + section citationAudit is always answerable. No “where did this come from?” moments.
Time to first token, p95< 2 sBranch staff do not abandon the assistant. This is the SLI; completion streams after it.
Time to first token, p99< 3.5 sTail bounded, no silent stragglers.
Cost per 1k queries (loaded)< $2Per-quarter budget review not triggered; unit economics stay honest.
Refusal rate5 to 10% (healthy band)Above 15% → retrieval issue. Below 2% → gate too permissive, hallucination risk. The gate has its own SLO.
PII leak rate (sampled traces)0 / 1000 sampledRedaction is doing its job.
Trace completeness100% of calls have full trace storedAudit invariant, non-negotiable.

Governance / Audit Considerations

AspectDetail
Primary frameworks (India anchor)RBI Master Direction on Outsourcing of IT Services (control, auditability and access to outsourced processing) · DPDP Act 2023 and the 2025 Rules (processing and consent, plus transfer restrictions on specified categories of personal data for Significant Data Fiduciaries, which a Tier-1 bank is). It is not a blanket localisation mandate, and it is not nothing either · RBI Information Technology Framework · RBI Cybersecurity Framework for Banks · internal model-risk policy
Mirror frameworks (geo extensions)UK: UK GDPR + PRA SS1/23 on model risk management · EU: GDPR + EU AI Act (limited-risk internal tool) · Singapore: PDPA + MAS FEAT principles
Risk classificationInternal staff-facing assistant · no customer-facing output · no automated decision-making · advisory not directive → limited-risk under EU AI Act analog · low-to-medium under internal model-risk taxonomy. Document the reasoning in the model card.
Audit artifacts producedLangfuse traces (query → chunks → prompt → output → cost → latency) · ADRs · model card · refusal log · lineage records (doc-version → chunk → answer) · monthly drift + precision reports
RetentionAll traces for the internal policy period of 7 years, stored in-region. Tamper-evidence is a mechanism, not an adjective: traces land in WORM object storage with a per-day hash chain whose root is written to a separate append-only ledger, so an altered record is detectable without trusting the store. Lineage records retained as long as any answer they fed remains queryable.
Data residencyInference + embeddings + logs + traces all in-region. Provable per component via cloud region pinning + storage location attestations.
Human oversightConfidence gate refusals route to a reviewer queue. Note the arithmetic before promising it: at 50k queries a day, a 5 percent refusal rate is 2,500 escalations a day, which is a staffed function and not a rounding error. Either the queue is sized and funded, or the gate threshold is tuned against a measured refusal rate. Sampled answers are reviewed weekly against a golden dataset, at a sample size the reviewing team can actually sustain.
Standards to align againstISO/IEC 42001 for the AI management system, which is the certification most likely to be named in a questionnaire, being the only AI-specific management-system standard, plus ISO/IEC 23894 for AI risk and NIST AI RMF as the mapping framework. Aligning early is cheaper than retrofitting evidence.
Adversarial surfaceA system that accepts free text and retrieves has one. Prompt injection through ingested documents is the material threat here, since a circular is untrusted input the moment anyone can add to the source share. Mitigations: retrieved content is never treated as instruction, the system prompt is fixed and versioned, and the OWASP LLM Top 10 is the checklist the red-team exercise runs against before each major release.
Release safetyNew retriever, new embedding model or new base model goes out in shadow mode first, scored against the golden dataset, then canary to a single region before general rollout. Champion and challenger run side by side until the challenger wins on groundedness, not on vibes.
Right to deletion / correctionDoc-version model + lineage store let the bank invalidate / re-embed when a policy is superseded or corrected, without rebuilding the whole index.

When NOT to Use This Pattern

  • Low-stakes, non-regulated internal tool, audit + residency overhead is not worth it. Naive RAG is fine for a marketing-copy assistant.
  • Tiny corpus that fits in context (<~100k tokens), skip retrieval entirely. Long-context prompting is simpler and often more accurate.
  • No ops capacity, if there is no platform team to run + monitor the stack, a managed bundle’s tradeoffs may genuinely win. Do not build what you cannot operate.
  • Read-only public data with no PII or regulatory weight, most of the compliance scaffolding (PII redaction, lineage, 7y retention) is dead weight. Use the boring stack.
  • Deterministic, structured-query problem, if the question is “what is policy X clause Y?”, a search index over structured metadata beats RAG. Do not reach for LLMs when SQL or a keyword index works.

Repository

A clean-room reference skeleton accompanies this pattern. Not anyone’s production code: a minimal build that demonstrates the shape end to end. The link lands here when the repository is public.

What the reference skeleton contains:

  • Reference stack. Orchestrator, PII redaction step, retriever, confidence gate, trace wiring, plus synthetic circular-style documents so it runs without any real policy text.
  • Diagram sources. The two flowcharts above, as editable source.
  • Cost model. Token cost per query, retrieval cost, loaded operations cost, and the break-even against a managed bundle.
  • ADR templates. One decision record per major choice: residency, model, gate threshold, retention.

The repository link lands here once it is public.


Forward Variants (this pattern extends to)

Same architecture; the data layer + one extra constraint changes:

VariantNew data layerNew constraint addedFuture CS
RM copilotCustomer profile + product holdings + KYCCross-customer isolation (one RM cannot retrieve another’s clients) + KYC residencyCS-01b (planned)
Credit-underwriting research assistantFinancial statements + news + internal credit memosNumber-level citation: every figure in the draft credit note carries its source locationCS-01c (planned)
EU insurer policy assistantInsurance product + claims policyGDPR + EU AI Act limited-risk + Solvency II artifactsCS-04 / CS-05 cover the EU layer

Further Reading

  • Anthropic, Contextual Retrieval, source of the retrieval-failure and index-cost figures quoted above
  • Microsoft Research, LazyGraphRAG, why the GraphRAG cost objection has expired, and why the audit objection has not
  • LangChain, retrieval orchestration docs
  • Langfuse, tracing + observability for LLM apps; audit-grade log format
  • RBI Master Direction on Outsourcing of IT Services (2023), the binding India constraint
  • DPDP Act 2023 and the 2025 Rules, processing, consent, and the transfer restrictions that attach to Significant Data Fiduciaries. Worth reading precisely, because it is routinely cited as a blanket localisation mandate, which it is not
  • EU AI Act, risk classification overview (covered in depth in CS-05)
  • NIST AI Risk Management Framework, useful even outside the US (covered in CS-10)
  • → Next in series: CS-02, LangGraph for Compliance-Aware Retrieval (multi-step retrieval where one shot is not enough)

About the Author

I architect production AI systems for regulated enterprises, with delivery experience in banking and logistics. Focus on MLOps maturity, data platforms, and compliance-aware system design.

→ If you are standing one of these up: book a 30-minute architecture review. No pitch. Bring the constraint that is blocking you and we will work out whether this pattern fits or whether something else does.

→ Also: how I work · get in touch

→ More case studies: the blog · LinkedIn · GitHub


Newsletter

New essays, straight to your inbox.

Occasional, in-depth writing on distributed systems, AI agent architecture, and engineering leadership. No spam, unsubscribe anytime.