TL;DR
Getting a data pipeline right is one job. Data engineering for AI agents is a different job. The agent has to actually consume that pipeline’s output, stay current with it, and treat every change as something worth auditing. Get that part right, and every decision the agent makes traces back to real, governed data instead of a plausible-looking guess. This article covers the three-layer pipeline we built to turn raw banking data into agent-ready features. It covers how caches, vector search, and configuration stay in sync within seconds of a change. And it covers what testing the design under real usage taught us.
Previous: Human-in-the-Loop Is an Architecture Decision, Not a Feature | Next: Governed Synthetic Data: The Pipeline You Didn’t Think You Needed
🏦 Why Data Engineering for AI Agents Is a Different Job
A traditional data pipeline serves dashboards and overnight reports. A pipeline feeding live AI agents serves decisions. Real ones, about real customers, in real time. That changes what “correct” means.
The familiar Bronze-Silver-Gold pattern, sometimes called medallion architecture (raw data, cleaned data, business-ready data), was built for analytics. There, a stale dashboard is an inconvenience someone notices and refreshes. Put an AI agent on the receiving end instead, and a stale or ungoverned feature stops being a reporting glitch. As a result, it becomes a decision made on the wrong grounds. In banking, under frameworks like SR 11-7 and the EU AI Act‘s rules for high-risk systems, a number an agent used to make a decision needs a traceable origin, a validation history, and a governed approval behind it. I’ve come to treat “the pipeline ran successfully last night” as a non-answer to “where did this specific number come from.” For an agent-facing pipeline, that second question has to be answerable every time.
Back to top . Next: The Three-Layer Pipeline That Feeds Our Agents
🔶 The Three-Layer Pipeline That Feeds Our Agents
Three layers, each with a single job, is what keeps a pipeline this consequential auditable end to end.
Raw, Untouched, and Append-Only
The first layer holds data exactly as it arrived. It never corrects anything in place. Instead, it flags and quarantines a bad record, never silently rewriting it. There’s always an unedited copy of what the source system actually sent.
Validated, Masked, and Typed
- We mask personal data before anything downstream sees it. Names and contact details never reach an agent or a later layer in readable form. That handles data-erasure obligations at the source, rather than as an afterthought.
- We band sensitive figures rather than keep them exact. An agent reasoning about a customer’s finances works with ranges, never a precise income figure.
- The layer validates every record against expected shape and type. Anything that fails stays here, with a record of why, instead of quietly flowing forward.
Agent-Ready Features
The final layer assembles one governed view per customer: score, risk tier, eligibility, and the demographic fields needed for fairness auditing. It rebuilds daily. An automated validation gate checks every rebuild before the result reaches anything live. Completion of that daily rebuild also triggers a fast cache warm-up. As a result, agents get sub-millisecond reads from it, without ever querying the underlying store directly.
Back to top . Next: Keeping Every Downstream System in Sync
🔄 Keeping Every Downstream System in Sync
A web app handles a config change with a restart. In a regulated agentic system, a restart is itself a compliance event. So the goal became zero-downtime propagation: a change lands once, and every system that depends on it knows within seconds, without anyone restarting anything.
Agentic systems carry a lot of dynamic configuration: which tools an agent may call, what autonomy level it’s operating under, which policy version governs a decision. All of it needs to reach every consumer reliably. We settled on two complementary mechanisms rather than one, because they solve different problems. Debezium streams every change durably to any number of consumers, while Postgres’s own built-in LISTEN/NOTIFY reaches a co-located process in milliseconds with no broker in between. In fact, it’s event-driven propagation we first explored in an earlier piece, Real-Time Intelligence: Kafka Streaming and Event-Driven Agent Triggers, applied here to a multi-agent, regulated setting.
| What we needed | Durable stream | Instant notification |
|---|---|---|
| Speed | Seconds to minutes | Milliseconds |
| Survives an outage | Yes: a consumer that was offline can catch up | No: a missed moment is gone |
| Reaches other services | Any number of them | Only the process listening directly |
| Best for | Audit trails, compliance dashboards, cross-service consumers | Warming an agent’s own in-memory cache |
Running both side by side closes the gap either one leaves alone. The instant path invalidates a co-located agent’s cache in milliseconds. The durable path feeds audit and monitoring tooling seconds later, with nothing lost if a consumer was briefly offline. This pairs with a governance pattern from Article 6, Compliance Infrastructure: Audit Trails, Policy-as-Code, and the Append-Only Principle. We write every change to the audit trail and the durable stream in the same transaction, so the two can never drift apart.
Back to top . Next: Making Sure the Vector Store Stays Current and Trustworthy
🛡️ Making Sure the Vector Store Stays Current and Trustworthy
Every vector-search tutorial covers embedding documents, storing vectors, and querying similarity. Then it stops. Production adds the requirement that actually decides whether the search is trustworthy: the vectors have to stay current as the source data changes, and every match has to be defensible if someone asks where it came from.
Refreshing Each Collection on Its Own Schedule
We treat each Qdrant vector collection as a materialised view with its own refresh cadence, rather than a one-time load. Our agents draw on a small set of purpose-built collections: reference material the compliance agent checks product structures against, product descriptions the credit agent draws on for offer narratives, and per-customer summaries built from already-masked data. Each refreshes on the schedule its source actually changes on. Specifically, reference material updates weekly, product descriptions monthly, and customer summaries daily, timed right after the features layer finishes its own daily validation.
Two Layers of Integrity Checking
Because a retrieval system is only as trustworthy as what’s sitting in it, we also run two layers of integrity checking. One watches the overall shape of each collection against its own established baseline and flags a statistically significant shift for review. That’s the kind of drift a poisoning attempt against retrieval systems produces, a technique MITRE ATLAS catalogs as a way to make a prohibited answer look compliant without a line of code changing. The full threat-modeling picture is coming in Article 12, Threat Modeling AI Systems: STRIDE, PASTA, and MITRE ATLAS, later in this series.
The other re-embeds a small, fixed set of known reference texts daily and checks the result hasn’t drifted. That catches a different failure. In practice, an embedding model can start producing subtly wrong output across the board, a problem the first check alone won’t see. Both checks fail open. They flag for review rather than halting an agent mid-decision.
Back to top . Next: Where the Design Got Tested
🔍 Where the Design Got Tested — and What We Changed
A pipeline that runs cleanly and an agent that’s genuinely reading from it are two separate claims. Three moments during build-out taught us to stop assuming the second follows automatically from the first.
Tracing a Decision Back to Its Source
The first came from tracing a single decision all the way back to its source, rather than trusting that pipeline tests passing and agent tests passing meant the seam between them was fine. It wasn’t. An early placeholder value had quietly outlived the real data it was supposed to be temporary for; we’d put it in place before the features layer even existed. Nothing about it looked broken from either side. The values landed in a plausible range, and every test kept passing. I’d have missed it too, if we hadn’t gone looking specifically for whether the agent was actually reading real numbers. In other words, we’d been trusting that it was, not verifying it.
The fix wasn’t a bug patch. It was a new kind of check. A new automated test, which runs on every release, verifies that a live decision traces back to its governed source. That’s different from just checking that the pipeline and the agent each work in isolation.
How Long “Instant” Actually Took
The second came out of an incident drill. We measured how long a safety-relevant change actually took to reach a running agent. The change was an agent’s autonomy level being tightened. The answer came close enough to a full minute that a request could complete under the old, looser setting before the new one even arrived. In fact, I think a minute sounds fine right up until you say it next to the word “safety-relevant.” We treated the gap as a compliance requirement rather than a performance detail to shave down later, and it’s the direct reason two propagation paths run side by side today instead of one.
The Collection Nobody Owned
The third was the vector collection we seeded once during early build-out and never went back to. A handful of illustrative examples were still doing real work months later. Nothing had assigned the collection an owner, or a refresh schedule, the way the rest of the pipeline had. It’s a useful reminder that a vector store needs the same operational discipline as any other production data asset.
Data Infrastructure — Technology Stack
Delta Lake + MinIO
Append-only, three-layer storage.
Apache Spark + dbt
Ingestion, masking, and the final feature build.
DuckDB
Lightweight reads straight from object storage.
Durable change streaming
Cross-service, replayable delivery.
In-process notification
Sub-second, pairs with the durable stream.
Vector search + local embeddings
Refreshed on a schedule, checked for integrity.
Great Expectations
Gates every rebuild before it goes live.

Back to top . Next: Key Lessons Learnt
🎯 Key Lessons Learnt
A pipeline being finished and a pipeline being genuinely connected to the agents that depend on it are two different claims, and the second one deserves its own test, never just an inference from the first. That’s the thread running through everything above, and a few more lessons from the wider build sit alongside it.
On Trusting What’s Connected
| Lesson | Recommendation |
|---|---|
| A system can pass every test on both sides and still not be doing what it’s supposed to. That happens when nobody checks that one part is actually reading from the other. | Before calling a pipeline “connected” to an agent, trace one real decision all the way back to where its numbers came from. Don’t just trust that each side works on its own. |
| A live agent making decisions needs fresher, better-checked data than a dashboard does. Someone notices and refreshes a dashboard’s stale number; an agent acts on its own stale number before anyone looks. | Treat every number an agent uses the way a bank treats a number in a report. That means a clear origin, a validation step behind it, and someone accountable for both. |
| Though a minute sounds fast, a safety-related setting that takes close to a minute to reach a live system is too slow. | Measure how long an important change actually takes to reach everything depending on it. Treat that number as a real requirement, rather than a performance detail to fix later. |
On Keeping a Knowledge Base Honest
| Lesson | Recommendation |
|---|---|
| Search results built from old information keep quietly feeding wrong answers to an agent, month after month, without ever looking broken. | Any knowledge base an agent searches needs someone responsible for keeping it current, same as any other data that matters. In other words, it can’t be a “set it up once” job. |
| Manually spotting a handful of odd, out-of-place entries is not something a person can keep doing reliably at scale. | Wherever an agent’s knowledge base could be tampered with or quietly go stale, add an automatic check for anything unusual, and route it to a person for review rather than trusting it to be noticed by eye. |
| It’s easy to watch whether a vector search returns fast results. It’s just as easy to miss whether anyone can see who’s reading from or writing to it. | Treat query and write logging for a vector store as a first-class requirement from day one, the same bar every other governed data store in the pipeline already has to meet. |
On the Infrastructure Underneath
| Lesson | Recommendation |
|---|---|
| A dead-letter-queue setting can be configured, look correct, and still quietly do nothing, because the mechanism behind it may only apply to one direction of a connector rather than both. | Never assume a documented feature covers your exact setup; verify it against the tool’s own behaviour for your specific connector type, and add an explicit status check for the direction it doesn’t cover. |
| A pipeline spanning multiple tools tends to grow more than one copy of the same table name or path. In particular, application code, SQL, and config files can’t all import the same source of truth. | Add an automated check that derives the canonical value from wherever it’s actually defined. Have it flag any copy that’s drifted, rather than trusting every copy to be updated by hand. |
| An append-only raw layer is only as safe as the storage controls sitting behind it. A scheduled backup job on its own isn’t independent protection against accidental deletion. | Turn on versioning and lifecycle protection at the storage layer itself for anything append-only and regulator-facing. Don’t rely on pipeline discipline alone. |
| An audit trail meant to retain data indefinitely eventually becomes a real capacity and query-performance question, not just a policy one. | Size that kind of fix correctly: partitioning and archival for a fast-growing, long-retention table is a deliberate design project, not a quick patch. Plan it as one. |
Thank You, Reader
The most useful moments in building this pipeline weren’t the clean runs. They were the times we traced one number all the way back to where it actually came from. That habit is worth more than any dashboard that says everything’s green.
Connect With Me
- LinkedIn: Connect on LinkedIn
- GitHub: github.com/neerajg5
- Blog: learnwithneeraj.com
Enjoyed this article?
Get notified when the next one is published.
We send one email per new article — no spam, unsubscribe any time.
⚠️ Disclaimer: The information provided on LearnWithNeeraj.com regarding Astrology, Numerology, and other topics is for educational and guidance purposes only.
Not Professional Advice: This content should not be used as a substitute for professional medical, legal, or financial advice. Always consult a certified professional for specific concerns.
Guest Authors: This site features articles by various contributors. The views and interpretations expressed are those of the individual authors and do not necessarily reflect the views of the website administrator.
Your destiny is in your hands. Use this information as a map, not a mandate.