9. Synthetic Data Governance for AI: Proving Data Provenance

TL;DR

Synthetic data governance for AI comes down to one question: can you prove the data your model saw is the data your generator produced? The question applies wherever a model’s output touches a person, whether that’s a loan, a diagnosis or a claim. This article walks through three controls we built to answer it: signed batches, an anomaly scorer and hash-locked dependencies. It borrows lessons from pharma, automotive and machine learning, and says where each control stops short today.

Previous: Data Engineering for AI Agents: Pipelines, Real-Time Sync, and Vector Search  |  Next: Model Risk, Agentic Risk, and the Observability Gap


🔍 Synthetic Data Governance for AI Starts With a Chain of Custody

Most generation guides fixate on realistic distributions and statistical fidelity. A regulator won’t ask about either.

They will ask how you know the data a model trained on is the data your generator produced. That gap, between generation and consumption, is where governance lives.

What Other Industries Already Learned

Other industries learned this the hard way. Pharmaceutical regulators rely on the ALCOA data-integrity principles, which say a record must be attributable, original and accurate, because inspectors kept finding batch records that had been replaced or backdated. Carmakers reached the same conclusion for software. The Uptane framework signs update metadata with separate keys, so one stolen key can’t push a bad update to a whole fleet.

Where Our Platform Had the Same Gap

In our platform the Generator API feeds Bronze, Bronze feeds Silver, Silver feeds Gold, and the agents read Gold. For a long time nothing along that path could say whether anyone had modified a batch, slipped crafted records into it, or built it with a tampered generator. MITRE ATLAS calls this ML Supply Chain Compromise. Corrupt the source, whether data or tooling, and everything downstream inherits the damage without a single anomalous record tripping a quality check.

Why does this matter for us specifically? Our credit agent’s decisions have to survive model risk review under SR 11-7, the CBUAE’s expectations and the EU AI Act, and each of those assumes you can say what data sat behind a decision. The same holds in any field where a model’s output reaches a person. Every byte in our platform is synthetic today, which makes it tempting to relax. I think that gets it backwards. A synthetic pipeline is exactly where you can build these habits cheaply, before real customer data, real regulators and real consequences all arrive in the same quarter. The gaps below came out of the audit described in The 57-Gap Audit: What “Done” Actually Means in Production AI.

Back to top . Next: Signing Every Batch


🔐 Signing Every Batch the Generator Produces

Software has signed its artifacts for decades. Data pipelines rarely bother. HMAC-SHA256 is the plain mechanism: compute a signature over the payload with a secret key, attach it, verify on receipt. Change one byte, or sign without the key, and verification fails.

How the Signature Works

Each batch gets a fresh ID. The Generator signs that ID together with a hash of the file contents. Binding the ID matters, because it stops someone from lifting a valid signature onto a different batch, which is the cheapest attack against any scheme that signs content alone and leaves the identity of the payload unprotected. The signature and record count travel in a small companion file next to the batch. A verification task in the Bronze pipeline runs first. It checks every batch against its companion file before Spark touches it. A missing file, or a signature that doesn’t match, counts as unsigned.

The signing key is one secret shared by the Generator and the verifier, and it can’t sit in a config file that anyone with deployment access can read. OpenBao, our Vault fork, is the intended production source. Today the code reads the key from an environment variable, and I haven’t verified the OpenBao path end to end.

What Happens to a Bad Batch

In production, a failed batch moves to a separate quarantine bucket. Spark never sees it. In development, the same failure logs a warning and lets the batch through. In practice, signing also switches itself off when no key is set. That keeps local work friction-free, but it means the protection comes from configuration rather than from the code.

Rotation has a sharp edge too. When the key changes, old signatures simply fail, get treated as unsigned and land in quarantine, with no overlap window. Rotating in a real deployment needs one.

Back to top . Next: Scoring Credit Requests


🔬 Scoring Credit Requests for Adversarial Intent

A valid signature says a batch wasn’t tampered with in transit. It says nothing about whether the content was crafted to push a model across a decision boundary.

Microsoft’s Tay chatbot is the famous version. It learned from whatever users typed, coordinated users fed it abusive content, and Microsoft shut it down about sixteen hours after launch, long before any review process could have caught what it had absorbed. Nothing was tampered with in transit. The input was simply crafted.

Why the Round-Number Rule Had to Go

Our first defence was a heuristic. It flagged any monthly income above 10,000 that was divisible by 1,000. It also flagged a debt burden ratio landing suspiciously close to the 45% limit. Anyone who had read our own code could walk around it. In other words, a crafted input simply avoids round numbers. Asking whether this was the standard way to catch adversarial inputs led us to anomaly detection, and to scikit-learn’s IsolationForest. The old rule still runs in the credit agent, demoted to a cheap secondary check.

How the Scorer Works

IsolationForest rests on a simple idea. Unusual points are easier to isolate than typical ones, and a request that sits on a decision boundary by design is unusual by construction. The model trains on an anonymised feature view from the Gold layer, fourteen features covering credit score, debt burden ratio, income band and payment history with no personal identifiers, and it retrains daily inside the Gold pipeline. MLflow holds it like any other model, with its own model card.

At request time the credit agent estimates a debt burden ratio from the income, debt and amount, scores the result, and maps it to a number from 0 to 100. Above 50, the request counts as suspected. I prefer a continuous score to a yes-or-no flag. A flag invites pressure to tune the threshold until it stops firing. Our red-teaming article covers the probes we run against systems like this one.

One thing I keep coming back to: the scorer’s training data is itself synthetic. It learns what normal looks like from a generator we wrote, so a blind spot in the generator becomes a blind spot in the detector. Until then, I read a low score as weakly reassuring rather than clean.

What a Flag Does, and Doesn’t Do

In practice, a flag lowers the decision’s confidence score by ten points, which can tip a borderline case into human review. It doesn’t block the request, and it doesn’t quarantine anything. If the model isn’t loaded, the scorer returns zero and the decision carries on. I’d rather ship a detector that fails open than one that takes credit decisions down with it. I still don’t love the trade-off.

The risk of failing open showed up in a naming audit. We had registered the scorer in MLflow under one name while every other file looked for another, so its promotion path silently did nothing. A quiet detector looks exactly like a healthy one, and that is why the fix mattered.

Back to top . Next: Locking Down the Code


🔗 Locking Down the Code the Pipeline Runs On

Signatures protect the data and the scorer watches the requests. That leaves the Python packages the pipeline itself imports.

ATLAS treats this as the software side of supply chain compromise. A tampered evidently release could quietly bend drift metrics, and a patched aif360 could report flattering fairness numbers while discriminatory approvals go through, and no test would notice, because the tests run against the same tampered package. It has happened to a mainstream framework. Between 25 and 30 December 2022, anyone installing PyTorch’s nightly build also pulled a malicious package called torchtriton, which won out over the legitimate one and stole system files and SSH keys. A hash-checked lockfile would have refused it, because the imposter’s bytes differ from the audited ones. These packages sit at the centre of governance. From an attacker’s view, there is no clean line between application code and governance code.

Hash Locking, the Way Package Managers Do It

Every published package has a deterministic SHA-256 hash. Record it when you resolve dependencies, verify it on install, and you get exactly the bytes you audited. pip-compile can turn each service’s declared dependencies into a lockfile that covers every transitive package, and pip’s hash-checking mode refuses to install anything that doesn’t match. Two make targets wrap the workflow. One regenerates the lockfiles, and the other checks that they exist.

Hash locking has its own honest limit. It is trust on first use: the first lockfile records whatever the package index served that day, so it proves later installs match the first one, not that the first was clean. That still turns a silent swap into a loud build failure, which is most of what you wanted from it.

What’s Built and What Isn’t

I’ll be plain about the state of this one. The tooling exists. The lockfiles haven’t been generated and committed yet, the image builds still install from version ranges rather than hashes, and no CI job runs the presence check, so a contributor could add a dependency, skip the lockfile entirely, and nothing in the workflow would object. Around it, a nightly workflow runs static analysis on the source and vulnerability scans on the container images, and a script audits packages for known CVEs on demand.

So the platform can detect a known-bad package today. However, it can’t yet prove that the installed ones are the audited ones. Once enforcement lands, one piece of friction is worth keeping: a dependency change should need a code review. For governance tooling in a regulated environment, that friction is the point.

Back to top . Next: Tracing One Decision


🗺️ Tracing One Decision Back to Its Source

Put the controls together and a single credit decision leaves a trail. The Generator signs a batch, and Bronze verifies it before ingesting. Silver masks personal data and runs quality checks, and only passing records reach Gold. dbt builds the agent-readiness table, with lineage visible end to end in OpenMetadata. The credit agent reads it through the MCP server, the scorer examines the request, and the finished decision lands in a trace linked to its credit and Sharia explanations.

Regulators want validation, monitoring and an explainable audit trail. None of that works if you can’t show what data stood behind a model. The chain above supplies the pieces, although I haven’t yet tested a walk from a decision all the way back to its signed batch, and I think that is the next thing worth building.

Everything still runs on synthetic data. When real core banking data arrives through CDC, the verification idea should carry over. I doubt file-level signing will, because a change stream has no batch to sign. As a result, that control needs a different shape.

Back to top . Next: Key Lessons Learnt


💡 Key Lessons Learnt

Each row pairs something the build taught us with what we’d recommend to anyone doing similar work.

Lessons From Signing and Keys

LessonRecommendation
A control that switches itself off when its key is missing only protects the environments where someone remembered to set the key.Make production refuse to start without the key, and test that the refusal actually fires.
Rotating a signing key without an overlap window turns every in-flight batch into a false alarm.Accept the old and new key together for a fixed period, retire the old one afterwards, and record which key verified each batch.
Flagging a suspicious input is different from stopping it.Decide up front what each score band does, whether that is lowering confidence, forcing review or blocking, and test the path with a crafted input rather than only the flag.

Lessons From Scoring and Detection

LessonRecommendation
Anyone who reads a rule built around a known threshold can beat it.Let a statistical detector carry the main signal, keep rules as cheap secondary checks, and ask whether there is a standard approach before writing a custom one.
Fail-open detectors go quiet without anyone noticing, and a model sitting in the registry under an unexpected name is one way that happens.Define model names in one shared place and add a check that every consumer resolves the model it expects.

Lessons From Enforcement and Audit

LessonRecommendation
Lockfile tooling in a repository is not lockfile enforcement in a build.Commit the lockfiles, install with hash checking in every image build, and fail pull requests when the lockfile drifts from the declared dependencies.
Vulnerability scans catch bad versions but not swapped bytes.Pair the CVE audit with hash verification, since each answers a different question.
Integrity failures that reach only an application log are invisible to the people who review audits.Emit them to the audit trail reviewers query, and confirm the event by triggering a failure on purpose.
Batch signing suits files, and it doesn’t transfer to a change stream.When moving to streaming sources, choose an integrity mechanism per event or per segment, and prove it before the real data arrives.

Back to top


Thank You, Reader

Thanks for reading. Tracing each control back through the code showed which parts are enforced today and which are tooling waiting to be switched on, and I’d rather you saw both. The next article builds a model risk framework around them.

Connect With Me

Enjoyed this article?

Get notified when the next one is published.

🔒 We send one email per new article — no spam, unsubscribe any time.

⚠️ Disclaimer: The information provided on LearnWithNeeraj.com regarding Astrology, Numerology, and other topics is for educational and guidance purposes only.

Not Professional Advice: This content should not be used as a substitute for professional medical, legal, or financial advice. Always consult a certified professional for specific concerns.

Guest Authors: This site features articles by various contributors. The views and interpretations expressed are those of the individual authors and do not necessarily reflect the views of the website administrator.

Your destiny is in your hands. Use this information as a map, not a mandate.

Related Posts

8. Data Engineering for AI Agents: Pipelines, Real-Time Sync, and Vector Search

TL;DR Getting a data pipeline right is one job. Data engineering for AI agents is a different job. The agent has to actually consume that pipeline’s output,…

Building Real Human-in-the-Loop with LangGraph Checkpointing

7. Building Real Human-in-the-Loop with LangGraph Checkpointing

TL;DR A review queue with a button looks like human oversight. It isn’t — the automated decision has already taken effect by the time anyone opens that…

Compliance Infrastructure: Audit Trails, Policy-as-Code, and the Append-Only Principle

6. Compliance Infrastructure: Audit Trails, Policy-as-Code, and the Append-Only Principle

TL;DR We built three pieces of audit trail architecture for a regulated AI platform: a transactional outbox so an audit event can’t be silently lost between Postgres…

Red-Teaming AI — OWASP LLM Top 10 and the Probes You Actually Need

5. Red-Teaming AI: OWASP LLM Top 10 and the Probes You Actually Need

TL;DR We had heard of the OWASP LLM Top 10, and we had implemented fewer than half of it. We had PyRIT installed, and we had almost…

Securing Agentic AI Authentication, Authorization, and PII

4. Securing Agentic AI: Authentication, Authorization, and PII

TL;DR Our five-agent banking AI platform had OPA wired into exactly one agent. The MCP server had no auth middleware. All five agents shared one API key….

The 57-Gap Audit — Gap Categories and Discovery Method

3. The 57-Gap Audit — What “Done” Actually Means in Production AI

TL;DR After weeks of building a multi-agent AI platform — five agents, full pipeline, red-team harness, control UI — the system looked done. It wasn’t. A config…