TL;DR
A review queue with a button looks like human oversight. It isn’t — the automated decision has already taken effect by the time anyone opens that queue. We built Human-in-the-Loop so it can’t work that way: the workflow stops dead the moment a case needs a person, and nothing moves until a named, authenticated reviewer makes the call. It’s built to survive a restart, too: we save the case’s state the moment it pauses, so killing the process mid-wait shouldn’t lose it — the case should pick back up exactly where it stopped. If the clock runs out with no answer, the system still won’t decide on its own: the case keeps waiting for a person, even when a human’s decision beats our automatic check by a single second. Regulators are converging on the same standard worldwide: a person has to act before an outcome takes effect, not just be able to. The EU AI Act’s Article 14 states it plainly; GDPR and similar rules elsewhere carry the same requirement in substance. That’s what we built: real oversight, not a review queue bolted on after the decision is already made.
Previous: Compliance Infrastructure: Audit Trails, Policy-as-Code, and the Append-Only Principle | Next: Data Engineering for AI Agents: Pipelines, Real-Time Sync, and Vector Search
📋 Where the Requirement Comes From
Regulation in this space has converged on one specific requirement, worded differently in each jurisdiction but consistent in substance: an automated decision with real consequences for a person needs a specific person answerable for it — not a process that merely permits review after the fact.
The EU AI Act’s Art.14 states it most directly — for high-risk systems, which financial credit decisions are, people must be able to “oversee, understand and, where necessary, intervene in, or interrupt, the system.” GDPR Art.22 makes a related point: a customer has the right not to be subject to a decision based solely on automated processing when it has legal or similarly significant effects. Gulf data-protection frameworks mirror the same floor.
That’s the requirement the rest of this piece is about actually meeting — not on paper, but in how the workflow itself behaves. The case we’ll keep coming back to: a murabaha financing request moving through our credit-decisioning workflow, and what has to happen the moment it needs a person instead of a model. A large exposure, or a financing structure with no close precedent to check it against, carries real consequences for the customer if the model gets it wrong — and a customer who nudges their numbers to land just under a review threshold is exactly the kind of pattern a model shouldn’t be trusted to resolve on its own.
Back to top . Next: What HITL Actually Has to Guarantee →
⏸️ What HITL Actually Has to Guarantee
HITL has to guarantee one specific thing: the workflow cannot reach a final answer without an explicit human action. Anything short of that is not oversight, no matter how it looks.
The common shortcut is a review queue with a button. An AI reaches a decision, the decision takes effect, and a record of it lands in a queue somebody can look at later. Or not, if the queue runs long and the deadline passes anyway. Nothing in that flow stops the outcome from happening. A person could intervene. Nothing requires it.
A real gate doesn’t produce an outcome and then invite review. It stops before producing one, and only a named person’s decision lets it continue. A human being able to intervene isn’t the same as a human being required to. Our own gap register has an entry for exactly this mistake — Learning L14: “HITL that is implemented as a UI button does not satisfy the regulatory requirement for mandatory human oversight.” We built the first version that way once. It looked like oversight and wasn’t.
In our credit-decisioning workflow, a large murabaha request that trips the review threshold doesn’t get an answer of any kind — it halts right there, and stays halted until a named reviewer signs off, however long that takes.
Back to top . Next: Where This Pattern Applies →
🌐 Where This Pattern Applies — Some Real Examples
This isn’t a banking-specific problem. It shows up anywhere an automated system reaches a conclusion that has a real consequence for a person, and it’s one of the most active design questions across the industry right now:
- Healthcare. A diagnostic or treatment-recommendation model flags a likely finding — it doesn’t reach the patient until a licensed clinician confirms it.
- Hiring and employment. Automated screening or a termination recommendation increasingly requires a documented human sign-off before it takes effect, not an after-the-fact audit trail.
- Content moderation. An account suspension or takedown above a certain severity goes to a human reviewer before it’s enforced, not logged for someone to check later.
- Criminal justice. A risk-assessment score can inform a bail or sentencing decision, but the decision is supposed to remain the judge’s, not the model’s.
- Autonomous and industrial systems. When a self-driving system or an industrial control loop hits a situation it can’t resolve confidently, it hands control back to a person instead of guessing.
Credit decisioning is one more version of the same problem. Somewhere in that process, something has to decide whether this particular case is one a machine is allowed to finish alone. That decision comes from more than one place — a policy and risk-tier evaluation, plus, specific to this platform, a check for whether the request is a genuinely novel financing structure with no close precedent. Any one of those can force the case to a human. A standard request under the risk threshold clears on its own; a large exposure, a risk-tier edge case, or a structure the rules engine hasn’t seen before goes to a person every time, by construction — not because someone remembered to check. That threshold is evaluated on every single request, not reviewed occasionally, which is exactly why the mechanism enforcing it has to be structural — in banking or in any of the domains above.
Back to top . Next: What We Built →
🏗️ What We Built
Five pieces make the guarantee real: the workflow actually pauses, the pause survives a restart, a timeout suspends rather than resolves, the resume is tied to a named, accountable person, and two detection layers catch what a distracted reviewer wouldn’t.
A Genuine Pause
The workflow’s own state now tracks whether a case needs human review, why, what the operator decided, and whether it timed out — with a genuine “pending human review” outcome as one of its real final states, not an afterthought. Nothing downstream runs until a human decision changes that state. Getting there took splitting the escalation into a one-time setup step and a separate, minimal pause step, since the engine re-runs a step’s full body on every resume; anything that should happen only once now runs before the pause, not inside it. The system also responds immediately once a case is escalated, instead of holding a connection open for minutes waiting on a poll the way the earlier version did.
Surviving a Restart
Every agent now saves its state to durable storage the moment it pauses, which is what lets the pause survive a restart — the earlier version had no such durability, so a restart mid-wait silently erased a case instead of suspending it. Resuming reloads that saved state and picks up where the case left off, triggered by either an operator’s decision or the timeout sweep below.

Handling a Timeout
Two options for an unanswered escalation feel reasonable, and both are wrong: auto-approve makes the review requirement pointless, and auto-reject is still a fully automated decision wearing a caution label. So a timeout does neither. The case sits suspended until a person acts, resolved only by a background sweep that never approves or rejects on its own — and that always defers if an operator’s own decision lands even a second before it does.
Who Can Approve a Case
An approval has to come from somewhere accountable. A real operator registry now backs every decision, with properly hashed passwords and a verified identity checked before anything is accepted — no more a shared operations key anyone with API access could use. The review console finally has a working login screen, instead of carrying unused authentication for months. Operators also get a third option beyond approve or reject: hold, which extends the deadline and logs a note without resuming the case.
Watching for Manipulation
A correct pause doesn’t help if someone can shape a request around it. A red-team exercise resubmitted the same applicant with small income tweaks and walked their debt ratio under the review threshold in four tries — one near-miss means nothing alone, but a pattern from the same customer now forces every later request from them to a human. A second check flags suspiciously round income figures and ratios sitting right at the ceiling; neither auto-rejects, both just route to a person. A quieter failure lives inside the automation itself: when the deterministic result and the model’s own explanation disagree — an optimistic writeup on a case that was actually declined — that mismatch is a warning sign worth escalating on its own, a real first line of defense against the kind of overreliance failure the OWASP Top 10 for LLM Applications tracks, the same category we probed for in Red-Teaming AI: OWASP LLM Top 10 and the Probes You Actually Need.
Back to top . Next: What We Learned →
💡 What We Learned
A guarantee that looks correct in a demo isn’t the same as one that’s structurally true. Seven things this rebuild taught us, and what we’d tell anyone building the same kind of system.
| What We Learned | Recommendation |
|---|---|
| A review queue with a button isn’t oversight. If a decision can complete on its own and a person merely gets a chance to look at it afterward, that satisfies nothing — the regulation requires a person to act, not to be able to. | Make the pause a real, final state inside the workflow itself, so nothing downstream can run until a person explicitly resolves it — not a queued notification sitting beside a decision the system already made. |
| A timeout has to suspend the decision, not make one. Auto-approve and auto-reject are both fully automated outcomes wearing different labels — neither actually leaves the outcome with a person. | Let a background process act only after the deadline passes, and only to confirm the case is still awaiting a person — never to approve or reject it outright — and always yield if a human’s own decision arrives first. |
| A pause that doesn’t survive a restart isn’t a guarantee, it’s a hope. Without durable state, an interrupted process silently loses the paused case instead of preserving it — and nothing says so afterward. | Save a paused case’s full state to durable storage before control returns to the caller, so a process killed mid-wait resumes from that exact point instead of losing the case outright. |
| Resuming a paused step can silently repeat the setup around it. Depending on the workflow engine, code meant to run once around a pause can fire again on every resume, double-counting or double-logging something that only happened once. | Put anything that must happen only once — writing the case, logging the request, updating a counter — in a setup step that runs before the pause, and keep the pause step doing nothing else. Then deliberately trigger a resume during testing, not just the first pass through. |
| A decision is only as accountable as the identity behind it. A shared credential lets anyone with access approve anything, which quietly defeats “a specific person is answerable for this” even when a human technically clicked approve. | Require a verified login for every reviewer, and record which specific person approved, rejected, or held each case — so no shared credential can ever stand in for an individual. |
| A single evaluation misses someone gaming the threshold over repeated attempts. One near-miss against a limit means nothing on its own; a pattern of near-misses from the same source is a signal a one-shot check can never see. | Keep a running count of how often the same customer lands near the threshold within a set time window, not just whether this one request crosses it. Once that count passes a limit, route every later request from them to a human automatically, regardless of what that request’s own score says. |
| Automation can disagree with itself, and that disagreement is a signal worth acting on. A model’s explanation and a deterministic result can diverge — an optimistic-sounding narrative on a declined case is itself evidence something needs a second look. | Compare a model’s written explanation against its own deterministic result, not just against the original input. Treat a mismatch — an upbeat writeup on a case that was actually declined — as a reason to escalate automatically, not a wording quirk to shrug off. |
Thank You, Reader
Rebuilding this forced us to actually define what oversight guarantees as a technical matter instead of leaving it as a checkbox next to a queue. My rule of thumb now: does the workflow need a person to act before it continues, or only before it stops? Getting that answer into the graph itself, rather than trusting a process document to catch it, is what changed for us — and if you’re building something with LangGraph’s orchestration layer, our earlier piece on multi-agent orchestration with LangGraph covers the layer this HITL design sits on top of.
Connect With Me
- LinkedIn: Connect on LinkedIn
- GitHub: github.com/neerajg5
- Blog: learnwithneeraj.com
Enjoyed this article?
Get notified when the next one is published.
We send one email per new article — no spam, unsubscribe any time.
⚠️ Disclaimer: The information provided on LearnWithNeeraj.com regarding Astrology, Numerology, and other topics is for educational and guidance purposes only.
Not Professional Advice: This content should not be used as a substitute for professional medical, legal, or financial advice. Always consult a certified professional for specific concerns.
Guest Authors: This site features articles by various contributors. The views and interpretations expressed are those of the individual authors and do not necessarily reflect the views of the website administrator.
Your destiny is in your hands. Use this information as a map, not a mandate.