Questions senior leaders ask us
We’ve gathered together answers to the questions senior leaders like you ask us – in boardrooms, in discovery workshops, and in the corridor afterwards.
Strategy – and getting started
Q. Where do I start with GenAI in a regulated business?
A. Not with the technology, and not with a workshop that produces forty use cases.
Both roads lead to the same place: AI bolted onto processes designed for a pre-AI world, which gets you a faster horse - the same thing you do today, slightly quicker, with the old assumptions locked in.
Start from the other end. Pick one material area of the business and answer five questions in order:
What are the functional outcomes this area must produce - not the process, the result?
What data is needed to get there?
Where must a human stay in the loop, and for what named reason?
What must be true for this to be safe, compliant and trustworthy?
And what do you know that a competitor with public data and off-the-shelf AI can't replicate?
Work through those five and you have a strategic design you can sequence and act on - and a defensible answer when your board or your regulator asks how you decided. We've published the full method as a working guide for leadership teams.
Check out our comprehensive GenAI Strategy Playbook for more on this topic.
Q. Should we fix the process before we automate it?
A. Neither, in the order that question assumes. Fix first and you've spent a year optimising a process designed for a pre-AI world - improvements AI would have made irrelevant. Automate first and you've baked today's workarounds and undocumented decisions into your future state as permanent inefficiencies. Both roads preserve the architecture, and the architecture is usually the problem.
The way out is to stop starting from the process at all. Define the outcome the process exists to produce, then design the process that outcome requires - with AI carrying what it should, and humans placed where their judgement changes the result. Done this way, the change is bigger than either fixing or automating: we've seen up to 90% of the steps in some processes eliminated, because most of them existed to serve the old method, not the outcome.
That's a harder conversation than "where can we bolt-on AI?" It's also the difference between transformation and ‘faster horses’.
We can work with you to deliver a transformational GenAI Strategy Programme.
Q. Should we pilot AI in our messiest process or our most consistent one?
A. It depends what the pilot has to teach you - so decide that first. The consistent process gives you a faster, cleaner proof, but risks a result that doesn't generalise and a business that shrugs. The messy one - where ten people give ten definitions of a good outcome - is a harder test and a bigger prize, and a better guarantee that what you learn carries elsewhere. If you can solve it there, the vanilla cases follow.
In practice the deciding factors are usually unglamorous: how easily you can get at the data, how much time you'll get with the people who do the work, and whether the capability you build - the way knowledge is captured, described and surfaced - is a foundation the rest of the business can inherit.
What we'd steer you away from is the third option many organisations drift into: the safe, peripheral area that touches nothing mission-critical. It feels prudent and it settles nothing - the impact is negligible, and the sceptics gain ground.
You could benefit from our expertise running experiments, or in baselining and prioritisation for AI work.
Q. How do we avoid ending up with ten different AI agents for ten different teams?
A. You avoid it the same way you avoid ten different core banking systems: by deciding what's foundation and what's application. The proliferation happens when teams solve locally - each builds a tool on its own definitions, its own data access, its own idea of what good looks like. Every one works in isolation; together they're a maintenance burden and a governance problem.
The foundation is the knowledge layer: how the organisation's data connects, how it's described, how institutional knowledge is captured and made available. Get that right once and each team-level application inherits it. This is why sequencing matters more than most roadmaps admit - the first thing you build should create capability the next five things stand on, which is rarely the most eye-catching use case on the list.
A practical test when prioritising: does this initiative create shared capability, or just a local win? A roadmap full of local wins is how you get to ten agents.
Check out our GenAI Strategy Programmes.
Q. Why are our AI pilots stalling?
A. Almost never for technical reasons. Across the financial services organisations we work with, three causes come up again and again. There's no owner - nobody with the authority to change policy or workflow inside a quarter, so the pilot succeeds and nothing moves. Governance can't say yes - explainability and auditability weren't designed in, so risk and compliance are being asked to approve something they can't see inside. Or adoption was assumed - and people do not change how they work because a system tells them to.
The order matters, because these get misdiagnosed into each other constantly. An ownership problem gets treated with more technology. An adoption problem gets treated with another framework. Meanwhile scepticism hardens - and when early changes fail to show value, that scepticism becomes the settled view, and the next attempt starts from a worse position than the last one did.
So before scaling anything, diagnose. Work out which constraint is in play, fix that one, and prove the fix with a tightly scoped test before committing the big budget.
Consider baselining and prioritisation followed by running rapid experiments.
Q. How much of our people's time will an AI programme actually take?
A. More than a vendor will tell you; less than you might fear if the engagement is designed around reality. The reality in question: your best people are on live client work, you don't carry spare capacity, and a request for a full-time secondee is a request you can't grant.
Ring-fenced focus time works better than full allocation - structured so it's predictable, with specific outputs to validate and sessions at times that suit the team's working rhythm. Where possible, one person close to the work partially seconded, because on-tap knowledge is the single biggest accelerator a programme can have. And discovery designed to extract maximum insight per hour of expert time, because that hour is the scarcest resource in the business.
We plan engagements this way from the start - which is also why our diagnostics are measured in weeks. The longer a programme runs, the more expert time it burns simply staying alive.
Judgement, governance and accountability
Q. Who's accountable when an AI system makes the decision?
A. You are (or one of your close colleagues). Consumer Duty and SM&CR make that personal, and no vendor contract changes it. Handing a decision to a system doesn't remove responsibility; it concentrates it - in the design choices made when the system and its workflows were built. Those are also the hardest decisions to roll back, which is why they deserve the most scrutiny at exactly the moment programmes are keenest to move fast.
The regulator's question has changed shape accordingly. They no longer want to read your policies; they want to see how you make decisions. How did you decide this use case was safe to put in front of a customer? What trade-offs did you weigh? Could you defend the decision afterwards - with evidence of the reasoning, not a framework document?
The practical response is to decide your limits in advance: which trade-offs you'll make, which calls are a hard no, who decides, and how the reasoning is recorded. Firms that do this move faster, not slower, because each new demand is weighed against a position already agreed instead of being re-argued from scratch.
Read about our Responsible Change practice, and download our report on Designing Limits that Hold.
Q. Is a human in the loop enough - or are they just a ‘human on the hook’?
A. Usually the latter, as commonly implemented. By the time a decision reaches the human, it has already been shaped by choices made when the system was designed: what information is surfaced, how options are framed, what the default is, how much time the reviewer has. A person approving outputs under those conditions isn't exercising judgement. They're absorbing liability.
A human in the loop is a real safeguard only when the design gives them the means to intervene: authority to say no, the information to see how the output was reached, the time to engage with it, and an operating model that treats their challenge as the system working rather than the system failing. Those are design decisions, and left unmade, the drift is always towards rubber-stamping.
So the question to ask of any AI-enabled process isn't "is there a human in the loop?" It's "what would it take for that human to stop the process - and when did they last do it?"
Again, check out our Responsible Change practice, and download our report on Designing Limits that Hold.
Q. Where should we keep friction in an AI-enabled process?
A. Good AI design adds friction as well as removing it. The instinct in most programmes is to strip out every step that slows things down, but some friction is where the thinking happens - and a human in the loop is only a safeguard if the human engages with the decision.
Our test has four factors. Keep humans - and the friction that makes their involvement real - where regulation requires it, where liability concentrates, where the cost of error is high, and where the relationship itself is the value. Remove friction everywhere else. Most processes have this the wrong way round: friction piled onto low-stakes steps through legacy sign-offs, and a dangerously smooth path through the judgements that matter.
One warning sign to watch for: if your safeguard is a person approving outputs they had no hand in shaping, at a pace that makes real scrutiny impossible, you don't have a human in the loop. You have a human on the hook.
Q. How do we evidence to a regulator what the fact base was when a decision was made?
A. AI makes information continuous. Facts assemble and update in real time, which is useful right up to the decision point, where it becomes a liability. At some point the parts have to stop moving. You need to be able to say: this was the fact base at the moment this decision was made, and here is the reasoning applied to it.
Treat that as a design requirement, not a compliance afterthought. A locked, timestamped fact base at each decision point; rationale recorded alongside it; and a clear line between what the system assembled and what a human judged. Systems designed this way are auditable by construction. Systems where explainability is retrofitted rarely satisfy anyone - including your own risk function.
This is also the strongest argument against black-box tooling in regulated work. If you can't trace how an answer was reached, you can't defend the decision built on it.
Read about how our internal AI tooling avoids the ‘black box’ problem.
Knowledge, quality and evidence
Q. How do we capture the knowledge that only exists in our experts' heads?
A. Start by accepting that tacit knowledge isn't one thing. What a viable deal looks like, how a relationship is really going, the early-warning signals in a sector, the sniff test a veteran applies before committee - these live in different places in your process, surface in different moments, and need different capture approaches. Treating "capture our tacit knowledge" as a single workstream is how programmes stall.
The practical method is to find where each kind of knowledge naturally appears - a meeting, a document, a decision, a correction - and capture it there, in the flow of work, rather than through interviews about work. Then structure it so it's usable somewhere else: machine-readable, attributable, maintained. Knowledge captured once and left to rot is worse than none, because people learn to distrust it.
There's a commercial reason to do this well beyond efficiency. Your institutional knowledge - the judgement, policy and history embedded in how your experts decide - is what a competitor with public data and off-the-shelf AI cannot replicate. It's the defensible part of your AI strategy. Most firms automate around it instead of building on it.
Check out layer 5 of our GenAI Strategy Playbook to explore this further.
Q. If AI drafts the document, where does the thinking happen?
A. The worry is well founded. For many artefacts in regulated work - a credit paper, an advice letter, a risk assessment - the value was never really the document. It was the discipline of producing it. Anyone who has written a hundred credit papers will tell you the act of writing surfaces things no diligence meeting caught.
So the answer isn't to automate the artefact and hope. It's to separate the thinking from the container. The document is a format; the decision quality is the value. Design the workflow so the interrogation still happens - where challenge is applied, by whom, at what point - and let AI carry the assembly and recall it is better at. Get this right and the conversation at the decision point becomes more substantive, not less: critical interrogation of the case, rather than "have you thought about this?"
Get it wrong - automate the container without redesigning where the thinking lives - and quality slips, and nobody notices until it costs you.
Consider running a GenAI Strategy Programme.
Q. Can we test AI against decisions we've already made?
A. Yes, and you should. Back-testing against historical decisions is one of the most useful validation tools available, and one of the cheapest. If a system assembling facts and surfacing risks would have flagged the deal that later went wrong, that's evidence worth having. If it wouldn't, better to learn that in a test than in production.
Two caveats. First, outcomes take time to mature: a lending decision looks fine until it doesn't, so recent decisions can't tell you much about decision quality yet. Second, hindsight bias - it's easy to build a test the system passes because you already know the answer. Design back-tests blind where you can.
The deeper point: “was this a good decision?” is hard to measure in advance, so define the leading indicators that stand in for it - consistency, breadth and depth of the evidence considered, quality of the rationale. Those you can measure from day one, on historical cases and live ones alike.
Q. How do we know an AI output is right?
A. You measure it - continuously, against a human expert standard. Assumed quality is the failure mode of most AI deployments in regulated work: the system was impressive in the demo, nobody defined what good looks like, and drift goes unnoticed until something reaches a customer.
Define what a good output is in testable terms. Grade every output against it - for accuracy, grounding in source material, and staying in scope. Validate the grader itself against your human experts, so you know its judgement tracks theirs. And re-test whenever the underlying model changes, because it will, and quality that held last quarter is not guaranteed to hold now.
We know this method works because we run it on our own platform. Our automated judge reaches near-perfect agreement with our human experts, and in every test to date it has never approved an answer they would have rejected - a bar our Head of Research has held it to for over a year. If a consultancy or vendor tells you their AI is accurate, ask them how they know. There should be a number.
Related - how we undertake traceable, evidenced AI-driven research.

