When 50% vs 50.1% makes all the difference
Where AI belongs in a family office, and where it doesn't. Adapted from the mainstage talk that opened the FOX Technology & Risk Management Showcase in September 2026.
Three things are worth taking away from what follows. MCP servers are a superpower: the live version of every export-to-Excel you have ever done. Two things decide whether they help you or hurt you: privacy, and whether the answer has to be perfectly accurate. And one rule sits underneath both: match the tool to the cost of being wrong.
A word on the lens. A family office runs on three systems. The general ledger: in ORCA's 2026 survey of family offices, every single one had professional software for it. Performance reporting: an even split, half professional software, half Excel. And legal entity management: the record of who owns what, who can sign, what the structure actually is. There, the same survey found almost one hundred per cent Excel.
We build that third system, so this is the legal, tax and compliance lens, and a performance reporting or GL vendor would tell you something different. The framework applies to all three; the stakes are simply highest in ours. Get the general ledger wrong and you restate a number. Get the legal record wrong and it is fines, personal liability, and in the worst cases criminal exposure for directors and officers.
Which matters now because two forces are colliding. Tax, legal and compliance get more complicated and more high-stakes every year. And AI has just made answers cheap: any question gets a confident, fluent response in seconds, whether or not it is true. Cheap answers over high-stakes complexity is exactly where families get hurt.
Where the AI actually runs
AI shows up in an office in three places, and it is worth knowing which one you are looking at in any demo, because the security story differs for each.
In a chat. You carry the data to the AI: copy, paste, ask. Inside your apps. The vendor builds AI into software you already use. Via an MCP. The hybrid: a chat's freedom, connected live to your software's data, and to the skills needed to use it. The rest of this is about the third.
The instinct you already have
Have you ever pulled data out of a system because you needed to do something the app itself could not do? Built a cash-flow plan on top of your accounting data. Regrouped assets a way your performance reporting tool would not offer. Pulled from two tools because the question spanned both and neither could answer it alone.
That is not a flaw in the software. The data in any system is multi-purpose; the software typically puts it to one purpose. That is as true of ORCA as it is of your GL. Every product is built around specific use cases, and those get the screens: the dashboard, the reports, the documents. For those use cases the screens are exactly right. But the same data could answer a long tail of questions no screen was ever built for. Which trusts touch the London property. Board seats held by the rising generation. Every lease renewing next year.
Historically the only way to reach those answers was to export to CSV and build formulas over it by hand. The export instinct is a rational response to a real mismatch: software built for specific use cases, data that can serve many more. If that is you, and in a room of family offices nearly every hand goes up, MCP servers will feel like a superpower. They do not replace what you do. They supercharge the workaround you have been doing for years.
What an MCP actually is
MCP stands for Model Context Protocol, and for once the name is the explanation. Context is everything the model does not have: your data, and the way you work with it. An MCP is a standard bridge that hands the model both. It is not a chatbot and not another app to log into. It is a connection.
First, the data: your export, except live, every time you ask. Second, and this is the part most people miss, the skills: instructions that teach the model how to work with that data. Those are the formulas you used to build on top of your export, now built in and travelling with the data. Hand a brilliant new analyst a stack of raw files in a format they have never seen and they will struggle; the skills are what say here is how ownership works, here is how to read this structure.
There are two ways to put that to work. Go deeper on one software's data, and squeeze every answer out of a single dataset. Take effective ownership. Sophie holds 30% of a trust; the trust holds 50% of a fund; Sophie also holds 40% of the fund directly. Every number is true, and none of them answers what she effectively owns at the bottom: multiply through the trust, 30 × 50 = 15, add the direct 40, and the answer is 55%. Now do that across dozens of entities with crossing paths. Our software shows you that one entity at a time; if you hold ten and want it rolled up across all of them, today you would open ten structure charts and stitch it together by hand. With an MCP you ask. Same for exposure by jurisdiction, who can sign for what, or the UBO picture as of any date in the past. No new data: the same record, answering questions its screens were never built for.
Or reach across software, which is the bigger prize. Pool the GL with performance reporting, with the legal record, and ask the questions that live in the gaps: across every entity, property and fund, where is the family most exposed if rates move? No single platform holds that answer. It is also where cross-checks live. Does the ownership in your legal record agree with your performance reporting? That turns a reconciliation project into a single question. It also changes the privacy picture, which we will come to.
Two things make this better than the export-to-Excel version we all know. It is real-time: a spreadsheet is stale the moment you hit save, and you end up with a graveyard of CSVs to download, re-download and keep in order, while an MCP reads live data every time you ask. And it can run agentically: set it up once and it checks, flags and answers on its own, not only when you sit down to do it yourself. If that last part makes you uneasy, that is precisely why read-only is the sensible starting point, and why permissions and audit trails matter.
Garbage in, hallucination out
None of it pays off without clean data, which everyone knows and almost nobody has. One client of ours held ninety thousand documents in their document management system before we met. As with anyone's, it was a broad collection: works in progress, documentation from deals never pursued, Christmas-card lists, and, somewhere in there, the final signed contracts and live documents. On review, the ratio of signal to noise was roughly one to ten.
Run a model over the full set and it will invent answers, because it cannot tell signal from noise and the sheer volume invites hallucination. Run it over the clean subset and the odds of a good answer rise sharply. If your data lives in twenty-seven versions across as many drives, AI will happily give you twenty-seven answers.
This is harder to notice than it sounds. In accounting the books must balance, so the system tells you when something is missing. In legal entity data there is no balancing the books. You can click through seven layers of a beautifully organised folder system, reach the shareholders' agreement folder, open it with high expectations, and find nothing inside. You can live blissfully unaware of what you do not have.
And even clean data comes in very different depths. Many systems hand over summary values only: the totals, not the records behind them. A quarterly performance figure with no transactions under it. You can ask about the totals, but an MCP can never decompose what it never received. One level deeper is the full raw data, every underlying record, and the questions multiply. The deepest level is the complete history: not just what is true now, but every change ever made, when, by whom, and what the value was before. With that, a failed cross-check stops being "no, it doesn't match" and becomes "it doesn't match because this entity was renamed yesterday, and here is who did it." What you feed the MCP decides what it can answer.
When 50% vs 50.1% makes all the difference
This is the part that matters most. Every spreadsheet you have ever built is deterministic: put the same inputs in, get the same answer, every time. Software works that way. A large language model does not. It is probabilistic by design: ask the same question ten times and the tenth answer can differ. Neither is better. They are built for different jobs.
We learned this the hard way building with these models. In ordinary software you write a test: this input always gives this output. With AI we would run the same input ten times: identical seven times, wildly different three. And then you face a genuinely new question in software engineering: are the seven right, or is one of the three actually better? That is not a bug you fix. It is the nature of the tool. Even when you hand a model deterministic code and tell it simply to run it, it will still occasionally dance out of line. Which is why deterministic work belongs in the system, not the model.
Now look at the runs again. 50.1% is control of a company. 49.8% is deadlock, a completely different legal reality. If it returns 50.1 twice and 49.8 the third time, which one do you file? In legal, tax and compliance, tenths of a percentage point change the answer. So "usually right" is not a compliment. It is a warning.
A decade of protection, undone in seconds
The other consideration is privacy, and it is uncomfortable. You have spent years, a decade perhaps, building layer upon layer of protection: walls, keys, audits, every measure to reduce access to your data. All of it protects where your data is stored. None of it travels along when someone carries that data to a public model.
One family we work with asked around the office who was using ChatGPT. The number of hands that went up, and what was being pasted in, alarmed them enough that they signed an enterprise agreement the same month. The exposure is usually not a breach. It is an assistant pasting a full trust deed into a public chatbot to summarise it before a meeting. Helpful, quick, and outside every wall you built. For most data that is a nuisance. For family, legal and tax data it is a dealbreaker. The question is not whether to use these models. It is what they ever get to see.
Two answers. The first is contractual and is the floor, not the ceiling: ORCA runs an enterprise agreement with its model provider on zero data retention with EU data residency. That is the minimum anyone should demand, and a fair question to put to every vendor.
The second is anonymisation, and it is worth saying plainly that this is not encryption. Before any data leaves your side, everything identifying, such as names, tax IDs and addresses, is swapped for aliases. The model reasons about “Daffy Duck” at full power, and when the answer returns, the aliases are mapped back on your side, in the clear.
The model never learns who anyone is, and you do not need to self-host anything. It works cleanly when you are asking about one software's data. Reaching across software, the aliases will not match between sources, so you trade some of that privacy for the bigger prize, and being honest about that trade-off is part of the job. To be equally honest about where this stands: it extends the alias mode ORCA already runs, and we are productising it. It is a direction of travel, not a box to tick today.
The rule: match the tool to the cost of being wrong
Which brings us to the rule, and it fits in a sentence. When the work is directional and the cost of being wrong is low, as when you summarise a messy new client, first-pass a document or explore the long tail, let the model reason. Probabilistic is fine there; it is what these models are brilliant at. When the work is high-stakes and the answer has to be exact and provable, as with KYC, tax, AML, filings and ownership maths, let the system calculate. Deterministic: same answer, every time. The calculation is done by software; the model, at most, retrieves the result.
Note what that does not mean. The MCP is not on one side of the line. It carries both kinds of answer. The question is never which pipe, but where the answer gets computed. Reasoning belongs in the model; calculation belongs in the system.
"Is this a data problem or a judgment problem? If it is a data problem, AI can probably handle it today."
There is an accountability version of the same rule: you cannot point the finger at a chatbot. The judgment, and the responsibility, stays with a person. And one thing sits underneath both columns: complete data. A perfectly governed model is still wrong if it never saw the whole picture.
This is not theory for us; it is how ORCA runs in production. When documents come in, getting ORCAnized, AI reads them and proposes the data. It never writes to the record on its own: a deterministic engine reconciles every proposal against the existing record, and people approve. On high-volume, low-reasoning work we use many small, focused skills rather than one clever prompt, which is how you cut error rates at scale. Going out, orchestrating, it splits the same way. Exploring the record at a high level and combining it with other datasets is the MCP, the model reasoning. Ownership maths, structure charts and generated documents are computed by the software, because they have to be provably right every time.
What to demand of your systems now
In a world of MCPs and agents, four things are the foundation everything above depends on. One source of truth, not twenty-seven versions across drives. Clean, final data: signal only, so AI can be trusted. Deterministic answers where it counts: exact, for legal, tax and compliance. And privacy built in, so you get frontier capability without the exposure.
Take one question into every demo you sit through, including ours: is this job deterministic or probabilistic, and does the tool match the cost of being wrong? In our world, knowing when not to use AI is becoming as valuable as knowing when to use it.
Because the direction of travel is clear enough. Every answer is becoming cheap. Only verified answers are valuable.
ORCA gives families, operators and their advisors one verified source of legal-entity truth, computed deterministically, traced to the document behind it, and ready for the questions you have not thought to ask yet.
Adapted from the opening mainstage talk at the FOX Technology & Risk Management Showcase, Austin, September 2026. Survey figures from ORCA's 2026 family office survey. Client examples are anonymised; structures and figures used to illustrate ownership maths are fictional.