Multi-Model Agents, Harness as a Platform (HaaP), Microsoft Autopilot
Few weeks ago I wrote about how using a single AI model for diagnostics is fundamentally biased:
It is a well established fact that a single-person view is inherently subjective. We all have our blind spots, no matter how much of an expert somebody might be. The solution is logical: assemble a team of experts to cover as much perspective as possible and challenge each others conclusions and thought process. The only issue with that solution is the economic imperative - it’s too costly to have multiple people looking at the same problem.
The same problem has been widely documented when it comes to diagnosing diseases and performing different assessments. Indeed, it is exactly the selling proposition of many AIs that they can now finally overcome that subjectivity thanks to being trained on thousands of images and thus reaching a level of expertise few humans are capable of reaching. And yet, despite often beating human performance benchmarks in lab conditions, they struggle to do so in real life.
The reasons are many - from differences in data produced using device A vs device B, to edge cases, to hallucinations - but the problem remains the same: no single model can produce reliably objective diagnosis. The solution should then follow the same line of thought - add more AIs, that are trained on different datasets, to analyse the test, provide a second opinion and then also work out what the takeaway should actually be. The difference with a human team - it’s considerably cheaper.
Using different models for the same task is a good idea beyond diagnostics. Applying probabilistic software for admin tasks can still produce errors, which if not caught on time can become deadly for the patient. And since a single model can’t both produce the work and reliably rate it’s quality, and catch mistakes, at the minimum a second model is required to supervise.
Harness as a Platform (HaaP)
To operate multiple models and provide them with the context, rules and permissions under which they should act, you need a system: you need harness. What is that? Put simply, if models are the brain of the agent, then the harness is all the other organ systems in the body - nervous, skeletal, muscular, etc. The brain comes up with an idea, but the body’s limits dictate what is actually possible to do.
So the harness layer becomes very central in the whole task execution, as it drives model and tool calls, manages conversation state and context, applies approval policies.
As such, it functions as a shared control layer for AI across teams within an organisation, which is another way of saying that it acts as a platform on which organisation-specific agents, apps and experiences are built.
Microsoft Autopilot
Enter Microsoft, from The Verge:
After teasing its new Copilot “super app” last month, Microsoft is officially unveiling it today. The redesigned Copilot app bundles three AI capabilities into a single interface of chat, coding, and agents…The Code tab is the surprise addition to this so-called “super app,” and not one you’d typically associate with an app designed for knowledge workers. It will let anyone create an app, tracker, dashboard, or automation and share the results with colleagues as cloud-hosted internal apps…
The real difference against some other AI agents is that Microsoft has made Autopilot enterprise grade, appealing to businesses that rely on its Office and identity management suites. “Autopilot lives in your tenant with its own identity, memory, computer, and workspace, and it’s built on Microsoft IQ so it understands how your organization actually works,” says Spataro. “It shows up where people already work — Teams, Outlook, chats, channels, and documents — so you can @mention it like a colleague, with permissions, audit, and governance behind it.”
Ben Thompson from Stratechery was on to it from the start:
I wrote years ago that the best way to understand Teams was as Microsoft’s OS for SaaS; the new Copilot is the obvious new iteration of that. The point has always been to own the interface, and to force everyone else to integrate into that interface on your terms.
The article is well worth a full read. Microsoft had already showed us the v1 of this with Dragon Copilot and is now deploying the full version. Microsoft has always done well as a platform and controlling the harness fits perfectly in that strategy. The agentic interface is becoming a reality and disruption is on the horizon.