Originally published on LinkedIn, November 15, 2024.
If your work is motivated by curiosity, you often have a question that you want to pursue wherever it may lead. Sometimes, though, the question turns into a whale and you become, as it were, a puppet at the other end of the harpoon.
For me, that Moby Dick question is this: "How should we trust autonomous AI agents?"
The question unfolds as a fractal, revealing new questions. What do "agency", "artificial," "intelligence," "autonomy," and "trust" really mean? Who is the "we" implied here? And why "should" rather than the simpler "can"? The question swims at the confluence of three deep currents: how we know (episteme), how we build (tekhne), and why we should (ethos).
All that Greek would take too long to unpack in a single post. For now, let's focus on one word: Trust.
Most books and papers on Trustworthy AI ignore the nature of trust and rush to a list of attributes — reliability, transparency, robustness, and so on. At which point, decoding trust reduces to an operational issue of benchmarking AI systems based on these attributes.
I want instead to excavate the root of trust and then climb up to branching questions. Does the nature of trust depend on the subject and the object of trust? Are there fundamental differences in the way we trust machines and persons? Does understanding these differences prompt us to rethink how we should build and operate AI agents that humans can trust? How could autonomous AI agents trust each other?
But first, what is this trust of which we speak?
Root of Trust
Trust is a word we use so often that it skips over our thoughts like a smooth pebble across water. Yet its ripples reach concepts like reliability, transparency, expectation, cooperation, and goodwill on one side of the river and ignorance, incompetence, malice, and betrayal on the other.
Economists and social scientists tend to see trust as a matter of rational self-interest: you trust someone if you think it's in their interest to help you. Philosophers see it as a matter of goodwill: you trust someone if you think they have good intentions towards you. Evolutionary psychologists see trust as reciprocal altruism — a stable reward strategy for everyone involved.
At a personal level, we usually trust the air that we breathe and the ground on which we stand. When we say that we trust a machine we use, like the car we drive, we refer to reliability. We trust things that work well consistently. We are disappointed when they don't.
But the word holds a richer complexity when we say that we trust a person. At an inter-personal level, when we trust another person we do more than just rely on their competence. We expect them to encapsulate our intentions and our interests in their behavior, even when we are not looking.
For example, when you hand over your car keys to a friend, you trust them not only to drive safely but also to be responsible for your property. This form of trust is an expectation that your friend will care about your interests just as you would.
When that expectation isn't met, we feel not only disappointed, as we would if our car broke down, but betrayed. The vulnerability to betrayal is what distinguishes trust from mere reliability.
Trust is manifest in language as well as action. We trust someone when they do something with competence and goodwill. We trust them when they speak with knowledge and sincerity. We trust someone who refuses to answer a question about which they have no certain knowledge but not someone who confabulates an answer and delivers it confidently. This contrast between refusal and confabulation ("hallucination") is critical when we consider trust in AI agents.
Interpersonal trust, we may infer, is an attitude based on the reasoned calculation of the rewards of cooperation against the risks of betrayal, accompanied by feelings of empathy and goodwill.
The Quality of Trust Is Not Strained
A critical issue for all of us at one point or another is whether to trust someone and why. What evidence makes trust reasonable? And does that burden of proof rest entirely with the one trusted?
We act as if the decision to trust someone is within our control and that it is based on the other party's trustworthiness. This is why the literature on the topic assumes that trustworthiness is a set of attributes associated with the object of trust.
But this is not how we trust a child alone with a marshmallow or a teenager with their first car or an employee on their first day. In many cases, trust operates like a belief. It's a value we project onto the other in good faith and then recalculate in light of evidence. Our investment of trust without evidence often inspires the other to live up to our belief in them.
Thus, rather than a property of the object alone, trust is a property of the relationship between subject and object. It recalls memories of past interactions. It involves the reputation of both parties based on the history of interactions with each other and with others. It refers to their antecedents and origins. And it requires an alignment of values between everyone involved. All this and more goes into figuring out whether to trust a person and why.
That steady look in the eyes, the firm shake of hands, and the warm smile of agreement are all well-established markers of trust between people. Without that trust, we could not collaborate with each other, cooperate in groups, maintain long-term relationships, or build complex flows of work and play.
AI agents are not like humans. They do not reason, do not have feelings, and most certainly do not have empathy for the unique human condition. Yet, because they present themselves as if they were people, fluent in natural language and conversant with publicly available knowledge, we tend to anthropomorphize them. And thus, we give them an undue credit of trust in advance of evidence.
This, then, is the central paradox of trust in AI agents. As humans, we are prone to trust AI agents the way we trust machines or humans. But we cannot trust agents the same way because they do not have machine-like reliability or human-like empathy.
Trust is an orientation by the truster towards the trusted, for a particular task in a particular context. This orientation creates an expectation and a vulnerability that the benefit of cooperation exceeds the risk of betrayal. Without this net-positive benefit, there would be no reason to cooperate. Thus, trust is a prerequisite for the evolution of cooperation among agents. Without cooperation, there would be no way to achieve complex large-scale goals that require collective effort over a long period of time. So if we hope to realize the transformative potential of AI, we must build and operate agents that humans can trust.
And so we're pulled back into the maelstrom by a fin of the question: How should we trust autonomous AI agents?
Stay tuned for the next post.



