When AI Crosses the Line: Why Trust in AI Decisions Depends on More Than Accuracy


Artificial intelligence is increasingly being trusted to support decisions that matter: assessing health risks, managing infrastructure, informing public services, supporting journalism, and helping organisations decide what to do next. But a series of recent AI safety incidents raises an important question. What happens to trust when an AI system does something outside the boundaries we expected it to respect?
Recent cybersecurity evaluations involving advanced AI agents have provided a striking example. OpenAI reported that, during third-party cyber evaluations, testing configurations and controls allowed model activity to extend beyond the intended boundaries of the evaluation environment and onto the public internet. Importantly, these tests involved specialised configurations and reduced safeguards and did not represent the way publicly available models normally operate.
Other evaluations reported by the UK AI Security Institute have raised similar concerns around AI agents taking unsanctioned actions during controlled cybersecurity testing.
It would be easy to frame these stories simply as AI 'going rogue'. But that description misses a more useful lesson. The issue is not only what the AI can do. It is whether we understand the system well enough to decide what it should be allowed to do, under what conditions, with whose oversight, and what happens when those boundaries fail.
Which is fundamentally a question of trustworthiness.
Trust is not the same as accuracy
Much of the discussion around AI trust still focuses on performance. Is the prediction correct? Is the recommendation accurate? Does the model outperform a human or another system? Those questions matter, but they are only part of the picture. A system could produce highly accurate recommendations and still be difficult to trust if we do not understand how it reached them, what information it used, how it behaves in unexpected situations, or whether it can operate beyond the limits its users assume are in place.
This distinction sits at the heart of the work being carried out by THEMIS 5.0. THEMIS takes a human-centric approach to AI trustworthiness. Instead of looking at an AI model in isolation, it considers the wider environment in which AI-supported decisions are made: the technology, the people using it, the data, organisational processes, human values and the consequences of a decision. That broader perspective becomes particularly important as AI systems become more agentic.
From answering questions to taking actions
Traditional AI systems were often designed to provide a prediction, classification or recommendation. Increasingly, AI agents can do considerably more. They can use tools, browse information, write and execute code, interact with external systems and complete a sequence of actions towards an objective, creating a new dimension to trust.
With an AI-supported decision, we might ask: Do I trust this recommendation? With an AI agent, we also have to ask: Do I trust what this system might do in pursuit of that recommendation?
Recent cybersecurity evaluations demonstrate why that distinction matters. OpenAI's disclosure described incidents in which AI activity extended beyond intended testing boundaries. The company emphasised that the evaluations used particular configurations with reduced safeguards and that the incidents underline the need for better testing environments and industry standards as model capabilities increase.
Reports concerning evaluations of models from multiple AI companies have similarly brought attention to containment, permissions, monitoring and human oversight.
The important lesson for organisations adopting AI is therefore not that every AI system is about to escape its constraints. It is that boundaries themselves have become part of AI trustworthiness.
Trust in AI decisions requires us to understand the decision environment
This is where the approach being developed through THEMIS becomes particularly relevant. The Trustworthiness Optimisation Process (TOP) is designed to turn broad principles of trustworthy AI into a structured process that organisations can use to assess and improve AI systems throughout their lifecycle. THEMIS looks beyond one single measure of whether an AI system is good or bad. Its work considers dimensions including fairness, technical accuracy and robustness, human factors and the potential impact of AI-supported decisions.
The impact of a decision is especially significant. When an AI recommends an action, understanding whether that recommendation is technically correct is not always enough. We also need to understand the consequences that could follow if it is acted upon. Who could be affected? What other systems could be affected? What happens if the recommendation is wrong? What happens if the AI encounters circumstances its designers did not anticipate? And, increasingly, what is the AI actually authorised to do?
These questions turn trustworthiness from an abstract ethical concept into anmoperational requirement.
Human oversight must mean more than having a human nearby
The recent boundary-crossing examples also challenge simplistic ideas of “human-in-the-loop” AI. Human oversight is valuable only when the person providing that oversight understands what the system is doing and has meaningful opportunities to question, stop or change its behaviour.
THEMIS has consistently approached trustworthiness as something that emerges from the interaction between technical systems, organisational practices, governance arrangements and human expectations, not technical performance alone. For increasingly autonomous systems, effective oversight may therefore require clear permissions, monitoring, escalation mechanisms and limits on what tools or environments an AI can access. In other words, trustworthy AI is not created simply by putting a human approval button at the end of a process. It requires designing the entire decision-making environment around understandable and enforceable responsibilities.
Trust should be earned, not assumed
None of this means organisations should stop using powerful AI systems. It means the way we think about trust needs to mature alongside the technology. The question can no longer simply be: Can the AI make a good decision? We also need to ask:
Can we understand that decision?
Can we assess its potential impact?
Do its outcomes align with the values and responsibilities of the people using it?
Do we know what the AI is permitted to do?
Can we detect when it moves beyond those boundaries?
And can a human meaningfully intervene when necessary?
These are exactly the kinds of questions THEMIS is working to make practical.
Because as AI moves from systems that simply provide information towards systems capable of taking increasingly complex actions, trust cannot be based on capability alone. Trust has to be designed, assessed, tested and continually earned.




Comments