Building the Scaling Trust portfolio
How the portfolio fits together and what comes next.
AI agents are increasingly writing software, making purchases, and operating robots and machinery. Yet they struggle to interact in multi-principal, multi-agent settings where some agents may be untrustworthy, information asymmetry exists, and parties may have opposing goals. Decades ago, technologies such as encryption and digital signatures provided the foundations for trust between parties online, enabling the digital economy to thrive. Now, new technologies are needed to build the trust infrastructure for the agentic era.
Scaling Trust is ARIA’s £50m programme to build them: the frontier infrastructure and fundamental research for AI agents to coordinate securely on our behalf, across digital and physical worlds.
Success could mean AI agents can be trusted to act reliably on behalf of people and businesses, from negotiating deals to coordinating supply chains, opening up economic and social opportunities that are currently too risky to pursue. It could help people and businesses delegate more to agents with confidence, open up new forms of trade and collaboration that are currently out of reach, such as cyber-physical markets, and distribute the benefits of artificial intelligence as widely as possible, while helping people retain choice over the infrastructure their agents depend on.
Today we announced our first Scaling Trust Creators. Below you’ll find more about the programme’s structure, a map of the portfolio, and some ideas on what we think is still missing from it. We fund new teams every quarter, so consider the map a snapshot!
Portfolio Structure
Scaling Trust has three connected tracks, moving between theory, implementation, and experimental evidence.
Track 1: the Arena. A cyber-physical environment for evaluating agents owned by different people as they pursue open-ended, long-horizon economic tasks under adversarial pressure. See our September update on the Arena.
Track 2: the cyber-physical agent stack. An open-source stack for trustworthy agents spanning the digital and physical worlds: agent harnesses, secure protocol reasoners, but also tools for agents to use such as secure hardware and tamper proof sensors.
Track 3: fundamental theory. The science beneath the stack, including formal AI security, autonomous protocol generation, protocols for physical verification (both using trusted hardware or harnessing the properties of the physical world, which we call nature cryptography).
We made a map of how the portfolio fits together, divided by the programme tracks. It shows the teams awarded grants from our first solicitation, and the gaps still open for future solicitations.
Arena partners
Environment & challenge design
Physical build & operations
Security & red teaming
Community partners
Industry partners
Digital partners
Open-source practice, use cases
Digital stack
Agents · agentic loop
Agents · new models
Agents · negotiation
Tools · TEE sandbox, auditing
Integration
Glues the stack together
Physical stack
Physical verification · bridges, sensors, actuators
Physical environments · evals, benchmarks, harness
Cyber-physical partners
Real-world data and environments
Theory of secure agent interaction
Generative security
Formal AI security
Secure agent interaction
Cyber-physical bridges
Physical verification theory
Secure hardware
Nature cryptography
Physical environments · evals, world models
Builds Arena v1 — the environment, scenarios, and mechanics agents are tested in — then supports launch, maintenance, and iteration.
Designs and operates the physical side of the Arena.
Sets the Arena's security criteria and operating model, then red-teams it: adversarial signal, challenge design, and community once it is live.
An agentic security stack that captures security and privacy risks, requirements and functionality boundaries, translates task requirements into formal specifications, selects or generates suitable cryptographic protocols (e.g. secure comms, MPC, ZK), and verifies their properties with Lean, ProVerif and similar tools. The goal is an end-to-end workflow in which agents, from orchestrators to sub-agents, can negotiate security needs and connect verified protocol descriptions to executable multi-agent interactions.
Team: Myrto Arapinis, Jianyi Cheng, Michele Ciampi, Wenda Li (University of Edinburgh)
An automated mechanism and protocol design engine. Agents specify desired properties, such as incentive compatibility, safety and liveness, and Dovetail designs an interaction protocol and proves in Lean that it satisfies those requirements. The team will also manage a cross-team repository of protocols, with the goal of speeding up automated protocol generation.
Team: Emanuele La Malfa, Angelo Huang (Institute for Decentralized AI); Mirco Giacobbe, Gabriel Santos (Zeroth Research)
ProtoSage is an agentic system for auditing cryptographic protocols, supported by ProtoBench’s corpus and evaluation tasks. The project builds on OpenSage and WireWatch to identify vulnerabilities, recover protocol specifications, and eventually produce machine-checkable models that experts and formal-verification tools can inspect.
Team: Mona Wang (RDI Foundation)
Evaluates how AI agents negotiate, procure, bid and cooperate under strategic and adversarial pressure. It combines an open game-based evaluation suite with a hardening pipeline intended to produce agents that preserve economic value, resist manipulation and avoid leaking private information.
Enables mutually untrusting AI agents, for example representing different companies, to establish trust by cryptographically proving facts about those companies and their physical-world processes to each other: that a product was produced to a particular specification, say, or under what conditions an agent would be willing to reach an agreement. CCTI is also the programme’s integration partner, mapping how Creators’ work fits together, co-developing shared building blocks with interested teams, and proposing open-source benchmarks and challenges to the Arena.
Team: grjte (Ink & Switch); Hossein Hafezi, Alireza Kavousi, Arman Kolozyan, Jessica Man (University of Cambridge); Jonathan Heiß, Ágnes Kiss, Aurel Stenzel (SINE); Daniel Hugenroth, Mario Lins (Light Squares)
Connects the programme to practical industrial needs. Drawing on its manufacturing expertise, the team will support the design and evaluation of cyber-physical trust tools that are relevant to real production environments, providing real-world data for other teams to use and keeping the programme grounded in operational reality as it explores how agents can interact securely across digital and physical systems.
Team: Kate Lucas, Ciara Gumsheimer, Ollie Rosen (Matta)
A tamper-resistant physical watermarking framework to verify the provenance and safety compliance of control policies driving embodied robots. By extending the Colored Noise Coherency (CoNoCo) construction, it lets independent parties remotely authenticate active controllers using commodity hardware, such as standard CCTV cameras or smartphones, without direct access to the robot or specialised sensing equipment.
Team: Manon Flageat, Mateusz Sypniewski, Sally Matthews (University of Cambridge)
Investigates whether the recurring structure of AI computations can support more efficient proofs than generic circuit-based approaches. It also explores self-proving models, and the boundary between computations that can remain black-box and those that must be decomposed for verification.
Privacy techniques suited to AI workloads, built with cryptography: private retrieval for RAG, semantic search, secure computation, and methods for concealing queries or embeddings. The aim is to let agents use shared memory and sensitive data without exposing commercially or personally confidential information.
Team: Nir Bitansky (NYU); Yuval Ishai (Technion); Ron Rothblum (Succinct/Technion); Sarah Meiklejohn (UCL/Google)
Commentary
Several bets on the digital stack
We do not yet know the best architecture for secure agent coordination, so we are backing several approaches in parallel. Some work on the protocols agents use, from specifying and designing them to auditing them; others work on the agents themselves:
- Agents that reason about security. The Agentic Security Reasoner turns task requirements into formal specifications, then selects or generates suitable protocols and verifies them.
- Protocols designed automatically. Dovetail takes the properties agents need, such as incentive compatibility, safety, and liveness, and designs a protocol proven in Lean to deliver them.
- Protocols audited by agents. ProtoSage and ProtoBench build and evaluate agents that audit cryptographic protocols and recover their specifications.
- Agents hardened under pressure. Project HAMMER tests how agents negotiate, bid, and cooperate under adversarial pressure, and hardens them against manipulation and leakage.
Integration across the stack
Ultimately, many different components are being built, and someone needs to glue them together in useful ways and work on their integration. That is the role of CCTI, our first integration partner: mapping how Creators’ work fits together, co-developing shared building blocks, and combining the pieces into an end-to-end application that shows where interfaces are missing.
Partners sit at either edge of the stack. Matta grounds the physical side in real manufacturing data; on the digital side, the slots for open-source practice and real use cases are still open. If you maintain open-source infrastructure, operate a real environment, or have a use case that needs agents to coordinate securely, reach out!
The Arena and the rest of the programme
We expect much of what Creators build to be put to work in the Arena, where agents run organisations that trade, make payments, and operate machines under adversarial pressure. In turn, the Arena will open up new needs for the rest of the programme.
Say an agent running one Arena organisation buys a part made by another. Using technology developed by CCTI, the seller could prove the part was produced to the agreed specification; with physical watermarking, the buyer could check which controller the robot making it was actually running.Martin Kleppmann describes this scenario, and how CCTI plans to make it work, in his announcement of CCTI. If red teams then trick the buyer’s agent into overpaying or leaking its budget, that failure points to what we should fund next.
Open source by default
Everything we fund is open source by default. People need to be able to inspect, test, and improve infrastructure they are being asked to trust, and to run, adapt, or replace the infrastructure their agents depend on. ARIA is funding the first versions, but open source lets a global community develop, maintain, and use them long after. We want to build this together.
What is still missing
Our first Creators are the nucleus of the portfolio, and we will keep building around them; the empty slots on the map show where there is room. We fund new teams every quarter through our rolling solicitation for Tracks 2 and 3, and the current round closes on 31 October.Teams funded through our joint call with Google DeepMind, the Cooperative AI Foundation and Schmidt Sciences will join this portfolio. Separately, ARIA’s Opportunity Seeds fund related projects outside the programme. We will announce both soon. If you see a slot you could fill, or one we have not drawn yet, we encourage you to apply.
Below are some directions we have been thinking about. They are not prescriptive, just areas we think could be interesting. Keep an eye out on our blog for other ideas as they come up!
Track 2: the cyber-physical agent stack
- Early demonstrations. Small, convincing cases of agents acting for different owners using cryptographic tools to do something useful.
- Safe access to physical environments. Enforceable permissions and bounded access to sensors and actuators, so a host can give an outside agent real control without losing control of the environment.
- Physical evaluations. Building on Physical Evals: environments where independently owned agents delegate or coordinate real-world work, exposing failures of coordination, security, and incentives.
Track 3: fundamental theory
- Generative cryptography. The research loops in Generative Cryptography: cryptographic libraries in Lean, formalised problems, verification tools, and better routes from specification to protocol.
- Cryptographic resilience. Verified implementations, protocols that can swap or combine assumptions, benchmarks that track AI’s ability to find weaknesses, and fallbacks if key assumptions fail.
- Physical verification. Early proof that delegated physical work actually happened, especially from protocols whose guarantees come from physical constraints.