<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Projects on Scaling Trust Community</title><link>https://scalingtrust.org.uk/categories/projects/</link><description>Recent content in Projects on Scaling Trust Community</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Wed, 07 Oct 2026 00:00:00 +0100</lastBuildDate><atom:link href="https://scalingtrust.org.uk/categories/projects/index.xml" rel="self" type="application/rss+xml"/><item><title>Agentic Security Reasoner</title><link>https://scalingtrust.org.uk/projects/agentic-security-reasoner/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/agentic-security-reasoner/</guid><description>Tariq Elahi · University of Edinburgh</description><content:encoded><![CDATA[<p><strong>Agentic Security Reasoner for Emerging Multi-agent Protocols</strong></p>
<p>An agentic security stack that captures security and privacy risks, requirements and functionality boundaries, translates task requirements into formal specifications, selects or generates suitable cryptographic protocols (e.g. secure comms, MPC, ZK), and verifies their properties with Lean, ProVerif and similar tools. The goal is an end-to-end workflow in which agents, from orchestrators to sub-agents, can negotiate security needs and connect verified protocol descriptions to executable multi-agent interactions.</p>
<p><strong>Team:</strong> Myrto Arapinis, Jianyi Cheng, Michele Ciampi, Wenda Li (University of Edinburgh)</p>
<h2 id="outputs">Outputs</h2>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
<h2 id="background-reading">Background reading</h2>
<ul>
<li><a href="https://arxiv.org/abs/2205.12615" target="_blank" rel="noopener noreferrer">Autoformalization with Large Language Models</a> — Yuhuai Wu, Albert Q. Jiang, Wenda Li et al. · NeurIPS 2022. LLMs translating natural-language statements into formal specifications.</li>
<li><a href="https://arxiv.org/abs/2210.12283" target="_blank" rel="noopener noreferrer">Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs</a> — Albert Q. Jiang, Sean Welleck, Wenda Li et al. · ICLR 2023. LLM-written proof sketches that guide automated provers, the verification side of the stack.</li>
</ul>
]]></content:encoded></item><item><title>Andon Labs</title><link>https://scalingtrust.org.uk/projects/andon-labs/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/andon-labs/</guid><description>Andon Labs · Arena partner</description><content:encoded><![CDATA[<p>Builds Arena v1 — the environment, scenarios, and mechanics agents are tested in — then supports launch, maintenance, and iteration.</p>
<p><a href="https://andonlabs.com/" target="_blank" rel="noopener noreferrer">Andon Labs</a> is an AI safety and real-world evaluations startup, the team behind Vending-Bench and many real-life autonomous organisations. Few teams have run as many autonomous businesses in the wild; that intuition and experience guide the Arena&rsquo;s design towards something that can demonstrate new findings.</p>
<p><em>Selected after a public RFP, subject to contract and negotiation. See our <a href="/blog/update-on-the-scaling-trust-arena/">September update on the Arena</a>.</em></p>
<h2 id="outputs">Outputs</h2>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
<h2 id="background-reading">Background reading</h2>
<ul>
<li><a href="https://andonlabs.com/evals/vending-bench-arena" target="_blank" rel="noopener noreferrer">Vending-Bench Arena</a> — Andon Labs · ongoing evaluation. Several AI models each run a vending business in the same market, trading, messaging and paying each other: a small-scale precursor of the Scaling Trust Arena.</li>
<li><a href="https://arxiv.org/abs/2502.15840" target="_blank" rel="noopener noreferrer">Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents</a> — Axel Backlund, Lukas Petersson · 2025. The original single-agent benchmark, in which a model runs a simulated vending business over long horizons.</li>
</ul>
]]></content:encoded></item><item><title>Foundations for AI-Native Proof Systems</title><link>https://scalingtrust.org.uk/projects/ai-native-proof-systems/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/ai-native-proof-systems/</guid><description>Alessandro Chiesa · EPFL</description><content:encoded><![CDATA[<p>This project will investigate whether the recurring structure of AI computations can support more efficient proofs than generic circuit-based approaches. It also explores self-proving models, and the boundary between computations that can remain black-box and those that must be decomposed for verification.</p>
<h2 id="outputs">Outputs</h2>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
<h2 id="background-reading">Background reading</h2>
<ul>
<li><a href="https://arxiv.org/abs/2405.15722" target="_blank" rel="noopener noreferrer">Models That Prove Their Own Correctness</a> — Noga Amit, Shafi Goldwasser, Orr Paradise, Guy N. Rothblum · NeurIPS 2025. Introduces self-proving models, which learn to produce checkable proofs alongside their answers.</li>
<li><a href="https://arxiv.org/abs/2404.16109" target="_blank" rel="noopener noreferrer">zkLLM: Zero Knowledge Proofs for Large Language Models</a> — Haochen Sun, Jason Li, Hongyang Zhang · ACM CCS 2024. Specialised proofs for attention and tensor operations: an example of exploiting the structure of AI computations instead of generic circuits.</li>
</ul>
]]></content:encoded></item><item><title>Advanced Cryptography for AI</title><link>https://scalingtrust.org.uk/projects/advanced-cryptography-for-ai/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/advanced-cryptography-for-ai/</guid><description>Tom Gur · University of Cambridge</description><content:encoded><![CDATA[<p>Privacy techniques suited to AI workloads, built with cryptography: private retrieval for RAG, semantic search, secure computation, and methods for concealing queries or embeddings. The aim is to let agents use shared memory and sensitive data without exposing commercially or personally confidential information.</p>
<p><strong>Team:</strong> Nir Bitansky (NYU); Yuval Ishai (Technion); Ron Rothblum (Succinct/Technion); Sarah Meiklejohn (UCL/Google)</p>
<h2 id="outputs">Outputs</h2>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
<h2 id="background-reading">Background reading</h2>
<ul>
<li><a href="https://eprint.iacr.org/2022/949" target="_blank" rel="noopener noreferrer">One Server for the Price of Two: Simple and Fast Single-Server Private Information Retrieval</a> — Alexandra Henzinger, Matthew M. Hong, Henry Corrigan-Gibbs, Sarah Meiklejohn, Vinod Vaikuntanathan · USENIX Security 2023. SimplePIR: practical private retrieval, a building block for private access to shared data.</li>
<li><a href="https://eprint.iacr.org/2023/1438" target="_blank" rel="noopener noreferrer">Private Web Search with Tiptoe</a> — Alexandra Henzinger, Emma Dauterman, Henry Corrigan-Gibbs, Nickolai Zeldovich · SOSP 2023. Private search over semantic embeddings: the closest existing system to concealing queries and embeddings in retrieval.</li>
</ul>
]]></content:encoded></item><item><title>Amodo Design</title><link>https://scalingtrust.org.uk/projects/amodo-design/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/amodo-design/</guid><description>Amodo Design · Arena partner</description><content:encoded><![CDATA[<p>Designs and operates the physical side of the Arena.</p>
<p><a href="https://amododesign.com/" target="_blank" rel="noopener noreferrer">Amodo Design</a> is a UK hardware engineering company and one of ARIA&rsquo;s Activation Partners. Hardware hackers who invent, design and build novel scientific equipment, and work on securing advanced AI systems in hardware (including the flexHEG architecture) across ARIA programmes.</p>
<p><em>Selected after a public RFP, subject to contract and negotiation. See our <a href="/blog/update-on-the-scaling-trust-arena/">September update on the Arena</a>.</em></p>
<h2 id="outputs">Outputs</h2>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
]]></content:encoded></item><item><title>Dovetail</title><link>https://scalingtrust.org.uk/projects/dovetail/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/dovetail/</guid><description>Samuele Marro · Institute for Decentralized AI</description><content:encoded><![CDATA[<p><strong>Dovetail: Automated interaction protocol design</strong></p>
<p>An automated mechanism and protocol design engine. Agents specify desired properties, such as incentive compatibility, safety and liveness, and Dovetail designs an interaction protocol and proves in Lean that it satisfies those requirements. The team will also manage a cross-team repository of protocols, with the goal of speeding up automated protocol generation.</p>
<p><strong>Team:</strong> Emanuele La Malfa, Angelo Huang (Institute for Decentralized AI); Mirco Giacobbe, Gabriel Santos (Zeroth Research)</p>
<h2 id="outputs">Outputs</h2>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
<h2 id="background-reading">Background reading</h2>
<ul>
<li><a href="https://arxiv.org/abs/2410.11905" target="_blank" rel="noopener noreferrer">A Scalable Communication Protocol for Networks of Large Language Models</a> — Samuele Marro, Emanuele La Malfa et al. · 2024. Agora, the team’s earlier meta-protocol in which LLM agents write and agree on their own interaction protocols.</li>
<li><a href="https://arxiv.org/abs/2605.08426" target="_blank" rel="noopener noreferrer">Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI</a> — Xuanqiang Angelo Huang, Samuele Marro, Emanuele La Malfa et al. · ICML 2026 workshop. Where classical mechanism design falls short for LLM agents, the problem an automated design engine has to face.</li>
</ul>
]]></content:encoded></item><item><title>BT6</title><link>https://scalingtrust.org.uk/projects/bt6/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/bt6/</guid><description>BT6 · Arena partner</description><content:encoded><![CDATA[<p>Sets the Arena&rsquo;s security criteria and operating model, then red-teams it: adversarial signal, challenge design, and community once it is live.</p>
<p><a href="https://bt6.gg" target="_blank" rel="noopener noreferrer">BT6</a> is a frontier-AI red team. Open-source advocates who have stress-tested every frontier model and operate a large community of security experts; their expertise and playfulness make sure the Arena&rsquo;s adversarial design is thought-out, that it doesn&rsquo;t fail in easy ways, and that safety concerns are caught early.</p>
<p><em>Selected after a public RFP, subject to contract and negotiation. See our <a href="/blog/update-on-the-scaling-trust-arena/">September update on the Arena</a>.</em></p>
<h2 id="outputs">Outputs</h2>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
]]></content:encoded></item><item><title>Project HAMMER</title><link>https://scalingtrust.org.uk/projects/project-hammer/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/project-hammer/</guid><description>Davide Crapis · Multiscalar Intelligence</description><content:encoded><![CDATA[<p>Evaluates how AI agents negotiate, procure, bid and cooperate under strategic and adversarial pressure. It combines an open game-based evaluation suite with a hardening pipeline intended to produce agents that preserve economic value, resist manipulation and avoid leaking private information.</p>
<h2 id="outputs">Outputs</h2>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
<h2 id="background-reading">Background reading</h2>
<ul>
<li><a href="https://arxiv.org/abs/2603.20925" target="_blank" rel="noopener noreferrer">Profit is the Red Team: Stress-Testing Agents in Strategic Economic Interactions</a> — Shouqiao Wang, Marcello Politi, Samuele Marro, Davide Crapis · 2026. A profit-maximising opponent exploits agents in economic scenarios, and the exploits are distilled into defences: the evaluation-plus-hardening loop HAMMER scales up.</li>
</ul>
]]></content:encoded></item><item><title>ProtoSage and ProtoBench</title><link>https://scalingtrust.org.uk/projects/protosage-protobench/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/protosage-protobench/</guid><description>Peter Gilbert · RDI Foundation</description><content:encoded><![CDATA[<p><strong>ProtoSage and ProtoBench: Developing and evaluating agent capabilities in cryptographic protocol reverse-engineering, auditing, and verification</strong></p>
<p>ProtoSage is an agentic system for auditing cryptographic protocols, supported by ProtoBench’s corpus and evaluation tasks. The project builds on OpenSage and WireWatch to identify vulnerabilities, recover protocol specifications, and eventually produce machine-checkable models that experts and formal-verification tools can inspect.</p>
<p><strong>Team:</strong> Mona Wang (RDI Foundation)</p>
<h2 id="outputs">Outputs</h2>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
<h2 id="background-reading">Background reading</h2>
<ul>
<li><a href="https://arxiv.org/abs/2602.16891" target="_blank" rel="noopener noreferrer">OpenSage: Self-programming Agent Generation Engine</a> — Hongwei Li, Zhun Wang et al., with Dawn Song · 2026. The agent framework ProtoSage builds on, with security tooling for finding and reproducing vulnerabilities.</li>
<li><a href="https://doi.org/10.1109/SP61157.2025.00224" target="_blank" rel="noopener noreferrer">WireWatch: Measuring the Security of Proprietary Network Encryption in the Global Android Ecosystem</a> — Mona Wang, Jeffrey Knockel, Zoë Reichert, Prateek Mittal, Jonathan Mayer · IEEE S&amp;P 2025. Recovering and breaking proprietary encryption protocols in popular apps: the work ProtoSage aims to automate.</li>
</ul>
]]></content:encoded></item><item><title>Centre for Cryptographic Trust Infrastructure (CCTI)</title><link>https://scalingtrust.org.uk/projects/ccti/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/ccti/</guid><description>Martin Kleppmann · University of Cambridge; Martin Pompéry · SINE Foundation</description><content:encoded><![CDATA[<p>Enables mutually untrusting AI agents, for example representing different companies, to establish trust by cryptographically proving facts about those companies and their physical-world processes to each other: that a product was produced to a particular specification, say, or under what conditions an agent would be willing to reach an agreement.</p>
<p>CCTI is also the programme’s integration partner, mapping how Creators’ work fits together, co-developing shared building blocks with interested teams, and proposing open-source benchmarks and challenges to the Arena.</p>
<p><strong>Team:</strong> grjte (Ink &amp; Switch); Hossein Hafezi, Alireza Kavousi, Arman Kolozyan, Jessica Man (University of Cambridge); Jonathan Heiß, Ágnes Kiss, Aurel Stenzel (SINE); Daniel Hugenroth, Mario Lins (Light Squares)</p>
<h2 id="outputs">Outputs</h2>
<ul>
<li><a href="https://martin.kleppmann.com/2026/10/07/centre-for-cryptographic-trust-infrastructure.html" target="_blank" rel="noopener noreferrer">Announcing the Centre for Cryptographic Trust Infrastructure (CCTI)</a> — Martin Kleppmann’s announcement: the supplier-verification problem CCTI starts from, and how attested facts can feed zero-knowledge proofs about physical production processes (7 October 2026).</li>
</ul>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
<h2 id="background-reading">Background reading</h2>
<ul>
<li><a href="https://doi.org/10.1145/3744255.3811720" target="_blank" rel="noopener noreferrer">Emission Impossible: Cryptographically Verifiable Carbon Emissions Reporting for Cloud Computing</a> — Jessica Man, Martin Kleppmann · ACM e-Energy 2026. A datacenter operator proves to each customer, in zero knowledge, that its emissions report is accurate: the line of work CCTI generalises to other physical-world claims.</li>
</ul>
]]></content:encoded></item><item><title>Physical Watermarking</title><link>https://scalingtrust.org.uk/projects/physical-watermarking/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/physical-watermarking/</guid><description>Amanda Prorok · University of Cambridge</description><content:encoded><![CDATA[<p><strong>Physical Watermarking: Robust policy certification for embodied multi-agent systems</strong></p>
<p>A tamper-resistant physical watermarking framework to verify the provenance and safety compliance of control policies driving embodied robots. By extending the Colored Noise Coherency (CoNoCo) construction, it lets independent parties remotely authenticate active controllers using commodity hardware, such as standard CCTV cameras or smartphones, without direct access to the robot or specialised sensing equipment.</p>
<p><strong>Team:</strong> Manon Flageat, Mateusz Sypniewski, Sally Matthews (University of Cambridge)</p>
<h2 id="outputs">Outputs</h2>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
<h2 id="background-reading">Background reading</h2>
<ul>
<li><a href="https://arxiv.org/abs/2512.15379" target="_blank" rel="noopener noreferrer">Remotely Detectable Robot Policy Watermarking</a> — Michael Amir, Manon Flageat, Amanda Prorok · ICLR 2026. Introduces Colored Noise Coherency (CoNoCo), the construction this project extends: a signal hidden in a robot’s natural motion that can be detected from video.</li>
</ul>
]]></content:encoded></item><item><title>Manufacturing Grounding for Cyber-Physical Trust</title><link>https://scalingtrust.org.uk/projects/manufacturing-grounding/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate><guid>https://scalingtrust.org.uk/projects/manufacturing-grounding/</guid><description>Sebastian Pattinson, Douglas Brion · Matta</description><content:encoded><![CDATA[<p>Connects the programme to practical industrial needs. Drawing on its manufacturing expertise, the team will support the design and evaluation of cyber-physical trust tools that are relevant to real production environments, providing real-world data for other teams to use and keeping the programme grounded in operational reality as it explores how agents can interact securely across digital and physical systems.</p>
<p><strong>Team:</strong> Kate Lucas, Ciara Gumsheimer, Ollie Rosen (Matta)</p>
<h2 id="outputs">Outputs</h2>
<p>Repositories, papers, and demos will be linked here as the work gets underway.</p>
<h2 id="background-reading">Background reading</h2>
<ul>
<li><a href="https://doi.org/10.1038/s41467-022-31985-y" target="_blank" rel="noopener noreferrer">Generalisable 3D printing error detection and correction via multi-head neural networks</a> — Douglas A. J. Brion, Sebastian W. Pattinson · Nature Communications 2022. A fleet of printers producing over a million automatically labelled images, used to detect and correct errors in real time: the kind of real production data Matta brings to the programme.</li>
</ul>
]]></content:encoded></item></channel></rss>