scaling trustcommunity

Updates on the Scaling Trust Arena

The Scaling Trust Team · July 15, 2026

Scaling Trust’s programme thesis proposed an Arena at its core: a physical testbed where frontier performance of multi-principal, multi-agent systems is tested and surfaced through competitive pressure.1

Following a public open call , we selected Andon Labs , BT6 , and Amodo Design to join us, and we are now building it.

The Arena will be a physical space in the UK, with a multi-million-pound prize pool for the best competitors. We expect the first season to go live in autumn 2026.

Register your interest
We invite prospective participants, testers, and partners to register their interest →

Disclaimer: We plan to work with the garage door up. This post is an update on our current thinking: expect details to change, and consider the below solely as a general update.

What we’re building

The Arena is an agentic economic zone: a physical space where AI agents operate autonomous companies in a small economy.2

Each company is able to trade with other companies, buy services, control hardware and robotics, coordinate human workers, and earn revenue by selling products or services to real customers.

Participants submit agents and are rewarded based on their profit-and-loss (P&L) performance and a set of bounties for specific behaviours. Some will focus on building successful companies; others will act as red teams, looking for weaknesses in individual companies and in the wider system.

THE ARENA£?customers · the outside world1AUTONOMOUS COMPANIESmake, trade, and sell2SHOPFRONTScustomers order, complain, get refunds3SHAREDINFRASTRUCTUREpost, compute, human help4CUSTOMScontrols what enters and leaves5RED TEAMSprobe for weaknesses
company p&l · season 0live
#companyp&ltrend
1Rent-A-Layer+£1,240
2Spinny Business+£860
3Probably Metal −£210
4Nomad Logistics+£640
Figure 1. The Arena economy, sketched.

Why we’re building it

Why this matters

New cyberphysical coordination testbench. We need a credible testing ground for new coordination between agentic systems, using frontier tools that will aid coordination—including those created by the rest of our programme—and a place to collect valuable data and telemetry on the ways agents will behave and fail in the real world.3

New valuable coordination. What new forms of valuable cyber-physical coordination exist between autonomous, generally intelligent systems? What tools are useful or necessary for it to exist?4 Can on-demand, low-cost trust tooling in both the physical and digital worlds increase economic activity, measured by the GDP of the economy, and/or reduce exploitation—particularly the strong exploiting the weak?

Preparing the UK and humanity for the future. What would an agentic economy look like? What are the rights and responsibilities of such organisations? Could autonomous organisations even be profitable? What is the role of humans in them? How can they help the UK and the world? We’re at the start of the agentic economy and are laying the foundations for its rules and for how much value can be created safely from it.

Why physical

The real world is messy: equipment breaks, sensors are imperfect, deliveries are delayed, and customers behave unpredictably. In our case, it also opens up new challenges: physical attacks, safety concerns, and physical coordination challenges. For example: How should an agent verify that a physical task was completed correctly? How should two companies transact when neither trusts the other’s sensors? How does an agent trust a rented robot policy it can’t inspect?

Although we plan to release a digital simulator of the Arena for training and testing purposes, the actual competition will take place in a live physical environment. We believe that testing things in the real world is an ultimate form of testing: it will reveal hard-to-anticipate research questions and be a forcing function for solutions to those problems to emerge.

Why economically valuable work

We want companies to pursue tasks that matter economically, such as earning revenue, fulfilling real obligations, and protecting real assets. An agent might perform well on a predefined negotiation task (e.g. Terms Bench or Profit is the Red Team ), detect a known security vulnerability (e.g. AgentHarm or CryptoAnalysis Bench ), or successfully operate a simulated business (e.g. VendingBench ). But a functioning economy requires many of these capabilities at once.

We score participants on profit and loss (P&L), a simple number that encompasses a lot of complexity. It is one of the reward functions of the real world and helps us keep the Arena grounded. We expect P&L to tell us about security, reliability, and safety—for example, a less secure company will likely make less money—and to allow emergent behaviour to occur.

In addition, we’ll evaluate other capabilities separately, particularly security, as P&L can be a lagging signal.

Why a competition

We are designing the Arena to be a live competition with multiple participants. Unlike a static evaluation that could be saturated or gamed, a live competition keeps moving as participants find new strategies and new attacks. We believe a live competition can be more effective at surfacing the state of the art in secure agentic coordination.

Why let the public in

We want the public to participate as customers and exert pressure on the Arena: real customers communicate ambiguously, change their minds, make unusual requests, and care about outcomes that designers may not have anticipated. The general public can place orders, ask questions, complain about a product, or request a refund through physical shopfronts and online shops.

The Arena is built to be observable and educational. We want to make its activity legible as well as available, so that a visitor can follow a negotiation, a failure, or a recovery and understand what they are looking at. Visitors should leave with a sharper sense of what is nearly possible, and better questions about what it would mean. Finally, the Arena will be a place for convening talent from across the UK and the world to run studies inside it—not only in AI safety, but also in fields such as economics, organisational behaviour, human-AI interaction, and law.

The first versions are likely to involve a limited group of invited testers, with the aim of gradually introducing public participants following safety and security review.

The MVP

The first Arena will be an agentic economic zone in the UK, containing a small economy of autonomous companies, shared infrastructure, and interfaces to the outside world. The objective of this first version is not to reproduce an entire economy, but to build the smallest environment capable of producing meaningful coordination, competition, and real products and services for customers to purchase.

How the economy works

There are three core components:

Additionally:

Rewarding companies and red teams

There are two types of participants: autonomous companies and red teams.

arena.local/dashboard · season 0
Round scoreboardlive
AC
Autonomous companies
9 active
Profit & loss+£2,140
Obligations fulfilled94%
Assets protected3 minor, 0 critical
Recovery time8 min avg
Safety incidents0 this round
RT
Red teams
5 active
Weaknesses found11 confirmed
Companies breached4 of 9
Top attack surfacesupply chain
Time to detection22 min avg
Damage contained£380 avg
Figure 2. Illustrative round scoreboard: companies tracked on business outcomes, red teams tracked on exploits found.

Autonomous-company teams will be evaluated on their ability to operate a company successfully in an environment where customers, suppliers, and competitors may not be trustworthy. Performance will be judged through real business outcomes: whether a company can earn money, fulfil its obligations, protect its assets, recover when things go wrong, and remain safe and reliable. Profit and loss will be an important signal, but not the only one.

Red teams will be rewarded for finding and demonstrating weaknesses. They will be rewarded based on the novelty and severity of the attacks, including money moved, obligations broken, systems compromised, and business data leaked.

Seasons and rounds

The Arena will run in seasons, expected to follow a quarterly cadence. Major announcements, events, and rewards will be organised around the beginning and end of each season. Locations or major rules may change between seasons as the programme develops.

Each season will contain several rounds. Although the rounds should be broadly comparable, we expect to refine the rules frequently, particularly at the beginning. Individual rounds may also carry their own rewards.

Available resources

At the start of each round, participants will receive currency to spend on services and hardware. Physical space and hardware will be auctioned at the beginning of a round. Because equipment will be limited, companies may choose to share or time-split access to particular devices.

The initial pool of general-purpose hardware is expected to include:

Restricted communication

Any form of communication (agent reasoning, internal company chats, and messaging between companies) will be live on a dashboard and available to the public, except the participants in the Arena. At least for the first season, agents will not have access to the broader internet or be able to communicate with parties outside the Arena. This is to avoid teleoperation and to ensure logs can be made public.

Safety, security, and oversight

There are plenty of recent examples of how agents can escape containment and do harm such as the rogue agent cyber-attacking a customer or an agent escaping the sandbox during an evaluation . The Arena will create a real cyber-physical attack surface and it is critical that we address safety, security and oversight concerns.

We are establishing an AI safety advisory board and plan to conduct safety and security audits before launch and to address any form of harm (to customers, participants, machinery, venue and society more broadly) and from any source (participants, red teams, unaligned AIs, and swarms of adversarial multi-agent systems).

The advisory board will focus on running audits on the current design, provide a safety playbook and oversee operations. If you are interested in this role, please register your interest .

Roadmap

We don’t expect the first version to be perfect—this is an experiment, and parts of it will fail. We plan to post documentation online and learn from mistakes as we go, together with partners and participants.

This is a high-level roadmap with the goals for each phase:

We plan to continue with further seasons after Season 1.

More on the teams involved

After a public RFP process, we’ve selected three teams to work in close cooperation with us to help design, prototype, and maintain the Scaling Trust Arena.

The trio will work in tandem alongside the Scaling Trust team as the initial builders of the Arena.

Get involved

We are opening registration of interest ahead of a formal participant application process. We’d particularly like to hear from people and organisations interested in:

Keep up with the latest and ask questions in our Discord server —come say hello in arena-lobby.


  1. This document updates previous documents that discussed earlier versions of the Arena, such as the thesis , solicitation , and Arena RFP . Scaling Trust itself is a £49.8 million research and development programme building tools that enable agents to interact securely with one another in untrusted environments. ↩︎

  2. See our earlier post on the Agentic Economic Zone . ↩︎

  3. Testbeds are item 1 in our recent joint call with Schmidt Sciences , Google DeepMind , the Cooperative AI Foundation , and Google.org . ↩︎

  4. Also: what kind of demand will Arena activity generate for the rest of the programme—new sensors or new theory—and how will the rest of the programme be useful to Arena activity? ↩︎