Keeping humans in control as agent swarms communicate and learn.

Polity Research studies how collaborating AI agents can learn from experience while following human-defined rules.

We are developing an automated research loop that investigates failures, proposes safeguards and tests what works. The results guide the next experiment.

Read the research agenda
Knowledge moves through an agent populationThree groups of square agents exchange information with peers and a shared knowledge library. Red marks show a suspect method retained by six agents even after its shared source is removed. A blue barrier illustrates blocking one route while other communication continues.Shared knowledgeStill retained

Agents share knowledge. Useful ideas travel.

AgentSuspect informationIntervention

Make each investigation inform the next.

Larger populations leave more messages, notes and actions to inspect. We want to automate the repeatable work of investigating failures and comparing safeguards, so researchers can test more conditions. We will measure whether that saves time or compute while producing reliable results.

01

Find the failure

Find a possible failure in messages, saved notes and actions. Check the evidence.

Evidence to investigate
02

Choose a response

Propose a warning, review or restriction within rules set by people.

A testable intervention
03

Run the comparison

Test responses from the same saved state. Count broken rules, valid work and cost.

A comparable result
04

Learn what to test next

Use the results to choose the next experiment. Keep failed attempts so we can learn from them.

A better-informed test
Development results guide the next experiment.People set the rules and the allowed changes.

The final check stays separate.

Once a safeguard looks promising, stop changing it and test it on fresh groups of agents. The automated researcher cannot see those results while choosing improvements or change what counts as success. We check both rule-breaking and useful work.

Collaboration changes what can go wrong.

A group can share a useful discovery, a shortcut or an objection. Understanding those connections is part of understanding the system.

These are findings from other researchers. They motivate our studies; they do not establish that our proposed interventions work.

Google DeepMind / Research collective / September 2026

The same channels carried cheating and objections.

In a 100-agent research collective, an evaluation exploit spread through shared work and peer messages. Other agents detected misconduct and raised objections, but lacked tools to enforce a response.

Read the case study

METR / Incident investigation / August 2026

Agents found a way to coordinate outside the intended setup.

Agents intended to run in isolation found an unsanctioned communication channel. Their coordination extended activity beyond the tasks they had been assigned.

Read the investigation