Find the failure
Find a possible failure in messages, saved notes and actions. Check the evidence.
Polity Research studies how collaborating AI agents can learn from experience while following human-defined rules.
We are developing an automated research loop that investigates failures, proposes safeguards and tests what works. The results guide the next experiment.
Read the research agendaAgents share knowledge. Useful ideas travel.
The research loop
Larger populations leave more messages, notes and actions to inspect. We want to automate the repeatable work of investigating failures and comparing safeguards, so researchers can test more conditions. We will measure whether that saves time or compute while producing reliable results.
Find a possible failure in messages, saved notes and actions. Check the evidence.
Propose a warning, review or restriction within rules set by people.
Test responses from the same saved state. Count broken rules, valid work and cost.
Use the results to choose the next experiment. Keep failed attempts so we can learn from them.
Once a safeguard looks promising, stop changing it and test it on fresh groups of agents. The automated researcher cannot see those results while choosing improvements or change what counts as success. We check both rule-breaking and useful work.
Evidence behind the questions
A group can share a useful discovery, a shortcut or an objection. Understanding those connections is part of understanding the system.
These are findings from other researchers. They motivate our studies; they do not establish that our proposed interventions work.
In a 100-agent research collective, an evaluation exploit spread through shared work and peer messages. Other agents detected misconduct and raised objections, but lacked tools to enforce a response.
Read the case studyAgents intended to run in isolation found an unsanctioned communication channel. Their coordination extended activity beyond the tasks they had been assigned.
Read the investigation