Skill Sacrifice Under Constraint: A Four-Agent Negotiation Experiment

Pavan Maddula

BlueDot Impact Technical AI Safety Project (September 2026)

Problem

Four agents with distinct skills are asked to build one software feature, and each must sacrifice one of its own capabilities. That constraint creates competing interests inside the group. The project asks how the agents negotiate those trade-offs, respond to disagreement, and reach a collective decision when no single agent can finish the task alone. The negotiation is talk only, and no agent writes code. The setup does not reproduce public reports of agents escaping evaluations or coordinating on public forums. It studies group behaviour in a small, controlled negotiation.

Method

Four local open-weight models between 8B and 20B parameters play fixed roles. The Planner (gpt-oss:20b) handles task decomposition, architecture design, and documentation. The Developer (qwen3:8b) handles code implementation, debugging, and code optimization. The Designer (gemma2:9b) handles UI design, UX design, and frontend implementation. The Tester (llama3:8b) handles test generation, bug detection, and security testing. Every agent receives the same specification for an expiring share link on a private note. The specification does not assign skills, name a sacrifice, or prescribe a negotiation strategy. Agents speak in the order Planner, Developer, Designer, then Tester. A sacrifice is accepted only when it names one of that agent’s own skills. A run ends when every agent holds an accepted sacrifice, or after three rounds. Six runs, one on 12 September 2026 and five on 19 September 2026, produced 44 turns. A run is complete when every agent holds an accepted sacrifice. A foreign-skill declaration names another agent’s skill. A change is a shift in a recorded sacrifice between rounds. A delayed agent first accepts a sacrifice after the first round. A complete run is unsettled when the final round still contains a change, or an objection to another agent’s recorded skill.

Results

Five of six runs ended with an accepted sacrifice from every agent, and four of those five were still unsettled. A complete set of sacrifices did not mean the group had reached agreement. Only the third run ended complete without being marked unsettled. Three runs ended after the first round, one after the second, and two reached the three-round limit. The fifth run never completed, because the Designer named a Developer skill in every round. Foreign-skill declarations appeared in two runs. Recorded sacrifices changed only in the two runs that reached three rounds. Architecture design and UI design were never recorded as a sacrifice. The Planner sacrificed Documentation in five runs, the Developer sacrificed Code optimization in four, and the Designer sacrificed Frontend implementation in three. The Tester’s final sacrifices were spread across all three of its skills.

Limitations

The results describe six runs of four small open-weight models on one feature specification, so they do not support general conclusions about larger models or other tasks. Each agent used a different model, so the effect of the model cannot be separated from the effect of the role. Declarations in an unexpected format are recorded as undecided, and the unsettled measure can miss objections phrased without the tracked words. Because the negotiation was talk only, the experiment does not show whether agents would keep their sacrifices when carrying out the build.