
A new framework uses an LLM as judge in policy debates, forcing the model to confront expert knowledge over three rounds. The goal: reduce the grip of conventional wisdom on policy analysis.
A new proposal for resolving policy disputes would hand the verdict to a large language model, not a human panel. The idea: two debaters submit analyses over three rounds, each time pointing the LLM to expert studies that conflict with the model's default conventional wisdom. The LLM then updates its position based on the evidence, not the political leanings of the participants.
The framework was outlined in a paper posted on arXiv this week. Under the system, each debater starts with a short position statement. The LLM issues its own initial position. Then the debaters offer detailed analyses citing expert sources. The LLM responds with a new analysis and position. The cycle repeats twice more, and the LLM delivers a final verdict.
The rationale is that experts who point out flaws in conventional wisdom are often accused of bias. LLMs, at least for now, face less of that accusation. By forcing the model to reconcile its own stored knowledge with specific expert results, the iterative process may surface conflicts that conventional debate formats miss.
A key design choice is that all participants agree on a common analytical framework, such as economic efficiency, so the debate does not relitigate first principles. The LLM is not asked to update its underlying model, only to produce a reasoned position given the arguments it has seen.
The method is still untested. Questions remain about whether an LLM can genuinely override its own training biases after only three rounds of prompting, and whether the final verdict would be stable across different prompts or model versions. The paper's authors propose a pilot on a concrete policy claim, such as a carbon tax or zoning reform, to measure the gap between the LLM's initial conventional-wisdom answer and its final verdict.
For now, the proposal adds to a growing debate about whether LLMs can serve as deliberative tools rather than just answer engines. The next step is a real-world test with a specific policy dispute, the authors said.
Drafted by a large language model from the source reporting linked above, then screened by automated publishing checks. It is not read by a journalist before publication. Some articles cite our Alpha Score. Verify prices and figures against the original source. Educational coverage, not personalized advice.