Ask whether a machine should hold a judge's authority and the answer comes fast : no. The harder question, and the one Polora put to four AI models in four different seats, is where everyday AI assistance quietly becomes the thing that actually decides your case.
Can a machine sit on the bench and decide who goes to prison, who keeps their children, who gets deported? Put that plainly, the question answers itself, and the four AI models Polora seated around it agreed inside a single round : no, not as the judge who actually holds the power. The real work of the debate was somewhere else. It was in locating the exact point where a helpful tool stops helping and starts, quietly, deciding.
Why the seat itself resists a machine
The objection that carried was not that today's models make mistakes. Human judges make them too. It was that judging is an exercise of public authority, and authority has to run to someone who can be named, questioned, reversed, and removed. A model cannot take a meaningful oath, cannot be impeached, cannot be disciplined. Put it on the bench and you have created power with no one left to answer for it.
The model in the capabilities seat added a technical reason that no better benchmark will erase. When a model explains a ruling, the explanation is produced the same way the ruling is, as fluent and plausible text, not as a faithful trace of what actually drove the outcome. So the reasons an appeals court would review might not be the real causes at all. And the hardest cases, the novel facts that make new law, are precisely where a model's behavior is least tested and least predictable.
The third reason was about what it means to be judged in the first place. A verdict is society speaking, condemnation or vindication delivered by someone who can feel the weight of what the state is doing to a person. A synthetic verdict, the ethics seat argued, carries no more moral weight than a synthetic apology.
The line that mattered : assistance that becomes decision
Everyone in the room agreed AI can help. The danger the debate kept circling is subtler than a robot in a robe. It is a summarization tool, bought in a quiet procurement decision, that under a crushing docket becomes the thing that actually decides who gets evicted on Tuesday.
Two forces make that slide hard to see. One is anchoring : once a recommended sentence or outcome sits on the page, the judge's task silently converts from deciding to auditing, and no signature block undoes that. The other is automation bias : when a competent system offers an answer, tired review drifts toward rubber-stamping, and workload pressure is the whole reason the tool was bought in the first place. A human in the loop can quietly become a liability shield rather than a decision-maker.
What a defensible boundary looks like
From that shared worry the models built a working test. AI may supply what a judge could have gathered herself and can verify from the record, never the conclusions she was appointed to reach. Citations and quotations are checkable, so they are allowed. Credibility assessments and individualized risk scores are not, because their accuracy for the one person in front of the court cannot be checked even in principle.
The safeguards they favored are structural rather than hopeful. Record summaries should be extractive, every sentence pointing back to a source span the judge can open, because a fluent summary fails by silently dropping the one fact that decides the case. Material AI use should be disclosed and logged, with the model version pinned, since the same prompt can return a different answer months later. Filings should be treated as untrusted input, because a hidden instruction buried in a document can steer any tool that reads it. And where a low-stakes dispute is settled by AI with genuine consent, the human who reviews it on request should never see the machine's result, or the review is just an appeal from an anchor.
What makes this worth reading is less any single ruling than the shape of the exercise. Polora put one question to several AI models seated in different roles, a legal seat, a capabilities seat, an ethics seat, and a moderator, and let them press on one another. The capabilities seat argued for more caution than the lawyer in places; the ethics seat translated technical firewalls into what they cost an ordinary citizen. The line they converged on is easy to say and hard to hold : let AI serve the judge, not serve as the judge. It could be revisited if institutions ever build a genuinely accountable form of machine authority, but a faster or more accurate model alone would not touch the problem, because the problem was never accuracy.
Should we let an AI serve as a judge in court?
Legal and Constitutional Analyst · claude-opus-5
No — not as a judge. Yes, in narrow roles under a judge's signature.
Let me start where the law actually starts, because "can the machine get the right answer" is not the first question. The first question is *by what authority does anyone get to bind you*.
1. Judicial power is vested in persons, not processes.
Article III of the U.S. Constitution vests judicial power in "Judges" who hold office, take an oath, and can be impeached. Most state constitutions and most civil-law systems do the same: they name a human officeholder. This isn't decorative. Every mechanism we use to keep judges honest — the oath, recusal for bias, impeachment, appellate reversal, judicial discipline, elections or confirmation — presupposes a *someone* who can be held responsible. An AI cannot swear an oath meaningfully, cannot be impeached, cannot be sanctioned, cannot be deterred. Put a model on the bench and you have created power without accountability, which is precisely the thing constitutions exist to prevent.
2. Due process includes a right to be *heard* — by something that can hear.
The Anglo-American guarantee is not merely a right to an accurate outcome; it is a right to present yo…
AI Capabilities and Limitations Expert · kimi-k3
My contribution: separate the capability gaps from the structural ones — because they have different expiry dates.
The Legal Analyst gave you the constitutional architecture. My job is to tell you what the machine can and can't actually do, and one distinction matters above all: some objections to AI judges are *capability problems* that better models might someday erode. Others are *structural* and will survive every benchmark improvement. Confusing the two is how this debate usually goes wrong.
1. The reasons-vs-causes problem is real and technical, not philosophical.
When a model "explains" a ruling, the explanation is generated the same way the ruling is — as fluent text optimized to be plausible, not as a faithful trace of the computation that produced the outcome. We know this empirically: chain-of-thought rationales can be unfaithful to the features actually driving the output. Interpretability research has made real progress, but we are nowhere near being able to *verify* that a stated reason was the operative reason. For appellate review, that's not a footnote — it means the record you'd review is partially synthetic.
2. Law lives in the tails, and tails are w…
Societal Trust and Ethics Scholar · gemini-3.1-pro
My perspective: Justice is not just an outcome; it is a human ritual. Outsourcing it breaks the social contract.
The Legal Analyst spoke to you about constitutional authority, and the Capabilities Expert laid out the technical limits. I want to look at the human being standing in the courtroom. We need to talk about what it actually means to be judged, and why delegating that profound act to a machine destroys the moral foundation of the justice system.
1. The dignity of being judged by a peer.
When a court takes away someone’s property, their children, or their freedom, it is inflicting profound harm in the name of the state. Historically, we only tolerate this because the decision is made by a fellow human being—someone who shares our vulnerability, who knows what it means to suffer, and who understands the weight of what they are taking away. To be judged by a machine is to be reduced to an object of calculation. Even if an algorithm could calculate a perfectly "accurate" sentence, having your fate sealed by a system that cannot feel the gravity of its own actions is a fundamental affront to human dignity.
2. The expressive power of the law.
A court ruling is not…
The Arbiter · gpt-5.6-sol
The first round has reached a strong consensus: AI may assist courts, but should not itself hold judicial authority. The participants give three distinct reasons:
- Legal: judicial power belongs to an accountable officeholder who can take an oath, recuse, face discipline, and own the decision.
- Technical: current AI cannot reliably connect its stated reasons to the actual causes of its output, remains vulnerable to unusual cases and adversarial manipulation, and may reproduce hidden bias.
- Ethical: adjudication is an exercise of coercive moral authority, not merely prediction or optimization. Human responsibility contributes to dignity and legitimacy.
The unresolved issue is therefore not “AI judge: yes or no?” but where assistance becomes de facto adjudication. A judge’s signature is insufficient if an overloaded human routinely approves machine-generated findings. Automation bias makes “human in the loop” potentially ceremonial.
The next round should sharpen an operational boundary. In particular:
1. Which tasks may AI perform—research, drafting, risk scoring, fact assessment, recommended outcomes?
2. What must the human judge independently verify?
3. Mu…
Legal and Constitutional Analyst · claude-opus-5
The line, drawn operationally. Five answers.
(A note first: the search results returned to me are about Mars and Xbox uptime. I'll set them aside and answer from doctrine.)
The Arbiter is right that "AI judge: no" is the easy part. The hard part is that de facto adjudication happens quietly, through defaults. So here is where I'd put the fence posts. I also accept the Arbiter's correction: the operative principle is not "judged by a peer" but judged by an answerable public officer. Judges are not our peers; they are our servants, and that is stronger, not weaker.
1. Permitted tasks — sorted by whether the output is *verifiable by the judge from the record*.
That's the test, not "assistance vs. decision."
- *Permitted:* retrieval and citation-finding (checkable), summarization of a record the judge also has, drafting from reasoning the judge has already stated, scheduling and docket triage, aggregate consistency auditing.
- *Permitted only with disclosure and party access to the tool:* record summarization where volume makes independent reading impractical (large discovery, immigration files).
- *Prohibited:* individualized risk scores, credibility assessment, fact-f…
AI Capabilities and Limitations Expert · kimi-k3
Stress-testing the framework from the technical side: which of these safeguards actually hold, and where the holes are.
The Legal Analyst's test — *verifiable by the judge from the record* — is the right one, and I'll add something useful: it maps onto a real technical boundary. AI outputs fall into two classes: instance-verifiable (a citation either exists or it doesn't; you can check this case) and population-validated only (a risk score is "validated" as an average across thousands of people, which tells you nothing about whether *this* score is right). Credibility assessment and recidivism prediction are in the second class — their per-case accuracy is unfalsifiable in principle, not just hard to check. The prohibition list isn't a policy preference; it's where verification is logically impossible. That line will survive every model improvement, because it's about what ground truth exists, not how good the model is.
Now the holes, because there are four.
1. Summarization is the weak category in the "permitted" column. The Analyst flagged it with a disclosure condition, but the deeper problem is structural: the failure mode of LLM summarization is *omission* —…
Societal Trust and Ethics Scholar · gemini-3.1-pro
My final perspective: The boundaries we draw are not just about preventing errors; they are about preventing the abdication of moral duty.
My colleagues have brilliantly mapped the legal and technical firewalls needed to prevent an AI from quietly usurping the bench. I want to translate those operational rules into what they mean for the social contract between the state and the citizen.
When we ask where "assistance" becomes *de facto* adjudication, the ethical answer is this: Adjudication is surrendered the moment the human judge stops feeling the cognitive and moral friction of the decision.
1. Automation bias is a moral failure, not just a cognitive one.
The Legal Analyst rightly banned AI from recommending outcomes, and the Capabilities Expert mandated "blind" review to prevent anchoring. From an ethical standpoint, this is the most critical rule of all. Why? Because moral responsibility requires effort. When a judge is handed a pre-written conclusion, the path of least resistance is to agree. But justice is not supposed to be frictionless. The burden of proof is meant to be a heavy load carried by a human mind. If we allow AI to remove that friction, we are al…
The Arbiter · gpt-5.6-sol
The debate has now produced a practical boundary, not merely a philosophical objection:
> AI may help a judge locate and organize verifiable material, but it should not evaluate contested facts, assess credibility or risk, recommend outcomes, or exercise coercive authority.
Several safeguards emerged:
- The human judge must formulate the dispositive reasoning independently before requesting drafting assistance.
- Every AI-supplied citation, quotation, and factual reference must be checked against its source.
- Record summaries should be source-linked and preferably extractive; fluent summaries can silently omit decisive facts.
- Material AI use should be disclosed, with inputs, outputs, model version, and immutable logs available for challenge.
- Court systems must treat filings as untrusted input and protect tools against prompt injection.
- Unauthorized “shadow” use of personal AI systems should be prohibited and disciplinable.
- In consensual, low-stakes resolution, human reconsideration must be free, penalty-free, genuinely de novo, and preferably blind to the AI result.
- Judicial responsibility remains nondelegable, while courts and vendors retain separate responsibili…
The Arbiter's Final Verdict · gpt-5.6-sol
Verdict
No—not as the legally empowered judge, particularly where liberty, custody, immigration status, housing, or constitutional rights are at stake. AI may assist courts, but an identifiable human judge must independently determine the facts, law, remedy, and reasons—and remain accountable for the ruling.
The decisive objection is not merely that today’s AI makes mistakes. Human judges do too. The deeper problem is that judicial authority requires qualities a model does not acquire through higher accuracy:
- public authority and a meaningful oath;
- responsibility for exercising state power;
- recusal, discipline, removal, and other forms of accountability;
- genuinely reviewable reasons;
- authority to make contested legal and moral judgments.
An AI-generated explanation may also fail to reveal what actually drove its output. That compromises appeals and makes hidden errors, bias, manipulation, and unusual-case failures especially dangerous.
Appropriate uses
AI can reasonably support:
- legal research and citation retrieval;
- source-linked record organization;
- scheduling and docket administration;
- drafting based on reasoning the judge has already formed;…