9 ms·
IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers u
by codingdave 7mo ago
IANAL, but this seems like an odd test to me. Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-and-white thinking.
Digging a bit deeper, the actual paper seems to agree: "For the sake of consistency, we define an “error” in the same way that Klerman and Spamann do in their original paper: a departure from the law. Such departures, however, may not always reflect true lawlessness. In particular, when the applicable doctrine is a standard, judges may be exercising the discretion the standard affords to reach a decision different from what a surface-level reading of the doctrine would suggest"
- latchkey 7mo agoIn 30 seconds, did the entire corpus of all the legal cases since the dawn of time agree with the judges opinion on my case? For the state of things in AI today, I'll take it as a great second opinion.
- doctorpangloss 7mo agothe LLMs are phenomenal judges, i am surprised people are skeptical of this result. their training regime is really similar to what a judge does. the reason people are talking about this is because they want AI LAWYERS, which is different than AI JUDGES.
- gowld 7mo agoA mistake isn't "judgment". These were technical rulings on matters of jurisdiction, not subjective judgments on fairness. "The consistency in legal compliance from GPT, irrespective of the selected forum, differs significantly from judges, who were more likely to follow the law under the rule than the standard (though not at a statistically significant level). The judges’ behavior in this experiment is consistent with the conventional wisdom that judges are generally more restrained by rules than they are by standards. Even when judges benefit from rules, however, they make errors while GPT does not.
- swalsh 7mo agoI believed that too until I watched the Karen Read Trials. The judge had a bias, and it was clear karen got justice despite the judge trying to put her finger on the scale.
- droidjj 7mo agoWhether it’s reassuring depends on your judicial philosophy, which is partly why this is so interesting.
- tylervigen 7mo agoYes, your view is commonly called "legal realism."
- scottLobster 7mo agoYeah, I'm reminded of the various child porn cases where the "perpetrator" is a stupid teenager who took nude pics of themselves and sent them to their boy/girlfriend. Many of those cases have been struck down by judges because the letter of the law creates a non-sequitur where the teenager is somehow a felon child predator who solely preyed on themselves, and sending them to jail and forcing them to sign up for a sex offender registry would just ruin their lives while protecting nobody and wasting the state's resources. I don't trust AI in its current form to make that sort of distinction. And sure you can say the laws should be written better, but so long as the laws are written by humans that will simply not be the case.
- rco8786 7mo agoI don't know if I'm comfortable with any of this at all, but seems like having AI do "front line" judgments with a thinner appeals layer available powered by human judges would catch those edge cases pretty well.
- gambiting 7mo agoTo get to an appeal means you obviously already have a judgement against you - and as you can imagine in the cases like the one above that's enough to ruin your life completely and forever, even if you win on appeal.
- jagged-chisel 7mo agoI don't follow your reasoning at all. Without a specific input stating that you can't be your own victim, how would the AI catch this? In what cases does that specific input even make sense? Attempted suicide removes one's own autonomy in the eyes of the law in many ways in our world - would the aforementioned specific input negate appropriate decisions about said autonomy? I don't see how an AI / LLM can cope with this correctly.
- deleted 7mo ago[deleted]
- conradev 7mo ago
- qwertox 7mo ago> If LLMs give only one answer, no matter what nuances are at play, that sounds like they are failing to judge and instead are diminishing the thought process down to black-and-white thinking. You can have a team of agents exchange views and maybe the protocol would even allow for settling the cases automatically. The more agents you have, the higher the nuances.
- jagged-chisel 7mo agoPresumably all these agents would have been trained on different data, with different viewpoints? Otherwise, what makes them different enough from each other that such a "conversation" would matter?
- qwertox 7mo agoDifferent skills or plugins, different views and different tools for the analysis of the same object. Then the debate starts.
- viraptor 7mo agoThen you'd need to provide them with access to the law, previous cases, to the news, to various data sources. And you'd have to decide how much each of those sources of information matter. And at that point, you've got people making the decision again instead of the ai in practice. And then there's the question of the model used. Turns out I've got preferences for which model I'd rather be judged by, and it's not Grok for example...
- deepsun 7mo agoThe main job of a judicial system is to appear just to people. As long as people think it's just -- everyone is happy. But if it's strictly by the law, but people consider it's unjust -- revolutions happen. In both cases, lawmakers must adapt the law to reflect what people think is "just". That's why there are jury duty in some countries -- to involve people to the ruling, so they see it's just.
- toolslive 7mo agoBeing just (as in the right thing happened) and being legal (as in the judicial system does not object) are 2 totally different things. They overlap, but less than people would like to believe.
- rootusrootus 7mo ago> The main job of a judicial system is to appear just to people. Agree 100%. This is also the only form of argument in favor of capital punishment that has ever made me stop and think about my stance. I.e. we have capital punishment because without it we may get vigilante justice that is much worse. Now, whether that's how it would actually play out is a different discussion, but it did make me stop and think for a moment about the purpose of a justice system.
- andyferris 7mo agoI’ve never heard of vigilante justice against someone already sentenced to prison for life, just because they were sentenced in a place without capital punishment? (I mean - people get killed in prison sometimes, I suppose, but it’s not really like vigilante justice on the streets is causing a breakdown in society in Australia, say…)
- shiroiuma 7mo agoIt's probably rather difficult and risky to enact vigilante justice against someone who's in prison. I think the problem is with places where they don't have life sentences at all, but rather let murderers back out into society after some time. I don't know if vigilante justice is a problem there in reality, but at least I can see it as a possibility: someone might still be angry that you murdered their relative after 20 years and come kill you when you're released.
- vjulian 7mo agoThe legal system leaves much to be desired in relation to fairness and equity. I’d much prefer a multi-staged approach with an 1) AI analysis, 2) judge review with high bar for analysis if in disagreement with the AI, 3) public availability of the deliberations, 4) an appeals process.
- jagged-chisel 7mo agoEven having a ready-made determination by an AI runs the risk of prejudicing judges and juries.
- arctic-true 7mo ago“Ladies and gentlemen of the jury, I actually asked ChatGPT and it said my client is not guilty.”
- lemming 7mo agoGiven TFA, it seems that having human determinations involved might run the risk of prejudicing the AI.
- deleted 7mo ago[deleted]
- qotgalaxy 7mo ago[dead]
- fluidcruft 7mo agoThere are findings of fact (what happened, context) and findings of law (what does the law mean given the facts). I don't think inconsistentcy in findings of law is acceptable, really. If laws are bad fix the laws or have precident applied uniformly rather than have individual random judges invent new laws from the bench. Sentencing is a different thing.
- deleted 7mo ago[deleted]
- Nursie 7mo agoLeeway for human interpretation of laws is not a bug, it's a feature. It doesn't make things bad laws. This was the whole problem with the ludicrous "code is law!" movement a handful of years ago. No, it's not, law is made for people, life is imprecise and fairness and decency are not easy to encode.
- vidarh 7mo agoReally, there are three parts to a judgement: facts, the law, and the application of them. There should be no leeway in determining what the law says about a given situation. If that is not decidable, it is a bug. However, what a fair judgement is given the facts and the law, is really a separate issue. You can introduce measures to give clear guidance what the law says, and still give judges flexibility. One of the upsides of "code is law" in that respect is being able to provide a clear statement of what the law says and require the judge to then explain in their judgement why that justifies or does not justify a given judgement. A lot of bad judgement might be a lot more blatant (or not happen) if the judge had to justify outright ignoring the law.
- Nursie 7mo ago'The law' is open to judicial and legal interpretation. There isn't always a single 'the law' to interpret in complex cases. While there are many, many rules, they are not as simple as code and they rely on deep layers of precedent. Common law is made up of case history more than statute. > One of the upsides of "code is law" in that respect is being able to provide a clear statement of what the law says No, "code is law" in fact always ignored what any actual law said, in favour of framing everything as a sort of contract, regardless of whether said contract was actually fair or legal, and it removed the human factor from the whole equation. It was a basic failure to understand law.
- bawolff 7mo ago> Judges do what their name implies - make judgment calls. I find it re-assuring that judges get different answers under different scenarios, because it means they are listening and making judgment calls. I disagree - law should be the same for everyone. Yes sometimes crimes have mitigating curcumstances and those should be taken into account. However that seems like a separate question of what is and is not illegal.
- NoahZuniga 7mo agoThe thing is, Laws do not forsee in all cases, and language is not completely objective, so you cannot avoid judgement calls. One example is computer hacking, which in many jurisdictions is specified in very vague terms.
- NoahZuniga 7mo agoAnother example is that in the Netherlands, there's a crime called "valsheid in geschriften" which exists to make it easy to prosecute fraud. It states that if you create a document with false information with the intent to use that document to deceive, you can get up to 5 years of jail time or some really big fine. Is lying on a paper insurance form to get a cheaper premium breaking this law? This doesn't seem clear cut to me.
- deleted 7mo ago[deleted]
- thaumasiotes 7mo ago> This doesn't seem clear cut to me. ...why not? By your wording, that would be one of the clearest-cut legal cases you could imagine.
- cucumber3732842 7mo agoThe law is rife with words and phrasing that make legality dependent upon those subjective mitigating factors.
- 7mo ago
- homeonthemtn 7mo agoI don't think a lot of people understand the grueling nature of a judge. Day in and out of cases over years are going to generate bias in the judge in one form or another. I wouldn't mind an AI check* to help them check that bias *A magically thorough, secure, and well tested AI
- snitty 7mo agoSo here the test was effectively given a set of relevant facts, can we influence the way a judge (or LLM) rules based on superfluous facts. The judges were either confused or swayed by the superfluous facts. The LLM was not. The matter was one where the outcome should have been determinative, not judgment-based, under US law.
- godelski 7mo agoIANAL. One thing I like to say is There is no rule that can be written so precisely that there are no exceptions, including this one. A joke[0], but one I think people should take seriously. Law would be easy if it weren't for all the edge cases. Most of the things in the world would be easy if it weren't for all the edge cases[1]. This can be seen just by contemplating whatever domain you feel you have achieved mastery over and have worked with for years. You likely don't actually feel you have achieved mastery because you're developed to the point where you know there is so much you don't know[2]. The reason I wouldn't want an LLM judge (or any algorithmic judge) is the same reason I despise bureaucracy. Bureaucracy fucks everything up because it makes the naive assumption that you can figure everything out from a spreadsheet. It is the equivalent of trying to plan a city from the view out of an airplane window. The perspective has some utility, but it is also disconnected from reality. I'd also say that this feature of the world is part of what created us and made us the way we are. Humans are so successful because of our adaptability. If this wasn't a useful feature we'd have become far more robotic because it would be a much easier thing for biology to optimize. So when people say bureaucracies are dehumanizing, I take it quite literally. There's utility to it, but its utility leads to its overuse and the bias is clear that it is much harder to "de"-implement something than to implement it. We should strongly consider that bias in society when making large decisions like implementing algorithmic judges. I'm sure they can be helpful in the courtroom, but to abdicate our judgements to them only results in a dehumanized justice system. There are multiple literal interpretations of that claim too. [0] You didn't look at my name, did you? [1] https://news.ycombinator.com/item?id=43087779 https://news.ycombinator.com/item?id=43087779 [2] Hell, I have a PhD and I forget I'm an expert in my domain because there's just so much I don't know I continue to feel pretty dumb (which is also a driving force to continue learning).
- raw_anon_1111 7mo agoYou have a lot more faith in judges not being biased than I do. I’m about to say something that really makes me throw up a little in my mouth because it harkens back to the forced banal DEI training I had to suffer through in 2020 at BigTech [1]… But judges have all sorts of biases both conscious and unconscious. Where little Jacob will get in trouble for mischief and little Jerome will do the same thing and Jacob is just “a kid being a kid”. But little Jerome is “a thug in training who we need to protect society from”. [1] yes I’m well aware that biases exist. Not only did my still living parents grow up in the Jim Crow South. We had a house built in an infamous what was a “sundown town” as recently as 1990. We have seen how quickly the BS corporate concern was just marketing when it was convenient.
- ralusek 7mo agoDisagree completely. Judgement of the sort you're describing should be done at the legislative phase (i.e. writing code). Inconsistent execution/application of the law is how bias happens. If a judgement done to the letter of the law feels unjust to you, change the letter of the law.
- 6LLvveMx2koXfwn 7mo ago> I find it re-assuring that judges get different answers under different scenarios Unfortunately, as the aptly titled 'Noise' [1] demonstrated o so clearly, judges tend to make different judgement calls in the same scenarios at different times. 1. Noise - https://en.wikipedia.org/wiki/Noise:_A_Flaw_in_Human_Judgment https://en.wikipedia.org/wiki/Noise:_A_Flaw_in_Human_Judgmen...
- vidarh 7mo agoEven in that case, if these systems can be proven to be good enough, rules that require them to be consulted, and for the judge to justify the deviation (if any) from the automated reasoning, might be good. To draw a parallel to a real system, in Norway a lot of cases are heard by panels of judges that include a majority (2 or 3 usually) lay judges and a minority (1 or 2 usually) of professional judges. The lay judges are people without legal training that effective function like a "mini jury", but unlike in a jury trial the lay judges deliberate with the professional judges. The professional judges in this system has the power to override if the lay judges are blatantly ignoring the law, but this is generally considered a last resort. That power requires the lay judges to justify themselves if they intend on making a call the professional judges disagree with. Despite that, it is not unusual for the lay judges to come to a judgement that is different from what the professional judges do, and fairly rare for their choices to be overridden. The end result is somewhere in the middle between a jury and "just" a judge. If proven - with far more extensive testing - that its reasoning is good enough, an LLM could serve a similar function of providing the assessment of what the law says about the specific case, and leave to humans to determine if and why a deviation is justified.
- catlover76 7mo ago[dead]