AI models vs. law graduates: who performed better on the Israel Bar Association exams?

The Israeli startup Lizzy AI tested popular artificial intelligence models on the Israel Bar Association exam. The results show that AI tools can outperform even the top law students.

GeektimeAuthor: Oshri Alexalsi
Source
AI models vs. law graduates: who performed better on the Israel Bar Association exams?
Photo: Geektime / מקור: Unsplash

Preparation for the Bar Exam, the qualifying test for law graduates to become lawyers, is a significant challenge for any Israeli law student. But what happens when you introduce artificial intelligence models to this rigorous process? An Israeli startup decided to test which popular models could pass this demanding exam and how their performance compares to that of human students.

Who is the winner?

As part of a comprehensive study conducted by the Israeli startup Lizzy AI, which specializes in the intersection of AI and law, three of the most popular models were tested: ChatGPT, Gemini, and Claude. Alongside these, three specialized legal AI tools were also evaluated: LawMate, Takdin AI, and Nevo AI. The models were tasked with answering all multiple-choice questions, which account for 80% of the final exam grade.

The results were striking: these AI tools outperformed even the best law faculties.

The clear winner was ChatGPT, with an impressive 94.1% success rate. It was followed by two Israeli tools, LawMate (93.1%) and Nevo AI (92.2%). In fourth place, we finally see human students: graduates of the Hebrew University Faculty of Law, who led the student rankings with 91.82%. Takdin AI took fifth place (85.9%), slightly ahead of Tel Aviv University graduates (85.71%). The remaining rankings included Google's Gemini (84.4%), University of Haifa graduates (82.61%), Bar-Ilan University students (81.38%), the College of Management (81.37%), and finally, Claude (77.5%).

Advanced versions and limitations

The test, conducted in March and April, utilized the most advanced versions available: ChatGPT (GPT-5.4 with Extended Thinking mode), Gemini (3.1 Pro), and Claude (Opus 4.6 in extended thinking mode).

Despite these impressive results, the road to practical legal application remains long. Maya Amir, a product manager at Lizzy AI, explains that a correct answer alone is insufficient for professional use: "ChatGPT provides references to sources upon request, but does not perform quality control on them." According to Amir, the models may rely on blogs or private publications rather than authoritative case law or legislation. Furthermore, the lack of specific citations within the full context of a source makes the verification process cumbersome.

Lizzy AI, founded by Arnon Katlan and Netzer Shlomai, develops a platform for analyzing and drafting legal documents. Used by over 8,000 lawyers and firms in Israel, the startup aims to be the Israeli answer to the American Legal-Tech giant Harvey, which was valued at 11 billion dollars in its latest funding round.

Related News