
Stop Asking If Legal AI Is Accurate. Start Asking If It’s Reproducible.
Run the same contract through the same AI system 20 times. Should a lawyer expect the same legal findings 20 times? In a reproducibility stress test from my research published at ICCS 2026, 15 documents were each analyzed 20 times. The pure large-language-model baseline produced an average of 18.3 distinct output sets per document across those 20 identical re-runs. The deterministic hybrid architecture produced one. That gap points to a question legal teams should be asking AI vendors far more








