⚠️ Testing prototype — not legal advice. These scores measure how closely a model's contract review matches what a lawyer already flagged on a fixed set of sample documents. They say nothing about performance on a real, unseen contract.

Model Evaluation

How each hosted model's contract review compares against a lawyer's own flags on the same documents.

Direct model calls (baseline)

Phase 1: each model reviews a document on its own, no retrieval involved — the same call Legal Assistant makes. Later phases will add their own sections here: RAG-augmented, then RAG-only strict mode.

Loading…

Re-runs the baseline evaluation against the lawyer-labeled document set. Admin sign-in required.