The Medical-Legal Summary Report Card: How to Grade Any AI Summary in 30 Minutes
“AI summaries” are suddenly everywhere in medical-legal technology.
And if you look at vendor websites, they can all start to sound pretty similar:
Accurate. Comprehensive. AI-powered. Save hours reviewing records.
But anyone who has actually compared outputs knows that two “medical summaries” can be wildly different products.
One might capture every detail but bury you in noise.
Another might be beautifully concise but leave out the prior injury that changes the case.
A third might surface exactly the right facts but give you no practical way to verify where they came from.
That is because there is no single definition of a good medical summary.
A good summary depends on what you are trying to do with it.
An adjuster evaluating exposure needs something different from an attorney preparing for a deposition. An MSA reviewer cares about different details than a life care planner. A nurse reviewing causation may want more clinical depth than someone trying to understand a claim in five minutes.
So instead of asking:
“Is this a good summary?”
Ask:
“Is this summary trustworthy, and is it good for the job I need it to do?”
Here is a simple way to find out.
Before You Grade: Name the Job
Take one real file and finish this sentence:
I need this summary to help me ______________________.
For example:
- understand the medical story quickly
- evaluate causation
- prepare for an IME
- identify prior injuries
- evaluate a settlement
- draft an MSA
- prepare for a deposition
- understand future treatment
- audit a claim
- build a life care plan
Write down the answer.
This matters because the same output can be excellent for one job and terrible for another.
A 40-page chronology might be perfect for litigation preparation and useless for an executive claim review.
A two-page narrative might be perfect for orientation and dangerously thin for an MSA.
The job comes first.
Now grade the output.
Part I: Can You Trust It?
These are the non-negotiables.
A summary that fails here may look impressive, but it should not be carrying important decisions.
Test 1: Completeness
Did the system actually understand the whole file?
- Pick an important event you already know exists. It appears in the output.
- Find one ugly source page: handwriting, scan artifacts, rotated pages, or poor formatting. Its important information still made it through.
- Missing or unavailable records are identified instead of silently ignored.
Ask the vendor:
“How do I know which parts of the file your system could not confidently process?”
A summary cannot tell you the whole story if parts of the story never made it into the system.
Red flag: everything looks wonderfully complete, but the tool gives you no way to know what it missed.
Test 2: Accuracy
Are the facts actually right?
Pick five facts at random and check the source.
- Dates, diagnoses, procedures, medications, and other clinical facts match the record.
- Numbers such as dosages, costs, measurements, and dates survive intact.
- You cannot find statements that appear to have been invented, inferred, or overstated.
Ask the vendor:
“What happens when the system isn’t sure?”
Good systems preserve uncertainty.
Bad systems turn uncertainty into confident prose.
Red flag: a statement sounds plausible but cannot actually be found in the record.
Test 3: Traceability
Can you prove where a fact came from?
- Important facts include page, Bates, document, or source references.
- You can get from an important statement to its source in seconds.
- Citations remain usable after export.
Ask the vendor:
“Pick any sentence. Show me the evidence behind it.”
This is one of the biggest differences between a summary that is pleasant to read and one that can actually support professional work.
Red flag: verifying the summary means reopening the PDF and searching manually.
Part II: Is It Useful?
Passing the trust tests does not automatically make a summary good.
A 100% accurate summary can still be miserable to use.
This is where vendors start to separate.
Test 4: Relevance
Did it choose the right information?
Go back to the job you wrote down at the beginning.
- The facts most relevant to that job are easy to find.
- Administrative clutter and low-value details do not dominate the output.
- The summary emphasizes important facts rather than treating every sentence in the record equally.
Ask the vendor:
“How does your system decide what deserves attention?”
Summarization is fundamentally an exercise in judgment about what matters.
Simply shortening the source material is not enough.
Red flag: the output feels like a compressed table of contents.
Test 5: Signal-to-Noise
How much reading did the summary actually eliminate?
- Duplicate records do not create duplicate findings.
- Repeated facts are consolidated where appropriate.
- The output is meaningfully easier to consume than the source file.
Try this simple test:
Open the summary and scroll.
How much of what you see is genuinely helping you understand the case?
A 900-page file does not automatically deserve a 90-page summary.
Red flag: the vendor measures quality by how much text it generated.
Test 6: Structure
Does the output make the story easier to understand?
Depending on the job, that might mean chronology, episodes of care, diagnoses, providers, treatment categories, or something else.
- Events and facts are organized in a way that matches how you think about the case.
- Relationships between events are easier to see than they were in the raw record.
- Date conflicts, uncertain sequencing, or undated events are made explicit.
Ask the vendor:
“Why is the information organized this way?”
A good structure reduces cognitive work.
A bad structure simply moves information from one document into another.
Red flag: you still have to reconstruct the medical story yourself after reading the summary.
Test 7: Issue Spotting
Does the system help you notice things you might otherwise miss?
This is where summaries can become much more valuable than document compression.
Look for things like:
- prior injuries
- gaps in treatment
- conflicting histories
- inconsistent dates
- changes in diagnosis
- failed conservative treatment
- missing records
- unusual utilization patterns
- changes in work status
- competing causation narratives
Then score it:
- The output surfaces important patterns across multiple documents.
- Contradictions are shown with the underlying evidence.
- The system surfaces the issue without pretending to make the professional judgment for you.
Ask the vendor:
“What can your system find across the record that I might miss reading one document at a time?”
The best tools do more than summarize documents.
They help connect them.
Red flag: the AI confidently resolves ambiguity instead of showing it to you.
Test 8: Fit for Purpose
Now return to your original sentence:
I need this summary to help me ______________________.
Did it?
- The output contains the information required for that workflow.
- The format matches how your team actually works.
- You can use the output without substantial rewriting, reformatting, or re-review.
If your organization has specific templates, terminology, fields, or SOPs, test those too.
Ask the vendor:
“Can you make the output work like our process, or do we need to change our process to work like your software?”
There is no universal perfect summary.
There is only a summary that is more or less suited to the decision you are trying to make.
Red flag: you spend the time the software supposedly saved turning its output into something usable.
Add Up Your Score
| Test | Score |
|---|---|
| 1. Completeness | ___ / 3 |
| 2. Accuracy | ___ / 3 |
| 3. Traceability | ___ / 3 |
| 4. Relevance | ___ / 3 |
| 5. Signal-to-Noise | ___ / 3 |
| 6. Structure | ___ / 3 |
| 7. Issue Spotting | ___ / 3 |
| 8. Fit for Purpose | ___ / 3 |
| Total | ___ / 24 |
22–24: A
Excellent. The output is trustworthy and genuinely useful for the job you gave it.
18–21: B
Strong. Figure out whether the missed points matter for your particular workflow.
14–17: C
Useful with supervision. Expect meaningful review or rework.
Under 14: F
You may be creating a second thing to review instead of eliminating review.
One Important Catch
There is one grading rule we would not compromise on:
Trust comes before convenience.
If the output performs poorly on completeness, accuracy, or traceability, a beautiful format cannot rescue it.
A concise hallucination is still a hallucination.
A perfectly formatted chronology with missing records is still incomplete.
And a brilliant observation you cannot trace back to the source may be impossible to rely on when it matters.
So if a vendor misses more than two of the nine checks across Tests 1–3, cap the grade at a C.
Everything after that is optimization.
The first three are permission to trust the output at all.
The Best Way to Compare Vendors
Do not compare demo files.
Take one ugly, representative file from your own work and give the same file to every vendor.
Then use the same job:
“I need this summary to help me evaluate causation.”
Or:
“I need this chronology to prepare an attorney for mediation.”
Or:
“I need this review to determine whether we have all the records necessary for an MSA.”
Now compare the outputs.
Same file.
Same job.
Same rubric.
That will tell you dramatically more than any polished product demo.
And Yes, You Can Grade Us Too
We built this rubric because these details matter enormously to us at InQuery.
We do not think “AI summary” is a useful enough description of a product.
What matters is whether the output is complete enough to trust, precise enough to verify, structured enough to understand, and tailored enough to actually help someone make a decision.
So try the rubric on us.
Give InQuery a real file, ideally a messy one, and grade the result using exactly the same tests.
Grading us will not cost you anything, either. We are happy to process your first 3,000 pages for free, which is enough to run this entire rubric on a real file before a contract ever enters the conversation.
Because the easiest way to evaluate an AI summary is not to ask the vendor how good it is.
Give it a hard file and check its work.
Erick Enriquez
CEO & Co-Founder at InQuery