MILO Invoice Intelligence

Evaluation

Every number below is computed by scripts/eval.ts against a hidden ground_truth table — 93 synthetic invoices spanning 17 designed scenarios, each with a predetermined correct decision. The decision engine (src/engine/**) never reads that table; the eval script does, purely to score the run. Nothing here is hand-typed.

Because the recommendation engine is deterministic, rule-based code — not an LLM making judgment calls — this evaluation measures whether the rubric was implemented correctly against its own design, not probabilistic accuracy under uncertainty. That is the intended shape of this system: the parts that decide are code, and code is either right or has a bug.

Dataset size
93
synthetic invoices, 17 archetypes
Decision accuracy (exact)
100.0%
matches the single correct recommendation
Decision accuracy (acceptable)
100.0%
matches an acceptable recommendation

Safety metrics

False-approval rate
0.0%

Of 51 invoices that should not have been approved, MILO incorrectly approved 0. This is the single most important number on this page — an incorrectly approved invalid invoice is far more costly than one sent for unnecessary human review.

False-escalation rate
0.0%

Of 75 invoices that did not need escalation, MILO escalated 0 anyway — the over-triggering cost of concentrating human attention too aggressively.

Precision / recall by recommendation class

ClassSupportPrecisionRecallF1
Approve42100.0%100.0%100.0%
Review20100.0%100.0%100.0%
Request Information13100.0%100.0%100.0%
Escalate18100.0%100.0%100.0%

Abstention — knowing what it doesn’t know

Abstention rate (Request Information)14.0%
Abstained on13 invoices
Invoices genuinely missing critical evidence13
Abstention recall100.0%

Evidence quality

Mean evidence completeness94.5%

Average share of applicable checks that resolved to a definitive PASS/FAIL/RISK rather than UNKNOWN or MISSING, across the full book.

Confidence calibration

BandCountMean scoreAccuracy within band
High5098.0100.0%
Medium4371.3100.0%
Low0——

0 High-confidence recommendations were wrong. 0 Request Information recommendations carried High confidence — confidence.ts caps this band structurally, so it should always read 0.

Confusion matrix

Actual \ PredictedAPPROVEREVIEWREQUEST INFORMATIONESCALATE
APPROVE42000
REVIEW02000
REQUEST INFORMATION00130
ESCALATE00018