Reserve the LLM for ambiguity
Resolve stable cases with deterministic rules and route only the remaining 5%. The tiered architecture reduced latency, token consumption, and audit risk.
Alternative considered: Send all 50K+ records through an LLM or keep the workflow fully manual.
Trade-off: Explicit rules required calibration and maintenance, but cut token consumption by 90%.