Document Compliance Engine
A compliance pipeline that reads a contract and a W-9, extracts structured data with an LLM, cross-validates every fact against external systems, and pauses durably for a human when something doesn't line up — evaluated at a 0% hallucination rate across 79 fixtures.
The problem
Vendor onboarding is high-volume and compliance-sensitive: someone reads a contract and a tax form, keys the data in, checks the Tax ID against a government registry, screens against sanctions lists, and catches the case where the name on the tax form doesn't match the contract.
Doing it manually is slow and error-prone. Doing it with a naive LLM pipeline is fast but worse — no audit trail, no way to prove it isn't hallucinating tax IDs, and no resumable path when a human is needed. It's fast right up until it's confidently wrong in a way nobody can catch.
The approach
The agent pipeline is a Reactor module tree running inside the same Phoenix release, not a separate service — Oban's job process is the only async boundary, so webhook handlers never block. Postgres is the shared source of truth for both status and workflow checkpoints.
Instructor extracts documents into Ecto schemas via GPT-4o-mini. Every extracted field is then validated by calling out to two tool servers — a Tax API and a Sanctions DB — that run as genuinely separate OTP applications exposing MCP over HTTP, so the pipeline can't quietly fold "external system" into "same process."
When something's ambiguous, the reactor halts and persists its state to Postgres. A human reviews and decides in a LiveView queue; resuming reads that decision from a downstream step rather than the halt itself, so a decision can't be lost.