Document classification softwarefor unstructured archives.
Scans, PDFs and mixed document dumps only become useful once someone sorts them. MiruIQ classifies against document types you define, applies the right schema to each, and routes every file to the right destination automatically.
Turn document dumps into routed, structured archives: classify against your own types, extract what each type needs and send every file where it belongs.
Before automation starts, someone sorts.
Most archives are piles: scanned mail, exported mailboxes, decades of PDFs. Every automation project stalls at the same step, because humans must first recognize, label and route each file, slowly and inconsistently.
Misfiled documents are the silent cost: the contract in the invoice folder is effectively lost. Unstructured data extraction only pays off when classification and routing come first.
Archive routing, step by step.
From dump to destination without a sorting team.
1. Ingest
Point a pipeline at the archive: S3 bucket, drive, SFTP or a one-time bulk upload.
2. Classify
Each file is classified against your document types using a classifier chain: filename rules, content checks, semantic conditions.
3. Extract
The matched type's schema extracts what that type needs: a contract yields parties and dates, an invoice yields totals.
4. Verify
Files the chain cannot place confidently park in the review queue instead of landing in the wrong folder.
5. Route
Every document flows to the right destination per type: DMS, database, S3 prefix or webhook.
What changes for the archive.
Sorting cost drops
The manual recognize-label-route step disappears for everything the chain places confidently.
Fewer misfiles
Classification is consistent by construction, and the uncertain cases go to a human instead of a guess.
Automation coverage grows
Once routing exists, every downstream workflow (extraction, validation, delivery) can start from clean inputs.
Numbers to hold the system to.
Routing quality is measurable from day one.
Auto-routing rate
Share of files routed without any human action.
Misroute rate
Share of files sent to the wrong destination, measured by sampling.
Review load
Sorting decisions a human still makes per day; it should shrink as the chain is tuned.
Archive routing questions
What teams with document backlogs ask most.
You define the document types, and a classifier chain checks each file against them in order: filename rules first, then content-based checks, then semantic conditions evaluated by an LLM. Types are yours, not a fixed catalog, so a niche document class you invented last week classifies as reliably as invoices do.
They park in the human review queue instead of being guessed into a folder. A reviewer assigns the type in one click, and persistent gaps tell you which classifier rule or document type definition to sharpen. Nothing is silently misfiled.
Yes, with the same configuration: a batch pipeline works through the historical archive, and a streaming pipeline watches the source for new arrivals. Both use the same types, schemas and routing rules, so the archive and the inbox stay consistent.
See it on your own documents.
The 14-day free trial covers 200 documents: enough to define your schemas, run real files and measure the results before anyone calls you.
