Intelligent document processing
Thirty years of paperwork, searchable in weeks.
Your PDFs, Word files and spreadsheets become one database anyone in the office can query, held in standard Postgres in your own account.
Built with Claude, Supabase, Next.js, Vercel, Cloudflare.
The situation you are probably in
The information is already yours. It is in folders of proposals, invoices, order forms and function sheets, in three file formats and four naming conventions. Answering a simple question means opening files one at a time, so the questions that would take an afternoon never get asked.
What we do about it
We read the documents and extract the fields into a database designed around the questions you need answered. The schema is the control: if there is no field for a passport number, the model has nowhere to write one down. Every extraction is checked against the source, anything uncertain is flagged to a person rather than filed quietly, and the result lands in standard Postgres you own.
Typical engagement. Three to six weeks for the first archive, depending on how many document layouts there are. Extraction of new documents as they arrive runs as a monthly line after that.
Where this already runs
Built for a long-established catering business: 719 documents, being 388 PDFs, 331 Word files and 25 spreadsheets covering thirty years of trading, turned into customers, events and a priced menu of 1,312 dishes the business now quotes from live. A plain-English question box answers from the same database, and incoming proposals are extracted as they arrive.
Most builds of this shape start with the two-week assessment: we map the process as it runs today, cost it, and write the scope. The fee is credited in full against the build.
Start with the two-week assessmentWhat you receive.
Concrete things, not outcomes on a slide. This list is the scope of the engagement.
- 1
Your document archive turned into one queryable database, in standard Postgres held in your account
- 2
A schema built around the questions you actually ask, and around the fields you have decided not to store
- 3
A confidence check on every extraction, with anything uncertain flagged to a person instead of filed quietly
- 4
A plain-English question box over the result, so the team asks rather than requests a report
- 5
The raw data exported to you, so the records are yours
- 6
A written runbook for putting the next batch of documents through
A good fit when
- The archive runs to hundreds of files or more
- The same fields appear in most of them, even in different layouts
- There is a question the business would ask weekly if the answer were quick
The right starting point when
New documents keep arriving. Extraction runs on a schedule once the archive is done, so incoming paperwork lands in the same database. We build that agent alongside the first pass.
The result has to reach a board pack every month. Add reporting automation on top. The report assembles itself from the same database, with every figure computed in code.
Two teams mean different things by the same field. The two-week assessment settles one definition per field before anything is extracted. The fee is credited in full against the build.
Not sure which of these is you? Describe the process and we will point you at the right one, on this page or another.
Questions buyers ask about this
Ask a different oneWhat documents can intelligent document processing read?
PDFs, Word files and spreadsheets, in whatever layouts and naming conventions your archive has grown into. The catering archive we processed was 388 PDFs, 331 Word files and 25 spreadsheets covering thirty years of trading.
How long does it take?
Three to six weeks for the first archive, depending on how many different document layouts there are. After that, new documents are extracted as they arrive, as a monthly line.
How do you know the extracted data is right?
Every extraction is checked against the source document, and anything uncertain is flagged to a person rather than filed quietly. The database is designed around the fields you need, so there is nowhere to store a field you decided not to keep.
Who owns the data afterwards?
You do. The result lands in standard Postgres held in your own account, and the raw data is exported to you.
How is it priced?
At a fixed fee, scoped in writing before work starts. If the archive or the definitions are not settled yet, the two-week assessment settles them first, and its fee is credited in full against the build.
Price it before you commission it
Put the hours the task takes today, the salary of the person doing it and any quote you hold through the payback calculator. It shows the yearly saving, the payback period and the arithmetic behind both. Free, and nothing is stored.
Run the payback numbersThings you can check.
Figures from our own systems that this service draws on. Each one names its source.
- 719
- documents from thirty years of trading, searchable in weeks
- 1,312
- prices a customer can quote themselves, at any hour
- 120+
- tables of business data you could export tomorrow
Every figure names its source. How we count. The investing system runs on a broker paper account, not real money.
Services
More in research and intelligence
Discuss a business process
Thirty minutes on the process you want to improve. You will know what the right solution looks like, what it would take to build, and what it should return.
Replies come from a named person in Dubai within one working day.