Intelligent document processing

Thirty years of paperwork, searchable in weeks.

Your PDFs, Word files and spreadsheets become one database anyone in the office can query, held in standard Postgres in your own account.

The situation you are probably in

The information is already yours. It is in folders of proposals, invoices, order forms and function sheets, in three file formats and four naming conventions. Answering a simple question means opening files one at a time, so the questions that would take an afternoon never get asked.

What we do about it

We read the documents and extract the fields into a database designed around the questions you need answered. The schema is the control: if there is no field for a passport number, the model has nowhere to write one down. Every extraction is checked against the source, anything uncertain is flagged to a person rather than filed quietly, and the result lands in standard Postgres you own.

Typical engagement. Three to six weeks for the first archive, depending on how many document layouts there are. Extraction of new documents as they arrive runs as a monthly line after that.

Where this already runs

Built for a long-established catering business: 719 documents, being 388 PDFs, 331 Word files and 25 spreadsheets covering thirty years of trading, turned into customers, events and a priced menu of 1,312 dishes the business now quotes from live. A plain-English question box answers from the same database, and incoming proposals are extracted as they arrive.

Most builds of this shape start with the two-week assessment: we map the process as it runs today, cost it, and write the scope. The fee is credited in full against the build.

Start with the two-week assessment

What you receive.

Concrete things, not outcomes on a slide. This list is the scope of the engagement.

  1. 1

    Your document archive turned into one queryable database, in standard Postgres held in your account

  2. 2

    A schema built around the questions you actually ask, and around the fields you have decided not to store

  3. 3

    A confidence check on every extraction, with anything uncertain flagged to a person instead of filed quietly

  4. 4

    A plain-English question box over the result, so the team asks rather than requests a report

  5. 5

    The raw data exported to you, so the records are yours

  6. 6

    A written runbook for putting the next batch of documents through

A good fit when

  • The archive runs to hundreds of files or more
  • The same fields appear in most of them, even in different layouts
  • There is a question the business would ask weekly if the answer were quick

The right starting point when

Not sure which of these is you? Describe the process and we will point you at the right one, on this page or another.

Price it before you commission it

Put the hours the task takes today, the salary of the person doing it and any quote you hold through the payback calculator. It shows the yearly saving, the payback period and the arithmetic behind both. Free, and nothing is stored.

Run the payback numbers

Things you can check.

Figures from our own systems that this service draws on. Each one names its source.

719
documents from thirty years of trading, searchable in weeks
1,312
prices a customer can quote themselves, at any hour
120+
tables of business data you could export tomorrow

Every figure names its source. How we count. The investing system runs on a broker paper account, not real money.

Make the next AI project one the business can measure.

Thirty minutes with the founder, no slides. You will know what the right solution looks like, what it would take to build, what it should return, and which part to start with.

Replies come from a named person in Dubai within one working day.