AutomateFrom R18 000 · then a few rand per document

Document extraction

Stop retyping invoices and forms. We turn your documents into clean data automatically.

What it is

In plain English

Document extraction is software that reads a document — a PDF invoice, a scanned delivery note, a filled-in form — and pulls out the information you care about as clean, structured data. The invoice number, the supplier, the line items, the totals, the date: captured automatically and dropped straight into your spreadsheet, accounting system or database.

It replaces the job someone currently does by hand: opening each document, reading it, and retyping the same fields into another system. That work is slow, easy to get wrong at the end of a long day, and a poor use of a capable person's time.

When you need it

Is this you?

You need this when your team processes the same kind of document over and over, and the volume has grown past the point where typing it in by hand makes sense.

  • Someone spends hours a week retyping supplier invoices into Xero, Sage or a spreadsheet.
  • Delivery notes, application forms or timesheets pile up and get captured in batches, days late.
  • Typing errors are creeping into your records, and catching them costs more than the typing did.
  • You're hiring — or thinking of hiring — mainly to keep up with data capture.

How we build it

What actually happens

  1. 1

    Map your documents

    You send us 10–20 real examples. We identify every field you need pulled out and exactly how it should look on the other side.

  2. 2

    Build the extractor

    We set up an extraction pipeline that reads each document and returns structured data, tuned against your actual examples rather than a generic template.

  3. 3

    Add a confidence check

    High-confidence results flow straight through. Anything the system is unsure about is flagged for a quick human glance, so a wrong number never slips into your books silently.

  4. 4

    Wire it into your tools

    The clean data lands where you actually work — Google Sheets, Xero, Sage, a database, or via a webhook into whatever you use.

  5. 5

    Test and hand over

    We run your real documents through, check the output together, and give you a simple way to upload new ones and see results.

Example

How it plays out in practice

A Johannesburg plumbing supplies company receives around 80 supplier invoices a month from 6 different suppliers — all with different layouts. Their bookkeeper spends roughly 6 hours a week retyping them into Sage One.

  1. 1

    Invoice arrives

    An email from Pipe & Fittings SA lands in a monitored inbox with a PDF attached. The system picks it up within seconds — no one has to do anything.

  2. 2

    Fields are read

    The extractor pulls out: supplier name (Pipe & Fittings SA), invoice number (INV-7741), date (15 March 2026), three line items with quantities and unit prices, subtotal R2 580, VAT R387, total R2 967.

  3. 3

    Confidence is checked

    Every field scores above 95% — the document is clear, printed, and consistent with previous invoices from this supplier. It passes straight through without flagging anyone.

  4. 4

    Sage is updated automatically

    A new supplier bill is created: correct supplier, correct account codes, correct amounts, INV-7741 as the reference. The bookkeeper’s capture queue is shorter by one.

  5. 5

    Notification sent

    A short message goes to the bookkeeper: “INV-7741 · Pipe & Fittings SA · R2 967.00 · captured.” She opens it, glances at the amounts, clicks confirm in 20 seconds.

80 invoices a month. At 8 minutes each that was 10 + hours of typing. The same 80 invoices now take under 30 minutes to review. The bookkeeper uses the rest of that time on work that actually needs a human.

What we'd need from you

To get started

  • 10–20 sample documents — the real ones you process, not blanks.
  • The list of fields you need extracted (e.g. invoice number, supplier, line items, amounts, dates).
  • Where the output should go — spreadsheet, accounting system, or database.
  • Rough monthly volume — how many documents you process.
  • Which fields are zero-tolerance (must be exact) versus ones you can eyeball.

Timeline

How long it takes

Most document-extraction builds go live in 2–4 weeks from the point we have your sample documents.

What can affect that:

  • Many different document layouts add time — we tune the reader for each one.
  • Handwritten fields are harder than printed text and usually need a review step.
  • A live connection into an accounting system takes longer than a spreadsheet or CSV output.

What affects the price

Why it costs what it costs

Every build is scoped to your situation. These are the things that move the number — so there are no surprises on the call.

Document complexity
A clean, consistent invoice layout is quick. Many different formats, dense tables, or handwriting each take more work to read reliably.
Number of fields
Pulling five fields is simpler than pulling fifty line items with per-row detail.
Accuracy requirements
Fields that must be perfect — amounts going into your books — get a confidence-and-review layer that fields you can eyeball don't need.
Where it connects
Dropping data into a spreadsheet is straightforward. A live connection into Xero, Sage or an ERP is more to build and test.
Volume
The build is once-off; after that you pay a few rand per document to cover the processing. Higher volume means a lower per-document rate.

Ready to build something that works?

Tell us what you need. We'll tell you honestly if and how we can help.

Book a discovery call