# Invoice and document intake

**Demo build by Kedgework. Not client work. Synthetic data.**

Supplier invoices arrive by email as PDFs. Someone opens each one, types the supplier, number, date and totals into a spreadsheet or the accounting system, and hopes nothing was typed twice or added up wrong. This workflow does that typing, checks the numbers with fixed rules, and only asks a person when something looks wrong. Every morning at 08:00 it sends one line: what it did, what is waiting for a person, and any sign it can see that something stopped working.

It runs in [n8n](https://n8n.io), in your own n8n account, and writes to a Google Sheet you own.

![How it works](diagram.png)

## Step by step

1. **An email arrives** in the mailbox you use for invoices. Only unread mail is picked up, and each email is marked as read once it has been taken.
2. **Each PDF attachment becomes one job.** An email with no PDF is not ignored: it goes to the review queue so a person sees it.
3. **Claude (the AI model by Anthropic) reads the PDF** and copies the fields into a fixed form: supplier, VAT number, invoice number, dates, currency, each line, net total, VAT and total. It is told to copy what is printed and never to correct it, so a wrong total on the invoice stays wrong and gets caught in the next step. Invoices in Portuguese, Spanish, French, German, Dutch and English are in the sample set. PDFs go to Claude one at a time; if a call fails (network, overload), that PDF is tried up to 3 times in total.
4. **Fixed rules check the result** (plain code, not AI, so they behave the same every time):

   | Check | What it catches |
   |---|---|
   | The lines add up to the net total | A line left out or misread |
   | Net plus VAT equals the total due | A wrong total on the invoice |
   | VAT matches the rates on the lines | VAT charged at the wrong rate |
   | Supplier VAT number has a valid format, including the check digit for Portugal and France | A mistyped or invented number |
   | Same supplier VAT and invoice number not already saved | The same invoice sent twice, or a reminder with a copy |
   | Invoice date is a real date, not in the future, not more than a year old; due date after invoice date | Misread dates, old invoices sent again |
   | Currency is on your list | An invoice in a currency you do not pay in |
   | It is an invoice, not a credit note, a contract or an order | Other documents sent to the same mailbox. A credit note gets its own reason, so it can be matched with the original invoice |
   | Required fields are present | A missing invoice number or tax number |

5. **Clean invoices are saved** as one row in the "Invoices" sheet, with the sender, subject and file name of the email they came from. Values are stored exactly as read: a subject that starts with "=" stays text and never becomes a spreadsheet formula, and invoice numbers keep their leading zeros.
6. **Anything doubtful goes to the "Review queue" sheet** with the reason in plain words. A person gets one short email per batch listing those documents. When they have dealt with one, they set its status to done.

## What it looks like

From the offline test run on the 10 sample invoices (synthetic data, full files in [`test/sample-output/`](test/sample-output/)).

"Invoices" sheet, 3 of the 7 saved rows (some columns left out here):

| supplier_name | supplier_vat | invoice_number | invoice_date | currency | net_total | vat_total | gross_total | file_name |
|---|---|---|---|---|---|---|---|---|
| Mondego Sample Supplies, Lda | PT532449878 | FT 2026/0141 | 2026-09-14 | EUR | 259.80 | 59.75 | 319.55 | FT-2026-0141.pdf |
| Ejemplo Logistica Demo, S.L. | ESB80114844 | A-2026-0388 | 2026-09-18 | EUR | 516.00 | 108.36 | 624.36 | A-2026-0388.pdf |
| Exemple Conseil Demo SARL | FR88876963025 | F2026-1093 | 2026-09-22 | EUR | 1,420.00 | 284.00 | 1,704.00 | Facture_F2026-1093.pdf |

"Review queue" sheet, one of the 4 rows:

| supplier_name | invoice_number | gross_total | reason_text | reasons | review_status |
|---|---|---|---|---|---|
| Exemple Transport Demo SAS | T-26-00417 | 316.00 | Net total plus VAT does not equal the total due | totals_mismatch:255+51=306 but total is 316 | open |

The email the reviewer gets for that batch:

```
Subject: Invoices: 4 document(s) need a check

These documents were not saved automatically. They are in the "Review queue" sheet.

1. T-26-00417.pdf from Exemple Transport Demo SAS: Net total plus VAT does not equal the total due
2. A-2026-0388-copia.pdf from Ejemplo Logistica Demo, S.L.: This invoice was already processed (same supplier VAT and invoice number)
3. fatura_LDM_56.pdf from Lousa Demo Manutencao, Lda: A required field is missing on the document
4. (no file) from someone@b.example: The email had no PDF attachment

When you have dealt with one, set review_status to done.
```

## What you see every morning

At 08:00 (Lisbon time) the workflow checks itself and sends one line by email. The email goes out first; the same line is then kept in a "Morning log" sheet, and a failure there (for example an expired Google login) cannot stop the email. Three lines produced by the offline test:

```
OK. Last 24 h: 2 invoice(s) saved, 0 sent to review. Open in review: 0. Failed runs: 0. Last document: 2026-10-05. Rules self-test: passed.
ATTENTION. Last 24 h: 3 invoice(s) saved, 1 sent to review. Open in review: 2 (oldest 50 h). Failed runs: 1. Last document: 2026-10-05. Rules self-test: passed. Needs a look: a review item has waited 50 h; 1 failed run(s).
ATTENTION. Last 24 h: 1 invoice(s) saved, 0 sent to review. Open in review: 0. Failed runs: 0. Last document: 2026-10-05. Rules self-test: FAILED. Needs a look: code changed since the tested build in: Check the invoice (run the offline test again).
```

What it checks:

- **The counts:** invoices saved and sent to review in the last 24 hours, review items still open, and failed runs of this workflow.
- **That documents are still arriving.** If no document at all came in for 3 full working days (Monday to Friday), the line says ATTENTION. A mailbox that stopped delivering looks exactly like a quiet day, so silence is treated as a warning, not as good news.
- **That the checking rules are the ones we tested.** It reads the workflow from your n8n and compares the code of each intake step with a fingerprint of the code that passed our offline test. Then it runs the rules on two built-in sample invoices, one clean and one with a wrong total. If a step was edited inside n8n, or the rules no longer catch the wrong total, the line says "Rules self-test: FAILED" and names the step.
- **That it could read what it needs.** If it cannot read a sheet, the run history or the workflow itself, the line says ATTENTION and names what it could not read. The morning run still finishes and still sends its line.

What it cannot see from the inside:

- **If n8n itself is down**, or the schedule was switched off, no line is sent at all. Rule of thumb: if the 08:00 email has not arrived by 08:15, something is down. To be alerted for that too, switch on the "Heartbeat ping (optional)" step and point it at a dead man's switch service (for example Healthchecks.io): such a service alerts you when the daily ping does not arrive.
- **A mailbox that stops for 1 or 2 working days** looks like a quiet day until the third. Public holidays count as working days, so after a long holiday the line may say ATTENTION once.
- **How well Claude reads your invoices.** The self-test proves the rules work, not the reading. Reading problems show up as review items and, in the first live week, in the comparison with the paper.

## Nothing disappears quietly

- No PDF in the email: goes to the review queue.
- The AI step fails (network, overload): that PDF is tried up to 3 times in total, then it goes to the review queue with the reason. The other PDFs of the same email are not sent again.
- Claude declines to read a document, or its answer is cut off: review queue.
- A run of the workflow fails completely: counted in the next morning line.
- The morning check cannot read a sheet or the run history: the line says ATTENTION instead of failing silently.
- The Google login expires: the morning line is still emailed, says ATTENTION and names the sheets it could not read. Only the copy in the "Morning log" sheet is missing that day.

## What was tested, and what was not

**Offline (on this machine, 5 October 2026):** the code inside the workflow, taken straight out of `workflow.json` and run on 10 synthetic invoices plus 18 edge cases and 11 morning check situations. Result: **101 of 101 checks passed**, see [`test/RESULT.txt`](test/RESULT.txt). The 10 invoices are real PDFs in 6 languages; 7 are clean, 3 have a problem on purpose (a wrong total, a duplicate, a missing tax number), and all 3 were sent to review with the right reason.

**The test catches broken rules:** [`test/mutation.mjs`](test/mutation.mjs) breaks the rules on purpose in 3 ways (totals and duplicate checks off in the checking step; totals check off in that step only; totals check off everywhere and the workflow rebuilt) and runs the same test on each broken copy. All 3 fail, as they should (90, 91 and 89 of 100 checks pass on them). See [`test/MUTATION-RESULT.txt`](test/MUTATION-RESULT.txt).

**Inside a real n8n:** [`test/run-in-n8n.mjs`](test/run-in-n8n.mjs) imports `workflow.json` into n8n 2.41.7 and runs copies of it from the command line. Result: **16 of 16 checks passed**, see [`test/N8N-RUN.txt`](test/N8N-RUN.txt). That file carries the sha256 of the `workflow.json` it tested, and the offline test checks that it matches the file shipped here.

- Import and export: the same 28 steps, settings, connections and credential placeholders come back.
- 9 emails with 10 PDFs (two PDFs in one email, a logo next to a PDF in another, plus one email with no PDF) go through the real steps: 7 saved, 3 to review with the right reason, plus the email without a PDF, and one reviewer email listing the 4.
- The real "Claude reads the PDF" step sends each PDF, one at a time, to a stand-in for the Anthropic API running on this machine, with the right headers, model, schema and the key taken from the n8n credential. A PDF that gets "overloaded" once is asked again (only that PDF); one that keeps failing goes to review after 3 tries.
- The morning check runs with the real n8n API step (against a local stand-in), the real email step (to a local mail server) and, in one case, real Google Sheets steps with a login that does not work: the line still goes out, says ATTENTION and names what it could not read.
- Nothing left the machine: a guard inside n8n refuses any connection to another host, and the run fails if it had to refuse one. It refused none.

How it was run, said plainly: the n8n on the test machine is an npx install that was cut off before it finished. A small loader kept outside this folder fills in two missing parts (the sqlite3 database driver and one package folder) without changing any n8n file; a normal install (Docker, n8n Cloud, `npm i -g n8n`) does not need it. The test loader [`test/n8n-test-hook.cjs`](test/n8n-test-hook.cjs) adds the network guard and lets the command line use pinned data, which `n8n execute` otherwise ignores. Google Sheets and the incoming mailbox cannot be pointed at a stand-in, so their data is pinned (fixed test data on the step, an n8n feature), and one sheet write is replaced by a step that passes its rows on. [`test/N8N-RUN.txt`](test/N8N-RUN.txt) lists exactly what ran for real.

**What the n8n run found, and what changed.** The same run on the previous build failed 5 of 16 checks, for two reasons:

1. If the Google login expired, the morning check stopped at "Write Morning log" and sent no line, on exactly the day something was wrong. Now the line is emailed first, and a failed log write does not stop the run.
2. n8n retries a step as a whole and only looks at its first item. With several PDFs in one email, a failure on the second PDF or later was never retried, and a failure on the first sent every PDF of that email to Claude again (paid twice). Now the PDFs go through the Claude step one at a time ("One PDF at a time" and "Keep file details"), so a retry repeats only the PDF that failed.

**Checked against n8n's own definitions:** every parameter of every step resolves in n8n's node definitions with 0 problems ([`test/n8n-params-check.json`](test/n8n-params-check.json)); the same check finds 4 of 4 mistakes planted in a copy. Picture of the workflow as drawn by the local n8n editor: [`test/n8n-canvas.png`](test/n8n-canvas.png). The red marks mean "credential not connected yet", which is expected with placeholders.

**Not tested:** a real Claude call, a real mailbox, a real Google Sheet, a real mail server, import through the n8n web screen, and n8n Cloud. In every test so far Claude's answer is the correct fields of each invoice, so the tests prove the checks and the routing, not how well Claude reads your documents. That is measured in two steps:

- [`test/live-eval.mjs`](test/live-eval.mjs) sends the 10 sample PDFs to the real Anthropic API exactly as the workflow builds them and reports, per field, on how many invoices Claude copied the printed value, plus tokens and cost. It is ready and has **not been run**: it costs an estimated USD 0.24 to 0.54 for the 10 invoices (up to 0.68 if every answer came from a fallback model), and it stops at a spending limit (USD 1 by default). Without `--run` it only prints the estimate.
- In the first live week, with your own invoices: every saved row keeps the number of tokens used, so accuracy and cost are visible per invoice.

## Live test with Claude

**Not run yet.** Run it from the `invoice-intake` folder:

```
# Put your Anthropic API key in a local .env file that is never committed (ANTHROPIC_API_KEY=...).
node --env-file=.env test/live-eval.mjs --run --max-usd 1.00
```

- Without `--run` it sends nothing. It only prints the estimate: about 0.24 to 0.54 USD for the 10 invoices, and up to 0.68 USD if every answer came from a fallback model.
- `--max-usd` is a spending limit, 1.00 USD by default. Before each call the script sets aside the most one attempt of it can cost (the whole max_tokens at the fallback price, input counted high: up to 0.46 USD here) and stops if that would pass the limit. A call that times out or loses the connection counts as its full reservation and is not retried. The one way to go over: when Anthropic answers with a fallback model it can bill two attempts for one call, so the total can pass the limit by up to one reservation. If every answer came from a fallback model, all 10 invoices need a limit of about 1.07 USD, so `--max-usd 1.10` lets the whole set run.
- It writes `test/LIVE-EVAL.txt` (fields copied right, final status per invoice, money spent, and for each invoice the model that answered and every attempt Anthropic reports) and Claude's raw answers in `test/live-eval/`. The API key is never printed or written to a file.

## Monthly running cost

You pay these directly, in your own accounts. Prices checked on 5 October 2026.

| Item | Cost | Source |
|---|---|---|
| n8n | Free if self-hosted on a server you already have (Community Edition), or n8n Cloud Starter at EUR 20 per month billed annually (more if paid monthly) for 2,500 runs | [n8n.io/pricing](https://n8n.io/pricing/) |
| Claude API (model Claude Opus 5.5) | USD 4 per million input tokens, USD 20 per million output tokens. Estimate: about USD 0.02 to 0.06 per one-page invoice | [Anthropic pricing](https://platform.claude.com/docs/en/about-claude/pricing), [PDF support](https://platform.claude.com/docs/en/build-with-claude/pdf-support) |
| Google Sheets, your mailbox | No extra cost if you already have them | |

How the per-invoice estimate is made: Anthropic says a PDF page uses about 1,500 to 3,000 text tokens plus the cost of the page as an image; with the instructions that comes to roughly 3,000 to 6,000 input tokens, plus 600 to 1,500 output tokens for the answer. This is an estimate, not a measurement.

If Claude Opus 5.5 declines a document, Anthropic can retry it on another Claude model (expected to be Claude Opus 5 or Claude Opus 4.8, at USD 5 / USD 25 per million tokens). The "model" column shows which model answered. The token columns count only the attempt that answered, not the declined one; the count per attempt is in Anthropic's response, which this workflow does not store, and your Anthropic usage page shows the total you are billed.

Example: 200 invoices a month on n8n Cloud Starter is EUR 20 plus about USD 4 to 12 for Claude. 1,000 invoices a month is EUR 20 (the run count still fits in 2,500) plus about USD 20 to 60. The model is one setting; a smaller Claude model costs less, and is worth testing once there is a week of real results to compare against.

## What we need from you

All of this is connected inside your own n8n. We never need your passwords.

- The mailbox where invoices arrive (IMAP access, usually an app password you create).
- A Google Sheet with three tabs: Invoices, Review queue, Morning log. The header rows are in [`sheet-template/`](sheet-template/).
- An Anthropic API key in your own Anthropic account, with a monthly spend limit you choose.
- An n8n API key, so the morning check can read the run history and the workflow. On n8n Cloud the API is part of the Starter, Pro and Enterprise plans, not of the free trial ([n8n docs](https://docs.n8n.io/connect/n8n-api/)).
- An email address for the reviewer and one for the morning line.
- 10 to 20 real invoices from your suppliers, to measure accuracy before switching it on.

## Settings you can change

- In the "Settings and known invoices" step: the Claude model, the allowed currencies, the rounding tolerance when totals are compared (2 cents by default), the maximum invoice age (365 days) and which document types are accepted. Changing this step does not trip the morning self-test.
- In the "Morning summary" step: how many quiet working days before ATTENTION (3) and how long a review item may wait (48 hours).

## Limits, said plainly

- The VAT check is a format and check-digit check. It does not ask a tax authority whether the number is registered. An EU VIES lookup can be added as a next step.
- Contracts, purchase orders and credit notes are recognised and sent to a person, not filed. Extracting contract fields is a separate job.
- The sample set is one-page, computer-made PDFs. Scanned, photographed or very long invoices need testing with your real files.
- Google Sheets is fine for a few thousand invoices a year. Beyond that, or to post straight into an accounting system, the last step changes to a database or your ERP; the reading and checking stay the same.
- The duplicate check uses supplier VAT number plus invoice number. An invoice re-sent after it was rejected (for example with the total corrected) is accepted, because rejected invoices are not counted as already processed.
- The morning check runs its two sample invoices with the default settings, not with the ones in your Settings step.

## Files

| Path | What it is |
|---|---|
| `workflow.json` | The n8n workflow, ready to import (28 steps: intake, morning check and notes) |
| `diagram.svg`, `diagram.png` | The picture above |
| `prompts/` | The instructions and the fixed form (JSON schema) given to Claude |
| `fixtures/` | The 10 synthetic invoices (`pdf/`), the correct answer for each (`expected/`) and the generator |
| `test/run-tests.mjs`, `test/RESULT.txt` | The offline test and its result |
| `test/mutation.mjs`, `test/MUTATION-RESULT.txt`, `test/mutation/` | The proof that the test fails when the rules are broken |
| `test/sample-output/` | The sheet rows, reviewer email and morning lines from the test run |
| `test/run-in-n8n.mjs`, `test/n8n-test-hook.cjs`, `test/N8N-RUN.txt` | The run inside a real n8n, its test loader (network guard, pinned data) and its result |
| `test/n8n-export.json`, `test/n8n-canvas.png` | The workflow as n8n exported it back after the import, and as the n8n editor draws it |
| `test/check-n8n-params.cjs`, `test/n8n-params-check.json` | The check of every parameter against n8n's own node definitions, and its output |
| `test/live-eval.mjs` | The paid accuracy check against the real Anthropic API (not run yet) |
| `sheet-template/` | Header rows for the three Google Sheet tabs |
| `src/` | The source of the workflow code; `build-workflow.mjs` builds `workflow.json` from it |
## For the technical person: install

1. In n8n, import `workflow.json` (Workflows, Import from file).
2. Create the credentials and select them on the steps: IMAP (Email Trigger), Header Auth with name `x-api-key` and your Anthropic key (HTTP Request "Claude reads the PDF"), Google Sheets OAuth2, SMTP (both email steps), and an n8n API key ("Failed runs" and "This workflow"). The n8n public API must be available on your plan (see above).
3. Replace `REPLACE_WITH_YOUR_SPREADSHEET_ID` on the six Google Sheets steps, and the email addresses on the two email steps. Optional: put your heartbeat URL in "Heartbeat ping (optional)" and switch that step on.
4. To change the code, edit `src/`, then run `node src/build-workflow.mjs` to rebuild `workflow.json`, then `node test/run-tests.mjs` and `node test/mutation.mjs` (and, with an n8n install, `N8N_BIN=<path to n8n/bin/n8n> node test/run-in-n8n.mjs`), and import the new file. If you edit a step inside n8n instead, the morning line will say the code changed since the tested build; copy the change back into `src/` and rebuild.
5. Send a few real invoices to the mailbox before activating, and compare the sheet with the paper.

The Claude call uses the Messages API with structured outputs (a strict JSON schema), effort set to low, and Anthropic's server-side fallback in case the model declines a document. PDFs go to Claude one at a time, so a large batch does not hit the API rate limit and a retry repeats only the PDF that failed. Successful runs are not stored in n8n's execution history (the sheet is the record), to keep invoice contents out of logs; failed runs are kept for debugging.
