Demo build by Kedgework. Not client work. Synthetic data.

hello@kedgework.com

All demo builds

Invoice and document intake

Supplier invoices are entered into a spreadsheet for you, the numbers are checked, and anything that looks wrong waits for a person instead of slipping through.

Supplier invoices arrive by email as PDFs. Someone opens each one, types the supplier, number, date and totals into a spreadsheet or the accounting system, and hopes nothing was typed twice or added up wrong. This workflow does that typing, checks the numbers with fixed rules, and only asks a person when something looks wrong. Every morning at 08:00 it sends one line: what it did, what is waiting for a person, and any sign it can see that something stopped working.

It runs in n8n, in your own n8n account, and writes to a Google Sheet you own.

Offline test
101 of 101 checks passed, run on 2026-10-05. The Claude answer is simulated with the expected fields of each fixture.
Inside a real n8n 2.41.7
16 passed, 0 failed, run on 2026-10-05. Nothing left the machine.
Live test with Claude
Not run yet. In every test so far Claude's answer is the correct fields of each invoice, so the tests prove the checks and the routing, not how well Claude reads your documents.
Workflow
n8n, 28 nodes. 12 credential slots, all placeholders until setup. Download

What it solves

Supplier invoices arrive by email as PDFs. Someone opens each one, types the supplier, number, date and totals into a spreadsheet or the accounting system, and hopes nothing was typed twice or added up wrong.

What the flow does

  1. An email arrives in the mailbox you use for invoices. Only unread mail is picked up, and each email is marked as read once it has been taken.
  2. Each PDF attachment becomes one job. An email with no PDF is not ignored: it goes to the review queue so a person sees it.
  3. Claude (the AI model by Anthropic) reads the PDF and copies the fields into a fixed form: supplier, VAT number, invoice number, dates, currency, each line, net total, VAT and total. It is told to copy what is printed and never to correct it, so a wrong total on the invoice stays wrong and gets caught in the next step. Invoices in Portuguese, Spanish, French, German, Dutch and English are in the sample set. PDFs go to Claude one at a time; if a call fails (network, overload), that PDF is tried up to 3 times in total.
  4. Fixed rules check the result (plain code, not AI, so they behave the same every time):
    CheckWhat it catches
    The lines add up to the net totalA line left out or misread
    Net plus VAT equals the total dueA wrong total on the invoice
    VAT matches the rates on the linesVAT charged at the wrong rate
    Supplier VAT number has a valid format, including the check digit for Portugal and FranceA mistyped or invented number
    Same supplier VAT and invoice number not already savedThe same invoice sent twice, or a reminder with a copy
    Invoice date is a real date, not in the future, not more than a year old; due date after invoice dateMisread dates, old invoices sent again
    Currency is on your listAn invoice in a currency you do not pay in
    It is an invoice, not a credit note, a contract or an orderOther documents sent to the same mailbox. A credit note gets its own reason, so it can be matched with the original invoice
    Required fields are presentA missing invoice number or tax number
  5. Clean invoices are saved as one row in the "Invoices" sheet, with the sender, subject and file name of the email they came from. Values are stored exactly as read: a subject that starts with "=" stays text and never becomes a spreadsheet formula, and invoice numbers keep their leading zeros.
  6. Anything doubtful goes to the "Review queue" sheet with the reason in plain words. A person gets one short email per batch listing those documents. When they have dealt with one, they set its status to done.

What it looks like

From the offline test run on the 10 sample invoices (synthetic data, full files in test/sample-output/).

"Invoices" sheet, 3 of the 7 saved rows (some columns left out here):

supplier_namesupplier_vatinvoice_numberinvoice_datecurrencynet_totalvat_totalgross_totalfile_name
Mondego Sample Supplies, LdaPT532449878FT 2026/01412026-09-14EUR259.8059.75319.55FT-2026-0141.pdf
Ejemplo Logistica Demo, S.L.ESB80114844A-2026-03882026-09-18EUR516.00108.36624.36A-2026-0388.pdf
Exemple Conseil Demo SARLFR88876963025F2026-10932026-09-22EUR1,420.00284.001,704.00Facture_F2026-1093.pdf

"Review queue" sheet, one of the 4 rows:

supplier_nameinvoice_numbergross_totalreason_textreasonsreview_status
Exemple Transport Demo SAST-26-00417316.00Net total plus VAT does not equal the total duetotals_mismatch:255+51=306 but total is 316open

The email the reviewer gets for that batch:

Subject: Invoices: 4 document(s) need a check

These documents were not saved automatically. They are in the "Review queue" sheet.

1. T-26-00417.pdf from Exemple Transport Demo SAS: Net total plus VAT does not equal the total due
2. A-2026-0388-copia.pdf from Ejemplo Logistica Demo, S.L.: This invoice was already processed (same supplier VAT and invoice number)
3. fatura_LDM_56.pdf from Lousa Demo Manutencao, Lda: A required field is missing on the document
4. (no file) from someone@b.example: The email had no PDF attachment

When you have dealt with one, set review_status to done.

What you see every morning

At 08:00 (Lisbon time) the workflow checks itself and sends one line by email. The email goes out first; the same line is then kept in a "Morning log" sheet, and a failure there (for example an expired Google login) cannot stop the email. Three lines produced by the offline test:

OK. Last 24 h: 2 invoice(s) saved, 0 sent to review. Open in review: 0. Failed runs: 0. Last document: 2026-10-05. Rules self-test: passed.
ATTENTION. Last 24 h: 3 invoice(s) saved, 1 sent to review. Open in review: 2 (oldest 50 h). Failed runs: 1. Last document: 2026-10-05. Rules self-test: passed. Needs a look: a review item has waited 50 h; 1 failed run(s).
ATTENTION. Last 24 h: 1 invoice(s) saved, 0 sent to review. Open in review: 0. Failed runs: 0. Last document: 2026-10-05. Rules self-test: FAILED. Needs a look: code changed since the tested build in: Check the invoice (run the offline test again).

What it checks:

What it cannot see from the inside:

Nothing disappears quietly

Diagram

Diagram. For every new email: the email arrives, the PDFs are taken, Claude reads each one, rules check it, then it goes to the Invoices sheet or to the Review queue. An email with no PDF also goes to the Review queue. Every morning at 08:00: the morning check, counts and checks, one line by email. The example numbers in the picture are illustrative.
How it works
The workflow as the n8n editor draws it. Intake row: new invoice email, split PDF attachments, has a PDF, read processed invoices, settings and known invoices, one PDF at a time, build Claude request, Claude reads the PDF, keep file details, check the invoice, is it clean, save to Invoices sheet; emails without a PDF and invoices with a problem are collected, added to the Review queue, and the reviewer gets one email. Morning check row, 08:00 Europe/Lisbon: read the Invoices sheet and the Review queue, failed runs and this workflow from the n8n API, morning summary, send the morning line, an optional heartbeat ping that is switched off, write the Morning log. Red marks show credentials that are not connected yet.
Picture of the workflow as drawn by the local n8n editor. The red marks mean "credential not connected yet", which is expected with placeholders.

How it was tested

Offline (on this machine, 5 October 2026): the code inside the workflow, taken straight out of workflow.json and run on 10 synthetic invoices plus 18 edge cases and 11 morning check situations. Result: 101 of 101 checks passed, see test/RESULT.txt. The 10 invoices are real PDFs in 6 languages; 7 are clean, 3 have a problem on purpose (a wrong total, a duplicate, a missing tax number), and all 3 were sent to review with the right reason.

The test catches broken rules: test/mutation.mjs breaks the rules on purpose in 3 ways (totals and duplicate checks off in the checking step; totals check off in that step only; totals check off everywhere and the workflow rebuilt) and runs the same test on each broken copy. All 3 fail, as they should (90, 91 and 89 of 100 checks pass on them). See test/MUTATION-RESULT.txt.

Inside a real n8n: test/run-in-n8n.mjs imports workflow.json into n8n 2.41.7 and runs copies of it from the command line. Result: 16 of 16 checks passed, see test/N8N-RUN.txt. That file carries the sha256 of the workflow.json it tested, and the offline test checks that it matches the file shipped here.

How it was run, said plainly: the n8n on the test machine is an npx install that was cut off before it finished. A small loader kept outside this folder fills in two missing parts (the sqlite3 database driver and one package folder) without changing any n8n file; a normal install (Docker, n8n Cloud, npm i -g n8n) does not need it. The test loader test/n8n-test-hook.cjs adds the network guard and lets the command line use pinned data, which n8n execute otherwise ignores. Google Sheets and the incoming mailbox cannot be pointed at a stand-in, so their data is pinned (fixed test data on the step, an n8n feature), and one sheet write is replaced by a step that passes its rows on. test/N8N-RUN.txt lists exactly what ran for real.

What the n8n run found, and what changed. The same run on the previous build failed 5 of 16 checks, for two reasons:

  1. If the Google login expired, the morning check stopped at "Write Morning log" and sent no line, on exactly the day something was wrong. Now the line is emailed first, and a failed log write does not stop the run.
  2. n8n retries a step as a whole and only looks at its first item. With several PDFs in one email, a failure on the second PDF or later was never retried, and a failure on the first sent every PDF of that email to Claude again (paid twice). Now the PDFs go through the Claude step one at a time ("One PDF at a time" and "Keep file details"), so a retry repeats only the PDF that failed.

Checked against n8n's own definitions: every parameter of every step resolves in n8n's node definitions with 0 problems (test/n8n-params-check.json); the same check finds 4 of 4 mistakes planted in a copy. Picture of the workflow as drawn by the local n8n editor: test/n8n-canvas.png. The red marks mean "credential not connected yet", which is expected with placeholders.

From the test files

Copied as they are, line by line. The full files are below.

test/RESULT.txt, line 4

Run at: 2026-10-05T19:43:43.999Z  (fixed "today" inside the test: 2026-10-05T08:00:00.000+01:00)

test/RESULT.txt, lines 6 to 7

Checks passed: 101 of 101
VERDICT: PASS

test/RESULT.txt, lines 11 to 13

What was NOT tested here: a real Claude API call (no key used, no money spent), a real mailbox, a real Google Sheet,
and how well Claude reads these PDFs. The run inside a real n8n is test/run-in-n8n.mjs (result: test/N8N-RUN.txt);
the paid accuracy check is test/live-eval.mjs (not run: it spends money).

test/N8N-RUN.txt, lines 2 to 3

Run in a real n8n: 2026-10-05T19:43:39.606Z, n8n 2.41.7, node v26.3.0, win32, command: node test/run-in-n8n.mjs
workflow.json sha256: 29c3ad8c4e30c28bbc613e363e1154ce04e9782dad835ee5d6e879fa4ec93f8c

test/N8N-RUN.txt, line 20

  Not tested here: a real Claude call, a real mailbox, a real Google Sheet, a real SMTP server, the n8n web editor and n8n Cloud.

test/N8N-RUN.txt, line 43

TOTAL: 16 passed, 0 failed
test/RESULT.txt, the full file (128 lines)
Kedgework demo: invoice and document intake. Offline test result
Demo build by Kedgework. Not client work. Synthetic data.

Run at: 2026-10-05T19:43:43.999Z  (fixed "today" inside the test: 2026-10-05T08:00:00.000+01:00)
Node: v26.3.0
Checks passed: 101 of 101
VERDICT: PASS

What was tested: the JavaScript inside the workflow Code nodes, taken out of workflow.json and run with
mocked n8n inputs. The Claude answer is simulated with the expected fields of each fixture.
What was NOT tested here: a real Claude API call (no key used, no money spent), a real mailbox, a real Google Sheet,
and how well Claude reads these PDFs. The run inside a real n8n is test/run-in-n8n.mjs (result: test/N8N-RUN.txt);
the paid accuracy check is test/live-eval.mjs (not run: it spends money).

Fixture outcome (status, supplier, total as read, reasons):
  inv-01  ok        Mondego Sample Supplies, Lda             319.55 EUR  -
  inv-02  ok        Ejemplo Logistica Demo, S.L.             624.36 EUR  -
  inv-03  ok        Exemple Conseil Demo SARL                  1704 EUR  -
  inv-04  ok        Beispiel Software Demo GmbH               571.2 EUR  -
  inv-05  ok        Sample Cloud Hosting Demo Ltd               306 GBP  -
  inv-06  ok        Voorbeeld Print Demo B.V.                 640.3 EUR  -
  inv-07  ok        Coimbra Demo Catering, Unipessoal Lda    244.59 EUR  -
  inv-08  exception Exemple Transport Demo SAS                  316 EUR  totals_mismatch:255+51=306 but total is 316
  inv-09  exception Ejemplo Logistica Demo, S.L.             624.36 EUR  duplicate:ESB80114844|A20260388
  inv-10  exception Lousa Demo Manutencao, Lda                221.4 EUR  missing_field:supplier_vat

All checks:
  [PASS] workflow: node names are unique
  [PASS] workflow: every connection points to an existing node
  [PASS] workflow: has a emailReadImap node
  [PASS] workflow: has a httpRequest node
  [PASS] workflow: has a googleSheets node
  [PASS] workflow: has a scheduleTrigger node
  [PASS] workflow: has a code node
  [PASS] workflow: has a if node
  [PASS] workflow: has a emailSend node
  [PASS] workflow: IMAP trigger reads only unread mail and keeps attachments
  [PASS] workflow: IMAP attachment prefix is a top-level parameter (n8n ignores it inside options)
  [PASS] workflow: all 3 sheet writes store values as sent (cellFormat RAW, so "=..." never becomes a formula)
  [PASS] workflow: emails go out without the n8n footer
  [PASS] workflow: review rows from both sources meet in one Merge, so the reviewer gets one email per batch
  [PASS] workflow: failed-runs lookup is limited to this workflow
  [PASS] workflow: morning reads never stop the morning run (errors come through as ATTENTION)
  [PASS] workflow: heartbeat ping exists and is off until a URL is set
  [PASS] workflow: Claude call goes to /v1/messages with anthropic-version 2023-06-01
  [PASS] workflow: Claude API key is a credential placeholder, not in the file
  [PASS] workflow: Claude call tries up to 3 times and never drops an item on error (an error goes on to the checks)
  [PASS] workflow: one PDF per pass through the Claude step (Loop Over Items, batch size 1), so a retry repeats only that PDF
  [PASS] workflow: the morning line is emailed first and does not depend on the Morning log write (which continues on error)
  [PASS] workflow: morning check scheduled at 08:00 Europe/Lisbon
  [PASS] workflow: all credentials are placeholders
  [PASS] workflow: no em or en dashes
  [PASS] workflow: carries the demo label
  [PASS] fixtures: 10 invoices, 3 with problems
  [PASS] fixtures: inv-01.pdf contains its expected values
  [PASS] fixtures: inv-02.pdf contains its expected values
  [PASS] fixtures: inv-03.pdf contains its expected values
  [PASS] fixtures: inv-04.pdf contains its expected values
  [PASS] fixtures: inv-05.pdf contains its expected values
  [PASS] fixtures: inv-06.pdf contains its expected values
  [PASS] fixtures: inv-07.pdf contains its expected values
  [PASS] fixtures: inv-08.pdf contains its expected values
  [PASS] fixtures: inv-09.pdf contains its expected values
  [PASS] fixtures: inv-10.pdf contains its expected values
  [PASS] fixtures: inv-10.pdf has no supplier NIF printed
  [PASS] schema: every object is closed and every field is required (nullable instead of optional)
  [PASS] node Split PDF attachments: 1 PDF item + 1 no-PDF item, logo ignored
  [PASS] node Settings and known invoices: attaches settings, known keys and keeps the PDF
  [PASS] node Build Claude request: one request per PDF
  [PASS] node Build Claude request: model, effort, strict JSON schema, server-side fallback
  [PASS] node Build Claude request: the PDF is sent byte for byte as a base64 document block
  [PASS] node Build Claude request: system prompt says to copy printed totals, not fix them
  [PASS] node Build Claude request: no thinking or temperature fields (not accepted the old way on this model)
  [PASS] node Keep file details: each answer is joined to the file and email of its own PDF, without the request
  [PASS] node Keep file details: an HTTP error after the retries is kept, so the PDF goes to review instead of disappearing
  [PASS] fixture inv-01: expected ok
  [PASS] fixture inv-02: expected ok
  [PASS] fixture inv-03: expected ok
  [PASS] fixture inv-04: expected ok
  [PASS] fixture inv-05: expected ok
  [PASS] fixture inv-06: expected ok
  [PASS] fixture inv-07: expected ok
  [PASS] fixture inv-08: expected exception (totals_mismatch)
  [PASS] fixture inv-09: expected exception (duplicate)
  [PASS] fixture inv-10: expected exception (missing_field)
  [PASS] batch: 7 saved to Invoices, 3 sent to Review queue
  [PASS] batch: review rows are marked open for a person
  [PASS] batch: saved rows carry the source email and file
  [PASS] batch: VAT numbers stored in one normal form
  [PASS] next day: inv-02 sent again is caught as a duplicate from the sheet
  [PASS] next day: inv-08 re-sent with the right total is accepted (rejected ones are not duplicates)
  [PASS] edge: invoice date in the future -> future_date
  [PASS] edge: invoice date older than 365 days -> too_old
  [PASS] edge: date that does not exist (30 Feb) -> bad_date
  [PASS] edge: due date before invoice date -> due_before_issue
  [PASS] edge: Portuguese NIF with a wrong check digit -> vat_format
  [PASS] edge: VAT number with an unknown country prefix -> vat_format
  [PASS] edge: currency not on the list -> currency_not_allowed
  [PASS] edge: lines do not add up to the net total -> lines_do_not_add_up
  [PASS] edge: VAT amount does not match the line rates -> vat_does_not_match_rates
  [PASS] edge: a contract, not an invoice -> not_an_invoice
  [PASS] edge: a credit note gets its own reason -> credit_note
  [PASS] edge: AI marked a field as unclear -> model_unsure
  [PASS] edge: invoice number missing -> missing_field
  [PASS] edge: no line items read -> no_line_items
  [PASS] edge: Claude declined (refusal) -> review queue (model_refused)
  [PASS] edge: answer cut off (max_tokens) -> review queue (extraction_incomplete)
  [PASS] edge: HTTP error after 3 tries (e.g. 529 overloaded) -> review queue (extraction_failed)
  [PASS] edge: answer is not JSON -> review queue (extraction_failed)
  [PASS] node No PDF: email without a PDF becomes an open review row
  [PASS] node Reviewer summary: one email listing the 3 documents
  [PASS] node Reviewer summary: no email when nothing needs a person
  [PASS] morning: counts last 24 h (3 saved, 1 to review), 2 open, 1 failed run of this workflow
  [PASS] morning: ATTENTION when a run failed or a review item waits more than 48 h
  [PASS] morning: rules self-test passes on the shipped build (live code matches the tested code)
  [PASS] morning: OK on a quiet clean day
  [PASS] morning: ATTENTION when it cannot read the run history (no false OK)
  [PASS] morning: ATTENTION when a sheet cannot be read, and the run still ends with a line
  [PASS] morning: ATTENTION when no document arrived for 3 full working days (Tue to Mon: Wed, Thu, Fri)
  [PASS] morning: a weekend with no mail is not an alarm (last document Thursday, now Monday)
  [PASS] morning: ATTENTION when no document has ever been received
  [PASS] morning: self-test FAILS when the totals check is removed from "Check the invoice" only
  [PASS] morning: self-test FAILS when the copy of the rules inside "Morning summary" is edited
  [PASS] morning: self-test is "unknown" and the line says ATTENTION when the workflow cannot be read
  [PASS] morning: changing the Settings node is allowed and does not raise ATTENTION
  [PASS] n8n: every node parameter resolves against the real n8n node definitions (no unknown names, modes or choices)  :: n8n-nodes-base 2.41.5, 28 nodes, 0 problems
  [PASS] n8n: control, the parameter check reports 4 of 4 planted mistakes (unknown option, two wrong choices, wrong locator mode)  :: Claude reads the PDF caught, Add to Review queue caught, New invoice email caught, Read Invoices sheet caught
  [PASS] n8n: this workflow.json was imported and run in a real n8n (test/N8N-RUN.txt has its sha256 and 0 failures)  :: n8n 2.41.7: 16 passed, 0 failed
test/N8N-RUN.txt, the full file (43 lines)
Kedgework demo build: invoice and document intake. Not client work. Synthetic data.
Run in a real n8n: 2026-10-05T19:43:39.606Z, n8n 2.41.7, node v26.3.0, win32, command: node test/run-in-n8n.mjs
workflow.json sha256: 29c3ad8c4e30c28bbc613e363e1154ce04e9782dad835ee5d6e879fa4ec93f8c

How n8n was started (said plainly):
  n8n binary: from an npx install (npm cache "_npx" folder), ../_n8n-home-docs/npm-cache/_npx/83f51bd5dfda7e85/node_modules/n8n/bin/n8n
  extra loader(s) given in NODE_OPTIONS by the caller: n8n-install-hook.cjs
  Not a normal install: n8n 2.41.7 came from an npx download (npm cache) that was cut off before it finished, so it has no sqlite3 binary for Node 26 and a half-unpacked @smithy/core. n8n-install-hook.cjs (kept outside the project) points sqlite3 to a separate sqlite3 5.1.7 copy with its prebuilt N-API binary and @smithy/core to the copy inside n8n-nodes-base. No n8n file is changed. A normal install (Docker image, n8n Cloud, or npm i -g n8n on Node 22 or 24) does not need it.
  test/n8n-test-hook.cjs (always loaded by this script): network guard, and pinned data for "n8n execute" (which ignores it otherwise).

What ran for real inside n8n, and what stood in for the outside world:
  Real n8n nodes: every Code, If, Merge and Loop step; "Claude reads the PDF" (HTTP Request) against a local Anthropic stand-in on 127.0.0.1;
    "Tell the reviewer" and "Send the morning line" (Send Email) against a local SMTP stand-in; "Failed runs" and "This workflow" (n8n node)
    against a local n8n API stand-in; the heartbeat (HTTP Request) in one case. Dummy credentials ("not-a-secret") went only to those stand-ins.
  Pinned data (n8n's own mechanism): the emails, on a Manual Trigger that replaces the IMAP trigger (made with nodemailer and parsed with
    mailparser, as the IMAP node does); the Google Sheets reads; "Save to Invoices sheet" and "Write Morning log" (their output is not used later).
  Replaced by a pass-through Code node: "Add to Review queue" (pinned data is fixed, and the next step needs the rows it was given).
  Real Google Sheets nodes that fail on purpose: the two morning reads and "Write Morning log" in the "Google connection broken" case
    (credential never connected, so n8n stops before any network call).
  Not tested here: a real Claude call, a real mailbox, a real Google Sheet, a real SMTP server, the n8n web editor and n8n Cloud.

PASS  workflow.json imports into n8n 2.41.7 as shipped  :: 6.7 s
PASS  export after import: same nodes, parameters, error settings, credential ids, connections and settings (saved as test/n8n-export.json)  :: 28 nodes
PASS  n8n accepts the placeholder credential ids and types (5 placeholders + 1 extra test credential)  :: imap, googleSheetsOAuth2Api, httpHeaderAuth, smtp, n8nApi
PASS  test copies import (7 workflows)  :: 5 s
PASS  in n8n: 9 emails, 10 PDFs (two in one email, one email also has a logo) -> 7 saved, 3 to review with the right reason, plus the email without a PDF  :: 7 saved, 4 to review: T-26-00417.pdf totals_mismatch, A-2026-0388-copia.pdf duplicate, fatura_LDM_56.pdf missing_field, (no file) no_pdf
PASS  in n8n: the real HTTP Request node sent one Claude request per PDF, one PDF per run, to the local stand-in, with the PDF bytes, headers, model, schema and key header  :: 10 requests in 10 runs of the step; every PDF matched its fixture by sha256; x-api-key came from the n8n credential
PASS  in n8n: saved rows keep the values read, the source email and file, the model and the tokens  :: e.g. FT-2026-0141.pdf, faturacao@mondego-sample.example, 319.55 EUR, 3001+700 tokens
PASS  in n8n: the real Send Email node sent one reviewer email listing the 4 documents (no n8n footer)  :: "Invoices: 4 document(s) need a check" to reviewer@replace-with-your-domain.com; 4 numbered lines
      nodes that ran (real unless marked): Split PDF attachments, Has a PDF?, Read processed invoices [pinned], Settings and known invoices, One PDF at a time, Build Claude request, Claude reads the PDF, Keep file details, Check the invoice, Is it clean?, Save to Invoices sheet [pinned], No PDF: review row, Collect review items, Add to Review queue [stub], Reviewer summary, Tell the reviewer
PASS  in n8n: first PDF of an email overloaded once -> that PDF is asked again, the other PDF is not sent twice, both saved  :: requests: inv-01 2, inv-02 1; saved 2; reviewer email not needed
PASS  in n8n: second PDF of an email always overloaded -> tried 3 times, then to the review queue as "AI step failed"; the first PDF is saved  :: requests: inv-01 1, inv-02 3; saved 1; review A-2026-0388.pdf extraction_failed:529 - "{\"type\":\"error\",\"error\":{\"ty; reviewer email sent
PASS  in n8n: morning check on a normal day -> OK line by email, Lisbon time, run history and live code read through the real n8n API node  :: OK. Last 24 h: 3 invoice(s) saved, 1 sent to review. Open in review: 1 (oldest 2 h). Failed runs: 0. Last document: 2026-10-05. Rules self-test: passed. | 2026-10-05T20:43:18.601+01:00
      nodes that ran (real unless marked): Read Invoices sheet [pinned], Read Review queue [pinned], Failed runs (n8n API), This workflow (n8n API), Morning summary, Send the morning line, Heartbeat ping (optional) [switched off, passed through], Write Morning log [pinned]
PASS  in n8n: Google connection broken (real Sheets nodes fail: credential not connected) -> ATTENTION email still sent, Morning log fails without stopping the run, heartbeat still sent  :: email "Invoice intake morning check: ATTENTION"; Morning log error: Unable to sign without access token; heartbeat pings: 1
PASS  in n8n: n8n API not reachable (real n8n nodes fail, connection refused on 127.0.0.1) -> ATTENTION, failed runs and self-test "unknown", email sent  :: ATTENTION. Last 24 h: 3 invoice(s) saved, 1 sent to review. Open in review: 1 (oldest 2 h). Failed runs: unknown. Last document: 2026-10-05. Rules self-test: unknown. Needs a look: could not read the n8n run history; could not read this workflow to compare its code with the tested build.
PASS  in n8n: totals check removed from "Check the invoice" in the live copy, and 1 failed run -> "Rules self-test: FAILED" names the step, ATTENTION email  :: ATTENTION. Last 24 h: 3 invoice(s) saved, 1 sent to review. Open in review: 1 (oldest 2 h). Failed runs: 1. Last document: 2026-10-05. Rules self-test: FAILED. Needs a look: 1 failed run(s); code changed since the tested build in: Check the invoice (run the offline test again).
PASS  network guard: no connection or DNS lookup to any host other than 127.0.0.1 during the whole run  :: guard loaded in 12 n8n processes, 0 refused
PASS  pinned data handed to every test run (7 runs)  :: 7 runs used pinned data

Local stand-ins: Anthropic on 127.0.0.1:52246, n8n API on 127.0.0.1:52247, SMTP on 127.0.0.1:52248. Full n8n output of each command: <N8N_USER_FOLDER>/kedgework-test/*.log on the test machine (not shipped).

TOTAL: 16 passed, 0 failed

Honest status

Not tested: a real Claude call, a real mailbox, a real Google Sheet, a real mail server, import through the n8n web screen, and n8n Cloud. In every test so far Claude's answer is the correct fields of each invoice, so the tests prove the checks and the routing, not how well Claude reads your documents. That is measured in two steps:

Limits, said plainly

Monthly cost

You pay these directly, in your own accounts. Prices checked on 5 October 2026.

ItemCostSource
n8nFree if self-hosted on a server you already have (Community Edition), or n8n Cloud Starter at EUR 20 per month billed annually (more if paid monthly) for 2,500 runsn8n.io/pricing
Claude API (model Claude Opus 5.5)USD 4 per million input tokens, USD 20 per million output tokens. Estimate: about USD 0.02 to 0.06 per one-page invoiceAnthropic pricing, PDF support
Google Sheets, your mailboxNo extra cost if you already have them

How the per-invoice estimate is made: Anthropic says a PDF page uses about 1,500 to 3,000 text tokens plus the cost of the page as an image; with the instructions that comes to roughly 3,000 to 6,000 input tokens, plus 600 to 1,500 output tokens for the answer. This is an estimate, not a measurement.

If Claude Opus 5.5 declines a document, Anthropic can retry it on another Claude model (expected to be Claude Opus 5 or Claude Opus 4.8, at USD 5 / USD 25 per million tokens). The "model" column shows which model answered. The token columns count only the attempt that answered, not the declined one; the count per attempt is in Anthropic's response, which this workflow does not store, and your Anthropic usage page shows the total you are billed.

Example: 200 invoices a month on n8n Cloud Starter is EUR 20 plus about USD 4 to 12 for Claude. 1,000 invoices a month is EUR 20 (the run count still fits in 2,500) plus about USD 20 to 60. The model is one setting; a smaller Claude model costs less, and is worth testing once there is a week of real results to compare against.

Download the workflow

Download workflow.jsonn8n workflow, 28 nodes, 127 KB

The n8n workflow, ready to import (28 steps: intake, morning check and notes)

Same file as tested: its sha256 is 29c3ad8c4e30c28bbc613e363e1154ce04e9782dad835ee5d6e879fa4ec93f8c, the value on line 3 of test/N8N-RUN.txt.

Download README.md (22 KB). Its links point to other files of the demo package (prompts/, fixtures/, sheet-template/, src/, test/), which are not published here.

Full README

Rendered from README.md. Links to files that are not published here are shown as plain text.

Show the full README.md (205 lines)

Invoice and document intake

Demo build by Kedgework. Not client work. Synthetic data.

Supplier invoices arrive by email as PDFs. Someone opens each one, types the supplier, number, date and totals into a spreadsheet or the accounting system, and hopes nothing was typed twice or added up wrong. This workflow does that typing, checks the numbers with fixed rules, and only asks a person when something looks wrong. Every morning at 08:00 it sends one line: what it did, what is waiting for a person, and any sign it can see that something stopped working.

It runs in n8n, in your own n8n account, and writes to a Google Sheet you own.

Diagram. For every new email: the email arrives, the PDFs are taken, Claude reads each one, rules check it, then it goes to the Invoices sheet or to the Review queue. An email with no PDF also goes to the Review queue. Every morning at 08:00: the morning check, counts and checks, one line by email. The example numbers in the picture are illustrative.
How it works

Step by step

  1. An email arrives in the mailbox you use for invoices. Only unread mail is picked up, and each email is marked as read once it has been taken.
  2. Each PDF attachment becomes one job. An email with no PDF is not ignored: it goes to the review queue so a person sees it.
  3. Claude (the AI model by Anthropic) reads the PDF and copies the fields into a fixed form: supplier, VAT number, invoice number, dates, currency, each line, net total, VAT and total. It is told to copy what is printed and never to correct it, so a wrong total on the invoice stays wrong and gets caught in the next step. Invoices in Portuguese, Spanish, French, German, Dutch and English are in the sample set. PDFs go to Claude one at a time; if a call fails (network, overload), that PDF is tried up to 3 times in total.
  4. Fixed rules check the result (plain code, not AI, so they behave the same every time):
    CheckWhat it catches
    The lines add up to the net totalA line left out or misread
    Net plus VAT equals the total dueA wrong total on the invoice
    VAT matches the rates on the linesVAT charged at the wrong rate
    Supplier VAT number has a valid format, including the check digit for Portugal and FranceA mistyped or invented number
    Same supplier VAT and invoice number not already savedThe same invoice sent twice, or a reminder with a copy
    Invoice date is a real date, not in the future, not more than a year old; due date after invoice dateMisread dates, old invoices sent again
    Currency is on your listAn invoice in a currency you do not pay in
    It is an invoice, not a credit note, a contract or an orderOther documents sent to the same mailbox. A credit note gets its own reason, so it can be matched with the original invoice
    Required fields are presentA missing invoice number or tax number
  5. Clean invoices are saved as one row in the "Invoices" sheet, with the sender, subject and file name of the email they came from. Values are stored exactly as read: a subject that starts with "=" stays text and never becomes a spreadsheet formula, and invoice numbers keep their leading zeros.
  6. Anything doubtful goes to the "Review queue" sheet with the reason in plain words. A person gets one short email per batch listing those documents. When they have dealt with one, they set its status to done.

What it looks like

From the offline test run on the 10 sample invoices (synthetic data, full files in test/sample-output/).

"Invoices" sheet, 3 of the 7 saved rows (some columns left out here):

supplier_namesupplier_vatinvoice_numberinvoice_datecurrencynet_totalvat_totalgross_totalfile_name
Mondego Sample Supplies, LdaPT532449878FT 2026/01412026-09-14EUR259.8059.75319.55FT-2026-0141.pdf
Ejemplo Logistica Demo, S.L.ESB80114844A-2026-03882026-09-18EUR516.00108.36624.36A-2026-0388.pdf
Exemple Conseil Demo SARLFR88876963025F2026-10932026-09-22EUR1,420.00284.001,704.00Facture_F2026-1093.pdf

"Review queue" sheet, one of the 4 rows:

supplier_nameinvoice_numbergross_totalreason_textreasonsreview_status
Exemple Transport Demo SAST-26-00417316.00Net total plus VAT does not equal the total duetotals_mismatch:255+51=306 but total is 316open

The email the reviewer gets for that batch:

Subject: Invoices: 4 document(s) need a check

These documents were not saved automatically. They are in the "Review queue" sheet.

1. T-26-00417.pdf from Exemple Transport Demo SAS: Net total plus VAT does not equal the total due
2. A-2026-0388-copia.pdf from Ejemplo Logistica Demo, S.L.: This invoice was already processed (same supplier VAT and invoice number)
3. fatura_LDM_56.pdf from Lousa Demo Manutencao, Lda: A required field is missing on the document
4. (no file) from someone@b.example: The email had no PDF attachment

When you have dealt with one, set review_status to done.

What you see every morning

At 08:00 (Lisbon time) the workflow checks itself and sends one line by email. The email goes out first; the same line is then kept in a "Morning log" sheet, and a failure there (for example an expired Google login) cannot stop the email. Three lines produced by the offline test:

OK. Last 24 h: 2 invoice(s) saved, 0 sent to review. Open in review: 0. Failed runs: 0. Last document: 2026-10-05. Rules self-test: passed.
ATTENTION. Last 24 h: 3 invoice(s) saved, 1 sent to review. Open in review: 2 (oldest 50 h). Failed runs: 1. Last document: 2026-10-05. Rules self-test: passed. Needs a look: a review item has waited 50 h; 1 failed run(s).
ATTENTION. Last 24 h: 1 invoice(s) saved, 0 sent to review. Open in review: 0. Failed runs: 0. Last document: 2026-10-05. Rules self-test: FAILED. Needs a look: code changed since the tested build in: Check the invoice (run the offline test again).

What it checks:

  • The counts: invoices saved and sent to review in the last 24 hours, review items still open, and failed runs of this workflow.
  • That documents are still arriving. If no document at all came in for 3 full working days (Monday to Friday), the line says ATTENTION. A mailbox that stopped delivering looks exactly like a quiet day, so silence is treated as a warning, not as good news.
  • That the checking rules are the ones we tested. It reads the workflow from your n8n and compares the code of each intake step with a fingerprint of the code that passed our offline test. Then it runs the rules on two built-in sample invoices, one clean and one with a wrong total. If a step was edited inside n8n, or the rules no longer catch the wrong total, the line says "Rules self-test: FAILED" and names the step.
  • That it could read what it needs. If it cannot read a sheet, the run history or the workflow itself, the line says ATTENTION and names what it could not read. The morning run still finishes and still sends its line.

What it cannot see from the inside:

  • If n8n itself is down, or the schedule was switched off, no line is sent at all. Rule of thumb: if the 08:00 email has not arrived by 08:15, something is down. To be alerted for that too, switch on the "Heartbeat ping (optional)" step and point it at a dead man's switch service (for example Healthchecks.io): such a service alerts you when the daily ping does not arrive.
  • A mailbox that stops for 1 or 2 working days looks like a quiet day until the third. Public holidays count as working days, so after a long holiday the line may say ATTENTION once.
  • How well Claude reads your invoices. The self-test proves the rules work, not the reading. Reading problems show up as review items and, in the first live week, in the comparison with the paper.

Nothing disappears quietly

  • No PDF in the email: goes to the review queue.
  • The AI step fails (network, overload): that PDF is tried up to 3 times in total, then it goes to the review queue with the reason. The other PDFs of the same email are not sent again.
  • Claude declines to read a document, or its answer is cut off: review queue.
  • A run of the workflow fails completely: counted in the next morning line.
  • The morning check cannot read a sheet or the run history: the line says ATTENTION instead of failing silently.
  • The Google login expires: the morning line is still emailed, says ATTENTION and names the sheets it could not read. Only the copy in the "Morning log" sheet is missing that day.

What was tested, and what was not

Offline (on this machine, 5 October 2026): the code inside the workflow, taken straight out of workflow.json and run on 10 synthetic invoices plus 18 edge cases and 11 morning check situations. Result: 101 of 101 checks passed, see test/RESULT.txt. The 10 invoices are real PDFs in 6 languages; 7 are clean, 3 have a problem on purpose (a wrong total, a duplicate, a missing tax number), and all 3 were sent to review with the right reason.

The test catches broken rules: test/mutation.mjs breaks the rules on purpose in 3 ways (totals and duplicate checks off in the checking step; totals check off in that step only; totals check off everywhere and the workflow rebuilt) and runs the same test on each broken copy. All 3 fail, as they should (90, 91 and 89 of 100 checks pass on them). See test/MUTATION-RESULT.txt.

Inside a real n8n: test/run-in-n8n.mjs imports workflow.json into n8n 2.41.7 and runs copies of it from the command line. Result: 16 of 16 checks passed, see test/N8N-RUN.txt. That file carries the sha256 of the workflow.json it tested, and the offline test checks that it matches the file shipped here.

  • Import and export: the same 28 steps, settings, connections and credential placeholders come back.
  • 9 emails with 10 PDFs (two PDFs in one email, a logo next to a PDF in another, plus one email with no PDF) go through the real steps: 7 saved, 3 to review with the right reason, plus the email without a PDF, and one reviewer email listing the 4.
  • The real "Claude reads the PDF" step sends each PDF, one at a time, to a stand-in for the Anthropic API running on this machine, with the right headers, model, schema and the key taken from the n8n credential. A PDF that gets "overloaded" once is asked again (only that PDF); one that keeps failing goes to review after 3 tries.
  • The morning check runs with the real n8n API step (against a local stand-in), the real email step (to a local mail server) and, in one case, real Google Sheets steps with a login that does not work: the line still goes out, says ATTENTION and names what it could not read.
  • Nothing left the machine: a guard inside n8n refuses any connection to another host, and the run fails if it had to refuse one. It refused none.

How it was run, said plainly: the n8n on the test machine is an npx install that was cut off before it finished. A small loader kept outside this folder fills in two missing parts (the sqlite3 database driver and one package folder) without changing any n8n file; a normal install (Docker, n8n Cloud, npm i -g n8n) does not need it. The test loader test/n8n-test-hook.cjs adds the network guard and lets the command line use pinned data, which n8n execute otherwise ignores. Google Sheets and the incoming mailbox cannot be pointed at a stand-in, so their data is pinned (fixed test data on the step, an n8n feature), and one sheet write is replaced by a step that passes its rows on. test/N8N-RUN.txt lists exactly what ran for real.

What the n8n run found, and what changed. The same run on the previous build failed 5 of 16 checks, for two reasons:

  1. If the Google login expired, the morning check stopped at "Write Morning log" and sent no line, on exactly the day something was wrong. Now the line is emailed first, and a failed log write does not stop the run.
  2. n8n retries a step as a whole and only looks at its first item. With several PDFs in one email, a failure on the second PDF or later was never retried, and a failure on the first sent every PDF of that email to Claude again (paid twice). Now the PDFs go through the Claude step one at a time ("One PDF at a time" and "Keep file details"), so a retry repeats only the PDF that failed.

Checked against n8n's own definitions: every parameter of every step resolves in n8n's node definitions with 0 problems (test/n8n-params-check.json); the same check finds 4 of 4 mistakes planted in a copy. Picture of the workflow as drawn by the local n8n editor: test/n8n-canvas.png. The red marks mean "credential not connected yet", which is expected with placeholders.

Not tested: a real Claude call, a real mailbox, a real Google Sheet, a real mail server, import through the n8n web screen, and n8n Cloud. In every test so far Claude's answer is the correct fields of each invoice, so the tests prove the checks and the routing, not how well Claude reads your documents. That is measured in two steps:

  • test/live-eval.mjs sends the 10 sample PDFs to the real Anthropic API exactly as the workflow builds them and reports, per field, on how many invoices Claude copied the printed value, plus tokens and cost. It is ready and has not been run: it costs an estimated USD 0.24 to 0.54 for the 10 invoices (up to 0.68 if every answer came from a fallback model), and it stops at a spending limit (USD 1 by default). Without --run it only prints the estimate.
  • In the first live week, with your own invoices: every saved row keeps the number of tokens used, so accuracy and cost are visible per invoice.

Live test with Claude

Not run yet. Run it from the invoice-intake folder:

# Put your Anthropic API key in a local .env file that is never committed (ANTHROPIC_API_KEY=...).
node --env-file=.env test/live-eval.mjs --run --max-usd 1.00
  • Without --run it sends nothing. It only prints the estimate: about 0.24 to 0.54 USD for the 10 invoices, and up to 0.68 USD if every answer came from a fallback model.
  • --max-usd is a spending limit, 1.00 USD by default. Before each call the script sets aside the most one attempt of it can cost (the whole max_tokens at the fallback price, input counted high: up to 0.46 USD here) and stops if that would pass the limit. A call that times out or loses the connection counts as its full reservation and is not retried. The one way to go over: when Anthropic answers with a fallback model it can bill two attempts for one call, so the total can pass the limit by up to one reservation. If every answer came from a fallback model, all 10 invoices need a limit of about 1.07 USD, so --max-usd 1.10 lets the whole set run.
  • It writes test/LIVE-EVAL.txt (fields copied right, final status per invoice, money spent, and for each invoice the model that answered and every attempt Anthropic reports) and Claude's raw answers in test/live-eval/. The API key is never printed or written to a file.

Monthly running cost

You pay these directly, in your own accounts. Prices checked on 5 October 2026.

ItemCostSource
n8nFree if self-hosted on a server you already have (Community Edition), or n8n Cloud Starter at EUR 20 per month billed annually (more if paid monthly) for 2,500 runsn8n.io/pricing
Claude API (model Claude Opus 5.5)USD 4 per million input tokens, USD 20 per million output tokens. Estimate: about USD 0.02 to 0.06 per one-page invoiceAnthropic pricing, PDF support
Google Sheets, your mailboxNo extra cost if you already have them

How the per-invoice estimate is made: Anthropic says a PDF page uses about 1,500 to 3,000 text tokens plus the cost of the page as an image; with the instructions that comes to roughly 3,000 to 6,000 input tokens, plus 600 to 1,500 output tokens for the answer. This is an estimate, not a measurement.

If Claude Opus 5.5 declines a document, Anthropic can retry it on another Claude model (expected to be Claude Opus 5 or Claude Opus 4.8, at USD 5 / USD 25 per million tokens). The "model" column shows which model answered. The token columns count only the attempt that answered, not the declined one; the count per attempt is in Anthropic's response, which this workflow does not store, and your Anthropic usage page shows the total you are billed.

Example: 200 invoices a month on n8n Cloud Starter is EUR 20 plus about USD 4 to 12 for Claude. 1,000 invoices a month is EUR 20 (the run count still fits in 2,500) plus about USD 20 to 60. The model is one setting; a smaller Claude model costs less, and is worth testing once there is a week of real results to compare against.

What we need from you

All of this is connected inside your own n8n. We never need your passwords.

  • The mailbox where invoices arrive (IMAP access, usually an app password you create).
  • A Google Sheet with three tabs: Invoices, Review queue, Morning log. The header rows are in sheet-template/.
  • An Anthropic API key in your own Anthropic account, with a monthly spend limit you choose.
  • An n8n API key, so the morning check can read the run history and the workflow. On n8n Cloud the API is part of the Starter, Pro and Enterprise plans, not of the free trial (n8n docs).
  • An email address for the reviewer and one for the morning line.
  • 10 to 20 real invoices from your suppliers, to measure accuracy before switching it on.

Settings you can change

  • In the "Settings and known invoices" step: the Claude model, the allowed currencies, the rounding tolerance when totals are compared (2 cents by default), the maximum invoice age (365 days) and which document types are accepted. Changing this step does not trip the morning self-test.
  • In the "Morning summary" step: how many quiet working days before ATTENTION (3) and how long a review item may wait (48 hours).

Limits, said plainly

  • The VAT check is a format and check-digit check. It does not ask a tax authority whether the number is registered. An EU VIES lookup can be added as a next step.
  • Contracts, purchase orders and credit notes are recognised and sent to a person, not filed. Extracting contract fields is a separate job.
  • The sample set is one-page, computer-made PDFs. Scanned, photographed or very long invoices need testing with your real files.
  • Google Sheets is fine for a few thousand invoices a year. Beyond that, or to post straight into an accounting system, the last step changes to a database or your ERP; the reading and checking stay the same.
  • The duplicate check uses supplier VAT number plus invoice number. An invoice re-sent after it was rejected (for example with the total corrected) is accepted, because rejected invoices are not counted as already processed.
  • The morning check runs its two sample invoices with the default settings, not with the ones in your Settings step.

Files

PathWhat it is
workflow.jsonThe n8n workflow, ready to import (28 steps: intake, morning check and notes)
diagram.svg, diagram.pngThe picture above
prompts/The instructions and the fixed form (JSON schema) given to Claude
fixtures/The 10 synthetic invoices (pdf/), the correct answer for each (expected/) and the generator
test/run-tests.mjs, test/RESULT.txtThe offline test and its result
test/mutation.mjs, test/MUTATION-RESULT.txt, test/mutation/The proof that the test fails when the rules are broken
test/sample-output/The sheet rows, reviewer email and morning lines from the test run
test/run-in-n8n.mjs, test/n8n-test-hook.cjs, test/N8N-RUN.txtThe run inside a real n8n, its test loader (network guard, pinned data) and its result
test/n8n-export.json, test/n8n-canvas.pngThe workflow as n8n exported it back after the import, and as the n8n editor draws it
test/check-n8n-params.cjs, test/n8n-params-check.jsonThe check of every parameter against n8n's own node definitions, and its output
test/live-eval.mjsThe paid accuracy check against the real Anthropic API (not run yet)
sheet-template/Header rows for the three Google Sheet tabs
src/The source of the workflow code; build-workflow.mjs builds workflow.json from it

For the technical person: install

  1. In n8n, import workflow.json (Workflows, Import from file).
  2. Create the credentials and select them on the steps: IMAP (Email Trigger), Header Auth with name x-api-key and your Anthropic key (HTTP Request "Claude reads the PDF"), Google Sheets OAuth2, SMTP (both email steps), and an n8n API key ("Failed runs" and "This workflow"). The n8n public API must be available on your plan (see above).
  3. Replace REPLACE_WITH_YOUR_SPREADSHEET_ID on the six Google Sheets steps, and the email addresses on the two email steps. Optional: put your heartbeat URL in "Heartbeat ping (optional)" and switch that step on.
  4. To change the code, edit src/, then run node src/build-workflow.mjs to rebuild workflow.json, then node test/run-tests.mjs and node test/mutation.mjs (and, with an n8n install, N8N_BIN=<path to n8n/bin/n8n> node test/run-in-n8n.mjs), and import the new file. If you edit a step inside n8n instead, the morning line will say the code changed since the tested build; copy the change back into src/ and rebuild.
  5. Send a few real invoices to the mailbox before activating, and compare the sheet with the paper.

The Claude call uses the Messages API with structured outputs (a strict JSON schema), effort set to low, and Anthropic's server-side fallback in case the model declines a document. PDFs go to Claude one at a time, so a large batch does not hit the API rate limit and a retry repeats only the PDF that failed. Successful runs are not stored in n8n's execution history (the sheet is the record), to keep invoice contents out of logs; failed runs are kept for debugging.