# AI support inbox triage

**Demo build by Kedgework. Not client work. Synthetic data.**
ExamplePay is a fictional payments company. Every message in this folder was invented for the demo.

![How a message moves through the flow](diagram.png)

## The problem

A payments company gets the same questions all day: fees, limits, "where is my transfer". Mixed in are a few that cannot wait, like an unrecognised payment or a legal threat, and they wait behind the easy ones.

## What the flow does

1. **A message arrives** by email, web chat or WhatsApp. The flow confirms receipt to your inbox tool at once. A copy with the same message id is dropped until the next 08:00 check clears the list, so a quick retry from your tool gets no second answer; a resend after 08:00 counts as new.
2. **Claude reads it**: topic, language (Spanish, Portuguese, English), mood, and a short draft reply, only when your approved answers cover the question.
3. **Fixed rules decide, not the AI.** A person always gets the message when:
   - it mentions fraud, a payment the customer did not make, the police or a lawyer (marked urgent, even when the AI call fails)
   - it is about identity checks, refunds or disputes
   - a payment question involves 1,000 USD or more
   - the customer is angry, or your inbox tool reports a third message on the same issue
   - the AI is unsure, or its draft contains a number that is not in your approved answers
4. **Payment questions with a reference** are checked in your payments system (read only) and answered from a fixed template, never from the AI.
5. **A person approves, then the reply goes out.** For the first weeks every draft is posted in Slack with two buttons, **Approve and send** or **Decline, answer by hand**. Declined drafts, and drafts with no click within 4 hours, go to your team to answer by hand.

![The approval message in Slack (mock-up)](images/approval-card.png)

*Mock-up of the approval message, built with n8n's own Slack message code from the workflow's template and the invented message fx01. Not a screenshot of Slack.*

**In this demo, replies go out by email.** For web chat and WhatsApp the approved reply is posted in Slack for a person to send, until we connect your chat provider. The payment check matches the sender's email with the payment owner, so there a payment question goes to a person unless your chat tool passes a verified email.

**If Slack is down**, alerts for urgent messages and for the team are tried 3 times, then sent by email to a team address you choose.

## What you import into n8n

![The workflow drawn by n8n](images/n8n-canvas.png)

*`workflow.json` opened in a local n8n 2.41.7. The red marks mean "credential not connected yet": every credential is a placeholder until setup.*

## Checked every morning at 08:00, in your time zone

- Two test messages go through the live flow. Five minutes later the log must show one answered and one flagged as urgent. Test messages never reach customers or your team.
- Runs that failed in the last 24 hours, read from n8n. If that list or the log cannot be read, the check says FAIL instead of reporting zero.
- How many messages were answered, checked, passed to people (and how many urgent), and how many AI answers were unusable. No messages at all means the inbox may be disconnected.
- The list of message ids from the day before is cleared (n8n caps its size).

You get one line in Slack, problems also by email, and a ping to an outside monitor in case the check itself does not run.

## What it costs to run each month

Example: 50 messages a day, about 1,600 runs a month including the morning checks.

| Item | Estimate | Source |
|---|---|---|
| Claude Opus 5.5 (Anthropic API) | about 20 to 40 USD | 4 USD per million input tokens, 20 USD per million output tokens, [Anthropic pricing](https://platform.claude.com/docs/en/about-claude/pricing), checked 2026-10-05. About 2,000 tokens in and 300 to 500 out per message. Opus 5.5 always thinks a little before answering (the flow sets effort to low), billed as output, hence the upper end. |
| n8n | from 20 EUR a month on n8n Cloud (Starter, 2,500 runs, billed yearly), or free if self-hosted | [n8n pricing](https://n8n.io/pricing/), checked 2026-10-05 |
| Google Sheets and Slack | your existing accounts | |

Caching, not counted above: the fixed instructions (about 1,800 tokens) are cached for an hour. The first message after a quiet hour writes them at twice the input price (8 USD per million, about 0.015 USD), the next ones read them at 0.20 USD per million. At 50 messages a day that part of the input costs about 2 USD a month instead of 11. At low volume, when most messages arrive more than an hour apart, each pays the write: about 0.007 USD more per message than without caching. Claude Sonnet 5.5 (2 and 10 USD per million, one field in the Settings node) would roughly halve the Claude cost. In the first week we measure the real token use and send you the number.

## What is stored, and where

- **The log sheet**: category, route, rule, confidence, token counts and errors. No message text and no email address.
- **n8n** keeps no copy of successful runs. Failed runs are kept so they can be fixed and may include the message; on a self-hosted n8n we delete them after 7 days, on n8n Cloud your plan's retention applies.
- **Slack messages and alert emails** to your team include the sender and a short preview, so keep those channels private.
- **The payment check** sends the customer's email in a request header, never in the web address.

## What we would need from you

- Where messages arrive today (mailbox, chat tool, WhatsApp provider), with a stable id per message (most tools send one).
- 30 to 50 past messages with personal data removed, and how your team answered them.
- Your approved answers (help articles or saved replies).
- A read-only API key for payment lookups, or we start without lookups.
- Five private Slack channels (#support-drafts, #support-team, #support-urgent, #sales, #support-daily) and a team email address for alerts when Slack is down.
- An Anthropic API key and an n8n account in your name, so you see every bill. We never ask for passwords: an invite with the smallest permissions is enough.

## Honest status

- **Tested offline:** the workflow file and every node setting (against n8n 2.41's own node definitions), every rule on 12 invented messages, the approval limit, the Slack fallback and the morning check. [test/RESULT.txt](test/RESULT.txt)
- **Tested in a real n8n 2.41.7 on our machine, no internet:** the file imports as shipped and exports back the same ([test/n8n-export.json](test/n8n-export.json)). Ten runs with invented messages reached the right place, including no click in time, the AI API overloaded (the fraud message stayed urgent) and Slack refusing an urgent alert (sent by email). The Claude and n8n API steps talked to local stand-ins; Slack, Google Sheets, email and payments were test stubs. A repeated message stopped at the repeat check, and the morning check caught a failed run and an unreadable n8n API. [test/N8N-RUN.txt](test/N8N-RUN.txt), [test/n8n-run.log](test/n8n-run.log)
- **Not tested yet:** live Claude, Slack (buttons and the 4-hour limit in a running n8n), Google Sheets, email and payments. Claude's accuracy on real messages needs real messages and an API key.

## Live test with Claude

**Not run yet.** The tests above use hand-written stand-in answers. This one sends the 12 invented messages to the real Claude API, exactly as the workflow builds the request, runs each answer through the workflow's own rules, and compares the category, the route and the rule with the expected ones. Two expected routes come from stand-in answers that are wrong on purpose (an invented fee, and an answer cut off halfway), so for those two messages only the category is compared.

Run it from the `support-triage` folder:

```
# Put your Anthropic API key in a local .env file that is never committed (ANTHROPIC_API_KEY=...).
node --env-file=.env test/live-eval.mjs --run --max-usd 1.00
```

- Without `--run` it sends nothing. It shows the plan and the estimated cost: about 0.11 to 0.43 USD for the 12 messages, and up to 0.64 USD if every answer came from a fallback model.
- `--max-usd` is a spending limit, 1.00 USD by default. Before each call the script sets aside the most one attempt of it can cost (the whole max_tokens at the fallback price, input counted high: up to 0.14 USD here) and stops if that would pass the limit. A call that times out or loses the connection counts as its full reservation and is not retried. The one way to go over: when Anthropic answers with a fallback model it can bill two attempts for one call, so the total can pass the limit by up to one reservation.
- It writes `test/LIVE-RESULT.txt` (what matched, money spent, and for each message the model that answered and every attempt Anthropic reports) and `test/live-responses.json` (Claude's raw answers). The API key is never printed or written to a file.

## For your technical team

| Path | What it is |
|---|---|
| `workflow.json` | n8n workflow, import from file. Every credential and outside address is a placeholder named `REPLACE_ME_...` (only the Anthropic API address is real). Starts in draft mode (`send_mode` in the Settings node). |
| `prompts/` | Instructions to Claude, approved answers, answer format, and the categories and rules in order (`categories.md`). |
| `fixtures/` | 12 invented messages with the expected result, and an invented log for the morning check. |
| `src/` | The rules as plain JavaScript. `build-workflow.mjs` copies them into the workflow, so the tested code is the code that runs. |
| `test/run-tests.mjs` | Offline test; `RESULT.txt` holds the last run. |
| `test/live-eval.mjs` | The live check against the real Claude API. Not run yet. Without `--run` it only shows the plan and the cost estimate. |
| `test/run-in-n8n.mjs` | Imports the workflow into a real n8n and runs test copies with `n8n execute`, using `local-stand-ins.cjs` on 127.0.0.1. Writes `N8N-RUN.txt`, `n8n-run.log` and `n8n-export.json`. |
| `images/`, `test/render_images.py`, `test/approval-card.cjs` | The two pictures above and how they were made. |

Notes for the build: repeats are dropped by n8n's Remove Duplicates node (up to 10,000 ids, cleared every morning). The approval is the Slack node's "send and wait" with link buttons, so whoever clicks must reach your n8n address. Set execution pruning for failed runs. The log tab grows by about 1,600 rows a month and the morning check reads all of it; if that gets slow, start a new tab.
