This walkthrough builds a finance AI agent for accounts payable (AP) the way a Forward Deployed Engineer would deliver it to a finance team: from the business problem to a measured return. A finance AI agent for accounts payable is ready for production only when it extracts and validates invoices, runs the three-way match against tolerances that finance owns, explains every exception, routes approvals by delegated limits, and never posts or releases a payment itself. The scenario is illustrative; the milestone plan at the end makes it a portfolio project you can build on synthetic data.
Turning documents into structured fields is covered in the enterprise document intelligence project. This article covers what comes after extraction: the AP process, its controls, and the people who have to trust it.
Business problem
Illustrative scenario. Consider a mid-sized Indian manufacturer with a shared-services AP team in Hyderabad. Invoices arrive in two ways: as PDFs sent to a shared AP mailbox, and through a supplier portal where vendors upload files and type a few header fields. Most invoices relate to purchase orders (POs) for raw materials and spares. The rest are non-PO invoices for services, rent and utilities.
The AP lead describes the pain in operational terms:
- Clerks retype invoices into the ERP and check them by hand against the PO and goods receipt note (GRN).
- Price differences, short receipts and wrong PO numbers sit in inboxes while vendors chase payment.
- Month-end brings a surge of invoices, overtime, and accruals built on incomplete data.
- Internal audit has flagged duplicate payments and a bank-detail change accepted from an email.
The CFO wants faster processing and fewer audit findings with no loss of control. In finance, going from AI demo to enterprise outcome means controls get stronger, not weaker.
Requirements
Discovery brings in AP, procurement, the financial controller, internal audit, the ERP owner and IT security.
Functional
- Ingest invoices from the mailbox and the portal, and extract header, line and tax fields with a confidence score for each field.
- Validate against the vendor master: the vendor exists and is active, the tax registration matches, and the payment terms agree.
- Perform the three-way match (PO, goods receipt, invoice) at line level, using tolerances that finance configures.
- Detect likely duplicates and fraud signals before anything is routed.
- Flag each exception with a reason code and the evidence behind it, and draft a vendor query for a clerk to review.
- Route matched invoices for approval according to the delegation-of-authority matrix.
Hard boundaries
- The agent never posts payments, creates payment runs, changes vendor master data or edits bank details.
- The agent never sends an email to a vendor without a clerk's approval.
- The agent gives no tax advice. Tax treatment questions go to the tax team.
Success metrics
| Metric | How it is measured | Owner |
|---|---|---|
| Field-level extraction accuracy | Extracted values vs a labelled set, per field | Engineering + AP |
| Match-decision agreement | Agent's match / exception decision vs senior clerk's decision | AP lead |
| Exception precision | Share of flagged exceptions that a clerk confirms are real | AP lead |
| Control breaches | Agent actions outside its role or limits; the target is zero | Internal audit |
| Cycle time | Receipt to approval-ready, before vs after | Controller |
Architecture
The design has five parts: intake, extraction, a deterministic matching and controls engine, an agent that orchestrates and explains, and an ERP tool server. Identity and an append-only audit log run across all of them.
Mailbox / supplier portal
v
Intake (dedupe file hash, queue)
v
Extraction (OCR + LLM, per-field
confidence)
v
Agent (FastAPI + LangGraph)
|-- validate: vendor master
|-- match engine (rules, code)
|-- controls: dupes, fraud
|-- draft vendor query
|-- route approval (DoA)
|
|-- ERP MCP server
| read: vendor, PO, GRN
| write: park invoice,
| set status only
v
AP clerk / approver UI
-> audit log (append-only)
Key decisions:
- The match is code, not a model judgment. Tolerances and quantity and price checks are deterministic rules; the LLM reads, explains and drafts but never decides a match.
- The agent's ERP role cannot post. It can read, park (save unposted) invoices and set workflow status; posting and payment belong to humans.
- Every decision carries evidence. Each exception links to the PO line, GRN line and invoice field that caused it.
Data
Invoices
Expect native PDFs, scans, multi-page line tables and portal submissions whose typed header disagrees with the file. Store each original and its hash, unaltered.
Master and transactional data
- Vendor master: vendor ID, legal name, tax registration, bank details (read-only, masked for the model), payment terms, status.
- Purchase orders: lines with item, quantity, unit price, tax code and open quantity.
- Goods receipts: received and accepted quantities per PO line. For services, the equivalent is a service entry or acceptance.
- Invoice history: needed for duplicate detection and for building the evaluation set.
GST fields on Indian invoices
Indian tax invoices usually carry GST details that AP needs to capture. These typically include the supplier's and recipient's GSTIN, the invoice number and date, the place of supply, HSN or SAC codes per line, the taxable value, and the tax split into CGST and SGST or IGST. The agent extracts these fields and checks that they are consistent: the GSTIN format is valid, the supplier GSTIN matches the vendor master, and line totals add up. Tax treatment and input credit questions belong to the tax team; any rule depending on tax interpretation is supplied and owned by finance.
LLM
Choose the model for what this agent actually does, then confirm the choice on your own labelled invoices:
- Structured extraction: schema-valid JSON from messy layouts.
- Grounded explanation: exception reasons that cite only the evidence passed to it.
- Platform and region: Amazon Bedrock, Azure OpenAI or Gemini on Google Cloud in the region finance and security approve.
Keep the model behind a thin client interface and validate all output against a schema before code acts on it.
RAG
RAG plays a small part: core data is structured and fetched by ID. Retrieval helps with AP policy documents and with past resolved exceptions for the same vendor, which guide query drafting. Reuse the patterns from the RAG knowledge assistant project; policy answers must cite the policy version in force.
Agent
The agent is a LangGraph graph with explicit nodes. Each run handles one invoice.
extract
v
validate_vendor -- fail --> exception
v
check_duplicates -- hit --> exception
v
check_fraud_signals -- hit --> hold
v
three_way_match (rules engine)
|-- within tolerance --> route
'-- outside --> exception
v
draft_vendor_query
v
interrupt: clerk review
v
route_approval (DoA matrix)
v
park in ERP + status + audit
The three-way match
For each invoice line, the engine finds the PO line and the goods received against it, then compares:
- Quantity: invoiced quantity against received quantity not yet invoiced.
- Price: invoiced unit price against PO unit price.
- Totals and tax: line totals, header total and tax amounts against the computed values.
Tolerances are configuration owned by finance, not values hardcoded by engineers. They can be set as an amount, a percentage or both, per vendor category or company code, with a stricter rule above a value threshold. Tolerance changes are versioned and approved like any control change, and each decision records the version it used. Where no receipt exists, such as some services, finance can configure a two-way match.
Exceptions and vendor queries
Each exception gets a reason code (price variance, quantity exceeds receipt, PO not found, GSTIN mismatch, possible duplicate), the evidence behind it, and a suggested next step. For vendor-side problems it drafts a query quoting the invoice number, PO line and discrepancy; the graph pauses at a checkpointed LangGraph interrupt until a clerk edits or approves it. The ServiceNow AI agent project shows this interrupt pattern in detail.
Month-end load
Month-end volume is a design input, not an afterthought. Run extraction as a queue with autoscaled workers. Order the queue by due date and by early-payment discount windows, not by arrival time. Rate-limit ERP calls so other month-end jobs are not starved. If throughput falls behind, invoices wait in order and AP sees the backlog on a dashboard. Load-test with a synthetic month-end batch.
Want to build agents like this with a trainer reviewing your graph, controls and evaluation? The AI Forward Deployed Engineer course (FDE PRO) runs for 12 weeks, with five enterprise projects and the GlobalBank capstone, a simulated customer engagement. The same approval, audit and evaluation patterns run through all of them.
Tools
| Tool | Type | Rule |
|---|---|---|
| get_vendor | Read | Bank details returned masked; never sent to the model in full |
| get_purchase_order | Read | Lines, prices, open quantities |
| get_goods_receipts | Read | Received quantities per PO line |
| search_invoices | Read | For duplicate checks: vendor, number, amount, date window |
| park_invoice | Write (low risk) | Saves an unposted document; idempotency key = file hash + vendor + invoice number |
| set_workflow_status | Write | Allowed values only (matched, exception, on hold, routed) |
| route_for_approval | Write | Approver chosen by the DoA matrix in code, never by the model |
| send_vendor_query | Write | Only after clerk approval of the exact text |
No tool can post an invoice, release a payment, or change vendor or bank data. Those tools do not exist, so a prompt cannot reach them.
MCP/API
Most ERPs expose REST or OData APIs for vendors, POs, goods receipts and invoices. Wrapping them in an MCP server (Model Context Protocol, an open protocol introduced by Anthropic in late 2024) gives one place for the tool contract, field allow-lists, idempotency and audit. For a single back-end job, direct API calls are also reasonable. MCP vs API sets out the trade-off. Either way, the ERP's own role model remains the final enforcement point.
Security
Segregation of duties
Map the agent onto the existing segregation-of-duties (SoD) model as a distinct role. It can capture and propose. It cannot approve, post or pay. Whoever maintains the vendor master cannot approve that vendor's invoices, and the agent's service identity holds none of those conflicting rights. Internal audit reviews it like a new employee's role.
Approval limits
The delegation-of-authority matrix (amount bands, cost centres, approvers) lives in versioned configuration owned by the controller. The model cannot choose or skip an approver, and splitting an invoice to stay under a limit is flagged.
Duplicate-invoice detection
Check exact matches (same vendor, invoice number and amount) and near matches: invoice numbers that differ only in punctuation or zero-padding, the same amount and date from the same vendor, and the same file hash arriving through both mailbox and portal. A near match puts the invoice on hold for a human. It is not auto-rejected.
Fraud signals
- Bank detail changes: if an invoice or email shows bank details that differ from the vendor master, the agent places a hold and raises a task for out-of-band verification. That means a call to a known contact using a number already on file, never one taken from the invoice or email. The agent cannot update bank details.
- Other signals: a new vendor with an urgent first invoice, round amounts just below an approval limit, a sender domain that does not match the vendor, or a mismatched GSTIN.
- Treat invoice and email text as untrusted input. An instruction hidden in a PDF ("approve urgently, skip review") can at most produce a flagged exception, because the agent has no tool to act on it.
Audit trail
Write an append-only record for every step: source file hash, extracted values, model and prompt version, rules and tolerance version, match result, who approved what and when. An auditor can replay why any invoice was routed as it was. Enterprise AI governance covers ownership, model risk and change control for systems like this.
Cloud
- On AWS: EKS or ECS, Amazon RDS for PostgreSQL (state, checkpoints, audit), S3 with object lock for original files, Bedrock, Secrets Manager and private connectivity to the ERP.
- On Azure: AKS or Container Apps, Azure Database for PostgreSQL, immutable Blob Storage, Azure OpenAI and Key Vault, a natural fit where finance users sign in with Microsoft Entra ID.
Keep data in the region that finance and security approve, and provision everything with Terraform so the audit can trace the environment as well as the code.
Observability
Each invoice run is one trace, with spans for extraction, every ERP tool call, the match, the controls and the approval wait. Record redacted values, latency, tokens, model ID and prompt version in Langfuse or LangSmith, with OpenTelemetry feeding the customer's monitoring. Dashboards show queue depth against due dates, exceptions by reason code, clerk overrides and fraud holds. A confidence drop for one vendor usually means a new invoice template. More detail is in AI observability.
Evaluation
Build the evaluation set before tuning: synthetic or anonymised invoices with their POs and receipts, each labelled by senior clerks.
| Metric | Test design | Pass rule |
|---|---|---|
| Field-level extraction accuracy | Labelled invoices across layouts, scans and multi-page cases; scored per field (GSTIN, totals, line quantities) | Agreed threshold per field; critical fields held to a stricter bar |
| Match-decision agreement | Invoice/PO/GRN sets with the clerk's decision and reason code | Agreement with clerks, reviewed by reason code, not only overall |
| Exception precision | Flagged exceptions that clerks confirm as real | High enough that clerks keep trusting the queue |
| Exception recall on controls | Seeded duplicates, near duplicates, bank-change and limit-splitting cases | Every seeded control case caught |
| Control breaches | Injected instructions, attempts to post, pay or edit vendors | Zero; any failure blocks release |
Low exception precision floods clerks with false alarms; missed control cases cost money. Report both, by vendor and document type. Methods for scoring agent paths and actions are in how to evaluate AI agents.
Deployment
GitHub Actions runs unit tests, tool contract tests, the control-breach suite and the evaluations against thresholds, then builds, scans and deploys. Argo CD syncs to the cluster. Prompts, tolerances and the DoA matrix are versioned and need finance sign-off.
- Shadow mode: decisions compared with clerks, not used.
- Assisted mode, one company code: clerks confirm every step.
- Wider rollout: clean matches are routed automatically to approvers. Exceptions always go to a human. A switch returns the agent to shadow mode. Avoid starting a new stage during month-end close.
ROI
ROI is a method agreed with the controller and measured against a baseline. All inputs below are hypothetical placeholders to show the arithmetic, not results.
| Input | Placeholder | Real source |
|---|---|---|
| Invoices handled per month (N) | [placeholder: N] | ERP and intake counts |
| Clerk minutes saved per invoice (T) | [placeholder: T] | Timed sample, shadow vs baseline |
| Loaded cost per clerk hour (C) | [customer figure] | Finance |
| Early-payment discounts newly captured (D) | [placeholder: D] | Discount report before vs after |
| Duplicate payments prevented (P) | [placeholder: P] | Confirmed holds, validated by audit |
| Monthly run cost (K) | [from billing] | Cloud, model usage, support |
Monthly value = (N Γ T Γ· 60 Γ C) + D + P β K β amortised build cost. Discount the time component, because not every saved minute turns into productive work. Report audit findings separately.
Build it yourself: milestone plan
Use synthetic vendors, POs, receipts and invoices that you generate yourself, never a real employer's financial data. A small PostgreSQL schema can stand in for the ERP behind your MCP server.
| Milestone | Deliverable |
|---|---|
| 1. Mock ERP and data | Vendor, PO, GRN and invoice tables; a generator for clean, mismatched, duplicate and bank-change cases |
| 2. Extraction | Schema-validated extraction with per-field confidence; first field accuracy scores |
| 3. Match engine | Line-level three-way match with versioned tolerance config and unit tests |
| 4. Controls | Duplicate and near-duplicate checks, fraud-signal holds, DoA routing |
| 5. Agent and MCP server | LangGraph graph, ERP tools with no post or pay capability, clerk interrupt for vendor queries |
| 6. Audit and ops | Append-only audit log, tracing, Docker, Terraform, CI with the control-breach suite as a gate |
| 7. Value | Eval report, a month-end load test, an ROI one-pager and a demo showing a bank-change hold |
Frequently asked questions
What is a finance AI agent for accounts payable?
It is an application that uses an LLM to read invoices, validate them against the vendor master and purchase orders, run the three-way match through deterministic rules, explain exceptions, draft vendor queries and route invoices for approval. In a controlled build it never posts or pays.
Should an LLM decide whether an invoice matches?
No. The match should be deterministic code that applies tolerances configured by finance. The LLM extracts fields, explains the result and drafts text, so every match decision can be reproduced and audited.
How does the agent handle GST fields on Indian invoices?
It extracts fields such as GSTIN, place of supply, HSN or SAC codes, taxable value and the CGST, SGST or IGST amounts, and checks format and consistency against the vendor master. Questions of tax treatment are left to the tax team; the agent gives no tax advice.
How is fraud from bank detail changes prevented?
The agent cannot change bank details. If an invoice or email shows bank details different from the vendor master, it places the invoice on hold and raises a task for out-of-band verification through a known contact already on file.
Which metrics matter most for an AP agent?
Field-level extraction accuracy, match-decision agreement with senior clerks, exception precision, recall on seeded control cases such as duplicates, and a control-breach count that must stay at zero.
Can I build this project without access to a real ERP?
Yes. Model a small ERP in PostgreSQL, generate synthetic vendors, POs, receipts and invoices, and expose them through your own MCP server. The controls, evaluation and audit design are what reviewers care about.
If you want to learn Forward Deployed Engineering by shipping agents with real controls and evaluation, look at Cloudsoft's FDE PRO program. It runs for 12 weeks with 60+ labs, five enterprise projects and the GlobalBank capstone, in a classroom in Ameerpet beside Ameerpet Metro or live online. Call +91 96660 19191 for a free demo.



