{
  "question_id": "CG-P1B-FULL-132",
  "slug": "how-the-handling-of-accounting-documents-is-automated",
  "display_title": "How is the handling of a business's accounting documents automated?",
  "format": "article-v2",
  "applies_to": {
    "countries": [
      "US"
    ],
    "frameworks": [],
    "tax_year": null,
    "platforms": []
  },
  "general_concept": true,
  "summary": "Automation splits document handling into steps, each with its own mechanism: capture routes, recognition and extraction, rule- or model-based classification, integration with the books, workflow routing and archiving. Extracted values can look right and be wrong, so they must be validated before posting; documents the flow cannot process must go to a queue a person works; and every automated action must be set up to leave a record, because a platform may log only what you tell it to.",
  "body": "## What steps does handling a document break into?\n\nEach step is automated, and fails, separately. This table sets out the steps in a typical order, though classification often runs on extracted text.\n\n| Step | Typical mechanism | What stays with a person |\n|---|---|---|\n| Arrival and capture | Intake email address, upload, scanner, phone camera | Choosing which channels are accepted |\n| Identification and classification | Rules on sender or text, or a trained model | Documents nothing recognizes |\n| Extraction of names, dates, numbers and amounts | Optical character recognition (OCR) and field extraction | Fields flagged as uncertain |\n| Validation | Arithmetic, duplicate and record checks; confidence thresholds | Deciding what a failed check means |\n| Filing and indexing | Naming rules fed by validated fields | The naming and index scheme |\n| Linking to the accounting record | A platform function or an integration | Reviewing and posting the entry |\n| Routing for action | Workflow rules | The action itself, which for bills belongs to the payables workflow |\n| Archiving and retrieval | Archive storage linked from the entry | Chasing entries that have no document |\n\n## What does each mechanism actually do?\n\n### What do capture, recognition and extraction do?\n\nCapture turns whatever arrived into a readable file. Microsoft's Business Central help page on incoming documents tells users to \"Upload the received files, or use your device's camera to take a photo\"; Business Central's OCR help page adds that a recognition service can offer to \"process files forwarded to a dedicated email address\". OCR then converts the image into characters, and extraction assigns them to fields such as sender, date, number and total. Neither proves its own output: Amazon's undated Textract best-practices page describes a confidence score as \"a number between 0 and 100 that indicates the probability that a given prediction is correct\".\n\n### How do rules and trained models classify differently?\n\nA rule applies a written condition, such as a sender or a phrase. Business Central's OCR help page describes a mapping used \"to define that a certain text on a vendor invoice received from the OCR service is mapped to a certain vendor account\". A rule answers the same way every time, so its errors are systematic and one edit fixes them. If a sender's text changes, the rule stops matching and the document should surface as unidentified. The silent failure is a false match: the same page says any part of an incoming document's description that exists as a mapping text fills in that mapping's vendor. Keep mapping texts unique to one sender and review what rules matched.\n\nA trained model predicts from examples, so it can absorb variation among layouts like those it learned from but is likely to err on a sender it has not seen, and its errors are scattered rather than systematic. Business Central's OCR help page says corrections sent back to the recognition service train it for \"the same vendor\".\n\nUse rules for stable, recurring senders, where a wrong answer posts to the wrong account; use a model where layouts vary too widely for rules, behind a confidence threshold that sends doubtful results to a person, expecting more of them from new senders.\n\n## Where can the automation live?\n\nWhere the automation runs changes how the same step behaves.\n\n| Where it lives | How it works | What breaks it | Who maintains it |\n|---|---|---|---|\n| Built into the accounting or document platform | The platform's own functions capture, read and link | Product changes the vendor releases | The vendor, with your settings |\n| A connected service between systems | A separate service processes files and returns results | A stopped job, lapsed connection or change on either side | Both vendors and whoever set up the connection |\n| Scripted or robotic automation | A script operates the existing screens as a person would | Changes to screen elements or values it matches exactly | Whoever wrote the script |\n\nThe paths also differ in what they leave on record: before relying on one for evidence, check that its actions appear in your platform's logs under an identifiable account.\n\nThe paths blur: Business Central's incoming-documents help page says \"The OCR feature is provided by external providers\", so a built-in function can rest on a connected service. Microsoft's Power Automate page on custom selectors says selectors describe \"the exact location of a component in the application or webpage\" and that hard-coded values, though effective in static applications, \"can be a barrier in dynamic applications\"; it suggests operators such as Contains instead.\n\n## What decides accuracy, and what must be checked before posting?\n\nThree things set how accurate identification and extraction will be:\n\n- **The sender's layout.** Business Central's OCR help page says the service is likely to misread characters when it first meets a vendor's documents, and might \"misinterpret the total amount on a receipt because of its layout\".\n- **The image.** The Tesseract OCR project's undated guide to output quality says the quality of its line segmentation \"reduces significantly if a page is too skewed, which severely impacts the quality of the OCR\", and Amazon's undated Textract page asks for \"a high quality image, ideally at least 150 DPI\".\n- **The threshold.** Amazon's Textract page tells applications sensitive to errors to \"enforce a minimum confidence score threshold\" and says they \"should discard results below that threshold or flag situations as requiring a higher level of human scrutiny\"; it adds that \"Business processes involving financial decisions might require thresholds of 90% or higher\".\n\nExtracted values look like verified data, so never post straight from them: a misread total surfaces only when something fails to reconcile. Before posting, run checks that catch recognition errors. The duplicate and date checks themselves use recognized values:\n\n- The line amounts and tax add up to the extracted total.\n- The sender matches an existing vendor or customer record; a new sender goes to a person.\n- No earlier entry or document from the same sender has the same number, or the same amount on a nearby date; any near match goes to a person.\n- The date falls in an open period and is not in the future.\n- A person has reviewed every field flagged as uncertain.\n\nBusiness Central's OCR help page offers that last check as a setting: where the service \"is set up to require manual verification of processed documents\", a person has to \"manually edit or enter values in fields that the OCR service tagged as uncertain\" before the document returns.\n\n### What must be standardized when documents arrive in mixed formats?\n\nWhen documents arrive as emailed PDFs, paper and phone photos, capture is the binding step, because later steps are only as good as the files they receive. Standardize these first:\n\n- One intake email address and one upload point, so nothing arrives unseen\n- One scan resolution, at or above what your recognition service documents\n- Photos taken flat, straight and evenly lit, one document per image\n- Digital originals passed on as received, never printed and rescanned\n\n## Which steps still need a person, and why?\n\nThree kinds of step resist automation by their nature, not because nobody has automated them yet:\n\n- **Judgment.** Whether a document belongs to the business, what a cost was for and which account and period it belongs to depend on facts outside the document. A rule can repeat a decision a person made once, as a text-to-account mapping does, but cannot make the first one or notice when circumstances change.\n- **Accountability.** An automated step cannot answer for an entry; a person has to. NIST's AI Risk Management Framework, in its appendix C, lists among issues meriting further consideration and research that human roles and responsibilities in decision making and overseeing AI systems \"need to be clearly defined and differentiated\". Beyond that, for rules as well as models, give each automated step a named owner and keep posting a decision a named person takes or signs off.\n- **An exception rate that never reaches zero.** Textract's confidence score only indicates the probability that a prediction is correct, so results above a threshold can still be wrong, and at the 90% or higher that Amazon says processes involving financial decisions might require, a flow keeps sending documents to a person. New senders, credits and one-off documents also keep arriving.\n\n## How does a handled document reach the books?\n\nA flow that files documents but never touches the ledger has automated storage, not bookkeeping. Microsoft's Business Central help shows two ways a document reaches the books, and your own platform's current documentation says which it supports:\n\n- **It becomes the entry.** Business Central's OCR help page says a recognized document becomes a record that can be converted into other records, among them \"a sales invoice, credit memo, or a journal entry\", and that text can be mapped to a bank account, which suits documents \"for expenses that are already paid\".\n- **It attaches to an entry.** Business Central's incoming-documents help page says files \"can be attached at any process stage, including to posted documents and to the resulting vendor, customer, and general ledger entries\".\n\nThe tie has to run both ways and be tested. The same incoming-documents page says users can view document records \"from any purchase and sales document or entry\" and \"find all general ledger entries without incoming document records\". It adds that files on a card's or document's Attachments tab \"aren't included on the Incoming Documents page\", so attach evidence as incoming document records if you rely on this search. Run it every period; it finds only entries without a document. For the reverse, check that each document marked processed in the period links to a posted entry, using whatever list or export your platform provides, and send any without one back to the exception queue.\n\n## What stops documents moving, and which exceptions must surface?\n\nA flow runs end to end only while every hand-over happens unattended; Business Central's OCR help page, for example, has files sent \"by the job queue according to the schedule, if no errors exist\". Chains stop when a job or connection stops, a step waits on a person nobody assigned, something upstream changes, such as a sender's layout or a screen value a script matches exactly, or a failed document has no route onward.\n\nDocuments the flow cannot process must surface rather than be filed out of the way, because a failure moved to a folder nobody opens is undetectable. Each exception class needs its own handling:\n\n| Exception | How it shows up | Handling it needs |\n|---|---|---|\n| Unreadable | Blank or garbled fields, low confidence on key fields | Back to capture for a rescan or clean copy; never marked processed |\n| Unidentified | No rule, model class or record matches the sender or type | A person assigns it; a recurring sender gets a rule |\n| Duplicate | Same sender with the same number, or the same amount on a nearby date, already recorded; or one file arriving twice | Held and compared with the existing entry; removed only once confirmed |\n| Mismatched | Amounts do not add up or disagree with a related record | A person resolves the difference before posting |\n| Out of scope | Not an accounting document, or not this business's | A person decides; it leaves with a note, never a silent delete |\n\nGive every document a status and every exception queue an owner, and check daily how many documents sit in a status that is not final. Business Central's auditing-changes help page says that from some pages users can view \"an activity log that shows the status and any errors from files that you export from or import into Business Central\", and that they can empty it or clear entries older than seven days. Keep exception history on the document record or in an export, and limit who can clear the log.\n\n## What changes for controls and the audit trail?\n\nAutomation replaces seeing each document with a record. NIST's computer-security glossary defines an audit trail as \"A chronological record that reconstructs and examines the sequence of activities surrounding or leading to a specific operation, procedure, or event in a security relevant transaction from inception to final result\". For a document, the equivalent is its path from arrival to entry. For each document, keep:\n\n- The time it arrived and the channel it came by\n- What each automated step produced, including class, values and confidence\n- The rule, mapping or model setting that acted on it\n- Any change a person made, with who made it and when\n- The entry it produced or supports\n\nDo not assume the platform keeps this unprompted. Business Central's auditing-changes help page says its change log tracks \"all direct modifications a user makes to data in the database\" and that \"You specify each table and field that you want the system to log, and then you activate the change log\". Check what your system logs and whether automated actions appear under an identifiable account.\n\nThe automation's settings are now controls. The PCAOB's auditing standard AS 2110 is written for auditors working under the Board's standards, but its examples of risks from information technology illustrate the point: \"Reliance on systems or programs that are inaccurately processing data, processing inaccurate data, or both\", \"Unauthorized changes to systems or programs\", \"Failure to make necessary changes to systems or programs\" and \"Inappropriate manual intervention\". So limit who can edit rules, mappings and thresholds, log every edit and override, and have someone other than the editor review those logs where staffing allows; a sole owner reviews them on a fixed schedule.\n\nFlows also drift. NIST's AI Risk Management Framework, in its appendix B, says AI systems \"may require more frequent maintenance and triggers for conducting corrective maintenance due to data, model, or concept drift\". Each month, look for a rising exception rate, more corrections for one sender, entries without documents and unexplained reconciliation differences.\n\nAsk any auditor, lender or funder early what they will require. Keep enough to reconstruct each document's path to its entry and back: the original file, corrected values, rule-change history and exception history.\n\n## How does one document move through an automated flow?\n\nA monthly internet bill, already paid by automatic debit from the business checking account, moves through these steps, each marked automatic or manual:\n\n1. **Capture (automatic).** The emailed PDF is forwarded to the recognition service's intake address, and a document record is created with the file attached.\n2. **Extraction (automatic).** The service reads the provider, invoice number NF-20931, a subtotal of 180.00, tax of 14.40 and a total of 184.40, flagging the total as uncertain.\n3. **Validation (automatic check, manual fix).** Because 180.00 plus 14.40 is 194.40, the arithmetic check fails and the bill goes to the exception queue. The bookkeeper reads 194.40 on the image, corrects the field and sends the correction back; the log records who changed it and when.\n4. **Classification (automatic, by rule).** A text-to-account mapping on the provider's name assigns the internet expense account and, because the bill is paid, the checking account.\n5. **Duplicate check (automatic).** The flow looks for anything from this provider with the same number, or the same amount on a nearby date, including a bank-statement entry for the automatic debit.\n6. **Linking (automatic draft, manual posting).** If the debit is already recorded, the file is attached to that entry and nothing new is posted. If not, a journal line is created with the file attached, and the bookkeeper reviews and posts it.\n7. **Filing and retrieval (automatic).** The file is indexed by provider, date, number and amount and marked processed, and opening the ledger entry opens the bill.\n\nWhere the debit was not yet recorded, step 6 posts this entry:\n\n| Account | Debit | Credit |\n|---|---|---|\n| Internet expense | 194.40 | |\n| Business checking | | 194.40 |\n\nIf the entry was posted from the bill, reconciliation later matches the bank debit to it instead of recording it again.\n\n## How do you assess your own flow and choose what to automate first?\n\nMap the flow as it runs today before changing any tool. For each type of document, work through this checklist:\n\n- **Arrival.** List every channel documents come in by, from inboxes and paper mail to phone photos and portals, with a monthly count for each.\n- **Waiting.** Note where documents wait for a person and for how long; long waits at rule-governed steps mark the best candidates.\n- **Repetition.** Mark steps done the same way every time by a rule you could write down; steps decided case by case stay with a person.\n- **Prerequisites.** For each candidate, name what must be true first, such as one intake channel, a scan standard, vendor and account records, and an exception owner.\n- **Exceptions.** Estimate how many documents a month would fail each step, in which class, and who would work them.\n- **Link to the books.** Check your platform's current documentation for whether a document can become an entry or attach to one, and whether unlinked entries can be listed.\n- **Evidence.** If an auditor, lender or funder reviews your controls, list the record each automated step must leave.\n- **Existing automation.** Note what already runs automatically; the next candidate is where documents now pile up.\n\nLow volume may justify automating only intake and naming; higher volume justifies extraction and rules, and brings more exceptions to work.\n\nAutomate in dependency order:\n\n1. Fix any step that already produces errors by hand, because automating it repeats the defect faster and removes the person who caught it.\n2. Standardize capture.\n3. Connect to the books, and start the periodic checks for unlinked entries and documents.\n4. Add extraction only together with validation checks and an exception queue.\n5. Add rules for stable senders, and a model only where layouts vary too much for rules.\n6. Automate filing, indexing and archiving from validated fields.\n7. Add routing last, handing any approval or payment step to the payables workflow.",
  "sources": [
    {
      "id": "REF::1",
      "url": "https://learn.microsoft.com/en-us/dynamics365/business-central/across-income-documents",
      "title": "Work with incoming documents (Business Central)",
      "publisher": "Microsoft",
      "published": "2025-10-16",
      "retrieved_at": "2026-09-27T14:48:54+00:00",
      "sha256": "f5f7ccacfd688bd5d71efb0aa44ab648c9b33d29f6b747a9bc27f24f48059748",
      "supports": [
        "C2",
        "C8",
        "C23",
        "C24",
        "C25",
        "C42"
      ]
    },
    {
      "id": "REF::2",
      "url": "https://learn.microsoft.com/en-us/dynamics365/business-central/across-how-use-ocr-pdf-images-files",
      "title": "Use OCR to turn PDF into e-invoices (Business Central)",
      "publisher": "Microsoft",
      "published": "2025-10-16",
      "retrieved_at": "2026-09-27T14:48:54+00:00",
      "sha256": "037cc17641e229eaadf4f7ee7b30ce47f56c4cf0c2f35321d28b2809330f3e9b",
      "supports": [
        "C3",
        "C4",
        "C6",
        "C7",
        "C11",
        "C12",
        "C18",
        "C19",
        "C21",
        "C22",
        "C26",
        "C37"
      ]
    },
    {
      "id": "REF::3",
      "url": "https://learn.microsoft.com/en-us/dynamics365/business-central/across-log-changes",
      "title": "Auditing changes (Business Central)",
      "publisher": "Microsoft",
      "published": "2026-06-17",
      "retrieved_at": "2026-09-27T14:48:54+00:00",
      "sha256": "f30ddb72bfd06aa096210a0af7da1d8542919649503f8c61ee685f45113c7c8b",
      "supports": [
        "C28",
        "C30",
        "C31",
        "C43"
      ]
    },
    {
      "id": "REF::4",
      "url": "https://learn.microsoft.com/en-us/power-automate/desktop-flows/build-custom-selectors",
      "title": "Build a custom selector (Power Automate)",
      "publisher": "Microsoft",
      "published": "2025-05-12",
      "retrieved_at": "2026-09-27T14:48:55+00:00",
      "sha256": "49dc8f4fb0314fcab1bb553fec7d7e2191a04f4582f0701f93aac7f2d71a8112",
      "supports": [
        "C9",
        "C10",
        "C38"
      ]
    },
    {
      "id": "REF::5",
      "url": "https://docs.aws.amazon.com/textract/latest/dg/textract-best-practices.html",
      "title": "Best Practices (Amazon Textract Developer Guide)",
      "publisher": "Amazon Web Services",
      "published": "undated",
      "retrieved_at": "2026-09-27T14:48:55+00:00",
      "sha256": "a2038e56c7141ad1dce44fbc031ed04b9b41152ea8c9bde843b6450b44aad9f6",
      "supports": [
        "C5",
        "C14",
        "C15",
        "C16",
        "C17",
        "C40",
        "C41"
      ]
    },
    {
      "id": "REF::6",
      "url": "https://tesseract-ocr.github.io/tessdoc/ImproveQuality.html",
      "title": "Improving the quality of the output",
      "publisher": "Tesseract OCR project",
      "published": "undated",
      "retrieved_at": "2026-09-27T14:48:55+00:00",
      "sha256": "deb71aacb8052a8bdd3c7472cd4745aef407a306687169fea8efecbe930554fb",
      "supports": [
        "C13"
      ]
    },
    {
      "id": "REF::7",
      "url": "https://airc.nist.gov/airmf-resources/airmf/appendices/app-b-how-ai-risks-differ-from-traditional-software-risks/",
      "title": "AI RMF 1.0, Appendix B: How AI Risks Differ from Traditional Software Risks",
      "publisher": "National Institute of Standards and Technology",
      "published": "AI RMF 1.0 (2023)",
      "retrieved_at": "2026-09-27T14:48:56+00:00",
      "sha256": "b6475d66436e25de47ec8c488989a53f21854ad26108faf090688af97caabfce",
      "supports": [
        "C36"
      ]
    },
    {
      "id": "REF::8",
      "url": "https://airc.nist.gov/airmf-resources/airmf/appendices/app-c-ai-risk-management-and-human-ai-interaction/",
      "title": "AI RMF 1.0, Appendix C: AI Risk Management and Human-AI Interaction",
      "publisher": "National Institute of Standards and Technology",
      "published": "AI RMF 1.0 (2023)",
      "retrieved_at": "2026-09-27T14:49:18+00:00",
      "sha256": "7d519cdde7663fb2bce45dfa4d8d7a30f672f7f9d32e479995894839de18adc7",
      "supports": [
        "C20",
        "C39"
      ]
    },
    {
      "id": "REF::9",
      "url": "https://pcaobus.org/oversight/standards/auditing-standards/details/AS2110",
      "title": "AS 2110: Identifying and Assessing Risks of Material Misstatement",
      "publisher": "Public Company Accounting Oversight Board",
      "published": "undated web page",
      "retrieved_at": "2026-09-27T14:49:40+00:00",
      "sha256": "59713e7ad96d0bcb0f2a9d4b15f12606fd0da837051ef12bb4e8df775d4c8190",
      "supports": [
        "C32",
        "C33",
        "C34",
        "C35",
        "C44"
      ]
    },
    {
      "id": "REF::10",
      "url": "https://csrc.nist.gov/glossary/term/audit_trail",
      "title": "audit trail (Glossary)",
      "publisher": "National Institute of Standards and Technology, Computer Security Resource Center",
      "published": "undated",
      "retrieved_at": "2026-09-27T14:49:40+00:00",
      "sha256": "5a96415b847360f426a4395cdde265299e81cc3f34c79036f84756bc3d5decac",
      "supports": [
        "C29"
      ]
    }
  ],
  "related": [
    {
      "question_id": "CG-P1B-FULL-001",
      "slug": "can-bookkeeping-be-automated-and-how-to-choose-a-tool-or-service",
      "display_title": "Can my bookkeeping be automated — what do automated bookkeeping software and services actually do, and how do I choose one?"
    },
    {
      "question_id": "CG-P1B-006",
      "slug": "how-to-choose-software-to-manage-and-store-accounting-documents",
      "display_title": "What software should a small business use to manage and store its accounting documents (invoices, receipts, statements)?"
    },
    {
      "question_id": "CG-P1B-009",
      "slug": "how-a-small-business-can-manage-and-automate-its-accounts-payable-workflow",
      "display_title": "How should a small business manage and automate its accounts-payable vendor-invoice and document workflow?"
    },
    {
      "question_id": "CG-P1B-FULL-070",
      "slug": "how-robotic-process-automation-works-in-accounts-payable",
      "display_title": "How is robotic process automation applied to the accounts payable process, and what does it do there?"
    }
  ],
  "review_class": "consequential",
  "review_class_trigger": "claim_level_review_required",
  "provenance": {
    "author_model": "claude-opus-5-5",
    "reviewer_model": "claude-opus-5-5",
    "review_verdict": "ACCEPT",
    "review_source": "closure",
    "review_verdict_on_sha256": "b1f3a45558c4eeaa8f498ea0d9b933465e1d1729606bd0afa125e6617a96ad46",
    "editorial_disposition": "ACCEPT",
    "corrections": 1,
    "approved_by": null,
    "approved_at": null,
    "article_sha256": "b1f3a45558c4eeaa8f498ea0d9b933465e1d1729606bd0afa125e6617a96ad46",
    "source_map_sha256": "c8ffe899a6354de3e2fa6adf3c84d82935310e3e6a1b05cdc79c793824e28282",
    "transform_sha256": "4cc39bd6f8dd432445d5f52ad491a9a663ce598305809b6f81952d5f2b3a9998"
  },
  "offer": "ask",
  "offer_id": null,
  "sample_target_id": null,
  "datePublished": "2026-09-27T21:10:04Z",
  "reviewed_at": "2026-09-27T21:10:04Z",
  "content_sha": "ce1f923507f99ed6447a937f97e1e1ae9ef6df9e124f6a622ba8b8cdd260f4ed",
  "release": "2.11.0",
  "slug_provenance": "minted at first publication",
  "question_text": "How is the handling of a business's accounting documents automated?",
  "jsonld_types": [
    "Article"
  ],
  "related_question_ids": [
    "CG-P1B-FULL-001",
    "CG-P1B-006",
    "Q-2003",
    "CG-P1B-009",
    "CG-P1B-FULL-070"
  ],
  "aliases": [],
  "alias_provenance": []
}
