AI for Small Business: Six Boring Wins and Three Ways to Waste Money
Francis Okafor
On this page
Most of what gets written about AI for small business is written by people selling AI to small businesses. That produces a strange genre of advice, one in which a two-person freight forwarding office in Lagos and a 400-person contract manufacturer in Dongguan are addressed as the same buyer. The useful version is narrower. About six tasks reliably pay for themselves. Three mistakes reliably cost money. Both lists have held steady for roughly two years.
Start with the honest baseline. The US Census Bureau's Business Trends and Outlook Survey asks a deliberately narrow question: did this firm use AI to produce goods or services in the last two weeks. As of 3 May 2026, 19.8% of US businesses said yes. Among firms with 250 or more employees the figure was 37%. Among firms with four or fewer employees it sat below 20%. The US Chamber of Commerce's fourth Empowering Small Business report, published 18 August 2025, found 58% of small businesses saying they use generative AI.
Those numbers do not contradict each other. One counts people who have opened a chatbot. The other counts processes. The gap between 58% and 19.8% is the gap between software being present in a business and software being load-bearing in a business. Most of the money that gets wasted is spent trying to jump that gap by buying something.
Where AI for small business actually pays
The reliable wins share a shape. The task repeats. The input has a known form. The output has a known form. And a named person currently spends an afternoon on it every week.
Quoting and proposal generation. A quote is a template, a price lookup and a judgement about the customer. The first two are mechanical. The third is not. The firms that win here generate the draft and keep a human on the number.
Customer response drafting. Drafting, not sending. The model writes, a person approves. That single distinction is the whole risk profile, and it is the reason drafting works while autoreply keeps producing incidents.
Document extraction. Invoices, packing lists, bills of lading, spec sheets. Google's Enterprise Document OCR is listed at $1.50 per 1,000 pages. PaddleOCR shipped version 3.7.0 on 11 June 2026 with PP-OCRv6, whose medium tier is 34.5 million parameters and reports a 5.1% recognition gain over the previous server model, and it runs on a laptop CPU for nothing.
Bookkeeping preparation. Preparation, not bookkeeping. Categorising transactions, matching receipts to line items, flagging the eleven entries that do not fit a pattern. Your accountant still signs.
Translation for cross-border trade. That one deserves its own section.
Inventory and demand signals. The weakest of the six and the most oversold. It works when you have three years of clean sales history against a stable SKU list. Most small businesses have neither, and the honest answer there is to fix the record-keeping first.
The token cost of any of this rounds to zero. DeepSeek lists deepseek-v4-flash at $0.22 per million input tokens on a cache miss and $0.66 per million output tokens at off-peak rates, with peak hours priced at double. A full supplier quotation, extracted and translated in both directions, is a few thousand tokens. The expense is never the inference. It is the process design, and that is exactly where the failures live.
The dominant document format in this trade is not PDF. It is a screenshot of an Excel sheet pasted into a WeChat message.
Document work for China trade is the case I would fund first
China and Africa traded $348.05 billion in goods in 2025, a record, up 17.7% on 2024 according to China's General Administration of Customs. Chinese exports accounted for $225.03 billion of that, up 25.8%. African exports to China were $123.02 billion, up 5.4%.
From 1 May 2026 China applies zero tariffs to all 53 African countries with which it holds diplomatic relations, extending to 20 non-LDC states an arrangement that has covered 33 least developed countries since 1 December 2024. Cocoa, coffee, citrus and tea had been facing 8 to 30 percent. The tariff went to zero. The paperwork did not.
Eight years in Shenzhen and the thing I still find hardest to explain to buyers abroad is the file format problem. The dominant document format in this trade is not PDF. It is a screenshot of an Excel sheet pasted into a WeChat message. Merged cells, three price tiers, a footnote in Chinese about minimum order quantity and a model number that differs from the one on the proforma invoice by a single character. Any extraction pipeline that assumes clean PDFs will fail in its first real week.
Then there is vocabulary that costs money. 纯铜 is pure copper. 铜包铝 is copper-clad aluminium. A translation that renders both as "copper wire" is not wrong in any way a general benchmark would flag, and it is the difference between cable that meets spec and a container of scrap you cannot sell. I have watched that exact confusion play out inside a purchase order. The buyer read the English. The English was fluent. The English was wrong.
The machine translation research community has said this out loud. The findings paper for the WMT25 general translation task carries the subtitle "Time to Stop Evaluating on Easy Test Sets". Thirty authors, 60 systems evaluated including 24 pulled from commercial LLMs and online translation providers, and the headline methodological change was building deliberately harder test sets because the old ones no longer separated strong systems from weak ones. General translation is good enough to be boring. Your terminology is not general.
Which is why the controls matter more than the model choice. Alibaba Cloud's Qwen-MT, fine-tuned from Qwen3, covers 92 languages in its plus, flash and turbo tiers, and it accepts three things a small trader actually has: a terminology glossary, a translation memory of previously approved sentence pairs and a domain hint. The glossary is the asset. The model underneath is a commodity that will be replaced twice a year.
The failure mode is old and documented. In March 2012 the Shanghai Maritime Court told China Daily that translation errors in contract terms, while under 5% of its caseload, produced avoidable losses. In one shipping contract "drydocking" had been rendered as "tank washing" and "except fuel used for domestic service" as "except fuel used for domestic flights". Humans made those errors. A model with an empty glossary makes the same class of error faster and across every document at once.
Three ways owners burn the budget
Buying a platform before defining a process. In July 2025 MIT Media Lab's Project NANDA published "The GenAI Divide: State of AI in Business 2025" and reported that 95% of enterprise generative AI pilots produced no measurable impact on profit and loss. The figure travelled further than the methodology deserved: 52 executive interviews, 153 survey responses, 300 public deployments, no peer review. But the sharpest criticism of that study is also its most useful finding. A large share of those pilots had no documented pre-deployment baseline. They could not demonstrate impact because nobody had written down what the task cost beforehand.
Automating a broken workflow. If quoting takes six days because the price list lives in one salesperson's head and he travels three weeks a month, no model fixes that. You will get wrong quotes in four seconds instead of correct quotes in six days. Automation is a multiplier. Apply a multiplier to a negative number and the number gets worse.
Paying per seat for something used twice a week. Microsoft 365 Copilot Business is listed at $21 per user per month billed yearly, discounted to $18 for eligible existing Business customers through 31 December 2026, and the enterprise add-on is $30. Those sit on top of a base licence. In a team of twelve where four people use it daily and eight open it when they remember, the eight cost $2,016 a year. Zylo's 2026 SaaS Management Index, built on more than 40 million licences and $75 billion in tracked spend, found organisations leave 36% of SaaS licences unused on average. ChatGPT is now the most expensed application in that dataset, and expense-based SaaS spending rose 267% year over year, which means a large share of AI spend never passes procurement at all. It arrives on somebody's personal card and gets reimbursed.
The method, in four steps
Find the repeated task. Not the most annoying task, the most repeated one. Have two people log a week in a shared sheet: what they did, how long it took, how many times. Anything appearing five or more times gets a line. The list will surprise you. It is almost never the thing the owner complains about at dinner.
Measure what it costs now. Minutes and errors, both, written down before anything changes. Twenty-two quotes last month, average forty minutes each, three sent with the wrong freight term. That is a baseline. Without one you will never know whether the software helped, and you will end up inside the 95%.
Automate the narrow version. Not "AI for quoting". Instead: extract these nine fields from supplier PDFs arriving from these four suppliers, into this spreadsheet, in this format. Narrow enough that you can list every failure mode on one page. Narrow enough that a competent person builds the first version in two days with an API key and no platform.
Verify for a month before trusting it. Run it in parallel. The person still does the work. The machine does it too, and someone compares every output for thirty days. This is dull, and it is the entire value of the exercise, because what you are measuring is not whether the model is clever. You are measuring your own error rate against its error rate on your actual documents, including the ugly ones, the handwritten ones and the one from the supplier who scans at an angle.
At the end of the month you have a number. If the machine is worse, you stop, and you have spent a month and almost no cash. If it is better, you switch and you keep spot-checking one output in twenty, permanently. That spot-check never ends. Anyone who tells you it can end is selling something.
The strongest argument against all of this
The case against incremental automation is real and I have heard it made well.
It runs like this. Careful narrow automation is how you optimise your way into third place. The businesses extracting serious value are not timing their quoting process with a stopwatch. They rebuilt the process. The same MIT report quoted for the 95% figure also found that the small group getting value tended to buy specialised tools from vendors and integrate them deeply into a workflow rather than run cautious internal pilots. Deep integration is the opposite of the narrow version. There is a timing argument stacked on top: spend a month verifying and the model you verified has been superseded, because Qwen, DeepSeek and the rest ship on a cadence measured in weeks.
Both points are correct. Neither changes the method for a business under fifty people.
Deep integration is available to firms that can survive a failed integration. If a rebuild costs two quarters of margin and does not work, a 300-person company absorbs it and writes a memo. A twelve-person importer closes. The narrow version is not a smaller ambition. It is the version where the downside has a floor.
On model churn, the verification month is not verifying the model. It is verifying the process. The glossary you assembled, the nine fields you decided actually matter, the thirty days of logged disagreements between human and machine, all of that survives a model swap and most of it improves when the model does. I have replaced the model underneath production systems several times in the last three years, sometimes because something cheaper appeared and sometimes because the previous one was deprecated with six weeks of notice. The evaluation set is what made each of those swaps a Tuesday afternoon rather than a crisis. And the evaluation set came out of the boring month nobody wants to run.
Both sides of the table
The tooling that makes a Lagos importer competitive against a much larger buyer is the same tooling sitting on the supplier's side of the table in Foshan, and it got there first.
I watch this every week. Factory sales staff who spoke no English in 2019 now answer enquiries in fluent English at eleven at night, in Portuguese, in Arabic, with the product photos already localised. The language gap that an entire layer of agents, sourcing consultants and trading companies was built to bridge is closing from both directions at once.
If your business exists because you can read Chinese and your customer cannot, that business has a shelf life, and it is shorter than the depreciation schedule on your office furniture. If your business exists because you know which of eleven suppliers actually ships to spec in December, when the factory is running three shifts before Chinese New Year and quality quietly slips, none of this touches you. Those are two different businesses. A lot of people currently believe they are in the second one.
Tools referenced
PaddleOCR, reviewed here: PaddleOCR review.
MinerU, reviewed here: MinerU review.
DeepSeek, reviewed here: DeepSeek review.
Dify, reviewed here: Dify review.
Promptfoo, reviewed here: Promptfoo review.
Langfuse, reviewed here: Langfuse review.
Sources
US Census Bureau, Large Firms With at Least 20 Employees Biggest AI Users (May 2026): https://www.census.gov/library/stories/2026/05/ai-use-businesses.html
US Chamber of Commerce, Empowering Small Business report, fourth edition: https://www.uschamber.com/technology/artificial-intelligence/u-s-chambers-latest-empowering-small-business-report-shows-majority-of-businesses-in-all-50-states-are-embracing-ai
Zylo, 2026 SaaS Management Index: https://zylo.com/news/2026-saas-management-index
Microsoft 365 Copilot Business pricing: https://www.microsoft.com/en-us/microsoft-365/copilot/business
DeepSeek API models and pricing: https://api-docs.deepseek.com/quick_start/pricing
NTU Centre for African Studies, China-Africa trade hits record US$348bn: https://www.ntu.edu.sg/cas/news-events/news/detail/china-africa-trade-hits-record-us-348bn-as-deficit-balloons
gov.cn, China implements zero tariffs for African nations with diplomatic ties: https://english.www.gov.cn/policies/policywatch/202605/01/content_WS69f45e35c6d00ca5f9a0ac01.html
Alibaba Cloud Model Studio, Qwen-MT translation model documentation: https://www.alibabacloud.com/help/en/model-studio/machine-translation
Frequently Asked Questions
How much does it cost a small business to start using AI?
Far less than platform pricing suggests, if you start with API access rather than seats. DeepSeek lists deepseek-v4-flash at $0.22 per million input tokens on a cache miss and $0.66 per million output tokens at off-peak rates, and Google's Enterprise Document OCR is listed at $1.50 per 1,000 pages. By contrast, Microsoft 365 Copilot Business is $21 per user per month billed yearly, currently discounted to $18 for eligible existing customers through 31 December 2026. A narrow extraction or drafting workflow for a small team usually costs single-digit dollars a month in tokens. The real cost is two days of setup plus a month of parallel verification, not the software.
What should a small business automate with AI first?
The most repeated task with a known input format and a known output format, not the most annoying task. In practice that means quoting and proposal drafts, customer reply drafts that a person approves before sending, field extraction from invoices and packing lists, bookkeeping preparation before your accountant signs and translation for cross-border trade. Log a week of actual work first. Anything that appears five or more times is a candidate. Anything that appears once is not, however irritating it felt.
Is AI translation good enough for contracts and specifications with Chinese suppliers?
For general correspondence, yes. For contract terms and technical specifications, only with a controlled glossary. General machine translation now scores well enough that the WMT25 shared task had to construct deliberately harder test sets, but that tells you nothing about your product vocabulary. 纯铜 (pure copper) and 铜包铝 (copper-clad aluminium) can both surface as "copper wire" in a perfectly fluent translation, and the price difference is real. Alibaba Cloud's Qwen-MT, covering 92 languages, accepts a terminology glossary, a translation memory and a domain hint for exactly this reason. Keep a bilingual human on anything that creates a legal obligation.