The Automation Project That Did Not Make Sense Six Months Ago. Today It Costs 80% Less.
Cost & ROI
· 8 min read
Author: Lucian Pampu
On 30 July 2026, OpenAI cut one model price by 80%. Google followed on 13 August. The AI price index sits 88% below 2023 levels. Projects that were not viable six months ago now are.
There is a conversation we have had several times over the past year, with different companies, but almost identical every time.
A director describes an internal process: thousands of documents processed monthly, hours of repetitive work, qualified people doing transcription work. We run the numbers. AI processing cost at that volume exceeded the benefit. Conclusion: it does not make sense now. Maybe later.
What Happened in Recent Weeks
On 30 July 2026, OpenAI cut the price of GPT-5.6 Luna by 80% — from $1 to $0.20 per million input tokens, and from $6 to $1.20 per million output tokens. The cut came just three weeks after the model launched. It was not a promotion: it is the new list price, with no expiry date.
On 13 August, Google launched Gemini 3.7 Flash at $0.75 input and $3.75 output per million tokens — roughly half the previous generation.
On 15 August, Anthropic made the promotional pricing of Claude Sonnet 5 permanent, at $2 and $10 per million tokens, cancelling the increase planned for 1 September.
And the longer trend: the price index for frontier models sits 88% below March 2023 levels. GPT-4 launched at $60 per million output tokens; models with comparable capabilities cost under $5 today. A reduction of over 90% in less than three years.
Why This Matters Concretely for a Process in Your Company
A one-page document averages around 800 input tokens. A structured response with extracted data — supplier, amount, VAT, date, number — is around 200 output tokens.
At January 2026 prices, processing 10,000 documents monthly cost somewhere around $20-25 in tokens. That sounds small, but at that price level, any process that also required analysis, validation and longer generated responses quickly reached hundreds of dollars monthly — enough to question the viability of a small project.
At August 2026 prices, the same volume costs between $3 and $5. The same 10,000 documents. The same extraction quality — in fact better, because current models are significantly more capable than those of six months ago.
It is the difference between "does not justify the investment" and "pays for itself in two months."
Which Projects Become Viable Now
The practical rule we apply: if a process has high volume and low value per unit, token pricing was the main obstacle. It no longer is.
- High-volume document processing — invoices, receipts, forms, delivery notes, transport documents.
- Automatic classification and routing of communication — thousands of emails monthly sorted to the right departments, with relevant context extracted.
- Continuous monitoring and analysis — daily checks across large data sets (orders, inventory, transactions) to detect anomalies.
- Enriching existing databases — cleaning, standardising and completing CRM or ERP data across tens or hundreds of thousands of records.
- Internal documentation support — an assistant answering the team questions from company documents, around the clock, at a monthly cost comparable to a daily coffee.
The Rule We Apply and Recommend
Model pricing now moves on a timescale of weeks, not years. Any cost assumption locked in during the first half of 2026 is probably wrong by now — in both directions.
"In both directions" matters. Not all prices are falling. DeepSeek, which triggered part of this competition with aggressive pricing, had to raise its rates on 16 August because demand outstripped its infrastructure capacity.
Why Model-Agnostic Architecture Matters More Than Ever
Models can increasingly be swapped behind standardised interfaces. A properly built system does not depend on a specific vendor. The model is a replaceable component, not the foundation.
- You can change vendor in a configuration, not in a refactor. If the model you use today becomes three times more expensive in four months, you migrate in hours.
- You can use different models for different tasks — a fast, cheap one for simple classification; a more powerful one for complex analysis.
- You can test alternatives without risk — run a new model on part of your traffic and compare results, without committing the whole company.
Companies that built systems rigidly tied to a single vendor in 2025 are paying that cost now. Those that built model-agnostic have benefited from every price cut automatically.
What We Do at Visual AI Labs
- Model-agnostic architecture from the start — the model is configurable, not hard-wired into the code.
- Intelligent task routing — each task goes to the model that fits it on price and capability.
- Real cost monitoring — you see exactly what each automated process costs monthly.
Plus everything that makes an implementation work: your own company context, direct integrations into existing systems, repeatable workflows and governance with complete audit trails. We deliver fast and without replacing what already works for you.
If You Postponed a Project on Cost Grounds
It is worth recalculating. The figures on which you said "not now" are almost certainly out of date. Write to us on the contact page with a short description of the process you had in mind. We will come back with an updated estimate — at today prices, not those from six months ago.
Write to us on the contact page →
Sources: CNBC (30 Jul. 2026), VentureBeat — AI price wars (30 Jul. 2026), Neowin — Google joins the AI model price war (13 Aug. 2026), Developers Digest — Frontier Model API Pricing (verified 15 Aug. 2026), BenchLM.ai — LLM API Pricing Trends (19 Aug. 2026), XenoSpectrum (Aug. 2026).