An AI app costs roughly $4,000–$30,000 to build when the AI is a single feature on top of a model API, $18,000–$60,000 for retrieval (RAG) over your own documents, $15,000–$75,000 for a single tool-using agent, and $100,000+ for multi-agent workflow platforms. The bigger swing is the monthly bill: at October 2026 list prices, the same 1,000 user requests cost $0.35 on the cheapest frontier-lab model and $1,240 on the most expensive set-up below. Architecture, not the model, decides which end of that range you land on.
This guide gives build and run numbers per architecture, using token prices taken from the Anthropic, OpenAI and Google pricing pages on 7 October 2026. If you want the price of an MVP package, see our fixed-price MVP cost page. For agents specifically, read how much an AI agent costs to build. For the long-run operating bill, see the real cost of running an AI product.
It depends mostly on which of six architectures you need. These are BeevR's planning ranges for a custom build by a senior team. They are not quotes, and onshore US rates will sit higher. The run column uses the mid-priced tier (Claude Sonnet 5.5 or GPT-6 Sol, both $2 per million input tokens and $10 per million output) with no caching, at the usage levels listed further down.
| Architecture | What it is | Typical build cost | Typical timeline | Token bill at 10K monthly users (mid tier, no caching) |
|---|---|---|---|---|
| AI feature via API | One model call per action: summarise, classify, extract, draft | $4K–$30K | 2–6 weeks | ~$2,100/mo |
| RAG over your documents | Search your data, then answer with citations | $18K–$60K | 6–10 weeks | ~$5,100/mo |
| GraphRAG or on-device AI | Entity graph over your data, or the model runs in the user's browser or device | $30K–$100K | 8–14 weeks | Cloud GraphRAG: similar to RAG plus indexing; on-device: $0 in tokens |
| Single tool-using agent | A model plans steps and calls your APIs or MCP tools | $15K–$75K | 6–12 weeks | ~$15,000/mo |
| Multi-agent workflow | Several agents hand work to each other across a long process | $100K–$200K+ | 3–6 months | ~$12,400/mo at low task volume |
| Voice agent | Speech in, speech out, usually with tools behind it | Agent cost plus telephony and latency work | 8–14 weeks | Priced per minute (see FAQ) |
Two rules hold across the table. First, regulated data moves the build number more than the architecture does. Our MVP cost page puts compliance at +15–25% when you design it in, against +40–80% when you bolt it on later. Second, every architecture below the first row needs evals: a test set that tells you whether answers got worse after you changed a model or a prompt. Teams that skip evals pay for it later, and that is one of the patterns behind why AI agent projects fail.

Frontier prices now run from $0.10 to $10 per million input tokens, a 100x spread. These are standard list prices per million tokens, checked on each vendor's official pricing page on 7 October 2026:
| Model | Input | Output | Cached input / cache read | Note |
|---|---|---|---|---|
| OpenAI GPT-6 Luna | $0.10 | $0.50 | $0.01 | Cheapest current GPT-6 tier |
| Google Gemini 3.8 Flash | $0.75 | $3.75 | $0.075 | Price runs to 31 Dec 2026; $1.50 / $7.50 from 1 Jan 2027 |
| Anthropic Claude Haiku 4.5 | $1 | $5 | $0.10 | Anthropic's small model |
| Anthropic Claude Sonnet 5.5 | $2 | $10 | $0.20 | Mid tier |
| OpenAI GPT-6 Sol / GPT-6.1 Sol | $2 | $10 | $0.20 / $0.10 | Mid tier |
| Google Gemini 3.1 Pro Preview | $2 (≤200K prompt) | $12 (≤200K prompt) | n/a here | $4 / $18 above 200K tokens |
| Anthropic Claude Opus 5.5 | $4 | $20 | $0.20 | Cache read is 5% of input |
| Anthropic Claude Fable 5.1 | $10 | $50 | $0.25 | Top tier |
| OpenAI GPT-6 Astra | $10 | $50 | $1.00 | Top tier |
Three details change real bills more than the headline numbers. Batch processing halves the price for work that can wait (Anthropic lists a 50% Batch API discount on input and output). Anthropic notes that Claude 4.7 and later models use a tokenizer that produces roughly 30% more tokens for the same text, so compare cost per task rather than price per token. And Gemini 3.8 Flash doubles on 1 January 2027, so a business case built on today's Flash price needs a 2027 line.

From under $1 to over $1,000, depending on architecture and model. The table multiplies the list prices above by typical token counts per request: 1,500 input and 400 output tokens for a simple AI feature, 6,000 and 500 for a RAG answer, 60,000 and 3,000 for a single agent task (around eight model calls as context grows), and 250,000 and 12,000 for a multi-agent workflow. No caching or batch discounts are applied.
| Per 1,000 requests | GPT-6 Luna | Gemini 3.8 Flash (2026 price) | Sonnet 5.5 / GPT-6 Sol | Opus 5.5 |
|---|---|---|---|---|
| AI feature | $0.35 | $2.63 | $7.00 | $14.00 |
| RAG answer | $0.85 | $6.38 | $17.00 | $34.00 |
| Single agent task | $7.50 | $56.25 | $150.00 | $300.00 |
| Multi-agent workflow | $31.00 | $232.50 | $620.00 | $1,240.00 |
Your own token counts will differ, so swap in the numbers from your logs. The shape is the point: an agent task costs about 20 times a RAG answer on the same model, because the agent re-reads its growing context on every step.

Monthly token spend grows in line with users, so a design choice that looks trivial at 1,000 users becomes a payroll-sized line at 100,000. Assumptions: 30 AI-feature or RAG requests per active user per month, 10 agent tasks, or 2 multi-agent workflows, on the mid tier ($2 / $10), no caching.
| Monthly token bill | 1,000 users | 10,000 users | 100,000 users |
|---|---|---|---|
| AI feature (30 requests/user) | $210 | $2,100 | $21,000 |
| RAG chat (30 answers/user) | $510 | $5,100 | $51,000 |
| Single agent (10 tasks/user) | $1,500 | $15,000 | $150,000 |
| Multi-agent (2 workflows/user) | $1,240 | $12,400 | $124,000 |
Tokens are only part of the run cost. Hosting, vector storage, monitoring, eval runs and human review sit on top, and our post on AI running costs covers those lines.
Five levers change the token bill more than switching vendors does:
The model call is the cheap part. The money goes into everything around it:
Run this checklist before you ask anyone for a quote. Any studio worth hiring will want the same answers.
BeevR publishes three fixed-price packages, and the AI architecture decides which one fits. From our pricing page: the Pitch Demo at $4K (about 10 days) covers one core workflow on real infrastructure. The Investor MVP at $18K (about 6 weeks) covers 3–5 core workflows with auth and role-based access, tested to survive due diligence. The Flagship Sprint at $38K (about 10 weeks) covers 5–8 workflows with a full test suite, load testing, automated deploy and rollback, and full observability.
As a rule of thumb, a single AI feature or one RAG flow fits the shape of a Pitch Demo. A product whose core is RAG or one agent, plus the surrounding app, fits the Investor MVP. Multi-agent workflows, voice, on-device AI and regulated data usually push into a Flagship Sprint or a phased plan. The exact fit depends on scope, so treat this as a starting point for the conversation, not a quote.
We keep our AI work inspectable, so you can check how we handle the cost and safety levers above:
A chatbot that answers from your own documents is a RAG app: plan on $18,000–$60,000 for a production build, less for a demo. At the mid model tier it costs about $17 per 1,000 answers in tokens before caching, or under $1 per 1,000 on GPT-6 Luna if answer quality holds up on your eval set.
Usually yes, when the AI is one feature that calls a model API: $4,000–$30,000 is typical, because auth, data and UI already exist. It gets expensive again when the feature needs access to data the app wasn't built to expose, or needs to act on the user's behalf.
On OpenAI's gpt-realtime-2.1 ($32 per million audio input tokens, $64 per million audio output), audio is billed at 1 token per 100 ms of user speech and 1 per 50 ms of model speech. That makes about $0.02 per minute of caller audio and $0.08 per minute of agent audio, before the conversation history that is re-billed each turn, tools and telephony. The mini model costs roughly a third of that.
Per token, self-hosting can be cheaper at high, steady volume. Below that, the GPU, ops and eval work usually cost more than API tokens, and $0.10-per-million API tiers have narrowed the gap. Running a small model on the user's device is the case where open weights clearly win on cost and privacy.
The trend has been down for most tiers, but not in a straight line. Some prices rise when promotions end, as with Gemini 3.8 Flash on 1 January 2027. Build so you can swap models behind an eval set, and re-price your bill every quarter.
BeevR is a senior, founder-led AI studio in Hanoi, Vietnam. We work at a fixed price per phase, you own the code and IP from day one, and we build AI for regulated industries. Compare the packages on our MVP cost page, or tell us what you're building. We'll name the architecture, estimate the monthly token bill with you, and give you a fixed number for the build.