Quick answer
Calculate cost per successful task by dividing all delivery costs for a defined workload by the number of tasks that meet your acceptance criteria. Include model inference, retrieval, tools, retries, hosting, human review and delivery-related support—not just tokens.
Track customer acquisition and general business expenses separately. Use the resulting delivery cost as a pricing constraint, not as proof that customers will pay your proposed price.
A founder’s useful pricing question is not “What does one model response cost?” It is “What does it cost us to deliver one acceptable customer outcome?”
Below is a practical costing method and a hypothetical spreadsheet example. Provider prices were checked on October 4, 2026. All workload volumes, labor rates, success rates and infrastructure budgets in the example are explicit assumptions, not observed product results.
1. Define a successful task before calculating its cost
Choose a unit that matches the customer’s job: an invoice processed, a support ticket resolved, or a document summary accepted.
Do not equate an API response with a completed task. For a hypothetical document-summary product, define success as:
- The summary includes the required fields.
- Its factual statements are supported by the source document.
- It passes the agreed review process.
- It reaches the customer within the promised delivery window.
These are proposed acceptance criteria, not a universal standard. Adapt them to your product’s risks and customer expectations.
Keep three counts separate:
- Submitted tasks: distinct customer jobs received.
- Model attempts: calls made while processing those jobs.
- Successful tasks: jobs that ultimately pass acceptance.
One task might require an initial generation, a validation call and a rewritten answer. That is one submitted task, three model attempts and—if accepted—one success.
Use:
Cost per successful task =
Total delivery cost / Successful tasks
Also report automated successes separately from human-assisted successes. Otherwise, an apparent improvement in product quality could conceal additional manual work.
2. Build a delivery-cost ledger beyond inference
Create a spreadsheet row for each cost category, its billing unit and the activity that drives it.
| Cost category | What to record | Suggested allocation |
|---|---|---|
| Model inference | Input, cached input and output tokens by model | Actual usage per task |
| Retrieval | Search calls, indexing and storage | Calls per task; storage by customer or workload |
| External tools | Paid searches, third-party API calls or compute sessions | Actual billed activity |
| Retries and validation | Additional model and tool calls | Attach to the original task |
| Application hosting | Compute, databases, queues, logs and network charges | Metered use plus an explicit shared-cost allocation |
| Human review | Review time and loaded labor rate | Minutes per reviewed task |
| Delivery support | Troubleshooting and correcting failed outputs | Time attributable to the workload |
For a concrete billing example, OpenAI lists Responses API file-search calls at $2.50 per 1,000 calls, with storage at $0.10 per GB per day and the first GB free. Tokens used by built-in tools are also billed at the selected model’s rates. Retrieval therefore should not be modeled as a single token expense. (developers.openai.com)
Tool definitions and results can also enlarge the model bill. Anthropic documents that tool-use requests include tokens from tool descriptions, schemas and results, while some server-side tools carry additional usage charges. (platform.claude.com)
Keep sales commissions, advertising, general product development and administrative expenses outside this delivery ledger. Maintain a separate business-expense view so that a positive delivery margin is not mistaken for company profitability.
This is a management costing model, not a prescription for financial-statement classification.

3. Translate provider billing rules into spreadsheet inputs
Record the provider, model, service tier, currency, pricing date and any applicable contract terms beside each rate.
For the example, use GPT-4.1 mini’s documented prices:
| Token category | Price per million tokens |
|---|---|
| Input | $0.40 |
| Cached input | $0.10 |
| Output | $1.60 |
These are published model rates, not a recommendation that this model suits every task. (developers.openai.com)
With no caching, calculate:
Inference cost per attempt =
(Input tokens × Input rate
+ Output tokens × Output rate) / 1,000,000
With caching, split input into uncached and cached tokens. Do not charge the entire input at the ordinary rate and then add cached-input charges again.
Do not assume caching rules transfer between providers. Anthropic, for example, documents separate cache-write and cache-read pricing, including write multipliers that depend on cache duration. (platform.claude.com)
Finally, distinguish quality retries from transport errors. A completed answer rejected by your application can still have consumed billable tokens. Do not assume every unsuccessful request is billed identically; reconcile reported usage against the provider invoice.
For retry controls, OpenAI recommends limiting both attempts and total retry time, accounting for SDK retries, and using appropriate backoff. Its documentation also notes that unsuccessful requests contribute to per-minute limits—this is a rate-limit rule, not a statement that every error incurs a token charge. (developers.openai.com)
4. Work through a hypothetical monthly workload
Assume an AI product drafts source-grounded document summaries.
The following inputs are invented planning assumptions, not measurements or market benchmarks:
| Input | Assumption |
|---|---|
| Submitted tasks per month | 10,000 |
| Average model attempts per task | 1.2 |
| Input tokens per attempt, including retrieved context | 3,000 |
| Output tokens per attempt | 600 |
| File-search calls per attempt | 1 |
| Final task success rate | 90% |
| Tasks receiving human review | 10% |
| Review time per reviewed task | 2 minutes |
| Loaded review/support labor rate | $30/hour |
| Delivery-support cases | 40 |
| Time per support case | 10 minutes |
| Shared application-hosting budget | $120/month |
| Billable file-search storage, after the free allowance | 2 GB for 30 days |
Assume every modeled attempt completes and incurs the stated token and search usage. There is no caching, no separate validation model and no additional paid tool.
Using the documented rates above:
Inference per attempt =
(3,000 × $0.40 + 600 × $1.60) / 1,000,000
= $0.00216
Monthly attempts = 10,000 × 1.2 = 12,000
Monthly inference = 12,000 × $0.00216 = $25.92
Monthly file-search calls = 12,000 × $0.0025 = $30.00
The complete monthly ledger becomes:
| Delivery expense | Calculation | Monthly cost |
|---|---|---|
| Inference | 12,000 × $0.00216 | $25.92 |
| File-search calls | 12,000 × $0.0025 | $30.00 |
| Billable storage | 2 × 30 × $0.10 | $6.00 |
| Application hosting | Assumed budget | $120.00 |
| Human review | 1,000 × 2 ÷ 60 × $30 | $1,000.00 |
| Delivery support | 40 × 10 ÷ 60 × $30 | $200.00 |
| Total | $1,381.92 |
Successful tasks equal:
10,000 × 90% = 9,000
Cost per successful task =
$1,381.92 / 9,000
= approximately $0.154
By comparison, inference alone costs $0.00288 per successful task in this scenario.
The difference does not establish that human review dominates every AI product. It shows why this hypothetical product should investigate review requirements before concentrating exclusively on token discounts.
5. Stress-test request length, failures and customer usage
Change one assumption at a time first. Then build a combined downside scenario.
The following calculations retain the example’s other assumptions:
| Scenario | Monthly delivery cost | Successful tasks | Cost per success |
|---|---|---|---|
| Baseline | $1,381.92 | 9,000 | $0.154 |
| Double input and output tokens per attempt | $1,407.84 | 9,000 | $0.156 |
| Increase attempts per task to 1.5 | $1,395.90 | 9,000 | $0.155 |
| Reduce final success rate to 75% | $1,381.92 | 7,500 | $0.184 |
| Increase human-review share to 20% | $2,381.92 | 9,000 | $0.265 |
The retry scenario assumes the final success rate remains unchanged. That isolates the cost of extra attempts; it is not a prediction that retries preserve quality.
Similarly, doubling tokens assumes no effect on success or review time. Test those relationships rather than treating them as known.
For customer-usage sensitivity, treat inference, searches, review and support as proportional to submitted tasks. Hold the $126 hosting-and-storage allocation fixed:
| Submitted tasks | Monthly delivery cost | Cost per success at 90% |
|---|---|---|
| 1,000 | $251.59 | $0.280 |
| 10,000 | $1,381.92 | $0.154 |
| 20,000 | $2,637.84 | $0.147 |
This scaling assumption is deliberately simple. Real hosting capacity, storage and staffing may change with volume.
Build a downside case combining longer requests, lower acceptance and greater review demand. Inspect customer-level costs as well as the product-wide average: a subscription must accommodate its actual usage distribution, not merely its average customer.
6. Use delivery cost to constrain pricing—not determine value
Choose a target delivery contribution margin explicitly. There is no universal target implied by this example.
For an illustrative 70% margin on revenue after modeled delivery costs:
Required revenue per successful task =
Cost per successful task / (1 − Target margin)
= ($1,381.92 / 9,000) / 0.30
= approximately $0.512
That is a planning threshold under these assumptions—not a recommended selling price. It excludes the separate business expenses discussed earlier.
Match the calculation to the billing unit:
- Per successful task: define acceptance, exclusions and dispute handling.
- Per submitted task: disclose what constitutes a submission and how failures are handled.
- Subscription: model included usage, expensive workflows and support commitments.
- Hybrid: consider a base fee for shared service costs plus metered usage.
If charging per submission, use cost per submission in the pricing calculation. Here that is $1,381.92 divided by 10,000, or about $0.138—not the $0.154 cost per success.
Before offering unrestricted usage, decide which limits protect delivery economics: document size, workflow steps, paid-tool activity or included review. Make those limits understandable to customers.
7. Check the model against production evidence
Use this checklist before treating the spreadsheet as a pricing foundation:
- Define acceptance: document the quality and delivery conditions required for success.
- Connect costs to tasks: attach model calls, tools and reviews to the original customer job.
- Include rejected work: keep delivery costs from failed tasks in the numerator.
- Avoid double counting: retry costs already captured in call logs need no additional blanket surcharge.
- Separate token categories: distinguish ordinary input, cached input and output.
- Time human work: measure review and correction rather than estimating from memory.
- Reconcile invoices: explain differences between internal estimates and billed usage.
- Refresh assumptions: rerun the model after changes to pricing, prompts, models, tools or acceptance criteria.
Common mistakes include counting “response received” as success, spreading shared costs over an optimistic future volume, and treating reduced review as a saving without checking quality.
Your next step is straightforward: build the ledger, run a representative pilot, and replace each assumption with observed usage where possible. Keep unknowns visible. A useful unit-economics model explains what drives delivery cost—and how that cost changes when the product fails, customers use it differently, or the workflow grows.