Uncategorized

AI Product Unit Economics: Calculating Cost per Successful Task

Multiple servers mounted vertically in an equipment rack.

Quick answer

Calculate cost per successful task by dividing all delivery costs for a defined workload by the number of tasks that meet your acceptance criteria. Include model inference, retrieval, tools, retries, hosting, human review and delivery-related support—not just tokens.

Track customer acquisition and general business expenses separately. Use the resulting delivery cost as a pricing constraint, not as proof that customers will pay your proposed price.

A founder’s useful pricing question is not “What does one model response cost?” It is “What does it cost us to deliver one acceptable customer outcome?”

Below is a practical costing method and a hypothetical spreadsheet example. Provider prices were checked on October 4, 2026. All workload volumes, labor rates, success rates and infrastructure budgets in the example are explicit assumptions, not observed product results.

1. Define a successful task before calculating its cost

Choose a unit that matches the customer’s job: an invoice processed, a support ticket resolved, or a document summary accepted.

Do not equate an API response with a completed task. For a hypothetical document-summary product, define success as:

  • The summary includes the required fields.
  • Its factual statements are supported by the source document.
  • It passes the agreed review process.
  • It reaches the customer within the promised delivery window.

These are proposed acceptance criteria, not a universal standard. Adapt them to your product’s risks and customer expectations.

Keep three counts separate:

  1. Submitted tasks: distinct customer jobs received.
  2. Model attempts: calls made while processing those jobs.
  3. Successful tasks: jobs that ultimately pass acceptance.

One task might require an initial generation, a validation call and a rewritten answer. That is one submitted task, three model attempts and—if accepted—one success.

Use:

Cost per successful task =
Total delivery cost / Successful tasks

Also report automated successes separately from human-assisted successes. Otherwise, an apparent improvement in product quality could conceal additional manual work.

2. Build a delivery-cost ledger beyond inference

Create a spreadsheet row for each cost category, its billing unit and the activity that drives it.

Cost category What to record Suggested allocation
Model inference Input, cached input and output tokens by model Actual usage per task
Retrieval Search calls, indexing and storage Calls per task; storage by customer or workload
External tools Paid searches, third-party API calls or compute sessions Actual billed activity
Retries and validation Additional model and tool calls Attach to the original task
Application hosting Compute, databases, queues, logs and network charges Metered use plus an explicit shared-cost allocation
Human review Review time and loaded labor rate Minutes per reviewed task
Delivery support Troubleshooting and correcting failed outputs Time attributable to the workload

For a concrete billing example, OpenAI lists Responses API file-search calls at $2.50 per 1,000 calls, with storage at $0.10 per GB per day and the first GB free. Tokens used by built-in tools are also billed at the selected model’s rates. Retrieval therefore should not be modeled as a single token expense. (developers.openai.com)

Tool definitions and results can also enlarge the model bill. Anthropic documents that tool-use requests include tokens from tool descriptions, schemas and results, while some server-side tools carry additional usage charges. (platform.claude.com)

Keep sales commissions, advertising, general product development and administrative expenses outside this delivery ledger. Maintain a separate business-expense view so that a positive delivery margin is not mistaken for company profitability.

This is a management costing model, not a prescription for financial-statement classification.

Multiple servers mounted vertically in an equipment rack.
Rack-mounted servers illustrate the infrastructure layer to include in an AI delivery-cost model. — Abigor. Own work Source CC BY-SA 3.0

3. Translate provider billing rules into spreadsheet inputs

Record the provider, model, service tier, currency, pricing date and any applicable contract terms beside each rate.

For the example, use GPT-4.1 mini’s documented prices:

Token category Price per million tokens
Input $0.40
Cached input $0.10
Output $1.60

These are published model rates, not a recommendation that this model suits every task. (developers.openai.com)

With no caching, calculate:

Inference cost per attempt =
(Input tokens × Input rate
 + Output tokens × Output rate) / 1,000,000

With caching, split input into uncached and cached tokens. Do not charge the entire input at the ordinary rate and then add cached-input charges again.

Do not assume caching rules transfer between providers. Anthropic, for example, documents separate cache-write and cache-read pricing, including write multipliers that depend on cache duration. (platform.claude.com)

Finally, distinguish quality retries from transport errors. A completed answer rejected by your application can still have consumed billable tokens. Do not assume every unsuccessful request is billed identically; reconcile reported usage against the provider invoice.

For retry controls, OpenAI recommends limiting both attempts and total retry time, accounting for SDK retries, and using appropriate backoff. Its documentation also notes that unsuccessful requests contribute to per-minute limits—this is a rate-limit rule, not a statement that every error incurs a token charge. (developers.openai.com)

4. Work through a hypothetical monthly workload

Assume an AI product drafts source-grounded document summaries.

The following inputs are invented planning assumptions, not measurements or market benchmarks:

Input Assumption
Submitted tasks per month 10,000
Average model attempts per task 1.2
Input tokens per attempt, including retrieved context 3,000
Output tokens per attempt 600
File-search calls per attempt 1
Final task success rate 90%
Tasks receiving human review 10%
Review time per reviewed task 2 minutes
Loaded review/support labor rate $30/hour
Delivery-support cases 40
Time per support case 10 minutes
Shared application-hosting budget $120/month
Billable file-search storage, after the free allowance 2 GB for 30 days

Assume every modeled attempt completes and incurs the stated token and search usage. There is no caching, no separate validation model and no additional paid tool.

Using the documented rates above:

Inference per attempt =
(3,000 × $0.40 + 600 × $1.60) / 1,000,000
= $0.00216

Monthly attempts = 10,000 × 1.2 = 12,000
Monthly inference = 12,000 × $0.00216 = $25.92
Monthly file-search calls = 12,000 × $0.0025 = $30.00

The complete monthly ledger becomes:

Delivery expense Calculation Monthly cost
Inference 12,000 × $0.00216 $25.92
File-search calls 12,000 × $0.0025 $30.00
Billable storage 2 × 30 × $0.10 $6.00
Application hosting Assumed budget $120.00
Human review 1,000 × 2 ÷ 60 × $30 $1,000.00
Delivery support 40 × 10 ÷ 60 × $30 $200.00
Total $1,381.92

Successful tasks equal:

10,000 × 90% = 9,000

Cost per successful task =
$1,381.92 / 9,000
= approximately $0.154

By comparison, inference alone costs $0.00288 per successful task in this scenario.

The difference does not establish that human review dominates every AI product. It shows why this hypothetical product should investigate review requirements before concentrating exclusively on token discounts.

5. Stress-test request length, failures and customer usage

Change one assumption at a time first. Then build a combined downside scenario.

The following calculations retain the example’s other assumptions:

Scenario Monthly delivery cost Successful tasks Cost per success
Baseline $1,381.92 9,000 $0.154
Double input and output tokens per attempt $1,407.84 9,000 $0.156
Increase attempts per task to 1.5 $1,395.90 9,000 $0.155
Reduce final success rate to 75% $1,381.92 7,500 $0.184
Increase human-review share to 20% $2,381.92 9,000 $0.265

The retry scenario assumes the final success rate remains unchanged. That isolates the cost of extra attempts; it is not a prediction that retries preserve quality.

Similarly, doubling tokens assumes no effect on success or review time. Test those relationships rather than treating them as known.

For customer-usage sensitivity, treat inference, searches, review and support as proportional to submitted tasks. Hold the $126 hosting-and-storage allocation fixed:

Submitted tasks Monthly delivery cost Cost per success at 90%
1,000 $251.59 $0.280
10,000 $1,381.92 $0.154
20,000 $2,637.84 $0.147

This scaling assumption is deliberately simple. Real hosting capacity, storage and staffing may change with volume.

Build a downside case combining longer requests, lower acceptance and greater review demand. Inspect customer-level costs as well as the product-wide average: a subscription must accommodate its actual usage distribution, not merely its average customer.

6. Use delivery cost to constrain pricing—not determine value

Choose a target delivery contribution margin explicitly. There is no universal target implied by this example.

For an illustrative 70% margin on revenue after modeled delivery costs:

Required revenue per successful task =
Cost per successful task / (1 − Target margin)

= ($1,381.92 / 9,000) / 0.30
= approximately $0.512

That is a planning threshold under these assumptions—not a recommended selling price. It excludes the separate business expenses discussed earlier.

Match the calculation to the billing unit:

  • Per successful task: define acceptance, exclusions and dispute handling.
  • Per submitted task: disclose what constitutes a submission and how failures are handled.
  • Subscription: model included usage, expensive workflows and support commitments.
  • Hybrid: consider a base fee for shared service costs plus metered usage.

If charging per submission, use cost per submission in the pricing calculation. Here that is $1,381.92 divided by 10,000, or about $0.138—not the $0.154 cost per success.

Before offering unrestricted usage, decide which limits protect delivery economics: document size, workflow steps, paid-tool activity or included review. Make those limits understandable to customers.

7. Check the model against production evidence

Use this checklist before treating the spreadsheet as a pricing foundation:

  • Define acceptance: document the quality and delivery conditions required for success.
  • Connect costs to tasks: attach model calls, tools and reviews to the original customer job.
  • Include rejected work: keep delivery costs from failed tasks in the numerator.
  • Avoid double counting: retry costs already captured in call logs need no additional blanket surcharge.
  • Separate token categories: distinguish ordinary input, cached input and output.
  • Time human work: measure review and correction rather than estimating from memory.
  • Reconcile invoices: explain differences between internal estimates and billed usage.
  • Refresh assumptions: rerun the model after changes to pricing, prompts, models, tools or acceptance criteria.

Common mistakes include counting “response received” as success, spreading shared costs over an optimistic future volume, and treating reduced review as a saving without checking quality.

Your next step is straightforward: build the ledger, run a representative pilot, and replace each assumption with observed usage where possible. Keep unknowns visible. A useful unit-economics model explains what drives delivery cost—and how that cost changes when the product fails, customers use it differently, or the workflow grows.