byASB Ankiit Singh Book a scoping week
← All writing/AI integration

AI Integration Cost: What Changes the Estimate

Two illustrative workloads show what belongs in a production estimate beyond model calls.

byASB engineering guides · Published

AI integration cost depends on the workflow around the model: data access, permissions, evaluation, user experience, deployment and operations. The cost of a model call is one line in the estimate. A production feature also has to handle incorrect outputs, outages and changes to the surrounding product.

This guide explains how to scope the work before requesting a quote. The scenarios and numbers below are illustrative planning assumptions dated 5 October 2026. They are not client project prices, current provider rates or a byASB price list. Actual prices need a defined scope and current supplier quotes.

Start with the intended product behavior

Describe the action the feature should help someone complete. “Add AI” leaves too many decisions open. “Help an operator draft a response from approved internal documents, with sources and a human approval step” identifies inputs, outputs and a consequence to test.

Decide what happens when the system cannot produce an acceptable result. Can the user continue manually? Should the feature decline to answer? Can it propose an action, or can it execute one? Expanding from suggestions to actions changes permissions, evaluation and recovery requirements.

Define the boundary of the first release. A single workflow with a named owner can be estimated more reliably than an organization-wide assistant whose users, tools and data sources are still changing. The scope should include acceptance criteria as well as a demonstration.

The main cost drivers

What belongs in the estimate
AreaQuestions that affect effort
Data and accessHow many sources exist? Who may see each record? How are updates and deletions reflected?
Product integrationWhere does the result appear? What can the user inspect, edit or approve?
EvaluationWhich examples represent the task? What failure is unacceptable? Who labels and reviews results?
ConsumptionHow many attempts, tokens, images or processing minutes does one completed task require?
DeploymentIs processing hosted, on-device or hybrid? What connectivity and device limits apply?
OperationsWho monitors failures, responds to changes and maintains the integration?

Separate one-time delivery effort from recurring charges. Then show which items are fixed, usage-dependent or unresolved. A quote can be fixed for a defined deliverable while ongoing provider consumption remains variable. Presenting those as one unexplained number makes later comparisons difficult.

Scenario 1: an internal document assistant

Assume 30 staff, two approved document collections and 20 working days per month. For a planning example, each person completes four assisted tasks per day. That is 30 × 20 × 4 = 2,400 tasks per month. If each task averages 1.5 model attempts because of retries or revisions, the estimate uses 3,600 attempts.

Suppose an average attempt processes 3,000 input tokens and produces 500 output tokens. Under those assumptions, the monthly estimate is 10.8 million input tokens and 1.8 million output tokens. These are workload calculations, not measured usage. A pilot should replace the assumed lengths and attempt rate with observed distributions.

For a supplier charging per million tokens, the model line becomes 10.8 × the input rate + 1.8 × the output rate. Add retrieval, hosting, indexing and any other charged services separately. Check the supplier's billing units and caching or batch rules before using its rates. No provider rate is supplied here because the intended model has not been selected.

The engineering scope also includes source ingestion, permission checks, citations, user feedback and handling stale content. Evaluation needs examples where documents conflict, no answer exists, a user lacks access or the request asks the assistant to ignore its rules. A working chat interface does not establish those behaviors.

This scenario excludes purchasing or cleaning a new dataset, integrations into additional business systems, legal review and a guaranteed service level. Add those explicitly if they are required. Expanding to more sources can change access control and synchronization work more than the model consumption line.

Scenario 2: a field vision feature

Assume a pilot across six camera locations, with a decision needed near each device. The estimate has to include capture quality, lighting, available compute, installation and how the operator responds to a detection. A low-cost board can still require significant engineering if the input varies in the field.

For a bandwidth illustration, a stream at 2 megabits per second generates approximately 21.6 decimal gigabytes over 24 hours, before transport overhead. Six continuously transmitted streams would therefore generate about 129.6 GB per day. That is an arithmetic assumption, not the measured traffic of a byASB deployment.

A local processor that sends selected events instead could reduce traffic, but its hardware, maintenance and event completeness must be evaluated. Compare those choices using the edge versus cloud AI guide and adjustable TCO calculator. The result depends on your actual workload and supplier costs.

The EV detector case shows why the surrounding pipeline deserves attention. Exposure, normalization and decision logic contributed to poor behavior; another model training cycle was not the only possible intervention. Budget for observing failures and evaluating the complete decision path, not just training.

This scenario excludes a full rollout, installation quotes, a formal safety assessment and a universal accuracy guarantee. A pilot can establish evidence for the next stage. It should state which operating conditions it covered and which remain untested.

Include evaluation and human review

Build a task-specific evaluation set before optimizing consumption. Include routine inputs and consequential failures. For a document assistant, that means unsupported answers and permission boundaries. For vision, it means missed objects, false alarms and behavior under changing illumination.

Price the process of reviewing outputs as well as generating them. If a person must approve every proposed action, estimate the review time and how exceptions reach that person. An apparently cheap model can create expensive downstream work if its errors are hard to detect.

State who accepts the result and how changes are evaluated after launch. A supplier update, a new source document or a camera replacement can change behavior. A useful handover includes the evaluation procedure, operating limits and a fallback path.

How to compare AI integration proposals

Give each supplier the same workflow, sources, permissions and acceptance criteria. Ask them to distinguish assumptions from commitments. A proposal for a prototype cannot be compared directly with a proposal that includes production access controls, operating documentation and post-launch support.

  • Which work is included in the first release, and what is explicitly excluded?
  • What evidence will demonstrate acceptable behavior?
  • Which recurring charges depend on usage?
  • Who owns the source, deployment and supplier accounts?
  • How can the team disable the feature and continue the workflow?
  • What changes require a new estimate?

Record the date and currency of vendor rates. Keep local taxes, foreign exchange and procurement costs separate if they apply. Teams in the US, UK and other markets can share the same technical scope while having different contractual or data-handling requirements; those need review for the actual engagement.

Use scoping to reduce the unknowns

A paid scoping week can define the first workflow, inspect available data and set evaluation and delivery boundaries. It should leave you with a more useful estimate and an explanation of the assumptions, including any work that still needs a technical experiment.

Use the AI integration consulting service to describe the intended behavior, existing product and known constraints. If the main uncertainty is the platform around the AI feature, the architecture review may be the more useful starting point.