Deep dive
Do not rebuild a general assistant
General-purpose assistants from the major AI labs are improving every few months and are hard to beat on breadth. A new assistant succeeds when it is clearly better for a specific group: lawyers drafting under one jurisdiction, doctors summarising consultations, engineers searching their own codebase, students preparing for one exam, a company's staff asking about internal policy. Narrow audiences let you build the knowledge, tools and interface that general apps cannot.
Write down the five jobs your users will do most and design the product for those jobs. Our AI consulting and product discovery work starts there, and often ends with a smaller, sharper first release than founders expect.
How an AI assistant app works
When a user sends a message, the backend checks their plan and limits, assembles a prompt from the system instructions, recent conversation and any retrieved knowledge, and sends it to the chosen model. The response streams back token by token to the app. Tool calls, such as a web search or a database lookup, are executed by your backend and fed back to the model before the final answer.
Everything is logged with traces: which model, which prompt version, how many tokens, how long it took and what it cost. Those traces power cost dashboards, debugging and evaluation. Without them, quality and margins drift silently.
- Version system prompts like code and test every change against an evaluation set.
- Set per-user and per-plan token budgets with clear messages when limits are reached.
- Keep a fallback model for outages and rate limits from any single provider.
Choosing and routing models
In 2026 the sensible default is to use several models. Frontier models from the GPT, Claude and Gemini families handle complex reasoning; smaller and cheaper models handle classification, summaries and simple replies; open-weight models such as Llama, Qwen, DeepSeek and Mistral can run on your own infrastructure when privacy or cost demands it. A routing layer picks the model per task and keeps the rest of the product independent of any one provider.
Model choice should come from your own evaluations, not from public leaderboards. Build a test set of real tasks with expected answers or grading rubrics, run candidate models against it and compare quality, latency and cost. Repeat whenever providers release new versions. Our comparisons of open-source vs proprietary LLMs and RAG vs fine-tuning help frame these choices.
Safety, privacy and regulation
Assistants can produce harmful, false or biased content, and they can be manipulated through prompt injection, especially when they read documents or web pages. Layered defences help: input and output filters, strict tool permissions, separating instructions from untrusted content, human approval for consequential actions and monitoring for abuse patterns.
Disclose clearly that users are talking to AI; under the EU AI Act this is a legal duty for chatbots from August 2026. Decide what you store, for how long and whether conversations are used to improve the product, and make that a user choice where the law requires it. Check each model provider's data retention terms and regional hosting options before sending user data. Our cybersecurity services team reviews these designs.
Unit economics and scaling
An assistant's gross margin depends on model costs per active user. Heavy users on a flat subscription can cost more than they pay, so plans need usage tiers, fair-use limits or cheaper models for routine work. Prompt caching, shorter context windows, summarising long conversations and retrieving only what is needed all reduce token spend.
At scale, the investments shift to agents that complete multi-step tasks, integrations through the Model Context Protocol and APIs, enterprise controls, and possibly self-hosted models. Maintenance for the application is typically 15-20% of the build cost per year, and model and infrastructure costs come on top, growing with usage.