How function calling works under the hood, walked through with an order-support agent, plus the validation, authorization and testing practices that make agents safe.
In this article
- 01What function calling actually is
- 02Anatomy of a tool definition
- 03The worked example: an order-support agent
- 04The agent loop in code terms
- 05Validation and authorization: never trust the arguments
- 06Designing tools the model can use well
- 07Handling errors and partial failures
- 08Testing an agent
- 09Standardizing tools with MCP
What function calling actually is
Function calling, also called tool calling or tool use, lets a language model ask your application to run a function. The model never executes anything itself. You describe the available tools; the model reads the conversation and, when a tool would help, responds with a structured request: the tool name and arguments as JSON. Your code runs the function, sends the result back, and the model continues with that new information.
That loop is what turns a chatbot into an AI agent: a system that can look things up, take actions and decide what to do next. Our explainer on function calling covers the concept briefly; this article walks through a real example.
Anatomy of a tool definition
Every major model API accepts tools in a similar shape: a name, a description and a parameters object written in JSON Schema. The description is not decoration. The model uses it to decide when to call the tool, so write it as you would explain the function to a new colleague, including when not to use it.
- name: a clear verb and noun, such as get_order or create_return
- description: what it does, when to use it and what it returns
- parameters: a JSON Schema object with typed properties
- required: the arguments the model must always supply
- enum: a fixed list of allowed values wherever possible, such as return reasons
The worked example: an order-support agent
Imagine a store's support assistant with three tools. get_order takes an order_id string and returns items, status and a tracking number. get_tracking takes a tracking_number and returns the carrier's latest events. create_return takes an order_id, an item_sku and a reason from an enum such as damaged, wrong_item or changed_mind, and creates a return request.
A signed-in customer writes: "Where is order 1042, and can I return the blue mug?" Here is what happens, message by message.
- The app sends the conversation and the three tool definitions to the model
- The model replies with a tool call: get_order with arguments {"order_id": "1042"}, plus a call ID
- The app checks that order 1042 belongs to the signed-in customer, runs the query and returns the result as a tool message referencing that call ID
- The model sees the tracking number and calls get_tracking with it
- With the tracking result, the model sees the mug's SKU but no reason for the return, so it asks the customer why
- The customer answers that it arrived chipped; the model calls create_return with reason damaged
- The app asks the customer to confirm, creates the return and the model summarizes delivery status and next steps
The agent loop in code terms
The control flow is a loop. Send messages and tools to the model. If the response contains tool calls, validate and execute each, append the results as tool messages and call the model again. If the response is plain text, show it to the user and stop. Many models can request several independent tool calls in one turn, which you can run in parallel.
Always cap the loop, for example at five to ten iterations, and set timeouts on each tool. Without limits, a confused model can call tools repeatedly and run up cost and latency.
Validation and authorization: never trust the arguments
Arguments come from a model that read user input, so treat them like any untrusted request body. Validate them against the schema with a library such as Ajv in JavaScript or Pydantic in Python, and return a clear error message as the tool result when validation fails. Models usually correct themselves when told what was wrong.
Authorization must come from your session, not from the model. In the example, the customer ID comes from the signed-in session; the model only supplies the order ID, and the tool checks ownership. If the model could pass a customer_id argument, a cleverly worded message could make it access someone else's data.
- Require explicit user confirmation before actions with side effects, such as refunds or emails
- Use idempotency keys so a repeated call does not create two returns
- Return errors as tool results instead of throwing, so the model can recover
- Log every tool call with its arguments, result and latency
Designing tools the model can use well
Fewer, well-named tools beat many overlapping ones. If two tools sound similar, the model will pick the wrong one some of the time. Return compact JSON with only the fields the model needs; huge payloads waste tokens and bury the useful data. Our JSON formatter is handy for checking tool schemas and sample results while you design them.
Prefer tools that match business actions, such as create_return, over thin wrappers on database tables. Business-level tools carry the rules, so the model cannot skip a step that matters.
Handling errors and partial failures
Tools fail: APIs time out, records are missing and downstream services return errors. Return failures to the model as structured tool results, such as {"error": "order_not_found", "message": "No order 1042 for this customer"}, rather than raising exceptions that end the conversation. The model can then ask the user to check the number or try another route. For transient errors, retry inside the tool with backoff before reporting failure, so the model is not left guessing whether to call again.
Testing an agent
Test tools like any other function, with unit tests for the logic and permissions. Then test the agent with a set of scripted conversations covering normal requests, missing information, ambiguous requests, attempts to access other customers' data and tool failures. Run the set whenever you change prompts, tools or models, and review traces of failures. Because model output varies, run each scenario more than once.
Standardizing tools with MCP
As the number of tools grows, the Model Context Protocol, an open standard now governed by the Agentic AI Foundation under the Linux Foundation, offers a standard way to expose tools and data sources to different AI applications through MCP servers, instead of writing custom integrations for each one. The same principles apply: clear descriptions, strict schemas and authorization enforced on the server.
Deciding whether you need an agent or a simpler chatbot is the first question; our AI agent vs chatbot comparison helps. Nexzem builds production agents with these safeguards through its AI agent development service.



