AI Agents and Function Calling: Building Tools for LLMs
You're building an AI agent that needs to fetch real-time data, perform calculations, or interact with external APIs. But how do you bridge the gap between the LLM's text generation and actual code execution? Function calling (also known as tool use) is the answer. It lets LLMs request specific actions, and your code executes them. This guide walks you through designing, implementing, and debugging function calling for AI agents.
What is Function Calling?
Function calling is a mechanism where an LLM can output a structured request to call a function you've defined. Instead of generating free text, the model returns a JSON object with the function name and arguments. Your application then executes the function and feeds the result back to the model. This enables agents to perform actions beyond text generation, such as querying databases, sending emails, or calling APIs.
Major LLM providers like OpenAI, Anthropic, and Google support function calling. The core idea is consistent: you describe available tools, the model decides when to use them, and you handle execution.
Designing Tools for LLMs
Well-designed tools are crucial for reliable agent behavior. Follow these principles:
- Clear names and descriptions: Use descriptive function names and detailed descriptions. The LLM relies on these to choose the right tool.
- Simple parameters: Keep parameters minimal and use standard types (string, number, boolean, array, object). Avoid complex nested structures when possible.
- Idempotency: Where possible, design tools to be idempotent (safe to retry) to handle failures gracefully.
- Error handling: Return informative error messages so the LLM can adjust its approach.
Example: Weather Tool
Here's a simple tool definition in JSON schema format, commonly used with OpenAI's API:
{
"name": "get_weather",
"description": "Get the current weather for a given city",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "The city name, e.g., San Francisco"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature unit"
}
},
"required": ["city"]
}
}
Implementing Function Calling: Step-by-Step
Let's build a minimal agent loop using OpenAI's API (the pattern applies to other providers).
1. Define Your Tools
Create a list of tool schemas and a mapping from function names to actual Python functions.
import json
import openai
# Tool schemas
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
},
"required": ["city"]
}
}
}
]
# Actual functions
def get_weather(city: str, unit: str = "celsius") -> dict:
# In reality, call a weather API
return {"city": city, "temperature": 22, "unit": unit, "condition": "sunny"}
# Map names to functions
function_map = {
"get_weather": get_weather
}
2. Create the Agent Loop
The agent loop sends messages to the LLM, checks for tool calls, executes them, and repeats until the model returns a final answer.
def run_agent(user_message: str):
messages = [{"role": "user", "content": user_message}]
while True:
response = openai.ChatCompletion.create(
model="gpt-4",
messages=messages,
tools=tools,
tool_choice="auto"
)
message = response.choices[0].message
messages.append(message)
# If no tool calls, return the content
if not message.get("tool_calls"):
return message["content"]
# Execute each tool call
for tool_call in message.tool_calls:
function_name = tool_call.function.name
arguments = json.loads(tool_call.function.arguments)
if function_name in function_map:
result = function_map[function_name](**arguments)
else:
result = {"error": f"Unknown function: {function_name}"}
# Append tool result to messages
messages.append({
"role": "tool",
"tool_call_id": tool_call.id,
"content": json.dumps(result)
})
This loop continues until the LLM produces a response without tool calls, indicating it has enough information.
3. Handle Errors and Edge Cases
Real-world agents must handle:
- Invalid arguments: Catch exceptions when parsing arguments or calling functions.
- Unknown tools: Return an error message to the LLM so it can correct itself.
- Timeouts: Set timeouts for external API calls to avoid hanging.
- Rate limits: Implement retries with exponential backoff.
Best Practices for Reliable Agents
- Limit tool count: Too many tools confuse the model. Group related functions or use a router.
- Validate inputs: Sanitize and validate all arguments before execution.
- Log everything: Log tool calls and results for debugging and auditing.
- Test with varied prompts: Ensure the agent selects the right tools in different scenarios.
- Provide fallbacks: If a tool fails, the agent should try alternatives or ask for clarification.
Comparison of Function Calling Support
| Provider | Feature Name | Format |
|---|---|---|
| OpenAI | Function Calling | JSON Schema |
| Anthropic | Tool Use | JSON Schema |
| Function Calling | OpenAPI Schema |
Advanced Patterns
As your agent grows, consider these patterns:
- Parallel tool calls: Some models can request multiple tools at once. Execute them concurrently for speed.
- Human-in-the-loop: For sensitive actions (e.g., sending money), require human approval before execution.
- Memory: Store conversation history and tool results to provide context in long sessions.
- Tool routing: Use a lightweight classifier to select relevant tools before calling the main LLM.
Debugging Function Calling
When things go wrong, check:
- Are tool descriptions clear and unambiguous?
- Are parameter names and types correct?
- Does the model have enough context to choose the right tool?
- Are you handling tool results correctly (e.g., JSON serialization)?
Use logging to capture the full message history and tool calls. Often, the issue is a mismatch between the expected and actual arguments.
FAQ
What is the difference between function calling and tool use?
They refer to the same concept. OpenAI calls it "function calling," while Anthropic uses "tool use." Both allow LLMs to request execution of external functions.
Can I use function calling with open-source models?
Yes, some open-source models like Llama 3.1 support function calling, and frameworks like LangChain provide abstractions. However, support varies, and you may need to fine-tune or use specific prompt formats.
How do I prevent the LLM from calling dangerous functions?
Never expose dangerous functions directly. Use allowlists, validate inputs, and implement permission checks. For sensitive operations, require human confirmation before execution.
Ready to build your own AI agent? Start by defining a simple tool and testing the agent loop. For more developer tools, check out our JSON Formatter to debug tool call payloads.