LLM APIs for Developers: Getting Started with Chat Completions
You have a great idea for an AI-powered feature, but when you start reading the API docs, you hit a wall of jargon: tokens, temperature, streaming, system prompts. It feels like you need a second degree just to send a message. The good news is that the core concept is simple: you send a list of messages, and the model sends back a reply. Everything else is tuning. This article walks you through the practical steps to integrate an LLM chat completion API into your application, with code you can adapt today.
What Is a Chat Completion API?
At its heart, a chat completion API takes a conversation history as input and returns the next message in the conversation. The history is an array of message objects, each with a role and content. Roles are typically system, user, and assistant. The system message sets the behavior of the assistant, the user message is what the person typed, and the assistant message is what the model previously replied. By sending the entire history, you give the model context to generate a coherent response.
Most providers (OpenAI, Anthropic, Google, and open-source alternatives) follow a similar pattern, though the exact endpoint and payload may differ. Once you understand one, you can adapt to others quickly.
Step-by-Step: Your First API Call
Let's walk through a minimal example using Python and the OpenAI SDK. You can install it with pip install openai. Set your API key as an environment variable to keep it out of your code.
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain what an API is in one sentence."}
]
)
print(response.choices[0].message.content)
That's it. The response object contains a list of choices; usually you take the first one. The message content is the model's reply. You can also access usage statistics to track token consumption.
Key Parameters You Should Tune
Beyond the messages, a few parameters control the model's output:
- model: Which model to use. Smaller models are cheaper and faster; larger models are more capable.
- temperature: Controls randomness. 0 is deterministic, 1 is creative. For factual tasks, use a low value; for brainstorming, use a higher one.
- max_tokens: Limits the length of the reply. Set this to avoid unexpectedly long (and costly) responses.
- top_p: An alternative to temperature for controlling diversity. Usually you adjust one or the other, not both.
- stream: When true, the API sends partial responses as they are generated, improving perceived latency.
Streaming Responses for Better UX
Waiting for a full response can feel slow, especially for long answers. Streaming lets you display text as it arrives. Here's how to handle a stream in Python:
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Tell me a story."}],
stream=True
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
Each chunk contains a delta with a piece of the response. You accumulate these pieces to build the full message. In a web app, you can forward these chunks to the browser using Server-Sent Events (SSE) or WebSockets.
Handling Errors and Rate Limits
APIs fail. Networks hiccup. Rate limits kick in. Your integration should handle these gracefully. Common errors include:
- 401 Unauthorized: Invalid API key. Check your environment variable.
- 429 Too Many Requests: You've hit a rate limit. Implement exponential backoff and retry.
- 500 Internal Server Error: Provider-side issue. Retry with a delay.
Always wrap API calls in try/except blocks and log failures. For production, consider using a library like tenacity for retries.
Managing Costs and Tokens
Tokens are the currency of LLM APIs. Both input and output tokens count toward your bill. To keep costs predictable:
- Use smaller models for simple tasks.
- Trim conversation history to the last few messages when full context isn't needed.
- Set
max_tokensto cap output length. - Cache frequent responses if your use case allows.
Monitor usage through the provider's dashboard or by logging token counts from each response.
Security and Privacy Considerations
When sending user data to an LLM API, you're trusting a third party with that data. Review the provider's data retention policies. Avoid sending sensitive information like passwords or personal identifiers. If you must, consider self-hosted models or providers with strong privacy guarantees. Also, never expose your API key in client-side code; always proxy requests through your backend.
FAQ
What is the difference between a chat completion and a regular completion?
Chat completions are designed for conversational interfaces. They accept a list of messages with roles, allowing the model to maintain context. Regular completions take a single prompt string and are less suited for multi-turn dialogues.
How do I choose the right model?
Start with a smaller, cheaper model for prototyping. If you need better reasoning or longer context, move to a larger model. Test both on your specific task to find the best cost-performance trade-off.
Can I use multiple LLM providers in one app?
Yes. Many developers abstract the API calls behind an interface and switch providers based on cost, latency, or features. Libraries like LiteLLM or LangChain can help unify different APIs.
Next Steps
Now that you can make a basic call, experiment with system prompts to steer the model's behavior. Try streaming to improve user experience. And always keep an eye on token usage. As you build, you might need to format JSON responses from the model for further processing. For that, a JSON formatter can help you validate and pretty-print the output quickly.