Function calling vs structured outputs: how to decide
Compare function calling and structured output modes in LLM APIs. Understand when each fits, the trade-offs in reliability, latency, and schema complexity.

If you've worked with LLM APIs to produce machine-readable output, you've likely faced the choice between function calling and structured outputs. Both promise structured JSON, but they serve fundamentally different purposes: one is for tool invocation, the other for data extraction. Picking the wrong mode can cost you extra latency, validation headaches, or brittle agent logic.
What each mode actually does
Function calling is an API feature where the model returns a structured object representing a function invocation — a tool call with a name and arguments (usually a JSON object). The model doesn't execute anything; it just decides which tool to call and with what parameters. The API contract is "here are my tools, pick one and give me its arguments." This is typically used in multi-step agent loops: the model calls a tool, your code runs it, then feeds the result back.
Structured outputs, by contrast, instruct the model to return JSON directly matching a provided schema (e.g., a JSON Schema or a simpler type definition). There are no tool semantics — the output is just a data record. Providers achieve this through constrained decoding (e.g., grammar-based sampling), prompt engineering, or a dedicated response_format parameter that enforces the schema on the model side.
Both modes aim to produce structured, machine-readable output, but the underlying contract differs. Function calling implies an action ("call function X with args Y"), while structured output implies a data record ("here is the extracted entity").
Here's a concrete example of each using the OpenAI API.
Function calling (tool use):
import openai
response = openai.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
tools=[{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
},
"required": ["location"]
}
}
}],
tool_choice="auto"
)
tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name) # "get_weather"
print(tool_call.function.arguments) # '{"location":"Tokyo","unit":"celsius"}'Structured output (JSON schema):
import openai
response = openai.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Extract the event details: Meeting on June 5th at 3pm in Room 4B."}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "event_info",
"strict": True,
"schema": {
"type": "object",
"properties": {
"title": {"type": "string"},
"date": {"type": "string"},
"time": {"type": "string"},
"location": {"type": "string"}
},
"required": ["title", "date", "time", "location"]
}
}
}
)
print(response.choices[0].message.content)
# {"title":"Meeting","date":"June 5th","time":"3pm","location":"Room 4B"}Notice the difference: the first response is a tool call object, the second is raw JSON in the content field.
When function calling is the right choice
Function calling shines when the model must choose among multiple tools or APIs in a single generation. The API natively supports selecting which function to call and its arguments — no manual parsing of a "tool selection" field from raw JSON.
Use it when your downstream logic depends on the model's decision to call a function (or not). This is the canonical agent pattern: the model decides to search a database, call a calculator, or reply with a message. Many SDKs provide built-in fallback mechanisms for missing or malformed function calls, which reduces boilerplate.
If you're building an agent that needs to chain multiple tool calls, function calling integrates naturally with the loop. For a deeper look at the failure modes and retry logic, see Building an agent loop: tool calls, retries, failure modes.
Function calling also works well when you have a small, well-defined set of tools (say, 5–10). The model can reliably pick the right one, and you get structured arguments without writing a separate schema for the output.
When structured outputs serve better
Structured outputs are the right choice when you need a single, complex JSON object that does not represent a function invocation. Think extracting a full customer profile with nested addresses, arrays of orders, and optional fields. Function calling would force you to flatten that into a function's arguments, which is awkward and limits expressiveness.
Latency-critical applications also benefit. Function calling often requires a second request to execute the tool and feed results back. Structured outputs can complete the task in one shot — no tool execution, no re-prompting. Constrained decoding can further reduce tokens by preventing the model from generating invalid branches.
If you require strict schema validation at generation time, some providers (e.g., OpenAI with response_format) enforce the schema on the model side, reducing malformed output. This is particularly valuable for production pipelines where retries are expensive. For a detailed guide on reliable schema enforcement, see Reliable Structured Outputs from LLMs Using JSON Schema.
Structured outputs also play nicely with client-side validation libraries like Zod or Ajv. You can define your schema once and use it both for prompting and for parsing the response.
Reliability and failure modes compared
Both modes can fail, but the failure patterns differ.
Function calling can hallucinate function names or arguments that don't exist, especially when you provide many functions. The model might invent a function name, or produce arguments that violate the schema (e.g., a string where an integer is required). Mitigations include limiting the number of functions, using strict argument schemas (OpenAI's strict: true helps here), and setting temperature to 0.
Structured outputs can produce valid JSON that violates the schema — e.g., missing required fields, extra fields, or wrong types — unless the provider enforces it. Even with enforcement, some providers (like OpenAI's earlier response_format with type: "json_object") only guarantee valid JSON, not schema compliance. Client-side validation and a retry loop with the validation error in the prompt are often necessary.
Both modes degrade under high temperature or ambiguous prompts. Temperature 0 is recommended for both when determinism matters.
Here's a quick comparison:
| Aspect | Function Calling | Structured Outputs |
|---|---|---|
| Primary use case | Tool invocation, multi-step agents | Data extraction, single JSON record |
| Schema expressiveness | Limited to function parameters (often flat) | Full JSON Schema (nested, oneOf, $ref) |
| Provider-side enforcement | Argument schema validation (varies) | Some providers enforce full schema |
| Typical latency | Higher (tool call + execution roundtrip) | Lower (single turn if no tool needed) |
| Token overhead | Function name and argument keys add tokens | Compact if schema is simple |
| Failure mode | Hallucinated function names or arguments | Valid JSON but schema violations |
| Retry strategy | Re-prompt with tool call error | Re-prompt with validation error |
Latency and token efficiency trade-offs
Function calling typically returns a tool call object plus optional text content. The response includes the function name and argument keys, which adds tokens. For a simple function with two parameters, the overhead is small. For a complex schema with many optional fields, it can bloat.
Structured outputs can be more compact if the schema is simple — raw JSON without extra metadata. However, if the schema is deeply nested, the JSON itself can be long regardless of mode.
Constrained decoding (used by some structured output implementations) can reduce tokens by preventing invalid branches during generation. But it may add per-token overhead depending on the implementation (e.g., grammar-based sampling can be slower per token than free generation).
For multi-turn agents, function calling often requires a second request to execute the tool and feed results back. Structured outputs can sometimes complete the task in a single turn if no tool execution is needed. That said, if your task requires multiple data extractions, you might end up with multiple structured output calls anyway.
A practical tip: measure token usage and latency with your actual schema and prompt. Don't rely on general rules. Token counting is wrong: measure actual LLM cost explains why.
Schema expressiveness and constraints
Function calling typically limits your output shape to a single level of nested arguments, depending on the provider. OpenAI's function calling, for example, supports nested objects and arrays, but the schema is defined as a JSON Schema for the parameters object. However, some providers (e.g., Anthropic's tool use) enforce a flat structure. Complex nested arrays or deeply nested objects may require flattening or multiple calls.
Structured outputs support arbitrary JSON Schema features: oneOf, allOf, patternProperties, $ref, etc. This makes them suitable for data extraction tasks with rich domain models — e.g., extracting a legal document with clauses, subclauses, and cross-references.
But there's a catch: some providers restrict response_format to a subset of JSON Schema. OpenAI's json_schema mode, for instance, does not support $ref or default. Always check provider docs before designing a complex schema. The official OpenAI Structured Outputs guide lists supported features. Similarly, the JSON Schema specification is the reference for what's theoretically possible.
Making the call: a decision framework
Here's a simple heuristic:
- If the output is a single data object with no tool semantics → structured outputs.
- If the model must decide among multiple actions or needs to call external services → function calling.
Consider your validation pipeline. Structured outputs pair well with client-side validation libraries (Zod, Ajv); function calling pairs with tool execution middleware. If you're already using a schema-driven validation approach, structured outputs integrate more cleanly.
When in doubt, prototype both with your actual schema and measure token count, latency, and failure rate. The right choice often depends on the exact provider and model version. For example, on older models, function calling may be more reliable than structured outputs; on newer models, the reverse can be true.
Also think about the downstream consumers of the output. If you're feeding the result into another system that expects a specific JSON structure, structured outputs give you direct control. If you're building an agent that needs to decide which tool to call next, function calling is the natural fit.
Key takeaways
- Function calling is for tool invocation and agent loops; structured outputs are for data extraction and single-record generation.
- Function calling adds token overhead and often requires a second request for tool execution; structured outputs can be more compact and faster.
- Schema expressiveness is higher with structured outputs (full JSON Schema), but provider support varies — always check the docs.
- Both modes need validation and retry logic; don't rely solely on the API to guarantee correctness.
- Prototype both with your actual use case; the optimal choice depends on your provider, model, and latency requirements.
Frequently asked questions
- Can I use both function calling and structured outputs in the same request?
- Not directly in most APIs. You can simulate it by wrapping a structured output inside a single function call, or by having the model return a JSON string inside a function argument. OpenAI's parallel tool calls allow multiple function calls, but each call still uses the function schema, not a separate JSON Schema.
- Which is more reliable for complex schemas?
- Structured outputs with provider-side schema enforcement (e.g., OpenAI's `response_format` with `json_schema`) generally produce fewer malformed outputs than function calling for deeply nested or highly constrained schemas. Function calling can struggle with many optional fields or nested arrays because the model must generate valid JSON inside a string argument.
- Does structured output work with all LLMs?
- No. Only a subset of providers offer native structured output modes (OpenAI, Anthropic, Google Gemini). For open-weight models, you need constrained decoding libraries like `outlines`, `lm-format-enforcer`, or `guidance`. Without such tooling, you rely on prompt engineering, which is less reliable.


