← Previous: Prompting Like a Pro (Not Like a Google Search) | Next: Prompt Templates and Reusable Prompt Libraries →
Last time we covered patterns like chain-of-thought and role prompting — techniques for getting better answers. This article is about something different: getting answers in a format your code can actually use. That sounds simpler than reasoning quality, and yet it trips up almost every beginner building their first real application.
The Problem: “Just Ask for JSON” Isn’t Enough
Here’s the naive first attempt, and why it eventually breaks:
“Give me the name, price, and category of this product as JSON.”
Most of the time, this works fine. Then, without warning, one response comes back like this:
“Sure! Here’s the JSON you asked for:
json { "name": ... }”
Your code tries to parse that as pure JSON, hits the leading sentence and the code fences, and crashes. This isn’t rare or edge-case behavior — it’s a direct consequence of something from Foundations article 2: the model is predicting plausible conversational text, and “Sure! Here’s the JSON you asked for” is a completely normal, plausible way to start a helpful response. Nothing in plain prompting stops it from doing that.
Three Ways to Actually Solve This
flowchart TD
A[Need structured output] --> B[Option 1: Prompt engineering<br/>+ manual parsing]
A --> C[Option 2: JSON mode /<br/>schema constraints]
A --> D[Option 3: Function /<br/>tool calling]
Option 1: Better prompting, plus defensive parsing. Explicitly instruct the model to output only JSON, nothing else, then write parsing code that strips out stray text or code fences before parsing. This is the weakest option — it reduces the problem but never eliminates it, because you’re still relying on the model to behave, not enforcing it structurally.
Option 2: Schema-constrained output (JSON mode). Most major model APIs now support a mode where you provide a schema — the exact fields, types, and structure you need — and the model’s output generation is directly constrained to match it. This isn’t the model “trying harder” to follow instructions; it’s the underlying token generation process being restricted so it literally cannot produce a token that would break the schema. Far more reliable than prompting alone.
Option 3: Function calling / tool calling. Instead of asking the model to describe structured data in its response, you define a function signature (name, parameters, types) and let the model “call” it. The model produces arguments matching that function’s schema, and your code receives them directly, no parsing required. This is the same mechanism underlying the agent tool-use loop from Foundations article 8 — turns out reliable structured output and agent tool calls are solving the same core problem.
Why Schema Constraints Beat Prompting Alone
The difference is where the guarantee lives. With prompting alone, the guarantee lives in the model’s willingness to follow instructions — which, as we covered with hallucinations, is never 100% reliable because nothing forces it. With schema constraints or function calling, the guarantee lives in the generation process itself — certain tokens simply aren’t available to pick from if they’d violate the schema. That’s a structural fix, not a behavioral request.
A Practical Pattern: Define, Validate, Retry
Even with schema constraints, production systems usually add one more layer:
- Define the exact schema you need (fields, types, required vs optional).
- Generate using schema-constrained output or function calling.
- Validate the result against the schema in your own code anyway — trust, but verify, especially for nested or complex structures.
- Retry with feedback if validation fails — pass the error back to the model and ask it to correct the specific issue, rather than starting over blind.
This layered approach shows up constantly in real systems, and it echoes something from the evaluation article in Foundations: don’t just assume the output is correct because it looks right — check it against a real standard.
Mental Model
Think of the difference between asking someone to “please only write in the boxes on this form” versus handing them a form where the boxes are the only places a pen can physically touch the paper. Prompting alone is the first case — a polite, usually-followed request. Schema-constrained output and function calling are the second case — the format isn’t a request, it’s a constraint baked into the mechanism itself. One relies on cooperation, the other doesn’t need to.
Key Takeaways
- Asking for JSON in plain language works most of the time, but reliably breaks under conversational habits baked into the model.
- Schema-constrained output (JSON mode) restricts what tokens the model can generate, making structure a guarantee, not a request.
- Function/tool calling goes further — the model returns arguments to a defined function signature, no text parsing needed at all.
- This is the same mechanism behind agent tool use from Foundations — structured output and agent actions are the same underlying problem.
- Even with strong guarantees, production systems still validate output and retry with feedback rather than trusting blindly — the same discipline from the Evaluation article, applied at the individual response level.
Next up: Prompt Templates and Reusable Prompt Libraries — moving from one-off prompts to versioned, parameterized systems you can actually maintain.