Handling and Processing API Responses
The Last Mile Problem
Building a great AI prompt is only half the job. The other half is reliably processing the output. AI API responses are often messy in production: the model might include extra text around a JSON object, format a list differently than expected, or occasionally return something that does not match your expected structure at all.
The last mile of AI engineering -- extracting, validating, and transforming the model's output into something your application can reliably use -- is where many AI systems succeed or fail in production.
Parsing Structured Output
JSON Responses
When you use response_format: json_object, the API guarantees valid JSON. Parse it with standard JSON parsing:
// Run in Node.js
const response = await client.chat.completions.create({
model: 'gpt-4o-mini',
messages: [
{ role: 'system', content: 'Extract name, email, and company from the text. Return JSON with keys: name, email, company.' },
{ role: 'user', content: 'Hi, I am Amara Nwosu from BuildRight Ltd. Contact me at amara@buildright.com.' }
],
response_format: { type: 'json_object' },
temperature: 0,
});
const raw = response.choices[0].message.content;
const extracted = JSON.parse(raw);
console.log(extracted.name); // "Amara Nwosu"
console.log(extracted.email); // "amara@buildright.com"
console.log(extracted.company); // "BuildRight Ltd"
Always wrap JSON parsing in a try-catch. Even with json_object mode, edge cases can occasionally produce unparseable output:
function safeParseJSON(text) {
try {
return { data: JSON.parse(text), error: null };
} catch (e) {
return { data: null, error: 'JSON parse failed: ' + e.message };
}
}
Extracting JSON from Mixed Text
When you are not using response_format (for example, with older models), the model might include text around the JSON. Use a regex to extract it:
function extractJSON(text) {
// Match the first { ... } block in the response
const match = text.match(/\{[\s\S]*\}/);
if (!match) return null;
try {
return JSON.parse(match[0]);
} catch {
return null;
}
}
Output Validation
Never trust AI output blindly. Validate that the output matches your expected schema before using it:
function validateExtraction(data) {
const errors = [];
if (!data.name || typeof data.name !== 'string') {
errors.push('name is missing or not a string');
}
if (!data.email || !data.email.includes('@')) {
errors.push('email is missing or invalid');
}
if (!data.company || typeof data.company !== 'string') {
errors.push('company is missing or not a string');
}
return errors;
}
const result = safeParseJSON(raw);
if (result.error) {
// Handle parse failure
console.error('Could not parse AI output');
} else {
const validationErrors = validateExtraction(result.data);
if (validationErrors.length > 0) {
// Handle validation failure -- re-run the prompt or flag for human review
console.error('Validation failed:', validationErrors);
} else {
// Safe to use result.data
console.log('Valid extraction:', result.data);
}
}
Classification Response Handling
For classification tasks, the model should return a single label. But it might return "Positive." with a period, or "POSITIVE" in caps, or "The sentiment is Positive." Normalise the output:
function normaliseLabel(raw, validLabels) {
// Remove punctuation, trim whitespace, lowercase
const cleaned = raw.replace(/[^a-zA-Z_\-]/g, '').trim().toLowerCase();
// Find the matching valid label (case-insensitive)
const match = validLabels.find(label => label.toLowerCase() === cleaned);
return match || null; // Return null if no valid label found
}
const rawLabel = response.choices[0].message.content;
const label = normaliseLabel(rawLabel, ['Positive', 'Negative', 'Neutral']);
if (!label) {
console.error('Unexpected label:', rawLabel);
// Fallback: trigger human review or use default
}
Handling Model Refusals
AI models sometimes refuse to complete a request due to safety policies. The response will contain a refusal message in choices[0].message.refusal:
const choice = response.choices[0];
if (choice.message.refusal) {
console.log('Model refused:', choice.message.refusal);
// Handle gracefully -- log, notify, or return a safe fallback
} else {
const content = choice.message.content;
// Process normally
}
Design your application to handle refusals gracefully. Do not let an unexpected refusal crash your system or expose a confusing error to users.
Retry Logic for Failures
Some failures are transient (temporary network issues, rate limits). Implement exponential backoff:
// Run in Node.js
async function callWithRetry(messages, maxRetries = 3) {
for (let attempt = 1; attempt <= maxRetries; attempt++) {
try {
const response = await client.chat.completions.create({
model: 'gpt-4o-mini',
messages,
temperature: 0,
});
return response.choices[0].message.content;
} catch (error) {
const isRetryable = error.status === 429 || error.status >= 500;
if (!isRetryable || attempt === maxRetries) {
throw error; // Give up on non-retryable errors or after max retries
}
const delayMs = Math.pow(2, attempt) * 500; // 1s, 2s, 4s
console.log('Attempt ' + attempt + ' failed. Retrying in ' + delayMs + 'ms...');
await new Promise(resolve => setTimeout(resolve, delayMs));
}
}
}
Fallback Strategies
When AI fails or produces invalid output, have a plan:
Option 1: Re-prompt with clarification If the output is invalid, send a follow-up message asking the model to fix the format:
messages.push({ role: 'user', content: 'Your previous response was not valid JSON. Please return only a JSON object with keys: name, email, company.' });
Option 2: Human review queue Route failed cases to a human review queue. Log the input and output, flag it, and process it manually.
Option 3: Default values For non-critical fields, accept null or a default value rather than failing the entire pipeline.
Option 4: Try a different model If gpt-4o-mini consistently fails on certain inputs, try gpt-4o which may handle edge cases better.
Monitoring Response Quality in Production
Set up monitoring to track:
- Parse success rate: what percentage of responses are valid JSON?
- Validation pass rate: what percentage pass your schema checks?
- Empty response rate: are any responses coming back blank?
- Average response length: sudden changes may indicate prompt or model changes
- Latency: track p50, p95, p99 response times
Use tools like Datadog, Grafana, or simple database logging to monitor these metrics over time. A drop in parse success rate is an early warning sign that something has changed -- either in your prompts, your inputs, or the model's behaviour.
Practice Exercise
Write a function that:
- Takes raw AI output text as input
- Attempts to extract a JSON object
- Validates the JSON has the required fields: title, category, and summary
- Returns either the valid data or a structured error object with details about what failed
Key Takeaways
- Always wrap JSON parsing in try-catch -- even responses from json_object mode can occasionally fail.
- Validate AI output against your expected schema before using it in downstream processing.
- Normalise classification labels to handle casing, punctuation, and extra text variations.
- Implement retry logic with exponential backoff for transient errors (429 rate limits, 500 server errors).
- Monitor parse success rates, validation pass rates, and response latency in production to catch degradation early.
Try it yourself
Key Takeaways
- Always wrap JSON parsing in try-catch and validate extracted data against your expected schema before use.
- Use regex extraction to pull JSON from mixed-text responses when json_object mode is not available.
- Normalise classification labels to handle casing, punctuation, and extra text that models commonly add.
- Implement retry logic with exponential backoff for 429 (rate limit) and 500 (server error) responses.
- Monitor parse success rate, validation pass rate, and response latency in production to catch degradation early.
Quick Quiz
1.Why should JSON parsing always be wrapped in a try-catch block?
2.What is label normalisation and why is it needed for classification tasks?
3.What is exponential backoff in the context of API retry logic?
4.Which production metric is the most useful early warning sign that your AI output processing is degrading?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx