API Parameters - Temperature, Tokens, and More
Why Parameters Matter
When you call an AI API, you do not just send a prompt. You also send parameters that control how the model generates its response. These parameters determine whether the output is creative or deterministic, how long it can be, and how many alternatives you get.
Understanding these parameters is the difference between an AI system that works reliably and one that produces unpredictable results. This lesson covers the most important parameters you will use in production.
Temperature
Temperature is the most important parameter to understand. It controls the randomness of the model's output.
Range: 0.0 to 2.0 (typically 0.0 to 1.0 in practice)
Temperature 0.0 (Deterministic)
At temperature 0, the model always selects the most probable next token. This makes the output maximally consistent and predictable. The same prompt will produce nearly identical results every time.
Best for: Classification, extraction, structured data generation, factual question answering, anything where consistency is critical.
Temperature 0.3 to 0.7 (Balanced)
Introduces some variation while keeping outputs coherent. This is the most common range for general-purpose AI applications.
Best for: Summarisation, analysis, general writing tasks.
Temperature 0.8 to 1.2 (Creative)
Introduces significant variation. Outputs will differ meaningfully between runs. The model takes more creative risks.
Best for: Creative writing, brainstorming, generating multiple diverse options.
Temperature above 1.2 (Erratic)
Outputs become increasingly random and often incoherent. Rarely useful in production.
The Rule of Thumb
For any system where consistency and reliability matter (which is most production AI automation), use temperature 0 to 0.3. Reserve higher temperatures for explicitly creative tasks.
max_tokens
This parameter limits the maximum number of tokens the model can generate in its response.
Key facts:
- Tokens are roughly 3/4 of a word in English
- 1,000 tokens is approximately 750 words
- GPT-4o has a context window of 128,000 tokens (input + output combined)
- Setting max_tokens too low truncates responses mid-sentence
How to set max_tokens
- For classification tasks: 10 to 50 tokens (just a label)
- For short summaries: 100 to 300 tokens
- For full documents or long analyses: 1,000 to 4,000 tokens
- Leave it unset or high for creative tasks where you do not know the length in advance
Important: You pay for tokens generated, not tokens allocated. Setting max_tokens to 1,000 does not cost 1,000 tokens if the response is 200 tokens.
top_p (Nucleus Sampling)
Top_p is an alternative to temperature for controlling randomness. Instead of adjusting the probability distribution for all tokens, it considers only the smallest set of tokens whose cumulative probability exceeds the top_p value.
Range: 0.0 to 1.0
- top_p = 0.1: Only consider the top 10% most likely tokens
- top_p = 0.9: Consider tokens until their combined probability reaches 90%
- top_p = 1.0: Consider all tokens (default)
Recommendation: Use either temperature or top_p, not both simultaneously. OpenAI recommends adjusting temperature and leaving top_p at its default.
n (Number of Completions)
The n parameter tells the API to generate multiple independent completions from the same prompt.
const response = await client.chat.completions.create({
model: 'gpt-4o-mini',
messages: [{ role: 'user', content: 'Give me a tagline for a productivity app.' }],
n: 3,
temperature: 0.9,
});
// Access all three options
response.choices.forEach((choice, i) => {
console.log('Option ' + (i+1) + ': ' + choice.message.content);
});
Use cases:
- Generating multiple creative options for a human to choose from
- A/B testing different phrasings
- Self-consistency checking (generate multiple answers, pick the most common)
Cost note: n=3 costs three times as much as n=1 since the model generates three full responses.
presence_penalty and frequency_penalty
These parameters reduce repetition in generated text:
frequency_penalty (0.0 to 2.0): Penalises tokens based on how often they have already appeared. Higher values discourage the model from repeating the same words and phrases.
presence_penalty (0.0 to 2.0): Penalises tokens that have appeared at all, regardless of frequency. Encourages the model to explore new topics.
When to use:
- For long-form content generation where you want variety: set frequency_penalty to 0.5 to 1.0
- For creative writing where you want diverse vocabulary: try presence_penalty 0.3 to 0.6
- For structured outputs (JSON, classification): keep both at 0 to avoid interfering with format
stop
The stop parameter lets you define one or more sequences that, when generated, cause the model to stop producing tokens.
const response = await client.chat.completions.create({
model: 'gpt-4o-mini',
messages: [{ role: 'user', content: 'List 5 benefits of exercise:' }],
stop: ['6.', '###', '
'],
});
This is useful when you want the model to stop at a specific delimiter, preventing it from generating beyond the point you care about.
seed
The seed parameter makes outputs reproducible. When you set seed to a specific integer and use temperature 0, you get nearly identical outputs each run.
const response = await client.chat.completions.create({
model: 'gpt-4o-mini',
messages: [{ role: 'user', content: 'Generate a product name for a budgeting app.' }],
seed: 42,
temperature: 0,
});
Use case: Testing and debugging, where you need reproducible outputs to compare prompt changes.
response_format
For structured data extraction, you can force the model to return valid JSON:
const response = await client.chat.completions.create({
model: 'gpt-4o-mini',
messages: [
{ role: 'system', content: 'Extract the name, email, and company from the text. Return valid JSON.' },
{ role: 'user', content: 'Hi, I am James Obi from TechCorp. You can reach me at james@techcorp.com.' }
],
response_format: { type: 'json_object' },
});
This ensures the output is always parseable JSON, eliminating format failures in production pipelines.
Practical Parameter Configurations
| Use Case | temperature | max_tokens | Notes |
|---|---|---|---|
| Email classification | 0 | 20 | Consistent, concise label |
| Data extraction to JSON | 0 | 500 | Use response_format: json_object |
| Customer email response | 0.3 | 200 | Consistent but slightly varied |
| Blog post generation | 0.7 | 1500 | Creative, varied |
| Brainstorming ideas | 1.0 | 400 | Maximum diversity |
Practice Exercise
Experiment with temperature by running the same creative writing prompt at temperature 0, 0.5, and 1.0. Observe how the outputs differ. Then try running the same classification prompt at temperature 0 and 0.8 and compare consistency.
Key Takeaways
- Temperature controls output randomness -- use 0 for classification and extraction, 0.3 to 0.7 for general tasks, and 0.8 to 1.2 for creative applications.
- max_tokens limits the response length -- set it based on the expected output size to control costs and prevent truncation.
- response_format: json_object forces the model to return valid, parseable JSON -- essential for production data pipelines.
- frequency_penalty and presence_penalty reduce repetition in long-form content generation.
- Use seed with temperature 0 for reproducible outputs during testing and debugging.
Try it yourself
Key Takeaways
- Temperature is the most important parameter -- use 0 for consistency (classification, extraction) and higher values for creative tasks.
- max_tokens controls response length -- set it based on your expected output size to avoid truncation and control costs.
- response_format: json_object forces valid JSON output, which is essential for production data pipelines.
- The n parameter generates multiple completions from the same prompt -- useful for options, A/B testing, and self-consistency.
- frequency_penalty and presence_penalty reduce repetition in long-form content, while seed enables reproducible outputs for testing.
Quick Quiz
1.What temperature setting should you use for a classification task where consistency is critical?
2.What does max_tokens control in an API call?
3.What does response_format: json_object do?
4.When would you use the n parameter in an API call?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx