Skip to main content

Overview

The chat completions endpoint provides OpenAI-compatible chat completion functionality with unified access to multiple LLM providers through the Helicone AI Gateway.

Authentication

All requests to the AI Gateway require authentication using your Helicone API key in the Authorization header:

Endpoint

Request Parameters

string
required
The model identifier to use for the completion (e.g., gpt-4, claude-3-opus-20240229)
array
required
Array of message objects representing the conversation history. Each message must have a role and content.Supported roles:
  • system - System instructions
  • user - User messages
  • assistant - Assistant responses
  • tool - Tool/function call results
  • function - Legacy function call results
  • developer - Developer-level instructions
number
Sampling temperature between 0 and 2. Higher values make output more random. Default varies by model.
integer
Maximum number of tokens to generate in the completion.
integer
Maximum number of completion tokens to generate (alternative to max_tokens).
number
Nucleus sampling parameter. Alternative to temperature. Value between 0 and 1.
number
Top-K sampling parameter for limiting token selection.
boolean
default:"false"
Whether to stream the response as Server-Sent Events (SSE).
object
Options for streaming responses.
string | array
Up to 4 sequences where the API will stop generating further tokens.
integer
default:"1"
Number of chat completion choices to generate (1-128).
number
default:"0"
Penalize new tokens based on whether they appear in the text so far (-2.0 to 2.0).
number
default:"0"
Penalize new tokens based on their frequency in the text so far (-2.0 to 2.0).
object
Modify the likelihood of specified tokens appearing in the completion.
boolean
default:"false"
Whether to return log probabilities of output tokens.
integer
Number of most likely tokens to return at each position (0-20). Requires logprobs: true.
object
Format for the model’s output.
array
List of tools the model can call. Use this for function calling.
string | object
Controls which tool the model should use. Options: none, auto, required, or specific tool.
boolean
default:"true"
Whether to enable parallel function calling.
string
Unique identifier for the end-user, for monitoring and abuse detection.
integer
Random seed for deterministic sampling.
string
Service tier to use. Options: auto, default, flex, scale, priority
string
Amount of reasoning effort for reasoning models. Options: minimal, low, medium, high
object
Options for reasoning models.
object
Custom metadata to attach to the request for tracking and filtering in Helicone.
object
Cache control settings for prompt caching.
string
Key for prompt caching to reuse previous prompts.

Response Format

Non-Streaming Response

string
Unique identifier for the completion.
string
Object type, always chat.completion.
integer
Unix timestamp of when the completion was created.
string
The model used for completion.
array
Array of completion choices.
object
Token usage information.

Streaming Response

When stream: true, the response is returned as Server-Sent Events (SSE). Each event contains a JSON object with:
The stream ends with a [DONE] message.

Error Responses

object
Error information when a request fails.

Example Responses

Advanced Features

Function Calling

Define tools that the model can use:

Vision (Image Input)

Include images in your messages:

JSON Mode

Force the model to output valid JSON:

Rate Limits

Rate limits are applied at the organization level and vary based on your Helicone plan. Monitor your usage through the Helicone dashboard.

Best Practices

  • Always include error handling for API calls
  • Use streaming for better user experience with long responses
  • Set appropriate max_tokens to control costs
  • Use metadata to track and filter requests in Helicone
  • Implement retry logic with exponential backoff for transient errors