Skip to main content
Learn how to quickly identify and fix issues in production LLM applications using Helicone’s debugging features.

What You’ll Learn

  • Set up comprehensive request logging
  • Use filters to isolate problematic requests
  • Debug errors and unexpected outputs
  • Track prompt performance over time
  • Identify and fix latency issues

Prerequisites

  • Helicone API key (get one here)
  • An LLM application in production
  • Basic understanding of your application’s architecture

Common Debugging Scenarios

This tutorial covers:
  1. Finding and fixing errors (4XX/5XX)
  2. Debugging unexpected model outputs
  3. Identifying latency bottlenecks
  4. Tracking down cost spikes
  5. Investigating user-reported issues

Step 1: Enable Comprehensive Logging

Add headers to capture debugging context:
Key Debugging Headers:
  • Helicone-User-Id: Identify which users experience issues
  • Helicone-Property-Feature: Isolate problems to specific features
  • Helicone-Property-Environment: Separate dev/staging/production issues
  • Helicone-Property-Version: Track which code version has problems

Scenario 1: Finding and Fixing Errors

Problem: Users reporting 500 errors

1

Navigate to Requests Dashboard

2

Filter for Errors

Apply filters:
3

Identify Patterns

Look at the error list:
  • Are errors concentrated on a specific feature?
  • Affecting specific users?
  • Started at a specific time?
Example findings:
4

Inspect Request Details

Click on an error to see:
  • Full request payload
  • Error response
  • Model used
  • Request headers
  • Timestamp
5

Fix the Issue

Based on findings:
6

Monitor the Fix

Set up an alert to catch future issues:
  1. Go to Settings → Alerts
  2. Create alert:
    • Metric: Error Rate
    • Threshold: > 5%
    • Time window: 10 minutes
    • Filter: Feature = “document-analysis”
  3. Add Slack/email notification

Scenario 2: Debugging Unexpected Outputs

Problem: Model generating incorrect format

1

Find Problematic Requests

Filter requests:
2

Review Request/Response

Click on a request to see:
3

Identify the Issue

The prompt is too vague. The model needs clearer instructions.
4

Test Fix in Dashboard

Use Helicone’s prompt testing feature or test locally:
5

Compare Versions

After deploying, compare old vs. new:

Scenario 3: Identifying Latency Issues

Problem: Slow response times

1

Filter Slow Requests

2

Analyze Patterns

Look for:
  • Specific models (GPT-4 vs. GPT-4o-mini)
  • Request size (token count)
  • Features with long prompts
Example findings:
3

Optimize

4

Set Latency Alert

Create alert:
  • Metric: Latency
  • Threshold: P95 > 5000ms
  • Time window: 1 hour
  • Feature: report-generation

Scenario 4: Investigating Cost Spikes

Problem: Unexpected $500 charge

1

View Cost Dashboard

Go to Helicone Dashboard and check:
  • Daily cost trend (when did spike occur?)
  • Cost by feature
  • Cost by user
2

Filter High-Cost Requests

Findings:
3

Investigate the Request

Click on expensive request:
User uploaded massive document without chunking.
4

Implement Safeguards

Scenario 5: User-Reported Issue

Problem: “User user-456 says chatbot gave wrong answer yesterday”

1

Find User's Requests

2

Review Conversation

If using sessions:
View entire conversation flow to understand context.
3

Identify Issue

Review the specific request/response:
  • Was context missing?
  • Did model hallucinate?
  • Was there a misunderstanding?
Share findings with user:

Advanced: Custom Request IDs

Correlate Helicone logs with your application logs:

Best Practices

Add context headers: Include user ID, feature, environment, and version in every request
Use sessions for multi-step flows: Group related requests to see full context
Set up alerts early: Don’t wait for users to report issues
Compare before/after: Use version tags to measure impact of changes
Remove sensitive information from prompts before logging. Consider using environment variables or secure vaults for API keys.

Debugging Checklist

When investigating an issue:
  • Filter by relevant properties (user, feature, environment)
  • Check error rates and status codes
  • Review request/response payloads
  • Look for patterns (time-based, user-based, feature-based)
  • Check related requests (sessions)
  • Compare with working requests
  • Test fix with version tracking
  • Set up alert to catch recurrence

Next Steps

Alerts

Set up proactive monitoring for errors and anomalies

Sessions

Track multi-step workflows for better context

Custom Properties

Add metadata for powerful filtering and debugging

Webhooks

Get notified immediately when issues occur