What You’ll Learn
- Set up comprehensive request logging
- Use filters to isolate problematic requests
- Debug errors and unexpected outputs
- Track prompt performance over time
- Identify and fix latency issues
Prerequisites
- Helicone API key (get one here)
- An LLM application in production
- Basic understanding of your application’s architecture
Common Debugging Scenarios
This tutorial covers:- Finding and fixing errors (4XX/5XX)
- Debugging unexpected model outputs
- Identifying latency bottlenecks
- Tracking down cost spikes
- Investigating user-reported issues
Step 1: Enable Comprehensive Logging
Add headers to capture debugging context:Key Debugging Headers:
Helicone-User-Id: Identify which users experience issuesHelicone-Property-Feature: Isolate problems to specific featuresHelicone-Property-Environment: Separate dev/staging/production issuesHelicone-Property-Version: Track which code version has problems
Scenario 1: Finding and Fixing Errors
Problem: Users reporting 500 errors
1
Navigate to Requests Dashboard
Go to Helicone Requests
2
Filter for Errors
Apply filters:
3
Identify Patterns
Look at the error list:
- Are errors concentrated on a specific feature?
- Affecting specific users?
- Started at a specific time?
4
Inspect Request Details
Click on an error to see:
- Full request payload
- Error response
- Model used
- Request headers
- Timestamp
5
Fix the Issue
Based on findings:
6
Monitor the Fix
Set up an alert to catch future issues:
- Go to Settings → Alerts
- Create alert:
- Metric: Error Rate
- Threshold: > 5%
- Time window: 10 minutes
- Filter: Feature = “document-analysis”
- Add Slack/email notification
Scenario 2: Debugging Unexpected Outputs
Problem: Model generating incorrect format
1
Find Problematic Requests
Filter requests:
2
Review Request/Response
Click on a request to see:
3
Identify the Issue
The prompt is too vague. The model needs clearer instructions.
4
Test Fix in Dashboard
Use Helicone’s prompt testing feature or test locally:
5
Compare Versions
After deploying, compare old vs. new:
Scenario 3: Identifying Latency Issues
Problem: Slow response times
1
Filter Slow Requests
2
Analyze Patterns
Look for:
- Specific models (GPT-4 vs. GPT-4o-mini)
- Request size (token count)
- Features with long prompts
3
Optimize
4
Set Latency Alert
Create alert:
- Metric: Latency
- Threshold: P95 > 5000ms
- Time window: 1 hour
- Feature: report-generation
Scenario 4: Investigating Cost Spikes
Problem: Unexpected $500 charge
1
View Cost Dashboard
Go to Helicone Dashboard and check:
- Daily cost trend (when did spike occur?)
- Cost by feature
- Cost by user
2
Filter High-Cost Requests
3
Investigate the Request
Click on expensive request:User uploaded massive document without chunking.
4
Implement Safeguards
Scenario 5: User-Reported Issue
Problem: “User user-456 says chatbot gave wrong answer yesterday”
1
Find User's Requests
2
Review Conversation
If using sessions:View entire conversation flow to understand context.
3
Identify Issue
Review the specific request/response:
- Was context missing?
- Did model hallucinate?
- Was there a misunderstanding?
Advanced: Custom Request IDs
Correlate Helicone logs with your application logs:Best Practices
Debugging Checklist
When investigating an issue:- Filter by relevant properties (user, feature, environment)
- Check error rates and status codes
- Review request/response payloads
- Look for patterns (time-based, user-based, feature-based)
- Check related requests (sessions)
- Compare with working requests
- Test fix with version tracking
- Set up alert to catch recurrence
Next Steps
Alerts
Set up proactive monitoring for errors and anomalies
Sessions
Track multi-step workflows for better context
Custom Properties
Add metadata for powerful filtering and debugging
Webhooks
Get notified immediately when issues occur