What You’ll Build
A RAG evaluation pipeline that:- Runs RAGAS metrics (faithfulness, answer relevancy, context precision)
- Reports scores to Helicone
- Tracks evaluation trends over time
- Identifies low-performing responses
Prerequisites
- Helicone API key (get one here)
- OpenAI API key
- Python 3.8+ with pip
- A RAG application making LLM calls
Step 1: Install Dependencies
RAGAS is an evaluation framework for RAG pipelines. Learn more at docs.ragas.io
Step 2: Set Up Helicone Client
Configure your LLM client to log requests to Helicone:Step 3: Build RAG Function with Tracking
Create a RAG function that tracks Helicone request IDs:Step 4: Implement RAGAS Evaluation
Run RAGAS metrics on your RAG responses:RAGAS Metrics Explained:
- Faithfulness: Does the answer contain only information from the contexts? (no hallucinations)
- Answer Relevancy: How relevant is the answer to the question?
- Context Precision: Are the retrieved contexts relevant to the question?
- Context Recall: Do the contexts cover the ground truth answer?
Step 5: Report Scores to Helicone
Send evaluation scores to Helicone for tracking:Step 6: Create End-to-End Pipeline
Put everything together:Step 7: Run Evaluation
Create test cases and run the pipeline:Expected Output
Step 8: Analyze Results in Helicone
1
View Requests
Navigate to Helicone Requests and filter by:
- Property:
Feature = rag-qa - Property:
Environment = evaluation
2
Check Scores
Click on individual requests to see:
- RAGAS evaluation scores
- Request/response details
- Context used
- Latency and cost
3
Track Trends
Use the dashboard to:
- Plot average scores over time
- Identify degrading metrics
- Compare different prompt versions
- Find low-scoring requests for analysis
Advanced: Automated Evaluation
Run evaluations automatically on production traffic:Best Practices
Troubleshooting
RAGAS evaluation fails
RAGAS evaluation fails
Common issues:
- Missing OpenAI API key for RAGAS’s internal LLM calls
- Invalid data format (ensure contexts is a list of strings)
- Empty or None values in question/answer/contexts
Scores not appearing in Helicone
Scores not appearing in Helicone
Verify:
- Request ID is correct (check response headers for
helicone-id) - Scores are integers, not floats (multiply by 100)
- API response shows 200 status code
- Wait 10 minutes for score aggregation
Low faithfulness scores
Low faithfulness scores
Low faithfulness usually indicates:
- The model is hallucinating information not in contexts
- Contexts don’t contain enough information to answer
- Model is using external knowledge instead of contexts
Next Steps
Scores Documentation
Learn more about evaluation scores in Helicone
Sessions
Track multi-step RAG workflows
Custom Properties
Segment evaluation by version, environment, or user type
Webhooks
Get notified when scores drop below thresholds