Skip to main content
Learn how to evaluate Retrieval-Augmented Generation (RAG) applications using RAGAS and report evaluation scores to Helicone for centralized observability.

What You’ll Build

A RAG evaluation pipeline that:
  • Runs RAGAS metrics (faithfulness, answer relevancy, context precision)
  • Reports scores to Helicone
  • Tracks evaluation trends over time
  • Identifies low-performing responses

Prerequisites

  • Helicone API key (get one here)
  • OpenAI API key
  • Python 3.8+ with pip
  • A RAG application making LLM calls

Step 1: Install Dependencies

RAGAS is an evaluation framework for RAG pipelines. Learn more at docs.ragas.io

Step 2: Set Up Helicone Client

Configure your LLM client to log requests to Helicone:

Step 3: Build RAG Function with Tracking

Create a RAG function that tracks Helicone request IDs:
To get the Helicone request ID from response headers, use the OpenAI client’s .with_response() method or inspect response headers directly.

Step 4: Implement RAGAS Evaluation

Run RAGAS metrics on your RAG responses:
RAGAS Metrics Explained:
  • Faithfulness: Does the answer contain only information from the contexts? (no hallucinations)
  • Answer Relevancy: How relevant is the answer to the question?
  • Context Precision: Are the retrieved contexts relevant to the question?
  • Context Recall: Do the contexts cover the ground truth answer?

Step 5: Report Scores to Helicone

Send evaluation scores to Helicone for tracking:
Important: Helicone scores must be integers or booleans. Convert RAGAS scores (0-1 floats) to integers (0-100) by multiplying by 100.

Step 6: Create End-to-End Pipeline

Put everything together:

Step 7: Run Evaluation

Create test cases and run the pipeline:

Expected Output

Step 8: Analyze Results in Helicone

1

View Requests

Navigate to Helicone Requests and filter by:
  • Property: Feature = rag-qa
  • Property: Environment = evaluation
2

Check Scores

Click on individual requests to see:
  • RAGAS evaluation scores
  • Request/response details
  • Context used
  • Latency and cost
3

Track Trends

Use the dashboard to:
  • Plot average scores over time
  • Identify degrading metrics
  • Compare different prompt versions
  • Find low-scoring requests for analysis

Advanced: Automated Evaluation

Run evaluations automatically on production traffic:

Best Practices

Start with a golden dataset: Create 20-50 high-quality test cases with ground truth answers
Run evaluations regularly: Set up automated evaluations to catch regressions early
Track score trends: Monitor how metrics change over time, especially after prompt changes
Investigate outliers: Low-scoring responses often reveal edge cases or data quality issues
RAGAS requires an LLM to calculate some metrics, which adds cost and latency. Consider evaluating a sample of production traffic rather than every request.

Troubleshooting

Common issues:
  • Missing OpenAI API key for RAGAS’s internal LLM calls
  • Invalid data format (ensure contexts is a list of strings)
  • Empty or None values in question/answer/contexts
Check RAGAS logs for specific errors.
Verify:
  • Request ID is correct (check response headers for helicone-id)
  • Scores are integers, not floats (multiply by 100)
  • API response shows 200 status code
  • Wait 10 minutes for score aggregation
Low faithfulness usually indicates:
  • The model is hallucinating information not in contexts
  • Contexts don’t contain enough information to answer
  • Model is using external knowledge instead of contexts
Review the actual responses to identify the issue.

Next Steps

Scores Documentation

Learn more about evaluation scores in Helicone

Sessions

Track multi-step RAG workflows

Custom Properties

Segment evaluation by version, environment, or user type

Webhooks

Get notified when scores drop below thresholds